Point cloud data transmitting device, point cloud data transmitting method, point cloud data receiving device, and point cloud data receiving method

The method enhances point cloud data transmission by encoding geometry data into prediction units and applying motion vectors, addressing latency and complexity issues to enable efficient and scalable point cloud services.

JP2025188208APending Publication Date: 2025-12-25LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025171870
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-07-05
Filing Date
2025-10-10
Publication Date
2025-12-25

AI Technical Summary

Technical Problem

Existing technologies face challenges in efficiently processing and transmitting massive amounts of point cloud data due to high latency and encoding/decoding complexity, particularly in applications like VR, AR, MR, and autonomous driving.

Method used

A method and apparatus for point cloud data transmission and reception that involve encoding geometry data into prediction units, applying motion vectors selectively, and signaling data to improve compression performance, using techniques such as inter-prediction compression and spatially adaptive partitioning.

Benefits of technology

This approach reduces encoding time and bitstream size, enabling high-quality point cloud services with improved parallel processing and scalability, supporting real-time capture, compression, transmission, and playback of point cloud data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025188208000001_ABST
    Figure 2025188208000001_ABST
Patent Text Reader

Abstract

To disclose a point cloud data transmitting method, cloud data transmitting device, cloud data receiving method, and cloud data receiving device according to embodiments.SOLUTION: A point cloud data transmitting method according to an embodiment includes the steps of encoding geometry data of point cloud data, encoding characteristic data of the point cloud data based on the geometry data, and transmitting the encoded geometry data, the encoded characteristic data, and signaling data. The geometry encoding step includes the steps of dividing the geometry data into one or more prediction units and selectively applying motion vectors to each of the divided prediction units to perform inter-prediction encoding of the geometry data.SELECTED DRAWING: Figure 35
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The embodiments relate to a method and apparatus for processing point cloud content. [Background technology]

[0002] Point cloud content is content expressed as a point cloud, which is a collection of points belonging to a coordinate system that represents three-dimensional space. Point cloud content can represent three-dimensional media and is used to provide various services such as VR (Virtual Reality), AR (Augmented Reality), MR (Mixed Reality, Extended Reality, XR (Extended Reality)), and autonomous driving services. However, expressing point cloud content requires tens of thousands to hundreds of thousands of point data. Therefore, a method for efficiently processing massive amounts of point data is required. Summary of the Invention [Problem to be solved by the invention]

[0003] The technical objective of the embodiments is to provide a point cloud data transmitting device, a transmitting method, a point cloud data receiving device and a receiving method for efficiently transmitting and receiving point clouds in order to solve the above-mentioned problems.

[0004] A technical problem according to the embodiments is to provide a point cloud data transmitting apparatus and method, and a point cloud data receiving apparatus and method that solve the latency and encoding / decoding complexity.

[0005] A technical objective of the present invention is to provide a point cloud data transmitting device and transmitting method, and a point cloud data receiving device and receiving method that improve the compression performance of point clouds by improving the coding technology of attribute information in geometry-based point cloud compression (G-PCC).

[0006] A technical problem according to the embodiment is to provide a point cloud data transmitting device, a transmitting method, and a point cloud data receiving device and a receiving method for efficiently compressing, transmitting, and receiving point cloud data captured by a LiDAR device.

[0007] A technical problem according to the embodiments is to provide a point cloud data transmitting device, a transmitting method, and a point cloud data receiving device and a receiving method for efficient inter-prediction compression of point cloud data captured by a lidar device.

[0008] A technical problem according to the embodiments is to provide a point cloud data transmitting device, a transmitting method, and a point cloud data receiving device and a receiving method that divide point cloud data into predetermined units for efficient inter-prediction compression of point cloud data captured by a lidar device.

[0009] A technical object of the present embodiment is to provide a point cloud data transmitting device, transmitting method, and point cloud data receiving device and receiving method that divide point cloud data into predetermined units and then selectively apply a motion vector to each of the divided predetermined units for efficient inter-prediction compression of point cloud data.

[0010] However, the scope of the invention is not limited to the above-mentioned technical problems, but can be extended to other technical problems that a person skilled in the art can derive based on all the contents described herein. [Means for solving the problem]

[0011] To achieve the above-mentioned objects and other advantages, a point cloud data transmission method according to an embodiment includes the steps of: encoding geometry data of the point cloud data; encoding attribute data of the point cloud data based on the geometry data; and transmitting the encoded geometry data, the encoded attribute data, and signaling data.

[0012] In one embodiment, the geometry encoding step includes a step of dividing the geometry data into one or more prediction units and a step of selectively applying a motion vector to each of the divided prediction units to perform inter-prediction encoding of the geometry data.

[0013] In one embodiment, the signaling data includes, for each prediction unit, information identifying whether a motion vector is applied or not.

[0014] In one embodiment, the motion vector is a global motion vector obtained by estimating the motion between successive frames.

[0015] In one embodiment, the point cloud data is captured by a lidar that includes one or more lasers.

[0016] In one embodiment, the dividing step divides the geometry data into one or more prediction units based on elevation or vertical.

[0017] In one embodiment, the signaling data further includes information for identifying a reference altitude size for prediction unit partitioning.

[0018] A point cloud data transmission device according to an embodiment includes a geometry encoder that encodes geometry data of the point cloud data, a feature encoder that encodes feature data of the point cloud data based on the geometry data, and a transmission unit that transmits the encoded geometry data, the encoded feature data, and signaling data.

[0019] In one embodiment, the geometry encoder includes a division unit that divides geometry data into one or more prediction units, and an inter-prediction unit that selectively applies a motion vector to each of the divided prediction units and performs inter-prediction encoding of the geometry data.

[0020] In one embodiment, the signaling data includes, for each prediction unit, information identifying whether a motion vector is applied or not.

[0021] In one embodiment, the motion vector is a global motion vector obtained by estimating the motion between successive frames.

[0022] In one embodiment, the point cloud data is captured by a lidar that includes one or more lasers.

[0023] In one embodiment, the dividing unit divides the geometry data into one or more prediction units based on elevation or vertical.

[0024] In one embodiment, the signaling data further includes information for identifying a reference altitude size for prediction unit partitioning.

[0025] A point cloud data receiving method according to an embodiment includes steps of receiving geometry data, feature data, and signaling data, decoding the geometry data based on the signaling data, decoding the feature data based on the signaling data and the decoded geometry data, and rendering reconstructed point cloud data based on the decoded geometry data and the decoded feature data.

[0026] In one embodiment, the geometry decoding step includes a step of dividing reference data of the geometry data into one or more prediction units based on signaling data, and a step of selectively applying a motion vector to each divided prediction unit based on the signaling data, and inter-prediction decoding the geometry data.

[0027] In one embodiment, the signaling data includes, for each prediction unit, information identifying whether a motion vector is applied or not.

[0028] In one embodiment, the motion vector is a global motion vector obtained by estimating the motion between successive frames at the transmitting side.

[0029] In one embodiment, the point cloud data is captured by a lidar that includes one or more lasers at the transmitting end.

[0030] In one embodiment, the dividing step divides the reference data into one or more prediction units based on elevation or vertical.

[0031] In an embodiment, the signaling data further includes information for identifying a reference altitude size for prediction unit partitioning. [Effects of the Invention]

[0032] The point cloud data transmitting method, transmitting device, point cloud data receiving method, and receiving device according to the embodiments provide high-quality point cloud services.

[0033] The point cloud data transmitting method, transmitting device, point cloud data receiving method, and receiving device according to the embodiment achieve various video codec methods.

[0034] The point cloud data transmission method, transmission device, point cloud data reception method, and reception device according to the embodiments provide general-purpose point cloud content such as for autonomous driving services.

[0035] The point cloud data transmission method, transmission device, point cloud data reception method, and reception device according to the embodiments provide improved parallel processing and scalability by performing spatially adaptive partitioning of point cloud data for independent encoding and decoding of point cloud data.

[0036] The point cloud data transmission method, transmission device, point cloud data receiving method, and receiving device according to the embodiments can improve the performance of encoding and decoding of point clouds by spatially dividing point cloud data into tile and / or slice units and encoding and decoding the data, and signaling the data necessary for this purpose.

[0037] The point cloud data transmitting method, transmitting device, point cloud data receiving method, and receiving device according to the embodiments support a method of dividing point cloud data into prediction units, LPU / PU (Largest Prediction Unit / Prediction Unit), reflecting the characteristics of the content, thereby enabling inter-prediction-based compression technology via reference frames to be applied to point clouds captured by LIDAR and having multiple frames. This expands the area that can be predicted using local motion vectors, eliminates the need for additional calculations, and reduces the encoding time for point cloud data.

[0038] The point cloud data transmission method, transmitting device, point cloud data receiving method, and receiving device according to the embodiments divide point cloud data into one or more prediction units based on elevation or vertical, and then signal whether a motion vector is applied to each divided prediction unit, thereby reducing the size of the geometry information bitstream and thereby efficiently supporting point cloud data capture / compression / transmission / restoration / playback services in real time. [Brief explanation of the drawings]

[0039] The drawings are attached to provide a further understanding of the embodiments and together with the description of the embodiments illustrate the embodiments.

[0040]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

Figure 35

Figure 36

Figure 37

[0041] Hereinafter, the embodiments described in this specification will be described in detail with reference to the accompanying drawings. Identical or similar components will be designated by the same reference numerals regardless of the drawing reference numerals, and redundant explanations will be omitted. The following embodiments are intended to embody the present invention and are not intended to limit or restrict the scope of the present invention. Anything that can be easily inferred by a person skilled in the art to which the present invention pertains from the detailed description and embodiments of the present invention is deemed to fall within the scope of the present invention.

[0042] The detailed description of this specification should not be construed as limiting in all respects, but should be considered as illustrative. The scope of the present invention should be determined based on reasonable interpretation of the appended claims, and all modifications within the equivalent scope of the present invention are included in the scope of the present invention.

[0043] Preferred embodiments will be described in detail with reference to the accompanying drawings. The following detailed description with reference to the accompanying drawings is intended to illustrate preferred embodiments rather than to illustrate only embodiments that can be implemented by the embodiments. The following detailed description includes details to provide a thorough understanding of the embodiments. However, it will be apparent to those skilled in the art that the embodiments can be practiced without such details. Most of the terms used in the embodiments are common terms widely used in the relevant field, but some have been arbitrarily selected by the applicant, and their meanings will be explained in detail below as necessary. Therefore, the embodiments should be understood based on the intended meaning of the terms, rather than the simple names or meanings of the terms. Furthermore, the following drawings and detailed description should not be interpreted as being limited to the specifically described embodiments, but should also be interpreted as including equivalents or alternatives to the embodiments described in the drawings and detailed description.

[0044] FIG. 1 is a diagram illustrating an example of a point cloud content providing system according to an embodiment.

[0045] The point cloud content providing system shown in Figure 1 includes a transmission device 10000 and a reception device 10004. The transmission device 10000 and the reception device 10004 are capable of wired and wireless communication to transmit and receive point cloud data.

[0046] According to an embodiment, the transmitting device 10000 acquires, processes, and transmits a point cloud video (or point cloud content). In the embodiment, the transmitting device 10000 includes a fixed station, a base transceiver system (BTS), a network, an AI (Artificial Intelligence) device and / or system, a robot, an AR / VR / XR device and / or server, etc. In the embodiment, the transmitting device 10000 also includes a device that communicates with a base station and / or other wireless devices using a wireless connection technology (e.g., 5G NR (New RAT), LTE (Long Term Evolution)), a robot, a vehicle, an AR / VR / XR device, a mobile device, a home appliance, an IoT (Internet of Things) device, an AI device / server, etc.

[0047] The transmitting device 10000 according to the embodiment includes a Point Cloud Video Acquisition unit 10001, a Point Cloud Video Encoder 10002, and / or a Transmitter (or communication module) 10003.

[0048] The point cloud video acquisition unit 10001 according to the embodiment acquires a point cloud video through a process such as capturing, synthesizing, or generating. The point cloud video is point cloud content represented by a point cloud, which is a collection of points located in a three-dimensional space, and is also called point cloud video data. The point cloud video according to the embodiment includes one or more frames. One frame represents a still image / picture. Therefore, the point cloud video includes a point cloud image / frame / picture, and is called any of a point cloud image, a frame, and a picture.

[0049] The point cloud video encoder 10002 according to the embodiment encodes the secured point cloud video data. The point cloud video encoder 10002 encodes the point cloud video data based on point cloud compression coding. The point cloud compression coding according to the embodiment includes Geometry-based Point Cloud Compression (G-PCC) coding and / or Video-based Point Cloud Compression (V-PCC) coding, or next-generation coding. Note that the point cloud compression coding according to the embodiment is not limited to the above-described embodiments. The point cloud video encoder 10002 can output a bitstream including encoded point cloud video data. The bitstream includes not only the encoded point cloud video data but also signaling information related to the encoding of the point cloud video data.

[0050] According to an embodiment, the transmitter 10003 transmits a bitstream including encoded point cloud video data. According to an embodiment, the bitstream is encapsulated into a file or a segment (e.g., a streaming segment) and transmitted via various networks such as a broadcast network and / or a broadband network. Although not shown, the transmitting device 10000 includes an encapsulation unit (or encapsulation module) that performs the encapsulation operation. In an embodiment, the encapsulation unit is included in the transmitter 10003. According to an embodiment, the file or segment is transmitted to the receiving device 10004 via a network or stored on a digital storage medium (e.g., USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.). According to an embodiment, the transmitter 10003 can communicate with the receiving device 10004 (or receiver 10005) via wired or wireless communication via a network such as 4G, 5G, or 6G. Furthermore, the transmitter 10003 can perform necessary data processing operations via a network system (e.g., a communication network system such as 4G, 5G, or 6G). The transmitting device 10000 can also transmit encapsulated data on an on-demand basis.

[0051] The receiving device 10004 according to the embodiment includes a receiver 10005, a point cloud video decoder 10006, and / or a renderer 10007. In the embodiment, the receiving device 10004 includes a device, robot, vehicle, AR / VR / XR device, mobile device, home appliance, IoT (Internet of Things) device, AI device / server, etc. that communicates with a base station and / or other wireless device using a wireless connection technology (e.g., 5G NR (New RAT), LTE (Long Term Evolution)).

[0052] The receiver 10005 according to the embodiment receives a bitstream containing point cloud video data or a file / segment in which the bitstream is encapsulated from a network or storage medium. The receiver 10005 performs data processing operations required by a network system (e.g., a communication network system such as 4G, 5G, or 6G). The receiver 10005 according to the embodiment decapsulates the received file / segment and outputs a bitstream. In addition, in the embodiment, the receiver 10005 includes a decapsulation unit (or decapsulation module) for performing the decapsulation operation. The decapsulation unit is embodied as an element (or component) separate from the receiver 10005.

[0053] The point cloud video decoder 10006 decodes a bitstream containing point cloud video data. The point cloud video decoder 10006 can decode the point cloud video data in the manner in which it was encoded (e.g., the reverse process of the operation of the point cloud video encoder 10002). Therefore, the point cloud video decoder 10006 can decode the point cloud video data by performing point cloud reconstruction coding, which is the reverse process of point cloud compression. Point cloud reconstruction coding includes G-PCC coding.

[0054] The renderer 10007 renders the decoded point cloud video data. The renderer 10007 renders not only the point cloud video data but also the audio data to output point cloud content. In an embodiment, the renderer 10007 includes a display for displaying the point cloud content. In an embodiment, the display is not included in the renderer 10007, but is embodied as a separate device or component.

[0055] In the drawing, dotted arrows indicate the transmission path of feedback information obtained by the receiving device 10004. The feedback information is information for reflecting interaction with a user consuming point cloud content and includes user information (e.g., head orientation information, viewport information, etc.). In particular, if the point cloud content is for a service that requires interaction with a user (e.g., an autonomous driving service), the feedback information may be transmitted to a content transmitting side (e.g., the transmitting device 10000) and / or a service provider. In an embodiment, the feedback information may be used not only by the transmitting device 10000 but also by the receiving device 10004, or may not be provided.

[0056] According to an embodiment, head orientation information is information regarding the position, direction, angle, movement, etc. of the user's head. According to an embodiment, the receiving device 10004 calculates viewport information based on the head orientation information. The viewport information is information regarding the area of ​​the point cloud video viewed by the user. The viewpoint is the point at which the user views the point cloud video and refers to the center of the viewport area. That is, the viewport is an area centered on the viewpoint, and the size, shape, etc. of the area are determined by the FOV (Field of View). Therefore, the receiving device 10004 can extract viewport information based on the vertical or horizontal FOV supported by the device in addition to the head orientation information. The receiving device 10004 also performs gaze analysis to determine the user's point cloud consumption method, the point cloud video area the user is gazing at, the gaze time, etc. In an embodiment, the receiving device 10004 can transmit feedback information including the results of the gaze analysis to the transmitting device 10000. According to an embodiment, the feedback information is obtained during the rendering and / or display process. In some embodiments, the feedback information is obtained by one or more sensors included in the receiving device 10004. In other embodiments, the feedback information is obtained by the renderer 10007 or another external element (or device, component, etc.).

[0057] The dotted lines in FIG. 1 indicate the transmission process of feedback information secured by the renderer 10007. The point cloud content providing system processes (encodes / decodes) point cloud data based on the feedback information. Therefore, the point cloud video data decoder 10006 can perform a decoding operation based on the feedback information. The receiving device 10004 can also transmit the feedback information to the transmitting device 10000. The transmitting device 10000 (or the point cloud video data encoder 10002) can perform an encoding operation based on the feedback information. Therefore, the point cloud content providing system does not process (encodes / decodes) all point cloud data, but can efficiently process necessary data (e.g., point cloud data corresponding to the user's head position) based on the feedback information to provide point cloud content to the user.

[0058] In an embodiment, the sending device 10000 may be referred to as an encoder, a sending device, a transmitter, etc., and the receiving device 10004 may be referred to as a decoder, a receiving device, a receiver, etc.

[0059] 1 according to an embodiment (processed through a series of steps of acquisition / encoding / transmission / decoding / rendering), the point cloud data may also be referred to as point cloud content data or point cloud video data. In an embodiment, the point cloud content data may be used as a concept including metadata or signaling information related to the point cloud data.

[0060] The elements of the point cloud content providing system shown in FIG. 1 may be implemented in hardware, software, a processor, and / or a combination thereof.

[0061] FIG. 2 is a block diagram illustrating the operation of providing point cloud content according to an embodiment.

[0062] Figure 2 is a block diagram showing the operation of the point cloud content providing system described in Figure 1. As described above, the point cloud content providing system processes point cloud data based on point cloud compression coding (e.g., G-PCC).

[0063] A point cloud content providing system (e.g., a point cloud transmitting device 10000 or a point cloud video acquiring unit 10001) according to an embodiment acquires a point cloud video (20000). The point cloud video is represented by a point cloud belonging to a coordinate system representing a three-dimensional space. The point cloud video according to an embodiment includes a Ply (Polygon File format or the Stanford Triangle format) file. If the point cloud video has one or more frames, the acquired point cloud video includes one or more Ply files. A Ply file includes point cloud data such as the geometry and / or attributes of points. The geometry includes the position of the point. The position of each point is expressed by parameters (e.g., values ​​on the X-axis, Y-axis, and Z-axis) indicating a three-dimensional coordinate system (e.g., a coordinate system consisting of X, Y, and Z axes). The attributes include the characteristics of the point (e.g., texture information of each point, color (YCbCr or RGB), reflectance (r), transparency, etc.). A point has one or more characteristics (or attributes). For example, a point can have one attribute of hue, or two attributes of hue and reflectance.

[0064] In embodiments, geometry may also be referred to as position, geometry information, geometry data, etc., and features may also be referred to as features, feature information, feature data, etc.

[0065] In addition, the point cloud content providing system (e.g., the point cloud transmitting device 10000 or the point cloud video acquiring unit 10001) can acquire point cloud data from information related to the point cloud video acquiring process (e.g., depth information, color information, etc.).

[0066] A point cloud content providing system (e.g., a transmitting device 10000 or a point cloud video encoder 10002) according to an embodiment encodes point cloud data (20001). The point cloud content providing system encodes point cloud data based on point cloud compression coding. As described above, point cloud data includes the geometry and attributes of points. Therefore, the point cloud content providing system can perform geometry coding to encode the geometry and output a geometry bitstream. The point cloud content providing system can perform attribute coding to encode the attributes and output a attribute bitstream. In an embodiment, the point cloud content providing system can perform attribute coding based on the geometry coding. The geometry bitstream and the attribute bitstream according to an embodiment are multiplexed and output as a single bitstream. The bitstream according to an embodiment further includes signaling information related to the geometry coding and the attribute coding.

[0067] A point cloud content providing system (e.g., transmitting device 10000 or transmitter 10003) according to an embodiment transmits encoded point cloud data (20002). As described in FIG. 1, the encoded point cloud data is represented by a geometry bitstream and a feature bitstream. The encoded point cloud data is transmitted in the form of a bitstream together with signaling information related to the encoding of the point cloud data (e.g., signaling information related to geometry encoding and feature encoding). The point cloud content providing system encapsulates the bitstream for transmitting the encoded point cloud data and transmits it in the form of a file or segment.

[0068] A point cloud content providing system (e.g., receiving device 10004 or receiver 10005) according to an embodiment receives a bitstream including encoded point cloud data, and can demultiplex the bitstream.

[0069] The point cloud content providing system (e.g., the receiving device 10004 or the point cloud video decoder 10005) decodes the encoded point cloud data (e.g., geometry bitstream, attribute bitstream) transmitted in a bitstream. The point cloud content providing system (e.g., the receiving device 10004 or the point cloud video decoder 10005) decodes the point cloud video data based on signaling information related to the encoding of the point cloud video data included in the bitstream. The point cloud content providing system (e.g., the receiving device 10004 or the point cloud video decoder 10005) decodes the geometry bitstream to restore the position (geometry) of the point. The point cloud content providing system decodes the attribute bitstream based on the restored geometry to restore the attribute of the point. The point cloud content providing system (e.g., the receiving device 10004 or the point cloud video decoder 10005) restores the point cloud video based on the position according to the restored geometry and the decoded attribute.

[0070] A point cloud content providing system (e.g., a receiving device 10004 or a renderer 10007) according to an embodiment renders the decoded point cloud data (20004). The point cloud content providing system (e.g., a receiving device 10004 or a renderer 10007) renders the geometry and characteristics decoded during the decoding process using various rendering methods. Points of the point cloud content are rendered as fixed points with a certain thickness, cubes with a predetermined minimum size centered at the position of the fixed points, or circles centered at the position of the fixed points. All or part of the region of the rendered point cloud content is provided to a user via a display (e.g., a VR / AR display, a general display, etc.).

[0071] A point cloud content providing system (e.g., receiving device 10004) according to the embodiment may obtain feedback information (20005). The point cloud content providing system encodes and / or decodes point cloud data based on the feedback information. The feedback information and operation of the point cloud content providing system according to the embodiment are the same as the feedback information and operation described in FIG. 1, so a detailed description will be omitted.

[0072] FIG. 3 illustrates an example of a point cloud video capturing process according to an embodiment.

[0073] FIG. 3 illustrates an example of a point cloud video capture process in the point cloud content providing system described in FIGS.

[0074] Point cloud content includes point cloud video (images and / or video) showing objects and / or environments located in various three-dimensional spaces (e.g., a three-dimensional space showing a real environment, a three-dimensional space showing a virtual environment, etc.). Accordingly, a point cloud content providing system according to an embodiment captures point cloud video using one or more cameras (e.g., an infrared camera capable of obtaining depth information, an RGB camera capable of extracting color information corresponding to the depth information, etc.), projectors (e.g., an infrared pattern projector for obtaining depth information), LiDAR, etc. to generate point cloud content. The point cloud content providing system according to an embodiment extracts a geometric form composed of points in three-dimensional space from the depth information and extracts characteristics of each point from the color information to obtain point cloud data. Images and / or video according to an embodiment are captured based on either an inward-facing approach or an outward-facing approach.

[0075] The left side of Figure 3 shows the inward-facing method. The inward-facing method is a method in which one or more cameras (or camera sensors) positioned around a central object capture the central object. The inward-facing method is used to generate point cloud content that provides the user with a 360-degree image of the core object (e.g., VR / AR content that provides the user with a 360-degree image of an object (e.g., a core object such as a character, player, item, or actor)).

[0076] The right side of Figure 3 shows an outward-facing approach. The outward-facing approach is a method in which one or more cameras (or camera sensors) positioned around a central object capture the environment of the central object, which is not the central object. The outward-facing approach is used to generate point cloud content to provide the surrounding environment from a user's perspective (e.g., content showing the external environment provided to a user of an autonomous vehicle).

[0077] As shown in FIG. 3, point cloud content is generated based on the capture operation of one or more cameras. In this case, since each camera has a different coordinate system, the point cloud content providing system calibrates one or more cameras to set a global coordinate system before the capture operation. The point cloud content providing system also generates point cloud content by combining images and / or video captured using the above capture method with an arbitrary image and / or video. When generating point cloud content representing a virtual space, the point cloud content providing system does not perform the capture operation described in FIG. 3. The point cloud content providing system according to the embodiment can also perform post-processing on the captured images and / or video. That is, the point cloud content providing system can remove unwanted areas (e.g., background) or fill any spatial holes by recognizing the space where the captured images and / or video are connected.

[0078] The point cloud content providing system can also generate a single point cloud content by performing coordinate system transformation on points in the point cloud video acquired from each camera. The point cloud content providing system performs coordinate system transformation on points based on the position coordinates of each camera. This allows the point cloud content providing system to generate content showing a single wide area or point cloud content with a high point density.

[0079] FIG. 4 is a diagram illustrating an example of a point cloud video encoder according to an embodiment.

[0080] 4 shows an example of the point cloud video encoder 10002 of FIG. 1. The point cloud video encoder performs encoding by reconstructing point cloud data (e.g., point positions and / or characteristics) to adjust the quality of the point cloud content (e.g., lossless, lossy, near-lossless) depending on the network conditions or application. If the overall size of the point cloud content is large (e.g., point cloud content of 60 Gbps for 30 fps), the point cloud content providing system cannot stream the content in real time. Therefore, the point cloud content providing system can reconstruct the point cloud content based on a maximum target bitrate to provide it according to the network environment, etc.

[0081] As shown in Figures 1 and 2, a point cloud video encoder can perform geometry encoding and feature encoding, where geometry encoding is performed before feature encoding.

[0082] The point cloud video encoder according to the embodiment includes a transformation coordinates unit 40000, a quantization unit 40001, an octree analysis unit 40002, a surface approximation analysis unit 40003, an arithmetic encoder 40004, a geometry reconstruction unit 40005, a color transformation unit 40006, an attribute transformation unit 40007, a region adaptive hierarchical transform (RAHT) unit 40008, an LOD generation unit 40009, a lifting transformation unit 40010, a coefficient quantization unit 40011 and / or an arithmetic encoder 40012.

[0083] The coordinate system transformation unit 40000, the quantization unit 40001, the octree analysis unit 40002, the surface approximation analysis unit 40003, the arithmetic encoder 40004, and the geometry reconstruction unit 40005 can perform geometry coding. Geometry coding according to the embodiment includes octree geometry coding, direct coding, trisoup geometry encoding, and entropy coding. Direct coding and trisoup geometry encoding are applied selectively or in combination. Note that geometry coding is not limited to the above examples.

[0084] As shown in the figure, a coordinate system conversion unit 40000 according to an embodiment receives a position and converts it into a coordinate system. For example, the position is converted into position information in a three-dimensional space (e.g., a three-dimensional space expressed in an XYZ coordinate system). The position information in the three-dimensional space according to an embodiment is also referred to as geometry information.

[0085] The quantizer 40001 according to the embodiment quantizes geometry. For example, the quantizer 40001 quantizes points based on the minimum position value of all points (e.g., the minimum value on each axis for the X, Y, and Z axes). The quantizer 40001 performs a quantization operation by multiplying the difference between the minimum position value and the position value of each point by a predetermined quantization scale value and then rounding down or up to find the nearest integer value. Therefore, one or more points may have the same quantized position (or position value). The quantizer 40001 according to the embodiment performs voxelization based on the quantized position to reconstruct the quantized points. Just as the smallest unit containing 2D image / video information is a pixel, points in the point cloud content (or 3D point cloud video) according to the embodiment are included in one or more voxels. A voxel is a combination of the words volume and pixel and refers to a 3D cubic space generated when a 3D space is divided into units (unit=1.0) based on the axes (e.g., X-axis, Y-axis, Z-axis) that represent the 3D space. The quantization unit 40001 can match a group of points in the 3D space with voxels. In some embodiments, a voxel may contain only one point. In some embodiments, a voxel may contain one or more points. To represent a voxel as a point, the center of the voxel may be set based on the positions of one or more points contained in the voxel. In this case, the characteristics of all points contained in a voxel are combined and assigned to the voxel.

[0086] The octree analysis unit 40002 according to the embodiment performs octree geometry coding (or octree coding) to represent voxels in an octree structure, which represents points matched to voxels based on an octet structure.

[0087] The surface approximation analysis unit 40003 according to the embodiment analyzes and approximates the octree. The octree analysis and approximation according to the embodiment is a process of analyzing a region containing a large number of points to voxelize it in order to efficiently provide an octree and voxelization.

[0088] The arithmetic encoder 40004 according to the embodiment entropy encodes the octree and / or the approximated octree. For example, the encoding method includes an arithmetic encoding method. As a result of the encoding, a geometry bitstream is generated.

[0089] The color transform unit 40006, the feature transform unit 40007, the RAHT transform unit 40008, the LOD generation unit 40009, the lift transform unit 40010, the coefficient quantization unit 40011, and / or the arithmetic encoder 40012 perform feature coding. As described above, one point has one or more features. Feature coding according to the embodiment is applied equally to all features of one point. However, if one feature (e.g., hue) includes one or more elements, independent feature coding is applied to each element. Feature coding according to the embodiment includes color transform coding, feature transform coding, RAHT (Region Adaptive Hierarchical Transform) coding, Interpolarization-based hierarchical nearest-neighbor prediction-Prediction Transform (Interpolarization-based hierarchical nearest-neighbor prediction with an update / lifting step (Lifting Transform)) coding. Depending on the point cloud content, the above-mentioned RAHT coding, predictive transform coding, and lift transform coding may be selectively used, or a combination of one or more coding methods may be used. Furthermore, the feature coding according to the embodiment is not limited to the above examples.

[0090] The color converter 40006 according to the embodiment performs color conversion coding to convert the color values ​​(or textures) included in the attributes. For example, the color converter 40006 converts the format of the hue information (e.g., converts from RGB to YCbCr). The operation of the color converter 40006 according to the embodiment is optionally applied depending on the color values ​​included in the attributes.

[0091] The geometry reconstruction unit 40005 according to the embodiment reconstructs (restores) the octree and / or the approximated octree. The geometry reconstruction unit 40005 reconstructs the octree / voxel based on the result of analyzing the distribution of points. The reconstructed octree / voxel is also called a reconstructed geometry (or restored geometry).

[0092] According to an embodiment, the feature converter 40007 performs feature conversion, converting features based on a position where geometry encoding has not been performed and / or reconstructed geometry. As described above, because features depend on geometry, the feature converter 40007 can convert features based on reconstructed geometry information. For example, the feature converter 40007 can convert features of points included in a voxel based on their position values. As described above, if the center point of a voxel is set based on the positions of one or more points included in the voxel, the feature converter 40007 converts features of one or more points. If trisoup geometry encoding is performed, the feature converter 40007 can convert features based on the trisoup geometry encoding.

[0093] The feature conversion unit 40007 performs feature conversion by calculating the average value of the features or feature values ​​(e.g., hue or reflectance of each point) of adjacent points within a specific position / radius from the position (or position value) of the center point of each voxel. When calculating the average value, the feature conversion unit 40007 applies a weight based on the distance from the center point to each point. Therefore, each voxel has a position and a calculated feature (or feature value).

[0094] The feature conversion unit 40007 searches for neighboring points within a specific position / radius from the center point of each voxel based on a KD tree or Moulton code. A KD tree supports a data structure that manages points based on their position, enabling a fast Nearest Neighbor Search (NNS) using a binary search tree. A Moulton code is generated by mixing bits, representing the coordinate values ​​(e.g., (x, y, z)) that indicate the 3D position of all points. For example, if the coordinate value indicating the point's position is (5, 9, 1), the bit values ​​of the coordinate value are (0101, 1001, 0001). Mixing the bit values ​​in the order of z, y, and x according to the bit index results in 010001000111. This value is expressed in decimal as 1095. In other words, the Moulton code value of a point with coordinate values ​​(5, 9, 1) is 1095. The feature conversion unit 40007 aligns points based on the Moulton code value and performs nearest neighbor search (NNS) using a depth-first traversal process. After the feature conversion operation, if nearest neighbor search (NNS) is required in other conversion processes for feature coding, a KD tree or Moulton code is used.

[0095] As shown, the transformed attributes are input to a RAHT transformer 40008 and / or an LOD generator 40009 .

[0096] The RAHT converter 40008 according to the embodiment performs RAHT coding to predict feature information based on the reconstructed geometry information. For example, the RAHT converter 40008 can predict feature information of a node at a higher level of the octree based on feature information associated with a node at a lower level of the octree.

[0097] According to the embodiment, the LOD generator 40009 generates LOD (Level of Detail). According to the embodiment, the LOD indicates the degree of detail of the point cloud content, and the smaller the LOD value, the lower the detail of the point cloud content, and the larger the LOD value, the higher the detail of the point cloud content. Points can be classified according to the LOD.

[0098] The lift transform unit 40010 according to the embodiment performs lift transform coding, which transforms the characteristics of the point cloud based on weights. As described above, the lift transform coding is selectively applied.

[0099] The coefficient quantization unit 40011 according to the embodiment quantizes the feature-coded feature based on the coefficients.

[0100] The arithmetic encoder 40012 according to the embodiment encodes the quantized characteristics based on arithmetic coding.

[0101] The elements of the point cloud video encoder of FIG. 4 may be implemented in hardware, software, firmware, or a combination thereof, including one or more processors or integrated circuits (not shown) configured to communicate with one or more memories included in the point cloud providing device. The one or more processors may perform any one of the operations and / or functions of the elements of the point cloud video encoder of FIG. 4 described above. The one or more processors may also operate or execute a software program and / or set of instructions to perform the operations and / or functions of the elements of the point cloud video encoder of FIG. 4. According to embodiments, the one or more memories may include high-speed random access memory or non-volatile memory (e.g., one or more magnetic disk storage devices, flash memory devices, or other non-volatile solid-state memory devices).

[0102] FIG. 5 is a diagram showing an example of a voxel according to the embodiment.

[0103] FIG. 5 shows voxels located in a three-dimensional space expressed by a coordinate system consisting of three axes, the X axis, the Y axis, and the Z axis. As shown in FIG. 4, a point cloud video encoder (e.g., the quantization unit 40001) performs voxelization. A voxel is a three-dimensional cubic space that is generated when the three-dimensional space is divided into units (unit=1.0) based on the axes (e.g., the X axis, the Y axis, and the Z axis) that represent the three-dimensional space. FIG. 5 shows two extreme points (0,0,0) and (2 d , 2 d , 2 d ) is an example of a voxel generated by an octree structure that recursively subdivides a bounding box (cubical axis-aligned bounding box) defined by the cubic axis-aligned bounding box (B). One voxel contains at least one point. The spatial coordinates of a voxel can be estimated from its positional relationship with other voxel groups. As mentioned above, a voxel has characteristics (such as color or reflectance) just like a pixel in a 2D image / video. A detailed explanation of voxels is omitted here as it has been explained in FIG. 4.

[0104] FIG. 6 is a diagram illustrating an example of an octree and occupancy code according to an embodiment.

[0105] As shown in Figures 1 to 4, the point cloud content providing system (point cloud video encoder 10002) or the octree analysis unit 40002 of the point cloud video encoder performs octree geometry coding (or octree coding) based on an octree structure to efficiently manage the area and / or position of voxels.

[0106] The upper part of Figure 6 shows an octree structure. The three-dimensional space of the point cloud content according to the embodiment is represented by the axes of a coordinate system (e.g., X-axis, Y-axis, Z-axis). The octree structure is defined by two extreme points (0,0,0) and (2 d , 2 d , 2 d ) is generated by recursively subdividing the cubic axis-aligned bounding box defined by the point cloud content (or point cloud video). 2d is set to the value that constitutes the smallest bounding box that encloses all points in the point cloud content (or point cloud video). The d value is determined by the following equation 1. In the following equation 1, (x int n , y int n , z int n ) indicates the quantized position (or position value) of the point.

[0107]

number

[0108] As shown in the upper center of Figure 6, the entire 3D space is divided into eight spaces through division. Each divided space is represented by a cube with six faces. As shown in the upper right of Figure 6, each of the eight spaces is again divided by the axes of the coordinate system (e.g., X-axis, Y-axis, Z-axis). Therefore, each space is again divided into eight smaller spaces. The divided smaller spaces are also represented by cubes with six faces. This division method is applied until the leaf nodes of the octree become voxels.

[0109] The bottom of Figure 6 shows an occupancy code for an octree. An occupancy code for an octree is generated to indicate whether each of the eight subspaces generated by dividing a space contains at least one point. Therefore, one occupancy code is represented by eight child nodes. Each child node indicates the occupancy of the divided space and has a 1-bit value. Therefore, the occupancy code is represented by an 8-bit code. That is, if the space corresponding to a child node contains at least one point, the corresponding node has a value of 1. If the space corresponding to a node does not contain any points (is empty), the corresponding node has a value of 0. The occupancy code shown in Figure 6 is 00100001, which indicates that the spaces corresponding to the third and eighth child nodes of the eight child nodes each contain at least one point. As shown in the figure, the third and eighth child nodes each have eight child nodes, and each child node is represented by an 8-bit occupancy code. In the drawing, the occupied code of the third child node is 10000111, and the occupied code of the eighth child node is 01001111. A point cloud video encoder (e.g., arithmetic encoder 40004) according to an embodiment can entropy encode the occupied code. To improve compression efficiency, the point cloud video encoder can also intra / inter-code the occupied code. A receiving device (e.g., receiving device 10004 or point cloud video decoder 10006) according to an embodiment reconstructs an octree based on the occupied code.

[0110] A point cloud video encoder (e.g., the octree analyzer 40002) according to an embodiment performs voxelization and octree coding to store the positions of points. However, since points in a 3D space are not always uniformly distributed, there may be certain areas where there are not many points. Therefore, performing voxelization on the entire 3D space is inefficient. For example, if there are almost no points in a certain area, there is no need to perform voxelization on that area.

[0111] Therefore, the point cloud video encoder according to the embodiment does not perform voxelization for the above-mentioned specific region (or nodes other than the leaf nodes of the octree), but performs direct coding, which directly codes the positions of points included in the specific region. The coordinates of the direct coding points according to the embodiment are called a direct coding mode (DCM). The point cloud video encoder according to the embodiment can also perform trisoup geometry encoding, which reconstructs the positions of points within a specific region (or node) based on voxels based on a surface model. Trisoup geometry encoding is a geometry encoding that represents an object as a series of triangle meshes. Therefore, the point cloud video decoder can generate a point cloud from the mesh surface. Direct coding and trisoup geometry encoding according to the embodiment can be selectively performed. Furthermore, direct coding and trisoup geometry encoding according to the embodiment can be combined with octree geometry coding (or octree coding).

[0112] To perform direct coding, a direct mode option for applying direct coding must be activated, and the node to which direct coding is applied must not be a leaf node, but must have points below a threshold within a specific node. Furthermore, the total number of points to be subjected to direct coding must not exceed a predetermined threshold. If these conditions are met, a point cloud video encoder (e.g., the computation encoder 40004) according to an embodiment can entropy code the positions (or position values) of the points.

[0113] A point cloud video encoder (e.g., the surface approximation analysis unit 40003) according to an embodiment can determine a specific level of the octree (if the level is smaller than the depth d of the octree) and perform trisoup geometry encoding (trisoup mode) from that level, which uses a surface model to reconstruct the positions of points within the node area based on voxels. The point cloud video encoder according to an embodiment can specify the level to which trisoup geometry encoding is applied. For example, if the specified level is the same as the depth of the octree, the point cloud video encoder does not operate in trisoup mode. That is, the point cloud video encoder according to an embodiment can operate in trisoup mode only when the specified level is smaller than the depth value of the octree. The 3D cubic area of ​​a node at a specified level according to an embodiment is called a block. One block includes one or more voxels. A block or voxel can also correspond to a brick. Geometry within each block is represented as a surface. According to an embodiment, a surface can intersect each edge of the block at most once.

[0114] Since one block has 12 edges, there are at least 12 intersections within one block. Each intersection is called a vertex. A vertex along an edge is detected if there is at least one occupied voxel adjacent to that edge among all blocks that share the edge. In this embodiment, an occupied voxel refers to a voxel that contains a point. The position of a vertex detected along an edge is the average position along the edge of all voxels adjacent to that edge among all blocks that share the edge.

[0115] When a vertex is detected, the point cloud video encoder according to the embodiment may entropy code the edge start point (x, y, z), the edge direction vector (Δx, Δy, Δz), and the vertex position value (relative position value within the edge). When trisoup geometry coding is applied, the point cloud video encoder according to the embodiment (e.g., the geometry reconstruction unit 40005) may perform triangle reconstruction, up-sampling, and voxelization processes to generate a restored geometry (reconstructed geometry).

[0116] The vertices located on the edge of a block determine the surface that passes through the block. In this embodiment, the surface is a non-planar polygon. In the triangulation process, the surface represented by triangles is reconstructed based on the start point of the edge, the edge direction vector, and the vertex position values. The triangulation process is as shown in Equation 2 below: (1) calculate the centroid value of each vertex, (2) subtract the centroid value from each vertex value, and (3) square the result and add up all the resulting values.

[0117]

number

[0118] Next, the minimum value of the added values ​​is found, and a projection process is performed along the axis where the minimum value is. For example, if the x element is the smallest, each vertex is projected onto the x-axis based on the center of the block, and then onto the (y,z) plane. If the value obtained by projecting onto the (y,z) plane is (ai,bi), the θ value is found using atan2(bi,ai), and the vertices are aligned based on the θ value. Table 1 below shows the vertex combinations used to create triangles depending on the number of vertices. Vertices are aligned in order from 1 to n. Table 1 below shows that for four vertices, two triangles are formed by combining the vertices. The first triangle is formed by the first, second, and third vertices of the aligned vertices, and the second triangle is formed by the third, fourth, and first vertices of the aligned vertices.

[0119] Table 1.Triangles formed from vertices ordered 1,...,n

[0120] [Table 1]

[0121] The upsampling process is performed to add intermediate points along the edges of triangles for voxelization. The additional points are generated based on the upsampling factor and the block width. The additional points are called refined vertices. A point cloud video encoder according to an embodiment can voxelize the refined vertices. The point cloud video encoder can also perform feature encoding based on the voxelized positions (or position values).

[0122] FIG. 7 is a diagram illustrating an example of an adjacent node pattern according to the embodiment.

[0123] To increase the compression efficiency of the point cloud video, the point cloud video encoder according to the embodiment performs entropy coding based on context adaptive arithmetic coding.

[0124] As described with reference to FIGS. 1 to 6, the point cloud content providing system or the point cloud video encoder 10002 of FIG. 2 or the point cloud video encoder or arithmetic encoder 40004 of FIG. 4 can immediately entropy code the occupied code. The point cloud content providing system or the point cloud video encoder can also perform entropy coding (intra coding) based on the occupied code of the current node and the occupied rate of neighboring nodes, or can perform entropy coding (inter coding) based on the occupied code of a previous frame. A frame according to the present embodiment refers to a collection of point cloud videos generated at the same time. The compression efficiency of intra coding / inter coding according to the present embodiment varies depending on the number of neighboring nodes referenced. Although the complexity increases as the number of bits increases, the compression efficiency can be improved by focusing on one side. For example, a 3-bit context requires eight coding methods, which is 2^3. The separately coded portion affects the complexity of the implementation. Therefore, it is necessary to balance compression efficiency and complexity at appropriate levels.

[0125] FIG. 7 shows a process for determining an occupancy pattern based on the occupancy of neighboring nodes. A point cloud video encoder according to an embodiment obtains a neighbor pattern value by determining the occupancy of neighboring nodes for each node in an octree. The neighboring node pattern is used to infer the occupancy pattern of the corresponding node. The left side of FIG. 7 shows a cube corresponding to the node (the cube located in the middle) and six cubes (neighboring nodes) that share at least one side with the corresponding cube. The illustrated nodes are nodes at the same depth. The illustrated numbers indicate weights (1, 2, 4, 8, 16, 32, etc.) associated with each of the six nodes. Each weight is assigned in order according to the position of the neighboring node.

[0126] The right side of FIG. 7 shows the neighboring node pattern value. The neighboring node pattern value is the sum of values ​​multiplied by the weight values ​​of occupied neighboring nodes (neighboring nodes with points). Therefore, the neighboring node pattern value ranges from 0 to 63. An neighboring node pattern value of 0 means that there are no nodes with points (occupied nodes) among the neighboring nodes of the corresponding node. An neighboring node pattern value of 63 means that all neighboring nodes are occupied nodes. As shown in the figure, neighboring nodes assigned weight values ​​of 1, 2, 4, and 8 are occupied nodes, so the neighboring node pattern value is 15, which is the sum of 1, 2, 4, and 8. The point cloud video encoder can perform coding according to the neighboring node pattern value (e.g., if the neighboring node pattern value is 63, 64 coding is performed). In an embodiment, the point cloud video encoder can reduce coding complexity by changing the neighboring node pattern value (e.g., based on a table that changes 64 to 10 or 6).

[0127] FIG. 8 is a diagram illustrating an example of a point configuration for each LOD according to the embodiment.

[0128] As described in Figures 1 to 7, before feature encoding, the coded geometry is reconstructed (restored). When direct coding is applied, the geometry reconstruction operation involves changing the placement of the direct coded points (e.g., placing the direct coded points in front of the point cloud data). When trisoup geometry encoding is applied, the geometry reconstruction process involves the processes of triangulation, upsampling, and voxelization. Since features are dependent on the geometry, feature encoding is performed based on the reconstructed geometry.

[0129] A point cloud video encoder (e.g., the LOD generator 40009) can reorganize or group points by LOD. FIG. 8 shows point cloud content corresponding to LOD. The leftmost part of FIG. 8 shows the original point cloud content. The second from the left in FIG. 8 shows the distribution of points with the lowest LOD, and the rightmost part in FIG. 8 shows the distribution of points with the highest LOD. That is, the points with the lowest LOD have a sparse distribution, and the points with the highest LOD have a fine distribution. That is, as the LOD increases along the arrow direction shown at the bottom of FIG. 8, the interval (or distance) between points becomes shorter.

[0130] FIG. 9 is a diagram illustrating an example of a point configuration for each LOD according to the embodiment.

[0131] As described in FIGS. 1 to 8, a point cloud content providing system or a point cloud video encoder (e.g., point cloud video encoder 10002 of FIG. 2, point cloud video encoder or LOD generator 40009 of FIG. 4) generates LOD. The LOD is generated by rearranging points into a set of refinement levels according to a set LOD distance value (or a set of Euclidean distances). The LOD generation process is performed not only in the point cloud video encoder but also in the point cloud video decoder.

[0132] The upper part of Figure 9 shows an example of points (P0 to P9) of point cloud content distributed in 3D space. The original order in Figure 9 shows the order of points P0 to P9 before LOD generation. The LOD-based order in Figure 9 shows the order of points after LOD generation. Points are re-sorted for each LOD. Also, higher LODs include points belonging to lower LODs. As shown in Figure 9, LOD0 includes P0, P5, P4, and P2. LOD1 includes points of LOD0, P1, P6, and P3. LOD2 includes points of LOD0, points of LOD1, and P9, P8, and P7.

[0133] As described in FIG. 4, the point cloud video encoder according to the embodiment can selectively or in combination perform LOD-based predictive transform coding, lift transform coding, and RAHT transform coding.

[0134] A point cloud video encoder according to an embodiment generates a predictor for each point and performs LOD-based predictive conversion coding to set a prediction characteristic (or prediction characteristic value) for each point. That is, N predictors are generated for N points. The predictor according to an embodiment can calculate a weight (=1 / distance) based on the LOD value of each point, index information for neighboring points within a distance set for each LOD, and the distance value to the neighboring point.

[0135] According to an embodiment, the predicted feature (or feature value) is set as the average value of the features (or feature values, e.g., hue, reflectance, etc.) of neighboring points set in the predictor of each point multiplied by a weight (or weight value) calculated based on the distance to each neighboring point. A point cloud video encoder (e.g., coefficient quantization unit 40011) according to an embodiment may quantize and inverse quantize a residual value (also called a residual feature, residual feature value, feature prediction residual value, prediction error feature value, etc.) of a corresponding point, which is obtained by subtracting the corresponding predicted feature (feature value) from the feature (i.e., original feature value) of the corresponding point. The quantization process performed by the transmitter on the residual feature value is shown in Table 2. The inverse quantization process performed by the receiver on the quantized residual feature value as shown in Table 2 is shown in Table 3.

[0136] [Table 2]

[0137] [Table 3]

[0138] A point cloud video encoder (e.g., the computation encoder 40012) according to an embodiment performs entropy coding on the quantized and dequantized residual feature values ​​as described above if there are neighboring points in the predictor of each point. If there are no neighboring points in the predictor of each point, the point cloud video encoder (e.g., the computation encoder 40012) according to an embodiment performs entropy coding on the feature of the corresponding point without performing the above-described process. A point cloud video encoder (e.g., the lift transform unit 40010) according to an embodiment generates a predictor for each point, sets the calculated LOD in the predictor, registers neighboring points, and sets weights according to the distances to the neighboring points to perform lift transform coding. Lift transform coding according to an embodiment is similar to the above-described lift transform coding, but differs in that weights are cumulatively applied to feature values. The process of cumulatively applying weights to feature values ​​according to an embodiment is as follows.

[0139] 1) Create an array QW (QuantizationWeight) that stores the weight value of each point. The initial value of all elements of QW is 1.0. The QW value of the predictor index of the adjacent node registered in the predictor is multiplied by the weight value of the predictor of the current point and added.

[0140] 2) Lift prediction process: To calculate the predicted attribute value, the attribute value of the point is multiplied by the weight and subtracted from the existing attribute value.

[0141] 3) Create temporary arrays called updateweight and update, and initialize the temporary arrays to 0.

[0142] 4) The calculated weights for all predictors are multiplied by the weights stored in the QW corresponding to the predictor index, and the resulting weights are accumulated and summed in the update weight array as the index of the adjacent node. The update array is accumulated and summed by multiplying the calculated weights by the characteristic values ​​of the index of the adjacent node.

[0143] 5) Lift update process: For every predictor, divide the feature value in the update array by the weight value in the update weight array of the predictor index, and add the result to the existing feature value again.

[0144] 6) For all predictors, the feature values ​​updated in the lift update process are further multiplied by the weights updated (stored in the QW) in the lift prediction process to calculate predicted feature values. According to an embodiment, a point cloud video encoder (e.g., coefficient quantization unit 40011) quantizes the predicted feature values. Furthermore, a point cloud video encoder (e.g., arithmetic encoder 40012) entropy codes the quantized feature values.

[0145] A point cloud video encoder (e.g., the RAHT transform unit 40008) according to an embodiment performs RAHT transform coding, which predicts the characteristics of higher-level nodes using characteristics associated with lower-level nodes in an octree. RAHT transform coding is an example of characteristic intra-coding using octree backward scanning. A point cloud video encoder according to an embodiment scans the entire region from a voxel, and at each step, repeats a merging process up to the root node while combining the voxels into larger blocks. The merging process according to an embodiment is performed only on occupied nodes. The merging process is not performed on empty nodes, but is performed on the nodes immediately above the empty nodes.

[0146] TIFF2025188208000007.tif47159

[0147]

number

[0148] TIFF2025188208000009.tif55160

[0149]

number

[0150] The gDC values ​​are also quantized and entropy coded like the high-pass coefficients.

[0151] FIG. 10 is a diagram illustrating an example of a point cloud video decoder according to an embodiment.

[0152] The point cloud video decoder shown in FIG. 10 is an example of the point cloud video decoder 10006 shown in FIG. 1 and performs operations that are the same as or similar to those of the point cloud video decoder 10006 described in FIG. 1. As shown, the point cloud video decoder receives a geometry bitstream and an attribute bitstream included in one or more bitstreams. The point cloud video decoder includes a geometry decoder and an attribute decoder. The geometry decoder performs geometry decoding on the geometry bitstream and outputs decoded geometry. The attribute decoder performs attribute decoding on the attribute bitstream based on the decoded geometry and outputs decoded attributes. The decoded geometry and decoded attributes are used to restore point cloud content (decoded point cloud).

[0153] FIG. 11 illustrates an example of a point cloud video decoder according to an embodiment.

[0154] The point cloud video decoder shown in FIG. 11 is an example of the point cloud video decoder described in FIG. 10, and performs a decoding operation that is the reverse process of the encoding operation of the point cloud video encoder described in FIGS.

[0155] As explained in Figures 1 and 10, the point cloud video decoder performs geometry decoding and feature decoding, where geometry decoding is performed before feature decoding.

[0156] The point cloud video decoder according to the embodiment includes an arithmetic decoder (11000), an octree synthesis unit (11001), a surface approximation synthesis unit (11002), a geometry reconstruction unit (11003), a coordinates inverse transformation unit (11004), an arithmetic decoder (11005), an inverse quantization unit (11006), a RAHT transform unit 11007, an LOD generation unit (11008), an inverse lifting unit (11009), and / or a color inverse transformation unit (11010).

[0157] The arithmetic decoder 11000, octree synthesis unit 11001, surface approximation synthesis unit 11002, geometry reconstruction unit 11003, and coordinate system inverse transformation unit 11004 perform geometry decoding. Geometry decoding according to this embodiment includes direct decoding and trisoup geometry decoding. Direct decoding and trisoup geometry decoding are selectively applied. Furthermore, geometry decoding is not limited to the above examples, and is performed by the reverse process of the geometry encoding described with reference to FIGS. 1 to 9.

[0158] The arithmetic decoder 11000 according to the embodiment decodes the received geometry bitstream based on arithmetic coding. The operation of the arithmetic decoder 11000 corresponds to the reverse process of the arithmetic encoder 40004.

[0159] The octree synthesis unit 11001 according to the embodiment obtains exclusive codes from the decoded geometry bitstream (or from the decoded result, information on the secured geometry) to generate an octree. Specific details regarding exclusive codes are as described in FIGS. 1 to 9.

[0160] The surface approximation synthesis unit 11002 according to the embodiment synthesizes a surface based on the decoded geometry and / or the generated octree if trisoup geometry encoding is applied.

[0161] According to the embodiment, the geometry reconstruction unit 11003 regenerates geometry based on the surface and / or decoded geometry. As described with reference to FIGS. 1 to 9, direct coding and trisoup geometry coding are selectively applied. Therefore, the geometry reconstruction unit 11003 directly retrieves and adds position information of points to which direct coding is applied. Also, when trisoup geometry coding is applied, the geometry reconstruction unit 11003 reconstructs geometry by performing reconstruction operations of the geometry reconstruction unit 40005, such as triangulation, upsampling, and voxelization operations. The detailed contents are the same as those described with reference to FIG. 6, and therefore will not be repeated. The reconstructed geometry includes a point cloud picture or frame that does not include features.

[0162] The coordinate system inverse transform unit 11004 according to the embodiment transforms the coordinate system based on the reconstructed geometry to obtain the position of the point.

[0163] The arithmetic decoder 11005, the inverse quantization unit 11006, the RAHT transform unit 11007, the LOD generation unit 11008, the inverse lift unit 11009, and / or the color inverse transform unit 11010 perform the feature decoding described in FIG. 10. Feature decoding according to the embodiment includes RAHT (Region Adaptive Hierarchical Transform) decoding, Interpolarization-based hierarchical nearest-neighbor prediction-Prediction Transform (Interpolarization-based hierarchical nearest-neighbor prediction with an update / lifting step (Lifting Transform)) decoding. The above three decoding methods may be used selectively, or a combination of one or more of them may be used. Also, feature decoding according to the embodiment is not limited to the above examples.

[0164] The arithmetic decoder 11005 according to the embodiment decodes the attribute bitstream into arithmetic coding.

[0165] The inverse quantization unit 11006 according to the embodiment inverse quantizes the decoded feature bitstream or the feature information obtained as a result of decoding, and outputs the inverse quantized feature (or feature value). The inverse quantization is selectively applied based on the feature encoding of the point cloud video encoder.

[0166] In some embodiments, the RAHT transform unit 11007, LOD generator 11008, and / or inverse lifting unit 11009 process the reconstructed geometry and dequantized features. As described above, the RAHT transform unit 11007, LOD generator 11008, and / or inverse lifting unit 11009 selectively perform corresponding decoding operations according to the encoding of the point cloud video encoder.

[0167] According to an embodiment, the color inverse transform unit 11010 performs inverse transform coding to inversely transform the color values ​​(or textures) included in the decoded features. The operation of the color inverse transform unit 11010 is selectively performed based on the operation of the color transform unit 40006 of the point cloud video encoder.

[0168] The elements of the point cloud video decoder of Figure 11 may be embodied in hardware, software, firmware, or a combination thereof, including one or more processors or integrated circuits (not shown) configured to communicate with one or more memories included in the point cloud providing device. The one or more processors perform any of the operations and / or functions of the elements of the point cloud video decoder of Figure 11 described above. The one or more processors also operate or execute a software program and / or set of instructions to perform the operations and / or functions of the elements of the point cloud video decoder of Figure 11.

[0169] FIG. 12 shows an example of a transmitting device according to the embodiment.

[0170] The transmitting device shown in Fig. 12 is an example of the transmitting device 10000 of Fig. 1 (or the point cloud video encoder of Fig. 4). The transmitting device shown in Fig. 12 performs any of the same or similar operations and methods as those of the operation and encoding method of the point cloud video encoder described with reference to Figs. 1 to 9. The transmitting device according to the embodiment includes a data input unit 12000, a quantization processing unit 12001, a voxelization processing unit 12002, an octree occupation code generation unit 12003, a surface model processing unit 12004, an intra / inter coding processing unit 12005, an arithmetic coder 12006, a metadata processing unit 12007, a hue conversion processing unit 12008, a characteristic conversion processing unit (or attribute conversion processing unit) 12009, a prediction / lift / RAHT conversion processing unit 12010, an arithmetic coder 12011, and / or a transmission processing unit 12012.

[0171] According to the embodiment, the data input unit 12000 receives or acquires point cloud data. The data input unit 12000 performs operations and / or acquisition methods that are the same as or similar to the operations and / or acquisition methods of the point cloud video acquisition unit 10001 (or the acquisition process 20000 shown in FIG. 2).

[0172] Geometry encoding is performed by a data input unit 12000, a quantization unit 12001, a voxelization unit 12002, an octree occupation code generation unit 12003, a surface model processing unit 12004, an intra / inter coding processing unit 12005, and an arithmetic coder 12006. The geometry encoding according to this embodiment is the same as or similar to the geometry encoding described with reference to Figures 1 to 9, and therefore a detailed description thereof will be omitted.

[0173] The quantization unit 12001 according to the embodiment quantizes geometry (e.g., position values ​​of points). The operation and / or quantization of the quantization unit 12001 is the same as or similar to the operation and / or quantization of the quantization unit 40001 shown in Fig. 4. The specific description is as described in Figs. 1 to 9.

[0174] The voxelization processing unit 12002 according to the embodiment voxels the position values ​​of the quantized points. The voxelization processing unit 12002 performs operations and / or processes that are the same as or similar to the operations and / or voxelization processes of the quantization unit 40001 shown in Fig. 4. Specific details are as described in Figs. 1 to 9.

[0175] The octree occupation code generator 12003 according to the embodiment performs octree coding on the positions of voxelized points based on an octree structure. The octree occupation code generator 12003 generates occupation codes. The octree occupation code generator 12003 performs operations and / or methods that are the same as or similar to the operations and / or methods of the point cloud encoder (or the octree analyzer 40002) described in FIGS. 4 and 6. Specific descriptions are as described in FIGS. 1 to 9.

[0176] The surface model processing unit 12004 according to this embodiment performs trisoup geometry encoding, which reconstructs the positions of points within a specific region (or node) on a voxel basis based on a surface model. The surface model processing unit 12004 performs operations and / or methods that are the same as or similar to those of the point cloud video encoder (e.g., the surface approximation analysis unit 40003) shown in Figure 4. Specific details are as described in Figures 1 to 9.

[0177] According to the embodiment, the intra / inter coding processor 12005 performs intra / inter coding of the point cloud data. The intra / inter coding processor 12005 performs coding that is the same as or similar to the intra / inter coding described in FIG. 7, as detailed in FIG. 7. In the embodiment, the intra / inter coding processor 12005 is included in the arithmetic coder 12006.

[0178] According to an embodiment, the arithmetic coder 12006 entropy encodes the octree and / or approximated octree of the point cloud data. For example, the encoding method includes an arithmetic encoding method. The arithmetic coder 12006 performs operations and / or methods that are the same as or similar to those of the arithmetic encoder 40004.

[0179] The metadata processing unit 12007 according to the embodiment processes metadata related to point cloud data, such as setting values, and provides the metadata to necessary processing steps such as geometry encoding and / or feature encoding. The metadata processing unit 12007 according to the embodiment also generates and / or processes signaling information related to geometry encoding and / or feature encoding. The signaling information according to the embodiment is encoded separately from the geometry encoding and / or feature encoding. The signaling information according to the embodiment may also be interleaved.

[0180] The hue conversion processor 12008, the feature conversion processor 12009, the prediction / lift / RAHT conversion processor 12010, and the arithmetic coder 12011 perform feature coding. The feature coding according to the embodiment is the same as or similar to the feature coding described with reference to FIGS. 1 to 9, so a detailed description thereof will be omitted.

[0181] The color conversion unit 12008 according to this embodiment performs color conversion coding to convert the hue value included in the feature. The color conversion unit 12008 performs color conversion coding based on the reconstructed geometry. The reconstructed geometry has been described with reference to FIGS. 1 to 9. The color conversion unit 12008 also performs operations and / or methods that are the same as or similar to the operations and / or methods of the color conversion unit 40006 described with reference to FIG. 4. Detailed description thereof will be omitted.

[0182] The feature conversion processor 12009 according to the embodiment performs feature conversion, which converts features based on positions where geometry coding has not been performed and / or reconstructed geometry. The feature conversion processor 12009 performs operations and / or methods identical to or similar to those of the feature conversion unit 40007 described in FIG. 4, detailed descriptions of which will be omitted. The prediction / lift / RAHT conversion processor 12010 according to the embodiment codes the converted features using any one or a combination of RAHT coding, predictive transformation coding, and lift transformation coding. The prediction / lift / RAHT conversion processor 12010 performs any one of operations identical to or similar to those of the RAHT conversion unit 40008, LOD generation unit 40009, and lift transformation unit 40010 described in FIG. 4. Since the predictive transformation coding, lift transformation coding, and RAHT transformation coding have been described with reference to FIGS. 1 to 9, detailed descriptions of these will be omitted.

[0183] The arithmetic coder 12011 according to the embodiment encodes the coded characteristic based on arithmetic coding. The arithmetic coder 12011 performs operations and / or methods that are the same as or similar to the operations and / or methods of the arithmetic encoder 40012.

[0184] The transmission processing unit 12012 according to the embodiment transmits each bitstream including the coded geometry and / or coded attribute and metadata information, or transmits the coded geometry and / or coded attribute and metadata information in one bitstream. When the coded geometry and / or coded attribute and metadata information according to the embodiment is configured in one bitstream, the bitstream includes one or more sub-bitstreams. The bitstream according to the embodiment includes signaling information including a Sequence Parameter Set (SPS) for sequence-level signaling, a Geometry Parameter Set (GPS) for signaling geometry information coding, an Attribute Parameter Set (APS) for signaling attribute information coding, and a Tile Parameter Set (TPS) for tile-level signaling, and slice data. The slice data includes information about one or more slices. According to the embodiment, one slice is one geometry bitstream (Geometry 0). 0 ) and one or more attribute bitstreams (Attr0 0 , Attr1 0 ) is included.

[0185] According to an embodiment, a TPS includes information about each of one or more tiles (for example, coordinate information and height / size information of a bounding box).

[0186] The geoheader bitstream includes a header and a payload. The header of the geometry bitstream according to the embodiment includes identification information of a parameter set included in the GPS (geom_parameter_set_id), a tile identifier (geom_tile_id), a slice identifier (geom_slice_id), and information about the data included in the payload. As described above, the metadata processing unit 12007 according to the embodiment can generate and / or process signaling information and transmit it to the transmission processing unit 12012. In the embodiment, the element performing geometry coding and the element performing attribute coding can share data / information with each other as indicated by the dotted lines. The transmission processing unit 12012 according to the embodiment performs an operation and / or a transmission method that is the same as or similar to the operation and / or transmission method of the transmitter 10003. Detailed description thereof will be omitted as it is the same as that described with reference to FIGS. 1 and 2.

[0187] FIG. 13 shows an example of a receiving device according to the embodiment.

[0188] The receiving device shown in Figure 13 is an example of the receiving device 10004 in Figure 1 (or the point cloud video decoder in Figures 10 and 11). The receiving device shown in Figure 13 performs any of the same or similar operations and methods as the operations and decoding methods of the point cloud video decoder described in Figures 1 to 11.

[0189] The receiving device according to the embodiment includes a receiving unit 13000, a receiving processing unit 13001, an arithmetic decoder 13002, an occupancy code-based octree reconstruction processing unit 13003, a surface model processing unit (triangle reconstruction, upsampling, voxelization) 13004, an inverse quantization processing unit 13005, a metadata analysis 13006, an arithmetic decoder 13007, an inverse quantization processing unit 13008, a prediction / lift / RAHT inverse transform processing unit 13009, a color inverse transform processing unit 13010, and / or a renderer 13011. Each component of the decoding according to the embodiment performs the reverse process of the component of the encoding according to the embodiment.

[0190] The receiving unit 13000 according to the embodiment receives point cloud data. The receiving unit 13000 performs operations and / or a receiving method that are the same as or similar to the operations and / or a receiving method of the receiver 10005 of Fig. 1. Detailed description thereof will be omitted.

[0191] The receiving unit 13001 according to the embodiment obtains a geometry bitstream and / or a characteristic bitstream from the received data. The receiving unit 13000 includes the receiving unit 13000.

[0192] Geometry decoding is performed by an arithmetic decoder 13002, an exclusive code based octree reconstruction processor 13003, a surface model processor 13004, and an inverse quantization processor 13005. The geometry decoding according to this embodiment is the same as or similar to the geometry decoding described with reference to Figures 1 to 10, so a detailed description thereof will be omitted.

[0193] The arithmetic decoder 13002 according to the embodiment decodes the geometry bitstream based on arithmetic coding. The arithmetic decoder 13002 performs operations and / or coding that are the same as or similar to the operations and / or coding of the arithmetic decoder 11000.

[0194] The exclusive code-based octree reconstruction processor 13003 according to the embodiment obtains exclusive codes from the decoded geometry bitstream (or information on the secured geometry as a result of decoding) and reconstructs an octree. The exclusive code-based octree reconstruction processor 13003 performs operations and / or methods identical to or similar to the operations and / or octree generation method of the octree synthesis unit 11001. The surface model processor 13004 according to the embodiment performs trisoup geometry decoding and associated geometry reconstruction (e.g., triangle reconstruction, upsampling, voxelization) based on a surface model scheme when trisoup geometry encoding is applied. The surface model processor 13004 performs operations identical to or similar to the operations of the surface approximation synthesis unit 11002 and / or the geometry reconstruction unit 11003.

[0195] The inverse quantization unit 13005 according to the embodiment inverse quantizes the decoded geometry.

[0196] According to an embodiment, the metadata analysis 13006 analyzes metadata, such as setting values, included in the received point cloud data. The metadata analysis 13006 transmits the metadata to geometry decoding and / or feature decoding. A detailed description of the metadata is omitted here as it has been described with reference to FIG.

[0197] The arithmetic decoder 13007, the inverse quantization processor 13008, the prediction / lift / RAHT inverse transform processor 13009, and the hue inverse transform processor 13010 perform feature decoding. The feature decoding is the same as or similar to the feature decoding described in Figures 1 and 10, so a detailed description thereof will be omitted.

[0198] The operation decoder 13007 according to the embodiment decodes the attribute bitstream into operation coding. The operation decoder 13007 performs decoding of the attribute bitstream based on the reconstructed geometry. The operation decoder 13007 performs operations and / or coding that are the same as or similar to the operations and / or coding of the operation decoder 11005.

[0199] The inverse quantization unit 13008 according to the embodiment inversely quantizes the decoded characteristic bitstream. The inverse quantization unit 13008 performs the same or similar operations and / or methods as the inverse quantization unit 11006.

[0200] The prediction / lift / RAHT inverse transform processing unit 13009 according to the embodiment processes the reconstructed geometry and dequantized features. The prediction / lift / RAHT inverse transform processing unit 13009 performs any of the same or similar operations and / or decoding as those of the RAHT transform unit 11007, the LOD generation unit 11008, and / or the inverse lift unit 11009. The hue inverse transform processing unit 13010 according to the embodiment performs inverse transform coding to inversely transform the color values ​​(or textures) included in the decoded features. The hue inverse transform processing unit 13010 performs the same or similar operations and / or inverse transform coding as those of the color inverse transform unit 11010. The renderer 13011 according to the embodiment renders point cloud data.

[0201] FIG. 14 illustrates an example of a structure that can be linked to a point cloud data transmission / reception method / apparatus according to an embodiment.

[0202] 14 illustrates a configuration in which any of a server 17600, a robot 17100, an autonomous vehicle 17200, an XR device 17300, a smartphone 17400, a home appliance 17500, and / or an HMD (Head-Mounted Display) 17700 is connected to a cloud network 17100. The robot 17100, the autonomous vehicle 17200, the XR device 17300, the smartphone 17400, or the home appliance 17500 may also be referred to as a device. The XR device 17300 corresponds to a point cloud data (PCC) device according to an embodiment or is linked to a PCC device.

[0203] Cloud network 17000 refers to a network that forms part of a cloud computing infrastructure or exists within a cloud computing infrastructure, where cloud network 17000 is configured using a 3G network, a 4G or LTE network, a 5G network, or the like.

[0204] The server 17600 is connected to any of the robot 17100, autonomous vehicle 17200, XR device 17300, smartphone 17400, home appliance 17500, and / or HMD 17700 via a cloud network 17000 and can assist with at least some of the processing of the connected devices 17100-17700.

[0205] An HMD (Head-Mounted Display) 17700 represents any of the types of XR devices and / or PCC devices according to the embodiment. An HMD type device according to the embodiment includes a communications unit, a control unit, a memory unit, an I / O unit, a sensor unit, and a power supply unit.

[0206] Various embodiments of the devices 17100 to 17500 to which the above-described technology is applied will be described below. Here, the devices 17100 to 17500 shown in Fig. 14 can be linked / coupled to the point cloud data transmitting / receiving device according to the above-described embodiments.

[0207] <PCC+XR>

[0208] The XR / PCC device 17300 may be implemented by applying PCC and / or XR (AR+VR) technology to an HMD (Head-Mounted Display), a HUD (Head-Up Display) installed in a vehicle, a TV, a mobile phone, a smartphone, a computer, a wearable device, a home appliance, a digital sign, a vehicle, a fixed robot, a mobile robot, etc.

[0209] The XR / PCC device 17300 can obtain information about the surrounding space or real objects by analyzing 3D point cloud data or image data acquired by various sensors or from an external device to generate position data and attribute data for 3D points, and can render and output the XR object to be output. For example, the XR / PCC device 17300 can output an XR object including additional information about the recognized object corresponding to the recognized object.

[0210] <PCC+Self-propelled+XR>

[0211] The autonomous vehicle 17200 is realized as a mobile robot, vehicle, unmanned aerial vehicle, etc. by applying PCC technology and XR technology.

[0212] The autonomous vehicle 17200 to which XR / PCC technology is applied refers to an autonomous vehicle equipped with a means for providing XR images, an autonomous vehicle that is the target of control / interaction within XR images, etc. In particular, the autonomous vehicle 17200 that is the target of control / interaction within XR images can be separated from the XR device 17300 and linked to each other.

[0213] The autonomous vehicle 17200, which is equipped with a means for providing XR / PCC images, obtains sensor information from sensors including a camera and outputs XR / PCC images generated based on the obtained sensor information. For example, the autonomous vehicle 17200 is equipped with a HUD and outputs XR / PCC images, thereby providing passengers with XR / PCC objects corresponding to real objects or objects on a screen.

[0214] At this time, when the XR / PCC object is output to the HUD, at least a portion of the XR / PCC object is output to overlap with an actual object toward which the passenger's gaze is directed. On the other hand, when the XR / PCC object is output to a display provided in the autonomous vehicle, at least a portion of the XR / PCC object is output to overlap with an object on the screen. For example, the autonomous vehicle 1220 may output XR / PCC objects corresponding to objects such as a roadway, another vehicle, a traffic light, a traffic sign, a motorcycle, a pedestrian, a building, etc.

[0215] The VR (Virtual Reality) technology, AR (Augmented Reality) technology, MR (Mixed Reality) technology, and / or PCC (Point Cloud Compression) technology according to the embodiments can be applied to various devices.

[0216] In other words, VR technology is a display technology that presents real objects and backgrounds only as CG images. In contrast, AR technology is a technology that displays virtual CG images on top of images of real things. MR technology is similar to AR technology in that it mixes virtual objects into the real world. However, AR technology clearly distinguishes between real objects and virtual objects made of CG images, and uses virtual objects to complement real objects, while MR technology is different from AR technology in that virtual objects and real objects are considered to have the same characteristics. More specifically, for example, hologram services are an application of the MR technology.

[0217] However, recently, rather than clearly distinguishing between VR, AR, and MR technologies, they are also referred to as XR (extended reality) technologies. Therefore, embodiments of the present invention can be applied to any of VR, AR, MR, and XR technologies. These technologies apply encoding / decoding based on PCC, V-PCC, and G-PCC technologies.

[0218] The PCC method / apparatus according to the embodiment can be applied to vehicles that provide autonomous driving services.

[0219] Vehicles that provide autonomous driving services are connected to the PCC device via wired or wireless communication.

[0220] When a point cloud data (PCC) transceiver according to an embodiment is connected to a vehicle via wired or wireless communication, it can receive / process content data related to AR / VR / PCC services that can be provided along with an autonomous driving service and transmit the content data to the vehicle. When the point cloud data transceiver is installed in a vehicle, it can receive / process content data related to AR / VR / PCC services according to a user input signal input through a user interface device and provide the content data to a user. The vehicle or user interface device according to an embodiment receives a user input signal. The user input signal according to an embodiment includes a signal instructing an autonomous driving service.

[0221] Meanwhile, as mentioned above, the point cloud content providing system can use one or more cameras (e.g., an infrared camera that obtains depth information, an RGB camera that extracts color information corresponding to the depth information, etc.), projectors (e.g., an infrared pattern projector that obtains depth information), LiDAR, etc. to generate point cloud content (or point cloud data).

[0222] Lidar, a device that measures distance by measuring the time it takes for light to reflect off an object and return, provides precise 3D information of the real world over a wide area and long distances as point cloud data. This large amount of point cloud data can be widely used in various fields that use computer vision technology, such as autonomous vehicles, robots, and 3D map creation. To generate point cloud content, Lidar equipment uses a radar system that emits a laser pulse, measures the time it takes for the pulse to reflect off an object (i.e., a reflector) and return, and determines the position coordinates of the reflector. According to an embodiment, depth information can be extracted by the Lidar equipment. Furthermore, point cloud content generated by the Lidar equipment may consist of multiple frames, and multiple frames may be integrated into a single content.

[0223] This lidar is located at different elevations θ(i) i=1,…,N The lasers spin along an azimuth angle φ based on the Z axis, and capture point cloud data as shown in Figure 15(a) and / or Figure 15(b). This type is called a spinning LiDAR model, and the point cloud content captured and generated by the spinning LiDAR model has angular characteristics.

[0224] 15(a) and 15(b) are diagrams illustrating an example of a spinning rider learning model according to an embodiment.

[0225] 15(a) and 15(b), a laser i hits an object M, and the position of M can be estimated as (x, y, z) on a Cartesian coordinate system. In this case, due to the fixed position of the laser sensor, its straight forward movement characteristics, and its characteristic of rotating at a predetermined azimuth angle, when the position of the object M is represented as (r, φ, i) on the Cartesian coordinate system instead of (x, y, z), a rule between points can be advantageously induced for compression.

[0226] Therefore, by leveraging such characteristics, in the case of data captured by a spinning lidar, if the angular mode is applied in the process of geometric encoding / decoding, the compression efficiency will be even higher. The angular mode is a method of compression using (r, φ, i) instead of (x, y, z). Here, r represents the radius, φ represents the azimuth or azimuthal angle, and i represents the i-th laser of the lidar (e.g., laser index). That is, the frames of the point cloud content generated by the lidar equipment are not combined but are individual frames, and since each origin is 0, 0, 0, it is possible to change to the spherical coordinate system and use the angular mode.

[0227] According to an embodiment, when a moving / stopped automobile captures a point cloud by a lidar, the angular mode (r, φ, i) can be utilized. In this case, for the same azimuth φ, the larger the radius r, the longer the arc. For example, as shown in Fig. 16(a), when the radius r1 < r2 for the same azimuth φ, the arc arc1 < arc2.

[0228] Figs. 16(a) and 16(b) are diagrams showing an example of comparing the lengths of arcs with the same azimuth from the center of an automobile according to an embodiment.

[0229] In other words, when using angle mode, the point cloud content captured by the lidar will remain within the same azimuth angle even if it moves farther away from the capture device, meaning that the movement of objects in closer areas can be captured better. In other words, objects in closer areas (i.e., objects closer to the center) can be captured better even if they move only slightly because the azimuth angle is larger. Also, objects farther from the center will appear to move only slightly even if they move a lot because the arc is larger.

[0230] In summary, objects moving within the same azimuth angle will have the same rate of change of arc, and the closer an object is to the center (i.e., the smaller the radius), the smaller the movement will appear to be in the azimuth angle number, and the farther an object is from the center (i.e., the larger the radius), the larger the movement will appear to be in the azimuth angle number.

[0231] According to the embodiment, this characteristic also varies depending on the accuracy of the lidar. The lower the accuracy (= the larger the φ angle per rotation), the more pronounced this characteristic becomes. That is, a larger rotation angle means a larger azimuth angle, and the larger the azimuth angle, the better the movement of objects in the nearby area can be captured.

[0232] For this reason, small movements of objects close to the automobile (i.e., rider equipment) may appear large and may become local motion vectors, while the same movements may not appear when moving far away from the automobile, and may be covered by global motion vectors without local motion vectors. Here, a global motion vector refers to the overall motion change vector obtained by comparing consecutive frames, for example, a reference frame (or previous frame) with a current frame, and a local motion vector refers to the motion change vector in a specific area.

[0233] Therefore, in order to apply inter-prediction-based compression technology via reference frames to point cloud data captured by lidar and having multiple frames, a method is required to divide the point cloud data into prediction units, LPUs (largest prediction units) and / or PUs (prediction units), which reflect the characteristics of the content.

[0234] This specification supports a method for dividing point cloud data into prediction units (LPUs and / or PUs) that reflect the characteristics of the content in order to perform inter-prediction via reference frames on point cloud data captured by a lidar and having multiple frames. In this way, this specification can expand the range that can be predicted using local motion vectors, eliminate the need for additional calculations, and reduce the encoding time of the point cloud data. For convenience of explanation, this specification may also refer to an LPU as a first prediction unit and a PU as a second prediction unit.

[0235] Furthermore, this specification predicts whether it is beneficial to apply a motion vector within a divided prediction unit by RDO (Rate-Distortion Optimization), and signals the prediction result. That is, whether to apply a motion vector is signaled for each divided prediction unit. Here, in one embodiment, the motion vector is a global motion vector. Alternatively, the motion vector may be a local motion vector. Alternatively, the motion vector may be both a global motion vector and a local motion vector.

[0236] For inter prediction according to the embodiment, the following definitions of terms are provided.

[0237] 1) I (Intra) frame, P (Predicted) frame, B (Bidirectional) frame

[0238] Frames to be coded / decoded are classified into I (Intra) frames, P (Predicted) frames, and B (Bidirectional) frames, and a frame is also called a picture.

[0239] For example, they are transmitted in the order of I frame → P frame → (B frame) → (I frame | P frame) → .... The B frame may be omitted.

[0240] 2) Frame of Reference

[0241] A reference frame is a frame involved in encoding / decoding the current frame.

[0242] The I frame or P frame immediately preceding the current P frame used as a reference for encoding / decoding can be called a reference frame. The I frame or P frame immediately preceding the current B frame and the I frame or P frame immediately following the current B frame used as a reference for encoding / decoding can be called reference frames.

[0243] 3) Frame and intra-predictive coding / inter-predictive coding

[0244] Intra-prediction coding is performed on I frames, and inter-prediction coding is performed on P and B frames.

[0245] Also, if a P frame has a change rate greater than a predetermined threshold compared to a previous reference frame, the P frame undergoes intra-prediction coding like an I frame.

[0246] 4) Criteria for determining I (intra) frames

[0247] Of the multi-frames, every kth frame may be determined as an I-frame, or the correlation between frames may be scored, and frames with high scores may be set as I-frames.

[0248] 5) I-frame encoding / decoding

[0249] When encoding / decoding point cloud data having multiple frames, the geometry of the I frame is encoded / decoded based on an octree or a predictive tree, and the feature information of the I frame is encoded / decoded based on the restored geometry information using a predictive / lifting transform scheme or a RAHT scheme.

[0250] 6) P-frame encoding / decoding

[0251] When encoding / decoding point cloud data having multiple frames according to an embodiment, encoding / decoding of P frames is performed based on a reference frame.

[0252] In this case, the coding unit for inter prediction of a P frame is a frame unit, a tile unit, a slice unit, or an LPU or PU. To this end, this specification divides (or separates) point cloud data, frames, tiles, or slices into LPUs and / or PUs. For example, this specification divides points divided into slices again into LPUs and / or PUs.

[0253] The point cloud content, frames, tiles, slices, etc. to be divided are also called point cloud data. In other words, the points belonging to the point cloud content, frames, tiles, and slices to be divided are also called point cloud data.

[0254] In this specification, one embodiment is partitioning or segmenting point cloud data based on elevation. In this specification, one embodiment is partitioning point cloud data into LPUs and / or PUs based on elevation. In this specification, elevation is also referred to as vertical. That is, in this specification, elevation and vertical are used interchangeably. In other words, in this specification, point cloud data is partitioned into LPUs and / or PUs based on vertical.

[0255] In this specification, one embodiment is to divide the point cloud data based on a radius. In this specification, one embodiment is to divide the point cloud data into LPUs and / or PUs based on a radius.

[0256] In this specification, one embodiment is to divide the point cloud data based on the azimuth. In this specification, one embodiment is to divide the point cloud data into LPUs and / or PUs based on the azimuth.

[0257] In one embodiment, this specification divides the point cloud data on an altitude (or vertical) basis, a radius basis, an azimuth angle basis, or a combination of two or more thereof. In one embodiment, this specification divides the point cloud data into LPUs and / or PUs on an altitude (or vertical) basis, a radius basis, an azimuth angle basis, or a combination of two or more thereof.

[0258] In this specification, one embodiment is to divide the point cloud data into LPUs on one or a combination of two or more of an elevation (or vertical) basis, a radius basis, and an azimuth basis.

[0259] In this specification, one embodiment is to divide point cloud data into PUs on an altitude (or vertical) basis, a radius basis, an azimuth angle basis, or a combination of two or more of these.

[0260] This specification takes as an example a method of dividing point cloud data into LPUs based on one or a combination of two or more of an altitude (or vertical) basis, a radius basis, and an azimuth angle basis, and then further dividing the data into one or more PUs based on one or a combination of two or more of an altitude (or vertical) basis, a radius basis, and an azimuth angle basis.

[0261] This specification takes as an example the division of a PU into smaller PUs.

[0262] In one embodiment, this specification determines whether to apply a motion vector for each region divided based on one or a combination of two or more of altitude (or vertical), radius, and azimuth. In one embodiment, this specification performs a rate distortion optimization (RDO) check for each region divided based on one or a combination of two or more of altitude (or vertical), radius, and azimuth, and determines whether to apply a motion vector for each region. In one embodiment, this specification signals whether to apply a motion vector for each region. Here, the divided region or divided block may be an LPU or a PU. Furthermore, the motion vector may be a global motion vector or a local motion vector. In one embodiment, this specification defines a global motion vector.

[0263] This specification takes as an example signaling the method used for LPU partitioning and / or PU partitioning.

[0264] In one embodiment, this specification describes determining whether to apply a motion vector to each divided region on an altitude (or vertical) basis. In another embodiment, this specification describes dividing point cloud data on an altitude (or vertical) basis, performing an RDO check for each divided region, and determining whether to apply a global motion vector to each region. In another embodiment, this specification describes signaling whether to apply a global motion vector to each region. Here, the divided region or divided block may be an LPU or a PU.

[0265] According to the embodiment, LPU / PU division and inter-prediction-based encoding (i.e., compression) is performed in the geometry encoder on the transmitting side, and LPU / PU division and inter-prediction-based decoding (i.e., restoration) is performed in the geometry decoder on the receiving side.

[0266] According to an embodiment, the transmitting geometry encoder signals whether or not to apply a motion vector to each divided LPU / PU, and the receiving geometry decoder performs motion compensation for the corresponding LPU / PU based on the signaling information including whether or not to apply a motion vector.

[0267] Below, we explain the LPU division method for point cloud data captured by LIDAR.

[0268] According to an embodiment, an LPU (Largest Prediction Unit) is the largest unit into which point cloud content (or a frame) is divided for inter-frame prediction (i.e., inter-prediction).

[0269] According to an embodiment, the multi-frames captured by the lidar have the following characteristics in terms of frame-to-frame variation:

[0270] That is, the closer to the center, the higher the probability of a local motion vector occurring. Also, the probability of a new point being generated is high in the farthest area among areas within a predetermined angle based on the global motion vector.

[0271] 17 is a diagram illustrating an example of radius-based LPU division and motion probability according to an embodiment. That is, FIG. 17 shows an example of dividing point cloud data captured by a lidar into five regions (also called blocks) based on radius.

[0272] As shown in Figure 17, when point cloud data is divided based on the radius, there are regions where local motion vectors are likely to occur based on the global motion vector, i.e., a region (50010) where a moving object exists and a region (50030) where a new object appears. Therefore, region (50030) is likely to contain additional points, and region (50010) is the region to which local motion vectors are applied. In other regions, the positions of points similar to those in the current frame can be obtained by prediction using only the global motion vector.

[0273] According to the embodiment, the LPU division criterion is specified based on the radius as shown in FIG. 17 or FIG.

[0274] 18 illustrates a specific example of radius-based LPU division of point cloud data according to an embodiment, where the radius size used as a reference for LPU division is r.

[0275] Figure 18 is an example to help those skilled in the art understand, and depending on the characteristics of the point cloud data (or point cloud content or frame), LPU division of the point cloud data may be performed on an azimuth basis or an elevation (or vertical) basis.

[0276] This specification divides point cloud data into one or more LPUs based on one or more of the following: radius, azimuth, and altitude. This expands the area that can be predicted using only the global motion vector, eliminating the need for additional calculations. This reduces the time required to encode point cloud data, i.e., speeds up the encoding process.

[0277] Below, we will explain how to divide point cloud data captured by a lidar or point cloud data divided into LPUs into PUs.

[0278] According to an embodiment, for inter-frame prediction (i.e., inter-prediction), the point cloud data (or point cloud content or region or block) divided into LPUs (Largest Prediction Units) is further divided into one or more PUs.

[0279] According to an embodiment, if the corresponding area is divided into smaller PUs according to the probability of the area in which the local motion vector occurs, the process of fine division and motion vector search due to the fine division can be reduced, further calculations are not required, and the encoding time can be reduced.

[0280] This specification applies the following characteristics of point cloud data (or point cloud content) to the PU division method:

[0281] 1) The higher the elevation, the lower the probability of local motion vectors occurring. This is because the higher the elevation, the higher the probability of static sky and buildings. In other words, the higher the probability of no local motion.

[0282] 2) When the altitude is very low, the probability of local motion vectors occurring is low, because when the altitude is very low, the probability of it being a road is high.

[0283] 3) There is a probability that an object exists within a specific azimuth angle in the divided LPU or PU. In this case, the azimuth angle for PU division (e.g., the azimuth angle size used as the reference for PU division) is set by experiment. Also, there may be an azimuth angle that includes a moving person with a difference of one frame, and there is a certain azimuth angle that includes a moving vehicle. According to the embodiment, by finding a typical azimuth angle through experiment, there is a high probability of isolating an area to which a local motion vector should be applied.

[0284] 4) There is a probability that an object exists within a specific radius in the divided LPU or PU. In this case, the radius for PU division (e.g., the radius size used as a reference for PU division) is set by experiment. Also, there may be a radius that includes a moving person with a difference of one frame, and there is a certain radius that includes a moving car. According to the embodiment, by finding a typical radius through experiment, there is a high probability of isolating an area to which a local motion vector should be applied.

[0285] Therefore, in this embodiment, after dividing point cloud data into LPUs, when the LPUs are further divided into one or more PUs, the blocks (or regions) divided into the LPUs are first further divided based on the motion block elevation (motion_block_pu_elevation) e. If a local motion vector cannot be matched to the further divided blocks (or regions), further division is performed. In this case, the corresponding blocks are further divided based on (or applied to) the motion block azimuth (motion_block_pu_azimuth) φ. However, if a local motion vector cannot be matched to the further divided blocks (or regions) based on the motion block azimuth φ, further division is performed based on the motion block radius (motion_block_pu_radius) r. Alternatively, the further division may be performed to half the size of the PU block (or region).

[0286] 19 is a diagram showing an example of PU division according to an embodiment. In this case, the PU division may be performed using one or a combination of two or more of the motion block elevation (motion_block_pu_elevation) e, the motion block azimuth (motion_block_pu_azimuth) φ, and the motion block radius (motion_block_pu_radius) r. Here, the motion block elevation (motion_block_pu_elevation) e indicates the size of the elevation (or vertical) that serves as the reference for PU division, the motion block azimuth (motion_block_pu_azimuth) φ indicates the size of the azimuth that serves as the reference for PU division, and the motion block radius (motion_block_pu_radius) r indicates the size of the radius that serves as the reference for PU division. In this case, the PU division is applied to frames, tiles, or slices.

[0287] Depending on the embodiment, when PU division is performed by combining two or more of the motion block altitude (motion_block_pu_elevation) e, the motion block azimuth angle (motion_block_pu_azimuth) φ, and the motion block radius (motion_block_pu_radius) r, the division can be performed in various orders. For example, PU division is performed in the order of altitude -> azimuth -> radius, altitude -> radius -> azimuth, azimuth -> altitude -> radius, azimuth -> radius -> altitude, radius -> altitude -> azimuth, radius -> altitude, altitude -> azimuth, altitude -> radius, azimuth -> altitude, azimuth -> radius, radius -> altitude, radius -> azimuth.

[0288] This allows this embodiment to expand the area that can be predicted using local motion vectors, eliminating the need for additional calculations and reducing coding time.

[0289] Hereinafter, a method for supporting LPU / PU division based on content characteristics based on an octree will be described.

[0290] In this specification, when performing octree-based geometry coding, in order to match LPU and PU division to octree occupied bits, an appropriate size is set by the following process.

[0291] That is, the size of the octree node that can be covered by the center-based motion block radius (motion_block_pu_radius) r is set as the motion block size (motion_block_size). Also, based on the set size, it is not necessary to divide into LPUs up to a certain octree level.

[0292] After dividing into LPUs, the axis order for dividing into PUs is determined. For example, the axis order is specified and applied in the order of xyz, xzy, yzx, yxz, zxy, or zyx.

[0293] This embodiment supports a method for applying both an octree structure and an LPU / PU division method that suits the characteristics of the content. The basic goal of LPU / PU division is to expand the area that can be predicted by local motion vectors as much as possible, eliminating the need for additional calculations and reducing the coding time.

[0294] FIG. 20 is a diagram illustrating another example of a point cloud transmitting device according to an embodiment.

[0295] The point cloud transmission device according to the embodiment includes a data input unit 51001, a coordinate system conversion unit 51002, a quantization processing unit 51003, a spatial division unit 51004, a signaling processing unit 51005, a geometry encoder 51006, a feature encoder 51007, and a transmission processing unit 51008. According to the embodiment, the coordinate system conversion unit 51002, the quantization processing unit 51003, the spatial division unit 51004, the geometry encoder 51006, and the feature encoder 51007 are referred to as a point cloud video encoder.

[0296] The point cloud transmitting device in Figure 20 corresponds to transmitting device 10000, point cloud video encoder 10002, transmitter 10003, acquire-encode-transmit 20000-20001-20002 in Figure 2, point cloud video encoder in Figure 4, transmitting device in Figure 12, device in Figure 14, etc. Each component in Figure 20 and corresponding figures corresponds to software, hardware, a processor in connection with memory, and / or combinations thereof.

[0297] The data input unit 51001 may perform some or all of the operations of the point cloud video acquisition unit 10001 in FIG. 1, or may perform some or all of the operations of the data input unit 12000 in FIG. 12. Furthermore, the coordinate system conversion unit 51002 may perform some or all of the operations of the coordinate system conversion unit 40000 in FIG. 4. Furthermore, the quantization processing unit 51003 may perform some or all of the operations of the quantization unit 40001 in FIG. 4, or may perform some or all of the operations of the quantization processing unit 12001 in FIG. 12. That is, the data input unit 51001 receives data to encode point cloud data. The data may be geometry data (also referred to as geometry, geometry information, etc.), characteristic data (also referred to as characteristic, characteristic information, etc.), parameter information indicating coding settings, etc.

[0298] The coordinate system conversion unit 51002 supports coordinate system conversion of point cloud data, such as changing the xyz axes or converting from an xyz Cartesian coordinate system to a spherical coordinate system.

[0299] The quantization processor 51003 quantizes the point cloud data. For example, the scale is adjusted by multiplying the x, y, and z position values ​​of the point cloud data by a scale (scale = geometry quantization value) setting. The scale value is either set according to the setting value or is included in the bitstream as parameter information and transmitted to the receiving side.

[0300] The spatial division unit 51004 performs spatial division into one or more 3D blocks based on a bounding box and / or a sub-bounding box, etc., on the point cloud data quantized and output by the quantization processing unit 51003. For example, the spatial division unit 51004 divides the quantized point cloud data into tiles or slices for access or parallel processing by content region. In addition, in one embodiment, signaling information for spatial division is entropy coded in the signaling processing unit 51005 and then transmitted via the transmission processing unit 51008 in the form of a bitstream.

[0301] In one embodiment, the point cloud content may represent one or more people, such as actors, or one or more objects. It may also represent a larger area, such as a map for autonomous driving or a map for indoor robot navigation. In this case, the point cloud content may be a huge amount of geographically linked data. Therefore, since the point cloud content cannot be encoded / decoded all at once, it is partitioned into tiles before being compressed. For example, room 101 in a building is divided into one tile and room 102 is divided into another tile. The divided tiles are then partitioned into slices to enable parallelization and faster encoding / decoding. This process is called slice partitioning.

[0302] That is, a tile refers to a region (e.g., a rectangular cube) of a three-dimensional space occupied by point cloud data according to an embodiment. A tile according to an embodiment includes one or more slices. A tile according to an embodiment is partitioned into one or more slices, and a point cloud video encoder encodes the point cloud data in parallel.

[0303] A slice refers to a unit of data (or bitstream) that is independently encoded in a point cloud video encoder according to an embodiment and / or a unit of data (or bitstream) that is independently decoded in a point cloud video decoder according to an embodiment. A slice according to an embodiment may refer to a set of data in a 3D space occupied by point cloud data, or may refer to a set of partial data of the point cloud data. A slice refers to a region of points or a set of points included in a tile according to an embodiment. A tile according to an embodiment is divided into one or more slices based on the number of points included in a tile. For example, one tile refers to a set of points divided by the number of points. A tile according to an embodiment is divided into one or more slices based on the number of points, and some data is split or merged during the division process. That is, a slice is also a unit that is independently coded within the tile. A tile divided into space in this way is further divided into one or more slices for fast and efficient processing.

[0304] The point cloud video encoder according to the embodiment encodes point cloud data in units of slices or tiles including one or more slices, and may perform different quantization and / or transformation for each tile or slice.

[0305] The positions of one or more 3D blocks (e.g., slices) spatially divided by the spatial division unit 51004 are output to a geometry encoder 51006, and feature information (also called features) is output to a feature encoder 51007. The positions are position information of points included in the divided units (boxes, blocks, tiles, tile groups, or slices), and are called geometry information.

[0306] The geometry encoder 51006 performs inter-prediction or intra-prediction based encoding on the positions output from the spatial division unit 51004 and outputs a geometry bitstream. For inter-prediction based encoding of a P frame, the geometry encoder 51006 divides the frame, tile, or slice into LPUs and / or PUs by applying the LPU / PU division method described above, and for motion compensation, applies or does not apply a motion vector to each division region (i.e., LPU or PU). The geometry encoder 51006 also signals whether or not a motion vector is applied to each division region. Here, the motion vector may be a global motion vector or a local motion vector. The geometry encoder 51006 also reconstructs the encoded geometry information and outputs it to the feature encoder 51007.

[0307] The feature encoder 51007 encodes (i.e., compresses) the feature (e.g., the divided feature original data) output from the spatial division unit 51004 based on the reconstructed geometry output from the geometry encoder 51006, and outputs a feature bitstream.

[0308] FIG. 21 is a diagram illustrating an example of the operation of the geometry encoder 51006 and the feature encoder 51007 according to an embodiment.

[0309] As one example, a quantization processing unit may be further provided between the spatial division unit 51004 and the voxelization processing unit 53001. The quantization processing unit quantizes the positions of one or more 3D blocks (e.g., slices) spatially divided by the spatial division unit 51004. In this case, the quantization unit may perform some or all of the operations of the quantization unit 40001 in Fig. 4, or may perform some or all of the operations of the quantization processing unit 12001 in Fig. 12. When a quantization processing unit is further provided between the spatial division unit 51004 and the voxelization processing unit 53001, the quantization processing unit 51003 in Fig. 20 may or may not be omitted.

[0310] According to an embodiment, the voxelization processor 53001 performs voxelization based on the positions of one or more spatially divided 3D blocks (e.g., slices) or quantized positions. Voxelization refers to the smallest unit representing position information in 3D space. That is, the voxelization processor 53001 supports a process of rounding the geometric position values ​​of scaled points to integers. According to an embodiment, points of point cloud content (or 3D point cloud video) are included in one or more voxels. According to an embodiment, one voxel includes one or more points. In one embodiment, if quantization is performed before voxelization, multiple points may belong to one voxel.

[0311] In this specification, when two or more points are contained in one voxel, these two or more points are called overlapping points (or duplicated points), i.e., duplicated points are generated by geometry quantization and voxelization during the geometry encoding process.

[0312] The voxelization processing unit 53001 according to the embodiment may output overlapping points belonging to one voxel as they are without merging them, or may merge the overlapping points into one point and output the merged points.

[0313] According to an embodiment, the geometry information intra prediction unit 53003 applies geometry intra prediction coding to geometry information of an I frame when the frame of the input point cloud data (i.e., the frame to which the input points belong) is an I frame. Intra prediction coding methods include octree coding, predictive tree coding, trisoup coding, etc.

[0314] For this purpose, a determination unit 53002 (or a determination unit) determines whether the points output from the voxelization processing unit 53001 belong to the I frame or the P frame.

[0315] According to an embodiment, when the frame confirmed by the discrimination unit 53002 is a P frame, the LPU / PU division unit 53004 divides the points divided into tiles or slices by the spatial division unit 51004 again into LPUs / PUs to support inter-prediction. As another embodiment, when the frame confirmed by the discrimination unit 53002 is a P frame, the LPU / PU division unit 53004 divides the points included in the frame into LPUs / PUs to support inter-prediction.

[0316] The method for dividing points of point cloud data (e.g., slices) into LPUs and / or PUs has been described in detail in Figures 15 to 19, so a description thereof will be omitted here to avoid duplication. Signaling related to LPU / PU division will be described later.

[0317] In this specification, if a P frame has a change rate greater than a predetermined threshold compared to a previous reference frame, intra-prediction coding is performed on the P frame, just like an I frame. For example, if the change rate of the entire frame is large and falls outside a predetermined threshold range, intra-prediction coding is performed on the P frame rather than inter-prediction coding. This is because intra-prediction coding is more accurate and efficient than inter-prediction coding when the change rate is large. Here, the previous reference frame is provided from a reference frame buffer 53009.

[0318] For this purpose, a determination unit 53005 checks whether the rate of change is greater than a threshold value.

[0319] If the discrimination unit 53005 determines that the rate of change between the P frame and the reference frame is greater than a threshold, the P frame is output to the geometry information intra prediction unit 53003 for intra prediction. On the other hand, if the discrimination unit 53005 determines that the rate of change is not greater than the threshold, the P frame divided into LPUs and / or PUs is output to the motion compensation application unit 53006 for inter prediction.

[0320] In one embodiment, the motion compensation application unit 53006 according to the embodiment determines whether to apply a motion vector to each divided LPU / PU and signals the result. For example, the RDO of a specific PU is checked to determine whether to apply a motion vector to the PU. In one embodiment, if applying a motion vector to the PU is beneficial, the motion vector is applied to the PU. In one embodiment, if applying a motion vector to the PU is not beneficial, the motion vector is not applied to the PU. Here, the benefit is determined by comparing the bitstream size when the motion vector is applied, etc. In another embodiment, information identifying whether a motion vector has been applied to the PU (e.g., pu_motion_compensation_type) is included in option information related to inter prediction (or information related to inter prediction). In this case, the motion vector applied to the PU may be a global motion vector calculated by overall motion estimation between frames, a local motion vector calculated for the PU, or both a global motion vector and a local motion vector.

[0321] That is, this specification divides point cloud data into prediction units (PUs), calculates a local motion vector for each PU, and then applies the local motion vector regardless of whether the coding unit or PU matches in order to apply it to all of octree-based geometry coding, prediction tree-based geometry coding, and trisoup-based geometry coding.

[0322] In addition, after applying the global motion vector on the LPU, local motion vectors are obtained according to PU division, and whether it is beneficial to apply the local motion vector within the PU, to apply only the global motion vector, or to use the previous frame as is is predicted by RDO and applied to the PU. That is, depending on the optimized application method, the global motion vector or the local motion vector is applied to the PU, or the previous frame is used as is. Here, using the previous frame as is means not using a motion vector.

[0323] According to an embodiment, the optimized application method, if there is a local motion vector, signals the local motion vector and then transmits it to the receiving side for decoding.

[0324] Therefore, on the receiving side, the signaling information can be used to check whether a motion vector (e.g., a global motion vector) has been applied to the PU, and if a global motion vector has been applied, the global motion vector is applied to the PU to perform motion compensation.

[0325] According to the embodiment, the LPU / PU division unit 53004 divides the point cloud data into LPUs and / or PUs, determines whether to apply a global motion vector to the LPU / PU, and signals the decision in the signaling information. Also, the motion compensation application unit 53006 performs motion compensation on the LPU / PU according to the signaling information.

[0326] The geometry information inter-prediction unit 53007 according to the embodiment performs octree-based inter-coding, predictive-tree-based inter-coding, or trisoup inter-coding based on the difference in geometry prediction values ​​between the current frame and a reference frame that has undergone motion compensation or a previous frame that has not undergone motion compensation.

[0327] The geometry information intra prediction unit 53003 according to the embodiment applies geometry intra prediction coding to the geometry information of the P frame input by the determination unit 53005. Methods of intra prediction coding include octree coding, predictive tree coding, trisoup coding, etc.

[0328] The geometry information entropy coding unit 53008 according to the embodiment performs entropy coding on the geometry information coded on an intra-prediction basis in the geometry information intra-prediction unit 53003 or on the geometry information coded on an inter-prediction basis in the geometry information inter-prediction unit 53007, and outputs a geometry bitstream (or referred to as a geometry information bitstream).

[0329] The geometry restoration unit according to the embodiment restores (or reconstructs) geometry information based on the changed position through intra-prediction-based coding or inter-prediction-based coding, and outputs the restored geometry information (also referred to as restored geometry) to the feature encoder 51007. That is, since feature information is dependent on geometry information (position), restored (or reconstructed) geometry information is necessary to compress the feature information. In addition, the restored geometry information is stored in the reference frame buffer 53009 to be used as a reference frame during inter-predictive coding of a P frame. The reference frame buffer 53009 also stores the feature information restored by the feature encoder 51007. That is, the restored geometry information and the restored feature information stored in the reference frame buffer 53009 are used as previous reference frames for geometry information inter-predictive coding and feature information inter-predictive coding in the geometry information inter-prediction unit 53007 of the geometry encoder 51006 and the feature information inter-prediction unit 55005 of the feature encoder 51007.

[0330] The hue conversion processing unit 55001 of the feature encoder 51007 corresponds to the color conversion unit 40006 in Fig. 4 or the hue conversion processing unit 12008 in Fig. 12. The hue conversion processing unit 55001 according to the embodiment performs hue conversion coding to convert the hue values ​​(or textures) included in the feature provided from the data input unit 51001 and / or the space division unit 51004. For example, the hue conversion processing unit 55001 converts the format of the hue information (e.g., converts from RGB to YCbCr). The operation of the hue conversion processing unit 55001 according to the embodiment is optionally applied depending on the hue values ​​included in the feature. In another embodiment, the hue conversion processing unit 55001 performs hue conversion coding based on the reconstructed geometry.

[0331] According to an embodiment, the feature encoder 51007 performs color readjustment depending on whether lossy coding is applied to the geometry information. To this end, a determination unit 55002 (also referred to as a determination unit) determines whether lossy coding is applied to the geometry information in the geometry encoder 51006.

[0332] For example, if the determination unit 55002 determines that lossy coding has been applied to the geometry information, the hue readjustment unit 55003 performs hue readjustment (or recoloring) to reset a characteristic (color) according to the lost point. That is, the hue readjustment unit 55003 searches for and sets an appropriate characteristic value at the position of the lost point from the original point cloud data. In other words, if a scale is applied to the geometry information and the position information value is changed, the hue readjustment unit 55003 predicts an appropriate characteristic value at the changed position.

[0333] According to an embodiment, the operation of the hue readjustment unit 55003 is applied selectively (optionally) depending on whether or not duplicated points are merged. In one embodiment, whether or not duplicated points are merged is determined by the voxelization processing unit 53001 of the geometry encoder 51006.

[0334] In this specification, one embodiment is when points belonging to one voxel are merged into one point in the voxelization processing unit 53001, and the hue readjustment unit 55003 performs hue readjustment (i.e., recoloring).

[0335] The hue readjustment unit 55003 performs operations and / or methods that are the same as or similar to the operations and / or methods of the characteristic conversion unit 40007 in FIG. 4 or the characteristic conversion processing unit 12009 in FIG.

[0336] If the discrimination unit 55002 determines that lossy coding has not been applied to the geometry information, the drawing reference numeral 55004 (also referred to as the discrimination unit) determines whether inter-prediction based coding has been applied to the geometry information.

[0337] If the determination unit 55004 determines that inter-prediction-based coding is not applied to the geometry information, the feature information intra prediction unit 55006 performs intra prediction coding on the input feature information. According to an embodiment, the intra prediction coding method performed by the feature information intra prediction unit 55006 includes predictive transform coding, lift transform coding, RAHT coding, etc.

[0338] If the determination unit 55004 determines that inter-prediction-based coding is applied to the geometry information, the feature information inter prediction unit 55005 performs inter-prediction coding on the input feature information. According to an embodiment, the feature information inter prediction unit 55005 includes a method of coding a residual value based on a difference between feature prediction values ​​of a current frame and a reference frame that has undergone motion compensation.

[0339] The feature information entropy coding unit 55008 according to the embodiment performs entropy coding on feature information coded based on intra prediction in the feature information intra prediction unit 55006, or feature information coded based on inter prediction in the feature information inter prediction unit 55005, and outputs a feature bit stream (or referred to as a feature information bit stream).

[0340] The feature restoration unit according to the embodiment restores (or reconstructs) feature information based on features changed by intra-predictive coding or inter-predictive coding, and stores the restored feature information (or referred to as restored features) in the reference frame buffer 53009. That is, the restored geometry information and restored feature information stored in the reference frame buffer 53009 are used as previous reference frames for geometry information inter-predictive coding and feature information inter-predictive coding in the geometry information inter-prediction unit 53007 and the feature information inter-prediction unit 55005 of the feature encoder 51007.

[0341] The LPU / PU division unit 53004 will be described below in relation to signaling.

[0342] That is, the LPU / PU splitter 53004 applies criterion type information (motion_block_lpu_split_type) for splitting point cloud data (e.g., points input in units of frames, tiles, or slices) into LPUs to the point cloud data, splits the point cloud data into LPUs, and then signals the applied type information. According to an embodiment, the criterion type information (motion_block_lpu_split_type) for splitting into LPUs may be radius-based, azimuth-based, altitude-based (or vertical-based), etc. In this specification, it is assumed that the criterion type information (motion_block_lpu_split_type) for splitting into LPUs is included in option information related to inter prediction (or referred to as information related to inter prediction) as an example.

[0343] When splitting point cloud data according to the criterion type information (motion_block_lpu_split_type) for splitting into LPUs, the LPU / PU splitter 53004 applies reference information (e.g., motion_block_lpu_radius, motion_block_lpu_azimuth, or motion_block_lpu_elevation) to the point cloud data, splits the data into LPUs, and then signals the applied value. According to an embodiment, the reference information for splitting into LPUs includes a radius size, an azimuth size, and an altitude (or vertical) size (e.g., motion_block_lpu_radius, motion_block_lpu_azimuth, or motion_block_lpu_elevation). In this specification, as an embodiment, the reference information for splitting into LPUs (e.g., motion_block_lpu_radius, motion_block_lpu_azimuth, or motion_block_lpu_elevation) is included in optional information related to inter prediction.

[0344] If a local motion vector corresponding to the LPU exists, the LPU / PU division unit 53004 signals the motion vector. Also, if the RDO of the predicted value is better after applying the global motion vector, the local motion vector does not need to be applied to the LPU.

[0345] According to an embodiment, information indicating whether a motion vector is present (referred to as motion_vector_flag or pu_has_motion_vector_flag, or information indicating whether an applicable motion vector is present) is signaled. In this specification, an embodiment is that the information indicating whether a motion vector is present (motion_vector_flag or pu_has_motion_vector_flag) is included in optional information related to inter prediction.

[0346] If a local motion vector corresponding to an LPU exists and there are various variations, the LPU / PU division unit 53004 further divides the LPU into one or more PUs and searches for a local motion vector for each PU. The LPU / PU division unit 53004 applies a global motion vector to each PU, calculates a gain, and then determines whether to apply the global motion vector to each PU. In one embodiment, information (pu_motion_compensation_type) identifying whether a motion vector (e.g., a global motion vector) is applied to the PU is included in option information related to inter prediction.

[0347] The LPU / PU splitter 53004 applies splitting criterion order type information (motion_block_pu_split_type) for splitting an LPU into one or more PUs to the LPU, splits the LPU into one or more PUs, and then signals the applied splitting criterion order type information (motion_block_pu_split_type). Splitting criterion order types include radius-based → azimuth-based → altitude (or vertical)-based splitting, radius-based → altitude (or vertical) → azimuth-based splitting, azimuth-based → radius-based → altitude (or vertical)-based splitting, azimuth-based → altitude (or vertical) → radius-based splitting, altitude (or vertical) → radius-based → azimuth-based splitting, and altitude (or vertical) → azimuth-based → radius-based splitting. In this specification, it is assumed that information on the splitting reference order type (motion_block_pu_split_type) for splitting into one or more PUs is included in option information related to inter prediction. In this specification, "altitude" is used to have the same meaning as "vertical" and can be used interchangeably.

[0348] When performing geometry coding based on an octree, the LPU / PU splitter 53004 applies octree-related base order type information (Motion_block_pu_split_octree_type) for splitting into PUs to the octree, splits the PUs, and then signals the applied type information. Split base order types include x→y→z-based split application, x→z→y-based split application, y→x→z-based split application, y→z→x-based split application, z→x→y-based split application, and z→y→x-based split application. In this specification, as an example, the base order type information (Motion_block_pu_split_octree_type) related to the octree for splitting into PUs is included in option information related to inter prediction.

[0349] The LPU / PU splitter 53004 applies information (e.g., motion_block_pu_radius, motion_block_pu_azimuth, motion_block_pu_elevation) that serves as a reference when splitting the point cloud data or the LPU into one or more PUs according to the criterion type information (motion_block_pu_split_type) for splitting into PUs, splits the data into one or more PUs, and then signals the applied value. The information that serves as a reference when splitting the data includes a radius size, an azimuth size, an altitude (or vertical) size, etc. Alternatively, the data may be split by reducing the current size by half for each step of splitting into PUs. In this specification, as an example, the information that serves as a reference when splitting the data into PUs (e.g., motion_block_pu_radius, motion_block_pu_azimuth, motion_block_pu_elevation) is included in optional information related to inter prediction.

[0350] If a local motion vector corresponding to the PU exists and there are various variations, the LPU / PU division unit 53004 divides the PU into one or more smaller PUs and searches for the local motion vector. At this time, information indicating whether the PU is further divided into one or more smaller PUs is signaled. In this specification, as an example, the information indicating whether the PU is further divided into one or more smaller PUs is included in option information related to inter prediction.

[0351] If a local motion vector corresponding to the PU exists, the LPU / PU division unit 53004 signals the motion vector (pu_motion_vector_xyz). It also signals information indicating whether or not a motion vector exists (pu_has_motion_vector_flag). In this specification, one embodiment is that the motion vector and / or the information indicating whether or not a motion vector exists (pu_has_motion_vector_flag) are included in optional information related to inter prediction.

[0352] The LPU / PU division unit 53004 signals whether a block (or region) corresponding to an LPU / PU is divided. In this specification, as an example, information indicating whether a block (or region) corresponding to an LPU / PU is divided is included in option information related to inter prediction.

[0353] The LPU / PU division unit 53004 receives minimum PU size information (motion_block_pu_min_radius, motion_block_pu_min_azimuth, motion_block_pu_min_elevation), performs division / local motion vector search up to that size, and signals the value. Here, in one embodiment, the value is included in option information related to inter prediction.

[0354] In this way, if the frame is a P frame, the LPU / PU division unit 53004 divides the points divided into slices into division regions such as LPUs / PUs to support inter-prediction, and searches for and assigns motion vectors corresponding to each division region. LPUs are divided based on radius, in which case motion_block_lpu_radius is signaled in optional information related to inter-prediction and transmitted to the receiving decoder. Alternatively, LPUs may be divided based on other criteria, in which case motion_block_lpu_split_type is applied, and motion_block_lpu_split_type is included in optional information related to inter-prediction and transmitted to the receiving decoder. PUs are preferentially divided based on altitude (also called vertical), and additional division may be performed based on radius and azimuth, and the division level may be changed depending on the settings. Alternatively, LPUs may be divided based only on altitude (also called vertical). Alternatively, the division order may be changed. In this case, it is applied by motion_block_pu_split_type, which is included in optional information related to inter prediction and transmitted to a decoder on the receiving side. For example, splitting may be performed in the order of azimuth angle -> altitude (or vertical) -> radius, and the splitting method and splitting reference values, motion_block_pu_elevation, motion_block_pu_azimuth, and motion_block_pu_radius, are signaled in optional information related to inter prediction.

[0355] In addition, if the frame is a P frame, the LPU / PU division unit 53004 divides the points divided into slices into division regions such as LPUs / PUs to support inter-prediction, and searches for and assigns motion vectors corresponding to each division region. Here, the LPU / PU division unit 53004 predicts whether it is beneficial to apply a local motion vector to the PU, whether it is beneficial to apply only a global motion vector, or whether it is beneficial to use the previous frame as is, using RDO, and sets the prediction result to the PU. For example, if it is most beneficial to apply a global motion vector to the PU, the LPU / PU division unit 53004 applies the global motion vector to the PU and signals information identifying the appropriateness (pu_motion_compensation_type) in option information related to inter-prediction, which is then transmitted to the decoder on the receiving side. That is, the LPU / PU division unit 53004 applies a motion vector to the PU according to the optimized application method. If the optimized application method and local motion vector exist, the local motion vector is signaled to the decoder.

[0356] In addition, the motion compensation application unit 53006 determines, based on option information related to inter prediction, whether to select a value with a global motion vector applied to the PU, a value with a local motion vector applied, or use the point of the previous frame as is, and performs motion compensation based on that decision.

[0357] In this specification, optional information related to inter prediction is signaled in a GPS, a TPS, a geometry slice header, etc. In this case, it is assumed that the optional information related to inter prediction is processed by the signaling processing unit 61002, as an example.

[0358] As described above, the optional information related to inter prediction includes information on the reference type for splitting into LPUs (motion_block_lpu_split_type), information used as a reference when splitting into LPUs (e.g., motion_block_lpu_radius, motion_block_lpu_azimuth, or motion_block_lpu_elevation), information indicating whether a motion vector is present (motion_vector_flag or pu_has_motion_vector_flag), information on the split reference order type for splitting into PUs (motion_block_pu_split_type), information on the reference order type related to the octree for splitting into PUs (motion_block_pu_split_octr The optional information related to inter prediction includes at least one of information for identifying a tile to which a PU belongs, information for identifying a slice to which a PU belongs, information about the number of PUs included in a slice, information for identifying each PU, etc. In this specification, the information included in the optional information related to inter prediction may be added, deleted, or modified according to the skill of those skilled in the art, and therefore the present invention is not limited to the above examples.

[0359] FIG. 22 is a block diagram illustrating an example of a method for LPU / PU division-based geometry encoding according to an embodiment.

[0360] 22, steps 57001 to 57003 indicate detailed operations of the LPU / PU division unit 53004, steps 57004 and 57005 indicate detailed operations of the motion compensation application unit 53007, and step 57006 indicates detailed operations of the geometry information inter prediction unit 53007.

[0361] That is, in step 57001, a global motion vector is found, and in step 57002, in order to apply the global motion vector found in step 57001, the point cloud data is divided into LPUs based on one or a combination of two or more of the radius, azimuth, and altitude. In step 57003, if a local motion vector corresponding to the LPU exists and there are various variations, the LPU is further divided into one or more PUs, and a local motion vector is found in each divided PU. In steps 57001 to 57003, RDO (Rate Distortion Optimization) is applied to select the best (i.e., optimal motion vector).

[0362] In addition, in steps 57001 to 57003, the RDO checks whether it is beneficial to apply a global motion vector within the LPU or PU, or whether it is beneficial not to apply it, and determines whether to apply a global motion vector to the LPU or PU, and signals the result (e.g., pu_motion_compensation_type) to the option information related to inter-prediction in the signaling information.

[0363] In step 57004, global motion compensation is performed by applying a global motion vector to the LPU or PU according to pu_motion_compensation_type. Alternatively, global motion compensation may be omitted for the LPU or PU according to pu_motion_compensation_type. In step 57005, local motion compensation is performed by applying a local motion vector to the divided PU. Alternatively, local motion compensation may be omitted for the PU. In step 57006, octree-based inter-coding, prediction tree-based inter-coding, or trisoup-based inter-coding is performed based on the difference (also called residual value) between the prediction value of the current frame and a motion-compensated reference frame (or a non-motion-compensated reference frame).

[0364] Meanwhile, the geometry bitstream compressed and output on an intra-prediction or inter-prediction basis by the geometry encoder 51006, and the feature bitstream compressed and output on an intra-prediction or inter-prediction basis by the feature encoder 51007 are output to the transmission processing unit 51008.

[0365] The transmission processing unit 51008 according to the embodiment may perform an operation and / or a transmission method that is the same as or similar to the operation and / or transmission method of the transmission processing unit 12012 in Fig. 12, or may perform an operation and / or a transmission method that is the same as or similar to the operation and / or transmission method of the transmitter 10003 in Fig. 1. For a detailed description, please refer to the description of Fig. 1 or Fig. 12 and will not be described here.

[0366] In this embodiment, the transmission processing unit 51008 may transmit the geometry bit stream output from the geometry encoder 51006, the characteristic bit stream output from the characteristic encoder 51007, and the signaling bit stream output from the signaling processing unit 51005 individually, or may multiplex them into a single bit stream and transmit them.

[0367] The transmission processing unit 51008 according to the embodiment may encapsulate the bitstream into a file or a segment (for example, a streaming segment), and then transmit it via various networks such as a broadcast network and / or a broadband network.

[0368] The signaling processing unit 51005 according to the embodiment generates and / or processes signaling information and outputs it in the form of a bitstream to the transmission processing unit 51008. The signaling information generated and / or processed in the signaling processing unit 51005 may be provided to the geometry encoder 51006, the attribute encoder 51007, and / or the transmission processing unit 51008 for geometry encoding, attribute encoding, and transmission processing, or the signaling processing unit 51005 may be provided with signaling information generated from the geometry encoder 51006, the attribute encoder 51007, and / or the transmission processing unit 51008.

[0369] In this specification, the signaling information is signaled and transmitted in units of parameter sets (SPS: sequence parameter set, GPS: geometry parameter set, APS: attribute parameter set, TPS: tile parameter set, etc.). It may also be signaled and transmitted in units of coding units of each video, such as slices or tiles. In this specification, the signaling information includes metadata (e.g., setting values, etc.) related to point cloud data, and is provided to the geometry encoder 51006, the attribute encoder 51007, and / or the transmission processing unit 51008 for geometry encoding, attribute encoding, and transmission processing. Depending on the application, the signaling information may also be defined at the system end, such as a file format, dynamic adaptive streaming over HTTP (DASH), or MPEG media transport (MMT), or at the wired interface end, such as HDMI (High Definition Multimedia Interface), Display Port, VESA (Video Electronics Standards Association), or CTA.

[0370] The method / apparatus according to the embodiment signals relevant information to add / perform the operation of the embodiment. The signaling information according to the embodiment is used in a transmitting device and / or a receiving device.

[0371] In this specification, it is assumed that optional information related to inter prediction used for inter prediction of geometry information is signaled in at least one of a geometry parameter set, a tile parameter set, and a geometry slice header, or in a separate PU header (referred to as geom_pu_header) in one embodiment.

[0372] FIG. 23 is a diagram illustrating another example of a point cloud receiving device according to an embodiment.

[0373] The point cloud receiving device according to the embodiment includes a receiving processing unit 61001, a signaling processing unit 61002, a geometry decoder 61003, a feature decoder 61004, and a post-processor 61005. According to the embodiment, the geometry decoder 61003 and the feature decoder 61004 are also called a point cloud video decoder. According to the embodiment, the point cloud video decoder is also called a PCC decoder, PCC decoding unit, point cloud decoder, point cloud decoding unit, etc.

[0374] The point cloud receiving device in Figure 23 corresponds to receiving device 10004, receiver 10005, point cloud video decoder 10006, transmit-decode-render 20002-20003-20004 in Figure 2, point cloud video decoder in Figure 11, receiving device in Figure 13, device in Figure 14, etc. Each component in Figure 23 and corresponding figures corresponds to software, hardware, a processor in connection with memory, and / or a combination thereof.

[0375] The receiving processing unit 61001 according to the embodiment may receive one bitstream, or may receive a geometry bitstream (or a geometry information bitstream), a feature bitstream (or a feature information bitstream), and a signaling bitstream. When a file and / or a segment is received, the receiving processing unit 61001 according to the embodiment decapsulates the received file and / or segment and outputs the decapsulated file and / or segment to a bitstream.

[0376] In this embodiment, when one bitstream is received (or decapsulated), the receiving processing unit 61001 demultiplexes the geometry bitstream, attribute bitstream, and / or signaling bitstream from the one bitstream, and outputs the demultiplexed signaling bitstream to the signaling processing unit 61002, the geometry bitstream to the geometry decoder 61003, and the attribute bitstream to the attribute decoder 61004.

[0377] In this embodiment, when a geometry bitstream, a feature bitstream, and / or a signaling bitstream is received (or decapsulated), the receiving processing unit 61001 transmits the signaling bitstream to the signaling processing unit 61002, the geometry bitstream to the geometry decoder 61003, and the feature bitstream to the feature decoder 61004.

[0378] The signaling processing unit 61002 parses and processes signaling information, such as information contained in the SPS, GPS, APS, TPS, metadata, etc., from the input signaling bitstream, and provides the parsed information to the geometry decoder 61003, the attribute decoder 61004, and the post-processing unit 61005. In another embodiment, the signaling information contained in the geometry slice header and / or the attribute slice header is also pre-parsed by the signaling processing unit 61002 before decoding the corresponding slice data. That is, when the point cloud data is divided into tiles and / or slices at the transmitting side, the TPS includes the number of slices included in each tile, so that the point cloud video decoder according to the embodiment can check the number of slices and quickly parse the information for parallel decoding.

[0379] Therefore, a point cloud video decoder according to this specification receives an SPS with a reduced data amount and quickly parses a bitstream containing point cloud data. The receiving device decodes a tile as soon as it receives the tile, and maximizes decoding efficiency by performing decoding on a slice-by-slice basis based on the GPS and APS included in each tile. Alternatively, the receiving device maximizes decoding efficiency by performing inter-prediction decoding on point cloud data for each PU based on optional information related to inter-prediction signaled in the GPS, TPS, geometry slice header, and / or PU header.

[0380] That is, the geometry decoder 61003 performs the reverse process of the geometry encoder 51006 in Fig. 20 on the compressed geometry bitstream based on signaling information (e.g., geometry-related parameters) to restore the geometry. The geometry restored (or reconstructed) by the geometry decoder 61003 is provided to the feature decoder 61004. Here, the geometry-related parameters include optional information related to inter-prediction used for inter-prediction restoration of the geometry information.

[0381] The feature decoder 61004 performs the reverse process of the feature encoder 51007 in Fig. 20 on the compressed feature bitstream based on the signaling information (e.g., feature-related parameters) and the reconstructed geometry to restore the feature. According to an embodiment, if the point cloud data is divided into tiles and / or slices at the transmitting side, the geometry decoder 61003 and the feature decoder 61004 perform geometry decoding and feature decoding on a tile and / or slice basis.

[0382] FIG. 24 is a diagram showing an example of the operation of the geometry decoder 61003 and the feature decoder 61004 according to the embodiment.

[0383] The geometry information entropy coding unit 63001, inverse quantization processing unit 63007, and coordinate system inverse transformation unit 63008 included in the geometry decoder 61003 in Fig. 24 may perform some or all of the operations of the arithmetic decoder 11000 and coordinate system inverse transformation unit 11004 in Fig. 11, or may perform some or all of the operations of the arithmetic decoder 13002 and inverse quantization processing unit 13005 in Fig. 13. The positions restored by the geometry decoder 61003 are output to a post-processing unit 61005.

[0384] According to an embodiment, if optional information related to inter-prediction for inter-prediction restoration of geometry information is signaled in at least one of the geometry parameter set (GPS), tile parameter set (TPS), geometry slice header, and geometry PU header, it is acquired by the signaling processing unit 61002 and provided to the geometry decoder 61003, or acquired directly by the geometry decoder 61003.

[0385] According to an embodiment, the optional information related to inter prediction includes information on a reference type for splitting into LPUs (motion_block_lpu_split_type), information used as a reference when splitting into LPUs (e.g., motion_block_lpu_radius, motion_block_lpu_azimuth, or motion_block_lpu_elevation), information indicating whether an applicable motion vector is available (motion_vector_flag or pu_has_motion_vector_flag), information on a split reference order type for splitting into PUs (motion_block_pu_split_type), information on a reference order type related to an octree for splitting into PUs (Motion_block_pu_split_octree), _type), information used as a reference when dividing into PUs (e.g., motion_block_pu_radius, motion_block_pu_azimuth, or motion_block_pu_elevation), local motion vector information corresponding to the PU, information identifying whether a motion vector (e.g., a global motion vector) is applied to the PU (pu_motion_compensation_type), information indicating whether a block (or region) corresponding to an LPU / PU is divided, and minimum PU size information (e.g., motion_block_pu_min_radius, motion_block_pu_min_azimuth, or motion_block_pu_min_elevation). In addition, the optional information related to inter prediction further includes information for identifying a tile to which the PU belongs, information for identifying a slice to which the PU belongs, information on the number of PUs included in the slice, information for identifying each PU, etc. In this specification, information included in the optional information related to inter prediction may be added, deleted, or modified by those skilled in the art, and the present invention is not limited to the above examples.

[0386] That is, the geometry information entropy decoding unit 63001 entropy decodes the input geometry bitstream.

[0387] According to an embodiment, when intra-prediction based coding is applied to geometry information at the transmitting side, the geometry decoder 61003 performs intra-prediction based reconstruction on the geometry information. Conversely, when inter-prediction based coding is applied to geometry information at the transmitting side, the geometry decoder 61003 performs inter-prediction based reconstruction on the geometry information.

[0388] For this purpose, a determination unit 63002 determines whether intra-prediction based coding or inter-prediction based coding is applied to geometry information.

[0389] If the determining unit 63002 determines that intra-prediction-based coding has been applied to the geometry information, the entropy-decoded geometry information is provided to a geometry information intra-prediction restoration unit 63003. Conversely, if the determining unit 63002 determines that inter-prediction-based coding has been applied to the geometry information, the entropy-decoded geometry information is output to an LPU / PU splitting unit 63004.

[0390] The geometry information intra-prediction restoration unit 63003 according to the embodiment decodes and restores geometry information based on an intra-prediction method. That is, the geometry information intra-prediction restoration unit 63003 restores geometry information predicted by geometry intra-prediction coding. Intra-prediction coding methods include octree coding, predictive tree coding, trisoup coding, etc.

[0391] According to an embodiment, when the frame of geometry information to be decoded is a P frame, the LPU / PU division unit 63004 divides the reference frame into LPUs / PUs using optional information related to inter-prediction signaled to support inter-prediction-based reconstruction and to indicate LPU / PU division.

[0392] According to the embodiment, the motion compensation application unit 63005 applies motion vectors (e.g., global motion vectors and / or local motion vectors) to the LPUs / PUs divided from the reference frame to generate predicted geometry information. Here, the motion vectors are received as part of signaling information.

[0393] The motion compensation application unit 63005 according to the embodiment performs motion compensation by applying a global motion vector to the PU in question, in accordance with pu_motion_compensation_type included in the option information related to inter prediction.

[0394] The motion compensation application unit 63005 according to the embodiment performs motion compensation by applying a local motion vector to the PU in question, in accordance with pu_motion_compensation_type included in the option information related to inter prediction.

[0395] The motion compensation application unit 63005 according to the embodiment may omit the motion compensation process for the PU in question, depending on pu_motion_compensation_type included in the option information related to inter prediction.

[0396] The geometry information inter-prediction restoration unit 63006 according to an embodiment decodes and restores geometry information based on an inter-prediction method. That is, geometry information subjected to geometry inter-prediction coding is restored based on geometry information of a reference frame that has undergone motion compensation (or a reference frame that has not undergone motion compensation). Inter-prediction coding methods according to an embodiment include octree-based inter-coding, predictive-tree-based inter-coding, trisoup-based inter-coding, etc.

[0397] The geometry information restored by the geometry information intra-prediction restoration unit 63003 or the geometry information restored by the geometry information inter-prediction restoration unit 63006 is input to the geometry information transformation and inverse quantization processing unit 63007 .

[0398] The geometry information inverse transform and inverse quantization unit 63007 according to this embodiment performs the inverse process of the transform performed by the geometry information transform and quantization processing unit 51003 of the transmitting device on the restored geometry information, and multiplies the result by a scale (=geometry quantization value) to generate restored geometry information that has been inversely quantized. That is, the geometry information transform and inverse quantization processing unit 63007 applies the scale (scale = geometry quantization value) included in the signaling information to the geometry position x, y, and z values ​​of the restored point to inverse quantize the geometry information.

[0399] The coordinate system inverse transformation unit 63008 performs the inverse process of the coordinate system transformation performed by the coordinate system transformation unit 51002 of the transmitting device on the dequantized geometry information. For example, the coordinate system inverse transformation unit 63008 restores the x, y, and z axes changed on the transmitting side, or inversely transforms the transformed coordinate system into an x, y, and z orthogonal coordinate system.

[0400] According to the embodiment, the geometry information dequantized by the geometry information conversion and dequantization processing unit 63007 undergoes a geometry restoration process, is stored in the reference frame buffer 63009, and is also output to the feature decoder 61004 for feature decoding.

[0401] According to the embodiment, the feature residual information entropy decoding unit 65001 of the feature decoder 61004 entropy decodes the input feature bitstream.

[0402] According to the embodiment, if intra-prediction based coding is applied to the feature information at the transmitting side, the feature decoder 61004 performs intra-prediction based reconstruction on the feature information. Conversely, if inter-prediction based coding is applied to the feature information at the transmitting side, the feature decoder 61004 performs inter-prediction based reconstruction on the feature information.

[0403] For this purpose, a determination unit 65002 (also referred to as a determination unit) determines whether intra-prediction based coding or inter-prediction based coding is applied to the characteristic information.

[0404] If the determining unit 65002 determines that intra-prediction based coding is applied to the feature information, the entropy decoded feature information is provided to the feature information intra-prediction restoration unit 65004. Conversely, if the determining unit 65002 determines that inter-prediction based coding is applied to the feature information, the entropy decoded feature information is provided to the feature information inter-prediction restoration unit 65003.

[0405] The feature information inter-prediction restoration unit 65003 according to the embodiment decodes and restores feature information based on the inter-prediction method, that is, restores feature information predicted by inter-prediction coding.

[0406] The feature information intra prediction restoration unit 65004 according to the embodiment decodes and restores feature information based on an intra prediction method, i.e., restores feature information predicted by intra prediction coding. Intra coding methods include predictive transform coding, lift transform coding, RAHT coding, etc.

[0407] According to the embodiment, the restored feature information is stored in the reference frame buffer 63009. The geometry information and feature information stored in the reference frame buffer 63009 are provided to the geometry information inter-prediction restoration unit 63003 and the feature information inter-prediction restoration unit 65003 as the previous reference frame.

[0408] According to an embodiment, the restored characteristic information is provided to a hue inverse conversion processor 65005 and restored to RGB hue. That is, the hue inverse conversion processor 65005 performs inverse conversion coding to inversely convert the color values ​​(or texture) included in the restored characteristic information, and outputs the result to the post-processing unit 61005. The hue inverse conversion processor 65005 performs operations and / or inverse conversion coding that are the same as or similar to the operations and / or inverse conversion coding of the color inverse conversion unit 11010 of Fig. 11 or the hue inverse conversion processor 13010 of Fig. 13.

[0409] The post-processing unit 61005 reconstructs point cloud data by matching the geometry information (i.e., position) restored and output by the geometry decoder 61003 with the feature information restored and output by the feature decoder 61004. In addition, if the reconstructed point cloud data is in units of tiles and / or slices, the post-processing unit 61005 performs the reverse process of spatial division on the transmitting side based on signaling information.

[0410] Hereinafter, the LPU / PU splitter 63004 of the geometry decoder 61003 will be described in relation to signaling. In this case, in one embodiment, the signaling processor 61002 restores option information related to inter prediction that is received and included in at least one of a GPS, a TPS, a geometry slice header, and / or a geometry PU header, and provides the restored option information to the LPU / PU splitter 63004.

[0411] The LPU / PU splitter 63004 applies criterion type information (motion_block_lpu_split_type) for splitting the reference frame into LPUs to the reference frame to split it into LPUs, and then restores the transmitted motion vectors. In this specification, one embodiment is that the criterion type information (motion_block_lpu_split_type) for splitting the reference frame into LPUs is received while being included in at least one of a GPS, a TPS, or a geometry slice header.

[0412] The LPU / PU splitter 63004 applies reference type information (motion_block_lpu_split_type) for splitting into LPUs, and applies information (e.g., motion_block_lpu_radius, motion_block_lpu_azimuth, or motion_block_lpu_elevation) that serves as a reference when splitting the reference frame to the reference frame to split into LPUs. According to an embodiment, the information that serves as a reference when splitting into LPUs includes a radius size, an azimuth size, and an altitude (or vertical) size (e.g., motion_block_lpu_radius, motion_block_lpu_azimuth, or motion_block_lpu_elevation). In this specification, in one embodiment, the information that serves as a reference when splitting into LPUs (e.g., motion_block_lpu_radius, motion_block_lpu_azimuth, or motion_block_lpu_elevation) is received while being included in at least one of a GPS, a TPS, or a geometry slice header.

[0413] In the LPU / PU division unit 63004, if the information indicating whether or not a motion vector corresponding to the LPU exists (motion_vector_flag or pu_has_motion_vector_flag) indicates that an applicable motion vector exists, the LPU / PU division unit 63004 restores the motion vector. In this specification, one embodiment is that the information indicating whether or not a motion vector corresponding to the LPU exists (motion_vector_flag or pu_has_motion_vector_flag) and the motion vector are received by being included in at least one of a GPS, a TPS, or a geometry slice header.

[0414] If the information indicating whether the LPU is divided into PUs indicates that the LPU is divided into PUs, the LPU / PU division unit 63004 additionally divides the LPU into one or more PUs.

[0415] The LPU / PU splitter 63004 splits the LPU into one or more PUs by applying reference order type information (motion_block_pu_split_type) for splitting into PUs to the LPU. Splitting reference order types include radius-based → azimuth-based → altitude (or vertical)-based splitting, radius-based → altitude (or vertical) → azimuth-based splitting, azimuth-based → radius-based → altitude (or vertical)-based splitting, azimuth-based → altitude (or vertical) → radius-based splitting, altitude (or vertical) → radius-based → azimuth-based splitting, and altitude (or vertical) → azimuth-based → radius-based splitting. In this specification, it is assumed that the reference order type information (motion_block_pu_split_type) for splitting into PUs is received while being included in at least one of a GPS, TPS, or geometry slice header.

[0416] When geometry coding is applied based on an octree, the LPU / PU splitter 63004 splits the octree structure into one or more PUs based on a base order type (motion_block_pu_split_octree_type) associated with the octree for splitting into PUs. Base order types associated with the octree for splitting into PUs include x->y->z-based split application, x->z->y-based split application, y->x->z-based split application, y->z->x-based split application, z->x->y-based split application, z->y->x-based split application, etc. In this specification, it is assumed in one embodiment that the base order type (motion_block_pu_split_octree_type) associated with the octree for splitting into PUs is received while being included in at least one of a GPS, TPS, or geometry slice header.

[0417] The LPU / PU splitter 63004 splits the LPU into one or more PUs by applying information (e.g., motion_block_pu_radius, motion_block_pu_azimuth, or motion_block_pu_elevation) that serves as a reference when splitting the LPU into PUs to the LPU according to the criterion type information (motion_block_pu_split_type) for splitting into PUs. The information that serves as a reference when splitting includes the radius size, azimuth size, and altitude (or vertical) size. In this specification, one embodiment is that the information that serves as a reference when splitting into PUs (e.g., motion_block_pu_radius, motion_block_pu_azimuth, or motion_block_pu_elevation) is received while being included in at least one of a GPS, a TPS, or a geometry slice header.

[0418] The LPU / PU division unit 63004 further divides the PUs by applying minimum PU size information (e.g., motion_block_pu_min_radius, motion_block_pu_min_azimuth, or motion_block_pu_min_elevation) to the PUs. In this specification, one embodiment is that the minimum PU size information (e.g., motion_block_pu_min_radius, motion_block_pu_min_azimuth, or motion_block_pu_min_elevation) is received while being included in at least one of a GPS, a TPS, or a geometry slice header.

[0419] In this specification, optional information related to inter prediction includes information on the reference type for splitting into LPUs (motion_block_lpu_split_type), information used as a reference when splitting into LPUs (e.g., motion_block_lpu_radius, motion_block_lpu_azimuth, or motion_block_lpu_elevation), information indicating whether an applicable motion vector is available (motion_vector_flag or pu_has_motion_vector_flag), information on the split reference order type for splitting into PUs (motion_block_pu_split_type), information on the reference order type related to the octree for splitting into PUs (motion_block_pu_split _octree_type), information used as a reference when dividing into PUs (e.g., motion_block_pu_radius, motion_block_pu_azimuth, or motion_block_pu_elevation), local motion vector information corresponding to the PU, information identifying whether a global motion vector is applied to the PU (pu_motion_compensation_type), information indicating whether a block (or region) corresponding to an LPU / PU is divided, and minimum PU size information (e.g., motion_block_pu_min_radius, motion_block_pu_min_azimuth, or motion_block_pu_min_elevation). In this specification, information included in optional information related to inter prediction may be added, deleted, or modified by those skilled in the art, and the present invention is not limited to the above examples.

[0420] The motion compensation application unit 63005 performs motion compensation according to pu_motion_compensation_type included in option information related to inter prediction. For example, the motion compensation application unit 63005 determines whether to select a value to which a global motion vector is applied to the corresponding PU, a value to which a local motion vector is applied, or to use a point from a previous frame as is, based on pu_motion_compensation_type, and performs motion compensation on the corresponding PU according to the determination result. That is, the motion compensation application unit 63005 applies motion vectors to the divided LPUs / PUs according to the optimized application method (pu_motion_compensation_type) to generate a predicted point cloud. This process may be performed before geometry coding, or may be performed together if the PU unit matches the execution unit of geometry coding.

[0421] FIG. 25 is a diagram illustrating an example of a geometry decoding method based on LPU / PU division according to an embodiment.

[0422] In Figure 25, step 67001 shows the detailed operation of the geometry information entropy encoding unit 63001, step 67003 shows the detailed operation of the LPU / PU splitting unit 63004, steps 67002 and 67004 show the detailed operation of the motion compensation application unit 67005, and step 67006 shows the detailed operation of the geometry information inter-prediction restoration unit 63006.

[0423] That is, entropy decoding is performed on the geometry bitstream in step 67001. An example of entropy decoding is arithmetic decoding.

[0424] Step 67002 applies a global motion vector to the entropy-decoded geometry information to perform global motion compensation. Step 67003 divides the entropy-decoded geometry information into LPUs / PUs. Step 67004 applies local motion vectors to the divided LPUs / PUs to perform local motion compensation. At this time, local motion compensation may be omitted. Alternatively, step 67004 applies a global motion vector to the LPUs / PUs to perform global motion compensation. At this time, global motion compensation may be omitted. At this time, whether global motion compensation is performed by applying a global motion vector to the LPUs and / or PUs is identified by pu_motion_compensation_type included in option information related to inter prediction. In this specification, it is assumed that the global motion vector and / or local motion vector are received by being included in at least one of the GPS, TPS, geometry slice header, and geometry PU header. Since LPU / PU division has been described above, a description thereof will be omitted here.

[0425] To perform global motion compensation, a previous reference frame (i.e., a reference point cloud) stored in a reference frame buffer is provided to step 67002.

[0426] For local motion compensation, either the world coordinates for which global motion compensation was performed in step 67002 or the vehicle coordinates of the previous reference frame (i.e., reference point cloud) are provided to step 67004.

[0427] The geometry information that has undergone local motion compensation in step 67004 is decoded and restored on an inter prediction basis in step 67005.

[0428] FIG. 26 is a diagram illustrating an example of a bitstream structure of point cloud data for transmission / reception according to an embodiment.

[0429] According to an embodiment, the term "slice" is also referred to as the term "data unit" in FIG.

[0430] 26, the abbreviations have the following meanings. Each abbreviation may also be referred to by other terms within the same range of meaning: SPS: Sequence Parameter Set, GPS: Geometry Parameter Set, APS: Attribute Parameter Set, TPS: Tile Parameter Set, Geometry (Geom: Geometry bitstream = geometry slice header + [geometry PU header + Geometry PU data] | geometry slice data), Attribute (Attr: Attribute bitstream = attribute data unit header + [attribute PU header + attribute PU data] | attribute data unit data).

[0431] This specification provides signaling information related to adding / executing the embodiments described above. The signaling information according to the embodiments is used by a point cloud video encoder at a transmitting end or a point cloud video decoder at a receiving end.

[0432] As described above, the point cloud video encoder according to the embodiment encodes the geometry information and the feature information to generate a bitstream as shown in Fig. 26. In addition, signaling information related to the point cloud data is generated and processed in at least one of the geometry encoder, the feature encoder, and the signaling processing unit of the point cloud video encoder and is included in the bitstream.

[0433] As an example, a point cloud video encoder that performs geometry encoding and / or feature encoding generates an encoded point cloud (or a bitstream including a point cloud) as shown in Figure 26. Furthermore, signaling information related to the point cloud data is generated and processed by a metadata processing unit of the point cloud data transmitting device and is included in the point cloud as shown in Figure 26.

[0434] The signaling information according to the embodiment is received / obtained in at least one of the geometry decoder, the feature decoder, and the signaling processing unit of the point cloud video decoder.

[0435] According to the embodiment, the bitstream may be divided into a geometry bitstream, a characteristic bitstream, and a signaling bitstream and transmitted / received, or may be merged into one bitstream and transmitted / received.

[0436] When the geometry bitstream, the attribute bitstream, and the signaling bitstream according to the embodiment are configured as one bitstream, the bitstream includes one or more sub-bitstreams. The bitstream according to the embodiment includes a Sequence Parameter Set (SPS) for sequence-level signaling, a Geometry Parameter Set (GPS) for signaling geometry information coding, one or more Attribute Parameter Sets (APS0, APS1) for signaling attribute information coding, a Tile Parameter Set (TPS) for tile-level signaling, and one or more slices (slice 0 to slice n). That is, the point cloud data bitstream according to the embodiment includes one or more tiles, and each tile is a group of slices including one or more slices (slice 0 to slice n). The TPS according to the embodiment includes information about each tile (e.g., bounding box coordinate value information and height / size information) for one or more tiles. Each slice contains one geometry bitstream (Geom0) and one or more attribute bitstreams (Attr0, Attr1).

[0437] The geometry bitstream in each slice (also called a geometry slice) consists of a geometry slice header and one or more geometry PUs (Geom PU0, Geom PU1). Each geometry PU consists of a geometry PU header and geometry PU data.

[0438] Each attribute bitstream (also called an attribute slice) in each slice consists of an attribute slice header and one or more attribute PUs (Attr PU0, Attr PU1). Each attribute PU consists of an attribute PU header (attr PU header) and attribute PU data (attr PU data).

[0439] Optional information related to inter prediction according to the embodiment is additionally signaled in the GPS and / or TPS.

[0440] Optional information related to inter prediction according to the embodiment is signaled in addition to the geometry slice header for each slice.

[0441] Optional information related to inter prediction according to the embodiment is signaled in the geometry PU header.

[0442] According to an embodiment, parameters necessary for encoding and / or decoding point cloud data are newly defined in a parameter set of the point cloud data (e.g., SPS, GPS, APS, and TPS (or referred to as tile inventory)) and / or a header of the slice. For example, when encoding and / or decoding geometry information, the parameters are added to the geometry parameter set (GPS), when encoding and / or decoding geometry information, the parameters are added to the tile (TPS) and / or slice header when encoding and / or decoding based on a tile, and when encoding and / or decoding based on a PU, the parameters are added to the geometry PU header and / or attribute PU header.

[0443] As shown in Figure 26, a bitstream of point cloud data is divided into tiles, slices, LPUs, and / or PUs so that the point cloud data is processed by dividing it into regions. According to the embodiment, each region of the bitstream has a different importance. Therefore, when the point cloud data is divided into tiles, a different filter (encoding method) and a different filter unit are applied to each tile. Also, when the point cloud data is divided into slices, a different filter and a different filter unit are applied to each slice. Also, when the point cloud data is divided into PUs, a different filter and a different filter unit are applied to each PU.

[0444] The transmitting apparatus according to the embodiment transmits point cloud data according to the bitstream structure shown in Fig. 26, thereby applying different encoding operations according to importance and providing a method for using a high-quality encoding method for important areas. It also supports efficient encoding and transmission according to the characteristics of point cloud data and provides characteristic values ​​according to user requirements.

[0445] A receiving device according to an embodiment receives point cloud data according to the bitstream structure shown in Fig. 26, and applies different filtering (decoding methods) to each region (region divided into tiles or slices) instead of using a complex decoding (filtering) method on the entire point cloud data depending on the processing capacity of the receiving device, thereby providing better image quality to regions important to the user and ensuring appropriate latency in the system.

[0446] As described above, tiles or slices are provided to process point cloud data by dividing it into regions. When dividing point cloud data into regions, an option is set to generate different sets of adjacent points for each region, providing a choice between low complexity but slightly reduced reliability, or conversely, high complexity but high reliability.

[0447] According to an embodiment, at least one of the GPS, the TPS, the geometry slice header, and the geometry PU header includes optional information related to inter prediction. According to an embodiment, the optional information related to inter prediction includes information on a criterion type for splitting into LPUs (motion_block_lpu_split_type), information on a criterion for splitting into LPUs (e.g., motion_block_lpu_radius, motion_block_lpu_azimuth, or motion_block_lpu_elevation), information indicating whether a motion vector is present (motion_vector_flag or pu_has_motion_vector_flag), information on a split criterion order type for splitting into PUs (motion_block_pu_split_type), and an octree related to splitting into PUs. The optional information related to inter prediction includes reference order type information (Motion_block_pu_split_octree_type), information used as a reference when dividing into PUs (e.g., motion_block_pu_radius, motion_block_pu_azimuth, or motion_block_pu_elevation), local motion vector information corresponding to the PU, information indicating whether a block (or region) corresponding to an LPU / PU is divided, and minimum PU size information (e.g., motion_block_pu_min_radius, motion_block_pu_min_azimuth, or motion_block_pu_min_elevation). In addition, the optional information related to inter prediction further includes information for identifying a tile to which a PU belongs, information for identifying a slice to which a PU belongs, information on the number of PUs included in the slice, information for identifying each PU, etc.

[0448] The term field, as used in the syntax of this specification below, is synonymous with parameter or element.

[0449] Figure 27 shows an example of the syntax structure of a sequence parameter set (seq_parameter_set_rbsp()) (SPS) according to this specification. The SPS contains sequence information for the point cloud data bitstream, and in particular, optional information related to neighbor point selection.

[0450] An SPS according to the embodiment includes a profile_idc field, a profile_compatibility_flags field, a level_idc field, an sps_bounding_box_present_flag field, an sps_source_scale_factor field, an sps_seq_parameter_set_id field, an sps_num_attribute_sets field, and an sps_extension_present_flag field.

[0451] The profile_idc field indicates the profile to which the bitstream conforms.

[0452] A value of 1 in the profile_compatibility_flags field indicates that the bitstream conforms to the profile indicated by the profile_idc field.

[0453] The level_idc field indicates the level to which the bitstream conforms.

[0454] The sps_bounding_box_present_flag field indicates whether source bounding box information is signaled to the SPS. The source bounding box information includes source bounding box offset and size information. For example, if the value of the sps_bounding_box_present_flag field is 1, source bounding box information is signaled to the SPS, and if it is 0, it is not signaled. The sps_source_scale_factor field indicates the scale factor of the source point cloud.

[0455] The sps_seq_parameter_set_id field provides an identifier for the SPS for reference by other syntax elements.

[0456] The sps_num_attribute_sets field indicates the number of coded attributes in the bitstream.

[0457] The sps_extension_present_flag field indicates whether the sps_extension_data syntax structure is present in the corresponding SPS syntax structure. For example, if the value of the sps_extension_present_flag field is 1, the sps_extension_data syntax structure is present in this SPS syntax structure, and if it is 0, it is not present (equal to 1 specifies that the sps_extension_data syntax structure is present in the SPS syntax structure. The sps_extension_present_flag field equal to 0 specifies that this syntax structure is not present. When not present, the value of the sps_extension_present_flag field is inferred to be equal to 0).

[0458] An SPS according to the embodiment further includes the following fields when the value of the sps_bounding_box_present_flag field is 1: sps_bounding_box_offset_x field, sps_bounding_box_offset_y field, sps_bounding_box_offset_z field, sps_bounding_box_scale_factor field, sps_bounding_box_size_width field, sps_bounding_box_size_height field, and sps_bounding_box_size_depth field.

[0459] The sps_bounding_box_offset_x field indicates the x-offset of the source bounding box in Cartesian coordinates. If there is no x-offset of the source bounding box, the value of the sps_bounding_box_offset_x field is 0.

[0460] The sps_bounding_box_offset_y field indicates the y-offset of the source bounding box in the Cartesian coordinate system. If there is no y-offset of the source bounding box, the value of the sps_bounding_box_offset_y field is 0.

[0461] The sps_bounding_box_offset_z field indicates the z-offset of the source bounding box in the Cartesian coordinate system. If there is no z-offset of the source bounding box, the value of the sps_bounding_box_offset_z field is 0.

[0462] The sps_bounding_box_scale_factor field indicates the scale factor of the source bounding box in Cartesian coordinates. If there is no source bounding box scale factor, the value of the sps_bounding_box_scale_factor field is 1.

[0463] The sps_bounding_box_size_width field indicates the width of the source bounding box in Cartesian coordinates. If the source bounding box width is not present, the value of the sps_bounding_box_size_width field is 1.

[0464] The sps_bounding_box_size_height field indicates the height of the source bounding box in Cartesian coordinates. If the source bounding box height is not present, the value of the sps_bounding_box_size_height field is 1.

[0465] The sps_bounding_box_size_depth field indicates the depth of the source bounding box in the Cartesian coordinate system. If the source bounding box depth does not exist, the value of the sps_bounding_box_size_depth field is 1.

[0466] An SPS according to the embodiment includes a repeat statement that is repeated the number of times equal to the value of the sps_num_attribute_sets field. In this example, i is initialized to 0 and increments by 1 each time the repeat statement is executed, and the repeat statement is repeated until the value of i becomes equal to the value of the sps_num_attribute_sets field. This repeat statement includes an attribute_dimension[i] field, an attribute_instance_id[i] field, an attribute_bitdepth[i] field, an attribute_cicp_colour_primaries[i] field, an attribute_cicp_transfer_characteristics[i] field, an attribute_cicp_matrix_coeffs[i] field, an attribute_cicp_video_full_range_flag[i] field, and a known_attribute_label_flag[i] field.

[0467] The attribute_dimension[i] field specifies the number of components of the i-th attribute.

[0468] The attribute_instance_id[i] field indicates the instance identifier of the ith attribute.

[0469] The attribute_bitdepth[i] field specifies the bitdepth of the i-th attribute signal(s).

[0470] The Attribute_cicp_colour_primaries[i] field indicates the chromaticity coordinates of the colour attribute source primaries of the i-th attribute.

[0471] The attribute_cicp_transfer_characteristics[i] field indicates the reference opto-electronic transfer characteristic function of the color attribute as a function of a source input linear optical intensity with a nominal real-valued range of 0 to 1 or indicates the inverse of the reference electro-optical transfer characteristic function as a function of an output linear optical intensity for the i-th attribute.

[0472] The attribute_cicp_matrix_coeffs[i] field describes the matrix coefficients used in deriving luma and chroma signals from the green, blue, and red (or Y, Z, and X primaries) of the i-th attribute.

[0473] The attribute_cicp_video_full_range_flag[i] field specifies indicates the black level and range of the luma and chroma signals as derived from E'Y, E'PB, and E'PR or E'R, E'G, and E'B real-valued component signals of the i-th attribute.

[0474] The known_attribute_label_flag[i] field indicates whether the known_attribute_label field or the attribute_label_four_bytes field is signaled for the i-th attribute. For example, a value of 1 in the known_attribute_label_flag[i] field indicates that the known_attribute_label field is signaled for the i-th attribute, and a value of 1 in the known_attribute_label_flag[i] field indicates that the attribute_label_four_bytes field is signaled for the i-th attribute.

[0475] The known_attribute_label[i] field indicates the type of attribute. For example, a value of 0 in the known_attribute_label[i] field indicates that the i-th attribute is color, a value of 1 in the known_attribute_label[i] field indicates that the i-th attribute is reflectance, and a value of 1 in the known_attribute_label[i] field indicates that the i-th attribute is frame index.

[0476] The attribute_label_four_bytes field indicates a known attribute type as a four-byte code.

[0477] In one embodiment, a value of 0 in the attribute_label_four_bytes field indicates color and a value of 1 indicates reflectance.

[0478] An SPS according to the embodiment may further include an SPS_extension_data_flag field if the value of the SPS_extension_present_flag field is 1.

[0479] The sps_extension_data_flag field can have any value.

[0480] Figure 28 illustrates one embodiment of the syntax structure of a geometry parameter set (geometry_parameter_set()) (GPS) according to this specification. A GPS according to an embodiment contains information about how to encode the geometry information of the point cloud data contained in one or more slices.

[0481] The GPS according to the embodiment includes a gps_geom_parameter_set_id field, a gps_seq_parameter_set_id field, a gps_box_present_flag field, a unique_geometry_points_flag field, a neighbor_context_restriction_flag field, an inferred_direct_coding_mode_enabled_flag field, a bitwise_occupancy_coding_flag field, an adjacent_child_contextualization_enabled_flag field, a log2_neighbour_avail_boundary field, a log2_intra_pred_max_node_size field, a log2_trisoup_node_size field, and a gps_extension_present_flag field.

[0482] The gps_geom_parameter_set_id field provides an identifier for the GPS for reference by other syntax elements.

[0483] The gps_seq_parameter_set_id field indicates the value of the seq_parameter_set_id field for the active SPS (specifies the value of sps_seq_parameter_set_id for the active SPS).

[0484] The gps_box_present_flag field indicates whether additional bounding box information is provided in the geometry slice header that references the current GPS. For example, a value of 1 in the gps_box_present_flag field indicates that additional bounding box information is provided in the geometry header that references the current GPS. Thus, if the value of the gps_box_present_flag field is 1, the GPS also includes a gps_gsh_box_log2_scale_present_flag field.

[0485] The gps_gsh_box_log2_scale_present_flag field indicates whether the gps_gsh_box_log2_scale field is signaled in each geometry slice header that references the current GPS. For example, a value of 1 in the gps_gsh_box_log2_scale_present_flag field indicates that the gps_gsh_box_log2_scale field is signaled in each geometry slice header that references the current GPS. As another example, a value of 0 in the gps_gsh_box_log2_scale_present_flag field indicates that the gps_gsh_box_log2_scale field is not signaled in each geometry slice header that references the current GPS, and a common scale for all slices is signaled in the gps_gsh_box_log2_scale field of the current GPS.

[0486] If the value of the gps_gsh_box_log2_scale_present_flag field is 0, the GPS also includes the gps_gsh_box_log2_scale field.

[0487] The gps_gsh_box_log2_scale field indicates the common scale factor of the bounding box origin for all slices that reference the current GPS.

[0488] The unique_geometry_points_flag field indicates whether all output points have unique positions. For example, if the value of the unique_geometry_points_flag field is 1, it indicates that all output points have unique positions. If the value of the unique_geometry_points_flag field is 0, it indicates that two or more output points may have the same positions (equal to 1 indicates that all output points have unique positions. unique_geometry_points_flag field equal to 0 indicates that the output points may have the same positions).

[0489] The neighbor_context_restriction_flag field indicates the context that octree occupancy coding uses. For example, a value of 0 in the neighbor_context_restriction_flag field indicates that octree occupancy coding uses contexts determined from six neighboring parent nodes. A value of 1 in the neighbor_context_restriction_flag field indicates that octree occupancy coding uses contexts determined from sibling nodes only.

[0490] The inferred_direct_coding_mode_enabled_flag field indicates whether the direct_mode_flag field exists in the corresponding geometry node syntax. For example, if the value of the inferred_direct_coding_mode_enabled_flag field is 1, it indicates that the direct_mode_flag field exists in the corresponding geometry node syntax. For example, if the value of the inferred_direct_coding_mode_enabled_flag field is 0, it indicates that the direct_mode_flag field does not exist in the corresponding geometry node syntax.

[0491] The bitwise_occupancy_coding_flag field indicates whether the geometry node occupancy is coded using the bitwise contextualization of its syntax element occupancy_map. For example, a value of 1 in the bitwise_occupancy_coding_flag field indicates that the geometry node occupancy is coded using the bit contextualization of its syntax element occupancy_map. For example, a value of 0 in the bitwise_occupancy_coding_flag field indicates that the geometry node occupancy is coded using the directory-coded syntax element occupancy_byte.

[0492] The adjacent_child_contextualization_enabled_flag field indicates whether the adjacent children of neighbouring octree nodes are used for bitwise occupancy contextualization. For example, if the value of the adjacent_child_contextualization_enabled_flag field is 1, it indicates that the adjacent children of neighbouring octree nodes are used for bitwise occupancy contextualization. For example, if the value of the adjacent_child_contextualization_enabled_flag field is 0, it indicates that the children of neighbouring octree nodes are not used for bitwise occupancy contextualization.

[0493] The log2_neighbour_avail_boundary field specifies the value of the variable NeighbAvailBoundary that is used in the decoding process as follows:

[0494] NeighbAvailBoundary = 2 log2_neighbour_avail_boundary

[0495] For example, if the value of the neighbour_context_restriction_flag field is 1, NeighbAvailabilityMask is set to 1. For example, if the value of the neighbour_context_restriction_flag field is 0, NeighbAvailabilityMask is set to 1 << log2_neighbour_avail_boundary.

[0496] The log2_intra_pred_max_node_size field specifies the octree nodesize eligible for occupancy intra prediction.

[0497] The log2_trisoup_node_size field specifies the variable TrisoupNodeSize as the size of the triangle nodes as follows:

[0498] TrisoupNodeSize=1< <log2_trisoup_node_size

[0499] The gps_extension_present_flag field indicates whether the gps_extension_data syntax structure exists in the corresponding GPS syntax. For example, if the value of the gps_extension_present_flag field is 1, it indicates that the gps_extension_data syntax structure exists in the corresponding GPS syntax. For example, if the value of the gps_extension_present_flag field is 0, it indicates that the gps_extension_data syntax structure does not exist in the corresponding GPS syntax.

[0500] In the embodiment, if the value of the gps_extension_present_flag field is 1, the GPS further includes a gps_extension_data_flag field.

[0501] The gps_extension_data_flag field may have any value. Its presence and value do not affect decoder conformance to profiles.

[0502] 29 is a diagram illustrating an example of a syntax structure of a geometry parameter set (geometry_parameter_set()) (GPS) including optional information related to inter prediction according to an embodiment. The names of the signaling information can be understood within the meaning and function of the signaling information.

[0503] In Figure 29, the gps_geom_parameter_set_id field provides an identifier for the GPS for reference by other syntax elements.

[0504] The gps_seq_parameter_set_id field indicates the value of the seq_parameter_set_id field for the active SPS (specifies the value of sps_seq_parameter_set_id for the active SPS).

[0505] The geom_tree_type field indicates the coding type of the geometry information. For example, if the value of the geom_tree_type field is 0, it indicates that the geometry information (i.e., position information) is coded using an octree, and if it is 1, it indicates that it is coded using a predictive tree.

[0506] The GPS according to the embodiment includes a motion_block_lpu_split_type field for each LPU.

[0507] The motion_block_lpu_split_type field specifies the type of LPU splitting criteria applied to a frame. For example, a value of 0 in the motion_block_lpu_split_type field indicates a radius-based LPU splitting method, a value of 1 indicates an azimuth-based LPU splitting method, and a value of 2 indicates an altitude (or vertical)-based LPU splitting method.

[0508] If the value of the motion_block_lpu_split_type field is 0, the GPS further includes a motion_block_lpu_radius field, which specifies the radius size used as the basis for LPU splitting applied to a frame.

[0509] If the value of the motion_block_lpu_split_type field is 1, the GPS further includes a motion_block_lpu_azimuth field, which specifies the azimuth angle size that is the basis for LPU splitting applied to a frame.

[0510] If the value of the motion_block_lpu_split_type field is 2, the GPS further includes a motion_block_lpu_elevation field, which specifies the altitude size used as a reference when splitting the LPU applied to the frame.

[0511] In this specification, the motion_block_lpu_radius field, motion_block_lpu_azimuth field, and motion_block_lpu_elevation field are referred to as information that serves as a reference when dividing the fields into LPUs.

[0512] The GPS according to the embodiment includes at least one of a motion_block_pu_split_octree_type field, a motion_block_pu_split_type field, a motion_block_pu_radius field, a motion_block_pu_azimuth field, a motion_block_pu_elevation field, a motion_block_pu_min_radius field, a motion_block_pu_min_azimuth field, and a motion_block_min_elevation field for each PU.

[0513] For example, if the value of the geom_tree_type field is 0 (ie, indicating that the geometry information (ie, location information) is coded using an octree), the GPS includes the motion_block_pu_split_octree_type field.

[0514] Also, if the value of the geom_tree_type field is 1 (i.e., indicating that the geometry information (i.e., location information) is coded using a prediction tree), the GPS includes the motion_block_pu_split_type field, motion_block_pu_radius field, motion_block_pu_azimuth field, motion_block_pu_elevation field, motion_block_pu_min_radius field, motion_block_pu_min_azimuth field, and motion_block_pu_min_elevation field.

[0515] The motion_block_pu_split_octree_type field indicates reference order type information related to an octree for splitting into PUs when geometry coding is performed based on an octree. That is, the motion_block_pu_split_octree_type field specifies the reference order type for splitting into PUs when geometry coding is applied based on an octree applied to a frame.

[0516] For example, if the value of the motion_block_pu_split_octree_type field is 0, it indicates the x→y→z base split application method, if it is 1, it indicates the x→z→y base split application method, if it is 2, it indicates the y→x→z base split application method, if it is 3, it indicates the y→z→x base split application method, if it is 4, it indicates the z→x→y base split application method, if it is 5, it indicates the z→y→x base split application method.

[0517] The motion_block_pu_split_type field is referred to as splitting criterion order type information for splitting an LPU into PUs, and specifies the criterion type for splitting into PUs applied to a frame.

[0518] For example, if the value of the motion_block_pu_split_type field is 0, it indicates the split application method of radius base → azimuth base → altitude base, if it is 1, it indicates the split application method of radius base → altitude base → azimuth base, if it is 2, it indicates the split application method of azimuth base → radius base → altitude base, if it is 3, it indicates the split application method of azimuth base → altitude base → radius base, if it is 4, it indicates the split application method of altitude base → radius base → azimuth base, and if it is 5, it indicates the split application method of altitude base → azimuth base → radius base.

[0519] The motion_block_pu_radius field specifies the radius size that is used as a reference when dividing a frame into PUs.

[0520] The motion_block_pu_azimuth field specifies the azimuth angle size that is used as the reference when dividing a frame into PUs.

[0521] The motion_block_pu_elevation field specifies the altitude size that is used as the reference when dividing a frame into PUs.

[0522] In this specification, the motion_block_pu_radius field, motion_block_pu_azimuth field, and motion_block_pu_elevation field are referred to as information that serves as a reference when dividing into PUs.

[0523] The motion_block_pu_min_radius field specifies the minimum radius size that is the basis for PU division applied to a frame. If the radius size of the PU block is smaller than the minimum radius size, it will not be divided any further.

[0524] The motion_block_pu_min_azimuth field specifies the minimum azimuth size that is the basis for PU division applied to a frame. If the azimuth size of a PU block is smaller than the minimum azimuth size, no further division is performed.

[0525] The motion_block_pu_min_elevation field specifies the minimum elevation size that is the basis for PU division applied to a frame. If the elevation value of the PU block is smaller than the minimum elevation size, it will not be divided any further.

[0526] In this specification, the motion_block_pu_min_radius field, the motion_block_pu_min_azimuth field, and the motion_block_pu_min_elevation field are referred to as minimum PU size information.

[0527] According to an embodiment, optional information related to inter prediction in FIG. 29 is included in any position of the GPS in FIG.

[0528] 30 illustrates an example of a syntax structure for a tile parameter set (tile_parameter_set()) (TPS) according to this specification. In some embodiments, the TPS (Tile Parameter Set) is also referred to as a tile inventory. The TPS according to some embodiments includes, for each tile, information related to each tile.

[0529] The TPS according to the embodiment includes a num_tiles field.

[0530] The num_tiles field indicates the number of tiles signaled for the bitstream. When no tiles are present, the value of the num_tiles field is 0.

[0531] The TPS according to the embodiment includes a loop statement that is repeated the number of times equal to the value of the num_tiles field. In this example, i is initialized to 0 and increments by 1 each time the loop statement is executed, and the loop statement is repeated until the value of i becomes the value of the num_tiles field. This loop statement includes a tile_bounding_box_offset_x[i] field, a tile_bounding_box_offset_y[i] field, a tile_bounding_box_offset_z[i] field, a tile_bounding_box_size_width[i] field, a tile_bounding_box_size_height[i] field, and a tile_bounding_box_size_depth[i] field.

[0532] The tile_bounding_box_offset_x[i] field indicates the x offset of the ith tile in the Cartesian coordinate system.

[0533] The tile_bounding_box_offset_y[i] field indicates the y offset of the ith tile in the Cartesian coordinate system.

[0534] The tile_bounding_box_offset_z[i] field indicates the z offset of the ith tile in the Cartesian coordinate system.

[0535] The tile_bounding_box_size_width[i] field indicates the width of the ith tile in the Cartesian coordinate system.

[0536] The tile_bounding_box_size_height[i] field indicates the height of the ith tile in the Cartesian coordinate system.

[0537] The tile_bounding_box_size_depth[i] field indicates the depth of the ith tile in the Cartesian coordinate system.

[0538] 31 is a diagram illustrating an example of a syntax structure of a tile parameter set (tile_parameter_set()) (TPS) including optional information related to inter prediction according to an embodiment. The names of the signaling information can be understood within the meaning and function of the signaling information.

[0539] In Figure 31, the explanations for the num_tiles field, tile_bounding_box_offset_x[i] field, tile_bounding_box_offset_y[i] field, etc. are the same as those in Figure 30, so they will be omitted here to avoid redundant explanation.

[0540] The TPS according to the embodiment includes a motion_block_lpu_split_type field for each LPU.

[0541] The motion_block_lpu_split_type field specifies the criterion type for splitting into LPUs applied to a tile. For example, a value of 0 in the motion_block_lpu_split_type field indicates a radius-based LPU split method, a value of 1 indicates an azimuth-based LPU split method, and a value of 2 indicates an altitude-based LPU split method.

[0542] If the value of the motion_block_lpu_split_type field is 0, the TPS further includes a motion_block_lpu_radius field, which specifies the radius size used as the basis for LPU splitting applied to a tile.

[0543] If the value of the motion_block_lpu_split_type field is 1, the TPS further includes a motion_block_lpu_azimuth field, which specifies the azimuth angle size that is used as the basis for LPU splitting applied to a tile.

[0544] If the value of the motion_block_lpu_split_type field is 2, the TPS further includes a motion_block_lpu_elevation field, which specifies the elevation size used as the reference when splitting the LPU applied to the tile.

[0545] In this specification, the motion_block_lpu_radius field, motion_block_lpu_azimuth field, and motion_block_lpu_elevation field are referred to as information that serves as a reference when dividing the fields into LPUs.

[0546] The TPS according to the embodiment includes at least one of a motion_block_pu_split_octree_type field, a motion_block_pu_split_type field, a motion_block_pu_radius field, a motion_block_pu_azimuth field, a motion_block_pu_elevation field, a motion_block_pu_min_radius field, a motion_block_pu_min_azimuth field, and a motion_block_min_elevation field for each PU.

[0547] For example, if the value of the geom_tree_type field is 0 (i.e., indicating that the geometry information (i.e., position information) is coded using an octree), the TPS includes a motion_block_pu_split_octree_type field.

[0548] Also, if the value of the geom_tree_type field is 1 (i.e., indicating that geometry information (i.e., position information) is coded using a prediction tree), the TPS includes a motion_block_pu_split_type field, a motion_block_pu_radius field, a motion_block_pu_azimuth field, a motion_block_pu_elevation field, a motion_block_pu_min_radius field, a motion_block_pu_min_azimuth field, and a motion_block_pu_min_elevation field.

[0549] The motion_block_pu_split_octree_type field indicates reference order type information related to an octree for splitting into PUs when geometry coding is performed based on an octree. That is, the motion_block_pu_split_octree_type field specifies the reference order type for splitting into PUs when geometry coding is applied based on an octree applied to a tile.

[0550] For example, if the value of the motion_block_pu_split_octree_type field is 0, it indicates an x→y→z-based split application method, if it is 1, it indicates an x→z→y-based split application method, if it is 2, it indicates a y→x→z-based split application method, if it is 3, it indicates a y→z→x-based split application method, if it is 4, it indicates a z→x→y-based split application method, and if it is 5, it indicates a z→y→x-based split application method.

[0551] The motion_block_pu_split_type field is referred to as splitting criterion order type information for splitting an LPU into PUs, and specifies the criterion type for splitting into PUs applied to a tile.

[0552] For example, if the value of the motion_block_pu_split_type field is 0, it indicates the split application method of radius base → azimuth base → altitude base, if it is 1, it indicates the split application method of radius base → altitude base → azimuth base, if it is 2, it indicates the split application method of azimuth base → radius base → altitude base, if it is 3, it indicates the split application method of azimuth base → altitude base → radius base, if it is 4, it indicates the split application method of altitude base → radius base → azimuth base, and if it is 5, it indicates the split application method of altitude base → azimuth base → radius base.

[0553] The motion_block_pu_radius field specifies the radius size that is used as the basis for the PU division applied to the tile.

[0554] The motion_block_pu_azimuth field specifies the azimuth angle size that is used as the reference when dividing a PU into tiles.

[0555] The motion_block_pu_elevation field specifies the altitude size that is used as the reference when dividing a PU into tiles.

[0556] In this specification, the motion_block_pu_radius field, motion_block_pu_azimuth field, and motion_block_pu_elevation field are referred to as information that serves as a reference when dividing into PUs.

[0557] The motion_block_pu_min_radius field specifies the minimum radius size that is the basis for PU division applied to the tile. If the radius size of the PU block is smaller than the minimum radius size, it will not be divided any further.

[0558] The motion_block_pu_min_azimuth field specifies the minimum azimuth size that is the basis for PU division applied to a tile. If the azimuth size of the PU block is smaller than the minimum azimuth size, no further division will be performed.

[0559] The motion_block_pu_min_elevation field specifies the minimum elevation size that is the basis for PU division applied to a tile. If the elevation value of the PU block is smaller than the minimum elevation size, it will not be divided any further.

[0560] In this specification, the motion_block_pu_min_radius field, the motion_block_pu_min_azimuth field, and the motion_block_pu_min_elevation field are referred to as minimum PU size information.

[0561] According to an embodiment, optional information related to inter prediction in FIG. 31 may be included at any position in the TPS in FIG.

[0562] FIG. 32 shows one embodiment of the syntax structure of a geometry slice bitstream() according to this specification.

[0563] The geometry slice bitstream (geometry_slice_bitstream()) according to the embodiment includes a geometry slice header (geometry_slice_header()) and geometry slice data (geometry_slice_data()).

[0564] FIG. 33 shows one embodiment of the syntax structure of a geometry slice header (geometry_slice_header()) according to this specification.

[0565] A bitstream transmitted by a transmitting device (or received by a receiving device) according to an embodiment includes one or more slices. Each slice includes a geometry slice and an attribute slice. A geometry slice includes a geometry slice header (GSH). An attribute slice includes an attribute slice header (ASH).

[0566] The geometry slice header (geometry_slice_header()) according to the embodiment includes a gsh_geom_parameter_set_id field, a gsh_tile_id field, a gsh_slice_id field, a gsh_max_node_size_log2 field, a gsh_num_points field, and a byte_alignment() field.

[0567] In the embodiment, the geometry slice header (geometry_slice_header()) further includes a gsh_box_log2_scale field, a gsh_box_origin_x field, a gsh_box_origin_y field, and a gsh_box_origin_z field when the value of the gps_box_present_flag field included in the geometry parameter set (GPS) is true (e.g., 1) and the value of the gps_gsh_box_log2_scale_present_flag field is true (e.g., 1).

[0568] The gsh_geom_parameter_set_id field specifies the value of the gps_geom_parameter_set_id of the active GPS.

[0569] The gsh_tile_id field indicates the identifier of the tile referenced by the geometry slice header (GSH).

[0570] The gsh_slice_id indicates the identifier of the slice for reference by other syntax elements.

[0571] The gsh_box_log2_scale field indicates the scaling factor of the bounding box origin for the slice.

[0572] The gsh_box_origin_x field indicates the x value of the bounding box origin scaled by the value of the gsh_box_log2_scale field.

[0573] The gsh_box_origin_y field indicates the y value of the bounding box origin scaled by the value of the gsh_box_log2_scale field.

[0574] The gsh_box_origin_z field indicates the z value of the bounding box origin scaled by the value of the gsh_box_log2_scale field.

[0575] The gsh_max_node_size_log2 field indicates the size of the root geometry octree node.

[0576] The gsh_points_number field indicates the number of coded points in the slice.

[0577] 34 is a diagram illustrating an example of a syntax structure of a geometry slice header (geometry_slice_header()) including optional information related to inter prediction according to an embodiment. The names of the signaling information can be understood within the meaning and function of the signaling information.

[0578] A bitstream transmitted by a transmitting device (or a bitstream received by a receiving device) according to an embodiment includes one or more slices.

[0579] In Figure 34, the gsh_geometry_parameter_set_id field, gsh_tile_id field, gsh_slice_id field, etc. are the same as those in Figure 33, so they will be omitted here to avoid redundant explanation.

[0580] The geometry slice header according to the embodiment includes a motion_block_lpu_split_type field for each LPU.

[0581] The motion_block_lpu_split_type field specifies the type of LPU splitting criteria applied to the slice. For example, a value of 0 in the motion_block_lpu_split_type field indicates the radius-based LPU splitting method, a value of 1 indicates the azimuth-based LPU splitting method, and a value of 2 indicates the altitude-based LPU splitting method.

[0582] If the value of the motion_block_lpu_split_type field is 0, the geometry slice header further includes a motion_block_lpu_radius field, which specifies the radius size that is used as the basis for LPU splitting applied to the slice.

[0583] If the value of the motion_block_lpu_split_type field is 1, the geometry slice header further includes a motion_block_lpu_azimuth field, which specifies the azimuth angle size that is used as the basis for LPU splitting applied to the slice.

[0584] If the value of the motion_block_lpu_split_type field is 2, the geometry slice header further includes a motion_block_lpu_elevation field, which specifies the elevation size that is used as the reference when LPU splitting is applied to the slice.

[0585] In this specification, the motion_block_lpu_radius field, motion_block_lpu_azimuth field, and motion_block_lpu_elevation field are referred to as information that serves as a reference when dividing the fields into LPUs.

[0586] The geometry slice header according to the embodiment includes at least one of a motion_block_pu_split_octree_type field, a motion_block_pu_split_type field, a motion_block_pu_radius field, a motion_block_pu_azimuth field, a motion_block_pu_elevation field, a motion_block_pu_min_radius field, a motion_block_pu_min_azimuth field, and a motion_block_min_elevation field for each PU.

[0587] For example, if the value of the geom_tree_type field is 0 (ie, indicating that the geometry information (ie, position information) is coded using an octree), the geometry slice header includes a motion_block_pu_split_octree_type field.

[0588] Also, if the value of the geom_tree_type field is 1 (i.e., indicating that the geometry information (i.e., position information) is coded using a prediction tree), the geometry slice header includes a motion_block_pu_split_type field, a motion_block_pu_radius field, a motion_block_pu_azimuth field, a motion_block_pu_elevation field, a motion_block_pu_min_radius field, a motion_block_pu_min_azimuth field, and a motion_block_pu_min_elevation field.

[0589] The motion_block_pu_split_octree_type field indicates reference order type information related to an octree for splitting into PUs when geometry coding is performed based on an octree. That is, the motion_block_pu_split_octree_type field specifies the reference order type for splitting into PUs when geometry coding is applied based on an octree applied to a slice.

[0590] For example, if the value of the motion_block_pu_split_octree_type field is 0, it indicates an x→y→z-based split application method, if it is 1, it indicates an x→z→y-based split application method, if it is 2, it indicates a y→x→z-based split application method, if it is 3, it indicates a y→z→x-based split application method, if it is 4, it indicates a z→x→y-based split application method, and if it is 5, it indicates a z→y→x-based split application method.

[0591] The motion_block_pu_split_type field is referred to as splitting criterion order type information for splitting an LPU into PUs, and specifies the criterion type for splitting into PUs applied to a slice.

[0592] For example, if the value of the motion_block_pu_split_type field is 0, it indicates the split application method of radius base → azimuth base → altitude base, if it is 1, it indicates the split application method of radius base → altitude base → azimuth base, if it is 2, it indicates the split application method of azimuth base → radius base → altitude base, if it is 3, it indicates the split application method of azimuth base → altitude base → radius base, if it is 4, it indicates the split application method of altitude base → radius base → azimuth base, and if it is 5, it indicates the split application method of altitude base → azimuth base → radius base.

[0593] The motion_block_pu_radius field specifies the radius size that is used as the reference when dividing a slice into PUs.

[0594] The motion_block_pu_azimuth field specifies the azimuth angle size that is used as the reference when dividing a slice into PUs.

[0595] The motion_block_pu_elevation field specifies the reference elevation size when dividing a slice into PUs.

[0596] In this specification, the motion_block_pu_radius field, motion_block_pu_azimuth field, and motion_block_pu_elevation field are referred to as information that serves as a reference when dividing into PUs.

[0597] The motion_block_pu_min_radius field specifies the minimum radius size that is the basis for PU division applied to a slice. If the radius size of the PU block is smaller than the minimum radius size, it will not be divided any further.

[0598] The motion_block_pu_min_azimuth field specifies the minimum azimuth size that is the basis for PU division applied to a slice. If the azimuth size of a PU block is smaller than the minimum azimuth size, no further division is performed.

[0599] The motion_block_pu_min_elevation field specifies the minimum elevation size that is the basis for PU division applied to a slice. If the elevation value of a PU block is smaller than the minimum elevation size, it will not be divided any further.

[0600] In this specification, the motion_block_pu_min_radius field, the motion_block_pu_min_azimuth field, and the motion_block_pu_min_elevation field are referred to as minimum PU size information.

[0601] According to an embodiment, optional information related to inter prediction in FIG. 34 may be included anywhere in the geometry slice header in FIG.

[0602] According to an embodiment, a slice is divided into one or more PUs. For example, a geometry slice consists of a geometry slice header and one or more geometry PUs. In this case, each geometry PU consists of a geometry PU header (geom PU header) and geometry PU data (geom PU data).

[0603] 35 is a diagram illustrating an example of a syntax structure of a geometry PU header (geom_pu_header()) including optional information related to inter prediction according to an embodiment. The names of the signaling information can be understood within the meanings and functions of the signaling information.

[0604] The geometry PU header according to the embodiment includes a pu_tile_id field, a pu_slice_id field, and a pu_cnt field.

[0605] The pu_tile_id field specifies a tile identifier (ID) for identifying the tile to which the PU belongs.

[0606] The pu_slice_id field specifies a slice identifier (ID) for identifying the slice to which the PU belongs.

[0607] The pu_cnt field specifies the number of PUs included in the slice identified by the value of the pu_slice_id field.

[0608] The geometry PU header according to the embodiment includes a repeat statement that is repeated the same number of times as the value of the pu_cnt field. In this example, puIdx is initialized to 0 and increments by 1 each time the repeat statement is executed, and the repeat statement is repeated until the puIdx value becomes equal to the value of the pu_cnt field. This repeat statement includes a pu_id[puIdx] field, a pu_split_flag[puIdx] field, a pu_motion_compensation_type[puIdx] field, and a pu_has_motion_vector_flag[puIdx] field.

[0609] The pu_id[puIdx] field specifies a PU identifier (ID) for identifying the PU corresponding to puIdx among the PUs included in the slice.

[0610] The pu_split_flag[puIdx] field indicates whether or not the PU corresponding to puIdx among the PUs included in the slice has been further split later.

[0611] The pu_motion_compensation_type[puIdx] field indicates whether a motion vector is applied to the PU corresponding to puIdx among the PUs included in the slice. According to an embodiment, the pu_motion_compensation_type[puIdx] field indicates whether a global motion vector is applied to the PU corresponding to puIdx among the PUs included in the slice. According to an embodiment, the pu_motion_compensation_type[puIdx] field indicates whether a local motion vector is applied to the PU corresponding to puIdx among the PUs included in the slice. According to an embodiment, the pu_motion_compensation_type[puIdx] field indicates that a motion vector is not applied to the PU corresponding to puIdx among the PUs included in the slice. For example, a value of 0 in the pu_motion_compensation_type[puIdx] field indicates that no motion vector is applied to the PU, a value of 1 indicates that a global motion vector is applied, and a value of 2 indicates that a local motion vector is applied.

[0612] Therefore, the geometry decoder on the receiving side determines that a global motion vector is not applied to the PU if the value of the pu_motion_compensation_type[puIdx] field is 0, and that a global motion vector is applied to the PU if the value is 1. Therefore, if the value of the pu_motion_compensation_type[puIdx] field is 1, the motion compensation application unit of the geometry decoder on the receiving side uses the point of the previous frame as is if the value of the pu_motion_compensation_type[puIdx] field is 0, selects a point to which a global motion vector is applied to the PU, and performs motion compensation if the value is 1, and selects a point to which a local motion vector is applied to the PU and performs motion compensation if the value is 2.

[0613] The pu_has_motion_vector_flag[puIdx] field indicates whether or not the PU corresponding to puIdx among the PUs included in the slice has a motion vector. That is, the pu_has_motion_vector_flag[puIdx] field indicates whether or not the PU corresponding to puIdx among the PUs included in the slice has a motion vector that can be applied to it.

[0614] For example, if the value of the pu_has_motion_vector_flag[puIdx] field is 1, it indicates that the PU has an applicable motion vector, and if it is 0, it indicates that the PU does not have an applicable motion vector.

[0615] According to an embodiment, if the value of the pu_has_motion_vector_flag[puIdx] field is 1, it indicates that the PU identified by the value of the pu_id[puIdx] field has an applicable motion vector, in which case the geometry PU header further includes a pu_motion_vector_xyz[pu_id][k] field.

[0616] The pu_motion_vector_xyz[pu_id][k] field specifies the motion vector applied to the kth PU identified by the pu_id field.

[0617] FIG. 36 is a diagram illustrating an example of a syntax structure of a feature slice bitstream() according to this specification.

[0618] An attribute slice bitstream (attribute_slice_bitstream()) according to the embodiment includes an attribute slice header (attribute_slice_header()) and attribute slice data (attribute_slice_data()).

[0619] FIG. 37 is a diagram illustrating an example of a syntax structure of an attribute slice header (attribute_slice_header()) according to this specification.

[0620] The attribute slice header (attribute_slice_header()) according to the embodiment includes an ash_attr_parameter_set_id field, an ash_attr_sps_attr_idx field, an ash_attr_geom_slice_id field, an ash_attr_layer_qp_delta_present_flag field, and an ash_attr_region_qp_delta_present_flag field.

[0621] In the embodiment, the attribute slice header (attribute_slice_header()) further includes an ash_attr_qp_delta_luma field if the value of the aps_slice_qp_delta_present_flag field of the attribute parameter set (APS) is positive (e.g., 1), and if the value of the attribute_dimension_minus1[ash_attr_sps_attr_idx] field is greater than 0, the attribute slice header further includes an ash_attr_qp_delta_chroma field.

[0622] The ash_attr_parameter_set_id field indicates the value of the aps_attr_parameter_set_id field of the currently active APS.

[0623] The ash_attr_sps_attr_idx field indicates the attribute set in the currently active SPS.

[0624] The ash_attr_geom_slice_id field indicates the value of the gsh_slice_id field in the current geometry slice header.

[0625] The ash_attr_qp_delta_luma field indicates the luma delta quantization parameter (qp) derived from the initial slice qp in the active attribute parameter set.

[0626] The ash_attr_qp_delta_chroma field indicates the chroma delta quantization parameter (qp) derived from the initial slice qp in the active attribute parameter set.

[0627] In this case, the variables InitialSliceQpY and InitialSliceQpC are derived as follows:

[0628] InitialSliceQpY=aps_attrattr_initial_qp+ash_attr_qp_delta_luma

[0629] InitialSliceQpC=aps_attrtr_initial_qp+aps_attr_chroma_qp_offset+ash_attr_qp_delta_chroma

[0630] The ash_attr_layer_qp_delta_present_flag field indicates whether the ash_attr_layer_qp_delta_luma and ash_attr_layer_qp_delta_chroma fields are present in the attribute slice header (ASH) for each layer. For example, if the value of the ash_attr_layer_qp_delta_present_flag field is 1, the ash_attr_layer_qp_delta_luma and ash_attr_layer_qp_delta_chroma fields are present in the attribute slice header, and if the value is 0, they are not present.

[0631] If the value of the ash_attr_layer_qp_delta_present_flag field is positive, the attribute slice header further includes an ash_attr_num_layer_qp_minus1 field.

[0632] The ash_attr_num_layer_qp_minus1 field plus 1 indicates the number of layers for which the ash_attr_qp_delta_luma and ash_attr_qp_delta_chroma fields are signaled. If the ash_attr_num_layer_qp field is not signaled, the value of the ash_attr_num_layer_qp field is 0. According to the embodiment, NumLayerQp, which specifies the number of layers, is calculated by adding 0 to the value of the ash_attr_num_layer_qp_minus1 field (NumLayerQp = ash_attr_num_layer_qp_minus1 + 1).

[0633] According to an embodiment, if the value of the ash_attr_layer_qp_delta_present_flag field is positive, the geometry slice header includes a repeat statement as many times as the value of NumLayerQp. In this case, i is initialized to 0 and increments by 1 each time the repeat statement is executed. In one embodiment, the repeat statement is repeated until the value of i becomes equal to the value of NumLayerQp. This repeat statement includes an ash_attr_layer_qp_delta_luma[i] field. Furthermore, if the value of the attribute_dimension_minus1[ash_attr_sps_attr_idx] field is greater than 0, the repeat statement further includes an ash_attr_layer_qp_delta_chroma[i] field.

[0634] The ash_attr_layer_qp_delta_luma field indicates the luma delta quantization parameter (qp) derived from InitialSliceQpY in each layer.

[0635] The ash_attr_layer_qp_delta_chroma field indicates the chroma delta quantization parameter (qp) derived from InitialSliceQpC for each layer.

[0636] The variables SliceQpY[i] and SliceQpC[i] with i = 0…NumLayerQPNumQPLayer−1 are derived as follows:

[0637] for(i=0;i <NumLayerQPNumQPLayer;i++){

[0638] SliceQpY[i]=InitialSliceQpY+ash_attr_layer_qp_delta_luma[i]

[0639] SliceQpC[i]=InitialSliceQpC+ash_attr_layer_qp_delta_chroma[i]

[0640] }

[0641] In the embodiment, the attribute slice header (attribute_slice_header()) indicates that ash_attr_region_qp_delta, region bounding box origin, and size are present in the current attribute slice header if the value of the ash_attr_region_qp_delta_present_flag field is 1. If the value of the ash_attr_region_qp_delta_present_flag field is 0, it indicates that ash_attr_region_qp_delta, region bounding box origin, and size are not present in the current attribute slice header.

[0642] That is, if the value of the ash_attr_layer_qp_delta_present_flag field is 1, the attribute slice header further includes an ash_attr_qp_region_box_origin_x field, an ash_attr_qp_region_box_origin_y field, an ash_attr_qp_region_box_origin_z field, an ash_attr_qp_region_box_width field, an ash_attr_qp_region_box_height field, an ash_attr_qp_region_box_depth field, and an ash_attr_region_qp_delta field.

[0643] The ash_attr_qp_region_box_origin_x field indicates the x offset of the region bounding box relative to slice_origin_x.

[0644] The ash_attr_qp_region_box_origin_y field indicates the y offset of the region bounding box relative to slice_origin_y.

[0645] The ash_attr_qp_region_box_origin_z field indicates the z offset of the region bounding box relative to slice_origin_z.

[0646] The ash_attr_qp_region_box_size_width field indicates the width of the region bounding box.

[0647] The ash_attr_qp_region_box_size_height field indicates the height of the region bounding box.

[0648] The ash_attr_qp_region_box_size_depth field indicates the depth of the region bounding box.

[0649] The ash_attr_region_qp_delta field indicates the delta qp from SliceQpY[i] and SliceQpC[i] for the region specified by the ash_attr_qp_region_box field.

[0650] According to an embodiment, a variable RegionboxDeltaQp, which specifies the region box delta quantization parameter, is set equal to the value of the ash_attr_region_qp_delta field (RegionboxDeltaQp=ash_attr_region_qp_delta).

[0651] As mentioned above, when capturing a point cloud using lidar equipment for a moving / stationary vehicle, the angle mode (r, φ, i) is used. In this case, the larger the radius r for the same azimuth, the longer the arc. Therefore, small movements of objects close to the vehicle are likely to appear large and become local motion vectors. Also, for objects far from the vehicle, even if they move in the same way as objects close to the vehicle, they may not appear, and they are likely to be covered by global motion vectors without local motion vectors. Also, moving objects are likely to be divided into PUs according to the main areas captured.

[0652] In order to apply reference frame-based inter-prediction compression to point clouds with such characteristics, this specification supports a method of dividing the content into prediction units (LPUs / PUs) that reflects the characteristics of the content.

[0653] Therefore, this specification can expand the predictable range using local motion vectors, eliminate the need for additional calculations, and reduce encoding time. In addition, although moving objects are not accurately divided, PU division is performed to achieve the effect of separating objects, which can improve compression efficiency for inter-prediction of point cloud data.

[0654] In this way, the transmitting method / device can efficiently compress point cloud data and transmit the data, and by transmitting signaling information for this purpose, the receiving method / device can also efficiently decode / restore the point cloud data.

[0655] Each of the above-mentioned parts, modules, or units is software, a processor, or a hardware part that performs a continuous execution process stored in a memory (or a storage unit). Each step described in the above embodiments is performed by a processor, software, or hardware part. Each of the modules / blocks / units described in the above embodiments operates as a processor, software, or hardware. Also, the methods presented in the embodiments are implemented as code. This code is written in a processor-readable storage medium and is thus read by the processor provided by the device.

[0656] Furthermore, throughout the specification, when a part "includes" a certain element, this does not mean excluding other elements, but means including other elements as well, unless otherwise specified. Furthermore, the term "part" or the like used in the specification means a unit that processes at least one function or operation, and this may be realized by hardware, software, or a combination of hardware and software.

[0657] For the sake of convenience, the drawings have been described separately, but it is possible to combine the embodiments described in the drawings to design a new embodiment. Furthermore, it is within the scope of the embodiments to design a computer-readable recording medium on which a program for executing the previously described embodiments is recorded, as needed by those skilled in the art.

[0658] As described above, the devices and methods according to the embodiments are not limited to the configurations and methods of the described embodiments, and the embodiments can be modified in various ways and can be configured by selectively combining all or part of each embodiment.

[0659] Although preferred embodiments of the invention have been shown and described, the invention is not limited to the specific embodiments described above, and various modifications may be made by a person having ordinary skill in the art to which the invention pertains without departing from the spirit of the embodiments claimed in the claims, and such modifications should not be understood individually from the technical ideas or perspectives of the invention.

[0660] Various components of the apparatus according to the embodiments may be implemented by hardware, software, firmware, or a combination thereof. Various components of the embodiments may be implemented by a single chip, e.g., a single hardware circuit. In some embodiments, components according to the embodiments may be implemented by individual chips. In some embodiments, any of the components of the apparatus according to the embodiments may be implemented by one or more processors capable of executing one or more programs, and the one or more programs include instructions for causing or causing any one or more of the operations / methods according to the embodiments to be performed. Executable instructions for performing the methods / operations of the apparatus according to the embodiments may be stored in a non-transitory CRM or other computer program product configured to be executed by one or more processors, or may be stored in a transient CRM or other computer program product configured to be executed by one or more processors. In addition, the term "memory" in the embodiments is used as a concept that encompasses not only volatile memory (e.g., RAM, etc.), but also non-volatile memory, flash memory, PROM, etc. It may also be implemented in the form of a carrier wave, such as transmission over the Internet. In addition, a processor-readable recording medium may be distributed among network-connected computer systems, and the processor-readable code may be stored and executed in a distributed manner. In this specification, " / " and "," are interpreted as "and / or." For example, "A / B" is interpreted as "A and / or B," and "A, B" is interpreted as "A and / or B." Furthermore, "A / B / C" means "any of A, B, and / or C." Also, "A, B, C" means "any of A, B, and / or C."

[0661] Furthermore, in this document, "or" is to be interpreted as "and / or." For example, "A or B" means 1) "A" only, 2) "B" only, or 3) "A and B." In other words, in this specification, "or" means "additionally or alternatively."

[0662] Various elements of the embodiments may be implemented in hardware, software, firmware, or a combination thereof. Various elements of the embodiments may be implemented on a single chip, such as a hardware circuit. In some embodiments, the embodiments may alternatively be implemented on separate chips. In some embodiments, at least one of the elements of the embodiments may be implemented by one or more processors that include instructions for performing operations according to the embodiments.

[0663] Furthermore, operations according to the embodiments described herein are performed by a transceiver device including one or more memories and / or one or more processors according to the embodiments. The one or more memories store programs for processing / controlling operations according to the embodiments, and the one or more processors control various operations described herein. The one or more processors may also be referred to as controllers, etc. In the embodiments, operations are performed by firmware, software, and / or a combination thereof, and the firmware, software, and / or a combination thereof are stored in the processor or memory.

[0664] Terms such as "first," "second," and the like are used to describe various components of the embodiments. However, the interpretation of the various components according to the embodiments should not be limited by these terms. These terms are merely used to distinguish one component from another. For example, a first user input signal can be referred to as a second user input signal. Similarly, a second user input signal can be referred to as a first user input signal. The use of such terms does not depart from the scope of the various embodiments. Although the first user input signal and the second user input signal are both user input signals, they do not refer to the same user input signal unless the context clearly indicates otherwise.

[0665] Terms used to describe the embodiments are used to describe particular embodiments and are not intended to limit the embodiments. As used in the description of the embodiments and in the claims, the singular includes the plural unless the context clearly dictates otherwise. The term "and / or" is used to include all possible combinations between terms. "Comprises" describes the presence of a feature, number, step, element, and / or component, but does not mean that additional features, numbers, steps, elements, and / or components are not present. Conditional expressions such as "if" and "when" used to describe the embodiments are only used in selective cases and are not interpreted as limiting. It is intended that when a specific condition is met, a related operation is performed in response to the specific condition, or a related definition is interpreted. Furthermore, operations according to the embodiments described in this specification are performed by a transceiver device including a memory and / or a processor according to the embodiments. The memory stores programs for processing / controlling operations according to the embodiments, and the processor controls various operations described in this specification. The processor is also referred to as a controller. In the embodiments, operations are performed by firmware, software, and / or a combination thereof, and the firmware, software, and / or a combination thereof are stored in the processor or memory.

[0666] [Mode for carrying out the invention] The best mode for carrying out the embodiment has been described above. [Industrial Applicability]

[0667] As described above, the embodiments can be applied in whole or in part to point cloud data transmitting and receiving devices and systems. Those skilled in the art can make various changes or modifications to the embodiments within the scope of the embodiments. The embodiments include modifications / modifications, and such modifications / modifications are within the scope and spirit of the claims.

Claims

1. 1. A method for encoding point cloud data, comprising: encoding geometry data of the point cloud data; encoding characteristic data of the point cloud data based on the geometry data; transmitting the encoded geometry data, the encoded attribute data, and signaling data; The step of encoding the geometry data comprises: dividing the geometry data into blocks for motion compensation based on a partitioning method; and inter-predictively encoding the geometry data by selectively applying the motion compensation to each of the blocks; the signaling data includes information for identifying the division method and information for identifying a size of a block divided based on the division method; the signaling data further includes information repeated as many times as the number of blocks, the information indicating whether the motion compensation has been applied to the corresponding block; The method, wherein the signaling data further includes type information for specifying whether the geometry data is coded based on an octree or a predictive tree.

2. The method of claim 1 , wherein the motion compensation is performed based on a global motion vector obtained by estimating motion between successive frames.

3. The method of claim 1 , wherein the point cloud data is captured by a LiDAR including one or more lasers.

4. The method of claim 1 , wherein the signaling data further includes information for identifying the number of blocks.

5. 1. An apparatus for encoding point cloud data, comprising: a geometry encoder configured to encode geometry data of the point cloud data; a feature encoder configured to encode feature data of the point cloud data based on the geometry data; a transmitter configured to transmit the encoded geometry data, the encoded attribute data, and signaling data; The geometry encoder a partitioning unit configured to partition the geometry data into blocks for motion compensation based on a partitioning method; an inter prediction unit configured to inter predictively code the geometry data by selectively applying the motion compensation to each of the blocks; the signaling data includes information for identifying the division method and information for identifying a size of a block divided based on the division method; the signaling data further includes information repeated as many times as the number of blocks, the information indicating whether the motion compensation has been applied to the corresponding block; The apparatus, wherein the signaling data further includes type information for specifying whether the geometry data is coded based on an octree or a predictive tree.

6. The apparatus of claim 5 , wherein the motion compensation is performed based on a global motion vector obtained by estimating motion between successive frames.

7. The apparatus of claim 5 , wherein the point cloud data is captured by a LiDAR including one or more lasers.

8. The apparatus of claim 5 , wherein the signaling data further includes information for identifying the number of blocks.

9. 1. A method for decoding point cloud data, comprising: receiving geometry data, attribute data, and signaling data; decoding the geometry data based on the signaling data; decoding the characteristic data based on the signaling data and the decoded geometry data; The step of decoding the geometry data comprises: dividing reference data for the geometry data into blocks for motion compensation based on a division method; and inter-predictively decoding the geometry data by selectively applying the motion compensation to each of the blocks based on the signaling data; the signaling data includes information for identifying the division method and information for identifying a size of a block divided based on the division method; the signaling data further includes information repeated as many times as the number of blocks, the information indicating whether the motion compensation has been applied to the corresponding block; The method, wherein the signaling data further includes type information for specifying whether the geometry data is coded based on an octree or a predictive tree.

10. The method of claim 9 , wherein the motion compensation is performed based on motion vectors obtained by estimating motion between successive frames at the transmitting side.

11. The method of claim 9 , wherein the point cloud data including the geometry data and the characteristic data is captured by a LiDAR including one or more lasers at a transmitting end.

12. The method of claim 9 , wherein the signaling data includes information to identify the number of blocks.