Apparatus and method for processing point cloud data
By encoding the geometric and attribute information of point cloud data and combining it with point cloud compression coding technology, the delay and complexity issues in point cloud data processing are resolved, and efficient point cloud services are implemented to support VR, AR, MR and self-driving applications.
Patent Information
- Application Number
- CN202511063476.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-21
- Filing Date
- 2020-05-29
- Publication Date
- 2025-10-10
AI Technical Summary
Existing technologies have problems with latency and high encoding/decoding complexity when processing point cloud data, especially since point cloud content requires the representation of tens to hundreds of thousands of point data, resulting in low processing efficiency.
A coding method based on geometric information and attribute information is used to encode and decode point cloud data, including the point cloud video acquisition, encoding, sending, receiving and rendering processes. Point cloud compression coding technologies such as G-PCC and V-PCC are used, combined with feedback information to optimize data processing.
It achieves efficient processing of point cloud data, provides high-quality point cloud services, supports VR, AR, MR and self-driving services, and reduces latency and encoding and decoding complexity.
Smart Images

Figure CN120769062A_ABST
Abstract
Description
[0001] This application is a divisional application of the original invention patent application with application number 202080044643.X (International application number: PCT / KR2020 / 006959, application date: May 29, 2020, invention name: Device and method for processing point cloud data). Technical Field
[0002] The present disclosure provides a method for providing point cloud content to provide users with various services such as virtual reality (VR), augmented reality (AR), mixed reality (MR), and self-driving services. Background Art
[0003] Point cloud content is represented by a point cloud, which is a collection of points belonging to a coordinate system representing three-dimensional space. Point cloud content can express media configured in three dimensions and is used to provide various services such as virtual reality (VR), augmented reality (AR), mixed reality (MR), and self-driving services. However, tens to hundreds of thousands of points of data are required to represent point cloud content. Therefore, a method for efficiently processing large amounts of point data is needed. Summary of the Invention
[0004] Technical issues
[0005] Embodiments provide an apparatus and method for efficiently processing point cloud data. Embodiments provide a method and apparatus for processing point cloud data to address latency and encoding / decoding complexity.
[0006] The technical scope of the embodiments is not limited to the above technical objectives, and can be extended to other technical objectives that can be inferred by those skilled in the art based on the entire content disclosed herein.
[0007] Technical Solution
[0008] To achieve these objectives and other advantages and in accordance with the present disclosure, in some embodiments, a method for transmitting point cloud data may include the following steps: encoding point cloud data including geometric information and attribute information and transmitting a bitstream including the encoded point cloud data. In some embodiments, the geometric information represents the positions of points in the point cloud data, and the attribute information represents the attributes of the points in the point cloud data.
[0009] In some embodiments, a method for processing point cloud data may include the following steps: receiving a bitstream comprising point cloud data. In some embodiments, the point cloud data includes geometric information and attribute information, wherein the geometric information represents the positions of points in the point cloud data, and the attribute information represents one or more attributes of the points in the point cloud data. The point cloud data processing method may include the following steps: decoding the point cloud data.
[0010] In some embodiments, a method for processing point cloud data can include the steps of receiving a bitstream including point cloud data, and decoding the point cloud data. In some embodiments, the point cloud data includes geometry information and attribute information, wherein the geometry information represents positions of points of the point cloud data, and the attribute information indicates one or more attributes of the points of the point cloud data.
[0011] In some embodiments, an apparatus for processing point cloud data can include a receiver configured to receive a bitstream including point cloud data, and a decoder configured to decode the point cloud data. In some embodiments, the point cloud data includes geometry information and attribute information, wherein the geometry information represents positions of points of the point cloud data, and the attribute information indicates one or more attributes of the points of the point cloud data.
[0012] Advantages
[0013] The apparatus and method according to embodiments can efficiently process point cloud data.
[0014] The apparatus and method according to embodiments can provide high-quality point cloud services.
[0015] The apparatus and method according to embodiments can provide point cloud content for providing general services such as VR services and self-driving services. BRIEF DESCRIPTION OF DRAWINGS
[0016] The accompanying drawings, which are included to provide a further understanding of the disclosure and are incorporated in and constitute a part of this application, illustrate embodiments of the disclosure and together with the description serve to explain the principles of the disclosure.
[0017] For a better understanding of the various embodiments described below, reference should be made to the Drawings in conjunction with the following description. In the Drawings:
[0018] Figure 1 An exemplary point cloud content providing system according to an embodiment is illustrated.
[0019] Figure 2 is a block diagram illustrating a point cloud content providing operation according to an embodiment.
[0020] Figure 3 An exemplary process of capturing a point cloud video according to an embodiment is illustrated.
[0021] Figure 4 An exemplary point cloud encoder according to an embodiment is illustrated.
[0022] Figure 5 An example of a voxel according to an embodiment is illustrated.
[0023] Figure 6Shown are examples of octrees and occupancy codes according to an embodiment.
[0024] Figure 7 An example of a neighbor node pattern according to an embodiment is shown.
[0025] Figure 8 An example of point arrangement in each LOD according to an embodiment is shown.
[0026] Figure 9 An example of point arrangement in each LOD according to an embodiment is shown.
[0027] Figure 10 An exemplary point cloud decoder according to an embodiment is shown.
[0028] Figure 11 An exemplary point cloud decoder according to an embodiment is shown.
[0029] Figure 12 An exemplary transmitting device according to an embodiment is shown.
[0030] Figure 13 An exemplary receiving device according to an embodiment is shown.
[0031] Figure 14 An architecture for streaming G-PCC based point cloud data is shown according to an embodiment.
[0032] Figure 15 An exemplary point cloud transmitting device according to an embodiment is shown.
[0033] Figure 16 An exemplary point cloud receiving device according to an embodiment is shown.
[0034] Figure 17 An exemplary structure operatively connectable with a method / apparatus for transmitting and receiving point cloud data according to an embodiment is shown.
[0035] Figure 18 A scalable representation according to an embodiment is shown.
[0036] Figure 19 Point cloud data based on a shaded octree according to an embodiment is shown.
[0037] Figure 20 Shown is a shading octree according to an embodiment.
[0038] Figure 21 is an exemplary flow chart of attribute encoding according to an embodiment.
[0039] Figure 22 An octree structure according to an embodiment is shown.
[0040] Figure 23 A shaded octree structure according to an embodiment is shown.
[0041] Figure 24 A shaded octree structure according to an embodiment is shown.
[0042] Figure 25 An exemplary syntax of an APS according to an embodiment is shown.
[0043] Figure 26 An exemplary syntax of an attribute slice bitstream according to an embodiment is shown.
[0044] Figure 27 is a block diagram illustrating the encoding operation of a point cloud encoder.
[0045] Figure 28 is a block diagram illustrating the decoding operation of a point cloud decoder.
[0046] Figure 29 Details of geometry and attributes according to scalable decoding according to an embodiment are shown.
[0047] Figure 30 is an exemplary flow chart of a method for processing point cloud data according to an embodiment.
[0048] Figure 31 is an exemplary flow chart of a method for processing point cloud data according to an embodiment. DETAILED DESCRIPTION
[0049] Reference will now be made in detail to preferred embodiments of the present disclosure, examples of which are illustrated in the accompanying drawings. The detailed description given below with reference to the accompanying drawings is intended to illustrate exemplary embodiments of the present disclosure and is not intended to illustrate the only embodiments that can be implemented according to the present disclosure. The following detailed description includes specific details in order to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure can be practiced without these specific details.
[0050] Although most of the terms used in this disclosure are selected from common terms widely used in the art, some terms are arbitrarily selected by the applicant and their meanings are explained in detail in the following description as needed. Therefore, this disclosure should be understood based on the intended meaning of the terms rather than their simple names or meanings.
[0051] Figure 1 An exemplary point cloud content providing system according to an embodiment is shown.
[0052] Figure 1 The illustrated point cloud content providing system may include a transmitting device 10000 and a receiving device 10004. The transmitting device 10000 and the receiving device 10004 may be capable of wired or wireless communication to transmit and receive point cloud data.
[0053] According to an embodiment, the point cloud data transmitting device 10000 can obtain and process point cloud video (or point cloud content) and transmit it. According to an embodiment, the transmitting device 10000 may include a fixed station, a base transceiver system (BTS), a network, an artificial intelligence (AI) device and / or system, a robot, an AR / VR / XR device and / or a server. According to an embodiment, the transmitting device 10000 may include a device configured to communicate with a base station and / or other wireless devices using a radio access technology (e.g., 5G New RAT (NR), Long Term Evolution (LTE)), a robot, a vehicle, an AR / VR / XR device, a portable device, a home appliance, an Internet of Things (IoT) device, and an AI device / server.
[0054] According to an embodiment, the sending device 10000 includes a point cloud video acquirer 10001, a point cloud video encoder 10002 and / or a transmitter (or communication module) 10003.
[0055] The point cloud video acquirer 10001 according to an embodiment acquires a point cloud video through a process such as capture, synthesis, or generation. A point cloud video is point cloud content represented by a point cloud. A point cloud is a collection of points located in 3D space and can be referred to as point cloud video data. A point cloud video according to an embodiment may include one or more frames. A frame represents a still image / screen. Therefore, a point cloud video may include point cloud images / frames / screens and may be referred to as a point cloud image, frame, or screen.
[0056] The point cloud video encoder 10002 according to an embodiment encodes the acquired point cloud video data. The point cloud video encoder 10002 may encode the point cloud video data based on point cloud compression coding. The point cloud compression coding according to an embodiment may include geometry-based point cloud compression (G-PCC) coding and / or video-based point cloud compression (V-PCC) coding or next-generation coding. The point cloud compression coding according to an embodiment is not limited to the above-mentioned embodiment. The point cloud video encoder 10002 may output a bit stream containing the encoded point cloud video data. The bit stream may include not only the encoded point cloud video data but also signaling information related to the encoding of the point cloud video data.
[0057] According to an embodiment, the transmitter 10003 transmits a bitstream containing encoded point cloud video data. According to an embodiment, the bitstream is encapsulated in a file or segment (e.g., a stream segment) and transmitted via various networks such as a broadcast network and / or a broadband network. Although not shown in the figure, the transmitting device 10000 may include an encapsulator (or encapsulation module) configured to perform an encapsulation operation. According to an embodiment, the encapsulator may be included in the transmitter 10003. According to an embodiment, the file or segment may be transmitted to the receiving device 10004 via a network or stored in a digital storage medium (e.g., USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.). According to an embodiment, the transmitter 10003 is capable of wired / wireless communication with the receiving device 10004 (or receiver 10005) via a 4G, 5G, 6G, etc. network. In addition, the transmitter may perform necessary data processing operations according to the network system (e.g., a 4G, 5G, or 6G communication network system). The transmitting device 10000 may transmit the encapsulated data on demand.
[0058] The receiving device 10004 according to an embodiment includes a receiver 10005, a point cloud video decoder 10006, and / or a renderer 10007. According to an embodiment, the receiving device 10004 may include a device configured to communicate with a base station and / or other wireless devices using a radio access technology (e.g., 5G New RAT (NR), Long Term Evolution (LTE)), a robot, a vehicle, an AR / VR / XR device, a portable device, a home appliance, an Internet of Things (IoT) device, and an AI device / server.
[0059] The receiver 10005 according to the embodiment receives a bitstream containing point cloud video data or a file / segment encapsulated with a bitstream from a network or a storage medium. The receiver 10005 may perform necessary data processing according to a network system (e.g., a communication network system of 4G, 5G, 6G, etc.). The receiver 10005 according to the embodiment may decapsulate the received file / segment and output a bitstream. According to the embodiment, the receiver 10005 may include a decapsulator (or decapsulation module) configured to perform a decapsulation operation. The decapsulator may be implemented as an element (or component) separate from the receiver 10005.
[0060] The point cloud video decoder 10006 decodes the bitstream containing the point cloud video data. The point cloud video decoder 10006 can decode the point cloud video data according to the method by which the point cloud video data was encoded (e.g., by performing the reverse process of the operation of the point cloud video encoder 10002). Therefore, the point cloud video decoder 10006 can decode the point cloud video data by performing point cloud decompression encoding (the reverse process of point cloud compression). Point cloud decompression encoding includes G-PCC encoding.
[0061] The renderer 10007 renders the decoded point cloud video data. The renderer 10007 may output point cloud content by rendering not only the point cloud video data but also the audio data. Depending on the embodiment, the renderer 10007 may include a display configured to display the point cloud content. Depending on the embodiment, the display may be implemented as a separate device or component rather than being included in the renderer 10007.
[0062] The arrow indicated by the dotted line in the figure represents the transmission path of the feedback information obtained by the receiving device 10004. Feedback information is information reflecting the interactivity with the user consuming the point cloud content, and includes information about the user (e.g., head orientation information, viewport information, etc.). Specifically, when the point cloud content is content for a service that requires interaction with the user (e.g., a self-driving service, etc.), the feedback information can be provided to the content sender (e.g., the sending device 10000) and / or the service provider. Depending on the embodiment, the feedback information may be used in the receiving device 10004 and the sending device 10000, or may not be provided.
[0063] According to an embodiment, head orientation information refers to information regarding the user's head position, orientation, angle, movement, and the like. According to an embodiment, the receiving device 10004 may calculate viewport information based on the head orientation information. Viewport information may be information regarding the area of the point cloud video being viewed by the user. The viewpoint is the point through which the user views the point cloud video and may refer to the center point of the viewport area. That is, the viewport is the area centered on the viewpoint, and the size and shape of the area may be determined by the field of view (FOV). Therefore, in addition to head orientation information, the receiving device 10004 may also extract viewport information based on the vertical or horizontal FOV supported by the device. Furthermore, the receiving device 10004 may perform gaze analysis, etc., to examine how the user consumes the point cloud, the area within the point cloud video the user is gazing at, the duration of the gaze, and other factors. According to an embodiment, the receiving device 10004 may transmit feedback information including the gaze analysis results to the sending device 10000. According to an embodiment, the feedback information may be obtained during the rendering and / or display process. According to an embodiment, the feedback information may be acquired by one or more sensors included in the receiving device 10004. Depending on the implementation, the feedback information may be obtained by the renderer 10007 or a separate external element (or device, component, etc.). Figure 1The dotted line in represents the process of sending the feedback information obtained by the renderer 10007. The point cloud content providing system can process (encode / decode) the point cloud data based on the feedback information. Therefore, the point cloud video decoder 10006 can perform a decoding operation based on the feedback information. The receiving device 10004 can send the feedback information to the sending device 10000. The sending device 10000 (or the point cloud video encoder 10002) can perform an encoding operation based on the feedback information. Therefore, the point cloud content providing system can effectively process necessary data (for example, point cloud data corresponding to the user's head position) based on the feedback information instead of processing (encoding / decoding) the entire point cloud data, and provide the point cloud content to the user.
[0064] Depending on the embodiment, the transmitting device 10000 may be referred to as an encoder, a transmitting device, a transmitter, etc., and the receiving device 10004 may be referred to as a decoder, a receiving device, a receiver, etc.
[0065] According to the embodiment Figure 1 Point cloud data processed in a point cloud content providing system (through a series of processes such as acquisition, encoding, transmission, decoding, and rendering) may be referred to as point cloud content data or point cloud video data. Depending on the embodiment, point cloud content data may be used as a concept encompassing metadata or signaling information related to point cloud data.
[0066] Figure 1 The elements of the illustrated point cloud content providing system may be implemented by hardware, software, a processor, and / or a combination thereof.
[0067] Figure 2 is a block diagram illustrating a point cloud content providing operation according to an embodiment.
[0068] Figure 2 The block diagram shows Figure 1 The point cloud content providing system described in
[0045] As described above, the point cloud content providing system may process point cloud data based on point cloud compression coding (eg, G-PCC).
[0069] According to an embodiment, a point cloud content providing system (e.g., a point cloud sending device 10000 or a point cloud video acquirer 10001) can acquire a point cloud video (20000). The point cloud video is represented by a point cloud belonging to a coordinate system for expressing a 3D space. According to an embodiment, a point cloud video may include a Ply (Polygon file format or Stanford Triangle format) file. When a point cloud video has one or more frames, the acquired point cloud video may include one or more Ply files. The Ply file contains point cloud data such as point geometry and / or attributes. Geometry includes the position of a point. The position of each point can be represented by parameters (e.g., X, Y, and Z axis values) representing a three-dimensional coordinate system (e.g., a coordinate system consisting of X, Y, and Z axes). Attributes include attributes of the point (e.g., information about the texture, color (YCbCr or RGB), reflectivity r, transparency, etc. of each point). A point has one or more attributes. For example, a point may have a color attribute or two attributes, color and reflectivity. Depending on the embodiment, geometry may be referred to as position, geometry information, geometry data, etc., and attribute may be referred to as attribute, attribute information, attribute data, etc. The point cloud content providing system (e.g., the point cloud transmitting device 10000 or the point cloud video acquirer 10001) may obtain point cloud data from information related to the point cloud video acquisition process (e.g., depth information, color information, etc.).
[0070] According to an embodiment, a point cloud content providing system (e.g., a transmitting device 10000 or a point cloud video encoder 10002) may encode point cloud data (20001). The point cloud content providing system may encode point cloud data based on point cloud compression coding. As described above, point cloud data may include geometry and attributes of points. Therefore, the point cloud content providing system may perform geometry coding for encoding the geometry and output a geometry bitstream. The point cloud content providing system may perform attribute coding for encoding the attributes and output an attribute bitstream. According to an embodiment, the point cloud content providing system may perform attribute coding based on geometry coding. The geometry bitstream and the attribute bitstream according to the embodiment may be multiplexed and output as one bitstream. The bitstream according to the embodiment may also include signaling information related to geometry coding and attribute coding.
[0071] The point cloud content providing system according to the embodiment (eg, the transmitting device 10000 or the transmitter 10003) may transmit the encoded point cloud data (20002). Figure 1 As shown, the encoded point cloud data can be represented by a geometry bitstream and an attribute bitstream. In addition, the encoded point cloud data can be sent in the form of a bitstream together with signaling information related to the encoding of the point cloud data (e.g., signaling information related to geometry encoding and attribute encoding). The point cloud content providing system can encapsulate the bitstream carrying the encoded point cloud data and send it in the form of a file or fragment.
[0072] According to an embodiment, a point cloud content providing system (e.g., receiving device 10004 or receiver 10005) may receive a bitstream containing encoded point cloud data. In addition, the point cloud content providing system (e.g., receiving device 10004 or receiver 10005) may demultiplex the bitstream.
[0073] The point cloud content providing system (e.g., the receiving device 10004 or the point cloud video decoder 10005) can decode the encoded point cloud data (e.g., geometry bitstream, attribute bitstream) sent in the bitstream. The point cloud content providing system (e.g., the receiving device 10004 or the point cloud video decoder 10005) can decode the point cloud video data based on the signaling information related to the encoding of the point cloud video data contained in the bitstream. The point cloud content providing system (e.g., the receiving device 10004 or the point cloud video decoder 10005) can decode the geometry bitstream to reconstruct the position (geometry) of the point. The point cloud content providing system can reconstruct the attributes of the point by decoding the attribute bitstream based on the reconstructed geometry. The point cloud content providing system (e.g., the receiving device 10004 or the point cloud video decoder 10005) can reconstruct the point cloud video based on the position according to the reconstructed geometry and the decoded attributes.
[0074] According to an embodiment, a point cloud content providing system (e.g., receiving device 10004 or renderer 10007) can render the decoded point cloud data (20004). The point cloud content providing system (e.g., receiving device 10004 or renderer 10007) can use various rendering methods to render the geometry and attributes decoded by the decoding process. Points in the point cloud content can be rendered as vertices with a specific thickness, cubes with a specific minimum size centered at the corresponding vertex position, or circles centered at the corresponding vertex position. All or part of the rendered point cloud content is provided to the user via a display (e.g., a VR / AR display, a general display, etc.).
[0075] According to an embodiment, the point cloud content providing system (e.g., receiving device 10004) may obtain feedback information (20005). The point cloud content providing system may encode and / or decode the point cloud data based on the feedback information. Figure 1 The feedback information and operations described are the same, so their detailed description is omitted.
[0076] Figure 3 An exemplary process of capturing point cloud video according to an embodiment is shown.
[0077] Figure 3 Show reference Figures 1 to 2 An exemplary point cloud video capture process of a point cloud content providing system is described.
[0078] Point cloud content includes point cloud videos (images and / or videos) representing objects and / or environments located in various 3D spaces (e.g., 3D spaces representing real environments, 3D spaces representing virtual environments, etc.). Therefore, a point cloud content providing system according to an embodiment may use one or more cameras (e.g., an infrared camera capable of obtaining depth information, an RGB camera capable of extracting color information corresponding to the depth information, etc.), a projector (e.g., an infrared pattern projector that obtains depth information), LiDAR, etc. to capture point cloud videos. A point cloud content providing system according to an embodiment may extract a geometric shape composed of points in a 3D space from the depth information and extract attributes of each point from the color information to obtain point cloud data. Images and / or videos according to an embodiment may be captured based on at least one of an inward-facing technique and an outward-facing technique.
[0079] Figure 3 The left portion of FIG shows an inward-facing technique. Inward-facing techniques refer to techniques for capturing images of a central object using one or more cameras (or camera sensors) positioned around the central object. Inward-facing techniques can be used to generate point cloud content that provides a 360-degree image of a key object to the user (e.g., VR / AR content that provides a 360-degree image of an object (e.g., a key object such as a character, player, object, or actor) to the user).
[0080] Figure 3 The right side of the figure shows outward-facing techniques. Outward-facing techniques utilize one or more cameras (or camera sensors) positioned around a central object to capture the surroundings of the central object rather than an image of the central object. Outward-facing techniques can be used to generate point cloud content that provides the surrounding environment as it appears from the user's perspective (e.g., content representing the external environment that can be provided to a user of a self-driving vehicle).
[0081] As shown in the figure, point cloud content may be generated based on a capture operation of one or more cameras. In this case, the coordinate system may be different between the cameras, so the point cloud content providing system may calibrate one or more cameras to set a global coordinate system before the capture operation. In addition, the point cloud content providing system may generate point cloud content by synthesizing arbitrary images and / or videos with the images and / or videos captured by the above-mentioned capture technology. The point cloud content providing system may not perform a point cloud operation when generating point cloud content representing a virtual space. Figure 3 The point cloud content providing system according to an embodiment may perform post-processing on the captured image and / or video. In other words, the point cloud content providing system may remove unwanted areas (e.g., background), identify the space to which the captured image and / or video is connected, and, when there is a space hole, perform an operation to fill the space hole.
[0082] The point cloud content providing system generates a piece of point cloud content by performing coordinate transformation on points in point cloud videos acquired from various cameras. The point cloud content providing system performs coordinate transformation on the points based on the position coordinates of each camera. This allows the point cloud content providing system to generate content representing a wide range or point cloud content with a high point density.
[0083] Figure 4 An exemplary point cloud encoder according to an embodiment is shown.
[0084] Figure 4 Show Figure 1 An example of a point cloud video encoder 10002 is provided. The point cloud encoder reconstructs and encodes point cloud data (e.g., the position and / or attributes of a point) to adjust the quality of the point cloud content (e.g., lossless, lossy, or near-lossless) according to network conditions or applications. When the total size of the point cloud content is large (e.g., 60 Gbps of point cloud content for 30 fps), the point cloud content providing system may not be able to stream the content in real time. Therefore, the point cloud content providing system may reconstruct the point cloud content based on a maximum target bit rate to provide the point cloud content according to the network environment, etc.
[0085] As reference Figure 1 and Figure 2 As described above, the point cloud encoder can perform both geometry encoding and attribute encoding. Geometry encoding is performed before attribute encoding.
[0086] The point cloud encoder according to an embodiment includes a coordinate transformer (transforming coordinates) 40000, a quantizer (quantizing and removing points (voxelization)) 40001, an octree analyzer (analyzing octree) 40002 and a surface approximation analyzer (analyzing surface approximation) 40003, an arithmetic encoder (arithmetic coding) 40004, a geometry reconstructor (reconstructing geometry) 40005, a color transformer (transforming color) 40006, an attribute transformer (transforming attributes) 40007, a RAHT transformer (RAHT) 40008, an LOD generator (generating LOD) 40009, a lifting transformer (lifting) 40010, a coefficient quantizer (quantizing coefficients) 40011 and / or an arithmetic encoder (arithmetic coding) 40012.
[0087] The coordinate transformer 40000, quantizer 40001, octree analyzer 40002, surface approximation analyzer 40003, arithmetic encoder 40004, and geometry reconstructor 40005 may perform geometry coding. Geometric coding according to embodiments may include octree geometry coding, direct coding, triplet geometry coding, and entropy coding. Direct coding and triplet geometry coding may be applied selectively or in combination. Geometric coding is not limited to the above examples.
[0088] As shown in the figure, the coordinate converter 40000 according to an embodiment receives a position and converts it into coordinates. For example, the position can be converted into position information in a three-dimensional space (e.g., a three-dimensional space represented by an XYZ coordinate system). The position information in the three-dimensional space according to an embodiment can be referred to as geometric information.
[0089] The quantizer 40001 according to an embodiment quantizes the geometry. For example, the quantizer 40001 may quantize the points based on the minimum position value of all points (e.g., the minimum value on each of the X, Y, and Z axes). The quantizer 40001 performs a quantization operation by multiplying the difference between the minimum position value and the position value of each point by a preset quantization scale value, and then finding the nearest integer value by rounding the value obtained by the multiplication. Therefore, one or more points may have the same quantized position (or position value). The quantizer 40001 according to an embodiment performs voxelization based on the quantized position to reconstruct the quantized points. As in the case of a pixel (the smallest unit containing 2D image / video information), the points of the point cloud content (or 3D point cloud video) according to an embodiment may be included in one or more voxels. As a composite of volume and pixel, the term voxel refers to the 3D cubic space generated when the 3D space is divided into units (unit = 1.0) based on the axes representing the 3D space (e.g., the X, Y, and Z axes). Quantizer 40001 can match point groups in 3D space to voxels. Depending on the embodiment, a voxel may include only one point. Depending on the embodiment, a voxel may include one or more points. To represent a voxel as a point, the position of the voxel's center can be set based on the positions of one or more points included in the voxel. In this case, the attributes of all positions included in a voxel can be combined and assigned to the voxel.
[0090] The octree analyzer 40002 according to an embodiment performs octree geometry encoding (or octree encoding) to represent voxels in an octree structure. The octree structure represents points that are matched to voxels based on an octal tree structure.
[0091] The surface approximation analyzer 40003 according to an embodiment may analyze and approximate an octree. The octree analysis and approximation according to an embodiment is a process of analyzing a region including a plurality of points to efficiently provide an octree and voxelization.
[0092] According to an embodiment, the arithmetic encoder 40004 performs entropy coding on the octree and / or approximate octree. For example, the coding scheme includes arithmetic coding. As a result of the coding, a geometry bitstream is generated.
[0093] The color converter 40006, attribute converter 40007, RAHT converter 40008, LOD generator 40009, lifting converter 40010, coefficient quantizer 40011, and / or arithmetic encoder 40012 perform attribute coding. As described above, a point may have one or more attributes. Attribute coding according to an embodiment is also applied to the attributes of a point. However, when an attribute (e.g., color) includes one or more elements, attribute coding is applied independently to each element. Attribute coding according to an embodiment includes color transform coding, attribute transform coding, region adaptive hierarchical transform (RAHT) coding, interpolation-based hierarchical nearest neighbor prediction (prediction transform) coding, and interpolation-based hierarchical nearest neighbor prediction (lifting transform) coding with an update / lifting step. Depending on the point cloud content, the above-mentioned RAHT coding, prediction transform coding, and lifting transform coding can be selectively used, or a combination of one or more coding schemes can be used. Attribute coding according to an embodiment is not limited to the above examples.
[0094] Color converter 40006 according to an embodiment performs color conversion encoding to convert the color value (or texture) included in the attribute. For example, color converter 40006 may convert the format of color information (e.g., from RGB to YCbCr). Alternatively, the operation of color converter 40006 according to an embodiment may be applied based on the color value included in the attribute.
[0095] The geometry reconstructor 40005 according to an embodiment reconstructs (decompresses) an octree and / or an approximate octree. The geometry reconstructor 40005 reconstructs the octree / voxel based on the result of analyzing the point distribution. The reconstructed octree / voxel may be referred to as reconstructed geometry (restored geometry).
[0096] The attribute converter 40007 according to an embodiment performs attribute conversion to convert attributes based on reconstructed geometry and / or locations where geometry encoding is not performed. As described above, since attributes depend on geometry, the attribute converter 40007 can convert attributes based on reconstructed geometry information. For example, based on the position value of a point included in a voxel, the attribute converter 40007 can convert the attributes of the point at that position. As described above, when the center position of a voxel is set based on the positions of one or more points included in the voxel, the attribute converter 40007 converts the attributes of one or more points. When triplet geometry encoding is performed, the attribute converter 40007 can convert attributes based on the triplet geometry encoding.
[0097] The attribute converter 40007 can perform attribute conversion by calculating the average of the attributes or attribute values (e.g., the color or reflectivity of each point) of neighboring points within a specific position / radius from the center position (or position value) of each voxel. The attribute converter 40007 can apply weights based on the distance from the center to each point when calculating the average. Thus, each voxel has a position and a calculated attribute (or attribute value).
[0098] The attribute converter 40007 can search for neighbor points within a specific position / radius from the center position of each voxel based on a KD tree or a Morton code. The KD tree is a binary search tree and supports a data structure capable of managing points based on position, so that a nearest neighbor search (NNS) can be performed quickly. The Morton code is generated by presenting the coordinates (e.g., (x, y, z)) representing the 3D positions of all points as bit values and mixing the bits. For example, when the coordinates representing the point position are (5, 9, 1), the bit values of the coordinates are (0101, 1001, 0001). Mixing the bit values in the order of z, y, and x according to the bit index produces 010001000111. This value is represented as a decimal number 1095. That is, the Morton code value of the point with coordinates (5, 9, 1) is 1095. The attribute converter 40007 can sort the points based on the Morton code value and perform NNS through a depth-first traversal process. After the attribute transformation operation, KD tree or Morton code is used when NNS is needed in another transformation process for attribute encoding.
[0099] As shown, the transformed attributes are input to the RAHT transformer 40008 and / or the LOD generator 40009.
[0100] The RAHT transformer 40008 according to an embodiment performs RAHT encoding for predicting attribute information based on the reconstructed geometric information. For example, the RAHT transformer 40008 may predict attribute information of a higher-level node in the octree based on attribute information associated with a lower-level node in the octree.
[0101] According to an embodiment, LOD generator 40009 generates a level of detail (LOD) to perform predictive transform coding. According to an embodiment, LOD is the degree of detail of the point cloud content. As the LOD value decreases, the detail of the point cloud content degrades. As the LOD value increases, the detail of the point cloud content increases. Points can be categorized by LOD.
[0102] The lifting transformer 40010 according to an embodiment performs lifting transform coding that transforms point cloud attributes based on weights. As described above, lifting transform coding can be optionally applied.
[0103] The coefficient quantizer 40011 according to the embodiment quantizes the attribute of the attribute encoding based on the coefficient.
[0104] The arithmetic encoder 40012 according to the embodiment encodes quantized properties based on arithmetic coding.
[0105] Although not shown in the figure, Figure 4 The elements of the point cloud encoder may be implemented by hardware, software, firmware, or a combination thereof including one or more processors or integrated circuits configured to communicate with one or more memories included in the point cloud providing device. The one or more processors may execute the above Figure 4 At least one of the operations and / or functions of the elements of the point cloud encoder. In addition, one or more processors may be operable or executable to perform Figure 4 The one or more memories of the embodiment may include high-speed random access memory, or include non-volatile memory (e.g., one or more magnetic disk storage devices, flash memory devices, or other non-volatile solid-state memory devices).
[0106] Figure 5 An example of voxels according to an embodiment is shown.
[0107] Figure 5 The voxels are shown in a 3D space represented by a coordinate system consisting of three axes (X axis, Y axis and Z axis). Figure 4 As described above, the point cloud encoder (e.g., quantizer 40001) may perform voxelization. A voxel refers to a 3D cubic space generated by dividing a 3D space into units (unit=1.0) based on axes representing the 3D space (e.g., X-axis, Y-axis, and Z-axis). Figure 5 An example of voxels generated by an octree structure is shown, where a cubic axis-aligned bounding box defined by two poles (0,0,0) and (2d,2d,2d) is recursively subdivided. A voxel includes at least one point. The spatial coordinates of the voxel can be estimated from the positional relationship with the voxel group. As described above, the voxel has properties similar to the pixels of a 2D image / video (e.g., color or reflectivity). The details of the voxel are similar to those of the reference Figure 4 Those described are the same, so their description is omitted.
[0108] Figure 6 Shown are examples of octrees and occupancy codes according to an embodiment.
[0109] As reference Figures 1 to 4 As described, the point cloud content providing system (point cloud video encoder 10002) or the point cloud encoder (e.g., octree analyzer 40002) performs octree geometry encoding (or octree encoding) based on the octree structure to efficiently manage the area and / or position of voxels.
[0110] Figure 6 The upper part of FIG shows the octree structure. The 3D space of the point cloud content according to the embodiment is represented by the axes (eg, X-axis, Y-axis, and Z-axis) of the coordinate system. The octree structure is represented by the two poles (0,0,0) and (2 d ,2 d ,2 d ) to create an octree structure. Here, 2 d It can be set to the value of the minimum bounding box that surrounds all points of the point cloud content (or point cloud video). Here, d represents the depth of the octree. The value of d is determined in the following formula. In the following formula, (x int n ,y int n ,z int n ) represents the position (or position value) of the quantized point.
[0111]
[0112] like Figure 6 As shown in the middle of the upper part of , the entire 3D space can be divided into eight spaces according to the partition. Each divided space is represented by a cube with six faces. Figure 6 As shown in the upper right portion of the octree, each of the eight spaces is further divided based on the axes of the coordinate system (e.g., X-axis, Y-axis, and Z-axis). Thus, each space is divided into eight smaller spaces. Each of the divided smaller spaces is also represented by a cube with six faces. This partitioning scheme is applied until the leaf nodes of the octree become voxels.
[0113] Figure 6 The lower part shows the octree occupancy code. The occupancy code of the octree is generated to indicate whether each of the eight divided spaces generated by dividing one space contains at least one point. Therefore, a single occupancy code is represented by eight child nodes. Each child node represents the occupancy of the divided space, and the child node has a 1-bit value. Therefore, the occupancy code is represented as an 8-bit code. That is, when the space corresponding to the child node contains at least one point, the node is assigned a value of 1. When the space corresponding to the child node does not contain a point (the space is empty), the node is assigned a value of 0. Since Figure 6The occupancy code shown is 00100001, indicating that the spaces corresponding to the third and eighth child nodes among the eight child nodes each contain at least one point. As shown in the figure, each of the third and eighth child nodes has eight child nodes, and the child nodes are represented by 8-bit occupancy codes. The accompanying figure shows that the occupancy code of the third child node is 10000111, and the occupancy code of the eighth child node is 01001111. The point cloud encoder according to the embodiment (e.g., the arithmetic encoder 40004) may perform entropy coding on the occupancy code. In order to increase compression efficiency, the point cloud encoder may perform intra-frame / inter-frame coding on the occupancy code. The receiving device according to the embodiment (e.g., the receiving device 10004 or the point cloud video decoder 10006) reconstructs the octree based on the occupancy code.
[0114] According to an embodiment of the point cloud encoder (e.g., Figure 4 The point cloud encoder or octree analyzer 40002 may perform voxelization and octree encoding to store point locations. However, points are not always evenly distributed in 3D space, so there may be specific areas with fewer points. Therefore, performing voxelization on the entire 3D space is inefficient. For example, when a specific area contains very few points, voxelization does not need to be performed in that specific area.
[0115] Therefore, for the above-mentioned specific area (or nodes other than the leaf nodes of the octree), the point cloud encoder according to the embodiment may skip voxelization and perform direct encoding to directly encode the point positions included in the specific area. The coordinates of the directly encoded points according to the embodiment are called direct coding mode (DCM). The point cloud encoder according to the embodiment may also perform triplet geometry encoding based on the surface model, which is to reconstruct the point positions in the specific area (or node) based on voxels. Triplet geometry encoding is a geometry encoding that represents an object as a series of triangular meshes. Therefore, the point cloud decoder can generate a point cloud from the mesh surface. Direct encoding and triplet geometry encoding according to the embodiment can be performed selectively. In addition, direct encoding and triplet geometry encoding according to the embodiment can be performed in combination with octree geometry encoding (or octree encoding).
[0116] To perform direct encoding, the option to use direct mode to apply direct encoding should be enabled. The node to which direct encoding is to be applied is not a leaf node, and there should be less than a threshold number of points within the specific node. In addition, the total number of points to which direct encoding is to be applied should not exceed a preset threshold. When the above conditions are met, the point cloud encoder (or arithmetic encoder 40004) according to the embodiment can perform entropy encoding on the point position (or position value).
[0117] A point cloud encoder according to an embodiment (e.g., a surface approximation analyzer 40003) may determine a specific level of the octree (a level less than the depth d of the octree), and may use a surface model starting from this level to perform triplet geometry encoding to reconstruct the point positions in the node area based on voxels (triplet mode). A point cloud encoder according to an embodiment may specify the level to which triplet geometry encoding is to be applied. For example, when the specific level is equal to the depth of the octree, the point cloud encoder does not operate in triplet mode. In other words, the point cloud encoder according to an embodiment may operate in triplet mode only when the specified level is less than the depth value of the octree. A 3D cubic area of a node at a specified level according to an embodiment is called a block. A block may include one or more voxels. A block or voxel may correspond to a block. The geometry is represented as a surface within each block. A surface according to an embodiment may intersect each edge of the block at most once.
[0118] A block has 12 edges, so there are at least 12 intersections within a block. Each intersection is called a vertex. Vertices along an edge are detected when there is at least one occupied voxel adjacent to the edge across all blocks that share the edge. An occupied voxel, according to embodiments, refers to a voxel containing a point. The vertex position detected along an edge is the average position of all voxels adjacent to the edge across all blocks along the edge.
[0119] Once the vertex is detected, the point cloud encoder according to an embodiment may perform entropy encoding on the edge start point (x, y, z), the edge direction vector (Δx, Δy, Δz), and the vertex position value (relative position value within the edge). When triplet geometry encoding is applied, the point cloud encoder according to an embodiment (e.g., geometry reconstructor 40005) may generate recovered geometry (reconstructed geometry) by performing triangle reconstruction, upsampling, and voxelization processes.
[0120] Vertices located at the edges of a block define a surface that passes through the block. The surface according to an embodiment is a non-planar polygon. During triangle reconstruction, the surface represented by the triangles is reconstructed based on the starting points of the edges, the direction vectors of the edges, and the position values of the vertices. The triangle reconstruction process is performed by 1) calculating the centroid value of each vertex, 2) subtracting the centroid value from each vertex value, and 3) estimating the sum of the squares of the values obtained by the subtraction.
[0121]
[0122] The minimum value of the sum is estimated, and the projection process is performed along the axis with the minimum value. For example, when the element x is minimum, each vertex is projected onto the x-axis relative to the center of the block and onto the (y,z) plane. When the value obtained by projection onto the (y,z) plane is (ai,bi), the value of θ is estimated by atan2(bi,ai), and the vertices are sorted based on the value of θ. The following table shows the vertex combinations that create triangles based on the number of vertices. The vertices are sorted from 1 to n. The following table shows that for four vertices, two triangles can be constructed based on the vertex combinations. The first triangle can be composed of vertices 1, 2, and 3 among the sorted vertices, and the second triangle can be composed of vertices 3, 4, and 1 among the sorted vertices.
[0123] Table. Triangles formed from vertices sorted 1
[0124] [Table 1]
[0125] n triangle 3 (1,2,3) 4 (1,2,3),(3,4,1) 5 (1,2,3),(3,4,5),(5,1,3) 6 (1,2,3),(3,4,5),(5,6,1),(1,3,5) 7 (1,2,3),(3,4,5),(5,6,7),(7,1,3),(3,5,7) 8 (1,2,3),(3,4,5),(5,6,7),(7,8,1),(1,3,5),(5,7,1) 9 (1,2,3),(3,4,5),(5,6,7),(7,8,9),(9,1,3),(3,5,7),(7,9,3) 10 (1,2,3),(3,4,5),(5,6,7),(7,8,9),(9,10,1),(1,3,5),(5,7,9),(9,1,5) 11 (1,2,3),(3,4,5),(5,6,7),(7,8,9),(9,10,11),(11,1,3),(3,5,7),(7,9,11),(11,3,7) 12 (1,2,3),(3,4,5),(5,6,7),(7,8,9),(9,10,11),(11,12,1),(1,3,5),(5,7,9),(9,11,1),(1,5,9)
[0126] An upsampling process is performed to add points in the middle along the edges of the triangle, and voxelization is performed. The added points are generated based on the upsampling factor and the width of the block. The added points are called refinement vertices. The point cloud encoder according to an embodiment may voxelize the refinement vertices. In addition, the point cloud encoder may perform attribute encoding based on the voxelized positions (or position values).
[0127] Figure 7 An example of a neighbor node pattern according to an embodiment is shown.
[0128] In order to increase the compression efficiency of the point cloud video, the point cloud encoder according to an embodiment may perform entropy coding based on context-adaptive arithmetic coding.
[0129] As reference Figures 1 to 6 As described, the point cloud content providing system or point cloud encoder (eg, point cloud video encoder 10002, Figure 4 The point cloud encoder or arithmetic encoder 40004) may immediately perform entropy coding on the occupancy code. In addition, the point cloud content providing system or the point cloud encoder may perform entropy coding (intra-frame coding) based on the occupancy code of the current node and the occupancy of the neighboring nodes, or perform entropy coding (inter-frame coding) based on the occupancy code of the previous frame. The frame representation according to the embodiment is a collection of point cloud videos generated simultaneously. The compression efficiency of the intra-frame coding / inter-frame coding according to the embodiment may depend on the number of neighboring nodes referenced. When the number of bits increases, the operation becomes complicated, but the coding can be biased to one side, which can increase the compression efficiency. For example, when a 3-bit context is given, 2 bits need to be used. 3 = 8 ways to perform encoding. The division of parts for encoding affects the implementation complexity. Therefore, it is necessary to meet the appropriate level of compression efficiency and complexity.
[0130] Figure 7 The process of obtaining an occupancy pattern based on the occupancy of neighbor nodes is shown. According to an embodiment, a point cloud encoder determines the occupancy of neighbor nodes of each node of an octree and obtains the value of a neighbor pattern. The neighbor node pattern is used to infer the occupancy pattern of the node. Figure 7 The left side of the diagram shows the cube corresponding to the node (the cube in the middle) and the six cubes (neighboring nodes) that share at least one face with it. The nodes shown in the diagram are at the same depth. The numbers in the diagram represent the weights associated with the six nodes (1, 2, 4, 8, 16, and 32), respectively. Weights are assigned sequentially based on the positions of neighboring nodes.
[0131] Figure 7 The right part of shows the neighbor node pattern value. The neighbor node pattern value is the sum of the values multiplied by the weights of the occupied neighbor nodes (neighbor nodes with points). Therefore, the neighbor node pattern value is 0 to 63. When the neighbor node pattern value is 0, it indicates that there is no node with a point among the neighbor nodes of the node (no occupied node). When the neighbor node pattern value is 63, it indicates that all neighbor nodes are occupied nodes. As shown in the figure, since the neighbor nodes assigned with weights 1, 2, 4 and 8 are occupied nodes, the neighbor node pattern value is 15 (the sum of 1, 2, 4 and 8). The point cloud encoder can perform encoding according to the neighbor node pattern value (for example, when the neighbor node pattern value is 63, 64 types of encoding can be performed). According to an embodiment, the point cloud encoder can reduce the encoding complexity by changing the neighbor node pattern value (for example, based on a table that changes 64 to 10 or 6).
[0132] Figure 8 An example of point arrangement in each LOD according to an embodiment is shown.
[0133] As reference Figures 1 to 7 As described, the encoded geometry is reconstructed (decompressed) before attribute encoding is performed. When direct encoding is applied, the geometry reconstruction operation may include changing the placement of directly encoded points (e.g., placing directly encoded points in front of the point cloud data). When triplet geometry encoding is applied, the geometry reconstruction process is performed through triangle reconstruction, upsampling, and voxelization. Since attributes depend on the geometry, attribute encoding is performed based on the reconstructed geometry.
[0134] The point cloud encoder (e.g., LOD generator 40009) can classify (reorganize) points by LOD. The figure shows the point cloud content corresponding to the LOD. The leftmost frame in the figure shows the original point cloud content. The second frame from the left in the figure shows the point distribution in the lowest LOD, and the rightmost frame in the figure shows the point distribution in the highest LOD. In other words, the points in the lowest LOD are sparsely distributed, while the points in the highest LOD are densely distributed. In other words, as the LOD increases in the direction indicated by the arrow at the bottom of the figure, the space (or distance) between points becomes narrower.
[0135] Figure 9 An example of point configuration for each LOD according to an embodiment is shown.
[0136] As reference Figures 1 to 8 As described, the point cloud content providing system or point cloud encoder (eg, point cloud video encoder 10002, Figure 4 The point cloud encoder or LOD generator 40009 can generate LODs. LODs are generated by reorganizing points into a set of refinement levels based on a set LOD distance value (or Euclidean distance set). The LOD generation process is performed not only by the point cloud encoder but also by the point cloud decoder.
[0137] Figure 9 The upper part of shows examples of points (P0 to P9) of point cloud contents distributed in 3D space. Figure 9 In , the original order represents the order of points P0 to P9 before LOD generation. Figure 9 In the LOD-based order, the order of points generated according to the LOD is shown. Points are reorganized by LOD. In addition, higher LODs contain points belonging to lower LODs. Figure 9 As shown, LOD0 contains P0, P5, P4, and P2. LOD1 contains the points of LOD0, P1, P6, and P3. LOD2 contains the points of LOD0, LOD1, P9, P8, and P7.
[0138] As reference Figure 4 As described, the point cloud encoder according to the embodiment may selectively or in combination perform predictive transform coding, lifting transform coding, and RAHT transform coding.
[0139] The point cloud encoder according to an embodiment may generate a predictor for each point to perform predictive transform coding for setting a predicted attribute (or predicted attribute value) for each point. That is, N predictors may be generated for N points. The predictor according to an embodiment may calculate a weight (=1 / distance) based on the LOD value of each point, index information about neighboring points within a set distance of each LOD, and the distance to the neighboring point.
[0140] The predicted attribute (or attribute value) according to the embodiment is set to the average of the values obtained by multiplying the attribute (or attribute value) (e.g., color, reflectivity, etc.) of the neighboring points set in the predictor of each point by the weight (or weight value) calculated based on the distance to each neighboring point. The point cloud encoder (e.g., coefficient quantizer 40011) according to the embodiment can quantize and inverse quantize the residual (which may be referred to as residual attribute, residual attribute value, or attribute prediction residual) obtained by subtracting the predicted attribute (attribute value) from the attribute (attribute value) of each point. The quantization process is configured as shown in the following table.
[0141] Attribute prediction residual quantization pseudocode
[0142] [Table 2]
[0143]
[0144] Attribute prediction residual inverse quantization pseudo code
[0145] [Table 3]
[0146]
[0147] When the predictor of each point has neighboring points, the point cloud encoder (e.g., the arithmetic encoder 40012) according to an embodiment may perform entropy encoding on the quantized and inverse quantized residual values as described above. When the predictor of each point has no neighboring points, the point cloud encoder (e.g., the arithmetic encoder 40012) according to an embodiment may perform entropy encoding on the attributes of the corresponding point without performing the above operation.
[0148] The point cloud encoder (e.g., lifting transformer 40010) according to an embodiment may generate a predictor for each point, set the calculated LOD, register neighboring points in the predictor, and set weights based on the distances to the neighboring points to perform lifting transform encoding. Lifting transform encoding according to an embodiment is similar to the above-described predictive transform encoding, but differs in that weights are cumulatively applied to attribute values. The process of cumulatively applying weights to attribute values according to an embodiment is configured as follows.
[0149] 1) Create an array called quantized weights (QW) to store the weight values for each point. All elements of QW are initially set to 1.0. Multiply the QW value of the predictor index of the neighboring node registered in the predictor by the weight of the predictor for the current point, and add the resulting values.
[0150] 2) Boosting prediction process: The value obtained by multiplying the attribute value of the point by the weight is subtracted from the existing attribute value to calculate the predicted attribute value.
[0151] 3) Create temporary arrays called updateweight and update, and initialize the temporary arrays to zero.
[0152] 4) The weight calculated by multiplying the weight calculated for all predictors by the weight stored in the QW corresponding to the predictor index is added to the updateweight array as the index of the neighbor node. The value obtained by multiplying the attribute value of the neighbor node index by the calculated weight is added to the update array.
[0153] 5) Boosting update process: Divide the attribute values of the update array of all predictors by the weight value of the updateweight array of the predictor index, and add the existing attribute value to the value obtained by the division.
[0154] 6) For all predictors, the predicted attribute is calculated by multiplying the attribute value updated by the lifting update process by the weight (stored in QW) updated by the lifting prediction process. The point cloud encoder (e.g., coefficient quantizer 40011) according to the embodiment quantizes the predicted attribute value. In addition, the point cloud encoder (e.g., arithmetic encoder 40012) performs entropy encoding on the quantized attribute value.
[0155] A point cloud encoder according to an embodiment (e.g., RAHT transformer 40008) may perform RAHT transform coding, in which attributes associated with nodes at a lower level in an octree are used to predict attributes of nodes at a higher level. RAHT transform coding is an example of intra-coding of attributes by scanning backward through an octree. A point cloud encoder according to an embodiment scans the entire area starting from voxels and repeats a merging process of merging voxels into larger blocks at each step until a root node is reached. The merging process according to an embodiment is performed only on occupied nodes. The merging process is not performed on empty nodes. The merging process is performed on the upper node directly above the empty node.
[0156] The following formula represents the RAHT transformation matrix. In this formula, Represents the average attribute value of voxels at level l. Can be based on and To calculate. and The weight of and
[0157]
[0158] here, is the low-pass value and is used during the next highest level of merging. Denotes the high-pass coefficient. The high-pass coefficient of each step is quantized and subjected to entropy coding (eg, by arithmetic coder 400012). The weight is calculated as pass and Create the root node as follows.
[0159]
[0160] Similar to the high-pass coefficients, the values of gDC are also quantized and subjected to entropy coding.
[0161] Figure 10 A point cloud decoder according to an embodiment is shown.
[0162] Figure 10 The point cloud decoder shown is Figure 1 An example of a point cloud video decoder 10006 described in Figure 1 The operations of the point cloud video decoder 10006 shown in FIG. 10006 are the same as or similar to those of the point cloud video decoder 10006 shown in FIG. As shown in the figure, the point cloud decoder can receive a geometry bitstream and an attribute bitstream contained in one or more bitstreams. The point cloud decoder includes a geometry decoder and an attribute decoder. The geometry decoder performs geometry decoding on the geometry bitstream and outputs decoded geometry. The attribute decoder performs attribute decoding based on the decoded geometry and attribute bitstreams and outputs decoded attributes. The decoded geometry and decoded attributes are used to reconstruct the point cloud content (decoded point cloud).
[0163] Figure 11 A point cloud decoder according to an embodiment is shown.
[0164] Figure 11 The point cloud decoder shown is Figure 10 An example of a point cloud decoder is shown, and a decoding operation can be performed, which is Figures 1 to 9 The inverse process of the encoding operation of the point cloud encoder is shown.
[0165] As reference Figure 1 and Figure 10 As described, the point cloud decoder can perform both geometry decoding and attribute decoding. Geometry decoding is performed before attribute decoding.
[0166] According to an embodiment, the point cloud decoder includes an arithmetic decoder (arithmetic decoding) 11000, an octree synthesizer (synthesize octree) 11001, a surface approximation synthesizer (synthesize surface approximation) 11002 and a geometry reconstructor (reconstruct geometry) 11003, an inverse coordinate transformer (inverse transform coordinates) 11004, an arithmetic decoder (arithmetic decoding) 11005, an inverse quantizer (inverse quantization) 11006, a RAHT transformer 11007, an LOD generator (generate LOD) 11008, an inverse lifting (inverse lifting) 11009 and / or an inverse color transformer (inverse transform color) 11010.
[0167] The arithmetic decoder 11000, the octree synthesizer 11001, the surface approximation synthesizer 11002, the geometry reconstructor 11003, and the coordinate inverse transformer 11004 may perform geometry decoding. The geometry decoding according to the embodiment may include direct encoding and triplet geometry decoding. Direct encoding and triplet geometry decoding are selectively applied. The geometry decoding is not limited to the above example, and as a reference Figures 1 to 9 The inverse process of the geometric encoding described is performed.
[0168] The arithmetic decoder 11000 according to the embodiment decodes the received geometry bitstream based on arithmetic coding. The operation of the arithmetic decoder 11000 corresponds to the inverse process of the arithmetic encoder 40004.
[0169] The octree synthesizer 11001 according to the embodiment can generate an octree by obtaining an occupancy code (or information about the geometry obtained as a result of decoding) from the decoded geometry bitstream. Figures 1 to 9 Describe the configuration in detail.
[0170] When triplet geometry encoding is applied, the surface approximation synthesizer 11002 according to an embodiment may synthesize a surface based on the decoded geometry and / or the generated octree.
[0171] The geometry reconstructor 11003 according to an embodiment may regenerate the geometry based on the surface and / or decoded geometry. Figures 1 to 9 As described, direct coding and triplet geometry coding are selectively applied. Therefore, the geometry reconstructor 11003 directly imports the position information about the points to which direct coding is applied and adds them. When triplet geometry coding is applied, the geometry reconstructor 11003 can reconstruct the geometry by performing the reconstruction operations (e.g., triangle reconstruction, upsampling, and voxelization) of the geometry reconstructor 40005. Details and References Figure 6 The reconstructed geometry may include a point cloud frame or screen that does not contain attributes.
[0172] The coordinate inverse transformer 11004 according to an embodiment may acquire a point position by transforming the coordinates based on the reconstructed geometry.
[0173] The arithmetic decoder 11005, the inverse quantizer 11006, the RAHT transformer 11007, the LOD generator 11008, the inverse lifter 11009 and / or the inverse color transformer 11010 may perform a reference Figure 10Attribute decoding described. Attribute decoding according to an embodiment includes region adaptive hierarchical transform (RAHT) decoding, interpolation-based hierarchical nearest neighbor prediction (prediction transform) decoding, and interpolation-based hierarchical nearest neighbor prediction (lifting transform) decoding with an update / lifting step. The above three decoding schemes can be used selectively, or a combination of one or more decoding schemes can be used. Attribute decoding according to an embodiment is not limited to the above examples.
[0174] The arithmetic decoder 11005 according to the embodiment decodes the attribute bit stream through arithmetic coding.
[0175] The inverse quantizer 11006 according to an embodiment inversely quantizes information about a decoded attribute bitstream or an attribute obtained as a result of decoding, and outputs the inversely quantized attribute (or attribute value). Inverse quantization may be selectively applied based on attribute encoding of the point cloud encoder.
[0176] Depending on the embodiment, the RAHT transformer 11007, the LOD generator 11008, and / or the inverse lifter 11009 may process the reconstructed geometry and inverse quantized attributes. As described above, the RAHT transformer 11007, the LOD generator 11008, and / or the inverse lifter 11009 may selectively perform a decoding operation corresponding to the encoding of the point cloud encoder.
[0177] The color inverse converter 11010 according to an embodiment performs inverse transform encoding to inversely transform the color value (or texture) included in the decoded attribute. The operation of the color inverse converter 11010 may be selectively performed based on the operation of the color converter 40006 of the point cloud encoder.
[0178] Although not shown in the figure, Figure 11 The elements of the point cloud decoder may be implemented by hardware, software, firmware, or a combination thereof including one or more processors or integrated circuits configured to communicate with one or more memories included in the point cloud providing device. The one or more processors may execute the above Figure 11 At least one or more of the operations and / or functions of the elements of the point cloud decoder. In addition, one or more processors may be operable or executed to perform Figure 11 A software program and / or instruction set for the operation and / or functionality of the elements of a point cloud decoder.
[0179] Figure 12 An exemplary transmitting device according to an embodiment is shown.
[0180] Figure 12 The sending device shown is Figure 1 The sending device 10000 (or Figure 4 An example of a point cloud encoder. Figure 12The sending device shown can be executed with reference to Figures 1 to 9 The transmitting apparatus according to the embodiment may include a data input unit 12000, a quantization processor 12001, a voxelization processor 12002, an octree occupancy code generator 12003, a surface model processor 12004, an intra / inter coding processor 12005, an arithmetic encoder 12006, a metadata processor 12007, a color transform processor 12008, an attribute transform processor 12009, a prediction / lifting / RAHT transform processor 12010, an arithmetic encoder 12011 and / or a transmission processor 12012.
[0181] The data input unit 12000 according to the embodiment receives or acquires point cloud data. The data input unit 12000 may perform the same operation and / or acquisition method as the point cloud video acquirer 10001 (or refer to Figure 2 The acquisition process 20000) described above has the same or similar operations and / or acquisition methods.
[0182] The data input unit 12000, the quantization processor 12001, the voxelization processor 12002, the octree occupancy code generator 12003, the surface model processor 12004, the intra / inter coding processor 12005 and the arithmetic encoder 12006 perform geometric coding. Figures 1 to 9 The geometric encoding described is the same or similar, so its detailed description is omitted.
[0183] The quantization processor 12001 according to an embodiment quantizes geometry (e.g., position values of points). The operation and / or quantization of the quantization processor 12001 are related to the reference Figure 4 The operation and / or quantization of the quantizer 40001 described above are the same or similar. Figures 1 to 9 Same as those described.
[0184] The voxelization processor 12002 according to the embodiment voxelizes the quantized position value of the point. The voxelization processor 120002 may perform the same operation as the reference. Figure 4 The operation and / or voxelization process of the quantizer 40001 described above is the same or similar to the operation and / or process described above. Figures 1 to 9 Same as those described.
[0185] The octree occupancy code generator 12003 according to the embodiment performs octree encoding on the voxelized position of the point based on the octree structure. The octree occupancy code generator 12003 can generate an occupancy code. The octree occupancy code generator 12003 can perform the octree encoding with reference to Figure 4 and Figure 6The operations and / or methods of the point cloud encoder (or octree analyzer 40002) described herein are the same or similar to the operations and / or methods described herein. Figures 1 to 9 Same as those described.
[0186] According to an embodiment, the surface model processor 12004 may perform triplet geometry encoding based on the surface model to reconstruct the point position in a specific area (or node) based on voxels. Figure 4 The operations and / or methods of the point cloud encoder (e.g., surface approximation analyzer 40003) described herein are the same or similar to the operations and / or methods described herein. Figures 1 to 9 Same as those described.
[0187] The intra-frame / inter-frame encoding processor 12005 according to the embodiment may perform intra-frame / inter-frame encoding on the point cloud data. Figure 7 The same or similar encoding as described for intra / inter encoding. Details and references Figure 7 According to an embodiment, the intra / inter encoding processor 12005 may be included in the arithmetic encoder 12006.
[0188] According to an embodiment, arithmetic encoder 12006 performs entropy encoding on an octree and / or approximate octree of point cloud data. For example, the encoding scheme includes arithmetic coding. Arithmetic encoder 12006 performs operations and / or methods that are the same as or similar to those of arithmetic encoder 40004.
[0189] The metadata processor 12007 according to an embodiment processes metadata (e.g., set values) about the point cloud data and provides it to necessary processing processes such as geometry coding and / or attribute coding. In addition, the metadata processor 12007 according to an embodiment may generate and / or process signaling information related to geometry coding and / or attribute coding. The signaling information according to an embodiment may be encoded separately from the geometry coding and / or attribute coding. The signaling information according to an embodiment may be interleaved.
[0190] The color conversion processor 12008, the attribute conversion processor 12009, the prediction / lifting / RAHT conversion processor 12010, and the arithmetic encoder 12011 perform attribute coding. Figures 1 to 9 The attribute codes described are the same or similar, so their detailed description is omitted.
[0191] The color transform processor 12008 according to an embodiment performs color transform encoding to transform the color value included in the attribute. The color transform processor 12008 may perform color transform encoding based on the reconstructed geometry. The reconstructed geometry is compared with the reference Figures 1 to 9 In addition, its execution is the same as that of reference Figure 4 The operations and / or methods of the color converter 40006 are the same as or similar to those described above, and detailed description thereof is omitted.
[0192] The attribute transformation processor 12009 according to an embodiment performs attribute transformation to transform attributes based on the reconstructed geometry and / or the location where geometry encoding is not performed. Figure 4 The operations and / or methods of the attribute converter 40007 described above are the same as or similar to those of the attribute converter 40007. Detailed description thereof is omitted. The prediction / lifting / RAHT transform processor 12010 according to the embodiment may encode the transformed attributes by any one or a combination of RAHT coding, prediction transform coding, and lifting transform coding. The prediction / lifting / RAHT transform processor 12010 performs the same as the reference. Figure 4 The operations of the RAHT transformer 40008, the LOD generator 40009 and the lifting transformer 40010 described above are identical or similar to at least one operation. In addition, the prediction transform coding, the lifting transform coding and the RAHT transform coding are the same as those of the reference Figures 1 to 9 Those described are the same, so detailed descriptions thereof are omitted.
[0193] The arithmetic encoder 12011 according to the embodiment may encode the properties of the encoding based on arithmetic coding. The arithmetic encoder 12011 performs the same or similar operations and / or methods as those of the arithmetic encoder 400012.
[0194] The transmission processor 12012 according to the embodiment may send individual bit streams containing coded geometry and / or coded attributes and metadata information, or send a bit stream configured with coded geometry and / or coded attributes and metadata information. When the coded geometry and / or coded attributes and metadata information according to the embodiment are configured as a bit stream, the bit stream may include one or more sub-bit streams. The bit stream according to the embodiment may include signaling information and slice data, and the signaling information includes a sequence parameter set (SPS) for sequence level signaling, a geometry parameter set (GPS) for geometry information coding signaling, an attribute parameter set (APS) for attribute information coding signaling, and a patch parameter set (TPS) for patch level signaling. The slice data may include information about one or more slices. A slice according to the embodiment may include a geometry bit stream Geom0 0 and one or more attribute bitstreams Attr0 0 and Attr1 0The TPS according to the embodiment can include information on each tile among one or more tiles (e.g., coordinate information on a bounding box and height / size information). The geometry bitstream can include a header and a payload. The header of the geometry bitstream according to the embodiment can include a parameter set identifier (geom_parameter_set_id), a tile identifier (geom_tile_id), and a slice identifier (geom_slice_id) included in the GPS, and information on data included in the payload. As described above, the metadata processor 12007 according to the embodiment can generate and / or process signaling information and transmit the same to the transmission processor 12012. According to the embodiment, elements performing geometry encoding and elements performing attribute encoding can share data / information with each other as indicated by dotted lines. The transmission processor 12012 according to the embodiment can perform the same or similar operations and / or transmission methods as those of the transmitter 10003 and / or the transmission method. Details are the same as those described with reference to Figure 1 and Figure 2 are omitted.
[0195] Figure 13 An exemplary reception apparatus according to the embodiment is illustrated.
[0196] Figure 13 The illustrated reception apparatus is an example of the reception apparatus 10004 (or Figure 1 a point cloud decoder) of Figure 10 and Figure 11 . Figure 13 The illustrated reception apparatus can perform one or more operations and methods the same as or similar to those of the point cloud decoder described with reference to Figures 1 to 11 .
[0197] The reception apparatus according to the embodiment includes a receiver 13000, a reception processor 13001, an arithmetic decoder 13002, an occupancy code-based octree reconstruction processor 13003, a surface model processor (triangle reconstruction, upsampling, voxelization) 13004, an inverse quantization processor 13005, a metadata parser 13006, an arithmetic decoder 13007, an inverse quantization processor 13008, a prediction / lifting / RAHT inverse transform processor 13009, a color inverse transform processor 13010, and / or a renderer 13011. Each decoding element according to the embodiment can perform an inverse process of the operation of the corresponding encoding element according to the embodiment.
[0198] The receiver 13000 according to the embodiment receives point cloud data. The receiver 13000 can perform the same or similar operations and / or reception methods as those of the receiver 10005 of Figure 1 . A detailed description thereof is omitted.
[0199] According to an embodiment, the reception processor 13001 may obtain a geometry bitstream and / or an attribute bitstream from the received data. The reception processor 13001 may be included in the receiver 13000.
[0200] The arithmetic decoder 13002, the octtree reconstruction processor 13003 based on the occupancy code, the surface model processor 13004 and the inverse quantization processor 13005 may perform geometric decoding. Figures 1 to 10 The geometric decoding described is the same or similar, so its detailed description is omitted.
[0201] The arithmetic decoder 13002 according to an embodiment may decode a geometry bitstream based on arithmetic coding. The arithmetic decoder 13002 performs operations and / or encoding that are the same as or similar to those of the arithmetic decoder 11000.
[0202] The octree reconstruction processor 13003 based on the occupancy code according to the embodiment can reconstruct the octree by obtaining the occupancy code from the decoded geometry bitstream (or information about the geometry obtained as a result of decoding). The octree reconstruction processor 13003 based on the occupancy code performs the same or similar operations and / or methods as the operations of the octree synthesizer 11001 and / or the octree generation method. When triplet geometry coding is applied, the surface model processor 13004 according to the embodiment can perform triplet geometry decoding and related geometry reconstruction (e.g., triangle reconstruction, upsampling, voxelization) based on the surface model method. The surface model processor 13004 performs the same or similar operations as the operations of the surface approximation synthesizer 11002 and / or the geometry reconstructor 11003.
[0203] The inverse quantization processor 13005 according to an embodiment may inversely quantize the decoded geometry.
[0204] The metadata parser 13006 according to an embodiment may parse metadata (e.g., setting values) contained in the received point cloud data. The metadata parser 13006 may pass the metadata to the geometry decoding and / or attribute decoding. Figure 12 The metadata described are the same, so a detailed description thereof is omitted.
[0205] The arithmetic decoder 13007, the inverse quantization processor 13008, the prediction / lifting / RAHT inverse transform processor 13009 and the color inverse transform processor 13010 perform attribute decoding. Figures 1 to 10 The attribute decoding described is the same or similar, so the detailed description thereof is omitted.
[0206] According to an embodiment, the arithmetic decoder 13007 can decode the attribute bitstream through arithmetic coding. The arithmetic decoder 13007 can decode the attribute bitstream based on the reconstructed geometry. The arithmetic decoder 13007 performs the same or similar operations and / or encoding as those of the arithmetic decoder 11005.
[0207] The inverse quantization processor 13008 according to an embodiment may inversely quantize the decoded attribute bitstream. The inverse quantization processor 13008 performs the same or similar operation and / or inverse quantization method as the inverse quantizer 11006.
[0208] The prediction / lifting / RAHT inverse transform processor 13009 according to an embodiment may process the reconstructed geometry and inverse quantized attributes. The prediction / lifting / RAHT inverse transform processor 13009 performs one or more operations and / or decoding that are the same or similar to the operations and / or decoding of the RAHT converter 11007, the LOD generator 11008, and / or the inverse lifter 11009. The color inverse transform processor 13010 according to an embodiment performs inverse transform encoding to inversely transform the color values (or textures) included in the decoded attributes. The color inverse transform processor 13010 performs operations and / or inverse transform encoding that are the same or similar to the operations and / or inverse transform encoding of the color inverse converter 11010. The renderer 13011 according to an embodiment may render point cloud data.
[0209] Figure 14 An architecture for streaming G-PCC based point cloud data is shown according to an embodiment.
[0210] Figure 14 The upper part shows the Figures 1 to 13 The sending device described in (for example, the sending device 10000, Figure 12 The process of processing and sending point cloud content (such as a sending device).
[0211] As reference Figures 1 to 13 As described above, the transmitting device can obtain the audio Ba (audio acquisition) of the point cloud content, encode the obtained audio (audio encoding), and output the audio bit stream Ea. In addition, the transmitting device can obtain the point cloud (or point cloud video) Bv (point acquisition) of the point cloud content, and perform point cloud encoding on the obtained point cloud to output the point cloud video bit stream Eb. The point cloud encoding of the transmitting device is similar to the reference point. Figures 1 to 13 The point cloud encoding described (e.g., Figure 4 The encoding of the point cloud encoder is the same or similar, so its detailed description will be omitted.
[0212] The sending device may encapsulate the generated audio bitstream and video bitstream into files and / or fragments (file / fragment encapsulation). The encapsulated file and / or fragment Fs,File may include a file in a file format such as ISOBMFF or DASH fragment. Point cloud related metadata according to an embodiment may be included in the encapsulated file format and / or fragment. The metadata may be included in boxes at various levels on the ISOBMFF file format, or may be included in a separate track within the file. According to an embodiment, the sending device encapsulates the metadata into a separate file. The sending device according to an embodiment may transmit the encapsulated file format and / or fragment via a network. The encapsulation and transmission processing method of the sending device is the same as that of the reference Figures 1 to 13 Same as described (eg, transmitter 10003, Figure 2 transmission step 20002, etc.), so its detailed description will be omitted.
[0213] Figure 14 The lower part shows the reference Figures 1 to 13 The receiving device described (eg, receiving device 10004, Figure 13 The process of processing and outputting point cloud content (such as a receiving device).
[0214] According to an embodiment, the receiving device may include a device configured to output final audio data and final video data (e.g., a speaker, headphones, a display) and a point cloud player configured to process point cloud content (point cloud player). The final data output device and the point cloud player may be configured as separate physical devices. The point cloud player according to an embodiment may perform geometry-based point cloud compression (G-PCC) encoding, video-based point cloud compression (V-PCC) encoding, and / or next-generation encoding.
[0215] The receiving device according to the embodiment can obtain the file and / or fragment F', Fs' contained in the received data (for example, a broadcast signal, a signal transmitted via a network, etc.) and decapsulate it (file / fragment decapsulation). Figures 1 to 13 Those described (e.g., receiver 10005, receiver 13000, reception processor 13001, etc.) are the same, so their detailed description will be omitted.
[0216] A receiving device according to an embodiment obtains an audio bitstream E'a and a video bitstream E'v contained in a file and / or a segment. As shown in the figure, the receiving device performs audio decoding on the audio bitstream to output decoded audio data B'a, and then renders the decoded audio data (audio rendering) to output final audio data A'a through speakers or headphones.
[0217] In addition, the receiving apparatus performs point cloud decoding on the video bitstream E'v and outputs the decoded video data B'v. The point cloud decoding according to the embodiments is the same as or similar to the decoding of the point cloud decoder described with reference to Figures 1 to 13 the point cloud decoder described with reference to Figure 11 , and thus a detailed description thereof will be omitted. The receiving apparatus can render the decoded video data and output final video data through a display.
[0218] The receiving apparatus according to the embodiments can perform at least one of de-encapsulation, audio decoding, audio rendering, point cloud decoding, and point cloud video rendering based on the transmitted metadata. Details of the metadata are the same as those described with reference to Figures 12 to 13 , and thus a description thereof will be omitted.
[0219] As indicated by dotted lines shown in the figure, the receiving apparatus (e.g., a point cloud player or a sensing / tracking unit in the point cloud player) according to the embodiments can generate feedback information (orientation, viewport). According to the embodiments, the feedback information can be used in the de-encapsulation process, the point cloud decoding process, and / or the rendering process of the receiving apparatus, or can be transmitted to the transmitting apparatus. Details of the feedback information are the same as those described with reference to Figures 1 to 13 , and thus a description thereof will be omitted.
[0220] Figure 15 An exemplary transmitting apparatus according to the embodiments is shown.
[0221] Figure 15 The transmitting apparatus described with reference to Figures 1 to 14 is an apparatus configured to transmit point cloud content, and corresponds to an example of the transmitting apparatus described with reference to Figure 1 , the transmitting apparatus 10000 described with reference to Figure 4 , the point cloud encoder described with reference to Figure 12 , the transmitting apparatus described with reference to Figure 14 , and the transmitting apparatus. Thus, Figure 15 the transmitting apparatus performs the same or similar operations as those of the transmitting apparatus described with reference to Figures 1 to 14 .
[0222] The transmitting apparatus according to the embodiments can perform one or more of point cloud acquisition, point cloud encoding, file / segment encapsulation, and transmission.
[0223] Since the operations of the point cloud acquisition and transmission shown in the figure are the same as those described with reference to Figures 1 to 14 , a detailed description thereof will be omitted.
[0224] As described above with reference to Figures 1 to 14As described, the transmitting device according to the embodiment may perform geometry encoding and attribute encoding. Geometry encoding may be referred to as geometry compression, and attribute encoding may be referred to as attribute compression. As described above, a point may have a geometry and one or more attributes. Therefore, the transmitting device performs attribute encoding on each attribute. The figure shows that the transmitting device performs one or more attribute compressions (attribute #1 compression, ..., attribute #N compression). In addition, the transmitting device according to the embodiment may perform auxiliary compression. Auxiliary compression is performed on metadata. Details of metadata and reference Figures 1 to 14 The transmitting device may also perform mesh data compression. The mesh data compression according to the embodiment may include referring to Figures 1 to 14 Triplet geometric encoding of descriptions.
[0225] According to an embodiment, a transmitting device may encapsulate a bitstream (e.g., a point cloud stream) output from point cloud encoding into a file and / or fragment. According to an embodiment, the transmitting device may perform media track encapsulation to carry data other than metadata (e.g., media data), and perform metadata track encapsulation to carry metadata. According to an embodiment, metadata may be encapsulated into a media track.
[0226] As reference Figures 1 to 14 As described, the sending device may receive feedback information (orientation / viewport metadata) from the receiving device and perform at least one of point cloud encoding, file / segment packaging, and transmission operations based on the received feedback information. Figures 1 to 14 Those described are the same, so their description will be omitted.
[0227] Figure 16 An exemplary receiving device according to an embodiment is shown.
[0228] Figure 16 The receiving device is a device for receiving point cloud content, and corresponds to the reference Figures 1 to 14 Examples of receiving devices described (e.g., Figure 1 The receiving device 10004, Figure 11 Point cloud decoder and Figure 13 receiving device, Figure 14 Therefore, Figure 16 The receiving device performs the following operations with reference to Figures 1 to 14 The receiving device described above operates the same or similarly. Figure 16 The receiving device can receive Figure 15 The sending device sends a signal and executes Figure 15 The reverse process of the operation of the sending device.
[0229] The receiving device according to the embodiment may perform at least one of transmission, file / segment decapsulation, point cloud decoding, and point cloud rendering.
[0230] Since the point cloud reception and point cloud rendering operations shown in the figure are similar to the reference Figures 1 to 14 Those described are the same, so detailed descriptions thereof will be omitted.
[0231] As reference Figures 1 to 14 As described above, according to an embodiment, a receiving device decapsulates files and / or segments obtained from a network or storage device. According to an embodiment, the receiving device may perform media track decapsulation to carry data other than metadata (e.g., media data), and perform metadata track decapsulation to carry metadata. According to an embodiment, if metadata is encapsulated in a media track, metadata track decapsulation is omitted.
[0232] As reference Figures 1 to 14 As described, the receiving device may perform geometry decoding and attribute decoding on the bit stream (e.g., point cloud stream) obtained by decapsulation. Geometry decoding may be referred to as geometry decompression, and attribute decoding may be referred to as attribute decompression. As described above, a point may have a geometry and one or more attributes, each of which is encoded by the transmitting device. Therefore, the receiving device performs attribute decoding on each attribute. The figure shows that the receiving device performs one or more attribute decompressions (attribute #1 decompression, ..., attribute #N decompression). The receiving device according to the embodiment may also perform auxiliary decompression. Auxiliary decompression is performed on metadata. Details of metadata and reference Figures 1 to 14 The receiving device may also perform mesh data decompression. The mesh data decompression according to the embodiment may include referring to Figures 1 to 14 The receiving device according to the embodiment can render the point cloud data output according to the point cloud decoding.
[0233] As reference Figures 1 to 14 As described, the receiving device may use a separate sensing / tracking element to obtain the orientation / viewport metadata and send feedback information including the same to the sending device (e.g., Figure 15 In addition, the receiving device may perform at least one of a receiving operation, file / segment decapsulation, and point cloud decoding based on the feedback information. Figures 1 to 14 Those described are the same, so their description will be omitted.
[0234] Figure 17 An exemplary structure operatively connectable with a method / apparatus for transmitting and receiving point cloud data according to an embodiment is shown.
[0235] Figure 17The structure of 1700 represents a configuration in which at least one of the server 1760, the robot 1710, the self-driving vehicle 1720, the XR device 1730, the smartphone 1740, the home appliance 1750, and / or the HMD 1770 is connected to the cloud network 1700. The robot 1710, the self-driving vehicle 1720, the XR device 1730, the smartphone 1740, or the home appliance 1750 is referred to as a device. In addition, the XR device 1730 may correspond to a point cloud data (PCC) device according to an embodiment or may be operatively connected to a PCC device.
[0236] The cloud network 1700 may represent a network that constitutes a part of a cloud computing infrastructure or exists in a cloud computing infrastructure. Here, the cloud network 1700 may be configured using a 3G network, a 4G or Long Term Evolution (LTE) network, or a 5G network.
[0237] The server 1760 may be connected to at least one of the robot 1710 , the self-driving vehicle 1720 , the XR device 1730 , the smart phone 1740 , the home appliance 1750 , and / or the HMD 1770 via the cloud network 1700 , and may assist in at least a portion of the processing of the connected devices 1710 to 1770 .
[0238] HMD 1770 represents one of the implementation types of an XR device and / or a PCC device according to an embodiment. According to an embodiment, the HMD type device includes a communication unit, a control unit, a memory, an I / O unit, a sensor unit, and a power supply unit.
[0239] Hereinafter, various embodiments of devices 1710 to 1750 to which the above-described technology is applied will be described. Figure 17 The illustrated devices 1710 to 1750 may be operatively connected / coupled to the point cloud data transmitting / receiving device according to the above-described embodiment.
[0240] <PCC+XR>
[0241] The XR / PCC device 1730 may employ PCC technology and / or XR (AR+VR) technology, and may be implemented as an HMD, a head-up display (HUD) provided in a vehicle, a television, a mobile phone, a smart phone, a computer, a wearable device, a home appliance, a digital signage, a vehicle, a fixed robot, or a mobile robot.
[0242] The XR / PCC device 1730 can analyze 3D point cloud data or image data acquired through various sensors or from external devices and generate position data and attribute data regarding 3D points. Thereby, the XR / PCC device 1730 can acquire information regarding a surrounding space or a real object and render and output an XR object. For example, the XR / PCC device 1730 can match an XR object including auxiliary information regarding an identified object with the identified object and output the matched XR object.
[0243] <PCC+自驾驶+XR>
[0244] The self-driving vehicle 1720 can be implemented as a mobile robot, a vehicle, an unmanned aerial vehicle, etc. by applying PCC technology and XR technology.
[0245] The self-driving vehicle 1720 to which XR / PCC technology is applied can denote an autonomous vehicle provided with a means for providing an XR image, or an autonomous vehicle that is a control / interaction target in an XR image. Specifically, as a control / interaction target in an XR image, the self-driving vehicle 1720 can be distinguished from and operatively connected with the XR device 1730.
[0246] The self-driving vehicle 1720 having a means for providing an XR / PCC image can acquire sensor information from sensors including a camera and output a generated XR / PCC image based on the acquired sensor information. For example, the self-driving vehicle 1720 can have a HUD and output an XR / PCC image thereto to provide a passenger with an XR / PCC object corresponding to a real object or an object presented on a screen.
[0247] In this case, when an XR / PCC object is output to a HUD, at least a part of the XR / PCC object can be output to overlap with a real object pointed by a passenger's eyes. On the other hand, when an XR / PCC object is output on a display provided inside a self-driving vehicle, at least a part of the XR / PCC object can be output to overlap with an object on a screen. For example, the self-driving vehicle 1220 can output an XR / PCC object corresponding to an object such as a road, another vehicle, a traffic light, a traffic sign, a two-wheeled vehicle, a pedestrian, and a building.
[0248] Virtual reality (VR) technology, augmented reality (AR) technology, mixed reality (MR) technology, and / or point cloud compression (PCC) technology according to an embodiment are applicable to various devices.
[0249] In other words, VR technology is a display technology that only provides CG images of real-world objects, backgrounds, etc. On the other hand, AR technology refers to a technology that displays a virtually created CG image on an image of a real object. MR technology is similar to the above-mentioned AR technology in that the virtual objects to be displayed are mixed and combined with the real world. However, MR technology is different from AR technology in that AR technology clearly distinguishes between real objects and virtual objects created as CG images and uses virtual objects as supplementary objects to real objects, while MR technology treats virtual objects as objects with properties equivalent to real objects. More specifically, an example of the application of MR technology is holographic services.
[0250] Recently, VR, AR, and MR technologies are often referred to as extended reality (XR) technologies, rather than being clearly distinguished from each other. Therefore, embodiments of the present disclosure are applicable to any of VR, AR, MR, and XR technologies. Encoding / decoding based on PCC, V-PCC, and G-PCC technologies are applicable to such technologies.
[0251] The PCC method / apparatus according to the embodiment may be applied to a vehicle providing a self-driving service.
[0252] Vehicles providing self-driving services are connected to the PCC device for wired / wireless communication.
[0253] When the point cloud data (PCC) transmitting / receiving device according to the embodiment is connected to the vehicle for wired / wireless communication, the device can receive / process content data related to AR / VR / PCC services (which can be provided together with self-driving services) and send it to the vehicle. When the PCC transmitting / receiving device is installed on the vehicle, the PCC transmitting / receiving device can receive / process content data related to AR / VR / PCC services and provide it to the user based on a user input signal input through a user interface device. The vehicle or the user interface device according to the embodiment can receive the user input signal. The user input signal according to the embodiment may include a signal indicating a self-driving service.
[0254] According to an embodiment, scalable decoding is performed by a receiving device (e.g., Figure 1 The receiving device 10004, Figure 10 and Figure 11 Point cloud decoder or Figure 13 The decoding of some or all geometries and / or attributes is selectively performed by a receiving device) according to the decoding performance of the receiving device. Some geometries and attributes according to the embodiment are referred to as partial geometries and partial attributes. The scalable decoding applied to geometry according to the embodiment is referred to as scalable geometry decoding or geometric scalable decoding. The scalable decoding applied to attributes according to the embodiment is referred to as scalable attribute decoding or attribute scalable decoding. Figures 1 to 17As described, the points of point cloud content are distributed in three-dimensional space, and the distributed points are represented by an octree structure. An octree structure is an octree structure whose depth increases from higher nodes to lower nodes. Depending on the embodiment, the depth is referred to as a level and / or layer. Therefore, to provide low-resolution point cloud content, a receiving device may perform geometry decoding (or geometry scalable decoding) on a portion of the geometry and / or attribute decoding (or attribute scalable decoding) on a portion of the attributes, from higher nodes to lower nodes corresponding to a particular depth or level in the octree structure. Alternatively, the receiving device may perform geometry and attribute decoding corresponding to the entire octree structure to provide high-resolution point cloud content. A receiving device according to an embodiment performs scalable point cloud representation, or scalable representation, which is an operation that displays point cloud data corresponding to some of the one or more layers (or levels) of the point cloud data that comprise the geometry and attribute decoding. The scalable representation according to an embodiment may be selectively performed by the receiving device when the performance of a display and / or renderer is lower than that of a point cloud decoder. The levels of the scalable representation according to an embodiment correspond to the depth of the octree structure. As the value of the level according to an embodiment increases, the resolution or detail increases.
[0255] Scalable decoding supports receiving devices exhibiting various performances and enables point cloud services to be provided even in an adaptive bitrate environment. However, since attribute decoding is performed based on geometric decoding, geometric information is required to perform accurate attribute decoding. For example, the transform coefficients of RAHT encoding are determined based on geometric distribution information (or geometric structure information (e.g., octree structure)). In addition, predictive transform coding and lifting transform coding require the entire geometric distribution information (or geometric structure information (e.g., octree structure)) in order to obtain points belonging to each LOD.
[0256] Therefore, the receiving device can receive and process all geometry to perform robust attribute decoding. However, transmitting and receiving geometry information that is not actually displayed depending on the performance of the receiving device is inefficient in terms of bit rate. Furthermore, decoding all geometry by the receiving device may cause delays in providing point cloud content services. Furthermore, if the decoder of the receiving device has low performance, not all geometry may be decoded.
[0257] Figure 18 A scalable representation according to an embodiment is shown.
[0258] Figure 18 A point cloud decoder according to an embodiment is shown (e.g., referring to Figure 10 Point cloud video decoder 10006 described or reference Figure 1118. Example 1810 of a scalable representation of a point cloud decoder (described in conjunction with a point cloud decoder). The arrows 1800 shown in the figure indicate the direction of increasing depth of an octree structure of a geometry. The highest node of the octree structure according to an embodiment corresponds to the minimum depth or first depth and is referred to as the root. The lowest node of the octree structure according to an embodiment corresponds to the maximum depth or last depth and is referred to as a leaf. The depth of the octree structure according to an embodiment increases in the direction from the root to the leaves.
[0259] The point cloud decoder according to the embodiment performs decoding 1811 for providing full-resolution point cloud content or decoding 1812 for providing low-resolution point cloud content according to its performance. In order to provide full-resolution point cloud content, the point cloud decoder decodes the geometry bitstream 1811-1 and the attribute bitstream 1812-1 corresponding to the entire octree structure (1811). In order to provide low-resolution point cloud content, the point cloud decoder decodes the partial geometry bitstream 1812-1 and the partial attribute bitstream 1812-2 corresponding to a specific depth of the octree structure (1812). According to the embodiment, attribute decoding is performed based on geometry decoding. Therefore, even when the point cloud decoder is to decode the attribute corresponding to the partial attribute bitstream 1812-2, the point cloud decoder should decode the entire geometry bitstream 1811-1 (from the root depth to the leaf depth). In other words, in the figure, the shaded portion 1811-3 corresponds to geometric information that is not displayed but is sent and decoded to decode the attribute corresponding to the partial attribute bitstream 1812-1.
[0260] In addition, according to the transmitting device of the embodiment (for example, refer to Figure 1 The transmitting device 10000 described or referenced Figure 12 The sending device described) or point cloud encoder ( Figure 1 Point cloud video encoder 10002, Figure 4 Point cloud encoder in, reference Figure 12 、 Figure 14 and Figure 15 The point cloud encoder described above may only send a portion of the geometry bitstream 1812-1 and a portion of the attribute bitstream 1812-2 corresponding to a specific depth of the octree structure. A point cloud decoder providing low-resolution point cloud content decodes the portion of the geometry bitstream 1812-1 and the portion of the attribute bitstream 1812-2 corresponding to the specific depth of the octree structure (1812).
[0261] Therefore, if Figure 18 As shown, in order for a receiving device to provide point cloud data of various resolutions, that is, to provide scalable representation, a processing procedure and signaling information for processing a portion of the geometry bitstream 1811-1 and the attribute bitstream 1812-1 are required.
[0262] The point cloud encoder according to an embodiment generates a shaded octree by matching attributes to a geometric structure. The shaded octree according to an embodiment is generated by matching nodes and attributes at each level among one or more levels (or depths) of an octree structure representing the geometry.
[0263] The point cloud encoder according to an embodiment performs attribute encoding based on the generated shaded octree. Furthermore, the point cloud encoder generates scalable representation information including information related to the shaded octree to enable scalable decoding and scalable representation on a receiving device. This generated information is transmitted in the bitstream along with the encoded geometry and encoded attributes.
[0264] As a reverse process to the operations of the transmitting device or point cloud encoder, the receiving device may generate a shaded octree based on the scalable representation information. As described above, the shaded octree represents attributes that match the geometric octree structure. Therefore, the receiving device may select a specific level based on the shaded octree and output or render low-resolution point cloud content based on the matching attributes. Specifically, the receiving device may provide point cloud content of various resolutions according to the performance of the receiving device without requiring a separate receiving or processing process. According to embodiments, both the transmitting device (or point cloud encoder) and the receiving device (or point cloud decoder) may generate a shaded octree. The process or method of generating a shaded octree according to embodiments may be referred to as octree shading. According to embodiments, the point cloud encoder may perform octree shading on the entire octree structure, from the highest node (lowest level) to the lowest node (highest level). In addition, according to embodiments, the point cloud encoder may perform octree shading on any depth segment of the octree structure (e.g., the segment from level n-1 to level n). According to embodiments, the point cloud decoder may perform octree shading based on the above-mentioned scalable coding information.
[0265] Figure 19 Point cloud data based on a shaded octree according to an embodiment is shown.
[0266] As reference Figure 18 As described, the point cloud encoder according to the embodiment (e.g., Figure 1 Point cloud video encoder 10002, Figure 4 Point cloud encoder, reference Figure 12 、 Figure 14 and Figure 15 Point cloud encoder described in the description)) or transmitting device (for example, referring to Figure 1 The transmitting device 10000 described or referenced Figure 12 The transmitting device described herein may perform geometry encoding and attribute encoding on point cloud data (referred to as source geometry and source attributes or geometry and attributes) to enable scalable representation (1900). The encoded geometry and encoded attributes according to an embodiment are encoded in a bitstream (e.g., with reference to Figure 1 described in a bit stream).
[0267] Depending on the implementation, partial geometry and partial attributes at the same level may be sent (1910), or full geometry and partial attributes may be sent (1920). Alternatively, partial geometry and full attributes may be sent (1930), or full geometry and full attributes may be sent (1940). As described above, scalable representation information is sent in the bitstream along with the coded geometry and coded attributes.
[0268] Therefore, according to the receiving device of the embodiment (for example, Figure 1 The receiving device 10004, Figure 10 and Figure 11 Point cloud decoder or Figure 13 The receiving device receives the bitstream and ensures scalable representation information. The receiving device can provide point cloud data of various resolutions by processing all or part of the geometry and / or attributes in the bitstream based on the scalable representation information (scalable representation).
[0269] Figure 20 Shown is a shading octree according to an embodiment.
[0270] According to an embodiment of the point cloud encoder (e.g., Figure 1 Point cloud video encoder 10002, Figure 4 Point cloud encoder or reference Figure 12 、 Figure 14 and Figure 15 The point cloud encoder described in the embodiment (e.g., referring to FIG. 1 ) may perform attribute encoding for encoding attributes and output an attribute bitstream. The point cloud encoder performs attribute encoding based on a geometric structure. That is, the point cloud encoder according to the embodiment (e.g., referring to FIG. 1 ) may perform attribute encoding for encoding attributes and output an attribute bitstream. Figure 4 The attribute transformation unit 40007 described in the embodiment can transform the attributes based on the geometric structure (or reconstructed geometric structure). The geometric structure includes an octree structure. Since the octree structure according to the embodiment is a structure representing a three-dimensional space, it does not represent the attributes of each node. In addition, each node has a position (or position information) based on the position of the point included in the 3D space corresponding to each node. The position of the node according to the embodiment may be the average position of the positions of the points included in the area corresponding to the node, the position corresponding to a specific vertex, a preset representative position, the average value of the positions of the points included in the area corresponding to the node, etc. Therefore, the position of the node represents a position that is approximate to the actual position information about the point. Details of the octree structure according to the embodiment are similar to those of the reference numerals. Figures 1 to 19 Those described are the same, so their description will be omitted.
[0271] Figure 20The left side of FIG shows attributes c1, c2, c3, and c4 corresponding to four leaf nodes of the geometric octree structure according to an embodiment. As described above, the leaf node corresponds to the maximum depth or highest level (e.g., level n) of the octree structure. The point cloud encoder according to an embodiment matches each leaf node with an attribute.
[0272] Figure 20 The right portion of shows an exemplary colored octree generated based on the attributes corresponding to the leaf nodes of the octree structure. A point cloud encoder according to an embodiment may generate a colored octree structure representing attributes that match a lower depth or lower level (e.g., level n-1) (i.e., a higher node) of the octree structure based on the attributes corresponding to the leaf nodes (e.g., level n). The point cloud encoder according to an embodiment matches the position of the octree node with the attribute or the predicted value of the attribute, or matches the point cloud data with the octree node. The colored octree structure according to an embodiment represents the attributes that match the root node corresponding to the lowest level (e.g., level 0). As shown in reference Figure 18 and Figure 19 As described, the point cloud encoder according to the embodiment will generate signaling information related to the generation of the colored octree structure (for example, referring to Figure 18 The scalable representation information described in the signaling information) and the octree structure are sent together with the encoded geometry and attributes. Therefore, the receiving device or point cloud decoder performs scalable decoding according to the decoding performance based on the signaling information (for example, referring to Figure 18 and Figure 19 scalable decoding described in detail) up to a specific level (e.g., level 1 or the highest level n) to provide point cloud content at various resolutions.
[0273] Figure 21 is an exemplary flow chart of attribute encoding according to an embodiment.
[0274] The point cloud encoder according to the embodiment can generate a reference Figure 20 Describes the shader octree structure used to perform attribute encoding. Figure 21 The flowchart shown is to generate reference Figure 20 An example 2100 of a process of shading an octree structure (or octree shading) is described.
[0275] The point cloud encoder according to the embodiment makes the octree structure (for example, referring to Figures 1 to 20 The leaf nodes of the octree structure described in the embodiment are matched with the attributes (2110). The leaf nodes according to the embodiment correspond to the highest level of the octree structure. The number of leaf nodes is an integer greater than or equal to 1. The point cloud encoder according to the embodiment matches the attributes with the upper node (i.e., the lower level) of the leaf node based on the attributes matched with the leaf nodes. According to the embodiment, each node has a position (or position information). As shown in FIG. Figure 6As described, the nodes of the octree structure correspond to spaces that divide the 3D space according to the levels of the octree structure, and have positions set based on the positions of the points included in each space. The point cloud encoder can match attributes with each node (2120). When each node does not match the attribute, the point cloud encoder generates an estimated attribute based on the attribute of a higher level (e.g., the attribute of a leaf node) and matches it with the node of the corresponding level (e.g., the level lower than the level of the leaf node) to generate a colored octree structure (2130). The point cloud encoder can repeat the same process until the root node level.
[0276] When matching attributes to each node, the point cloud encoder may match the node's position or the position of the actual point cloud data with the attribute (2140). The point cloud encoder generates a colored octree structure by matching the position of each node with the attribute (2150). For example, the point cloud encoder may match the attribute of any one of one or more child nodes of each node with the node. The point cloud encoder may repeat the same process until the root node level.
[0277] The point cloud encoder generates a shaded octree structure (2160) by matching the actual point cloud data with the nodes. According to an embodiment, the leaf node corresponds to a voxel, which is the smallest segmentation unit. Therefore, the leaf node includes at least one point. Therefore, the position of the leaf node is consistent with the position of the corresponding point. However, in the octree structure, the position of the node other than the leaf node may not be completely consistent with the position of one or more points corresponding to the actual node in the 3D space. Therefore, for one or more child nodes of each node, the point cloud encoder may select the position and attributes of a pair of child nodes and match them with the node, or may select the average value or mean of the positions of the child nodes and match the attributes of the child node whose selected value and position are closest to the selected value with the node.
[0278] and generate Figure 21 The signaling information related to the process of coloring the octree structure is carried in the relevant signaling information to allow reference Figure 18 The receiving device described performs scalable representation. Therefore, the receiving device or point cloud decoder performs scalable decoding and scalable representation according to the decoding performance based on the signaling information (for example, referring to Figure 18 and Figure 19 The scalable representation described in the present invention) up to a specific level (e.g., level 1 or the highest level n) to provide point cloud content at various resolutions.
[0279] Figure 22 An octree structure according to an embodiment is shown.
[0280] As reference Figure 6As described above, the three-dimensional space of the point cloud content is represented by the axes of a coordinate system (e.g., X-axis, Y-axis, and Z-axis). An octree structure is generated by recursively subdividing the bounding box (or cube-axis-aligned bounding box). This subdivision method is applied until the leaf nodes of the octree become voxels. The levels increase from the highest node to the leaf node (the lowest node) in the octree structure, and each leaf node corresponds to a voxel.
[0281] Figure 22 The example of shows a three-level octree structure representing a point cloud. Each node in the octree structure has a position represented as a coordinate value in a three-dimensional coordinate system. As described above, each leaf node has the position of an actual point. Therefore, in the octree structure shown in the figure, the leaf node 2200 corresponding to level 3 (the highest level) has positions represented as (0, 2, 0), (1, 2, 0), (0, 3, 0), (1, 3, 0), (3, 2, 2), (2, 2, 1), (3, 2, 3), (2, 3, 3) and (3, 3, 3). That is, each leaf node may include points located at various positions. The two occupied nodes (00100001) 2210 corresponding to level 2 of the octree structure have positions represented as (0, 1, 0) and (1, 1, 1), respectively.
[0282] Figure 23 A shaded octree structure according to an embodiment is shown.
[0283] Figure 23 Examples 2300, 2310, and 2320 illustrate examples based on reference Figure 22 An example of a shaded octree structure generated by the three-level octree structure described.
[0284] Figure 23 Example 2300 shows an example of an octree structure (e.g., referring to Figures 1 to 20 The attributes of the leaf nodes of the octree structure described by . Figure 23 Example 2300 corresponds to reference Figure 21 The described point cloud encoder performs operations (e.g., 2110) to match leaf nodes with attributes.
[0285] As mentioned above, since leaf nodes correspond to voxels, the position of each leaf node is the same as the position of the actual point. That is, the positions of the leaf nodes are represented by coordinate values (0,2,0), (1,2,0), (0,3,0), (1,3,0), (3,2,2), (2,2,1), (3,2,3), (2,3,3), and (3,3,3). Each coordinate value is the position of a point in three-dimensional space. Each leaf node matches the attributes of the point at the corresponding position. In the figure, c1, c2, c3, c4, c5, c6, c7, c8, and c9 represent the attributes of each point. Figure 23An example 2300 of a point (e.g., a point located at (0, 2, 0)) having an attribute (e.g., c1) is shown. The relationship between an attribute and a point according to embodiments is not limited to this example. An attribute value of a location (x, y, z) is denoted as Attr(x, y, z). Thus, an attribute value matching an occupancy leaf node among 16 leaf nodes is expressed as follows.
[0286] c1 = Attr(0, 2, 0), c2 = Attr(1, 2, 0), c3 = Attr(0, 3, 0), c4 = Attr(1, 3, 0),
[0287] c5 = Attr(3, 2, 2), c6 = Attr(2, 2, 1), c7 = Attr(3, 2, 3), c8 = Attr(2, 3, 3), c9 = Attr(3, 3, 3)
[0288] An area corresponding to an upper node (parent node) of an octree leaf node is 8 times larger (2 times larger in each of width, length, and height) than an area corresponding to the leaf node. A size of an area corresponding to an upper node in an octree structure according to embodiments is expressed as 8^(an octree level different from a leaf node).
[0289] An area corresponding to an upper node according to embodiments includes one or more points. A location of an upper node according to embodiments can be set or denoted as an average value or average location of locations of child nodes of the node. Thus, a location of an upper node can not match a location of a point in an area corresponding to the upper node. In other words, since there is no actual point corresponding to a location of an upper node, a point cloud encoder cannot match an attribute corresponding to the location of the upper node. Thus, a point cloud encoder according to embodiments can match any attribute to a node. A point cloud encoder according to embodiments can detect a neighbor node and match it to any attribute. A neighbor node according to embodiments corresponds to an occupancy node among child nodes of a node. Figure 1 A point cloud video encoder 10002 of FIG. 1, Figure 4 A point cloud encoder of FIG. 1 or referring to Figure 12 , Figure 14 and Figure 15 A point cloud encoder described above can match any attribute to a node. A point cloud encoder according to embodiments can detect a neighbor node and match it to any attribute. A neighbor node according to embodiments corresponds to an occupancy node among child nodes of a node.
[0290] A point cloud encoder according to embodiments matches any attribute to an upper node (e.g., a node corresponding to level n-1) of a leaf node (e.g., level n) based on an attribute of the leaf node. The point cloud encoder generates a colored octree by repeating the same process until reaching a highest node. A colored octree according to embodiments is referred to as an attribute-paired octree.
[0291] Figure 23The upper right portion of example 2310 is a colored octree indicating the estimated attribute matching the upper node. As described above, the point cloud encoder according to the embodiment defines the position of the octree node and matches it with the estimated attribute to allow the receiving device to perform operations such as scalable coding, point cloud subsampling, and providing low-resolution point cloud content.
[0292] In the octree according to the embodiment, the position of the node can be expressed as a Morton code position corresponding to each node or a position in a three-dimensional coordinate system. The point cloud encoder can use estimated attributes to express the attributes of the upper node higher than the leaf node, and the estimated attributes are values that can represent the attributes of the child nodes of the node (for example, weighted average, mean, etc.). That is, the position and attributes of the node according to the embodiment may not be the same as the position and attributes of the actual point cloud data included in the area corresponding to each node. However, the receiving device can provide various versions of point cloud content (for example, low-resolution point cloud content or approximate point cloud content) according to the decoding performance or network environment based on the above-mentioned colored octree structure.
[0293] In the figure, p0 represents the estimated attribute matching the highest node (level 1), and p1 and p2 represent the estimated attributes matching the upper node. The point cloud encoder according to an embodiment can calculate p12315-1 based on the attributes c1, c2, c3 and c4 of the child node 2315.
[0294] The following equation represents the estimated properties (p(x,y,z)) of the upper node with position value (x,y,z).
[0295] [Formula 1]
[0296]
[0297] In this formula, Attr(xn,yn,yn) represents the attributes of the node's children. W represents the weight of the neighbor node. i, j, and k are parameters used to define the position of the neighbor node relative to the previous node's position (x, y, z). N represents the number of neighbor nodes.
[0298] Figure 23An example 2320 in the lower right of FIG. 23 indicates a colored octree indicating attributes matching the upper node. The point cloud encoder according to the embodiment can select one of the attributes of the child nodes of the node (e.g., the attribute of the first child node among the child nodes sorted in ascending order) in order to define the attribute of the parent node of the leaf node. Since the upper node according to the embodiment has the node position and the actual attribute, the receiving apparatus can provide a more approximate point cloud content (low resolution point cloud content) even when performing scalable decoding. In the figure, p0 indicates an estimated attribute matching the highest node (level 1), and p1 and p2 indicate attributes matching the upper node. As shown in the figure, p1 is equal to c4 among the attributes of the child nodes, and p2 is equal to c6 among the attributes of the child nodes. The attribute p0 matching the top node is equal to c4 among the lower attributes.
[0299] The following equation indicates an estimated attribute p(x, y, z) matching the upper node having a position value (x, y, z).
[0300] [Equation 2]
[0301]
[0302] In the above equation, Cp indicates an attribute (e.g., a color value) matching an occupied node among the child nodes. Attr(xn, yn, yn) indicates an attribute of a neighbor node around the node. M indicates a sum of weights of the neighbor nodes. That is, the equation is used to determine an attribute Cp minimizing a difference from the neighbor nodes as an estimated attribute (i.e., a prediction value) matching the upper node.
[0303] Figure 24 A colored octree structure according to the embodiment is illustrated.
[0304] Figure 24 Examples 2400, 2410, and 2420 shown are colored octree structures generated based on a three-level octree structure of leaf nodes having attributes matching the reference Figure 22 An example of a colored octree structure generated from the example 2300 of a three-level octree structure of leaf nodes having attributes matching the reference Figure 22 and Figure 23 As described above, each node (excluding the leaf nodes) in the octree structure has a position. The position of each node can not be the same as the position of the point included in the region corresponding to the node. The point cloud encoder according to the embodiment can generate a colored octree structure by matching not only the attributes of the reference Figure 23 described above, but also the positions of the nodes of the octree structure. That is, the actual points match the nodes of the colored octree according to the embodiment. The colored octree according to the embodiment is referred to as a point-paired octree. Therefore, the receiving apparatus can provide a point cloud content close to the original based on the colored octree even when performing scalable decoding.
[0305] In order to match the actual point with the upper node (excluding the leaf node) in the octree structure, the point cloud encoder ( Figure 1 Point cloud video encoder 10002, Figure 4 Point cloud encoder, reference Figure 12 、 Figure 14 and Figure 15 The point cloud encoder described in the foregoing (such as the point cloud encoder described in the foregoing) may select point cloud data close to the position of the node. For example, as the position of the node, the point cloud encoder may select the position of the node that has the smallest distance to the neighboring nodes. In addition, the point cloud encoder may match the attribute corresponding to the geometric centroid with the node. The geometric centroid according to the embodiment corresponds to the point at which the distance to all neighboring nodes (e.g., nodes at the same level and / or child nodes of the node) around the octree node is minimized. The formula given below represents the attribute of the geometric centroid matched to the node with the position (x, y, z).
[0306] [Formula 3]
[0307]
[0308] (x_p, y_p, z_p) represents the position of the node with the minimum distance to the neighbor node. (x_n, y_n, z_n) represents the neighbor node. The neighbor node according to the embodiment includes a node at the same level as the node and / or child node. (x_p, y_p, z_p) and (x_n, y_n, z_n) are included in the set of peripheral nodes (or neighbor nodes (NEIGHB(x,y,z)) around the node with value (x,y,z). Attr(xn,yn,yn) represents the attributes of the neighbor nodes around the node. M represents the sum of the weights of the peripheral nodes (or neighbor nodes). That is, the above formula is used to determine the attribute of the geometric centroid that minimizes the distance difference with the peripheral node as the estimated attribute (i.e., the predicted value) matched with the upper node.
[0309] Figure 24 The example 2400 shown at the top of FIG is a colored octree in which the position and attributes of any child node match those of the parent node. The point cloud encoder according to an embodiment selects the position corresponding to the geometric centroid as the position of the parent node (e.g., the parent node of a leaf node, etc.). For the position corresponding to the geometric centroid according to an embodiment, the point cloud data of a specific node among the child nodes of the node arranged in a fixed order (e.g., the first node among the child nodes arranged in ascending order) may be selected.
[0310] As shown in example 2400, the leaf nodes 2401 have attributes c1, c2, c3, and c4, respectively. The point cloud encoder according to the embodiment matches the position and the attribute c1 of the first child node to the parent node 2402 of the leaf nodes 2401. The leaf nodes 2403 have attributes c5, c6, c7, c8, and c9, respectively. Thus, the point cloud encoder matches the position and the attribute c5 of the first child node to the parent node 2404 of the leaf nodes 2403.
[0311] In the same way, the point cloud encoder matches the highest node 2405 to the position and the attribute c1 of the first node 2402 among the child nodes 2402 and 2404 of the node.
[0312] Thus, the points of the child nodes matched with the parent node are expressed by the following equation.
[0313] [Equation 4]
[0314] [x p ,y p ,z p ] T =[x k ,y k ,z k ] T
[0315] where [x k , y k , z k ] T represents the kth point in NEIGHBOR in ascending order.
[0316] In this equation, (x_p, y_p, z_p) is the position for obtaining the above-mentioned geometric centroid, and (x_k, y_k, z_k) represents a point matched with a child node corresponding to a fixed order position k.
[0317] The point cloud encoder according to the embodiment calculates an average position of positions of occupied nodes among child nodes, and matches point data (or points) of a child node whose position is closest to the average position to an upper node to generate a colored octree.
[0318] The equation given below expresses a method of calculating an average position of positions of occupied nodes among child nodes.
[0319] [Equation 5]
[0320]
[0321] In this formula, (x_p, y_p, z_p) represents the position of the parent node, and weight represents the weight value, which can indicate the occupancy of the child node (for example, the weight is set to 1 for an occupied node and 0 for an unoccupied node), or can indicate the weight in a specific direction. NEIGHBOR represents the neighbor node, that is, the child node of the parent node. M represents the sum of the weights.
[0322] Figure 24 The illustrated example 2410 is a colored octree obtained by calculating the average position of the positions of the occupied nodes among the child nodes and matching the position and attributes of the child node whose position is closest to the average position with the node above it.
[0323] As shown in example 2410, leaf nodes 2411 have attributes c1, c2, c3, and c4, respectively. The point cloud encoder according to an embodiment calculates the mean of the positions of the leaf nodes 2411. The positions and attributes of the leaf nodes 2411 according to an embodiment are represented as follows.
[0324] c1=Attr(0,2,0), c2=Attr(1,2,0), c3=Attr(0,3,0), c4=Attr(1,3,0)
[0325] The mean of the positions of the leaf nodes 2411 according to the embodiment is expressed as follows.
[0326] Mean=(0.4,2,0)
[0327] Therefore, since the position closest to the average value is c1, the point cloud encoder according to an embodiment matches the position and attributes of c1 with the parent node 2412 of the leaf node 2411 (e.g., the node of level n-1).
[0328] The leaf nodes 2413 have attributes c5, c6, c7, c8, and c9, respectively. The point cloud encoder according to the embodiment calculates the mean of the positions of the leaf nodes 2413. The positions and attributes of the leaf nodes 2413 according to the embodiment are represented as follows.
[0329] c5=Attr(3,2,2), c6=Attr(2,2,1), c7=Attr(3,2,3), c8=Attr(2,3,3), c9=Attr(3,3,3)
[0330] The mean of the positions of the leaf nodes 2413 according to the embodiment is expressed as follows.
[0331] Mean=(2.6,2.4,2.4)
[0332] The position closest to the mean is c5, so the point cloud encoder according to an embodiment matches the position and attributes of c5 with the parent node 2414 of the leaf node 2413 (e.g., the node at level n-1).
[0333] In the same manner, the point cloud encoder calculates the average position of the positions of nodes 2412 and 2414 (eg, nodes of level n-1). The positions and attributes of nodes 2412 and 2414 according to an embodiment are given as follows.
[0334] c1=Attr(0,2,0), c5=Attr(3,2,2)
[0335] The average position of the positions of the node 2412 and the node 2414 according to the embodiment is expressed as follows.
[0336] Mean=(1.5,2,1)
[0337] The position closest to the average is c1, so the point cloud encoder according to an embodiment matches the position and attributes of c1 with the parent node 2415 (e.g., the highest level node) of the nodes 2412 and 2414.
[0338] The point cloud encoder according to an embodiment calculates the median position of the occupied nodes among the child nodes, and matches the position and attributes of the child node closest to the median position with the parent node to generate a colored octree. The median position according to an embodiment may represent the mean value in Morton code order. In addition, when the number of occupied nodes is an even number, the point cloud encoder according to an embodiment may calculate the median position.
[0339] The formula given below represents a method of calculating the median position of the positions of occupied nodes among child nodes.
[0340] [Formula 6]
[0341] [x p ,y p ,z p ] T =MEDIAN NEIGHB(x,y,z) {weight(x n ,y n , z n |x,y,z)·[x n ,y n , z n ] T}
[0342] In this equation, (x_p, y_p, z_p) denotes the position of the parent node, weight denotes a weight value, can indicate the occupancy of a child node (for example, weight is set to 1 for an occupied node and 0 for an unoccupied node), or can indicate the weight of a specific direction. NEIGHBOR denotes a neighbor node, i.e., a child node of the parent node.
[0343] Figure 24 The example 2420 shown is a colored octree obtained by calculating the median position of the positions of the occupied nodes among the child nodes and matching the position and attributes of the child node whose position is closest to the median position with the upper node.
[0344] As shown in the example 2420, the leaf nodes 2421 have attributes c1, c2, c3, and c4, respectively. The point cloud encoder according to the embodiment calculates the median position of the positions of the leaf nodes 2421. The position closest to the median position is node c2, and thus the point cloud encoder according to the embodiment matches the position and attributes of c2 with the parent node 2422 (for example, a node of level n-1) of the leaf nodes 2421.
[0345] The leaf nodes 2423 have attributes c5, c6, c7, c8, and c9, respectively. The point cloud encoder according to the embodiment calculates the median position of the positions of the leaf nodes 2423. The position closest to the median position is node c7, and thus the point cloud encoder according to the embodiment matches the position and attributes of c7 with the parent node 2424 (for example, a node of level n-1) of the leaf nodes 2423.
[0346] In the same manner, the point cloud encoder calculates the median position of the positions of the nodes 2422 and 2424 (for example, nodes of level n-1). Since the position closest to the median position is c2, the point cloud encoder according to the embodiment matches the position and attributes of c2 with the parent node 2425 (for example, a top node) of the nodes 2422 and 2424.
[0347] As described with reference to Figure 23 In order to limit the attributes of the upper nodes higher than the leaf nodes, the point cloud encoder according to the embodiment can select one of the attributes of the child nodes (neighbor nodes) of the node (example 2320). Figure 23 The point cloud encoder according to the embodiment can match the position and attributes of the selected child node. The following equation is an embodiment of the equation corresponding to the example 2320. Figure 23
[0348] [Equation 7]
[0349]
[0350] In this formula, Attr(xp, yp, zp) represents the attributes of the parent node at position (xp, yp, zp), and Attr(xn, yn, zn) represents the attributes of the neighboring nodes. M represents the sum of the weights of the peripheral nodes (or neighboring nodes). In other words, the above formula is used to determine the attribute of the neighboring node that has the smallest difference from the attribute of the parent node as the estimated attribute (i.e., the predicted value) that matches the attribute of the parent node.
[0351] As described above, the point cloud data transmitting device (for example, referring to Figure 1 、 Figure 12 、 Figure 14 and Figure 15 The point cloud data transmitting apparatus described herein may transmit the encoded point cloud data in the form of a bitstream 2600. The bitstream 2600 according to an embodiment may include one or more sub-bitstreams.
[0352] Point cloud data transmitting device (for example, see Figure 1 、 Figure 12 、 Figure 14 and Figure 15 The point cloud data sending device described herein) may divide the image of the point cloud data into one or more packets in consideration of errors in the transmission channel, and send it via the network. According to an embodiment, the bit stream 2600 may include one or more packets (e.g., network abstraction layer (NAL) units). Therefore, even when some packets are lost under a poor network environment, the point cloud data receiving device may use the remaining packets to reconstruct the image. The point cloud data may be divided into one or more slices or one or more patches to be processed. The patches and slices according to the embodiment are areas where point cloud compression encoding is performed by segmenting the screen of the point cloud data. The point cloud data sending device may provide high-quality point cloud content by processing data corresponding to each area according to the importance of each segmented area of the point cloud data. That is, the point cloud data sending device according to the embodiment may perform point cloud compression encoding with better compression efficiency and appropriate delay on the data corresponding to the area important to the user.
[0353] A patch according to an embodiment represents a cuboid in a three-dimensional space (e.g., a bounding box) in which point cloud data is distributed. A slice according to an embodiment is a series of syntactic elements representing some or all of the encoded point cloud data, and represents a collection of points that can be independently encoded or decoded. According to an embodiment, a slice may include data sent via a packet, and may include one geometry data unit and zero or more attribute data units. According to an embodiment, a patch may include one or more slices.
[0354] The bitstream 2600 according to an embodiment may include signaling information and one or more slices, the signaling information including a sequence parameter set (SPS) for sequence level signaling, a geometry parameter set (GPS) for geometry information coding signaling, an attribute parameter set (APS) for attribute information coding signaling, and a patch parameter set (TPS) for patch level signaling.
[0355] The SPS according to an embodiment is encoding information about the entire sequence including a profile and a level, and may include comprehensive information about the entire file, such as picture resolution and video format.
[0356] According to an embodiment, a slice (e.g., Figure 30 The slice 0) includes a slice header and slice data. The slice data may include a geometry bit stream (Geom0 0 ) and one or more attribute bitstreams (Attr0 0 、Attr1 0 ). The geometry bitstream may include a header (e.g., a geometry slice header) and a payload (e.g., geometry slice data). The header of the geometry bitstream according to an embodiment may include identification information (geom_geom_parameter_set_id) of a parameter set included in GPS, a tile identifier (geom_tile id), a slice identifier (geom_slice_id), and information related to data included in the payload. The attribute bitstream may include a header (e.g., an attribute slice header or an attribute block header) and a payload (e.g., attribute slice data or attribute block data).
[0357] As reference Figures 18 to 24 As described, the point cloud encoder according to the embodiment can generate a colored octree and generate scalable representation information including signaling information related to the colored octree generation. Therefore, the bitstream includes the scalable representation information. In order to perform scalable representation, the point cloud decoder generates a reference Figures 18 to 24 Therefore, the point cloud decoder according to the embodiment acquires the scalable representation information included in the bitstream and performs scalable decoding based on the scalable representation information.
[0358] The scalable representation information included in the bitstream according to an embodiment may be provided by a metadata processor or a transport processor (e.g., Figure 12 Transport processor 12012), metadata processor or elements or references in the transport processor Figure 18 The attribute encoder 1840 generates the scalable representation information according to the embodiment. The scalable representation information may be generated based on the result of the attribute encoding.
[0359] The scalable representation information according to the embodiments can be included in the APS and the attribute slice. In addition, the scalable representation information according to the embodiments can be defined in association with attribute encoding (attribute encoding and attribute decoding), or can be defined independently. When defined in association with attribute scalable decoding and geometry scalable decoding, the scalable representation information can be included in the GPS. When the scalable representation information is applied to one or more point cloud bitstreams or applied per tile, it can be included in the SPS or the tile parameter set. The transmission location of the scalable representation information in the bitstream according to the embodiments is not limited to this example.
[0360] Figure 25 An example syntax of the APS according to the embodiments is shown.
[0361] In Figure 25 , the example syntax of the APS according to the embodiments can include the following information (or field, parameter, etc.).
[0362] The aps_attr_parameter_set_id indicates an identifier of the APS for other syntax elements to refer to. The value of the aps_attr_parameter_set_id should be in the range of 0 to 15. Since one or more attribute bitstreams are included in the bitstream, a field (for example, ash_attr_parameter_set_id) having the same value as the aps_attr_parameter_set_id can be included in the header of each attribute bitstream.
[0363] The point cloud decoder according to the embodiments can ensure the APS corresponding to each attribute bitstream based on the aps_attr_parameter_set_id and process the corresponding attribute bitstream.
[0364] The aps_seq_parameter_set_id specifies the value of the sps_seq_parameter_set_id of the active SPS. The value of the aps_seq_parameter_set_id should be in the range of 0 to 15.
[0365] The scalable encoding information 2500 according to the embodiments is described below.
[0366] The scalable_representation_available_flag indicates whether scalable decoding (or scalable representation) is available. As described with reference to FIG. 25, the scalable_representation_available_flag can be included in the attribute header or the attribute slice. Figures 18 to 24As described, a point cloud encoder according to an embodiment may generate a shaded octree to enable scalable representation and process attributes based on the shaded octree. When scalable_representation_available_flag is equal to 1, scalable_representation_available_flag indicates that the decoded point cloud data (decoded attributes) has a structure (shaded octree structure) for which scalable representation is available. Therefore, the receiving device may generate a shaded octree based on this information to provide scalable representation. When scalable_representation_available_flag is equal to 0, scalable_representation_available_flag indicates that the decoded point cloud data does not have a structure for which scalable representation is available.
[0367] octree_colorization_type indicates the type of shading octree or the method of generating the shading octree. octree_colorization_type equal to 0 indicates that the octree is paired according to the attribute (for example, refer to Figure 23 octree_colorization_type is equal to 1 to indicate that the octree is generated according to the attribute pairing octree described in Figure 24 The shading octree is generated by using the point pairing octree generation method described in
[15] .
[0368] The relevant parameters given when octree_colorization_type is equal to 0 (i.e., when the shading octree is an attribute-paired octree) are disclosed below.
[0369] Matched_attribute_type indicates the type of attribute that is matched to the octree node. When matched_attribute_type is equal to 0, matched_attribute_type indicates that the attribute is an estimated attribute (e.g., an attribute estimated based on the attributes of neighboring nodes or child nodes). Since the process of calculating the estimated attribute and matching the estimated attribute is the same as that of the reference Figure 23 The process described in Example 2310 is the same or similar, so its detailed description will be omitted. When Matched_attribute_type is equal to 1, Matched_attribute_type indicates that the attribute is an actual attribute (e.g., an attribute of a child node). Since the actual attribute matching process is the same as the reference attribute matching process, the matching process is repeated. Figure 23 The processes described (e.g., Example 2320) are the same or similar, so their detailed description will be omitted.
[0370] attribute_selection_type indicates the method for matching attributes to octree nodes. When attribute_selection_type is equal to 0, the average value of the attributes corresponding to the child nodes is estimated. When attribute_selection_type is equal to 1, the mean value of the attributes corresponding to the child nodes is estimated. When attribute_selection_type is equal to 2, the attributes correspond to the attributes of the child nodes in a fixed order (for example, the attribute of the first child node or the attribute of the second child node among the child nodes sorted in ascending order).
[0371] The following describes the relevant parameters given when octree_colorization_type is equal to 1 (ie, the shading octree is a point pair octree).
[0372] point_data_selection_type indicates the method or type of selecting point data to be matched with an octree node.
[0373] When point_data_selection_type is equal to 0, point_data_selection_type indicates a method (eg, Figure 24 Example 2400). When point_data_selection_type is equal to 1, point_data_selection_type indicates a method of selecting a point of a child node whose position is closest to an average position of positions of occupied nodes among child nodes (eg, Figure 24 Example 2410). When point_data_selection_type is equal to 2, point_data_selection_type indicates a method of selecting a point of a child node whose position is closest to the mean position of the positions of the occupied nodes among the child nodes (eg, Figure 24 Example 2420).
[0374] The following describes the relevant parameters provided when point_data_selection_type is equal to 0 or 3. point_cloud_geometry_info_present_flag indicates whether geometric information about the point data (or points) matching the octree node is directly provided. When point_cloud_geometry_info_present_flag is equal to 1, geometric information (e.g., position) about the point data (or points) matching the octree node is sent together. When point_cloud_geometry_info_present_flag is equal to 0, geometric information about the point data (or points) matching the octree node is not sent together.
[0375] Figure 26 An exemplary syntax of an attribute slice bitstream according to an embodiment is shown.
[0376] Figure 26 The first syntax 2600 shown represents an example of the syntax of an attribute slice bitstream according to an embodiment.The attribute slice bitstream includes an attribute slice header (attribute_slice_header) and attribute slice data (attribute_slice_data).
[0377] Figure 26 The second syntax 2610 shown is an example of the syntax of the attribute header according to an embodiment. The syntax of the attribute header may include the following information (or fields, parameters, etc.).
[0378] ash_attr_parameter_set_id has the same value as aps_attr_parameter_set_id of the active APS (e.g., refer to Figure 25 aps_attr_parameter_set_id included in the syntax of the APS described.
[0379] ash_attr_sps_attr_idx identifies an attribute set included in the active SPS. The value of ash_attr_sps_attr_idx falls within the range of 0 to the value of sps_num_attribute_sets included in the active SPS.
[0380] Figure 26 The third syntax 2620 shown is an example of the syntax of attribute slice data according to an embodiment. The syntax of attribute slice data may include the following information.
[0381] dimension=attribute_dimension[ash_attr_sps_attr_idx] indicates the attribute dimension (attribute_dimension) of the attribute set identified by ash_attr_sps_attr_idx. Attribute_dimension indicates the number of components that constitute the attribute. The attributes according to the embodiment represent reflectivity, color, etc. Therefore, the number of components possessed by the attribute is different. For example, the attribute corresponding to color may have three color components (e.g., RGB). The attribute corresponding to reflectivity may be a one-dimensional attribute, and the attribute corresponding to color may be a three-dimensional attribute. The attributes according to the embodiment may be attribute-encoded on a per-dimension basis. For example, the attribute corresponding to reflectivity and the attribute corresponding to color may be attribute-encoded separately. The attributes according to the embodiment may be attribute-encoded regardless of the dimension. For example, the attribute corresponding to reflectivity and the attribute corresponding to color may be attribute-encoded together.
[0382] When scalable attribute decoding according to an embodiment is applied to each slice, the syntax of the attribute slice data includes a bitstream according to the attribute coding type. The APS according to the embodiment may include attr_coding_type. attr_coding_type indicates the attribute coding type. In the bitstream according to the embodiment, attr_coding_type is equal to any one of 0, 1 or 2. Other values of attr_coding_type may be reserved for future use by ISO / IEC. Therefore, the point cloud decoder according to the embodiment may ignore values of attr_coding_type other than 0, 1 and 2. 0 indicates that the attribute coding type is prediction weight boosting transform coding, and 1 indicates that the attribute coding type is RAHT transform coding. 2 indicates that the attribute coding type is fixed weight boosting.
[0383] When attr_coding_type is equal to 0, the attribute coding type is prediction weight lifting transform coding. Therefore, the syntax of the attribute slice data includes PredictingWeight_Lifting bitstream (PredictingWeight_Lifting_bitstream(dimension)).
[0384] When attr_coding_type is equal to 1, the attribute coding type is RAHT transform coding. Therefore, the syntax of attribute slice data includes RAHT bitstream (RAHT_bitstream(dimension)).
[0385] When attr_coding_type (for example, refer to Figure 23When the attribute coding type is equal to 2, the attribute coding type is fixed prediction weight lifting transform coding. Therefore, the syntax of the attribute slice data includes FixedWeight_Lifting bitstream (FixedWeight_Lifting_bitstream(dimension)).
[0386] In addition, as reference Figure 25 As described, the APS according to the embodiment includes point_cloud_geometry_info_present_flag. When point_cloud_geometry_info_present_flag is equal to 1, the syntax of the attribute slice data also includes a shaded octree position bitstream.
[0387] The fourth syntax 2630 shown in the figure is an example of the syntax of the colored octree position bitstream. The syntax of the colored octree position bitstream includes the following parameters.
[0388] colorization_start_depth_level indicates the starting octree level or the octree depth level to which octree shading is applied for generating a shading octree. colorization_end_depth_level indicates the ending octree level or the octree depth level to which octree shading is applied.
[0389] Therefore, the total number of octree depth levels to which octree shading is applied is expressed as a value obtained by adding 1 to the difference between the level indicated by colorization_end_depth_level and the level indicated by colorization_start_depth_level.
[0390] numOctreeDepthLevel=colorization_end_depth_level-colorization_start_depth_level+1
[0391] The following is information about each octree depth level to which octree shading is applied. num_colorized_nodes[i] indicates the number of nodes to which octree shading is applied for the i-th octree depth level. i has a value greater than or equal to 0 and less than the number indicated by numOctreeDepthLevel. Each octree depth level (the (i+colorization_start_depth_level)th octree depth level) is equal to the sum of the value indicated by i and the value indicated by colorization_start_depth_level.
[0392] position_index[i][j] indicates the position of the j-th node at the (i+colorization_start_depth_level)-th octree level. j represents the index of each node. j has a value greater than or equal to 0 and less than the value of num_colorized nodes. The position according to the embodiment includes a coordinate value in a three-dimensional coordinate system consisting of x, y, and z axes and / or a sequential position of a corresponding point in a Morton code order. position_index[i][j] according to the embodiment may be sent for each node. When a method (e.g., point cloud data (or points) of a node in a fixed order (e.g., the first node in ascending order among the child nodes) is selected according to point_data_selection_type, Figure 24 When the same information is repeated in the full octree structure or the corresponding octree level as in Example 2400), position_index[i][j] according to the implementation may be sent or its transmission position may be changed only at the repetition time (e.g., the syntax may be changed).
[0393] Figure 27 is a block diagram illustrating the encoding operation of a point cloud encoder.
[0394] According to an embodiment, the point cloud encoder 2700 (e.g., Figure 1 Point cloud video encoder 10002, Figure 4 Point cloud encoder, reference Figure 12 、 Figure 14 and Figure 15 Point cloud encoder described in Figures 1 to 26 Describes encoding operations (including octree shading).
[0395] Point cloud (PCC) data or point cloud compression (PCC) data is input data of the point cloud encoder 2700 and may include geometry and / or attributes. Geometry according to an embodiment is information indicating the position (e.g., location) of a point and may be expressed as parameters of a coordinate system such as orthogonal coordinates, cylindrical coordinates, or spherical coordinates. Attributes according to an embodiment indicate attributes of a point (e.g., color, transparency, reflectivity, grayscale, etc.). Geometry may be referred to as geometric information (or geometric data), and attributes may be referred to as attribute information (or attribute data).
[0396] The point cloud encoder 2700 according to the embodiment performs octree generation 2710, geometric prediction 2720, and entropy encoding 2730 to perform reference Figures 1 to 26 The geometry encoding described herein and the output geometry bitstream. The octree generation 2710, geometry prediction 2720 and entropy encoding 2730 according to the embodiment are the same as those of the reference Figure 4The operations and / or references of the coordinate transformer 40000, quantizer 40001, octree analyzer 40002, surface approximation analyzer 40003, arithmetic encoder 40004, and geometry reconstructor 40005 are described. Figure 12 The operations of the described data input unit 12000, quantization processor 12001, voxelization processor 12002, octree occupancy code generator 12003, surface model processor 12004, intra-frame / inter-frame coding processor 12005, arithmetic encoder 12006, and metadata processor 12007 are the same or similar, so their detailed description will be omitted.
[0397] To perform the reference Figures 1 to 26 Attribute encoding is described, and the point cloud encoder 2700 according to an embodiment performs shading octree generation 2740, attribute prediction 2750, transformation 2760, quantization 2770, and entropy encoding 2780.
[0398] According to the embodiment, shading octree generation 2740, attribute prediction 2750, transformation 2760, quantization 2770 and entropy coding 2780 are compared with reference Figure 4 The operation of and / or reference to the geometry reconstructor 40005, color transformer 40006, attribute transformer 40007, RAHT transformer 40008, LOD generator 40009, lifting transformer 40010, coefficient quantizer 40011 and / or arithmetic encoder 40012 described herein Figure 12 The operations of the described color transform processor 12008, attribute transform processor 12009, prediction / lifting / RAHT transform processor 12110 and arithmetic encoder 12011 are the same or similar, so their detailed description will be omitted.
[0399] The point cloud encoder 2700 according to the embodiment generates a shaded octree based on the octree structure or information about the octree structure generated in the octree generation 2710 (2740). Figures 18 to 24The octree coloring described is the same or similar. The point cloud encoder 2700 according to the embodiment predicts (estimates) and removes the similarity between the attributes of each depth level based on the colored octree (attribute prediction 2750). The point cloud encoder 2700 according to the embodiment performs attribute prediction 2750 based on the spatial distribution of adjacent data. The point cloud encoder 2700 transforms the predicted attributes into a format suitable for transmission or a domain with high compression efficiency (2760). The point cloud encoder 2700 may perform or not perform transformation based on various transformation methods (e.g., DCT type transformation, lifting transformation, RAHT, wavelet transformation, etc.) according to the data type. Thereafter, the point cloud encoder 2700 quantizes the transformed attributes (2770) and performs entropy coding to transform them into bit-unit data for transmission (2780).
[0400] The point cloud encoder 2700 according to an embodiment may perform attribute encoding based on a colored octree structure.
[0401] The point cloud encoder 2700 according to the embodiment detects neighbor nodes. Figures 18 to 26 As described above, neighbor nodes are child nodes belonging to the same parent node. In the octree structure according to the embodiment, the child nodes are likely to have adjacent positions. Therefore, under the assumption that the child nodes are adjacent to each other in the three-dimensional space composed of the x, y and z axes, the point cloud encoder 2700 performs prediction on the current node among the child nodes. The point cloud encoder 2700 according to the embodiment can obtain prediction values for each child node. In addition, the point cloud encoder 2700 defines the child nodes belonging to the same parent node as sibling nodes and defines the sibling nodes as having the same prediction value to reduce the number of coefficients required for encoding of each child node and perform efficient encoding. In addition, the same prediction value can be used as an attribute of the occupied node and can be used to predict attributes that match the parent node. The formula given below represents the prediction attribute of level 1 and the prediction attribute of level 1-1.
[0402] [Formula 8]
[0403]
[0404] In this formula, represents the quotient obtained when a is divided by b. In this formula, Cl represents the lth attribute. k, l, and m are parameters that define the position of the peripheral node (or neighbor node) relative to the position (x, y, z).
[0405] The point cloud encoder 2700 according to an embodiment calculates the attribute prediction error of each child node based on the predicted attribute. The following formula represents the residual (rl(x, y, z)) between the original attribute value and the predicted attribute value as a method of calculating the attribute prediction error of each child node.
[0406] [Formula 9]
[0407] r l (x,y,z)=g{c l (x,y,z),p l (x,y,z)}=c l (x,y,z)-p l (x,y,z)
[0408] In this formula, rl(x,y,z) denotes a residual of the l-th node having a position (x,y,z). cl(x,y,z) denotes an actual attribute (e.g., color) of the l-th node having a position (x,y,z). pl(x,y,z) denotes an estimated attribute (or prediction value) of the l-th node having a position (x,y,z).
[0409] The method of calculating an attribute prediction error according to the embodiments is not limited to this example, and can be based on various methods (e.g., weighted difference, weighted average difference, etc.).
[0410] As described above, in the octree structure which is an eight-ary tree structure, the depth increases from an upper node to a lower node. The depth according to the embodiments is referred to as a level and / or a layer. Accordingly, the upper node corresponds to a lower level, and the lower node corresponds to a higher level (e.g., the highest node corresponds to level 0, and the lowest node corresponds to level n). According to the embodiments, the level can be set to decrease from the upper node to the lower node (e.g., the highest node is level n, and the lowest node is level 0). The point cloud encoder 2700 according to the embodiments can signal a prediction attribute value of the highest level, and can signal an attribute prediction error of a lower level. In addition, the point cloud encoder 2700 according to the embodiments can determine a data (attribute) transmission order of each level (level in the octree structure) in consideration of a decoding process. For example, the point cloud encoder 2700 transmits data in ascending order of level (e.g., data transmission starts from a prediction attribute value of an upper node, and proceeds to a lower node). At the same level, data can be transmitted in ascending order of coordinate values on x, y, and z axes (e.g., in Morton order based on a Morton code). According to the embodiments, the point cloud encoder 2700 can perform rearrangement. In addition, the point cloud encoder 2700 according to the embodiments performs quantization on a prediction attribute value and an attribute prediction error, which is represented by the following formula.
[0411] [Formula 10]
[0412] d′ l (x,y,z)=Q{d l (x,y,z)}=round[d l (x,y,z) / q]
[0413] In this formula, q is a quantization coefficient, and the degree of quantization is determined according to the quantization coefficient. dl(x, y, z) represents the data of the lth node with position (x, y, z) to be quantized. dl'(x, y, z) represents the quantized data.
[0414] The point cloud encoder 2700 according to an embodiment may use different quantization coefficients according to the predicted attribute value and the attribute prediction error or for each level.
[0415] Information related to the geometry encoding and attribute encoding of the point cloud encoder 2700 according to the embodiment (eg, scalable representation information) can be obtained by referring to Figure 25 and Figure 26 The described bit stream is sent to the receiving device.
[0416] Figure 28 is a block diagram illustrating the decoding operation of a point cloud decoder.
[0417] The point cloud decoder 2800 according to the embodiment performs the Figure 27 The point cloud decoder 2800 (e.g., referring to FIG. 2700 ) according to an embodiment of the present invention is a point cloud decoder 2800 (e.g., referring to FIG. 2700 ). Figure 10 Point cloud video decoder 10006 described, refer to Figure 11 Point cloud decoder described, reference Figure 19 The point cloud decoder described above can be executed Figures 1 to 25 Describes the decoding operation.
[0418] The point cloud decoder 2800 according to an embodiment receives a bit stream. The bit stream according to an embodiment (e.g., referring to Figure 25 and Figure 26 The bitstream described in the embodiment includes a geometry bitstream and an attribute bitstream. The geometry bitstream according to the embodiment may include full geometry or partial geometry. Since the bitstream according to the embodiment includes a reference Figure 25 and Figure 26 The scalable representation information described, so the point cloud decoder 2800 can perform reference based on the scalable representation information Figures 18 to 26 Scalable representation of descriptions.
[0419] The point cloud decoder 2800 according to an embodiment performs entropy decoding 2810 on the geometry bitstream to perform reference Figures 1 to 26 The described geometry encoding (or geometry decoding) is performed, and octree reconstruction 2820 is performed to output geometry data and octree structure.
[0420] Entropy decoding 2810 and octree reconstruction 2820 according to an embodiment and reference Figure 11The operations of the described arithmetic decoder 11000, octree synthesizer 11001, surface approximation synthesizer 11002, geometry reconstructor 11003 and inverse coordinate transformer 11004 are the same or similar.
[0421] Therefore, a detailed description thereof will be omitted.
[0422] The point cloud decoder 2800 according to the embodiment performs reference decoding by performing entropy decoding 2830, inverse quantization 2840, inverse transformation 2850, attribute reconstruction 2860, and shading octree generation 2870 on the attribute bitstream. Figures 1 to 26 The attribute encoding (or attribute decoding) described in the embodiment is performed and the attribute data is output. Figure 11 The operations of the arithmetic decoder 11005, the inverse quantizer 11006, the RAHT transformer 11007, the LOD generator 11008, the inverse lifter 11009 and / or the color inverse transformer 11010 described are the same or similar, so their detailed descriptions will be omitted. Figure 27 The operation of the point cloud encoder 2700 performs or does not perform inverse quantization 2840 and inverse transform 2850 depending on the implementation.
[0423] The point cloud decoder 2800 according to the embodiment performs attribute reconstruction 2860 based on the octree structure and the scalable representation depth level (or scalable representation level) output from the octree reconstruction 2820. The scalable representation depth level according to the embodiment may be determined by the receiving device according to performance. In addition, the scalable representation depth level according to the embodiment may be determined by the point cloud encoder (e.g., the point cloud encoder 2700) and transmitted together with the encoded point cloud data. Similar to the reference Figure 27 The point cloud encoder 2700 and the point cloud decoder 2800 described above detect neighbor nodes based on the octree structure for reconstructing position information. The point cloud decoder 2800 according to the embodiment may detect neighbor nodes according to the definition of neighbor nodes signaled by signaling information included in the bitstream (e.g., referring to Figure 27The point cloud decoder 2800 according to an embodiment may predict attributes in the reverse order of the attribute prediction performed by the cloud encoder. For example, when the point cloud encoder performs attribute prediction in descending order of levels (for example, in the direction from the leaf node to the root node), the point cloud decoder 2800 predicts attributes in ascending order of levels (for example, in the direction from the root node to the leaf node). The point cloud decoder 2800 according to an embodiment uses the reconstructed attributes of the parent node as the predicted value of the child node in the same manner as the attribute prediction of the point cloud encoder. According to an embodiment, when one or more prediction methods are used, the above-mentioned bitstream may include information about the one or more prediction methods. The point cloud decoder 2800 may predict attributes based on signaling information sent through the bitstream. The following formula represents the attribute prediction process of the point cloud decoder 2800.
[0424] [Equation 11]
[0425]
[0426] The point cloud decoder 2800 according to the embodiment may reconstruct the attributes of each child node based on the predicted attributes (2860). The attribute reconstruction 2860 of the point cloud decoder 2800 corresponds to the reference Figure 27 The inverse process of the attribute prediction 2750 of the point cloud encoder 2700 is described. Figure 27 As described above, when the point cloud encoder 2700 generates an attribute prediction error, the point cloud decoder 2800 can reconstruct the attribute by adding the predicted attribute to the decoded attribute prediction error. The following formula represents the attribute reconstruction process.
[0427] [Equation 12]
[0428]
[0429] The shading octree generation 2870 according to an embodiment is performed for scalable decoding or scalable representation. The point cloud decoder 2800 according to an embodiment generates a shading octree (shading octree generation 2870) based on the octree structure output in the octree reconstruction 2820 and the scalable representation depth level (or scalable representation level).
[0430] According to the embodiment, the shaded octree is generated 2870 and referenced Figures 18 to 24 The octree shading described is the same or similar. Therefore, the receiving device or point cloud decoder 2800 can provide point cloud content of various resolutions.
[0431] Figure 29 Details of geometry and attributes according to scalable decoding according to an embodiment are shown.
[0432] Figure 29The upper portion of the octree is an example 2900 showing the details of geometry according to scalable decoding. A first arrow 2910 indicates the direction from the upper node to the lower node in the octree. As shown in the figure, as scalable decoding progresses from the upper node to the lower node in the octree, more points are present, and thus the detail of the geometry increases. The leaf nodes of the octree structure correspond to the top level of detail of the geometry.
[0433] Figure 29 The lower portion of FIG. 29 is an example 2920 showing details of attributes according to scalable decoding. A second arrow 2930 indicates the direction from the upper node to the lower node in the octree. As shown in the figure, as scalable decoding progresses from the upper node to the lower node in the octree, the details of the attribute increase.
[0434] Figure 30 is an exemplary flow chart of a method for processing point cloud data according to an embodiment.
[0435] Figure 30 Flowchart 3000 shows a method of processing a point cloud data by a point cloud data processing device (e.g., referring to Figure 1 、 Figure 11 、 Figure 14 、 Figure 15 and Figure 18 Described sending device or reference Figure 27 The point cloud data processing method described in the point cloud data encoder 2700) is executed. The point cloud data processing device according to the embodiment can be executed with reference to Figures 1 to 27 The encoding operations described are the same or similar operations.
[0436] The point cloud data processing apparatus according to an embodiment may encode point cloud data including geometric information and attribute information (3010). The geometric information according to an embodiment indicates the position of a point in the point cloud data. The attribute information according to an embodiment indicates the attribute of a point in the point cloud data.
[0437] The point cloud data processing apparatus according to the embodiment can encode the geometric information and encode the attribute information. Figures 1 to 27 In addition, the point cloud data processing device performs the same or similar operations as those described in the geometric information encoding. Figures 1 to 27 The point cloud data processing apparatus according to the embodiment receives an octree structure of geometric information. The octree structure is represented by one or more levels. The point cloud data processing apparatus according to the embodiment generates a colored octree by matching one or more attributes with each level of the octree structure (or by matching attributes with nodes at each level). The colored octree according to the embodiment is used to encode attribute information to perform scalable representation of some or all attribute information.
[0438] Since the geometric information encoding and attribute information encoding according to the embodiment are related to the reference Figures 1 to 27 The geometric information encoding and attribute information encoding described are the same or similar, so their detailed description will be omitted.
[0439] The point cloud data processing apparatus according to the embodiment may transmit a bit stream including encoded point cloud data ( 3020 ).
[0440] Since the structure of the bitstream according to the embodiment is related to the reference Figure 25 and Figure 26 The bitstream according to the embodiment may include scalable representation information (e.g., referring to Figures 18 to 26 Scalable presentation information according to the embodiment may be as described in Figure 25 and Figure 26 The attribute data transmitted via APS and attribute slices is not limited to the above example.
[0441] The scalable representation information according to an embodiment includes information indicating whether scalable representation is available (e.g., referring to Figure 25 When the information indicates that scalable representation is available, the scalable representation information also includes information indicating a method of generating a shading octree structure (e.g., referring to Figure 25 The information indicating the method of generating the colored octree structure according to the embodiment indicates that the method of generating the colored octree structure is a method of matching one or more arbitrary attributes with each level of the octree structure (for example, octree_colorization_type is equal to 0, and the colored octree is an attribute pairing octree (for example, referring to Figure 23 The attribute pairing octree described in
[15] ) and the method to match the actual attributes and positions of the lower levels of the level with the various levels of the octree structure (e.g., octree_colorization_type is equal to 1 and the shading octree is a point-pairing octree (e.g., refer to Figure 24 At least one of the point pairing octrees described herein. Figures 18 to 26 Those described are the same, so their description will be omitted.
[0442] Figure 31 is an exemplary flow chart of a method for processing point cloud data according to an embodiment.
[0443] Figure 31 Flowchart 3100 shows a method for processing point cloud data by a point cloud data processing device (e.g., referring to Figure 1 、 Figure 13 、 Figure 14 、 Figure 16 and Figure 25 Point cloud data receiving device described or Figure 28 The point cloud data processing method performed by the point cloud data decoder 2800)) is as follows. The point cloud data processing device according to the embodiment can be executed with reference to Figures 1 to 28 The decoding operations described are the same or similar operations.
[0444] The point cloud data processing apparatus according to an embodiment receives a bit stream including point cloud data (3110). The geometric information according to an embodiment indicates the position of a point of the point cloud data. The attribute information according to an embodiment indicates the attribute of a point of the point cloud data. The structure and reference of the bit stream according to an embodiment Figure 25 and Figure 26 The descriptions are the same, so their detailed descriptions will be omitted.
[0445] The point cloud data processing apparatus according to the embodiment decodes the point cloud data ( 3120 ).
[0446] The point cloud data processing device according to the embodiment can decode the geometric information and decode the attribute information. Figures 1 to 28 In addition, the point cloud data processing device performs the same or similar operation as that described in the geometric information decoding. Figures 1 to 28 The attribute information described decodes the same or similar operations.
[0447] A bitstream according to an embodiment may include scalable representation information (e.g., referring to Figure 25 and Figure 26 Scalable presentation information according to the embodiment may be as described in Figure 25 and Figure 26 The attribute data transmitted via APS and attribute slices is not limited to the above example.
[0448] The scalable representation information according to an embodiment includes information indicating whether scalable representation is available (e.g., referring to Figure 25 When the information indicates that scalable representation is available, the scalable representation information also includes information indicating a method of generating a shading octree structure (e.g., referring to Figure 25The information indicating the method of generating the colored octree structure according to the embodiment indicates that the method of generating the colored octree structure is a method of matching one or more arbitrary attributes with each level of the octree structure (for example, octree_colorization_type is equal to 0, and the colored octree is an attribute pairing octree (for example, referring to Figure 23 The attribute pairing octree described in
[15] ) and the method to match the actual attributes and positions of the lower levels of the level with the various levels of the octree structure (e.g., octree_colorization_type is equal to 1 and the shading octree is a point-pairing octree (e.g., refer to Figure 24 At least one of the point pairing octrees described herein. Figures 18 to 26 Those described are the same, so their description will be omitted.
[0449] The point cloud data processing apparatus receives the decoded octree structure of geometric information. The octree structure according to the embodiment is represented by one or more levels. The point cloud data processing apparatus according to the embodiment generates a shaded octree by matching one or more attributes with each level of the octree structure (or matching attributes to nodes of each level) based on the scalable representation information. Figures 1 to 28 Described,shading octrees for scalable representation.
[0450] According to the reference Figures 1 to 31 The components of the point cloud data processing apparatus of the described embodiments may be implemented as hardware, software, firmware, or a combination thereof, including one or more processors coupled to a memory. The components of the apparatus according to the embodiments may be implemented as a single chip, such as a single hardware circuit. Alternatively, the components of the point cloud data processing apparatus according to the embodiments may be implemented as separate chips. In addition, at least one component of the point cloud data processing apparatus according to the embodiments may include one or more processors capable of executing one or more programs, wherein the one or more programs may include executing or being configured to execute reference Figures 1 to 31 Instructions for one or more operations / methods of a point cloud data processing apparatus are described.
[0451] Although the drawings are described separately for simplicity, new embodiments can be designed by combining the embodiments shown in the various figures. Designing a recording medium that can be read by a computer and records a program for executing the above-mentioned embodiments according to the needs of those skilled in the art also falls within the scope of the attached claims and their equivalents. The apparatus and method according to the embodiment may not be limited to the configuration and method of the above-mentioned embodiment. Various modifications can be made to the embodiment by selectively combining all or some of the embodiments. Although preferred embodiments have been described with reference to the drawings, it will be understood by those skilled in the art that various modifications and changes can be made to the embodiment without departing from the spirit or scope of the present disclosure described in the attached claims. These modifications should not be understood separately from the technical ideas or viewpoints of the embodiment.
[0452] The descriptions of the methods and apparatuses can be applied to complement each other. For example, the point cloud data transmission method according to the embodiment can be performed by the point cloud data transmitting apparatus according to the embodiment or a component included in the point cloud data transmitting apparatus. In addition, the point cloud data receiving method according to the embodiment can be performed by the point cloud data receiving apparatus according to the embodiment or a component included in the point cloud data receiving apparatus.
[0453] The various elements of the device of the embodiment can be implemented by hardware, software, firmware or a combination thereof. The various elements in the embodiment can be implemented by a single chip (e.g., a single hardware circuit). According to the embodiment, the components according to the embodiment can be implemented as separate chips respectively. According to the embodiment, at least one or more components of the device according to the embodiment may include one or more processors capable of executing one or more programs. One or more programs can execute any one or more operations / methods according to the embodiment or include instructions for executing them. The executable instructions for executing the method / operation of the device according to the embodiment can be stored in a non-transitory CRM or other computer program product configured to be executed by one or more processors, or can be stored in a temporary CRM or other computer program product configured to be executed by one or more processors. In addition, the memory according to the embodiment can be used as not only covering volatile memory (e.g., RAM), but also covering the concept of non-volatile memory, flash memory and PROM. In addition, it can also be implemented in the form of a carrier wave (e.g., transmission via the Internet). In addition, the processor-readable recording medium can be distributed to computer systems connected via a network so that the processor-readable code can be stored and executed in a distributed manner.
[0454] In this specification, the terms " / " and "," should be interpreted as indicating "and / or". For example, the expression "A / B" may mean "A and / or B". In addition, "A, B" may mean "A and / or B". In addition, "A / B / C" may mean "at least one of A, B and / or C". In addition, "A / B / C" may mean "at least one of A, B and / or C". In addition, in this specification, the term "or" should be interpreted as indicating "and / or". For example, the expression "A or B" may mean 1) only A, 2) only B, or 3) both A and B. In other words, the term "or" used in this document should be interpreted as indicating "in addition or alternatively".
[0455] Terms such as first and second may be used to describe various elements of an embodiment. However, the various components according to the embodiment should not be limited by the above terms. These terms are only used to distinguish one element from another. For example, a first user input signal may be referred to as a second user input signal. Similarly, a second user input signal may be referred to as a first user input signal. The use of these terms should not be interpreted without departing from the scope of the various embodiments. Both the first user input signal and the second user input signal are user input signals, but unless the context clearly dictates otherwise, they do not refer to the same user input signal.
[0456] The terms used to describe the embodiments are used to describe specific embodiments and are not intended to limit the embodiments. As used in the description of the embodiments and in the claims, unless the context clearly dictates otherwise, the singular includes the plural referents. The expression "and / or" is used to include all possible combinations of terms. Terms such as "including" or "having" are intended to indicate the presence of graphics, quantities, steps, elements and / or components, and should be understood as not excluding the possibility of additional graphics, quantities, steps, elements and / or components. As used herein, conditional expressions such as "if" and "when" are not limited to optional cases and are intended to be interpreted as performing the relevant operation when a specific condition is met, or interpreting the relevant definition based on the specific condition.
[0457] Public mode
[0458] As described above, the relevant contents are described in the best mode for carrying out the embodiment.
[0459] Industrial Applicability
[0460] It will be apparent to those skilled in the art that various changes or modifications may be made to the embodiments within the scope of the embodiments. Thus, the embodiments are intended to cover the modifications and variations of this disclosure as long as they fall within the scope of the appended claims and their equivalents.
Claims
1. A method comprising the following steps: Encode the geometric information of point cloud data based on occupancy tree; as well as encoding the attribute information based on a level of detail of the attribute information of the point cloud data; generating syntax element information, the syntax element information including: scalability information for indicating reconstruction of the attribute information for a partial occupancy tree and type information for indicating a type of encoding of the attribute information; and Wherein, for the scalability information, the points of the level of detail are assigned to nodes of the occupancy tree.
2. The method according to claim 1, in, Select a centroid point of the geometric information.
3. The method according to claim 1, wherein The scalability information indicates that decoding of the attribute information is based on partially reconstructed geometric information, and the encoded attribute information is partially decoded using an index of a point indicated by the partially reconstructed geometric information.
4. A method comprising the following steps: Obtaining syntax element information in a bitstream, the syntax element information including scalability information indicating that attribute information in the bitstream is reconstructed for a partial occupancy tree and type information indicating a type of decoding the attribute information, wherein for the scalability information, a point of a level of detail for the attribute information is assigned to a node of the occupancy tree; decoding geometric information in the bitstream based on the occupancy tree; and The attribute information is decoded based on the level of detail.
5. The method according to claim 4, in, Select a centroid point of the geometric information.
6. The method according to claim 4, wherein: The scalability information indicates that decoding of the attribute information is based on partially reconstructed geometric information, and the attribute information is partially decoded using an index of a point indicated by the partially reconstructed geometric information.
7. A device comprising: Memory; as well as at least one processor connected to the memory, the at least one processor configured to: Obtaining syntax element information in a bitstream, the syntax element information including scalability information indicating that attribute information in the bitstream is reconstructed for a partial occupancy tree and type information indicating a type of decoding the attribute information, wherein for the scalability information, a point of a level of detail for the attribute information is assigned to a node of the occupancy tree; decoding geometric information in the bitstream based on the occupancy tree; and The attribute information is decoded based on the level of detail.
8. The device according to claim 7, in, Select a centroid point of the geometric information.
9. The device according to claim 7, wherein The scalability information indicates that decoding of the attribute information is based on partially reconstructed geometric information, and the attribute information is partially decoded using an index of a point indicated by the partially reconstructed geometric information.
10. A method comprising the steps of: Generate a bitstream for point cloud data, wherein the bitstream is generated by the following steps: encoding geometric information of the point cloud data based on an occupancy tree; and encoding attribute information based on a level of detail of attribute information for the point cloud data; and generating syntax element information including scalability information indicating that the attribute information is reconstructed for a partial occupancy tree and type information indicating a type of encoding of the attribute information, and wherein, for the scalability information, the level-of-detail points are assigned to nodes of the occupancy tree; and Data including the bitstream is transmitted.