Point cloud data processing method and apparatus
The method and apparatus efficiently process point cloud data by encoding and decoding geometry and attribute information, addressing latency and complexity issues to provide high-quality VR and autonomous driving services.
Patent Information
- Application Number
- JP2024033719
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-07-02
- Filing Date
- 2024-03-06
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2040-06-11
AI Technical Summary
Existing technologies face challenges in efficiently processing massive amounts of point cloud data required for VR, AR, MR, and autonomous driving services due to high latency and encoding/decoding complexity.
A method and apparatus for processing point cloud data by encoding geometry and attribute information, transmitting a bitstream, and decoding it efficiently using geometry and attribute information.
The solution enables high-efficiency processing of point cloud data, providing high-quality services such as VR and autonomous driving with reduced latency and improved encoding/decoding complexity.
Smart Images

Figure 0007775351000030 
Figure 0007775351000031 
Figure 0007775351000032
Abstract
Description
[Technical Field]
[0001] The embodiment provides a method for providing point cloud content to provide users with various services such as VR (Virtual Reality), AR (Augmented Reality), MR (Mixed Reality) and autonomous driving services. [Background technology]
[0002] Point cloud content is content expressed as a point cloud, which is a collection of points belonging to a coordinate system that represents three-dimensional space. Point cloud content can represent three-dimensional media and is used to provide various services such as VR (Virtual Reality), AR (Augmented Reality), MR (Mixed Reality), and autonomous driving services. However, expressing point cloud content requires tens of thousands to hundreds of thousands of point data. Therefore, a method for efficiently processing massive amounts of point data is required. Summary of the Invention [Problem to be solved by the invention]
[0003] SUMMARY OF THE INVENTION The embodiments provide an apparatus and method for efficiently processing point cloud data.The embodiments provide a method and apparatus for processing point cloud data to address latency and encoding / decoding complexity.
[0004] However, the scope of the invention is not limited to the above-mentioned technical problems, but can be extended to other technical problems that a person skilled in the art can derive based on all the contents described herein. [Means for solving the problem]
[0005] Therefore, in order to efficiently process point cloud data, a point cloud data processing method according to an embodiment includes encoding point cloud data including geometry information and attribute information, and transmitting a bitstream including the encoded point cloud data. According to an embodiment, the geometry information is information indicating the positions of points in the point cloud data, and the attribute information according to an embodiment is information indicating the attributes of points in the point cloud data.
[0006] A point cloud data processing device according to an embodiment includes an encoder for encoding point cloud data including geometry information and attribute information, and a transmitter for transmitting a bitstream including the encoded point cloud data. The geometry information according to an embodiment is information indicating positions of points in the point cloud data, and the attribute information according to an embodiment is information indicating attributes of points in the point cloud data.
[0007] A method for processing point cloud data according to an embodiment includes receiving a bitstream including point cloud data and decoding the point cloud data. The point cloud data according to an embodiment includes geometry information and attribute information. The geometry information according to an embodiment is information indicating positions of points of the point cloud data, and the attribute information according to an embodiment is information indicating one or more attributes of the points of the point cloud data.
[0008] A point cloud data processing device according to an embodiment includes a receiver for receiving a bitstream including point cloud data and a decoder for decoding the point cloud data. The point cloud data according to an embodiment includes geometry information and attribute information. The geometry information according to an embodiment is information indicating positions of points of the point cloud data, and the attribute information according to an embodiment is information indicating one or more attributes of the points of the point cloud data. [Effects of the Invention]
[0009] The apparatus and method according to the embodiment can process point cloud data with high efficiency.
[0010] The apparatus and method according to the embodiment can provide a high-quality point cloud service.
[0011] The apparatus and method according to the embodiment can provide point cloud content for providing general-purpose services such as VR services and autonomous driving services. [Brief explanation of the drawings]
[0012] The accompanying drawings illustrate embodiments of the invention and, together with the description, serve to explain the principles of the invention.
[0013] The drawings are attached for a better understanding of the embodiments and together with the description relating to the embodiments illustrate the embodiments.
[0014]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33
Figure 34
Figure 35
Figure 36
Figure 37
Figure 38
Figure 39
[0015] Preferred embodiments will be described in detail with reference to the accompanying drawings. The following detailed description with reference to the accompanying drawings is intended to illustrate preferred embodiments rather than merely showing embodiments that can be implemented by the embodiments. The following detailed description includes details to provide a thorough understanding of the embodiments. However, it will be apparent to those skilled in the art that embodiments can be practiced without such details.
[0016] Although most of the terms used in the examples are common and widely used in the relevant fields, some of them have been arbitrarily selected by the applicant, and their meanings will be explained in detail below as necessary. Therefore, the examples should be understood based on the intended meaning of the terms, rather than the simple names or meanings of the terms.
[0017] FIG. 1 is a diagram illustrating an example of a point cloud content providing system according to an embodiment.
[0018] The point cloud content providing system shown in Figure 1 includes a transmission device 10000 and a reception device 10004. The transmission device 10000 and the reception device 10004 are capable of wired and wireless communication to transmit and receive point cloud data.
[0019] According to an embodiment, the transmitting device 10000 acquires, processes, and transmits a point cloud video (or point cloud content). In the embodiment, the transmitting device 10000 includes a fixed station, a base transceiver system (BTS), a network, an AI (Artificial Intelligence) device and / or system, a robot, an AR / VR / XR device and / or server, etc. In the embodiment, the transmitting device 10000 also includes a device that communicates with a base station and / or other wireless devices using a wireless connection technology (e.g., 5G NR (New RAT), LTE (Long Term Evolution)), a robot, a vehicle, an AR / VR / XR device, a mobile device, a home appliance, an IoT (Internet of Things) device, an AI device / server, etc.
[0020] The transmitting device 10000 according to the embodiment includes a Point Cloud Video Acquisition unit 10001, a Point Cloud Video Encoder 10002, and / or a Transmitter (or communication module) 10003.
[0021] The point cloud video acquisition unit 10001 according to the embodiment acquires a point cloud video through a process such as capturing, synthesizing, or generating. The point cloud video is point cloud content represented by a point cloud, which is a collection of points located in a three-dimensional space, and is also called point cloud video data. The point cloud video according to the embodiment includes one or more frames. One frame represents a still image / picture. Therefore, the point cloud video includes a point cloud image / frame / picture, and is called any of a point cloud image, a frame, and a picture.
[0022] The point cloud video encoder 10002 according to the embodiment encodes the secured point cloud video data. The point cloud video encoder 10002 encodes the point cloud video data based on point cloud compression coding. The point cloud compression coding according to the embodiment includes Geometry-based Point Cloud Compression (G-PCC) coding and / or Video-based Point Cloud Compression (V-PCC) coding, or next-generation coding. Note that the point cloud compression coding according to the embodiment is not limited to the above-described embodiments. The point cloud video encoder 10002 can output a bitstream including encoded point cloud video data. The bitstream includes not only the encoded point cloud video data but also signaling information related to the encoding of the point cloud video data.
[0023] According to an embodiment, the transmitter 10003 transmits a bitstream including encoded point cloud video data. According to an embodiment, the bitstream is encapsulated into a file or a segment (e.g., a streaming segment) and transmitted via various networks such as a broadcast network and / or a broadband network. Although not shown, the transmitting device 10000 includes an encapsulation unit (or encapsulation module) that performs the encapsulation operation. In an embodiment, the encapsulation unit is included in the transmitter 10003. According to an embodiment, the file or segment is transmitted to the receiving device 10004 via a network or stored on a digital storage medium (e.g., USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.). According to an embodiment, the transmitter 10003 can communicate with the receiving device 10004 (or receiver 10005) via wired or wireless communication via a network such as 4G, 5G, or 6G. Furthermore, the transmitter 10003 can perform necessary data processing operations via a network system (e.g., a communication network system such as 4G, 5G, or 6G). The transmitting device 10000 can also transmit encapsulated data on an on-demand basis.
[0024] The receiving device 10004 according to the embodiment includes a receiver 10005, a point cloud video decoder 10006, and / or a renderer 10007. In the embodiment, the receiving device 10004 includes a device, robot, vehicle, AR / VR / XR device, mobile device, home appliance, IoT (Internet of Things) device, AI device / server, etc. that communicates with a base station and / or other wireless device using a wireless connection technology (e.g., 5G NR (New RAT), LTE (Long Term Evolution)).
[0025] The receiver 10005 according to the embodiment receives a bitstream containing point cloud video data or a file / segment in which the bitstream is encapsulated from a network or storage medium. The receiver 10005 performs data processing operations required by a network system (e.g., a communication network system such as 4G, 5G, or 6G). The receiver 10005 according to the embodiment decapsulates the received file / segment and outputs a bitstream. In addition, in the embodiment, the receiver 10005 includes a decapsulation unit (or decapsulation module) for performing the decapsulation operation. The decapsulation unit is embodied as an element (or component) separate from the receiver 10005.
[0026] The point cloud video decoder 10006 decodes a bitstream containing point cloud video data. The point cloud video decoder 10006 can decode the point cloud video data in the manner in which it was encoded (e.g., the reverse process of the operation of the point cloud video encoder 10002). Therefore, the point cloud video decoder 10006 can decode the point cloud video data by performing point cloud reconstruction coding, which is the reverse process of point cloud compression. Point cloud reconstruction coding includes G-PCC coding.
[0027] The renderer 10007 renders the decoded point cloud video data. The renderer 10007 renders not only the point cloud video data but also the audio data to output point cloud content. In an embodiment, the renderer 10007 includes a display for displaying the point cloud content. In an embodiment, the display is not included in the renderer 10007, but is embodied as a separate device or component.
[0028] In the drawing, dotted arrows indicate the transmission path of feedback information obtained by the receiving device 10004. The feedback information is information for reflecting interaction with a user consuming point cloud content and includes user information (e.g., head orientation information, viewport information, etc.). In particular, if the point cloud content is for a service that requires interaction with a user (e.g., an autonomous driving service), the feedback information may be transmitted to a content transmitting side (e.g., the transmitting device 10000) and / or a service provider. In an embodiment, the feedback information may be used not only by the transmitting device 10000 but also by the receiving device 10004, or may not be provided.
[0029] According to an embodiment, head orientation information is information regarding the position, direction, angle, movement, etc. of the user's head. According to an embodiment, the receiving device 10004 calculates viewport information based on the head orientation information. The viewport information is information regarding the area of the point cloud video viewed by the user. The viewpoint is the point at which the user views the point cloud video and refers to the center of the viewport area. That is, the viewport is an area centered on the viewpoint, and the size, shape, etc. of the area are determined by the FOV (Field of View). Therefore, the receiving device 10004 can extract viewport information based on the vertical or horizontal FOV supported by the device in addition to the head orientation information. The receiving device 10004 also performs gaze analysis to determine the user's point cloud consumption method, the point cloud video area the user is gazing at, the gaze time, etc. In an embodiment, the receiving device 10004 can transmit feedback information including the results of the gaze analysis to the transmitting device 10000. According to an embodiment, the feedback information is obtained during the rendering and / or display process. According to this embodiment, feedback information is obtained by one or more sensors included in the receiving device 10004. In another embodiment, feedback information is obtained by the renderer 10007 or another external element (or device, component, etc.). The dotted lines in FIG. 1 indicate the transmission process of feedback information obtained by the renderer 10007. The point cloud content providing system processes (encodes / decodes) point cloud data based on the feedback information. Therefore, the point cloud video data decoder 10006 can perform a decoding operation based on the feedback information. Furthermore, the receiving device 10004 can transmit the feedback information to the transmitting device 10000. The transmitting device 10000 (or point cloud video data encoder 10002) can perform an encoding operation based on the feedback information.Therefore, the point cloud content providing system does not process (encode / decode) all point cloud data, but can efficiently process necessary data (e.g., point cloud data corresponding to the user's head position) based on feedback information to provide point cloud content to the user.
[0030] In an embodiment, the sending device 10000 may be referred to as an encoder, a sending device, a transmitter, etc., and the receiving device 10004 may be referred to as a decoder, a receiving device, a receiver, etc.
[0031] 1 according to an embodiment (processed through a series of steps of acquisition / encoding / transmission / decoding / rendering), the point cloud data may also be referred to as point cloud content data or point cloud video data. In an embodiment, the point cloud content data may be used as a concept including metadata or signaling information related to the point cloud data.
[0032] The elements of the point cloud content providing system shown in FIG. 1 may be implemented in hardware, software, a processor, and / or a combination thereof.
[0033] FIG. 2 is a block diagram illustrating the operation of providing point cloud content according to an embodiment.
[0034] Figure 2 is a block diagram showing the operation of the point cloud content providing system described in Figure 1. As described above, the point cloud content providing system processes point cloud data based on point cloud compression coding (e.g., G-PCC).
[0035] A point cloud content providing system (e.g., a point cloud transmitting device 10000 or a point cloud video acquiring unit 10001) according to an embodiment acquires a point cloud video (20000). The point cloud video is represented by a point cloud belonging to a coordinate system representing a three-dimensional space. The point cloud video according to an embodiment includes a Ply (Polygon File format or the Stanford Triangle format) file. If the point cloud video has one or more frames, the acquired point cloud video includes one or more Ply files. A Ply file includes point cloud data such as the geometry and / or attributes of points. The geometry includes the position of the point. The position of each point is expressed by parameters (e.g., values on the X-axis, Y-axis, and Z-axis) indicating a three-dimensional coordinate system (e.g., a coordinate system consisting of X, Y, and Z axes). The attributes include the characteristics of the point (e.g., texture information of each point, color (YCbCr or RGB), reflectance (r), transparency, etc.). A point has one or more characteristics (or attributes). For example, one point may have one characteristic of hue, or two characteristics of hue and reflectance. In the embodiment, geometry may be referred to as position, geometry information, geometry data, etc., and characteristics may be referred to as characteristics, characteristic information, characteristic data, etc. In addition, the point cloud content providing system (e.g., the point cloud transmitting device 10000 or the point cloud video acquiring unit 10001) may acquire point cloud data from information related to the point cloud video acquiring process (e.g., depth information, color information, etc.).
[0036] A point cloud content providing system (e.g., a transmitting device 10000 or a point cloud video encoder 10002) according to an embodiment encodes point cloud data (20001). The point cloud content providing system encodes point cloud data based on point cloud compression coding. As described above, point cloud data includes the geometry and attributes of points. Therefore, the point cloud content providing system can perform geometry coding to encode the geometry and output a geometry bitstream. The point cloud content providing system can perform attribute coding to encode the attributes and output a attribute bitstream. In an embodiment, the point cloud content providing system can perform attribute coding based on the geometry coding. The geometry bitstream and the attribute bitstream according to an embodiment are multiplexed and output as a single bitstream. The bitstream according to an embodiment further includes signaling information related to the geometry coding and the attribute coding.
[0037] A point cloud content providing system (e.g., transmitting device 10000 or transmitter 10003) according to an embodiment transmits encoded point cloud data (20002). As described in FIG. 1, the encoded point cloud data is represented by a geometry bitstream and a feature bitstream. The encoded point cloud data is transmitted in the form of a bitstream together with signaling information related to the encoding of the point cloud data (e.g., signaling information related to geometry encoding and feature encoding). The point cloud content providing system encapsulates the bitstream for transmitting the encoded point cloud data and transmits it in the form of a file or segment.
[0038] A point cloud content providing system (e.g., receiving device 10004 or receiver 10005) according to an embodiment receives a bitstream including encoded point cloud data, and can demultiplex the bitstream.
[0039] The point cloud content providing system (e.g., the receiving device 10004 or the point cloud video decoder 10005) decodes the encoded point cloud data (e.g., geometry bitstream, attribute bitstream) transmitted in a bitstream. The point cloud content providing system (e.g., the receiving device 10004 or the point cloud video decoder 10005) decodes the point cloud video data based on signaling information related to the encoding of the point cloud video data included in the bitstream. The point cloud content providing system (e.g., the receiving device 10004 or the point cloud video decoder 10005) decodes the geometry bitstream to restore the position (geometry) of the point. The point cloud content providing system decodes the attribute bitstream based on the restored geometry to restore the attribute of the point. The point cloud content providing system (e.g., the receiving device 10004 or the point cloud video decoder 10005) restores the point cloud video based on the position according to the restored geometry and the decoded attribute.
[0040] A point cloud content providing system (e.g., a receiving device 10004 or a renderer 10007) according to an embodiment renders the decoded point cloud data (20004). The point cloud content providing system (e.g., a receiving device 10004 or a renderer 10007) renders the geometry and characteristics decoded during the decoding process using various rendering methods. Points of the point cloud content are rendered as fixed points with a certain thickness, cubes with a predetermined minimum size centered at the position of the fixed points, or circles centered at the position of the fixed points. All or part of the region of the rendered point cloud content is provided to a user via a display (e.g., a VR / AR display, a general display, etc.).
[0041] A point cloud content providing system (e.g., receiving device 10004) according to the embodiment may obtain feedback information (20005). The point cloud content providing system encodes and / or decodes point cloud data based on the feedback information. The feedback information and operation of the point cloud content providing system according to the embodiment are the same as the feedback information and operation described in FIG. 1, so a detailed description will be omitted.
[0042] FIG. 3 illustrates an example of a point cloud video capturing process according to an embodiment.
[0043] FIG. 3 illustrates an example of a point cloud video capture process in the point cloud content providing system described in FIGS.
[0044] Point cloud content includes point cloud video (images and / or video) showing objects and / or environments located in various three-dimensional spaces (e.g., a three-dimensional space showing a real environment, a three-dimensional space showing a virtual environment, etc.). Accordingly, a point cloud content providing system according to an embodiment captures point cloud video using one or more cameras (e.g., an infrared camera capable of obtaining depth information, an RGB camera capable of extracting color information corresponding to the depth information, etc.), projectors (e.g., an infrared pattern projector for obtaining depth information), LiDAR, etc. to generate point cloud content. The point cloud content providing system according to an embodiment extracts a geometric form composed of points in three-dimensional space from the depth information and extracts characteristics of each point from the color information to obtain point cloud data. Images and / or video according to an embodiment are captured based on either an inward-facing approach or an outward-facing approach.
[0045] The left side of Figure 3 shows the inward-facing method. The inward-facing method is a method in which one or more cameras (or camera sensors) positioned around a central object capture the central object. The inward-facing method is used to generate point cloud content that provides the user with a 360-degree image of the core object (e.g., VR / AR content that provides the user with a 360-degree image of an object (e.g., a core object such as a character, player, item, or actor)).
[0046] The right side of Figure 3 shows an outward-facing approach. The outward-facing approach is a method in which one or more cameras (or camera sensors) positioned around a central object capture the environment of the central object, which is not the central object. The outward-facing approach is used to generate point cloud content to provide the surrounding environment from a user's perspective (e.g., content showing the external environment provided to a user of an autonomous vehicle).
[0047] As shown in the figure, point cloud content is generated based on the capture operation of one or more cameras. In this case, since each camera has a different coordinate system, the point cloud content providing system calibrates one or more cameras to set a global coordinate system before the capture operation. The point cloud content providing system also generates point cloud content by combining images and / or video captured using the above capture method with an arbitrary image and / or video. When generating point cloud content representing a virtual space, the point cloud content providing system does not perform the capture operation described in FIG. 3. The point cloud content providing system according to the embodiment can also perform post-processing on the captured images and / or video. That is, the point cloud content providing system can remove unwanted areas (e.g., background) or fill in any spatial holes by recognizing the space where the captured images and / or video are connected.
[0048] The point cloud content providing system can also generate a single point cloud content by performing coordinate system transformation on points in the point cloud video acquired from each camera. The point cloud content providing system performs coordinate system transformation on points based on the position coordinates of each camera. This allows the point cloud content providing system to generate content showing a single wide area or point cloud content with a high point density.
[0049] FIG. 4 is a diagram illustrating an example of a point cloud encoder according to an embodiment.
[0050] 4 shows an example of the point cloud video encoder 10002 of FIG. 1. The point cloud encoder performs encoding by reconstructing point cloud data (e.g., point positions and / or characteristics) to adjust the quality of the point cloud content (e.g., lossless, lossy, near-lossless) depending on the network conditions or application. If the overall size of the point cloud content is large (e.g., point cloud content of 60 Gbps for 30 fps), the point cloud content providing system cannot stream the content in real time. Therefore, the point cloud content providing system can reconstruct the point cloud content based on a maximum target bitrate to provide it according to the network environment, etc.
[0051] As shown in Figures 1 and 2, a point cloud encoder can perform geometry encoding and feature encoding, where geometry encoding is performed before feature encoding.
[0052] The point cloud encoder according to the embodiment includes a coordinate system transformation unit (Transformation Coordinates) 40000, a quantization unit (Quantize and Remove Points (Voxelize)) 40001, an octree analysis unit (Analyze Octree) 40002, a surface approximation analysis unit (Analyze Surface Approximation) 40003, an arithmetic encoder (Arithmetic Encode) 40004, a geometry reconstruction unit (Reconstruct Geometry) 40005, a color transformation unit (Transform Colors) 40006, a characteristic transformation unit (Transfer Attributes) 40007, a RAHT transformation unit 40008, an LOD generation unit (Generated LOD) 40009, a lift transformation unit (Lifting) 40010, a coefficient quantization unit (Quantize Coefficients) 40011 and / or an arithmetic encoder (Arithmetic Encode) 40012.
[0053] The coordinate system transformation unit 40000, the quantization unit 40001, the octree analysis unit 40002, the surface approximation analysis unit 40003, the arithmetic encoder 40004, and the geometry reconstruction unit 40005 can perform geometry coding. Geometry coding according to the embodiment includes octree geometry coding, direct coding, trisoup geometry encoding, and entropy coding. Direct coding and trisoup geometry encoding are applied selectively or in combination. Note that geometry coding is not limited to the above examples.
[0054] As shown in the figure, a coordinate system conversion unit 40000 according to an embodiment receives a position and converts it into a coordinate system. For example, the position is converted into position information in a three-dimensional space (e.g., a three-dimensional space expressed in an XYZ coordinate system). The position information in the three-dimensional space according to an embodiment is also referred to as geometry information.
[0055] The quantizer 40001 according to the embodiment quantizes geometry. For example, the quantizer 40001 quantizes points based on the minimum position value of all points (e.g., the minimum value on each axis for the X, Y, and Z axes). The quantizer 40001 performs a quantization operation by multiplying the difference between the minimum position value and the position value of each point by a predetermined quantization scale value and then rounding down or up to find the nearest integer value. Therefore, one or more points may have the same quantized position (or position value). The quantizer 40001 according to the embodiment performs voxelization based on the quantized position to reconstruct the quantized points. Just as the smallest unit containing 2D image / video information is a pixel, points in the point cloud content (or 3D point cloud video) according to the embodiment are included in one or more voxels. A voxel is a combination of the words volume and pixel and refers to a 3D cubic space generated when a 3D space is divided into units (unit=1.0) based on the axes (e.g., X-axis, Y-axis, Z-axis) that represent the 3D space. The quantization unit 40001 can match a group of points in the 3D space with voxels. In some embodiments, a voxel may contain only one point. In some embodiments, a voxel may contain one or more points. To represent a voxel as a point, the center of the voxel may be set based on the positions of one or more points contained in the voxel. In this case, the characteristics of all points contained in a voxel are combined and assigned to the voxel.
[0056] The octree analysis unit 40002 according to the embodiment performs octree geometry coding (or octree coding) to represent voxels in an octree structure, which represents points matched to voxels based on an octet structure.
[0057] The surface approximation analysis unit 40003 according to the embodiment analyzes and approximates the octree. The octree analysis and approximation according to the embodiment is a process of analyzing a region containing a large number of points to voxelize it in order to efficiently provide the octree and voxelization.
[0058] The arithmetic encoder 40004 according to the embodiment entropy encodes the octree and / or the approximated octree. For example, the encoding method includes an arithmetic encoding method. As a result of the encoding, a geometry bitstream is generated.
[0059] The color transform unit 40006, the feature transform unit 40007, the RAHT transform unit 40008, the LOD generation unit 40009, the lift transform unit 40010, the coefficient quantization unit 40011, and / or the arithmetic encoder 40012 perform feature coding. As described above, one point has one or more features. Feature coding according to the embodiment is applied equally to all features of one point. However, if one feature (e.g., hue) includes one or more elements, independent feature coding is applied to each element. Feature coding according to the embodiment includes color transform coding, feature transform coding, RAHT (Region Adaptive Hierarchical Transform) coding, Interpolarization-based hierarchical nearest-neighbor prediction-Prediction Transform (Interpolarization-based hierarchical nearest-neighbor prediction with an update / lifting step (Lifting Transform)) coding. Depending on the point cloud content, the above-mentioned RAHT coding, predictive transform coding, and lift transform coding may be selectively used, or a combination of one or more coding methods may be used. Furthermore, the feature coding according to the embodiment is not limited to the above examples.
[0060] The color converter 40006 according to the embodiment performs color conversion coding to convert the color values (or textures) included in the attributes. For example, the color converter 40006 converts the format of the hue information (e.g., converts from RGB to YCbCr). The operation of the color converter 40006 according to the embodiment is optionally applied depending on the color values included in the attributes.
[0061] The geometry reconstruction unit 40005 according to the embodiment reconstructs (restores) the octree and / or the approximated octree. The geometry reconstruction unit 40005 reconstructs the octree / voxel based on the result of analyzing the distribution of points. The reconstructed octree / voxel is also called a reconstructed geometry (or restored geometry).
[0062] According to an embodiment, the feature converter 40007 performs feature conversion, converting features based on a position where geometry encoding has not been performed and / or reconstructed geometry. As described above, because features depend on geometry, the feature converter 40007 can convert features based on reconstructed geometry information. For example, the feature converter 40007 can convert features of points included in a voxel based on their position values. As described above, if the center point of a voxel is set based on the positions of one or more points included in the voxel, the feature converter 40007 converts features of one or more points. If trisoup geometry encoding is performed, the feature converter 40007 can convert features based on the trisoup geometry encoding.
[0063] The feature conversion unit 40007 performs feature conversion by calculating the average value of the features or feature values (e.g., hue or reflectance of each point) of adjacent points within a specific position / radius from the position (or position value) of the center point of each voxel. When calculating the average value, the feature conversion unit 40007 applies a weight based on the distance from the center point to each point. Therefore, each voxel has a position and a calculated feature (or feature value).
[0064] The feature conversion unit 40007 searches for neighboring points within a specific position / radius from the center point of each voxel based on a KD tree or Moulton code. A KD tree supports a data structure that manages points based on their position, enabling a fast Nearest Neighbor Search (NNS) using a binary search tree. A Moulton code is generated by mixing bits, representing the coordinate values (e.g., (x, y, z)) that indicate the 3D position of all points. For example, if the coordinate value indicating the point's position is (5, 9, 1), the bit values of the coordinate value are (0101, 1001, 0001). Mixing the bit values in the order of z, y, and x according to the bit index results in 010001000111. This value is expressed in decimal as 1095. In other words, the Moulton code value of a point with coordinate values (5, 9, 1) is 1095. The feature conversion unit 40007 aligns points based on the Moulton code value and performs nearest neighbor search (NNS) using a depth-first traversal process. After the feature conversion operation, if nearest neighbor search (NNS) is required in other conversion processes for feature coding, a KD tree or Moulton code is used.
[0065] As shown, the transformed attributes are input to a RAHT transformer 40008 and / or an LOD generator 40009 .
[0066] The RAHT converter 40008 according to the embodiment performs RAHT coding to predict feature information based on the reconstructed geometry information. For example, the RAHT converter 40008 can predict feature information of a node at a higher level of the octree based on feature information associated with a node at a lower level of the octree.
[0067] According to an embodiment, the LOD generator 40009 generates a Level of Detail (LOD) for predictive coding. The LOD according to an embodiment indicates the level of detail of the point cloud content, where a smaller LOD value indicates less detail of the point cloud content and a larger LOD value indicates more detail of the point cloud content. Points can be classified according to the LOD.
[0068] The lift transform unit 40010 according to the embodiment performs lift transform coding, which transforms the characteristics of the point cloud based on weights. As described above, the lift transform coding is selectively applied.
[0069] The coefficient quantization unit 40011 according to the embodiment quantizes the feature-coded feature based on the coefficients.
[0070] The arithmetic encoder 40012 according to the embodiment encodes the quantized characteristics based on arithmetic coding.
[0071] The elements of the point cloud encoder of FIG. 4 may be implemented in hardware, software, firmware, or a combination thereof, including one or more processors or integrated circuits (not shown) configured to communicate with one or more memories included in the point cloud providing device. The one or more processors may perform any one of the operations and / or functions of the elements of the point cloud encoder of FIG. 4 described above. The one or more processors may also operate or execute a software program and / or set of instructions to perform the operations and / or functions of the elements of the point cloud encoder of FIG. 4. According to embodiments, the one or more memories may include high-speed random access memory or non-volatile memory (e.g., one or more magnetic disk storage devices, flash memory devices, or other non-volatile solid-state memory devices).
[0072] FIG. 5 is a diagram showing an example of a voxel according to the embodiment.
[0073] FIG. 5 shows voxels located in a three-dimensional space expressed by a coordinate system consisting of three axes: X, Y, and Z. As shown in FIG. 4, a point cloud encoder (e.g., quantization unit 40001) performs voxelization. A voxel is a three-dimensional cubic space that is generated when the three-dimensional space is divided into units (unit=1.0) based on the axes (e.g., X, Y, and Z) that represent the three-dimensional space. FIG. 5 shows two extreme points (0,0,0) and (2 d , 2 d , 2 d) is an example of a voxel generated by an octree structure that recursively subdivides a bounding box (cubical axis-aligned bounding box) defined by the cubic axis-aligned bounding box (B). One voxel contains at least one point. The spatial coordinates of a voxel can be estimated from its positional relationship with other voxel groups. As mentioned above, a voxel has characteristics (such as color or reflectance) just like a pixel in a 2D image / video. A detailed explanation of voxels is omitted here as it has been explained in FIG. 4.
[0074] FIG. 6 is a diagram illustrating an example of an octree and occupancy code according to an embodiment.
[0075] As shown in Figures 1 to 4, a point cloud content providing system (point cloud video encoder 10002) or a point cloud encoder (e.g., octree analysis unit 40002) performs octree geometry coding (or octree coding) based on an octree structure to efficiently manage the area and / or position of voxels.
[0076] The upper part of Figure 6 shows an octree structure. The three-dimensional space of the point cloud content according to the embodiment is represented by the axes of a coordinate system (e.g., X-axis, Y-axis, Z-axis). The octree structure is defined by two extreme points (0,0,0) and (2 d , 2 d , 2 d ) is generated by recursively subdividing the cubic axis-aligned bounding box defined by the point cloud content (or point cloud video). 2d is set to the value that constitutes the smallest bounding box that encloses all points in the point cloud content (or point cloud video). d indicates the depth of the octree. The d value is determined by the following formula: In the following formula, (x int n , y int n , z int n) indicates the quantized position (or position value) of the point.
[0077] JPEG0007775351000001.jpg9121
[0078] As shown in the upper center of Figure 6, the entire 3D space is divided into eight spaces through division. Each divided space is represented by a cube with six faces. As shown in the upper right of Figure 6, each of the eight spaces is again divided by the axes of the coordinate system (e.g., X-axis, Y-axis, Z-axis). Therefore, each space is again divided into eight smaller spaces. The divided smaller spaces are also represented by cubes with six faces. This division method is applied until the leaf nodes of the octree become voxels.
[0079] The bottom of Figure 6 shows an occupancy code for an octree. An occupancy code for an octree is generated to indicate whether each of the eight subspaces generated by dividing a space contains at least one point. Therefore, one occupancy code is represented by eight child nodes. Each child node indicates the occupancy of the divided space and has a 1-bit value. Therefore, the occupancy code is represented by an 8-bit code. That is, if the space corresponding to a child node contains at least one point, the corresponding node has a value of 1. If the space corresponding to a node does not contain any points (is empty), the corresponding node has a value of 0. The occupancy code shown in Figure 6 is 00100001, which indicates that the spaces corresponding to the third and eighth child nodes of the eight child nodes each contain at least one point. As shown in the figure, the third and eighth child nodes each have eight child nodes, and each child node is represented by an 8-bit occupancy code. In the figure, the occupied code of the third child node is 10000111, and the occupied code of the eighth child node is 01001111. A point cloud encoder (e.g., arithmetic encoder 40004) according to an embodiment can entropy encode the occupied code. To improve compression efficiency, the point cloud encoder can also intra / inter-code the occupied code. A receiving device (e.g., receiving device 10004 or point cloud video decoder 10006) according to an embodiment reconstructs an octree based on the occupied code.
[0080] A point cloud encoder according to an embodiment (e.g., the point cloud encoder or octree analyzer 40002 in FIG. 4) performs voxelization and octree coding to store point positions. However, since points in a three-dimensional space are not always uniformly distributed, there may be certain areas where there are not many points. Therefore, performing voxelization on the entire three-dimensional space is inefficient. For example, if there are almost no points in a certain area, there is no need to perform voxelization on that area.
[0081] Therefore, the point cloud encoder according to the embodiment does not perform voxelization for the specific region (or nodes excluding leaf nodes of the octree) described above, but performs direct coding, which directly codes the positions of points included in the specific region. The coordinates of direct coding points according to the embodiment are called Direct Coding Mode (DCM). The point cloud encoder according to the embodiment can also perform trisoup geometry encoding, which reconstructs the positions of points within a specific region (or node) based on voxels based on a surface model. Trisoup geometry encoding is a geometry encoding that represents an object as a series of triangle meshes. Therefore, the point cloud decoder can generate a point cloud from the mesh surface. Direct coding and trisoup geometry encoding according to the embodiment can be performed selectively. Furthermore, direct coding and trisoup geometry encoding according to the embodiment can be combined with octree geometry coding (or octree coding).
[0082] To perform direct coding, the option to use direct mode for applying direct coding must be activated, and the node to which direct coding is applied must not be a leaf node, but a specific node must have points below a threshold. Also, the total number of points to be subjected to direct coding must not exceed a predetermined threshold. If the above conditions are met, the point cloud encoder (or operation encoder 40004) according to the embodiment can entropy code the positions (or position values) of the points.
[0083] A point cloud encoder (e.g., the surface approximation analysis unit 40003) according to an embodiment can determine a specific level of the octree (if the level is smaller than the depth d of the octree) and perform trisoup geometry encoding (trisoup mode) from that level, which uses a surface model to reconstruct the positions of points within the node area based on voxels. The point cloud encoder according to an embodiment can specify the level to which trisoup geometry encoding is applied. For example, if the specified level is the same as the depth of the octree, the point cloud encoder does not operate in trisoup mode. In other words, the point cloud encoder according to an embodiment can operate in trisoup mode only when the specified level is smaller than the depth value of the octree. The 3D cubic area of a node at a specified level according to an embodiment is called a block. One block contains one or more voxels. A block or voxel can also correspond to a brick. Within each block, geometry is represented as a surface. According to an embodiment, a surface can intersect each edge of the block at most once.
[0084] Since one block has 12 edges, there are at least 12 intersections within one block. Each intersection is called a vertex. A vertex along an edge is detected if there is at least one occupied voxel adjacent to that edge among all blocks that share the edge. In this embodiment, an occupied voxel refers to a voxel that contains a point. The position of a vertex detected along an edge is the average position along the edge of all voxels adjacent to that edge among all blocks that share the edge.
[0085] When a vertex is detected, the point cloud encoder according to the embodiment can entropy code the edge start point (x, y, z), the edge direction vector (Δx, Δy, Δz), and the vertex position value (relative position value within the edge). When trisoup geometry encoding is applied, the point cloud encoder according to the embodiment (e.g., the geometry reconstruction unit 40005) can perform triangle reconstruction, up-sampling, and voxelization processes to generate a restored geometry (reconstructed geometry).
[0086] JPEG0007775351000002.jpg33170
[0087] JPEG0007775351000003.jpg22150
[0088] The minimum value of the added values is found, and a projection process is performed along the axis where the minimum value is. For example, if the x element is the smallest, each vertex is projected onto the x-axis based on the center of the block, and then onto the (y,z) plane. If the value obtained by projecting onto the (y,z) plane is (ai,bi), the θ value is found using atan2(bi,ai), and the vertices are aligned based on the θ value. The following table shows the vertex combinations required to create a triangle depending on the number of vertices. Vertices are aligned in order from 1 to n. The following table shows that for four vertices, two triangles are formed by combining the vertices. The first triangle is formed by the 1st, 2nd, and 3rd vertices of the aligned vertices, and the second triangle is formed by the 3rd, 4th, and 1st vertices of the aligned vertices.
[0089] Table.Triangles formed from vertices ordered 1
[0090] [Table 1]
[0091] The upsampling process is performed to add intermediate points along the edges of triangles for voxelization. The additional points are generated based on the upsampling factor and the block width. These additional points are called refined vertices. A point cloud encoder according to an embodiment can voxelize the refined vertices. The point cloud encoder can also perform feature encoding based on the voxelized positions (or position values).
[0092] FIG. 7 is a diagram illustrating an example of an adjacent node pattern according to the embodiment.
[0093] To increase the compression efficiency of point cloud videos, the point cloud encoder according to the embodiment performs entropy coding based on context adaptive arithmetic coding.
[0094] As described with reference to FIGS. 1 to 6, a point cloud content providing system or a point cloud encoder (e.g., point cloud video encoder 10002, point cloud encoder or arithmetic encoder 40004 in FIG. 4) can immediately entropy code the occupied code. The point cloud content providing system or the point cloud encoder can also perform entropy coding (intra coding) based on the occupied code of the current node and the occupied rate of neighboring nodes, or can perform entropy coding (inter coding) based on the occupied code of a previous frame. A frame according to the present embodiment refers to a collection of point cloud videos generated at the same time. The compression efficiency of intra coding / inter coding according to the present embodiment varies depending on the number of neighboring nodes to reference. Although the complexity increases as the number of bits increases, the compression efficiency can be improved by focusing on one side. For example, a 3-bit context requires eight coding methods, which is equal to the cube of two. The separately coded portion affects the complexity of the implementation. Therefore, it is necessary to balance compression efficiency and complexity at an appropriate level.
[0095] FIG. 7 shows a process for determining an occupancy pattern based on the occupancy of neighboring nodes. The point cloud encoder according to the embodiment obtains a neighbor pattern value by determining the occupancy of neighboring nodes for each node in an occupancy tree. The neighboring node pattern is used to infer the occupancy pattern of the corresponding node. The left side of FIG. 7 shows a cube corresponding to the node (the cube located in the middle) and six cubes (neighboring nodes) that share at least one side with the corresponding cube. The nodes shown are nodes at the same depth. The numbers shown indicate the weights (1, 2, 4, 8, 16, 32, etc.) associated with each of the six nodes. Each weight is assigned in order according to the position of the neighboring node.
[0096] The right side of FIG. 7 shows the adjacent node pattern value. The adjacent node pattern value is the sum of values multiplied by the weight values of occupied adjacent nodes (adjacent nodes with points). Therefore, the adjacent node pattern value ranges from 0 to 63. An adjacent node pattern value of 0 means that there are no nodes with points (occupied nodes) among the adjacent nodes of the corresponding node. An adjacent node pattern value of 63 means that all adjacent nodes are occupied nodes. As shown in the figure, adjacent nodes assigned weight values of 1, 2, 4, and 8 are occupied nodes, so the adjacent node pattern value is 15, which is the sum of 1, 2, 4, and 8. The point cloud encoder can perform coding according to the adjacent node pattern value (e.g., if the adjacent node pattern value is 63, 64 coding is performed). In an embodiment, the point cloud encoder can reduce coding complexity by modifying the adjacent node pattern value (e.g., based on a table that modifies 64 to 10 or 6).
[0097] FIG. 8 is a diagram illustrating an example of a point configuration for each LOD according to the embodiment.
[0098] As described in Figures 1 to 7, before feature encoding, the coded geometry is reconstructed (restored). When direct coding is applied, the geometry reconstruction operation involves changing the placement of the direct coded points (e.g., placing the direct coded points in front of the point cloud data). When trisoup geometry encoding is applied, the geometry reconstruction process involves the processes of triangulation, upsampling, and voxelization. Since features are dependent on the geometry, feature encoding is performed based on the reconstructed geometry.
[0099] A point cloud encoder (e.g., LOD generator 40009) reorganizes points by LOD. The diagram shows point cloud content corresponding to LOD. The left side of the diagram shows the original point cloud content. The second from the left in the diagram shows the distribution of points with the lowest LOD, and the rightmost side shows the distribution of points with the highest LOD. That is, points with the lowest LOD have a sparse distribution, and points with the highest LOD have a fine distribution. That is, as the LOD increases along the arrow direction shown at the bottom of the diagram, the spacing (or distance) between points becomes shorter.
[0100] FIG. 9 is a diagram illustrating an example of a point configuration for each LOD according to the embodiment.
[0101] As described in Figures 1 to 8, a point cloud content providing system or point cloud encoder (e.g., point cloud video encoder 10002, point cloud encoder or LOD generator 40009 in Figure 4) generates LOD. LOD is generated by rearranging points into a set of refinement levels according to a set LOD distance value (or a set of Euclidean distances). The LOD generation process is performed not only by the point cloud encoder but also by the point cloud decoder.
[0102] The upper part of Figure 9 shows an example of points (P0 to P9) of point cloud content distributed in 3D space. The original order in Figure 9 shows the order of points P0 to P9 before LOD generation. The LOD-based order in Figure 9 shows the order of points after LOD generation. Points are re-sorted for each LOD. Also, higher LODs include points belonging to lower LODs. As shown in Figure 9, LOD0 includes P0, P5, P4, and P2. LOD1 includes points of LOD0, P1, P6, and P3. LOD2 includes points of LOD0, points of LOD1, and P9, P8, and P7.
[0103] As described in FIG. 4, the point cloud encoder according to the embodiment can selectively or in combination perform predictive transform coding, lift transform coding, and RAHT transform coding.
[0104] The point cloud encoder according to the embodiment generates a predictor for each point and performs predictive transformation coding to set a prediction characteristic (or a prediction characteristic value) for each point. That is, N predictors are generated for N points. The predictor according to the embodiment can calculate a weight (=1 / distance) based on the LOD value of each point, index information for neighboring points within a distance set for each LOD, and the distance value to the neighboring point.
[0105] According to an embodiment, the predicted feature (or feature value) is set as the average value of the feature (or feature value, e.g., hue, reflectance, etc.) of neighboring points set in the predictor of each point multiplied by a weight (or weight value) calculated based on the distance to each neighboring point. A point cloud encoder (e.g., coefficient quantization unit 40011) according to an embodiment may quantize and inverse quantize residuals (also called residual features, residual feature values, feature prediction residual values, etc.) obtained by subtracting the predicted feature (feature value) from the feature (feature value) of each point. The quantization process is shown in the following table.
[0106] Attribute prediction residuals quantization pseudo code
[0107] [Table 2]
[0108] Attribute prediction residuals inverse quantization pseudo Code
[0109] [Table 3]
[0110] The point cloud encoder (e.g., the arithmetic encoder 40012) according to the embodiment performs entropy coding on the quantized and dequantized residual values as described above if there are adjacent points in the predictor of each point. If there are no adjacent points in the predictor of each point, the point cloud encoder (e.g., the arithmetic encoder 40012) according to the embodiment does not perform the above process and performs entropy coding on the characteristics of the corresponding point.
[0111] The point cloud encoder (e.g., lift transform unit 40010) according to the embodiment generates a predictor for each point, sets the calculated LOD in the predictor, registers neighboring points, and sets weights according to the distance to the neighboring points to perform lift transform coding. The lift transform coding according to the embodiment is similar to the above-mentioned point cloud transform coding, but differs in that weights are cumulatively applied to feature values. The process of cumulatively applying weights to feature values according to the embodiment is as follows.
[0112] 1) Create an array QW (QuantizationWeight) that stores the weight value of each point. The initial value of all elements of QW is 1.0. The QW value of the predictor index of the adjacent node registered in the predictor is multiplied by the weight value of the predictor of the current point and added.
[0113] 2) Lift prediction process: To calculate the predicted attribute value, the attribute value of the point is multiplied by the weight and subtracted from the existing attribute value.
[0114] 3) Create temporary arrays called updateweight and update, and initialize the temporary arrays to 0.
[0115] 4) The calculated weights for all predictors are multiplied by the weights stored in the QW corresponding to the predictor index, and the resulting weights are accumulated and summed in the update weight array as the index of the adjacent node. The update array is accumulated and summed by multiplying the calculated weights by the characteristic values of the index of the adjacent node.
[0116] 5) Lift update process: For every predictor, divide the feature value in the update array by the weight value in the update weight array of the predictor index, and add the result to the existing feature value again.
[0117] 6) For all predictors, predicted feature values are calculated by multiplying the feature values updated in the lift update process by the weights updated (stored in the QW) in the lift prediction process. According to an embodiment, a point cloud encoder (e.g., coefficient quantizer 40011) quantizes the predicted feature values. Furthermore, a point cloud encoder (e.g., arithmetic encoder 40012) entropy codes the quantized feature values.
[0118] A point cloud encoder (e.g., the RAHT transform unit 40008) according to an embodiment performs RAHT transform coding, which predicts the characteristics of higher-level nodes using characteristics associated with lower-level nodes in an octree. RAHT transform coding is an example of characteristic intra-coding using octree backward scanning. The point cloud encoder according to an embodiment scans the entire region from a voxel, and at each step, repeats a merging process up to the root node while combining the voxels into larger blocks. The merging process according to an embodiment is performed only on occupied nodes. The merging process is not performed on empty nodes, but is performed on the nodes immediately above the empty nodes.
[0119] The following equation shows the RAHT transformation matrix: JPEG0007775351000007.jpg1023 denotes the average quality value of the voxels at level l. JPEG0007775351000008.jpg922 JPEG0007775351000009.jpg1031 and Calculated from JPEG0007775351000010.jpg1137. JPEG0007775351000011.jpg1135 and The weighting of JPEG0007775351000012.jpg1130 is JPEG0007775351000013.jpg944 and JPEG0007775351000014.jpg1048.
[0120] JPEG0007775351000015.jpg3993
[0121] JPEG0007775351000016.jpg826 is a low-pass value that will be used in the merging process at the next higher level. JPEG0007775351000017.jpg1025 is a high-pass coefficient, and the high-pass coefficient at each step is quantized and entropy coded (for example, the encoding of the computational encoder 400012). The weight value is JPEG0007775351000018.jpg1288 is calculated. The root node is the last JPEG0007775351000019.jpg1122 and JPEG0007775351000020.jpg1021 generates the following:
[0122] JPEG0007775351000021.jpg2899
[0123] The gDC values are also quantized and entropy coded like the high-pass coefficients.
[0124] FIG. 10 is a diagram illustrating an example of a point cloud decoder according to an embodiment.
[0125] The point cloud decoder shown in FIG. 10 is an example of the point cloud video decoder 10006 shown in FIG. 1 and performs operations that are the same as or similar to those of the point cloud video decoder 10006 described in FIG. 1. As shown, the point cloud decoder receives a geometry bitstream and an attribute bitstream included in one or more bitstreams. The point cloud decoder includes a geometry decoder and an attribute decoder. The geometry decoder performs geometry decoding on the geometry bitstream and outputs decoded geometry. The attribute decoder performs attribute decoding based on the decoded geometry and attribute bitstream and outputs decoded attributes. The decoded geometry and decoded attributes are used to restore point cloud content (decoded point cloud).
[0126] FIG. 11 is a diagram illustrating an example of a point cloud decoder according to an embodiment.
[0127] The point cloud decoder shown in FIG. 11 is an example of the point cloud decoder described in FIG. 10, and performs a decoding operation that is the reverse process of the encoding operation of the point cloud encoder described in FIGS.
[0128] As explained in Figures 1 and 10, the point cloud decoder performs geometry decoding and feature decoding, with geometry decoding occurring before feature decoding.
[0129] The point cloud decoder according to the embodiment includes an arithmetic decoder 11000, an octree synthesis unit 11001, a surface approximation synthesis unit 11002, a geometry reconstruction unit 11003, an inverse transform coordinates unit 11004, an arithmetic decoder 11005, an inverse quantization unit 11006, an RAHT transform unit 11007, a LOD generation unit 11008, an inverse lifting unit 11009 and / or an inverse color transformation unit 11010.
[0130] The arithmetic decoder 11000, octree synthesis unit 11001, surface approximation synthesis unit 11002, geometry reconstruction unit 11003, and coordinate system inverse transformation unit 11004 perform geometry decoding. Geometry decoding according to the embodiment includes direct coding and trisoup geometry decoding. Direct coding and trisoup geometry decoding are selectively applied. Furthermore, geometry decoding is not limited to the above examples and is performed in the reverse process of the geometry encoding described with reference to FIGS. 1 to 9.
[0131] The arithmetic decoder 11000 according to the embodiment decodes the received geometry bitstream based on arithmetic coding. The operation of the arithmetic decoder 11000 corresponds to the reverse process of the arithmetic encoder 40004.
[0132] The octree synthesis unit 11001 according to the embodiment obtains exclusive codes from the decoded geometry bitstream (or from the decoded result, information on the secured geometry) to generate an octree. Specific details regarding exclusive codes are as described in FIGS. 1 to 9.
[0133] The surface approximation synthesis unit 11002 according to the embodiment synthesizes a surface based on the decoded geometry and / or the generated octree if trisoup geometry encoding is applied.
[0134] According to the embodiment, the geometry reconstruction unit 11003 regenerates geometry based on the surface and / or decoded geometry. As described with reference to FIGS. 1 to 9, direct coding and trisoup geometry coding are selectively applied. Therefore, the geometry reconstruction unit 11003 directly retrieves and adds position information of points to which direct coding is applied. Also, when trisoup geometry coding is applied, the geometry reconstruction unit 11003 reconstructs geometry by performing reconstruction operations of the geometry reconstruction unit 40005, such as triangulation, upsampling, and voxelization operations. The detailed contents are the same as those described with reference to FIG. 6, and therefore will not be repeated. The reconstructed geometry includes a point cloud picture or frame that does not include features.
[0135] The coordinate system inverse transform unit 11004 according to the embodiment transforms the coordinate system based on the reconstructed geometry to obtain the position of the point.
[0136] The arithmetic decoder 11005, the inverse quantization unit 11006, the RAHT transform unit 11007, the LOD generation unit 11008, the inverse lift unit 11009, and / or the color inverse transform unit 11010 perform the feature decoding described in FIG. 10. Feature decoding according to the embodiment includes RAHT (Region Adaptive Hierarchical Transform) decoding, Interpolarization-based hierarchical nearest-neighbor prediction-Prediction Transform (Interpolarization-based hierarchical nearest-neighbor prediction with an update / lifting step (Lifting Transform)) decoding. The above three decoding methods may be used selectively, or a combination of one or more of them may be used. Also, feature decoding according to the embodiment is not limited to the above examples.
[0137] The arithmetic decoder 11005 according to the embodiment decodes the attribute bitstream into arithmetic coding.
[0138] The inverse quantization unit 11006 according to the embodiment inverse quantizes the decoded feature bitstream or the feature information obtained as a result of decoding, and outputs the inverse quantized feature (or feature value). The inverse quantization is selectively applied based on the feature encoding of the point cloud encoder.
[0139] In some embodiments, the RAHT transform unit 11007, LOD generator 11008, and / or inverse lifting unit 11009 process the reconstructed geometry and dequantized features. As described above, the RAHT transform unit 11007, LOD generator 11008, and / or inverse lifting unit 11009 selectively perform decoding operations corresponding to the encoding of the point cloud encoder.
[0140] According to an embodiment, the color inverse transform unit 11010 performs inverse transform coding to inversely transform the color values (or textures) included in the decoded features. The operation of the color inverse transform unit 11010 is selectively performed based on the operation of the color transform unit 40006 of the point cloud encoder.
[0141] The elements of the point cloud decoder of Figure 11 may be embodied in hardware, software, firmware, or a combination thereof, including one or more processors or integrated circuits (not shown) configured to communicate with one or more memories included in the point cloud providing device. The one or more processors perform any of the operations and / or functions of the elements of the point cloud decoder of Figure 11 described above. Additionally, the one or more processors operate or execute a software program and / or set of instructions to perform the operations and / or functions of the elements of the point cloud decoder of Figure 11.
[0142] FIG. 12 shows an example of a transmitting device according to the embodiment.
[0143] The transmitting device shown in Fig. 12 is an example of the transmitting device 10000 of Fig. 1 (or the point cloud encoder of Fig. 4). The transmitting device shown in Fig. 12 performs any of the same or similar operations and methods as the operations and encoding methods of the point cloud encoder described with reference to Figs. 1 to 9. The transmitting device according to the embodiment includes a data input unit 12000, a quantization processing unit 12001, a voxelization processing unit 12002, an octree occupation code generation unit 12003, a surface model processing unit 12004, an intra / inter coding processing unit 12005, an arithmetic coder 12006, a metadata processing unit 12007, a hue conversion processing unit 12008, a characteristic conversion processing unit (or attribute conversion processing unit) 12009, a prediction / lift / RAHT conversion processing unit 12010, an arithmetic coder 12011, and / or a transmission processing unit 12012.
[0144] According to the embodiment, the data input unit 12000 receives or acquires point cloud data. The data input unit 12000 performs operations and / or acquisition methods that are the same as or similar to the operations and / or acquisition methods of the point cloud video acquisition unit 10001 (or the acquisition process 20000 shown in FIG. 2).
[0145] Geometry encoding is performed by a data input unit 12000, a quantization unit 12001, a voxelization unit 12002, an octree occupation code generation unit 12003, a surface model processing unit 12004, an intra / inter coding processing unit 12005, and an arithmetic coder 12006. The geometry encoding according to this embodiment is the same as or similar to the geometry encoding described with reference to Figures 1 to 9, and therefore a detailed description thereof will be omitted.
[0146] The quantization unit 12001 according to the embodiment quantizes geometry (e.g., position values of points). The operation and / or quantization of the quantization unit 12001 is the same as or similar to the operation and / or quantization of the quantization unit 40001 shown in Fig. 4. The specific description is as described in Figs. 1 to 9.
[0147] The voxelization processing unit 12002 according to the embodiment voxels the position values of the quantized points. The voxelization processing unit 12002 performs operations and / or processes that are the same as or similar to the operations and / or voxelization processes of the quantization unit 40001 shown in Fig. 4. Specific details are as described in Figs. 1 to 9.
[0148] The octree occupation code generator 12003 according to the embodiment performs octree coding on the positions of voxelized points based on an octree structure. The octree occupation code generator 12003 generates occupation codes. The octree occupation code generator 12003 performs operations and / or methods that are the same as or similar to the operations and / or methods of the point cloud encoder (or the octree analyzer 40002) described in FIGS. 4 and 6. Specific descriptions are as described in FIGS. 1 to 9.
[0149] The surface model processing unit 12004 according to the embodiment performs trisoup geometry encoding, which reconstructs the positions of points within a specific region (or node) on a voxel basis based on a surface model. The surface model processing unit 12004 performs operations and / or methods that are the same as or similar to the operations and / or methods of the point cloud encoder (e.g., the surface approximation analysis unit 40003) shown in Figure 4. Specific descriptions are as described in Figures 1 to 9.
[0150] According to the embodiment, the intra / inter coding processor 12005 performs intra / inter coding of the point cloud data. The intra / inter coding processor 12005 performs coding that is the same as or similar to the intra / inter coding described in FIG. 7, as detailed in FIG. 7. In the embodiment, the intra / inter coding processor 12005 is included in the arithmetic coder 12006.
[0151] According to an embodiment, the arithmetic coder 12006 entropy encodes the octree and / or approximated octree of the point cloud data. For example, the encoding method includes an arithmetic encoding method. The arithmetic coder 12006 performs operations and / or methods that are the same as or similar to those of the arithmetic encoder 40004.
[0152] The metadata processing unit 12007 according to the embodiment processes metadata related to point cloud data, such as setting values, and provides the metadata to necessary processing steps such as geometry encoding and / or feature encoding. The metadata processing unit 12007 according to the embodiment also generates and / or processes signaling information related to geometry encoding and / or feature encoding. The signaling information according to the embodiment is encoded separately from the geometry encoding and / or feature encoding. The signaling information according to the embodiment may also be interleaved.
[0153] The hue conversion processor 12008, the feature conversion processor 12009, the prediction / lift / RAHT conversion processor 12010, and the arithmetic coder 12011 perform feature coding. The feature coding according to the embodiment is the same as or similar to the feature coding described with reference to FIGS. 1 to 9, so a detailed description thereof will be omitted.
[0154] The color conversion unit 12008 according to this embodiment performs color conversion coding to convert the hue value included in the feature. The color conversion unit 12008 performs color conversion coding based on the reconstructed geometry. The reconstructed geometry has been described with reference to FIGS. 1 to 9. The color conversion unit 12008 also performs operations and / or methods that are the same as or similar to the operations and / or methods of the color conversion unit 40006 described with reference to FIG. 4. Detailed description thereof will be omitted.
[0155] The feature conversion processor 12009 according to the embodiment performs feature conversion, which converts features based on positions where geometry coding has not been performed and / or reconstructed geometry. The feature conversion processor 12009 performs operations and / or methods identical to or similar to those of the feature conversion unit 40007 described in FIG. 4, detailed descriptions of which will be omitted. The prediction / lift / RAHT conversion processor 12010 according to the embodiment codes the converted features using any one or a combination of RAHT coding, predictive transformation coding, and lift transformation coding. The prediction / lift / RAHT conversion processor 12010 performs any one of operations identical to or similar to those of the RAHT conversion unit 40008, LOD generation unit 40009, and lift transformation unit 40010 described in FIG. 4. Since the predictive transformation coding, lift transformation coding, and RAHT transformation coding have been described with reference to FIGS. 1 to 9, detailed descriptions of these will be omitted.
[0156] The arithmetic coder 12011 according to the embodiment encodes the coded characteristic based on arithmetic coding. The arithmetic coder 12011 performs operations and / or methods that are the same as or similar to the operations and / or methods of the arithmetic encoder 400012.
[0157] The transmission processing unit 12012 according to the embodiment transmits each bitstream including the coded geometry and / or coded attribute and metadata information, or transmits the coded geometry and / or coded attribute and metadata information in one bitstream. When the coded geometry and / or coded attribute and metadata information according to the embodiment is configured in one bitstream, the bitstream includes one or more sub-bitstreams. The bitstream according to the embodiment includes signaling information including a Sequence Parameter Set (SPS) for sequence-level signaling, a Geometry Parameter Set (GPS) for signaling geometry information coding, an Attribute Parameter Set (APS) for signaling attribute information coding, and a Tile Parameter Set (TPS) for tile-level signaling, and slice data. The slice data includes information about one or more slices. According to the embodiment, one slice is one geometry bitstream (Geometry 0). 0 ) and one or more attribute bitstreams (Attr0 0 , Attr1 0) according to the embodiment. A TPS according to the embodiment includes information about each tile (e.g., coordinate value information and height / size information of a bounding box, etc.) for one or more tiles. A geometry bitstream includes a header and a payload. The header of a geometry bitstream according to the embodiment includes identification information of a parameter set included in the GPS (geom_parameter_set_id), a tile identifier (geom_tile_id), a slice identifier (geom_slice_id), and information about data included in the payload. As described above, the metadata processing unit 12007 according to the embodiment can generate and / or process signaling information and transmit it to the transmission processing unit 12012. In the embodiment, an element that performs geometry coding and an element that performs attribute coding can share data / information with each other, as shown by the dotted lines. The transmission processing unit 12012 according to the embodiment performs an operation and / or a transmission method that is the same as or similar to the operation and / or transmission method of the transmitter 10003. As it is the same as that described with reference to FIGS. 1 and 2, a detailed description thereof will be omitted.
[0158] FIG. 13 shows an example of a receiving device according to the embodiment.
[0159] The receiving device shown in Fig. 13 is an example of the receiving device 10004 in Fig. 1 (or the point cloud decoder in Figs. 10 and 11). The receiving device shown in Fig. 13 performs any of the same or similar operations and methods as the operations and decoding methods of the point cloud decoder described in Figs. 1 to 11.
[0160] The receiving device according to the embodiment includes a receiving unit 13000, a receiving processing unit 13001, an arithmetic decoder 13002, an occupancy code-based octree reconstruction processing unit 13003, a surface model processing unit (triangle reconstruction, upsampling, voxelization) 13004, an inverse quantization processing unit 13005, a metadata analysis 13006, an arithmetic decoder 13007, an inverse quantization processing unit 13008, a prediction / lift / RAHT inverse transform processing unit 13009, a color inverse transform processing unit 13010, and / or a renderer 13011. Each component of the decoding according to the embodiment performs the reverse process of the component of the encoding according to the embodiment.
[0161] The receiving unit 13000 according to the embodiment receives point cloud data. The receiving unit 13000 performs operations and / or a receiving method that are the same as or similar to the operations and / or a receiving method of the receiver 10005 of Fig. 1. Detailed description thereof will be omitted.
[0162] The receiving unit 13001 according to the embodiment obtains a geometry bitstream and / or a characteristic bitstream from the received data. The receiving unit 13000 includes the receiving unit 13000.
[0163] Geometry decoding is performed by an arithmetic decoder 13002, an exclusive code based octree reconstruction processor 13003, a surface model processor 13004, and an inverse quantization processor 13005. The geometry decoding according to this embodiment is the same as or similar to the geometry decoding described with reference to Figures 1 to 10, so a detailed description thereof will be omitted.
[0164] The arithmetic decoder 13002 according to the embodiment decodes the geometry bitstream based on arithmetic coding. The arithmetic decoder 13002 performs operations and / or coding that are the same as or similar to the operations and / or coding of the arithmetic decoder 11000.
[0165] The exclusive code-based octree reconstruction processor 13003 according to the embodiment obtains exclusive codes from the decoded geometry bitstream (or information on the secured geometry as a result of decoding) and reconstructs an octree. The exclusive code-based octree reconstruction processor 13003 performs operations and / or methods identical to or similar to the operations and / or octree generation method of the octree synthesis unit 11001. The surface model processor 13004 according to the embodiment performs trisoup geometry decoding and associated geometry reconstruction (e.g., triangle reconstruction, upsampling, voxelization) based on a surface model scheme when trisoup geometry encoding is applied. The surface model processor 13004 performs operations identical to or similar to the operations of the surface approximation synthesis unit 11002 and / or the geometry reconstruction unit 11003.
[0166] The inverse quantization unit 13005 according to the embodiment inverse quantizes the decoded geometry.
[0167] According to an embodiment, the metadata analysis 13006 analyzes metadata, such as setting values, included in the received point cloud data. The metadata analysis 13006 transmits the metadata to geometry decoding and / or feature decoding. A detailed description of the metadata is omitted here as it has been described with reference to FIG.
[0168] The arithmetic decoder 13007, the inverse quantization processor 13008, the prediction / lift / RAHT inverse transform processor 13009, and the hue inverse transform processor 13010 perform feature decoding. The feature decoding is the same as or similar to the feature decoding described in Figures 1 and 10, so a detailed description thereof will be omitted.
[0169] The operation decoder 13007 according to the embodiment decodes the attribute bitstream into operation coding. The operation decoder 13007 performs decoding of the attribute bitstream based on the reconstructed geometry. The operation decoder 13007 performs operations and / or coding that are the same as or similar to the operations and / or coding of the operation decoder 11005.
[0170] The inverse quantization unit 13008 according to the embodiment inversely quantizes the decoded characteristic bitstream. The inverse quantization unit 13008 performs the same or similar operations and / or methods as the inverse quantization unit 11006.
[0171] The prediction / lift / RAHT inverse transform processing unit 13009 according to the embodiment processes the reconstructed geometry and dequantized features. The prediction / lift / RAHT inverse transform processing unit 13009 performs any of the same or similar operations and / or decoding as those of the RAHT transform unit 11007, the LOD generation unit 11008, and / or the inverse lift unit 11009. The hue inverse transform processing unit 13010 according to the embodiment performs inverse transform coding to inversely transform the color values (or textures) included in the decoded features. The hue inverse transform processing unit 13010 performs the same or similar operations and / or inverse transform coding as those of the color inverse transform unit 11010. The renderer 13011 according to the embodiment renders point cloud data.
[0172] FIG. 14 is a diagram illustrating an architecture for G-PCC-based point cloud content streaming according to an embodiment.
[0173] The upper part of FIG. 14 illustrates a process in which the transmitting device described in FIGS. 1 to 13 (eg, the transmitting device 10000, the transmitting device in FIG. 12, etc.) processes and transmits point cloud content.
[0174] As described in FIGS. 1 to 13, the transmitting device acquires audio (Ba) of the point cloud content (Audio Acquisition), encodes the acquired audio (Audio Encoding), and outputs an audio bitstream (Ea). The transmitting device also acquires point cloud (Bv) (or point cloud video) of the point cloud content (Point Acqusition), performs point cloud encoding on the acquired point cloud, and outputs a point cloud video bitstream (Eb). The point cloud encoding of the transmitting device is the same as or similar to the point cloud encoding described in FIGS. 1 to 13 (e.g., encoding by the point cloud encoder in FIG. 4), so a detailed description thereof will be omitted.
[0175] The transmitting device encapsulates the generated audio and video bitstreams into files and / or segments (File / segment encapsulation). The encapsulated files and / or segments (Fs, File) include files in file formats such as ISOBMFF or DASH segments. According to an embodiment, point cloud-related metadata is included in the encapsulated file formats and / or segments. The metadata may be included in boxes at various levels in the ISOBMFF file format or in separate tracks within the file. In an embodiment, the transmitting device may encapsulate the metadata itself in a separate file. According to an embodiment, the transmitting device transmits the encapsulated file formats and / or segments via a network. The encapsulation and transmission processing method of the transmitting device is the same as that described in FIGS. 1 to 13 (e.g., transmitter 10003, transmission step 20002 of FIG. 2, etc.), so detailed description thereof will be omitted.
[0176] The lower part of FIG. 14 shows a process in which the receiving device described in FIGS. 1 to 13 (for example, receiving device 10004, receiving device of FIG. 13, etc.) processes and outputs point cloud content.
[0177] In an embodiment, the receiving device includes a device (e.g., loudspeakers, headphones, display) that outputs final audio data and final video data, and a point cloud player that processes point cloud content. The final data output device and the point cloud player are configured as separate physical devices. The point cloud player in this embodiment performs Geometry-based Point Cloud Compression (G-PCC) coding and / or Video-based Point Cloud Compression (V-PCC) coding and / or next-generation coding.
[0178] A receiving device according to an embodiment secures and decapsulates files and / or segments (F', Fs') included in received data (e.g., broadcast signals, signals transmitted over a network, etc.). The receiving and decapsulation methods of the receiving device are the same as those described in Figures 1 to 13 (e.g., receiver 10005, receiving unit 13000, receiving processing unit 13001, etc.), so detailed description thereof will be omitted.
[0179] A receiving device according to an embodiment of the present invention obtains an audio bitstream (E'a) and a video bitstream (E'v) included in a file and / or a segment. As shown in the figure, the receiving device performs audio decoding on the audio bitstream, outputs decoded audio data (B'a), and performs audio rendering on the decoded audio data to output final audio data (A'a) via a speaker or headphones.
[0180] The receiving device also performs point cloud decoding on the video bitstream (E'v) and outputs decoded video data (B'v). The point cloud decoding according to the embodiment is the same as or similar to the point cloud decoding described in Figures 1 to 13 (e.g., the decoding of the point cloud decoder in Figure 11), so a detailed description will be omitted. The receiving device renders the decoded video data and outputs the final video data to a display.
[0181] The receiving device according to the embodiment performs one of decapsulation, audio decoding, audio rendering, point cloud decoding, and rendering operations based on the transmitted metadata. The description of the metadata is omitted here as it is the same as that described in Figures 12 and 13.
[0182] As shown by the dotted lines, a receiving device according to an embodiment (e.g., a point cloud player or a sensing / tracking unit within the point cloud player) generates feedback information (orientation, viewport). The feedback information according to the embodiment is used in the decapsulation, point cloud decoding and / or rendering processes of the receiving device and can also be transmitted to the transmitting device. The description of the feedback information is omitted as it has been described in Figures 1 to 13.
[0183] FIG. 15 is a diagram illustrating an example of a transmitting device according to an embodiment.
[0184] The transmitting device in Fig. 15 is a device for transmitting point cloud content, and corresponds to the examples of the transmitting devices described in Fig. 1 to Fig. 14 (for example, the transmitting device 10000 in Fig. 1, the point cloud encoder in Fig. 4, the transmitting device in Fig. 12, the transmitting device in Fig. 14, etc.). Therefore, the transmitting device in Fig. 15 performs the same or similar operations as the transmitting devices described in Fig. 1 to Fig. 14.
[0185] A transmitting device according to an embodiment may perform one or more of point cloud acquisition, point cloud encoding, file / segment encapsulation, and delivery.
[0186] The point cloud acquisition and transmission operations shown in the figures are the same as those described with reference to FIGS. 1 to 14, so a detailed description thereof will be omitted.
[0187] As described with reference to FIGS. 1 to 14, the transmitting device according to the embodiment performs geometry encoding and attribute encoding. Geometry encoding according to the embodiment is also called geometry compression, and attribute encoding is also called attribute compression. As described above, one point has one geometry and one or more attributes. Therefore, the transmitting device performs attribute encoding for each attribute. The drawings show an example in which the transmitting device performs one or more attribute compressions (Attribute #1 compression, ... Attribute #N compression). The transmitting device according to the embodiment can also perform auxiliary compression. The auxiliary compression is performed on metadata. A description of the metadata is omitted as it has been described with reference to FIGS. 1 to 14. The transmitting device can also perform mesh data compression. Mesh data compression according to the embodiment includes the trisoup geometry encoding described with reference to FIGS. 1 to 14.
[0188] In an embodiment, a sending device encapsulates a bitstream (e.g., a point cloud stream) output by point cloud encoding into files and / or segments. In an embodiment, the sending device performs media track encapsulation to carry data other than metadata (e.g., media data) and metadata track encapsulation to carry metadata. In an embodiment, metadata is encapsulated in a media track.
[0189] As described in Figures 1 to 14, the transmitting device receives feedback information (orientation / viewport metadata) from the receiving device, and performs one of point cloud encoding, file / segment encapsulation, and transmission operations based on the received feedback information. Detailed descriptions are omitted as they are the same as those described in Figures 1 to 14.
[0190] FIG. 16 is a diagram illustrating an example of a receiving device according to an embodiment.
[0191] The receiving device of Figure 16 is a device for receiving point cloud content, and corresponds to an example of the receiving device described in Figures 1 to 14 (e.g., receiving device 10004 of Figure 1, point cloud decoder of Figure 11, receiving device of Figure 13, receiving device of Figure 14, etc.). Therefore, the receiving device of Figure 16 performs the same or similar operations as the receiving devices described in Figures 1 to 14. In addition, the receiving device of Figure 16 can receive signals transmitted by the transmitting device of Figure 15 and perform the reverse process of the operations of the transmitting device of Figure 15.
[0192] A receiving device according to an embodiment performs one or more of delivery, file / segment decapsulation, point cloud decoding, and point cloud rendering.
[0193] The illustrated point cloud receiving and point cloud rendering operations are the same as those described with reference to FIGS. 1 to 14, and therefore detailed description thereof will be omitted.
[0194] As illustrated in Figures 1 to 14, a receiving device according to an embodiment performs decapsulation on files and / or segments obtained from a network or storage device. In an embodiment, the receiving device may perform media track decapsulation, which carries data other than metadata (e.g., media data), and may perform metadata track decapsulation, which carries metadata. In an embodiment, if metadata is encapsulated in a media track, metadata track decapsulation is omitted.
[0195] As illustrated in FIGS. 1 to 14, the receiving device performs geometry decoding and attribute decoding on the bitstream (e.g., point cloud stream) secured by decapsulation. Geometry decoding according to the embodiment is also referred to as geometry decompression, and attribute decoding is also referred to as attribute decompression. As described above, one point has one geometry and one or more attributes, which are encoded separately. Therefore, the receiving device performs attribute decoding for each attribute. The drawings illustrate an example in which the receiving device performs one or more attribute decompressions (Attribute #1 decompression, ..., Attribute #N decompression). The receiving device according to the embodiment can also perform auxiliary decompression. The auxiliary decompression is performed on metadata. The description of the metadata is omitted here, as it has been described with reference to FIGS. 1 to 14. The receiving device also performs mesh data decompression. Mesh data decompression according to the embodiment includes the trisoup geometry decoding described with reference to FIGS. 1 to 14. The receiving device according to the embodiment renders the point cloud data output by point cloud decoding.
[0196] As described in Figures 1 to 14, the receiving device can obtain orientation / viewport metadata using a separate sensing / tracking element, etc., and transmit feedback information including the same to the transmitting device (e.g., the transmitting device in Figure 15). The receiving device can also perform one of a receiving operation, file / segment decapsulation, and point cloud decoding based on the feedback information. Detailed descriptions are omitted here as they are the same as those described in Figures 1 to 14.
[0197] FIG. 17 is a diagram illustrating an example of a structure that can be linked to a method / apparatus for transmitting and receiving point cloud data according to an embodiment.
[0198] 17 illustrates a configuration in which any of a server 1760, a robot 1710, an autonomous vehicle 1720, an XR device 1730, a smartphone 1740, a home appliance 1750, and / or an HMD 1770 are connected to a cloud network 1710. The robot 1710, the autonomous vehicle 1720, the XR device 1730, the smartphone 1740, or the home appliance 1750 may also be referred to as a device. The XR device 1730 may correspond to or be linked to a point cloud data (PCC) device according to an embodiment.
[0199] Cloud network 1700 refers to a network that forms part of a cloud computing infrastructure or exists within a cloud computing infrastructure, where cloud network 1700 is configured using a 3G network, a 4G or LTE network, a 5G network, or the like.
[0200] The server 1760 is connected to any of the robot 1710, autonomous vehicle 1720, XR device 1730, smartphone 1740, home appliance 1750, and / or HMD 1770 via the cloud network 1700 and can assist with at least some of the processing of the connected devices 1710-1770.
[0201] An HMD (Head-Mounted Display) 1770 represents any type of XR device and / or PCC device according to an embodiment. An HMD type device according to an embodiment includes a communications unit, a control unit, a memory unit, an I / O unit, a sensor unit, and a power supply unit.
[0202] Various embodiments of the devices 1710 to 1750 to which the above-described technology is applied will be described below. Here, the devices 1710 to 1750 shown in Fig. 17 can be linked / coupled to the point cloud data transmitting / receiving device according to the above-described embodiments.
[0203] <PCC+XR>
[0204] The XR / PCC device 1730 may be implemented by applying PCC and / or XR (AR+VR) technology to a head-mounted display (HMD), a head-up display (HUD) installed in a vehicle, a TV, a mobile phone, a smartphone, a computer, a wearable device, a home appliance, a digital sign, a vehicle, a fixed robot, a mobile robot, etc.
[0205] The XR / PCC device 1730 can obtain information about the surrounding space or real objects by analyzing 3D point cloud data or image data acquired by various sensors or from an external device to generate position data and attribute data for 3D points, and can render and output the XR object to be output. For example, the XR / PCC device 1730 can output an XR object including additional information about the recognized object corresponding to the recognized object.
[0206] <PCC+Self-propelled+XR>
[0207] The autonomous vehicle 1720 is realized as a mobile robot, vehicle, unmanned aerial vehicle, etc. by applying PCC technology and XR technology.
[0208] The autonomous vehicle 1720 to which XR / PCC technology is applied refers to an autonomous vehicle equipped with a means for providing XR images, an autonomous vehicle that can be controlled / interacted with within the XR images, etc. In particular, the autonomous vehicle 1720 that can be controlled / interacted with within the XR images can be separated from the XR device 1730 and can be linked to each other.
[0209] The autonomous vehicle 1720, which is equipped with a means for providing XR / PCC images, obtains sensor information from sensors including a camera and outputs XR / PCC images generated based on the obtained sensor information. For example, the autonomous vehicle 1720 may be equipped with a HUD and output XR / PCC images, thereby providing passengers with XR / PCC objects corresponding to real objects or objects on a screen.
[0210] At this time, when the XR / PCC object is output to the HUD, at least a portion of the XR / PCC object is output to overlap with an actual object toward which the passenger's gaze is directed. On the other hand, when the XR / PCC object is output to a display provided in the autonomous vehicle, at least a portion of the XR / PCC object is output to overlap with an object on the screen. For example, the autonomous vehicle 1220 may output XR / PCC objects corresponding to objects such as a roadway, another vehicle, a traffic light, a traffic sign, a motorcycle, a pedestrian, a building, etc.
[0211] The VR (Virtual Reality) technology, AR (Augmented Reality) technology, MR (Mixed Reality) technology, and / or PCC (Point Cloud Compression) technology according to the embodiments can be applied to various devices.
[0212] In other words, VR technology is a display technology that presents real objects and backgrounds only as CG images. In contrast, AR technology is a technology that displays virtual CG images on top of images of real things. MR technology is similar to AR technology in that it mixes virtual objects into the real world. However, AR technology clearly distinguishes between real objects and virtual objects made of CG images, and uses virtual objects to complement real objects, while MR technology is different from AR technology in that virtual objects and real objects are considered to have the same characteristics. More specifically, for example, hologram services are an application of the MR technology.
[0213] However, recently, rather than clearly distinguishing between VR, AR, and MR technologies, they are also referred to as XR (extended reality) technologies. Therefore, embodiments of the present invention can be applied to any of VR, AR, MR, and XR technologies. These technologies apply encoding / decoding based on PCC, V-PCC, and G-PCC technologies.
[0214] The PCC method / apparatus according to the embodiment can be applied to vehicles that provide autonomous driving services.
[0215] Vehicles that provide autonomous driving services are connected to the PCC device via wired or wireless communication.
[0216] When a point cloud data (PCC) transceiver according to an embodiment is connected to a vehicle via wired or wireless communication, it can receive / process content data related to AR / VR / PCC services that can be provided along with an autonomous driving service and transmit the content data to the vehicle. When the point cloud data transceiver is installed in a vehicle, it can receive / process content data related to AR / VR / PCC services according to a user input signal input through a user interface device and provide the content data to a user. The vehicle or user interface device according to an embodiment receives a user input signal. The user input signal according to an embodiment includes a signal instructing an autonomous driving service.
[0217] FIG. 18 is a block diagram illustrating an example of a point cloud encoder.
[0218] A point cloud encoder 1800 according to an embodiment (e.g., the point cloud video encoder 10002 of FIG. 1, the point cloud encoder of FIG. 4, or the point cloud encoders described in FIGS. 12, 14, and 15) performs the encoding operations described in FIGS. 1 to 17. The point cloud encoder 1800 according to an embodiment includes a spatial divider 1810, a geometry information encoder 1820, and an attribute information encoder (or characteristic information encoder) 1830. The point cloud encoder 1800 according to an embodiment further includes one or more elements, not shown in FIG. 18, for performing the encoding operations described in FIGS. 1 to 17.
[0219] Point Cloud Compression (PCC) data (or PCC data, point cloud data) is input data for the Point Cloud Encoder 1800 and includes geometry and / or features. According to an embodiment, geometry is information indicating the location (e.g., position) of a point and is expressed using parameters in a coordinate system such as a Cartesian coordinate system, a cylindrical coordinate system, or a spherical coordinate system. According to an embodiment, geometry is referred to as geometry information, and features are referred to as feature information.
[0220] The space division unit 1810 according to the embodiment generates geometry and characteristics of the point cloud data. The space division unit 1810 according to the embodiment divides the point cloud data into one or more 3D blocks in a 3D space to store point information of the point cloud data. According to the embodiment, a block indicates one of a tile group, a tile, a slice, a coding unit (CU), a prediction unit (PU), and a transformation unit (TU). According to the embodiment, the space division unit 1810 performs a division operation based on one of an octree, a quadtree, a binary tree, a triple tree, and a kd tree. One block includes one or more points. According to the embodiment, the block is a hexahedral block having predetermined width, length, and height values. The size of the block according to the embodiment is variable and is not limited to the above example. The spatial division unit 1810 according to the embodiment generates geometry information about one or more points included in the block.
[0221] The geometry information encoding unit 1820 (or geometry information encoder) according to the embodiment performs geometry encoding to generate a geometry bitstream and reconstructed geometry information. In the geometry encoding according to the embodiment, the reconstructed geometry information is input to a feature information encoding unit (or feature encoder) 1830. The geometry information encoding unit 1820 according to the embodiment performs any of the operations of the coordinate system conversion unit 40000, the quantization 40001, the octree analysis unit 40002, the surface approximation analysis unit 40003, the arithmetic encoder 40004, and the geometry reconstruction unit (Reconstruct Geometry) 40005 described in FIG. 4. In addition, the geometry information encoding unit 1820 according to the embodiment performs any one of the operations of the data input unit 12000, the quantization processing unit 12001, the voxelization processing unit 12002, the octree occupation code generation unit 12003, the surface model processing unit 12004, the intra / inter coding processing unit 12005, the arithmetic coder 12006, and the metadata processing unit 12007 described in FIG. 12 .
[0222] The feature information encoder 1830 according to the embodiment generates a feature information bitstream (or feature bitstream) based on the reconstructed geometry information and features.
[0223] A point cloud encoder according to an embodiment transmits a geometry information bitstream and a feature information bitstream, or a bitstream in which the geometry information bitstream and the feature information bitstream are multiplexed. As described above, the bitstream further includes signaling information associated with the geometry information and feature information, signaling information associated with coordinate system transformation, etc. Furthermore, the point cloud encoder according to an embodiment encapsulates the bitstream and transmits it in the form of a segment and / or a file.
[0224] FIG. 19 is a block diagram showing an example of a geometry information encoder.
[0225] A geometry information encoder 1900 (or geometry encoder) according to the embodiment is an example of the geometry information encoder 1820 of FIG. 18 and performs geometry encoding. The geometry encoding according to the embodiment is the same as or similar to the geometry encoding described with reference to FIGS. 1 to 18 , and therefore a detailed description thereof will be omitted. As shown in the figure, the geometry information encoder 1900 includes a coordinate system converter 1910, a geometry information transform and quantizer 1920, a residual geometry information quantizer 1930, a geometry information entropy encoder 1940, a residual geometry information dequantizer 1950, a filtering unit 1960, a memory 1970, and a geometry information predictor 1980. Although not shown in FIG. 19, the geometry information encoder 1900 according to the embodiment further includes one or more elements for performing geometry encoding described with reference to FIGS. 1 to 18 .
[0226] The coordinate system conversion unit 1910 according to the embodiment converts the received geometry information into information on a coordinate system in order to express the position of each point indicated by the input geometry information as a position in a three-dimensional space. The coordinate system conversion unit 1910 performs the same or similar operation as the coordinate system conversion unit 40000 described in FIG. 4. The coordinate system according to the embodiment includes, but is not limited to, the above-mentioned three-dimensional Cartesian coordinate system, cylindrical coordinate system, spherical coordinate system, etc. The coordinate system conversion unit 1910 according to the embodiment can convert a set coordinate system into another coordinate system.
[0227] The coordinate system conversion unit 1910 according to the embodiment may perform coordinate system conversion for units such as a sequence, a frame, a tile, a slice, a block, etc. Whether or not a coordinate system conversion is performed according to the embodiment and information related to the coordinate system and / or the conversion may be signaled in units of a sequence, a frame, a tile, a slice, a block, etc. Therefore, the point cloud data receiving apparatus according to the embodiment may obtain information related to the coordinate system and / or the conversion based on whether or not a coordinate system conversion is performed for a neighboring block, the block size, the number of points, the quantization value, the block division depth, the position of the unit, the distance between the unit and the origin, etc.
[0228] The geometry information transform and quantization unit 1920 according to the embodiment quantizes geometry information expressed in a coordinate system to generate transformed and quantized geometry information. The geometry information transform and quantization unit 1920 according to the embodiment applies one or more transformations, such as position transformation and / or rotation transformation, to the positions of points indicated by the geometry information output from the coordinate system transformation unit 1910, and quantizes the transformed geometry information by dividing it into quantization values. The geometry information transform and quantization unit 1920 performs operations that are the same as or similar to the operations of the quantization unit 40001 of FIG. 4 and / or the quantization processing unit 12001 of FIG. 12. The quantization value according to the embodiment varies based on the distance between the coding unit (e.g., tile, slice, etc.) and the origin of the coordinate system or the angle from the reference direction. The quantization value according to the embodiment is a preset value.
[0229] The geometry information prediction unit 1930 according to the embodiment calculates a predicted value (or predicted geometry information) based on the quantization value of the surrounding coding unit.
[0230] The residual geometry information quantization unit 1940 receives the transformed and quantized geometry information and the residual geometry information obtained by subtracting the predicted value, and quantizes the residual geometry information into quantized values to generate quantized residual geometry information.
[0231] The geometry information entropy coding unit 1950 entropy codes the quantized residual geometry information. Examples of entropy coding include Exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC).
[0232] The residual geometry information inverse quantization unit 1960 restores the residual geometry information by scaling the quantized geometry information to a quantized value. The restored residual geometry information and the predicted geometry information are combined to generate restored geometry information.
[0233] The filtering unit 1970 filters the restored geometry information. The filtering unit 1970 according to the embodiment includes a deblocking filter unit, an offset correction unit, etc. For geometry information in which two different coding units are transformed into different coordinate systems, the filtering unit 1970 according to the embodiment may perform additional filtering at the boundary between the two coding units.
[0234] The memory 1980 stores the reconstructed geometry information (or reconstructed geometry information). The stored geometry information is provided to the geometry information prediction unit 1930. The reconstructed geometry information stored in the memory is also provided to the characteristic information encoding unit 1830 described in FIG. 18.
[0235] FIG. 20 shows an example of a feature encoder according to an embodiment.
[0236] The feature information encoder 2000 shown in Figure 20 is an example of the feature information encoder 1830 described in Figure 18, and performs feature encoding. The feature encoding according to the embodiment is the same as or similar to the feature encoding described in Figures 1 to 17, so a detailed description will be omitted. As shown, the feature information encoder 2000 according to the embodiment includes a feature characteristic converter 2010, a geometry information mapping unit 2020, a feature information converter 2030, a feature information quantizer 2040, and a feature information entropy encoder 2050.
[0237] According to the embodiment, the characteristic conversion unit 2010 receives characteristic information and converts the characteristics (e.g., color) of the received characteristic information. For example, if the characteristic information includes hue information, the characteristic conversion unit 2010 can convert the color space of the characteristic information (e.g., from RGB to YCbCr). Alternatively, the characteristic conversion unit 2010 may selectively not convert the characteristics of the characteristic information. The characteristic conversion unit 2010 performs operations that are the same as or similar to the operations of the characteristic conversion unit 40007 and / or the hue conversion processing unit 12008.
[0238] According to the embodiment, the geometry information mapping unit 2020 maps the feature information output from the feature property conversion unit 2010 and the received restored geometry information to generate reconstructed feature information. The geometry information mapping unit 2020 generates feature information by reconstructing feature information of one or more points based on the restored geometry information. As described above, the geometry information of one or more points included in one voxel is reconstructed around the median of the voxel. Since the feature information depends on the geometry information, the geometry information mapping unit 2020 reconstructs the feature information based on the restored geometry information. The geometry information mapping unit 2020 performs operations that are the same as or similar to those of the attribute conversion processing unit 12009.
[0239] The attribute information converter 2030 according to the embodiment may receive and convert the reconstructed attribute information. The attribute information converter 2030 according to the embodiment may predict the attribute information and convert the residual attribute information using one or more transform types (e.g., DCT, DST, SADCT, RAHT) in response to the residual between the received reconstructed attribute information and the predicted attribute information.
[0240] The feature information quantization unit 2040 according to the embodiment receives the transformed residual feature information and generates transformed and quantized residual feature information based on the quantization value.
[0241] The characteristic information entropy encoder 2050 according to an embodiment receives the transformed and quantized residual characteristic information, performs entropy encoding on it, and outputs a characteristic information bitstream. The entropy encoding according to an embodiment may include any one of Exponential Golomb, CAVLC (Context-Adaptive Variable Length Coding), and CABAC (Context-Adaptive Binary Arithmetic Coding), and is not limited to the above-mentioned embodiments. The characteristic information entropy encoder 2050 performs the same or similar operation as the arithmetic coder 12011.
[0242] FIG. 21 shows an example of a feature encoder according to an embodiment.
[0243] The feature information encoder 2100 shown in Figure 21 corresponds to an example of the feature information encoding unit 1830 described in Figure 18 and the feature information encoder 2000 described in Figure 20. The feature information encoder 2100 according to this embodiment includes a feature characteristic transform unit 2110, a geometry information mapping unit 2120, a feature information prediction unit 2130, a residual feature information transform unit 2140, a residual feature information inverse transform unit 2145, a residual feature information quantization unit 2150, a residual feature information inverse quantization unit 2155, a filtering unit 2160, a memory 2170, and a feature information entropy encoding unit 2180. The feature information encoder 2100 shown in FIG. 21 differs from the feature information encoder 2000 shown in FIG. 20 in that it further includes a residual feature information transform unit 2140, a residual feature information inverse transform unit 2145, a residual feature information quantization unit 2150, a residual feature information inverse quantization unit 2155, a filtering unit 2160, and a memory 2170.
[0244] The feature characteristic converter 2110 and geometry information mapping unit 2120 according to the embodiment perform operations that are the same as or similar to the feature characteristic converter 2010 and geometry information mapping unit 2020 described in FIG. 20. The feature information predictor 2130 according to the embodiment generates predicted feature information. The residual feature information converter 2140 receives residual feature information generated by subtracting the reconstructed feature information and the predicted feature information output from the geometry information mapping unit 2120. The residual feature information converter 2140 converts the received residual 3D block including the residual feature information into one or more transform types (e.g., predictive transform, lifting transform, DCT, DST, SADCT, RAHT, etc.).
[0245] The residual characteristic information quantization unit 2150 according to the embodiment transforms the input transformed residual characteristic information based on the quantization value. The transformed residual characteristic information is input to the residual characteristic information inverse quantization unit 2155. The residual characteristic information inverse quantization unit 2155 according to the embodiment transforms the transformed and quantized residual characteristic information based on the quantization value to generate transformed residual characteristic information. The transformed residual characteristic information generated by the residual characteristic information inverse quantization unit 2155 is input to the residual characteristic inverse transform unit 2145. The residual characteristic inverse transform unit 2145 according to the embodiment inverse transforms the residual 3D block including the transformed residual characteristic information using one or more transform types (e.g., predictive transform, lifting transform, DCT, DST, SADCT, RAHT, etc.). The restored characteristic information according to the embodiment is generated by combining the inverse transformed residual characteristic information and the predicted characteristic information output from the characteristic information prediction unit 2130. According to the embodiment, the reconstruction feature information is generated by combining non-inverse-transformed residual feature information and predicted feature information. The reconstruction feature information is input to a filtering unit 2160. According to the embodiment, the feature information prediction unit 2130, the residual feature information conversion unit 2140, and / or the residual feature information quantization unit 2150 perform operations that are the same as or similar to those of the prediction / lift / RAHT transformation processing unit 12010.
[0246] The filtering unit 2160 according to the embodiment filters the reconstruction characteristic information. The filtering unit 2160 according to the embodiment includes a deblocking filter, an offset compensation unit, an adaptive loop filter (ALF), etc. The filtering unit 2160 performs an operation that is the same as or similar to the operation of the filtering unit 1970 of FIG.
[0247] The memory 2170 according to the embodiment stores the reconstruction feature information output from the filtering unit 2160. The stored reconstruction feature information is provided as input data for the prediction operation of the feature information prediction unit 2130. The feature information prediction unit 2130 generates predicted feature information based on the reconstruction feature information of the point. Although the memory 2170 is shown as a single block in the drawing, it may be configured as one or more physical memories. The feature information entropy coding unit 2180 according to the embodiment performs the same or similar operation as the feature information entropy coding unit 2050 described in FIG. 20.
[0248] FIG. 22 illustrates an example of a characteristic information prediction unit according to an embodiment.
[0249] The feature information prediction unit 2200 shown in FIG. 22 corresponds to an example of the feature information prediction unit 2130 described in FIG. 21. The feature information prediction unit 2200 performs operations identical or similar to those of the feature information prediction unit 2130. The feature information prediction unit 2200 according to the embodiment includes an LOD construction unit 2210 and a neighboring point set construction unit 2220. The LOD construction unit 2210 performs operations identical or similar to those of the LOD generation unit 40009. That is, as shown in the figure, the LOD construction unit 2210 receives features and reconstructed geometry and constructs one or more LODs based on the received features and reconstructed geometry. As described with reference to FIGS. 4 and 8, LODs are generated by realigning points distributed in a 3D space with a set of refinement levels. According to the embodiment, the LOD includes one or more points distributed at regular intervals. As described above, the LOD according to the embodiment indicates the level of detail of the point cloud content. Therefore, a lower level of LOD (or LOD value) indicates less detail in the point cloud content, while a higher level of LOD indicates more detail in the point cloud content. That is, a higher LOD includes points that are more closely spaced. A point cloud encoder (e.g., the point cloud encoder of FIG. 4) and a point cloud decoder (e.g., the point cloud decoder of FIG. 11) according to embodiments can generate LOD to improve feature compression ratio. This is because points with similar features are likely to be adjacent to a target point, and therefore the residual value between the predicted feature of a target point and the feature of a neighboring point with similar attributes is likely to be close to 0. Therefore, a point cloud encoder and a point cloud decoder according to embodiments can generate LOD to select appropriate neighboring points that can be used when predicting a feature.
[0250] The LOD construction unit 2210 according to the embodiment constructs the LOD using one or more methods. As described above, a point cloud decoder (e.g., the point cloud decoder described in FIGS. 10 and 11) must also generate the LOD. Therefore, information related to the LOD construction method (or LOD generation method) according to the embodiment, or LOD construction method information, is transmitted to a receiving device (e.g., the receiving device 10004 in FIG. 1 or the point cloud decoder in FIGS. 10 and 11) via a bitstream including encoded point cloud video data. Therefore, the receiving device can generate the LOD based on the LOD construction method information.
[0251] The adjacent point set constructor 2220 according to the embodiment generates each LOD (or LOD set) and then performs the LOD l Search or find one or more neighboring points of a point in the set. The number of one or more neighboring points is represented as X, where X is an integer greater than 0. In the embodiment, the neighboring points are located at LOD in 3D space. l Nearest Neighbor (NN) points located closest to a point in the set, with a target LOD (e.g., LOD l ) or a set of LODs with a lower LOD level than the target LOD (e.g., LOD l-1 , LOD l-2 , ..., LOD0). According to an embodiment, the neighboring point set constructing unit 2220 may register one or more searched neighboring points in the predictor as a neighboring point set. According to an embodiment, the number of neighboring points is the maximum number of neighboring points, and may be set by a user input signal or may be preset to a specific value according to a neighboring point search method.
[0252] According to an embodiment, the neighbor point set construction unit 2220 may search for neighbor points of point P3 belonging to LOD1 shown in FIG. 9 at LOD0 and LOD1. As shown in FIG. 9, LOD0 includes P0, P5, P4, and P2. LOD2 includes points at LOD0, LOD1, and P9, P8, and P7. When the number of neighbor points X is 3, the neighbor point set construction unit 2220 searches for three neighbor points closest to P3 from points belonging to LOD0 or LOD1 in the 3D space shown in the upper part of FIG. 9. That is, the neighbor point set construction unit 2220 may search for P6, which belongs to LOD1, which is the same LOD level, and P2 and P4, which belong to LOD0, which is a lower LOD level, as neighbor points of P3. Although P7 is close to P3 in the 3D space, it is not searched for as a neighbor point because it is at a higher LOD level. The adjacent point set construction unit 2220 registers the searched adjacent points (P2, P4, P6) as an adjacent point set in the predictor of P3. The adjacent point set generation method according to the embodiment is not limited to this example. Furthermore, information related to the adjacent point set generation method according to the embodiment (hereinafter, adjacent point set generation information) is included in a bitstream including the above-mentioned encoded point cloud video data and transmitted to a receiving device (for example, the receiving device 10004 in FIG. 1 or the point cloud decoder in FIGS. 10 and 11).
[0253] As described above, every point has one predictor. The point cloud encoder according to the embodiment applies the predictor to encode the feature value of the corresponding point to generate a predicted feature (or predicted feature value). The predictor according to the embodiment is generated based on neighboring points searched after LOD generation. The predictor is used to predict the feature of the target point. Therefore, the predictor can generate a predicted feature by applying a weight to the feature of neighboring points.
[0254] For example, the predictor calculates and registers weights for the set of neighboring points based on the distance values (e.g., 1 / 2 distance) between the target point (e.g., P3) and each of the neighboring points. As described above, the set of neighboring points of P3 is P2, P4, and P6, so the point cloud encoder (or predictor) according to the embodiment calculates weights based on the distance values between P3 and each of the neighboring points. Therefore, the weights of each neighboring point are JPEG0007775351000022.jpg20163. When the neighboring point set of the predictor is set, the point cloud encoder according to the embodiment normalizes the weight value of the neighboring point by the sum of all the weight values of the neighboring points. For example, the sum of the weight values of all neighboring points in the neighboring point set of node P3 is expressed as follows: JPEG0007775351000023.jpg15161. Furthermore, the normalized weight value generated by dividing the sum of the weight values (total_weight) by the weight value of each adjacent point is expressed as follows: JPEG0007775351000024.jpg15167.
[0255] A point cloud encoder (or a feature information predictor) according to an embodiment predicts features using a predictor. The predicted feature (or predicted feature information) according to an embodiment is an average value of values obtained by multiplying the features of registered adjacent points by a calculated weight, or a value obtained by multiplying the feature of a specific point by a weight, as described in FIG. 9. In an embodiment, the point cloud encoder pre-calculates compressed results and then selectively uses a predicted feature value that can generate the smallest stream (a predicted feature value with the highest compression efficiency) from among the above feature values. Note that the method of predicting features is not limited to the above example.
[0256] A point cloud encoder (e.g., coefficient quantization unit 40011) according to an embodiment encodes the feature values of the points, the residuals of the predicted feature values, and information about the selected predicted feature (or information about the method for selecting the predicted feature), and transmits the encoded information to a receiving device (e.g., receiving device 10004 in FIG. 1 or point cloud decoders in FIGS. 10 and 11). The receiving device according to an embodiment performs the same processes as the transmitting device, including LOD generation, neighbor point set generation, neighbor point weight normalization, and feature prediction. The receiving device predicts the feature in the same manner as the transmitting device, based on the information about the selected predicted feature. The receiving device decodes the received residual values and adds the decoded residual values to the predicted feature values to restore the feature values.
[0257] As described above, the LOD construction unit 2210 according to an embodiment generates LODs based on one or more of various LOD generation methods.
[0258] According to an embodiment, the LOD construction unit 2210 generates LODs based on distance (distance-based LOD generation method). The LOD construction unit 2210 sets a distance (e.g., Euclidean distance) between at least two points for each LOD, calculates the distance between all points distributed in 3D space, and generates LODs based on the calculation results. As described above, each LOD includes points distributed at regular intervals according to the level indicated by the LOD. Therefore, the LOD construction unit 2210 must calculate the distance between all points. That is, the point cloud encoder and point cloud decoder must calculate the distance between all points each time an LOD is generated, which imposes an unnecessary burden on the process of distance processing point cloud data. However, when the density of point cloud content is high, densely packed points have a high geometry-based nearby relationship, and therefore are more likely to have similar characteristics. Therefore, the LOD construction unit 2210 according to an embodiment can calculate the distance between points for densely packed point cloud content and generate LODs.
[0259] Furthermore, the LOD construction unit 2210 according to the embodiment can construct an LOD based on the Moulton code of the points (sampling LOD generation method based on the Moulton order). As described above, the Moulton code is generated by expressing the coordinate values (e.g., (x, y, z)) indicating the three-dimensional positions of all points as bit values and mixing the bits.
[0260] The LOD construction unit 2210 according to the embodiment generates a Moulton code for each point based on the reconstructed geometry and sorts the points in ascending order based on the Moulton code. The order of points sorted in ascending order of the Moulton code is called a Moulton order. The LOD construction unit 2210 performs sampling on the points sorted in the Moulton order to construct an LOD. The LOD construction unit 2210 according to the embodiment may perform sampling in various ways. For example, the LOD construction unit 2210 according to the embodiment selects points in order based on the gap in the sampling rate according to the Moulton order of the points included in each region corresponding to the node. The sampling rate (for example, k l ) is automatically changed according to the distribution of the point cloud and the content of the point cloud, or changed according to user input. In addition, the sampling rate according to the embodiment has a fixed value. The LOD construction unit 2210 selects points (e.g., the 1st to 5th points when the sampling rate is 5) aligned at a position that is the sampling rate (e.g., 5) away from the first aligned point (0th point) according to the Moulton order, and for the remaining points, selects one out of every 5 points (e.g., the point located 5th from the 5th point) and sets the selected point as the current LOD (e.g., LOD l ) smaller than LOD(LOD l-1) are classified into LODs. Therefore, the LOD construction unit 2210 can construct LODs without calculating the distances between all points. That is, the point cloud encoder and the point cloud decoder do not need to calculate the distances between all points every time an LOD is generated, and therefore, can process point cloud data more quickly. Furthermore, the LOD construction unit 2210 according to the embodiment can perform different sampling for each LOD.
[0261] However, when the density of point cloud content is low, the geometry-based nearby relationship between points is low, and therefore the probability that distributed points have similar characteristics is low. Therefore, to improve the accuracy of characteristic information encoding and decoding, the LOD construction unit 2210 according to the embodiment generates LOD by performing sampling based on one of a fixed sampling range, an octree-based fixed sampling range, and an octree-based dynamic sampling range. Information regarding the sampling range and sampling according to the embodiment is included in the above-mentioned LOD construction method (or LOD generation method) information and transmitted to a receiving device (e.g., the receiving device 10004 of FIG. 1 or the point cloud decoders of FIGS. 10 and 11) via a bitstream including the encoded point cloud video data. Therefore, the receiving device can generate LOD based on the LOD construction method information.
[0262] According to the embodiment, the LOD construction unit 2210 performs sampling based on a fixed sampling range. According to the embodiment, the sampling rate varies for each region in the 3D space depending on the density of the point cloud. Furthermore, according to the embodiment, the sampling rate varies for each LOD. According to the embodiment, the sampling range is fixed.
[0263] Figure 23 shows an example of the LOD generation process.
[0264] FIG. 23 illustrates an example LOD generation process (2300). FIG. 23 illustrates the sampling rate (e.g., kl 2 illustrates a process for generating an LOD based on a fixed sampling range when the sampling rate (expressed as 4) is 4. The description of the fixed sampling range according to the embodiment is omitted here as it has been described with reference to FIG. 22. The sampling rate according to the embodiment is changed depending on the distribution of the point cloud, user input, etc. The sampling rate according to the embodiment has a fixed value.
[0265] The upper part of Figure 23 shows the points contained in each of two spaces created by dividing a three-dimensional space. As shown, the first space 23010 contains five points P0, P1, P5, P6, and P9, and the second space 2302 contains five points P2, P3, P4, P7, and P8.
[0266] The first index 2310 in FIG. 23 indicates the original order of 10 points distributed in a three-dimensional space. As described above, the LOD construction unit (e.g., the LOD construction unit 2210 in FIG. 22) calculates the Moulton code of each point and sorts each point in ascending order (I) based on the calculated Moulton code. l ) The second index 2320 shown in FIG. 23 indicates the Moulton order of 10 points. Since the sampling rate in the example is 4, an example 2330 of points selected by sampling every 4 points, namely points P5, P1, and P7, is shown. The points not selected are the LOD. l The third index 2340 shown in Figure 23 is included only in the LOD l The index of the point is shown in the figure. The LOD component according to the embodiment l-1 The selected points P5, P1, and P7 are again sampled (or subsampled) to generate the LOD. Figure 23 shows an example 2350 of a point selected by sampling, namely P5. The points not selected (e.g., P1, P7) are l-1The fourth index 2350 shown indicates the LOD order the point belongs to. As a result of sampling, LOD0 contains P5, LOD1 contains points P5, P1, and P7, and LOD2 contains the entire point 2360.
[0267] FIG. 24 illustrates an example of a Moulton cord-based sampling process according to an embodiment.
[0268] The left side of FIG. 24 shows the Moulton order of points in one space obtained by dividing a 3D space (2400). As described above, an LOD construction unit (e.g., the LOD construction unit 2210) according to an embodiment calculates a Moulton code for each point and sorts the points in ascending order based on the calculated Moulton code. The numbers shown indicate the Moulton order of points sorted in ascending order based on the Moulton code. The right side of FIG. 24 shows an example of an LOD (2410) generated by sampling the Moulton order of points in multiple spaces generated by dividing a 3D space. The example 2410 in FIG. 24 is an example in which the sampling rate is 4. Points 2415 expressed in the coordinate system indicate points selected during the sampling process. Lines 2418 displayed in the coordinate system indicate the point selection process. As described above, an LOD construction unit (e.g., the LOD construction unit 2210) constructs an LOD by sampling the points sorted in the Moulton order. As shown in FIG. 24, the Moulton order is represented by a zigzag line. That is, the distance between points aligned in Moulton order is not constant. Also, points are not uniformly distributed in a three-dimensional area (e.g., a three-dimensional area represented by the X, Y, and Z axes). Therefore, even if sampling is performed on points aligned in Moulton order, the sampling results cannot guarantee the minimum and maximum distance between points, which affects the generation of a set of adjacent points.
[0269] According to an embodiment, the LOD constructor performs sampling based on a fixed sampling range based on an octree. According to an embodiment, the sampling rate varies depending on the depth of the octree. According to an embodiment, the sampling range is fixed. As described with reference to FIGS. 5 and 6, according to an embodiment, the octree is generated by recursively dividing the 3D space of the point cloud content into eight equal parts. According to an embodiment, the recursively divided regions have the shape of a cube or a rectangular parallelepiped with the same volume. According to an embodiment, the octree has an occupancy code indicating whether each of the eight divided spaces generated by dividing one space contains at least one point. The occupancy code includes multiple nodes, and each node indicates whether a point exists in each divided space. For example, if at least one point is included in a divided space, the node corresponding to the space indicates that the point exists (e.g., is assigned a value of 1). The depth of each node corresponds to at least one level indicated by the LOD.
[0270] The node region corresponding to each level guarantees the maximum distance between points within the LOD. In this embodiment, the maximum distance between points selected from points within a node is the node size, and the maximum distance between points selected from between nodes is limited to the node size. Therefore, sampling according to this embodiment can maintain a constant maximum distance between points, allowing the LOD construction unit to reduce the complexity of LOD generation and generate LOD with high compression efficiency.
[0271] In addition, the LOD construction unit according to the embodiment performs sampling based on an octree-based dynamic sampling range. In the embodiment, the sampling rate is set based on the LOD level and the depth of the octree. In addition, the sampling range in the embodiment is not fixed but is calculated based on the LOD level and the Moulton code value. The LOD construction unit according to the embodiment does not group (select) a number of points corresponding to each sampling rate value for sampling, but groups points for sampling based on the octree. In an octree structure, points belonging to the same parent node are obviously adjacent to each other. Therefore, the LOD construction unit according to the embodiment does not need to further calculate the distance between points to determine whether the selected points are actually adjacent. In addition, the LOD construction unit according to the embodiment checks the density of the point cloud content based on whether nodes in the occupancy code of the octree are assigned, and selects points for sampling based on the density. For example, if all nodes are assigned, it indicates that the density of points within the area indicated by each node is high, and if nodes are not used, it indicates that the density of points is low. Since densely packed points have more similar characteristics, the LOD component selects fewer points for sampling than sparsely packed points.
[0272] Based on the fixed sampling range according to the embodiment, sampling involves repeatedly selecting the same number of points according to the sampling rate value, regardless of the density of the points. The repeated point selection process is an unnecessary burden on the LOD generation process.
[0273] FIG. 25 shows an example of the LOD generation process according to an embodiment.
[0274] 25 illustrates an example of a process for generating an LOD based on an octree-based dynamic sampling range (2500). The octree-based dynamic sampling range according to the embodiment is the same as that described in FIG. 24, so a detailed description will be omitted. The octree-based dynamic sampling range according to the embodiment is set based on child points having the same parent node in the octree.
[0275] The maximum level of LOD in the embodiment is expressed as lmax. As described above, since the density of point cloud content varies, the setting of the maximum level of LOD varies for each point cloud content. Therefore, each point cloud content in the embodiment has the same or different octree depth for the maximum level of LOD. The depth of the octree in the embodiment is expressed as d. The maximum level of LOD in the embodiment is smaller than the depth of the octree. JPEG0007775351000025.jpg14151
[0276] LOD according to the embodiment (e.g., LOD l ) the maximum sampling rate is expressed as follows:
[0277]
number
[0278] k represents the maximum sampling rate. The LOD constructor according to the embodiment can change the sampling ratio according to the point density of the point cloud content. Therefore, the above equation 1 indicates sampling performed at the depth level of the octave tree (e.g., d-1 level). This is because if points are densely packed, there is a high possibility that all nodes at the d-1 level will be assigned. However, since the density differs depending on the point cloud content, the LOD constructor according to the embodiment can set a value for constructing the LOD at a specific depth of the octave tree. The LOD according to the embodiment (e.g., LOD l ) the maximum sampling rate is expressed as follows:
[0279]
number
[0280] X is a value for configuring LOD at a specific depth of the octree and has a value greater than 0. In this embodiment, X is set differently for each point cloud content.
[0281] 25 shows nodes to be sampled when X is 3 and when X is 1. When the X value is 3, the LOD construction unit according to the embodiment samples from node 2510 corresponding to depth 3 (e.g., LOD value is 0). When the X value is 1, the LOD construction unit according to the embodiment samples from node 2520 corresponding to depth 1. The number of child nodes according to the embodiment is expressed as pow(8, x) 2530.
[0282] The sampling range for each LOD according to the embodiment is expressed as follows:
[0283]
number
[0284] P i. mc indicates the Moulton code value of the i-th point. The sampling range according to the embodiment is limited to the Moulton code range that has the same upper parent node in the occupancy code of the octree. For example, the sampling range for a point with a Moulton code of 2 among points arranged in Moulton order at the maximum LOD level is set to the range including points with Moulton codes 0 to 7 (because the occupancy code has eight nodes at each depth). The LOD construction unit according to the embodiment determines whether the Moulton code value of the point is the maximum value of the sampling range (e.g., If the image size is larger than the specified size (JPEG0007775351000029.jpg1260), the next sampling range is calculated again.
[0285] Therefore, if the points are organized in an octree structure, the maximum sampling ratio for LOD0 is 8 (8 0+1 The sampling range of LOD0 is Moulton code 0 to 7, 8 to 15, 16 to 23, etc. The maximum sampling ratio of LOD1 in this example is 64 (8 1+1 =64), and the sampling range for LOD1 is 0 to 63, 64 to 127, 128 to 192, etc.
[0286] The LOD construction unit (e.g., the LOD construction unit 2210 in FIG. 22) according to the embodiment selects one of the points within the sampling range and sets the selected point to the current LOD (e.g., LOD l ) level lower than the LOD (e.g., LOD (0~-l-1) ) and the remaining unselected points are registered to the current LOD.
[0287] The LOD constructor according to an embodiment selects points based on one or more methods to generate good neighbor sets for feature predictive coding and to improve compression efficiency.
[0288] For example, the LOD construction unit according to the embodiment selects the Nth point within the sampling range. The value of N according to the embodiment is set by a user input signal. The value of N according to the embodiment is also set based on a Rate Distortion Optimization (RDO) method.
[0289] The LOD construction unit according to the embodiment selects a point having a Moulton code closest to the mean or mean Moulton code value of points belonging to the sampling range. The more similar the Moulton code values, the higher the probability that the distance between points is close. This is because a point having a Moulton code closest to the mean Moulton code value of points belonging to the same sampling range is likely to have a characteristic (e.g., color, reflectance value, etc.) that represents points belonging to the corresponding sampling range or a characteristic that has a relatively large influence on points belonging to the corresponding sampling range. In particular, when points representing peripheral values are selected as the neighbor node set, compression efficiency can be improved. In addition, the LOD construction unit according to the embodiment selects the Nth point from the point order sorted based on the mean Moulton code value. In the embodiment, N is a preset value. If there is no preset N value, the LOD construction unit according to the embodiment selects the 0th point from the sorted point order.
[0290] The LOD construction unit according to the embodiment selects a point corresponding to a median index value from among points belonging to a sampling range. The median index value according to the embodiment is highly likely to be the median value of the corresponding region. Therefore, the LOD construction unit can perform fast data processing with high compression efficiency without performing additional calculation processes. The median index value and median value according to the embodiment may be replaced with other terms having the same meaning. If points belonging to a sampling range are not uniformly distributed (e.g., if distributed to one side), the median Moulton code value of the point may not accurately represent the corresponding sampling range. Therefore, the LOD construction unit according to the embodiment selects a point having a Moulton code closest to the ideal median Moulton code value of the sampling range. The LOD construction unit according to the embodiment calculates the ideal median Moulton code value assuming that all child nodes of a parent node in an octree structure are assigned, and selects a point having a Moulton code closest to the ideal median Moulton code value. The LOD construction unit according to the embodiment selects the Nth point from the points sorted based on the ideal median Moulton code value. In the embodiment, N is a preset value. If there is no set N value, the LOD construction unit according to the embodiment selects the 0th point in the sorted point order.
[0291] According to an embodiment, the LOD construction unit selects one or more points within one sampling range. The number of points to be selected is preset. When selecting the Nth point within the sampling range, the LOD construction unit repeatedly selects the Nth point for each point within the sampling range. For example, if N is 2, the LOD construction unit repeatedly selects the second point, the fourth point, etc. within the sampling range a predetermined number X of times. When selecting a point having a Moulton code closest to the median Moulton code value of a point within the sampling range, the LOD construction unit repeatedly selects the Nth point in the order sorted by the median Moulton code value a predetermined number X of times. When selecting a point having a Moulton code closest to the ideal median Moulton code value of a point within the sampling range, the LOD construction unit repeatedly selects the Nth point in the order sorted by the ideal median Moulton code value a predetermined number X of times.
[0292] In this embodiment, a sampling range contains only one point. A point contained in a sampling range is called an isolated point. To improve compression efficiency, the LOD construction unit processes isolated points in one or more ways.
[0293] For example, the LOD component samples isolated points, selecting one point from a set of k consecutive isolated points at the fixed sampling rate mentioned above.
[0294] The LOD construction unit according to the embodiment separates isolated points into candidate groups of LODs lower than the current LOD. l During the generation of LOD, the isolated points are not selected for the sampling range that has the isolated points. (0~-l-1) The LOD configuration unit according to the embodiment is l-1-αWhen generating , a larger sampling range and sampling rate ratio are applied to select isolated points.
[0295] The LOD construction unit according to the embodiment registers the isolated point to the current LOD. For example, the LOD construction unit l During the generation of LOD, if an isolated point is generated, the isolated point is l This is because isolated points are unlikely to have characteristics representative of surrounding points, and therefore do not need to be selected as part of the set of neighboring points.
[0296] According to an embodiment, the LOD construction unit may merge isolated points with adjacent points. According to an embodiment, points are sorted in ascending order based on their Moulton code values and divided into one or more sampling ranges based on the sampling rate. As described above, points corresponding to each sampling range belong to the same parent node in the occupancy code of the octree. Therefore, the number of points included in one sampling range may be variable. According to an embodiment, the LOD construction unit may determine whether a point is an isolated point based on the sampling range, i.e., the number of points included in the sampling range. For example, if the number of points included in a sampling range is less than or equal to a predetermined value, the LOD construction unit may set the corresponding sampling range or the points included in the corresponding sampling range as isolated points. According to an embodiment, the LOD construction unit may merge (or combine) a sampling range including an isolated point with a next sampling range (e.g., a sampling range adjacent to the sampling range including the isolated point). If the number of points in the merged sampling range is less than or equal to a predetermined value, the LOD construction unit may merge the merged sampling range with a next sampling range (e.g., a sampling range adjacent to the merged sampling range). Alternatively, if the number of points in the integrated sampling range is greater than a predetermined value, the LOD construction unit sets the integrated sampling range as one sampling range.
[0297] Information about the isolated point processing method according to the embodiment is included in the above-mentioned LOD construction method (or LOD generation method) information and transmitted to a receiving device (e.g., receiving device 10004 in FIG. 1 or point cloud decoder in FIGS. 10 and 11) in a bitstream containing encoded point cloud video data. Thus, the receiving device generates LOD based on the LOD construction method information.
[0298] The point selection method according to the embodiment is also applied when selecting the above-mentioned isolated points. Information about the point selection method according to the embodiment is included in the above-mentioned LOD construction method (or LOD generation method) information and transmitted to a receiving device (e.g., the receiving device 10004 in FIG. 1 or the point cloud decoder in FIGS. 10 and 11) in a bitstream including encoded point cloud video data. Therefore, the receiving device generates LOD based on the LOD construction method information.
[0299] FIG. 26 is an example of an isolated point processing method according to an embodiment.
[0300] As described in FIG. 25 , the LOD construction unit (e.g., the LOD construction unit 2210) may merge a sampling range including an isolated point with the next sampling range. According to an embodiment, the LOD construction unit defines a predetermined value a to determine a sampling range including an isolated point. According to an embodiment, a is an integer greater than 0. According to an embodiment, if the number of points included in a sampling range is less than or equal to a, the LOD construction unit sets the corresponding sampling range or a point included in the corresponding sampling range as an isolated point. According to an embodiment, the LOD construction unit merges (or merges) the sampling range including the isolated point with the next sampling range (e.g., a sampling range adjacent to the sampling range including the isolated point). According to an embodiment, the LOD construction unit sets a maximum value r of the sampling range for sampling integration. According to an embodiment, r is an integer greater than 0 and has a value less than or equal to a. According to an embodiment, the LOD construction unit merges the integrated sampling range with the next sampling range (e.g., a sampling range adjacent to the integrated sampling range) if the number of points in the integrated sampling range is less than or equal to a. The LOD construction unit repeats the integration of the sampling ranges if the integrated sampling range is smaller than a. If the number of points in the integrated sampling range is greater than a, the LOD construction unit sets the integrated sampling range as one sampling range.
[0301] FIG. 26 illustrates an example of an isolated point processing method (2600) when a is 4 and r is 4. As described above, points are sorted in ascending order based on their Moulton code values. The illustrated index 2610 indicates the Moulton code value of a point. According to an embodiment, the LOD construction unit divides a plurality of points into one or more sampling ranges. According to an embodiment, the one or more sampling ranges include a first sampling range 2620, a second sampling range 2630, and a third sampling range 2640. As illustrated, the first sampling range 2620 includes six points (P0, P1, P2, P3, P4, and P5) that belong to the same parent node. The number of points included in the first sampling range 2620 is six, which is greater than the value of a, which is 4. Therefore, the LOD construction unit does not designate any points included in the first sampling range 2620 as isolated points, but instead selects one of the points in the first sampling range 2620 based on the sampling rate. The remaining points excluding the selected point are registered in the current LOD, and the selected point is set as a candidate point for an LOD lower than the current LOD. As shown, the second sampling range 2630 includes one point (P6). The number of points included in the second sampling range 2630 is 1, which is less than the a value of 4. Therefore, the LOD construction unit sets the points included in the second sampling range 2630 as isolated points. The LOD construction unit merges the second sampling range 2630 with the third sampling range 2640. As shown, the third sampling range 2640 includes four points (P7, P8, P9, and P10) that belong to the same parent node. The number of points in the merged sampling range obtained by merging the second sampling range 2630 and the third sampling range 2640 is 5. That is, the number of points in the merged sampling range, 5, is greater than the maximum sampling range value r, 4, for merging. Therefore, the LOD construction unit selects one point from the merged sampling range based on the sampling rate.The remaining points excluding the selected point are registered in the current LOD, and the selected point is set as a candidate point for an LOD lower than the current LOD. Although not shown, the third sampling range 2640 includes six points (P11, P12, P13, ...). The number of points included in the third sampling range 2640 is 6, which is greater than a. Therefore, the LOD construction unit does not set the points included in the third sampling range 2640 as isolated points, but instead selects one point from the points in the third sampling range 2640 according to the sampling rate.
[0302] FIG. 27 is an example of an isolated point processing method according to an embodiment.
[0303] The isolated point processing method example 2700 shown in FIG. 27 is an embodiment of the isolated point processing method example 2600 described in FIG.
[0304] FIG. 27 illustrates an example of an isolated point processing method (2700) when a is 4 and r is 4, similar to example 2600 in FIG. 26 . The illustrated index 2710 indicates the Moulton code value of the point. According to an embodiment, the LOD construction unit divides a plurality of points into one or more sampling ranges. According to an embodiment, the one or more sampling ranges include a first sampling range 2720, a second sampling range 2730, a third sampling range 2740, a fourth sampling range 2750, and a fifth sampling range 2760. As illustrated, the first sampling range 2720 includes five points (P0, P1, P2, P3, and P4). The number of points included in the first sampling range 2720 is five, which is greater than the value of a, which is 4. Therefore, the LOD construction unit does not designate the points included in the first sampling range 2720 as isolated points, but instead selects one of the points in the first sampling range 2720 based on the sampling rate. The remaining points excluding the selected point are registered in the current LOD, and the selected point is set as a candidate point for an LOD lower than the current LOD. As shown, the second sampling range 2730 includes one point (P5). Therefore, the LOD construction unit sets the points included in the second sampling range 2730 as isolated points because the number of points included in the second sampling range 2730 is 1, which is less than the a value of 4. The LOD construction unit merges the second sampling range 2730 with the third sampling range 2740. As shown, the third sampling range 2740 includes one point (P6). Therefore, the number of points in the merged sampling range obtained by merging the second sampling range 2730 and the third sampling range 2740 is 2. That is, the number of points in the merged sampling range, 2, is less than the maximum sampling range value r for the merger, 4. Therefore, the LOD construction unit merges the merged sampling range with the fourth sampling range 2750, which is the sampling range next to the merged sampling range. As shown, the fourth sampling range 2750 includes one point (P7).That is, the number of points in the integrated sampling range is 3, which is less than the maximum value r of the sampling range for integration, which is 4. Therefore, the LOD construction unit integrates the integrated sampling range with the fifth sampling range 2760, which is the sampling range next to the integrated sampling range. As shown in the figure, the fifth sampling range 2760 includes two points (P8 and P9). That is, the number of points in the integrated sampling range is 5, which is greater than the maximum value r of the sampling range for integration, which is 4. Therefore, the LOD construction unit does not set the points included in the integrated sampling range as isolated points and aborts the integration operation. The LOD construction unit selects one point from the points in the integrated sampling range based on the sampling rate. The remaining points excluding the selected point are registered in the current LOD, and the selected point is set as a candidate point for an LOD lower than the current LOD.
[0305] FIG. 28 shows an example of a structure diagram of a point cloud compression (PCC) bitstream.
[0306] As described above, a point cloud data transmitting device (e.g., the point cloud data transmitting device described in FIGS. 1, 11, 14, and 1) transmits encoded point cloud data in the form of a bitstream 2800. The bitstream 2800 according to the embodiment includes one or more sub-bitstreams.
[0307] A point cloud data transmitting apparatus (e.g., the point cloud data transmitting apparatus described in FIGS. 1, 11, 14, and 1) divides a point cloud data image into one or more packets and transmits the packets over a network, taking into account transmission channel errors. A bitstream 2800 according to an embodiment includes one or more packets (e.g., Network Abstraction Layer (NAL) units). Therefore, even if some packets are lost in a poor network environment, the point cloud data receiving apparatus can reconstruct the corresponding image using the remaining packets. Point cloud data can be divided into one or more slices or one or more tiles for processing. Tiles and slices according to an embodiment are regions for partitioning a picture of point cloud data and performing point cloud compression coding processing. The point cloud data transmitting apparatus processes data corresponding to each region of the divided point cloud data according to the importance of each region, thereby providing high-quality point cloud content. That is, the point cloud data transmitting apparatus according to an embodiment can perform point cloud compression coding processing with better compression efficiency and appropriate delay for data corresponding to regions important to a user.
[0308] In one embodiment, a tile refers to a rectangular parallelepiped in three-dimensional space (e.g., a bounding box) in which point cloud data is distributed. In another embodiment, a slice refers to a set of points that are independently coded or decoded, and is a series of syntax elements that represent part or all of the coded point cloud data. In one embodiment, a slice includes data transmitted in a packet, and includes one geometry data unit and zero or more attribute data units of the same size. In one embodiment, a tile includes one or more slices.
[0309] The bitstream 2800 according to the embodiment includes signaling information including a Sequence Parameter Set (SPS) for sequence-level signaling, a Geometry Parameter Set (GPS) for signaling geometry information coding, an Attribute Parameter Set (APS) for signaling attribute information coding, and a Tile Parameter Set (TPS) for tile-level signaling, and one or more slices.
[0310] The SPS according to the embodiment is coding information for the entire sequence, such as profile and level, and includes comprehensive information for the entire file, such as picture resolution and video format.
[0311] According to the embodiment, one slice (e.g., slice0 in FIG. 28) includes a slice header and slice data. The slice data includes one geometry bitstream (Geom0 0 ) and one or more attribute bitstreams (Attr0 0 , Attr1 0 ). The geometry bitstream includes a header (e.g., geometry slice header) and a payload (e.g., geometry slice data). The header of the geometry bitstream according to the embodiment includes identification information of the parameter set included in the GPS (geom_geom_parameter_set_id), a tile identifier (geom_tile id), a slice identifier (geom_slice_id), and information about the data included in the payload. The attribute bitstream includes a header (e.g., attribute slice header or attribute brick header) and a payload (e.g., attribute slice data or attribute brick data).
[0312] As described in Figures 22 to 27, the point cloud encoder and the point cloud decoder according to the embodiment generate LOD for feature prediction. Therefore, the bitstream shown in Figure 28 includes the LOD construction method information described in Figures 18 to 27. The point cloud decoder according to the embodiment generates LOD based on the LOD construction method information.
[0313] The signaling information included in the bitstream according to the embodiment is generated by a metadata processing unit or a transmission processing unit (e.g., the transmission processing unit 12012 in FIG. 12) included in the point cloud encoder or an element within the metadata processing unit or the transmission processing unit. The signaling information according to the embodiment is generated based on the results of geometry encoding and attribute encoding.
[0314] FIG. 29 shows an example of syntax for APS according to an embodiment.
[0315] FIG. 29 is an example of syntax for an APS according to an embodiment, which includes the following information (or fields, parameters, etc.):
[0316] aps_attr_parameter_set_id indicates the identifier of the APS for reference by other syntax elements. aps_attr_parameter_set_id has a value in the range of 0 to 15. As shown in Figure 30, since one or more attribute bitstreams are included in a bitstream, the header of each attribute bitstream contains a field (e.g., ash_attr_parameter_set_id) with the same value as aps_attr_parameter_set_id.
[0317] The point cloud decoder according to the embodiment secures an APS corresponding to each feature bitstream based on aps_attr_parameter_set_id and processes the feature bitstream.
[0318] aps_seq_parameter_set_id indicates the value of sps_seq_parameter_set_id for the active SPS. aps_seq_parameter_set_id has a value in the range of 0 to 15.
[0319] attr_coding_type indicates the coding type for a given value of attr_coding_type. The value of attr_coding_type is equal to 0, 1, or 2 in the bitstream according to the embodiment. Other values of attr_coding_type may be used in the future by ISO / IEC. Therefore, a point cloud decoder according to the embodiment may ignore values of attr_coding_type other than 0, 1, or 2. If the value is 0, the coding type is predictive weight lifting transform coding; if the value is 1, the coding type is RAHT transform coding; and if the value is 2, the coding type is fixed weight lifting transform coding.
[0320] The following are the parameters when the characteristic coding type is 0 or 2:
[0321] num_pred_nearest_neighbours indicates the maximum number of nearest neighbors used for prediction, and has a value ranging from 0 to a specified value.
[0322] max_num_direct_predictors indicates the number of predictors used for direct prediction conversion. max_num_direct_predictors has a value in the range from 0 to the value of num_pred_nearest_neighbours. The values of the variable MaxNumPredictors used in the decoding operation are as follows:
[0323] MaxNumPredictors=max_num_direct_predicots+1
[0324] lifting_search_range indicates the search range for lifting.
[0325] lifting_quant_step_size indicates the quantization step size for the first component of the attribute. The value of lifting_quant_step_size ranges from 1 to xx (any value).
[0326] lifting_quant_step_size_chroma indicates the quantization step size for the chroma component of the attribute if the attribute is color. lifting_quant_step_size_chroma has a range from 1 to xx (any value).
[0327] lod_binary_tree_enabled_flag indicates whether binary trees are applied in the LOD generation process.
[0328] num_detail_levels_minus1 indicates the number of LODs for feature coding. num_detail_levels_minus1 has a value in the range of 0 to xx (any value). The following for loop is information about each LOD.
[0329] sampling_distance_squared[idx] indicates the square of the sampling distance for each LOD represented by idx. In the embodiment, idx has a value from 0 to the number of LODs for the feature coding indicated by num_detail_levels_minus1. sampling_distance_squared has a value in the range of 0 to xx (any value).
[0330] When the same sampling information is applied to point cloud data transmitted by one bitstream, the syntax for APS according to the embodiment includes LOD configuration information.
[0331] The following is LOD configuration information 2900 according to an embodiment.
[0332] lod_generation_type is information indicating the type of method for constructing (generating) an LOD set. lod_generation_type has a value of either 1 or 2, and the value of lod_generation_type is not limited to this example. When the value of lod_generation_type is 1, lod_generation_type indicates that the LOD set generation method is a distance-based LOD generation method (for example, the distance-based LOD generation method described in FIG. 22). When the value of lod_generation_type is 2, lod_generation_type indicates that the LOD set generation method is a Moulton order-based sampling LOD generation method (for example, the Moulton order-based sampling LOD generation method described in FIG. 22). The LOD set generation method is as described in FIG. 22, so a detailed description thereof will be omitted.
[0333] The following is information related to sampling when the value of lod_generation_type is 2:
[0334] sampling_range_type indicates the sampling range and sampling rate, or the type of sampling range. sampling_range_type has a value of 1, 2, or 3, and the value of sampling_range_type is not limited to this example. When the value of sampling_range_type is 1, sampling_range_type indicates that the sampling range is a fixed sampling range (e.g., the fixed sampling range described in FIGS. 22 and 23). When the value of sampling_range_type is 2, sampling_range_type indicates that the sampling range is an octree-based fixed sampling range (e.g., the octree-based fixed sampling range described in FIGS. 22 and 24). When the value of sampling_range_type is 3, sampling_range_type indicates that the sampling range is an octree-based dynamic sampling range (e.g., the octree-based dynamic sampling range described in FIGS. 22 and 25). The sampling range and sampling rate according to the embodiment are as described in FIGS. 22 to 25, so detailed description will be omitted.
[0335] sampling_rate indicates the sampling rate. The value of sampling_rate is an integer greater than 0.
[0336] sampling_select_type indicates a point selection method for selecting a point within the sampling range. sampling_select_type has one of the values 1, 2, 3, and 4, and the value of sampling_select_type is not limited to this example. When the value of sampling_select_type is 1, sampling_select_type indicates that the point selection method is a method for selecting the Nth point within the sampling range (e.g., a method for selecting the Nth point within the sampling range described in FIG. 25). When the value of sampling_select_type is 2, sampling_select_type indicates that the point selection method is a method for selecting a point having a Moulton code closest to the median Moulton code value of a point belonging to the sampling range (e.g., a method for selecting a point having a Moulton code closest to the median Moulton code value of a point belonging to the sampling range described in FIG. 25). When the value of sampling_select_type is 3, sampling_select_type indicates that the point selection method is a method for selecting a point corresponding to an intermediate index value among points belonging to the sampling range (e.g., a method for selecting a point corresponding to an intermediate index value among points belonging to the sampling range described in FIG. 25). If the value of sampling_select_type is 4, sampling_select_type indicates that the point selection method is a method of selecting a point having a Moulton code closest to the ideal median Moulton code value of the sampling range (for example, a method of selecting a point having a Moulton code closest to the ideal median Moulton code value of the sampling range described in FIG. 25). The point selection method according to this embodiment is the same as the point selection method described in FIG. 25, so a detailed description will be omitted.
[0337] sampling_select_idx indicates the fixed index of the point to be selected. For example, if the value of sampling_select_type is 1, sampling_select_idx indicates the index of the Nth point selected from the points aligned within the sampling range. Also, if the value of sampling_select_type is 2 or 3, sampling_select_idx indicates the index of the Nth point selected from the points aligned based on the median Moulton code value or the ideal median Moulton code value.
[0338] sampling_select_max_num_of_points indicates the maximum number of points that can be selected within one sampling range.
[0339] The following shows the sampling information when the value of sampling_select_type is greater than or equal to 2. sampling_begin_depth indicates the depth of the octree at which LOD0 set generation begins. sampling_begin_depth has a positive integer value.
[0340] sampling_isolated_point_mininum_num_of_elements indicates the number of points that can be defined as isolated points (for example, the value a that is preset to determine the sampling range including the isolated points described in Figures 26 and 27). In other words, sampling_isolated_point_mininum_num_of_elements indicates the minimum number of points that belong to the sampling range.
[0341] sampling_isolated_point_sampling_type indicates the method for processing isolated points. sampling_isolated_point_sampling_type indicates one of the following: sampling isolated points, merging isolated points with adjacent points, separating isolated points into a candidate group with an LOD lower than the current LOD and processing them, or registering isolated points to the current LOD. The method for sampling isolated points is to select one point from k consecutive isolated points at a fixed sampling rate. The method for merging isolated points with adjacent points is to compare the number of points included in the sampling range with the number of points that can be defined as isolated points indicated by sampling_isolated_point_mininum_num_of_elements, and if it is an isolated point, to merge the sampling range containing the isolated point with the next sampling range. The method for separating isolated points into a candidate group with an LOD lower than the current LOD and processing them is to select one point from k consecutive isolated points at a fixed sampling rate. l During the generation of LOD, the isolated points are not selected for the sampling range that has the isolated points. (0~-l-1) Separate into LOD l-1-α This is a method to select isolated points when generating. The method to register isolated points to the current LOD is l During the generation of LOD, if an isolated point is generated, the isolated point is l The isolated point processing method is the same as the isolated point processing method explained with reference to Figures 25 to 27, so a detailed explanation will be omitted.
[0342] The sampling_isolated_point_maximum_merge_range indicates the maximum sampling range for sampling integration (for example, the maximum value r of the sampling range for sampling integration described in FIGS. 26 and 27).
[0343] The following are the relevant parameters when the Attribute Coding Type value is 0:
[0344] adaptive_prediction_threshold indicates the prediction threshold.
[0345] The following are the relevant parameters when the Attribute Coding Type value is 1:
[0346] raht_depth indicates the number of LODs for RAHT conversion. depthRAHT has a value in the range of 1 to xx (any value).
[0347] rath_quant_step_size indicates the quantization step size for the first component of the attribute. rath_quant_step_size has a value in the range of 1 to xx (any value).
[0348] raht_quant_step_size_chroma indicates the quantization step size for the chroma component of a quality when applying RAHT.
[0349] aps_extension_present_flag is a flag having a value of 0 or 1.
[0350] A value of 1 in aps_extension_present_flag indicates the presence of an aps_extension_data syntax structure within the APS RBSP syntax structure. A value of 0 in aps_extension_present_flag indicates the absence of this syntax structure. If the syntax structure does not exist, the value of aps_extension_present_flag is inferred to be 0.
[0351] aps_extension_data_flag can have any value. The presence and value of this field does not affect decoder performance according to the embodiment.
[0352] FIG. 30 shows an example of syntax for APS according to an embodiment.
[0353] FIG. 30 shows an example of the syntax for the APS described in FIG.
[0354] Figure 30 shows an example of syntax for APS when equal or different sampling information is applied to each LOD of point cloud data transmitted by one bitstream. Therefore, the same information and / or parameters as those described in Figure 29 will not be described again.
[0355] An LOD configuration unit (e.g., LOD configuration unit 2210) according to an embodiment performs equal or different sampling for each LOD. Therefore, in this case, the syntax for APS according to an embodiment further includes the following information 3000 related to sampling when the value of lod_generation_type is 2:
[0356] Sampling_attrs_per_lod_flag is a flag indicating whether the sampling execution method differs for each LOD. As described in FIG. 22, the LOD construction unit (e.g., the LOD construction unit 2210) performs different sampling for each LOD, so the LOD generation information according to the embodiment includes sampling information for each LOD. Therefore, when the value of sampling_attrs_per_lod_flag is 1, the following for loop is information 3010 related to sampling for each LOD. The illustrated idx indicates each LOD. The sampling related information 3010 according to the embodiment is the same as the sampling related information described in FIG. 29, so a detailed description will be omitted.
[0357] According to the embodiment, the LOD construction unit (e.g., the LOD construction unit 2210) may apply a different LOD generation method for each tile or slice. As described above, since the point cloud data receiving apparatus also generates LOD, the bitstream according to the embodiment further includes signaling information related to the LOD generation method for each region.
[0358] FIG. 31 shows an example of syntax for a TPS according to an embodiment.
[0359] If the LOD configuration unit according to the embodiment applies different LOD generation methods to each tile, the TPS according to the embodiment further includes LOD configuration information (e.g., LOD configuration information 2900, 3000, 3010 described in FIGS. 29 and 30). Figure 31 shows an example of syntax for a TPS according to the embodiment, which includes the following information (or fields, parameters, etc.):
[0360] num_tiles indicates the number of tiles signaled for the bitstream. If there are no tiles signaled for the bitstream, the value of this information is inferred to be 0. The signaling parameters for each tile are as follows:
[0361] tile_bounding_box_offset_x[i] indicates the x offset of the ith tile in the Cartesian coordinate system. If this parameter is not present, the value of tile_bounding_box_offset_x[0] is assumed to be the value of sps_bounding_box_offset_x contained in the SPS.
[0362] tile_bounding_box_offset_y[i] indicates the y offset of the ith tile in the Cartesian coordinate system. If this parameter is not present, the value of tile_bounding_box_offset_y[0] is assumed to be the value of sps_bounding_box_offset_y contained in the SPS.
[0363] tile_bounding_box_offset_z[i] indicates the z offset of the ith tile in the Cartesian coordinate system. If this parameter is not present, the value of tile_bounding_box_offset_z[0] is assumed to be the value of sps_bounding_box_offset_z.
[0364] tile_bounding_box_scale_fwactor[i] indicates the scale factor associated with the ith tile in the Cartesian coordinate system. If this parameter is not present, the value of tile_bounding_box_scale_fwactor[0] is assumed to be the value of sps_bounding_box_scale_fwactor.
[0365] tile_bounding_box_size_width[i] indicates the width of the ith tile in the Cartesian coordinate system. If this parameter is not present, the value of tile_bounding_box_size_width[0] is assumed to be the value of sps_bounding_box_size_width.
[0366] tile_bounding_box_size_height[i] indicates the height of the ith tile in the Cartesian coordinate system. If this parameter is not present, the value of tile_bounding_box_size_height[0] is assumed to be the value of sps_bounding_box_size_height.
[0367] tile_bounding_box_size_depth[i] indicates the depth of the ith tile in the Cartesian coordinate system. If this parameter is not present, the value of tile_bounding_box_size_depth[0] is assumed to be the value of sps_bounding_box_size_depth.
[0368] As shown, the TPS according to the embodiment includes LOD configuration information 2900. LOD configuration information 3100 according to the embodiment is applied to each tile. The LOD configuration information 3100 is the same as the LOD configuration information 2900 described in FIG. 29, so a detailed description thereof will be omitted.
[0369] FIG. 32 shows an example of syntax for a TPS according to an embodiment.
[0370] Figure 32 shows an example of syntax for the TPS described in Figure 30. Figure 32 shows an example of syntax for the TPS when equal or different sampling information is applied to each LOD of each tile of point cloud data transmitted by one bitstream. Descriptions of information and / or parameters that are the same as those described in Figure 30 will be omitted.
[0371] An LOD construction unit (e.g., LOD construction unit 2210) according to an embodiment performs equal or different sampling for each LOD. Therefore, in this case, the syntax for a TPS according to an embodiment further includes sampling_attrs_per_lod_flag 3200 related to sampling when the value of lod_generation_type is 2. Since sampling_attrs_per_lod_flag according to an embodiment is the same as sampling_attrs_per_lod_flag 3000 described with reference to FIG. 30, a detailed description thereof will be omitted. Also, when the value of sampling_attrs_per_lod_flag is 1, the syntax for a TPS according to an embodiment includes a for loop indicating information 3210 related to sampling for each LOD. Since information 3210 related to sampling for each LOD according to an embodiment is the same as information 3010 related to sampling for each LOD described with reference to FIG. 30, a detailed description thereof will be omitted.
[0372] FIG. 33 shows an example of syntax for an attribute header according to an embodiment.
[0373] The syntax for the attribute header in FIG. 33 is an example of the syntax of information conveyed by the header included in the attribute bitstream described in FIG.
[0374] If the LOD construction unit according to the embodiment (e.g., the LOD construction unit 2110) applies a different neighbor point set generation method for each slice, the attribute header according to the embodiment further includes LOD configuration information 3300 (e.g., the LOD configuration information 2900 described in FIG. 29 or the LOD configuration information 3100 described in FIG. 31). An example of syntax for the attribute header according to the embodiment shown in FIG. 33 includes the following information (or fields, parameters, etc.):
[0375] The ash_attr_parameter_set_id has the same value as the aps_attr_parameter_set_id of the active APS.
[0376] ash_attr_sps_attr_idx indicates the value of sps_seq_parameter_set_id for the active SPS. The value of ash_attr_sps_attr_idx ranges from 0 to the value of sps_num_attribute_sets contained within the active SPS.
[0377] ash_attr_geom_slice_id indicates the value of the geometry slice ID (e.g., geom_slice_id).
[0378] As shown, the feature header according to the embodiment includes LOD configuration information 3300. The LOD configuration information 3300 according to the embodiment is applied to each feature bitstream (or feature slice data) belonging to each slice. The LOD configuration information 3300 is the same as the LOD configuration information 2900 and 3100 described in FIGS. 29 and 31, so a detailed description thereof will be omitted.
[0379] FIG. 34 shows an example of syntax for an attribute header according to an embodiment.
[0380] Figure 34 shows an example of the syntax for the attribute header described in Figure 33. Figure 34 shows an example of the syntax for the attribute header when equal or different sampling information is applied for each LOD corresponding to each slice.
[0381] An LOD construction unit (e.g., the LOD construction unit 2210) according to an embodiment performs equal or different sampling for each LOD. Therefore, in this case, the syntax for the attribute header according to an embodiment further includes sampling_attrs_per_lod_flag 3400 related to sampling when the value of lod_generation_type is 2. Since sampling_attrs_per_lod_flag according to an embodiment is the same as sampling_attrs_per_lod_flag 3000 and 3200 described with reference to FIGS. 30 and 32, detailed description thereof will be omitted. Also, when the value of sampling_attrs_per_lod_flag is 1, the syntax for the attribute header according to an embodiment includes a for loop indicating information 3410 related to sampling for each LOD. Since the information 3410 related to sampling for each LOD according to an embodiment is the same as the information 3010 and 3210 related to sampling for each LOD described with reference to FIGS. 30 and 32, detailed description thereof will be omitted.
[0382] FIG. 35 is a block diagram illustrating an example of a point cloud decoder.
[0383] The point cloud decoder 3500 according to the embodiment performs a decoding operation that is the same as or similar to the decoding operation of the decoders described in FIGS. 1 to 17 (e.g., the point cloud decoders described in FIGS. 1, 10, 11, 13, 14, and 16). The point cloud decoder 3500 also performs a decoding operation that is the reverse of the encoding operation of the point cloud encoder 1800 described in FIG. 18. The point cloud decoder 3500 according to the embodiment includes a spatial divider 3510, a geometry information decoder 3520 (or geometry information decoder), and a feature information decoder (or feature decoder) 3530. The point cloud decoder 3300 according to the embodiment further includes one or more elements for performing the decoding operations described in FIGS. 1 to 17, although these elements are not shown in FIG.
[0384] The space divider 3510 according to the embodiment divides the space based on signaling information (e.g., information on the division operation performed by the space divider 1810 described in FIG. 18) received from a point cloud data transmitting apparatus according to the embodiment (e.g., the point cloud data transmitting apparatus described in FIGS. 1, 11, 14, and 1) or division information derived (generated) by the point cloud decoder 3500. As described above, the division operation of the space divider 1810 of the point cloud encoder 1800 is based on any one of an octree, a quadtree, a binary tree, a triple tree, and a kd tree.
[0385] The geometry information decoding unit 3520 according to the embodiment decodes the input geometry bitstream to restore geometry information. The restored geometry information is input to the characteristic information decoding unit 3530. The geometry information decoding unit 3520 according to the embodiment performs the operations of the arithmetic decoder 12000, the octree synthesis unit 12001, the surface approximation synthesis unit 12002, the geometry reconstruction unit 12003, and the inverse transform coordinates unit 12004 described in FIG. 12. The geometry information decoding unit 3520 according to the embodiment also performs the operations of the arithmetic decoder 13002, the proprietary code-based octree reconstruction processing unit 13003, the surface model processing unit (triangle reconstruction, upsampling, voxelization) 13004, and the inverse quantization processing unit 13005 described in FIG. 13. Alternatively, the geometry information decoding unit 3520 according to the embodiment performs the operation of the point cloud decoding unit described with reference to FIG.
[0386] The feature information decoding unit 3530 according to the embodiment decodes feature information based on the feature information bitstream and the reconstructed geometry information. The feature information decoding unit 3530 according to the embodiment performs operations identical to or similar to the operations of the arithmetic decoder 11005, the inverse quantization unit 11006, the RAHT transform unit 11007, the LOD generation unit 11008, the inverse lift unit 11009, and / or the color inverse transform unit 11010 included in the point cloud decoder of Fig. 11. The feature information decoding unit 3530 according to the embodiment performs operations identical to or similar to the operations of the arithmetic decoder 13007, the inverse quantization processing unit 13008, the prediction / lift / RAHT inverse transform processing unit 13009, and the hue inverse transform processing unit 13010 included in the receiving device of Fig. 13.
[0387] The point cloud decoder 3500 outputs the final PCC data based on the recovered geometry information and the recovered feature information.
[0388] FIG. 36 is a block diagram showing an example of a geometry information decoder.
[0389] The geometry information decoder 3600 according to the embodiment is an example of the geometry information decoding unit 3520 of Figure 35 and performs operations that are the same as or similar to those of the geometry information decoding unit 3520. The geometry information decoder 3600 according to the embodiment performs a decoding operation that corresponds to the inverse process of the encoding operation of the geometry information encoder 1900 described in Figure 19. The geometry information decoder 3600 according to the embodiment includes a geometry information entropy decoding unit 3610, a geometry information inverse quantization unit 3620, a geometry information prediction unit 3630, a filtering unit 3640, a memory 3650, a geometry information inverse transformation / inverse quantization unit 3660, and a coordinate system inverse transformation unit 3670. Although not shown in Figure 36, the geometry information decoder 3600 according to the embodiment further includes one or more elements for performing the geometry decoding operations described in Figures 1 to 35.
[0390] The geometry information entropy decoding unit 3610 according to the embodiment entropy decodes the geometry information bitstream to generate quantized residual geometry information. The geometry information entropy decoding unit 3610 performs an entropy decoding operation, which is the reverse process of the entropy encoding operation performed by the geometry information entropy encoding unit 1905 described in FIG. 19. As described above, the entropy encoding operation according to the embodiment includes Exponential Golomb, CAVLC (Context-Adaptive Variable Length Coding), CABAC (Context-Adaptive Binary Arithmetic Coding), etc., and the entropy decoding operation includes Exponential Golomb, CAVLC (Context-Adaptive Variable Length Coding), CABAC (Context-Adaptive Binary Arithmetic Coding), etc. corresponding to the entropy encoding operation. The geometry information entropy decoding unit 3610 according to the embodiment decodes information related to geometry coding included in the geometry information bitstream, such as information related to predicted geometry information generation, information related to quantization (e.g., quantization values, etc.), and signaling information related to coordinate system transformation.
[0391] The residual geometry information inverse quantization unit 3620 according to the embodiment performs an inverse quantization operation on the quantized residual geometry information based on information related to quantization to generate residual geometry information or geometry information.
[0392] The geometry information prediction unit 3630 according to the embodiment generates predicted geometry information based on information related to generation of predicted geometry information output from the geometry entropy decoding unit 3610 and previously decoded geometry information stored in the memory 3650. The geometry information prediction unit 3630 according to the embodiment includes an inter prediction unit and an intra prediction unit. The inter prediction unit according to the embodiment performs inter prediction on the current prediction unit based on information included in one of a previous space and a subsequent space of the current space (e.g., a frame, a picture, etc.) including the current prediction unit, based on information necessary for inter prediction of the current prediction unit (e.g., a node, etc.) provided by a geometry information encoder (e.g., the geometry information encoder 1900). The intra prediction unit according to the embodiment generates predicted geometry information based on geometry information of a point in the current space based on information related to intra prediction of the prediction unit provided by the geometry information encoder 1900.
[0393] The filtering unit 3640 according to the embodiment filters reconstructed geometry information generated by combining predicted geometry information generated based on filtering-related information and reconstructed residual geometry information. The filtering-related information according to the embodiment may be signaled from the geometry information encoder 1900. Alternatively, the geometry information decoder 3600 according to the embodiment may derive and calculate the filtering-related information during a decoding process.
[0394] The memory 3650 according to the embodiment stores the reconstructed geometry information. The geometry inverse transform and quantization unit 3660 according to the embodiment inverse transforms and quantizes the reconstructed geometry information stored in the memory 3650 based on the quantization-related information.
[0395] The coordinate system inverse transformation unit 3670 according to the embodiment inversely transforms the coordinate system of the inversely transformed and quantized geometry information based on the coordinate system transformation related information provided from the geometry information entropy decoding unit 3610 and the restored geometry information stored in the memory 3750, and outputs the geometry information.
[0396] FIG. 37 is a block diagram illustrating an example of a characteristic information decoder.
[0397] The feature information decoder 3700 according to the embodiment is an example of the feature information decoding unit 3530 of Figure 35, and performs operations that are the same as or similar to those of the feature information decoding unit 3530. The feature information decoder 3700 according to the embodiment performs a decoding operation that corresponds to the inverse process of the encoding operation of the feature information encoder 2000 and the feature information encoder 2100 described with reference to Figures 20 and 21. The feature information decoder 3700 according to the embodiment includes a feature information entropy decoding unit 3710, a geometry information mapping unit 3720, a residual feature information dequantization unit 3730, a residual feature information inverse transform unit 3740, a memory 3750, a feature information prediction unit 3760, and a feature information inverse transform unit 3770. The feature information decoder 3700 according to the embodiment further includes one or more elements for performing feature decoding operations described with reference to Figures 1 to 36, although these elements are not shown in Figure 37.
[0398] The feature information entropy decoding unit 3710 receives the feature information bitstream and performs entropy decoding to generate transformed and quantized feature information.
[0399] The geometry information mapping unit 3720 maps the transformed and quantized feature information and the restored geometry information to generate residual feature information.
[0400] The residual characteristic information inverse quantization unit 3730 inverse quantizes the residual characteristic information based on the quantization value.
[0401] The residual feature information inverse transform unit 3740 inverse transforms the residual 3D block including the inversely quantized feature information by performing transform coding such as DCT, DST, SADCT, RAHT, etc.
[0402] The memory 3750 stores the inversely transformed attribute information together with the predicted attribute information output by the attribute information prediction unit 3760. Alternatively, the memory 3750 does not inversely transform the attribute information but stores it together with the predicted attribute information.
[0403] The feature information prediction unit 3760 generates predicted feature information based on the feature information stored in the memory 3750. The feature information prediction unit 3760 generates predicted feature information by performing entropy decoding. The feature information prediction unit 3760 also performs operations that are the same as or similar to the operations of the feature information prediction unit 2200 described in Figure 22. The feature prediction unit 3760 also generates predicted feature information by generating the LOD or LOD set described in Figures 22 to 27 based on the LOD configuration information described in Figures 29 to 34.
[0404] The feature information inverse transform unit 3770 receives the feature information type and transformation information from the feature information entropy decoding unit 3710 and performs various color inverse transform coding.
[0405] FIG. 38 is an example flowchart of a method for processing point cloud data according to an embodiment.
[0406] Flowchart 3800 in Figure 38 shows a point cloud data processing method of a point cloud data processing device (for example, the point cloud data transmission device or point cloud data encoder described in Figures 1, 11, 14-15, and 18-22). The point cloud data processing device according to the embodiment performs operations that are the same as or similar to the encoding operations described in Figures 1 to 37.
[0407] The point cloud data processing device according to the embodiment encodes 3810 the point cloud data including geometry information and attribute information. The geometry information according to the embodiment is information indicating the positions of the points of the point cloud data. The attribute information according to the embodiment is information indicating the attributes of the points of the point cloud data.
[0408] A point cloud data processing device according to an embodiment encodes geometry information and encodes feature information. The point cloud data processing device according to an embodiment performs an operation that is the same as or similar to the geometry information encoding operation described with reference to FIGS. 1 to 37. The point cloud data processing device also performs an operation that is the same as or similar to the feature information encoding operation described with reference to FIGS. 1 to 37. The point cloud data processing device according to an embodiment generates at least one LOD by dividing points. The LOD generation method according to the embodiment is the same as or similar to the LOD generation method described with reference to FIGS. 22 to 27, and therefore a detailed description thereof will be omitted.
[0409] An embodiment of a point cloud data processing apparatus transmits 3820 a bitstream including the encoded point cloud data.
[0410] The structure of the bitstream according to the embodiment is the same as that described in Fig. 28, and therefore a detailed description thereof will be omitted. The bitstream according to the embodiment includes LOD configuration information (e.g., the LOD configuration information described in Figs. 29 to 34). Furthermore, the LOD configuration information according to the embodiment is transmitted to the receiving device by the APS, TPS, attribute header, etc., as described in Figs. 29 to 34.
[0411] According to an embodiment, the LOD configuration information includes type information indicating the type of LOD generation method (e.g., LOD_generation_type described in FIGS. 27 to 32). The type information according to an embodiment indicates either a first type that generates LOD based on the distance between points (e.g., the distance-based LOD generation method described in FIG. 22) or a second type that generates LOD by sampling based on the Moulton code of points (e.g., the Moulton order-based sampling LOD generation method described in FIG. 22). The LOD set generation method is as described in FIG. 22, so a detailed description thereof will be omitted.
[0412] When the type information according to the embodiment indicates the second type, the LOD configuration information includes sampling range type information (e.g., sampling_range_type described in FIGS. 29 to 34), sampling rate information (e.g., sampling_rate described in FIGS. 29 to 34), fixed index information of points selected by sampling (e.g., sampling_select_idx described in FIGS. 29 to 34), point selection method information for selecting points within the sampling range (e.g., sampling_select_type described in FIGS. 29 to 34), and information on the maximum number of points selected by sampling (e.g., sampling_select_max_num_of_points described in FIGS. 27 to 32). The sampling range type information according to the embodiment has any of values from 1 to 4. If the value of the sampling range type information according to the embodiment is greater than or equal to 2, the LOD configuration information further includes sampling information related to one or more points defined as isolated points (e.g., sampling_begin_depth, sampling_isolated_point_minimum_num_of_elements, and sampling_isolated_point_sampling_type described in FIG. 29). The LOD configuration information according to the embodiment is the same as that described in FIGS. 29 to 34, and therefore detailed description thereof will be omitted.
[0413] FIG. 39 is an example of a flowchart of a method for processing point cloud data according to an embodiment.
[0414] Flowchart 3900 in Figure 39 shows a point cloud data processing method of a point cloud data processing device (e.g., a point cloud data receiving device or a point cloud data decoder described in Figures 1, 13, 14, 16, and 35 to 37). The point cloud data processing device according to the embodiment performs operations that are the same as or similar to the decoding operations described in Figures 1 to 35.
[0415] An apparatus for processing point cloud data according to an embodiment receives 3910 a bitstream including point cloud data.
[0416] The point cloud data processing device according to the embodiment decodes the point cloud data (3920). The decoded point cloud data according to the embodiment includes geometry information and attribute information. The geometry information is information indicating the positions of the points of the point cloud data. The attribute information according to the embodiment is information indicating the attributes of the points of the point cloud data. The structure of the bitstream according to the embodiment is the same as that described in FIG. 28, so a detailed description will be omitted.
[0417] The point cloud data processing device according to the embodiment decodes geometry information and decodes feature information. The point cloud data processing device according to the embodiment performs an operation that is the same as or similar to the geometry information decoding operation described with reference to FIGS. 1 to 37. The point cloud data processing device also performs an operation that is the same as or similar to the feature information decoding operation described with reference to FIGS. 1 to 37. The point cloud data processing device according to the embodiment divides points to generate at least one LOD. The LOD generation method according to the embodiment is the same as or similar to the LOD generation method described with reference to FIGS. 22 to 27, so detailed description thereof will be omitted.
[0418] The structure of the bitstream according to the embodiment is the same as that described in Figure 28, so a detailed description will be omitted. The bitstream according to the embodiment includes LOD configuration information (for example, the LOD configuration information described in Figures 29 to 34). In addition, the neighbor point set generation information according to the embodiment is transmitted to the receiving device by the APS, TPS, characteristic header, etc., as described in Figures 29 to 34.
[0419] According to an embodiment, the LOD configuration information includes type information indicating the type of LOD generation method (e.g., LOD_generation_type described in FIGS. 29 to 34). The type information according to an embodiment indicates either a first type that generates LOD based on the distance between points (e.g., the distance-based LOD generation method described in FIG. 22) or a second type that generates LOD by sampling based on the Moulton code of points (e.g., the Moulton order-based sampling LOD generation method described in FIG. 22). The LOD set generation method is as described in FIG. 22, so a detailed description thereof will be omitted.
[0420] When the type information according to the embodiment indicates the second type, the LOD configuration information includes sampling range type information (e.g., sampling_range_type described in FIGS. 29 to 34), sampling rate information (e.g., sampling_rate described in FIGS. 29 to 34), fixed index information of points selected by sampling (e.g., sampling_select_idx described in FIGS. 29 to 34), point selection method information for selecting points within the sampling range (e.g., sampling_select_type described in FIGS. 29 to 34), and information on the maximum number of points selected by sampling (e.g., sampling_select_max_num_of_points described in FIGS. 29 to 34). The sampling range type information according to the embodiment has any of values from 1 to 4. If the value of the sampling range type information according to the embodiment is greater than or equal to 2, the LOD configuration information further includes sampling information related to one or more points defined as isolated points (e.g., sampling_begin_depth, sampling_isolated_point_minimum_num_of_elements, and sampling_isolated_point_sampling_type described in FIG. 29). The LOD configuration information according to the embodiment is the same as that described in FIGS. 29 to 34, and therefore detailed description thereof will be omitted.
[0421] The components of the point cloud data processing apparatus according to the embodiments described in FIGS. 1 to 39 may be implemented as hardware, software, firmware, or a combination thereof, including one or more processors coupled to a memory. The components of the device according to the embodiments may be implemented as a single chip, for example, a single hardware circuit. The components of the point cloud data processing apparatus according to the embodiments may be implemented as separate chips. Any of the components of the point cloud data processing apparatus according to the embodiments may be configured with one or more processors capable of executing one or more programs, and the one or more programs may include instructions for causing or executing any one or more operations / methods of the point cloud data processing apparatus described in FIGS. 1 to 39.
[0422] For convenience of explanation, the figures have been described separately. However, it is possible to combine the embodiments described in the figures to realize a new embodiment. Furthermore, designing a computer-readable recording medium on which a program for executing the previously described embodiments is recorded, as needed by those skilled in the art, is also within the scope of the embodiments. As described above, the apparatus and method according to the embodiments are not limited to the configurations and methods of the described embodiments. The embodiments may be configured by selectively combining all or part of each embodiment, allowing for various modifications. While preferred embodiments have been shown and described, the embodiments are not limited to the specific embodiments described above. Various modifications may be made by those skilled in the art to which the invention pertains without departing from the spirit and scope of the embodiments claimed in the claims. Such modifications should not be interpreted as being separate from the technical ideas and perspectives of the embodiments.
[0423] The descriptions of the apparatus and method according to the embodiments may be applied to complement each other. For example, the point cloud data transmitting method according to the embodiments is performed by the point cloud data transmitting apparatus according to the embodiments or components included in the point cloud data transmitting apparatus. Also, the point cloud data receiving method according to the embodiments is performed by the point cloud data receiving apparatus according to the embodiments or components included in the point cloud data receiving apparatus.
[0424] Various components of the apparatus according to the embodiments may be implemented by hardware, software, firmware, or a combination thereof. Various components of the embodiments may be implemented by a single chip, e.g., a single hardware circuit. In some embodiments, components according to the embodiments may be implemented by individual chips. In some embodiments, any of the components of the apparatus according to the embodiments may be implemented by one or more processors capable of executing one or more programs, and the one or more programs include instructions for causing or causing any one or more of the operations / methods according to the embodiments to be performed. Executable instructions for performing the methods / operations of the apparatus according to the embodiments may be stored in a non-transitory CRM or other computer program product configured to be executed by one or more processors, or may be stored in a transient CRM or other computer program product configured to be executed by one or more processors. In addition, the term "memory" in the embodiments is used as a concept that encompasses not only volatile memory (e.g., RAM, etc.), but also non-volatile memory, flash memory, PROM, etc. It may also be implemented in the form of a carrier wave, such as transmission over the Internet. In addition, a processor-readable recording medium may be distributed among network-connected computer systems, and the processor-readable code may be stored and executed in a distributed manner.
[0425] In this specification, " / " and "," are interpreted as "and / or." For example, "A / B" is interpreted as "A and / or B," and "A, B" is interpreted as "A and / or B." Furthermore, "A / B / C" means "any of A, B, and / or C." Also, "A, B, C" means "any of A, B, and / or C." Furthermore, in this document, "or" is interpreted as "and / or." For example, "A or B" means 1) only "A," 2) only "B," or 3) "A and B." In other words, in this specification, "or" means "additionally or alternatively."
[0426] Terms such as "first," "second," and the like are used to describe various components of the embodiments. However, the interpretation of the various components according to the embodiments should not be limited by these terms. These terms are merely used to distinguish one component from another. For example, a first user input signal can be referred to as a second user input signal. Similarly, a second user input signal can be referred to as a first user input signal. The use of such terms does not depart from the scope of the various embodiments. Although the first user input signal and the second user input signal are both user input signals, they do not refer to the same user input signal unless the context clearly indicates otherwise.
[0427] Terms used to describe the embodiments are used to describe particular embodiments and are not intended to limit the embodiments. As used in the description of the embodiments and in the claims, the singular includes the plural unless the context clearly dictates otherwise. The term "and / or" is used to include all possible combinations between terms. "Comprises" describes the presence of features, numbers, steps, elements, and / or components, but does not imply the absence of additional features, numbers, steps, elements, and / or components. Conditional expressions such as "if" and "when" used to describe the embodiments are only intended to be selective and not limiting. It is intended that when a specific condition is met, a related action is performed in response to a specific condition, or a related definition is interpreted.
[0428] The best mode for carrying out the invention will be described in detail below. [Industrial Applicability]
[0429] It will be apparent to those skilled in the art that various modifications and variations of the present invention are possible within the spirit and scope of the present invention. Thus, the present invention covers the modifications and variations of the present invention provided within the scope of the appended claims and their equivalents.
Claims
1. 1. A method for encoding point cloud data, comprising: encoding geometry information in the point cloud data, the geometry information indicating positions of points in the point cloud data; encoding quality information in the point cloud data, the quality information indicating a quality of the points of the point cloud data; the step of encoding the characteristic information includes generating at least one Level of Detail (LOD); generating at least one LOD includes grouping points based on a number of points in a sampling range in an octree, and selecting a point based on the grouping; the grouping operation includes grouping the sampling range with at least one other sampling range based on the number of points in the sampling range; The method, wherein the selection operation samples one point within the grouped sampling ranges instead of the sampling range.
2. the grouping operation includes comparing the number of points in the sampling range with a specified minimum number; If the number of points in the sampling range is less than the specified minimum number, the sampling range is grouped with the at least one other sampling range; The method of claim 1 , wherein the sampling range corresponds to at least one node of the octree.
3. 1. A method for decoding point cloud data, comprising: decoding geometry information in the point cloud data, the geometry information indicating positions of points in the point cloud data; decoding characteristic information in the point cloud data, the characteristic information indicating characteristics of the points of the point cloud data; the step of decoding the characteristic information includes generating at least one Level of Detail (LOD); generating at least one LOD includes grouping points based on a number of points in a sampling range in an octree, and selecting a point based on the grouping; the grouping operation includes grouping the sampling range with at least one other sampling range based on the number of points in the sampling range; The method, wherein the selection operation samples one point within the grouped sampling ranges instead of the sampling range.
4. the grouping operation includes comparing the number of points in the sampling range with a specified minimum number; If the number of points in the sampling range is less than the specified minimum number, the sampling range is grouped with the at least one other sampling range; The method of claim 3 , wherein the sampling range corresponds to at least one node of the octree.
5. 1. A method for transmitting a bitstream, comprising: encoding point cloud data including geometry information and feature information, the geometry information indicates positions of points in the point cloud data; the characteristic information indicates characteristics of the points of the point cloud data; encoding the point cloud data includes encoding the geometry information and encoding the characteristic information; encoding the characteristic information includes generating at least one Level of Detail (LOD); generating at least one LOD includes a grouping operation based on a number of points in a sampling range, and selecting one point based on the grouping operation; the grouping operation includes grouping the sampling range with at least one other sampling range based on the number of points in the sampling range; The selection operation samples one point within the grouped sampling range; transmitting the bitstream including the encoded point cloud data.
6. the grouping operation includes comparing the number of points in the sampling range with a specified minimum number; If the number of points in the sampling range is less than the specified minimum number, the sampling range is grouped with the at least one other sampling range; The method of claim 5 , wherein the sampling range corresponds to at least one node of an octree.
7. the bitstream includes information indicating a type of sampling method for generating the at least one LOD; the information includes a first type of sampling method representing the sampling range based on an octree; The method of claim 5 , wherein if the information includes the first type of sampling method, one point is sampled from the sampling range to generate the at least one LOD.
Citation Information
Patent Citations
JPP7451576B