Point cloud data transmission device, point cloud data transmission method, point cloud data reception device, and point cloud data reception method
The method and apparatus address latency and complexity issues in processing point cloud data by employing G-PCC and V-PCC coding, enabling efficient and high-quality point cloud services for VR, AR, and autonomous driving.
Patent Information
- Application Number
- JP2025076035
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-04-21
- Filing Date
- 2025-05-01
- Publication Date
- 2025-07-17
- Estimated Expiration
- 2042-01-26
AI Technical Summary
Existing technologies face challenges in efficiently processing large volumes of point cloud data required for VR, AR, MR, and autonomous driving services, leading to latency and encoding/decoding complexity issues.
A method and apparatus for encoding and decoding point cloud data using point cloud compression techniques, including G-PCC and V-PCC coding, to facilitate efficient transmission and rendering of point cloud content.
The solution enables high-quality point cloud services with reduced latency and improved encoding/decoding efficiency for applications such as VR, AR, and autonomous driving.
Smart Images

Figure 2025107332000001_ABST
Abstract
Description
Technical Field
[0001] The embodiments relate to a method and an apparatus for processing point cloud content.
Background Art
[0002] Point cloud content is content represented by a point cloud, which is a set of points (points) belonging to a coordinate system representing a three-dimensional space. Point cloud content can represent three-dimensional media and is used to provide various services such as VR (Virtual Reality), AR (Augmented Reality), MR (Mixed Reality), and autonomous driving services. However, in order to represent point cloud content, tens of thousands to hundreds of thousands of point data are required. Therefore, a method for efficiently processing a huge amount of point data is required.
Summary of the Invention
Problems to be Solved by the Invention
[0003] The embodiments provide an apparatus and a method for efficiently processing point cloud data. The embodiments provide a point cloud data processing method and an apparatus for solving latency and encoding / decoding complexity.
[0004] However, the scope of the rights of the embodiments is not limited to only the above-described technical problems, and can be extended to other technical problems derived by those skilled in the art based on all the described contents.
Means for Solving the Problems
[0005] The point cloud data transmission method according to the embodiment includes a step of encoding point cloud data and a step of transmitting a bit stream including the point cloud data. The point cloud data reception method according to the embodiment includes a step of receiving a bit stream including the point cloud data and a step of decoding the point cloud data.
Advantages of the Invention
[0006] The device and method according to the embodiment can process point cloud data with high efficiency.
[0007] The device and method according to the embodiment can provide a high-quality point cloud service.
[0008] The device and method according to the embodiment can provide point cloud content for general services such as VR services and autonomous driving services.
Brief Description of the Drawings
[0009] The accompanying drawings are for helping to understand the embodiment, and show the embodiment together with the description related to the embodiment. For a more appropriate understanding of various embodiments to be described later, the following description of the embodiment must be referred to in connection with the following drawings including parts corresponding to similar reference numerals in the accompanying drawings.
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15a
Figure 15b
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Best Mode for Carrying Out the Invention
[0011] The preferred embodiments will be specifically described with reference to the accompanying drawings. The following detailed description with reference to the accompanying drawings is for explaining the preferred embodiments rather than showing only the embodiments that can be implemented by the examples. The following detailed description includes details to provide a thorough understanding of the embodiments. However, it is obvious to those skilled in the art that the embodiments can be implemented without such details.
[0012] Most of the terms used in the embodiments are common ones widely used in the relevant field, but some are arbitrarily selected by the applicant, and their meanings will be explained in detail below if necessary. Therefore, the embodiments should be understood based on the intended meanings of the terms rather than the simple names and meanings of the terms.
[0013] FIG. 1 is a diagram showing an example of a point cloud content providing system according to an embodiment.
[0014] The point cloud content providing system shown in FIG. 1 includes a transmission device 10000 and a reception device 10004. The transmission device 10000 and the reception device 10004 can perform wired and wireless communication to transmit and receive point cloud data.
[0015] The transmission device 10000 according to the embodiment secures, processes, and transmits point cloud video (or point cloud content). In the embodiment, the transmission device 10000 includes a fixed station, a base transceiver system (BTS), a network, an artificial intelligence (AI) device and / or system, a robot, an AR / VR / XR device and / or server, etc. Also in the embodiment, the transmission device 10000 is a device that communicates with a base station and / or other wireless devices using a wireless connection technology (e.g., 5G New Radio (NR), Long Term Evolution (LTE)), and includes a robot, a vehicle, an AR / VR / XR device, a mobile device, a household appliance, an Internet of Things (IoT) device, an AI device / server, etc.
[0016] The transmission device 10000 according to the embodiment includes a point cloud video acquisition unit 10001, a point cloud video encoder 10002, and / or a transmitter (or communication module) 10003.
[0017] The point cloud video acquisition unit 10001 according to the embodiment acquires a point cloud video through a processing process such as capture, synthesis, or generation. The point cloud video is point cloud content represented by a point cloud that is a set of points located in a three-dimensional space, and is also called point cloud video data, etc. The point cloud video according to the embodiment includes one or more frames. One frame represents a still video / picture. Therefore, the point cloud video includes point cloud videos / frames / pictures, and is called any one of the point cloud video, frame, and picture.
[0018] The point cloud video encoder 10002 according to the embodiment encodes the acquired point cloud video data. The point cloud video encoder 10002 encodes the point cloud video data based on point cloud compression (PCC) coding. The point cloud compression coding according to the embodiment includes G-PCC (Geometry-based Point Cloud Compression) coding and / or V-PCC (Video based Point Cloud Compression) coding or next-generation coding. Note that the point cloud compression coding according to the embodiment is not limited to the above-described embodiment. The point cloud video encoder 10002 outputs a bitstream including the encoded point cloud video data. The bitstream includes not only the encoded point cloud video data but also signaling information related to the encoding of the point cloud video data.
[0019] The transmitter 10003 according to the embodiment transmits a bitstream including encoded point cloud video data. The bitstream according to the embodiment is encapsulated into a file or a segment (e.g., a streaming segment) and transmitted by various networks such as a broadcast network and / or a broadband network. Although not shown, the transmitting device 10000 includes an encapsulation unit (or an encapsulation module) that performs an encapsulation operation. Also, in an embodiment, the encapsulation unit is included in the transmitter 10003. In an embodiment, the file or the segment is transmitted to the receiving device 10004 by a network or stored in a digital storage medium (e.g., USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.). The transmitter 10003 according to the embodiment can communicate wirelessly or by wire with the receiving device 10004 (or the Receiver 10005) via a network such as 4G, 5G, 6G, etc. Also, the transmitter 10003 performs necessary data processing operations by a network system (e.g., a communication network system such as 4G, 5G, 6G, etc.). Also, the transmitting device 10000 can transmit the encapsulated data by an On Demand method.
[0020] The receiving device 10004 according to the embodiment includes a Receiver 10005, a Point Cloud Decoder 10006, and / or a Renderer 10007. In an embodiment, the receiving device 10004 uses a wireless connection technology (e.g., 5G NR (New RAT), LTE (Long Term Evolution)) to communicate with a base station and / or other wireless devices, including devices, robots, vehicles, AR / VR / XR devices, mobile devices, home appliances, IoT (Internet of Thing) devices, AI devices / servers, etc.
[0021] The receiver 10005 according to the embodiment receives from a network or a storage medium a bitstream including point cloud video data or a file / segment in which the bitstream is encapsulated. The receiver 10005 performs necessary data processing operations by means of a network system (for example, a communication network system such as 4G, 5G, 6G, etc.). The receiver 10005 according to the embodiment decapsulates the received file / segment and outputs a bitstream. Also, in the embodiment, the receiver 10005 includes a decapsulation unit (or decapsulation module) for performing the decapsulation operation. Also, the decapsulation unit is implemented as an element (or component) separate from the receiver 10005.
[0022] The point cloud video decoder 10006 decodes a bitstream including point cloud video data. The point cloud video decoder 10006 can decode in a manner in which the point cloud video data is encoded (for example, a process reverse to the operation of the point cloud video encoder 10002). Therefore, the point cloud video decoder 10006 can perform point cloud restoration coding, which is a reverse process of point cloud compression, to decode the point cloud video data. The point cloud restoration coding includes G-PCC coding.
[0023] The renderer 10007 renders the decoded point cloud video data. The renderer 10007 renders not only the point cloud video data but also audio data to output point cloud content. In the embodiment, the renderer 10007 includes a display for displaying the point cloud content. In the embodiment, the display is not included in the renderer 10007 but is implemented as another device or component.
[0024] In the drawings, the arrow indicated by the dotted line indicates the transmission path of the feedback information obtained by the receiving device 10004. The feedback information is information for reflecting the interaction with the user who consumes the point cloud content, and includes user information (for example, head orientation information, viewport information, etc.). In particular, when the point cloud content is for a service that requires interaction with the user (for example, an autonomous driving service), the feedback information can be transmitted to the content transmission side (for example, the transmitting device 10000) and / or the service provider. In an embodiment, the feedback information can be used not only by the transmitting device 10000 but also by the receiving device 10004, and can also be not provided.
[0025] The head orientation information according to the embodiment is information regarding the position, direction, angle, movement, etc. of the user's head. The receiving device 10004 according to the embodiment calculates viewport information based on the head orientation information. The viewport information is information regarding the area of the point cloud video that the user is viewing. The viewpoint is the point where the user is viewing the point cloud video and means the center of the viewport area. That is, the viewport is an area centered on the viewpoint, and the size, shape, etc. of the area are determined by the FOV (Field Of View). Therefore, in addition to the head orientation information, the receiving device 10004 can extract viewport information based on the vertical or horizontal FOV supported by the device. The receiving device 10004 also performs Gaze Analysis, etc. to confirm the user's point cloud consumption method, the point cloud video area that the user gazes at, the gazing time, etc. In the embodiment, the receiving device 10004 transmits feedback information including the result of the gaze analysis to the transmitting device 10000. The feedback information according to the embodiment is obtained in the rendering and / or display process. The feedback information according to the embodiment is ensured by one or more sensors included in the receiving device 10004. Also, in the embodiment, the feedback information is ensured by the renderer 10007 or another external element (or device, component, etc.). The dotted line shown in FIG. 1 shows the transmission process of the feedback information ensured by the renderer 10007. The point cloud content providing system processes (encodes / decodes) the point cloud data based on the feedback information. Therefore, the point cloud video data decoder 10006 can perform the decoding operation based on the feedback information. Also, the receiving device 10004 can transmit the feedback information to the transmitting device 10000. The transmitting device 10000 (or the point cloud video data encoder 10002) performs the encoding operation based on the feedback information.Therefore, the point cloud content providing system can efficiently process the necessary data (for example, the point cloud data corresponding to the user's head position) based on the feedback information without processing (encoding / decoding) all the point cloud data, and provide the point cloud content to the user.
[0026] In the embodiment, the transmitting device 10000 is called an encoder, a transmitting device, a transmitter, etc., and the receiving device 10004 is called a decoder, a receiving device, a receiver, etc.
[0027] The point cloud data processed (processed in a series of processes of acquisition / encoding / transmission / decoding / rendering) by the point cloud content providing system of FIG. 1 according to the embodiment is also called point cloud content data or point cloud video data. In the embodiment, the point cloud content data is used as a concept including metadata or signaling information related to the point cloud data.
[0028] The elements of the point cloud content providing system shown in FIG. 1 are implemented by hardware, software, a processor, and / or a combination thereof, etc.
[0029] FIG. 2 is a block diagram showing the operation of point cloud content providing according to the embodiment.
[0030] FIG. 2 is a block diagram showing the operation of the point cloud content providing system described in FIG. 1. As described above, the point cloud content providing system processes the point cloud data based on point cloud compression coding (for example, G-PCC).
[0031] In the point cloud content providing system according to the embodiment (for example, the point cloud transmission device 10000 or the point cloud video acquisition unit 10001), a point cloud video is acquired (20000). The point cloud video is represented by a point cloud belonging to a coordinate system representing a three-dimensional space. The point cloud video according to the embodiment includes a Ply (Polygon File format or the Stanford Triangle format) file. When the point cloud video has one or more frames, the acquired point cloud video includes one or more Ply files. The Ply file includes point cloud data such as point geometry and / or attributes. The geometry includes the position of the point. The position of each point is represented by parameters (for example, values of the X-axis, Y-axis, and Z-axis respectively) indicating a three-dimensional coordinate system (for example, a coordinate system composed of the XYZ axes). The attribute includes point attributes (for example, texture information of each point, hue (YCbCr or RGB), reflectivity (r), transparency, etc.). One point has one or more attributes (or characteristics). For example, one point can have one attribute of hue or two attributes of hue and reflectivity. In the embodiment, the geometry is also called position, geometry information, geometry data, etc., and the attribute is also called attribute, attribute information, attribute data, etc. Further, the point cloud content providing system (for example, the point cloud transmission device 10000 or the point cloud video acquisition unit 10001) can secure point cloud data from information related to the acquisition process of the point cloud video (for example, depth information, hue information, etc.).
[0032] The point cloud content providing system according to an embodiment (e.g., the transmission device 10000 or the point cloud video encoder 10002) encodes point cloud data (20001). The point cloud content providing system encodes point cloud data based on point cloud compression coding. As described above, point cloud data includes point geometry and characteristics. Therefore, the point cloud content providing system performs geometry encoding for encoding geometry to output a geometry bitstream. The point cloud content providing system performs characteristic encoding for encoding characteristics to output a characteristic bitstream. In an embodiment, the point cloud content providing system performs characteristic encoding based on geometry encoding. The geometry bitstream and the characteristic bitstream according to the embodiment are multiplexed and output as one bitstream. The bitstream according to the embodiment further includes signaling information related to geometry encoding and characteristic encoding.
[0033] The point cloud content providing system according to an embodiment (e.g., the transmission device 10000 or the transmitter 10003) transmits the encoded point cloud data (20002). As described with reference to FIG. 1, the encoded point cloud data is represented by a geometry bitstream and a characteristic bitstream. Further, the encoded point cloud data is transmitted in the form of a bitstream together with signaling information related to the encoding of the point cloud data (e.g., signaling information related to geometry encoding and characteristic encoding). Also, the point cloud content providing system encapsulates the bitstream for transmitting the encoded point cloud data and transmits it in the form of a file or a segment.
[0034] A point cloud content providing system according to an embodiment (for example, the receiving device 10004 or the receiver 10005) receives a bitstream including encoded point cloud data. Also, the point cloud content providing system (for example, the receiving device 10004 or the receiver 10005) demultiplexes the bitstream.
[0035] A point cloud content providing system (for example, the receiving device 10004 or the point cloud video decoder 10005) decodes the encoded point cloud data (for example, the geometry bitstream, the attribute bitstream) transmitted in the bitstream. The point cloud content providing system (for example, the receiving device 10004 or the point cloud video decoder 10005) decodes the point cloud video data based on the signaling information related to the encoding of the point cloud video data included in the bitstream. The point cloud content providing system (for example, the receiving device 10004 or the point cloud video decoder 10005) decodes the geometry bitstream to restore the positions (geometry) of the points. The point cloud content providing system decodes the attribute bitstream based on the restored geometry to restore the attributes of the points. The point cloud content providing system (for example, the receiving device 10004 or the point cloud video decoder 10005) restores the point cloud video based on the positions by the restored geometry and the decoded attributes.
[0036] The point cloud content providing system according to the embodiment (for example, the receiving device 10004 or the renderer 10007) renders the decoded point cloud data (20004). The point cloud content providing system (for example, the receiving device 10004 or the renderer 10007) renders the geometry and attributes decoded in the decoding process by various rendering methods. The points of the point cloud content are rendered as fixed points having a certain thickness, cubes having a predetermined minimum size centered at the position of the corresponding fixed point, or circles centered at the position of the fixed point. All or part of the area of the rendered point cloud content is provided to the user by a display (for example, a VR / AR display, a general display, etc.).
[0037] The point cloud content providing system according to the embodiment (for example, the receiving device 10004) can secure feedback information (20005). The point cloud content providing system encodes and / or decodes the point cloud data based on the feedback information. Since the operation of the feedback information and the point cloud content providing system according to the embodiment is the same as the feedback information and the operation described in FIG. 1, a detailed description is omitted.
[0038] FIG. 3 is a diagram showing an example of the point cloud video capture process according to the embodiment.
[0039] FIG. 3 shows an example of the point cloud video capture process of the point cloud content providing system described in FIGS. 1 and 2.
[0040] Point cloud content includes point cloud videos (images and / or videos) that represent objects and / or environments located in various three-dimensional spaces (e.g., a three-dimensional space representing a real environment, a three-dimensional space representing a virtual environment, etc.). Therefore, the point cloud content providing system according to the embodiments captures point cloud videos using one or more cameras (e.g., an infrared camera capable of obtaining depth information, an RGB camera capable of extracting hue information corresponding to the depth information, etc.), a projector (e.g., an infrared pattern projector for obtaining depth information), LiDAR, etc. to generate point cloud content. The point cloud content providing system according to the embodiments extracts a form of geometry composed of points on a three-dimensional space from the depth information, and extracts the characteristics of each point from the hue information to obtain point cloud data. The images and / or videos according to the embodiments are captured based on either an inward-facing method or an outward-facing method.
[0041] The left side of FIG. 3 shows the inward-facing method. The inward-facing method is a method in which one or more cameras (or camera sensors) located surrounding the central object capture the central object. The inward-facing method is used to generate point cloud content that provides a 360° image of the core object to the user (e.g., VR / AR content that provides a 360° image of an object (e.g., a core object such as a character, athlete, item, actor, etc.) to the user).
[0042] The right side of FIG. 3 shows the outward-facing method. The outward-facing method is a method in which one or more cameras (or camera sensors) located surrounding the central object capture the environment of the central object rather than the central object itself. The outward-facing method is used to generate point cloud content for providing the surrounding environment from the user's perspective (e.g., content showing the external environment provided to the user of an autonomous vehicle).
[0043] As shown, the point cloud content is generated based on the capture operations of one or more cameras. In this case, since the coordinate systems of the respective cameras are different, the point cloud content providing system performs calibration of one or more cameras in order to set a global coordinate system before the capture operation. Also, the point cloud content providing system generates point cloud content by synthesizing the images and / or videos captured by the above-described capture method and any images and / or videos. Further, when the point cloud content providing system generates point cloud content indicating a virtual space, it does not perform the capture operation described in FIG. 3. The point cloud content providing system according to the embodiment can also perform post-processing on the captured images and / or videos. That is, the point cloud content providing system can remove an unwanted area (for example, the background), or recognize the space where the captured images and / or videos are connected and perform an operation to fill it if there is a spatial hole.
[0044] Also, the point cloud content providing system can perform coordinate system conversion on the points of the point cloud videos obtained from the respective cameras to generate one point cloud content. The point cloud content providing system performs coordinate system conversion of the points based on the position coordinates of the respective cameras. Thereby, the point cloud content providing system can generate content indicating one wide range or generate point cloud content with a high point density.
[0045] FIG. 4 is a diagram showing an example of a point cloud encoder according to an embodiment.
[0046] FIG. 4 shows an example of the point cloud video encoder 10002 of FIG. 1. The point cloud encoder reconstructs point cloud data (e.g., the position and / or characteristics of points) to perform an encoding operation in order to adjust the quality of point cloud content (e.g., lossless, lossy, near-lossless) according to network conditions or applications. When the overall size of the point cloud content is large (e.g., for point cloud content that is 60 Gbps at 30 fps), the point cloud content providing system cannot stream the corresponding content in real time. Therefore, the point cloud content providing system can reconstruct the point cloud content based on the maximum target bitrate for providing according to the network environment, etc.
[0047] As shown in FIGS. 1 and 2, the point cloud encoder can perform geometric encoding and characteristic encoding. Geometric encoding is performed prior to characteristic encoding.
[0048] *
[0049] The point cloud encoder according to the embodiment includes a coordinate transformation unit (Transformation Coordinates) 40000, a quantization unit (Quantize and Remove Points (Voxelize)) 40001, an octree analysis unit (Analyze Octree) 40002, a surface approximation analysis unit (Analyze Surface Approximation) 40003, an arithmetic encoder (Arithmetic Encode) 40004, a geometry reconstruction unit (Reconstruct Geometry) 40005, a color conversion unit (Transform Colors) 40006, an attribute conversion unit (Transfer Attributes) 40007, a RAHT conversion unit 40008, a LOD generation unit (Generated LOD) 40009, a lifting conversion unit (Lifting) 40010, a coefficient quantization unit (Quantize Coefficients) 40011 and / or an arithmetic encoder (Arithmetic Encode) 40012.
[0050] The coordinate transformation unit 40000, the quantization unit 40001, the octree analysis unit 40002, the surface approximation analysis unit 40003, the arithmetic encoder 40004 and the geometry reconstruction unit 40005 perform geometry encoding. The geometry encoding according to the embodiment includes octree geometry coding, direct coding, trisoup geometry encoding and entropy encoding. Direct coding and trisoup geometry encoding are selectively or combinatorially applied. Note that the geometry encoding is not limited to the above examples.
[0051] As shown in the figure, the coordinate transformation unit 40000 according to the embodiment receives a position and converts it into a coordinate system. For example, the position is converted into position information in a three-dimensional space (such as a three-dimensional space represented by an XYZ coordinate system). The position information of the three-dimensional space according to the embodiment is also referred to as geometry information.
[0052] The quantization unit 40001 according to the embodiment quantizes the geometry. For example, the quantization unit 40001 quantizes points based on the minimum position values of the overall points (for example, the minimum values on each axis with respect to the X-axis, Y-axis, and Z-axis). The quantization unit 40001 performs a quantization operation of multiplying the difference between the minimum position value and the position value of each point by a predetermined quantization scale value, and then rounding down or up to find the nearest integer value. Therefore, one or more points can have the same quantized position (or position value). The quantization unit 40001 according to the embodiment performs voxelization based on the quantized positions in order to reconstruct the quantized points. The minimum unit including 2D image / video information is like a pixel, and the points of the point cloud content (or 3D point cloud video) according to the embodiment are included in one or more voxels. A voxel is a word combining volume and pixel, and it means a 3D cubic space generated when dividing the 3D space into units of unit = 1.0 based on the axes representing the 3D space (for example, the X-axis, Y-axis, and Z-axis). The quantization unit 40001 matches groups of points in the 3D space with voxels. In an embodiment, one voxel includes only one point. In an embodiment, one voxel includes one or more points. Also, in order to represent one voxel with one point, the position of the center point of the corresponding voxel can be set based on the position of one or more points included in one voxel. In this case, the characteristics of all the positions included in one voxel are combined and assigned to the corresponding voxel.
[0053] The octree analysis unit 40002 according to the embodiment performs octree geometric coding (or octree coding) for representing voxels in an octree structure. The octree structure represents points matched to voxels based on an eight-division structure.
[0054] The surface approximation analysis unit 40003 according to the embodiment analyzes and approximates the octree. The octree analysis and approximation according to the embodiment is a process of performing analysis to voxelize a region containing a large number of points in order to efficiently provide an octree and voxelization.
[0055] The arithmetic encoder 40004 according to the embodiment entropy-encodes the octree and / or the approximated octree. For example, the encoding method includes the arithmetic encoding method. As a result of the encoding, a geometry bitstream is generated.
[0056] The color conversion unit 40006, the texture conversion unit 40007, the RAHT conversion unit 40008, the LOD generation unit 40009, the lift conversion unit 40010, the coefficient quantization unit 40011 and / or the arithmetic encoder 40012 perform texture encoding. As described above, one point has one or more textures. The texture encoding according to the embodiment is equally applied to the textures possessed by one point. However, when one texture (for example, hue) includes one or more elements, independent texture encoding is applied to each element. The texture encoding according to the embodiment includes color conversion coding, texture conversion coding, RAHT (Region Adaptive Hierarchial Transform) coding, prediction transform (Interpolaration-based hierarchical nearest-neighbour prediction - Prediction Transform) coding, and lift transform (interpolation-based hierarchical nearest-neighbour prediction with an update / lifting step (Lifting Transform)) coding. Depending on the point cloud content, the above-described RAHT coding, prediction transform coding, and lift transform coding are selectively used, or a combination of one or more codings is used. Also, the texture encoding according to the embodiment is not limited to the above examples.
[0057] The color conversion unit 40006 according to the embodiment performs color conversion coding for converting the color values (or textures) included in the features. For example, the color conversion unit 40006 converts the format of the hue information (for example, converts from RGB to YCbCr). The operation of the color conversion unit 40006 according to the embodiment is optionally applied according to the color values included in the features.
[0058] The geometry reconstruction unit 40005 according to the embodiment reconstructs (restores) the octree and / or the approximated octree. The geometry reconstruction unit 40005 reconstructs the octree / voxel based on the result of analyzing the point distribution. The reconstructed octree / voxel is also called the reconstructed geometry (or restored geometry).
[0059] The feature conversion unit 40007 according to the embodiment performs feature conversion for converting features based on the positions where geometry coding has not been performed and / or the reconstructed geometry. As described above, since the features are subordinate to the geometry, the feature conversion unit 40007 can convert the features based on the reconstructed geometry information. For example, the feature conversion unit 40007 converts the features of the points at that position based on the position values of the points included in the voxel. As described above, when the position of the central point of the corresponding voxel is set based on the position of one or more points included in one voxel, the feature conversion unit 40007 converts the features of one or more points. When trisoup geometry coding is performed, the feature conversion unit 40007 converts the features based on the trisoup geometry coding.
[0060] The feature conversion unit 40007 calculates the average value of the features or feature values (for example, the hue or reflectivity of each point, etc.) of the points adjacent within a specific position / radius from the position (or position value) of the central point of each voxel to perform feature conversion. When calculating the average value, the feature conversion unit 40007 applies a weighting value according to the distance from the central point to each point. Therefore, each voxel has a position and a calculated feature (or feature value).
[0061] The feature transformation unit 40007 searches for adjacent points existing within a specific position / radius from the position of the center point of each voxel based on a K-D tree or a Morton code. The K-D tree is a binary search tree that supports a data structure for managing points based on their positions so as to enable rapid Nearest Neighbor Search (NNS). The Morton code represents the coordinate values (e.g., (x, y, z)) indicating the three-dimensional positions of all points as bit values and is generated by mixing the bits. For example, if the coordinate value indicating the position of a point is (5, 9, 1), the bit values of the coordinate values are (0101, 1001, 0001). Mixing the bit values in the order of z, y, x according to the bit indices results in 010001000111. When this value is represented in decimal, it becomes 1095. That is, the Morton code value of the point with the coordinate value (5, 9, 1) is 1095. The feature transformation unit 40007 aligns the points based on the Morton code values and performs Nearest Neighbor Search (NNS) through a depth-first traversal process. After the feature transformation operation, if Nearest Neighbor Search (NNS) is also required in other transformation processes for feature coding, the K-D tree or the Morton code is utilized.
[0062] As shown in the illustration, the transformed features are input to the RAHT transformation unit 40008 and / or the LOD generation unit 40009.
[0063] The RAHT transformation unit 40008 according to the embodiment performs RAHT coding for predicting feature information based on the reconstructed geometry information. For example, the RAHT transformation unit 40008 can predict the feature information of a node at a higher level of the octree based on the feature information associated with a node at a lower level of the octree.
[0064] The LOD generation unit 40009 according to the embodiment generates an LOD (Level of Detail) for performing predictive transform coding. The LOD according to the embodiment indicates the degree of detail of the point cloud content. The smaller the LOD value, the lower the detail of the point cloud content, and the larger the LOD value, the higher the detail of the point cloud content. Points can be classified by LOD.
[0065] The lift conversion unit 40010 according to the embodiment performs lift conversion coding that converts the characteristics of the point cloud based on the weighting value. As described above, lift conversion coding is selectively applied.
[0066] The coefficient quantization unit 40011 according to the embodiment quantizes the trait-coded traits based on the coefficients.
[0067] The arithmetic encoder 40012 according to the embodiment encodes the quantized traits based on arithmetic coding.
[0068] The elements of the point cloud encoder in FIG. 4 are implemented by hardware, software, firmware, or a combination thereof, including one or more processors or integrated circuits communicably configured with one or more memories included in the point cloud providing device, although not shown in the figure. One or more processors can perform any one of the operations and / or functions of the elements of the point cloud encoder in FIG. 4 described above. Also, one or more processors can operate or execute a software program and / or a set of instructions for performing the operations and / or functions of the elements of the point cloud encoder in FIG. 4. One or more memories according to the embodiment include high-speed random access memory or non-volatile memory (e.g., one or more magnetic disk storage devices, flash memory devices, or other non-volatile solid-state memory devices, etc.).
[0069] FIG. 5 is a diagram showing an example of a voxel according to an embodiment.
[0070] FIG. 5 shows voxels located in a three-dimensional space represented by a coordinate system composed of three axes: the X-axis, Y-axis, and Z-axis. As shown in FIG. 4, a point cloud encoder (e.g., quantization unit 40001, etc.) performs voxelization. A voxel means a three-dimensional cubic space generated when dividing the three-dimensional space into units (unit = 1.0) based on axes representing the three-dimensional space (e.g., the X-axis, Y-axis, and Z-axis). FIG. 5 shows two extreme points (0, 0, 0) and (2 d , 2 d , 2 d ) and shows an example of a voxel generated by an octree structure that recursively divides the cubical axis-aligned bounding box defined by them. One voxel contains at least one or more points. A voxel can estimate spatial coordinates from its positional relationship with a voxel group. As described above, a voxel has characteristics (such as hue or reflectance) similar to pixels of a two-dimensional image / video. Specific explanations for voxels are as described in FIG. 4 and are thus omitted.
[0071] FIG. 6 is a diagram showing an example of an octree and an occupancy code according to an embodiment.
[0072] As shown in FIGS. 1 to 4, a point cloud content providing system (point cloud video encoder 10002) or a point cloud encoder (e.g., octree analysis unit 40002) performs octree geometric coding (or octree coding) based on an octree structure in order to efficiently manage the regions and / or positions of voxels.
[0073] The upper part of FIG. 6 shows an octree structure. The three-dimensional space of the point cloud content according to the embodiment is represented by the axes of the coordinate system (for example, the X-axis, Y-axis, and Z-axis). The octree structure is generated by recursively subdividing a boundary box (cubical axis-aligned bounding box) defined by two extreme points (0, 0, 0) and (2 d , 2 d , 2 d ). 2d is set to the value that constitutes the smallest boundary box surrounding all the points of the point cloud content (or point cloud video). d indicates the depth of the octree. The d value is determined by the following formula. In the following formula, (x int n , y int n , z int n ) indicates the position (or position value) of the quantized point.
[0074]
Number
[0075] As shown in the upper center of FIG. 6, by the division, the entire three-dimensional space is divided into eight spaces. Each of the divided spaces is represented by a cube having six faces. As shown in the upper right of FIG. 6, the eight spaces are each further divided by the axes of the coordinate system (for example, the X-axis, Y-axis, and Z-axis). Thus, each space is again divided into eight small spaces. The divided small spaces are also represented by cubes having six faces. Such a division method is applied until the leaf nodes of the octree become voxels.
[0076] The lower side of FIG. 6 shows the occupancy code of the octree. The occupancy code of the octree is generated to indicate whether each of the eight divided spaces generated by dividing one space contains at least one point. Therefore, one occupancy code is represented by eight child nodes. Each child node indicates the occupancy of the divided space, and the child node has a value of 1 bit. Therefore, the occupancy code is represented by an 8-bit code. That is, if at least one point is included in the space corresponding to the child node, the corresponding node has a value of 1. If no point is included in the space corresponding to the node (empty), the corresponding node has a value of 0. Since the occupancy code shown in FIG. 6 is 00100001, among the eight child nodes, it indicates that the spaces corresponding to the third child node and the eighth child node each contain at least one point. As shown in the figure, the third child node and the eighth child node each have eight child nodes, and each child node is represented by an 8-bit occupancy code. In the drawing, it shows that the occupancy code of the third child node is 10000111 and the occupancy code of the eighth child node is 01001111. The point cloud encoder according to an embodiment (for example, the arithmetic encoder 40004) entropy-encodes the occupancy code. Also, in order to improve the compression efficiency, the point cloud encoder intra / inter-codes the occupancy code. The receiving device according to an embodiment (for example, the receiving device 10004 or the point cloud video decoder 10006) reconstructs the octree based on the occupancy code.
[0077] The point cloud encoder according to an embodiment (for example, the point cloud encoder in FIG. 4, or the octree analysis unit 40002) performs voxelization and octree coding to store the positions of points. However, since points in a three-dimensional space are not always evenly distributed, there may be specific regions where there are few points. Therefore, it is inefficient to perform voxelization on the entire three-dimensional space. For example, if there are almost no points in a specific region, there is no need to perform voxelization up to that region.
[0078] Therefore, the point cloud encoder according to the embodiment does not perform voxelization on the above-described specific regions (or nodes excluding the leaf nodes of the octree), but performs direct coding that directly codes the positions of the points included in the specific regions. The coordinates of the direct coding points according to the embodiment are called the Direct Coding Mode (DCM). Further, the point cloud encoder according to the embodiment can perform trisoup geometry encoding that reconstructs the positions of the points in a specific region (or node) based on a voxel based on a surface model. Trisoup geometry encoding is geometry encoding that represents the expression of an object in a series of triangle meshes. Therefore, the point cloud decoder can generate a point cloud from the mesh surface. The direct coding and trisoup geometry encoding according to the embodiment are selectively performed. Further, the direct coding and trisoup geometry encoding according to the embodiment can be combined with octree geometry coding (or octree coding).
[0079] In order to perform direct coding, it is necessary to activate the use option of the direct mode for applying direct coding, and the node to which direct coding is applied is not a leaf node, and there must be points below a threshold within a specific node. Also, the total number of points targeted for direct coding must not exceed a predetermined threshold. When the above conditions are met, the point cloud encoder (or the arithmetic encoder 40004) according to the embodiment can entropy code the position (or position value) of the point.
[0080] The point cloud encoder according to the embodiment (e.g., the surface approximation analysis unit 40003) can perform trisoup geometry encoding that determines a specific level of the octree (when the level is smaller than the depth d of the octree) and reconstructs the positions of the points within the node region based on voxels using the surface model from that level (trisoup mode). The point cloud encoder according to the embodiment can specify the level to which trisoup geometry encoding is applied. For example, if the specified level is the same as the depth of the octree, the point cloud encoder does not operate in trisoup mode. That is, the point cloud encoder according to the embodiment can operate in trisoup mode only when the specified level is smaller than the depth value of the octree. The three-dimensional cubic region of the node at the specified level according to the embodiment is called a block. One block contains one or more voxels. A block or voxel can also correspond to a brick. In each block, the geometry is represented as a surface. The surface according to the embodiment can intersect each edge of the block at most once.
[0081] Since one block has 12 edges, there are at least 12 intersection points within one block. Each intersection point is called a vertex. A vertex existing along an edge is detected when there is at least one occupied voxel adjacent to that edge among all the blocks sharing the corresponding edge. The occupied voxel according to the embodiment means a voxel containing a point. The position of the vertex detected along the edge is the average position along the edge of all voxels adjacent to the corresponding edge among all the blocks sharing the corresponding edge (the average position along the edge of all voxels).
[0082] When a vertex is detected, the point cloud encoder according to the embodiment entropy-codes the start point (x, y, z) of the edge, the direction vector (Δx, Δy, Δz) of the edge, and the vertex position value (relative position value within the edge). When trisoup geometry encoding is applied, the point cloud encoder according to the embodiment (for example, the geometry reconstruction unit 40005) performs triangle reconstruction, up-sampling, and voxelization processes to generate the restored geometry (reconstructed geometry).
[0083] Vertices located at the edges of the block determine the surface passing through the block. The surface according to the embodiment is a non-planar polygon. In the process of triangle reconstruction, a surface represented by a triangle is reconstructed based on the start point of the edge, the direction vector of the edge, and the position value of the vertex. The process of triangle reconstruction is as follows. (1) Calculate the centroid value of each vertex, (2) subtract the centroid value from the value of each vertex, and (3) square the result and sum up all the values to obtain the combined value.
[0084]
Number
[0085] Find the minimum value of the added values and perform a projection process along the axis where the minimum value exists. For example, if the x element is the minimum, project each vertex onto the x-axis with the center of the block as the reference and project it onto the (y, z) plane. If the value obtained by projecting onto the (y, z) plane is (ai, bi), find the θ value by atan2(bi, ai) and align the vertices based on the θ value. The following table shows the combinations of vertices for generating triangles according to the number of vertices. The vertices are aligned in order from 1 to n. The following table shows that for four vertices, two triangles are formed by the combinations of vertices. The first triangle is composed of the 1st, 2nd, and 3rd vertices among the aligned vertices, and the second triangle is composed of the 3rd, 4th, and 1st vertices among the aligned vertices.
[0086] Table 2-1 Triangles formed from vertices ordered 1,…,n
[0087] n triangles
[0088] 3(1,2,3)
[0089] 4(1,2,3), (3,4,1)
[0090] 5(1,2,3), (3,4,5), (5,1,3)
[0091] 6(1,2,3), (3,4,5), (5,6,1), (1,3,5)
[0092] 7(1,2,3), (3,4,5), (5,6,7), (7,1,3), (3,5,7)
[0093] 8(1,2,3), (3,4,5), (5,6,7), (7,8,1), (1,3,5), (5,7,1)
[0094] 9(1, 2, 3), (3, 4, 5), (5, 6, 7), (7, 8, 9), (9, 1, 3), (3, 5, 7), (7, 9, 3)
[0095] 10(1, 2, 3), (3, 4, 5), (5, 6, 7), (7, 8, 9), (9, 10, 1), (1, 3, 5), (5, 7, 9), (9, 1, 5)
[0096] 11(1, 2, 3), (3, 4, 5), (5, 6, 7), (7, 8, 9), (9, 10, 11), (11, 1, 3), (3, 5, 7), (7, 9, 11), (11, 3, 7)
[0097] 12(1, 2, 3), (3, 4, 5), (5, 6, 7), (7, 8, 9), (9, 10, 11), (11, 12, 1), (1, 3, 5), (5, 7, 9), (9, 11, 1), (1, 5, 9)
[0098] The upsampling process is performed to voxelize by adding points in the middle along the edges of the triangle. The additional points are generated based on the upsampling factor and the width of the block. The additional points are called refined vertices. The point cloud encoder according to the embodiment can voxelize the refined vertices. Also, the point cloud encoder can perform feature encoding based on the voxelized positions (or position values).
[0099] FIG. 7 is a diagram showing an example of an adjacent node pattern according to an embodiment.
[0100] To increase the compression efficiency of the point cloud video, the point cloud encoder according to the embodiment performs entropy coding based on context adaptive arithmetic coding.
[0101] As described with reference to FIGS. 1 to 6, a point cloud content providing system or a point cloud encoder (e.g., the point cloud video encoder 10002, the point cloud encoder or the arithmetic encoder 40004 in FIG. 4) immediately entropy-codes the occupancy code. Further, the point cloud content providing system or the point cloud encoder performs entropy coding (intra coding) based on the occupancy code of the current node and the occupancy rate of the adjacent nodes, or performs entropy coding (inter coding) based on the occupancy code of the previous frame. A frame according to an embodiment means a set of point cloud videos generated at the same time. The compression efficiency of the intra coding / inter coding according to the embodiment varies depending on the number of adjacent nodes to be referred to. Although it becomes more complex as the number of bits increases, the compression efficiency can be increased by tilting to one side. For example, with a 3-bit context, coding is performed in 8 ways, which is 2 to the power of 3. The portion to be coded separately affects the complexity of the implementation. Therefore, it is necessary to balance the appropriate levels of compression efficiency and complexity.
[0102] FIG. 7 shows the process of obtaining an occupancy pattern based on the occupancy rate of adjacent nodes. The point cloud encoder according to the embodiment determines the occupancy rate of the adjacent nodes of each node of the octree to obtain a neighbor pattern value. The neighbor pattern is used to infer the occupancy pattern of the corresponding node. The left side of FIG. 7 shows the cube corresponding to the node (the cube located in the middle) and the six cubes (adjacent nodes) that share at least one face with the corresponding cube. The illustrated nodes are nodes at the same depth. The illustrated numbers indicate the weighting values (1, 2, 4, 8, 16, 32, etc.) associated with the six nodes respectively. Each weighting value is sequentially assigned according to the position of the adjacent node.
[0103] The right side of FIG. 7 shows the adjacent node pattern value. The adjacent node pattern value is the sum of the values obtained by multiplying the weighted values of the occupied adjacent nodes (adjacent nodes having points). Therefore, the adjacent node pattern value has a value from 0 to 63. The fact that the adjacent node pattern value is 0 means that among the adjacent nodes of the corresponding node, there is no node having a point (occupied node). The fact that the adjacent node pattern value is 63 means that all the adjacent nodes are occupied nodes. As shown in the figure, since the adjacent nodes to which the weighted values 1, 2, 4, and 8 are assigned are occupied nodes, the adjacent node pattern value is 15, which is the combined value of 1, 2, 4, and 8. The point cloud encoder can perform coding based on the adjacent node pattern value (for example, when the adjacent node pattern value is 63, 64 codings are performed). In an embodiment, the point cloud encoder can reduce the coding complexity by changing the adjacent node pattern value (for example, based on a table that changes 64 to 10 or 6).
[0104] FIG. 8 is a diagram showing an example of the point configuration for each LOD according to an embodiment.
[0105] As described with reference to FIGS. 1 to 7, before the feature coding is performed, the encoded geometry is reconstructed (restored). When direct coding is applied, the operation of geometry reconstruction includes changing the arrangement of the directly coded points (for example, arranging the directly coded points in front of the point cloud data). When trisoup geometry coding is applied, the process of geometry reconstruction includes the processes of triangle reconstruction, upsampling, and voxelization. Since the feature is subordinate to the geometry, the feature coding is performed based on the reconstructed geometry.
[0106] The point cloud encoder (e.g., the LOD generation unit 40009) classifies (reorganizes) points for each LOD. The drawing shows the point cloud content corresponding to the LOD. In the figure, the left side shows the original point cloud content. The second from the left in the figure shows the distribution of the points of the lowest LOD, and the rightmost side shows the distribution of the points of the highest LOD. That is, the points of the lowest LOD have a sparse distribution, and the points of the highest LOD have a fine distribution. That is, as the LOD increases along the arrow direction shown on the lower side of the drawing, the interval (or distance) between points becomes shorter.
[0107] FIG. 9 is a diagram showing an example of the point configuration for each LOD according to an embodiment.
[0108] As described with reference to FIGS. 1 to 8, the point cloud content providing system or the point cloud encoder (e.g., the point cloud video encoder 10002, the point cloud encoder in FIG. 4, or the LOD generation unit 40009) generates an LOD. The LOD is generated by rearranging points in a set of refinement levels according to a set LOD distance value (or a set of Euclidean Distances). The LOD generation process is performed not only by the point cloud encoder but also by the point cloud decoder.
[0109] The upper side of FIG. 9 shows an example of points (P0 to P9) of the point cloud content distributed in the three-dimensional space. The original order in FIG. 9 shows the order of the points P0 to P9 before LOD generation. The LOD based order in FIG. 9 shows the order of the points by LOD generation. The points are rearranged for each LOD. Also, a higher LOD includes the points belonging to a lower LOD. As shown in FIG. 9, LOD0 includes P0, P5, P4, and P2. LOD1 includes the points of LOD0 and P1, P6, and P3. LOD2 includes the points of LOD0, the points of LOD1, and P9, P8, and P7.
[0110] As described with reference to FIG. 4, the point cloud encoder according to the embodiment can perform prediction conversion coding, lift conversion coding, and RAHT conversion coding selectively or in combination.
[0111] The point cloud encoder according to the embodiment performs prediction conversion coding for generating a predictor for each point and setting the prediction feature (or prediction feature value) of each point. That is, N predictors are generated for N points. The predictor according to the embodiment can calculate a weight value (=1 / distance) based on the index information for adjacent points existing within the distance set for each LOD value and LOD of each point and the distance value to the adjacent points.
[0112] The prediction feature (or feature value) according to the embodiment is set as the average value of the values obtained by multiplying the features (or feature values, such as hue, reflectance, etc.) of the adjacent points set in the predictor of each point by the weight (or weight value) calculated based on the distance to each adjacent point. The point cloud encoder according to the embodiment (for example, the coefficient quantization unit 40011) can quantize and inverse-quantize the residual value (also referred to as residuals, residual features, residual feature values, feature prediction residuals, etc.) obtained by subtracting the prediction feature (feature value) from the feature (feature value) of each point. The quantization process is as shown in the following table.
[0113] Table. Attribute prediction residuals quantization pseudo code
[0114] int PCCQuantization(inT value, inT quantStep) {
[0115] if( value >=0) {
[0116] return floor(value / quantStep + 1.0 / 3.0);
[0117] } else {
[0118] return -floor(-value / quantStep + 1.0 / 3.0);
[0119] }
[0120] }
[0121] Table. Attribute prediction residuals inverse quantization pseudo Code
[0122] int PCCInverseQuantization(inT value, inT quantStep) {
[0123] if( quantStep ==0) {
[0124] return value;
[0125] } else {
[0126] return value*quantStep;
[0127] }
[0128] }
[0129] The point cloud encoder according to the embodiment (e.g., the arithmetic encoder 40012) entropy-codes the quantized and inverse-quantized residual values as described above if there are points adjacent to the predictor of each point. The point cloud encoder according to the embodiment (e.g., the arithmetic encoder 40012) does not perform the above-described process if there are no points adjacent to the predictor of each point, and entropy-codes the characteristics of the corresponding point.
[0130] The point cloud encoder according to the embodiment (for example, the lift conversion unit 40010) generates a predictor for each point, sets the LOD calculated for the predictor, registers adjacent points, sets a weight value based on the distance to the adjacent points, and performs lift conversion coding. The lift conversion coding according to the embodiment is similar to the above-described measurement conversion coding, but differs in that a weight value is cumulatively applied to the characteristic value. The process of cumulatively applying the weight value to the characteristic value according to the embodiment is as follows.
[0131] 1) Generate an array QW (QuantizationWieght) for storing the weight value of each point. The initial value of all elements of QW is 1.0. Add the value obtained by multiplying the weight value of the predictor of the current point by the QW value of the predictor index of the adjacent node registered in the predictor.
[0132] 2) Lift prediction process: To calculate the predicted characteristic value, subtract the value obtained by multiplying the weight value by the characteristic value of the point from the existing characteristic value.
[0133] 3) Generate temporary arrays called updateweight and update, and initialize the temporary arrays to 0.
[0134] 4) For all predictors, accumulate and sum up the weight value calculated by further multiplying the weight value calculated for each predictor by the weight value stored in QW corresponding to the predictor index as the index of the adjacent node in the updateweight array. In the update array, accumulate and sum up the value obtained by multiplying the weight value calculated by the characteristic value of the index of the adjacent node.
[0135] 5) Lift update process: For all predictors, divide the characteristic value of the update array by the weight value of the updateweight array of the predictor index, and add the divided value to the existing characteristic value again.
[0136] 6) For all predictors, calculate the predicted trait values by further multiplying the trait values updated in the lift update process by the weight values (stored in QW) updated in the lift prediction process. The point cloud encoder according to the embodiment (for example, the coefficient quantization unit 40011) quantizes the predicted trait values. Also, the point cloud encoder (for example, the arithmetic encoder 40012) performs entropy coding on the quantized trait values.
[0137] The point cloud encoder according to the embodiment (for example, the RAHT transform unit 40008) performs RAHT transform coding that predicts the traits of the upper-level nodes using the traits associated with the lower-level nodes of the octree. RAHT transform coding is an example of trait intra-coding by octree backward scan. The point cloud encoder according to the embodiment scans from the voxels to the entire region and repeatedly performs the merging process to the root node while fitting each step's voxel to a larger block. The merging process according to the embodiment is performed only for the occupied nodes. For the empty nodes, the merging process is not performed, and the merging process is performed for the node immediately above the empty node.
[0138] The following formula shows the RAHT transform matrix. JPEG2025107332000004.jpg11104 represents the average trait value of the voxel at level l. The weight value of JPEG2025107332000005.jpg10125 is JPEG2025107332000006.jpg1173.
[0139]
Number
[0140] JPEG2025107332000008.jpg953 is a low-pass value and is used in the merging process at the next upper level. JPEG2025107332000009.jpg1177 are high-pass coefficients, and the high-pass coefficients at each step are quantized and entropy-coded (e.g., encoding by the computing encoder 400012). The weighting values are calculated by JPEG2025107332000010.jpg1387. The root node is the last generated as follows by JPEG2025107332000011.jpg1273.
[0141] [Number]
[0142] FIG. 10 is a diagram showing an example of a Point Cloud Decoder according to an embodiment.
[0143] The point cloud decoder shown in FIG. 10 is an example of the point cloud video decoder 10006 shown in FIG. 1, and performs operations identical or similar to those of the in-cloud video decoder 10006 described in FIG. 1. As shown in the figure, the point cloud decoder receives a geometry bitstream and an Attribute bitstream included in one or more bitstreams. The point cloud decoder includes a geometry decoder and an Attribute decoder. The geometry decoder performs geometry decoding on the geometry bitstream and outputs the decoded geometry. The Attribute decoder performs Attribute decoding based on the decoded geometry and the Attribute bitstream and outputs the decoded Attributes. The decoded geometry and the decoded Attributes are used to restore the point cloud content (decoded point cloud).
[0144] FIG. 11 is a diagram showing an example of a point cloud decoder according to an embodiment.
[0145] The point cloud decoder shown in FIG. 11 is an example of the point cloud decoder described in FIG. 10, and performs a decoding operation which is the reverse process of the encoding operation of the point cloud encoder described in FIGS. 1 to 9.
[0146] As described in FIGS. 1 and 10, the point cloud decoder performs geometry decoding and feature decoding. The geometry decoding is performed prior to the feature decoding.
[0147] The point cloud decoder according to the embodiment includes an arithmetic decoder 11000, a synthesize octree 11001, a synthesize surface approximation 11002, a reconstruct geometry 11003, an inverse transform coordinates 11004, an arithmetic decoder 11005, an inverse quantize 11006, a RAHT transform 11007, a generate LOD 11008, an Inverse lifting 11009 and / or an inverse transform colors 11010.
[0148] The arithmetic decoder 11000, octree synthesis unit 11001, surface approximation synthesis unit 11002, geometry reconstruction unit 11003, and coordinate system inverse conversion unit 11004 perform geometry decoding. The geometry decoding according to the embodiment includes direct coding and trisoup geometry decoding. The direct coding and trisoup geometry decoding are selectively applied. Also, the geometry decoding is not limited to the above examples and is performed in the reverse process of the geometry encoding described in FIGS. 1 to 9.
[0149] The arithmetic decoder 11000 according to the embodiment decodes the received geometry bit stream based on arithmetic coding. The operation of the arithmetic decoder 11000 corresponds to the reverse process of the arithmetic encoder 40004.
[0150] The octree synthesis unit 11001 according to the embodiment obtains occupancy codes from the decoded geometry bit stream (or the decoded result, information regarding the secured geometry) and generates an octree. The specific description of the occupancy codes is as described in FIGS. 1 to 9.
[0151] The surface approximation synthesis unit 11002 according to the embodiment synthesizes a surface based on the decoded geometry and / or the generated octree when trisoup geometry coding is applied.
[0152] The geometry reconstruction unit 11003 according to the embodiment regenerates the geometry based on the surface and / or the decoded geometry. As described with reference to FIGS. 1 to 9, direct coding and trisoup geometry coding are selectively applied. Therefore, the geometry reconstruction unit 11003 directly obtains and adds the position information of the points to which direct coding is applied. Further, when trisoup geometry coding is applied, the geometry reconstruction unit 11003 performs the reconstruction operations of the geometry reconstruction unit 40005, for example, triangle reconstruction, upsampling, and voxelization operations, to restore the geometry. Since the specific content is as described in FIG. 6, it is omitted. The restored geometry includes a point cloud picture or a frame without features.
[0153] The coordinate system inverse conversion unit 11004 according to the embodiment converts the coordinate system based on the restored geometry to obtain the positions of the points.
[0154] The arithmetic decoder 11005, the inverse quantization unit 11006, the RAHT conversion unit 11007, the LOD generation unit 11008, the inverse lifting unit 11009, and / or the color inverse conversion unit 11010 perform the feature decoding described in FIG. 10. The feature decoding according to the embodiment includes RAHT (Region Adaptive Hierarchial Transform) decoding, prediction transform (Interpolaration-based hierarchical nearest-neighbour prediction-Prediction Transform) decoding, and lifting transform (interpolation-based hierarchical nearest-neighbour prediction with an update / lifting step (Lifting Transform)) decoding. The above three decodings are selectively used, or a combination of one or more decodings is used. Further, the feature decoding according to the embodiment is not limited to the above-described examples.
[0155] The arithmetic decoder 11005 according to the embodiment decodes the feature bitstream into arithmetic coding.
[0156] The inverse quantization unit 11006 according to the embodiment inverse quantizes the decoded feature bit stream or the information regarding the features for which the decoding result is ensured, and outputs the inverse quantized features (or feature values). The inverse quantization is selectively applied based on the feature encoding of the point cloud encoder.
[0157] In the embodiment, the RAHT transform unit 11007, the LOD generation unit 11008, and / or the inverse lifting unit 11009 processes the reconstructed geometry and the inverse quantized features. As described above, the RAHT transform unit 11007, the LOD generation unit 11008, and / or the inverse lifting unit 11009 selectively performs the corresponding decoding operation by the encoding of the point cloud encoder.
[0158] The color inverse conversion unit 11010 according to the embodiment performs inverse conversion coding for inverse converting the color values (or textures) included in the decoded features. The operation of the color inverse conversion unit 11010 is selectively performed based on the operation of the color conversion unit 40006 of the point cloud encoder.
[0159] The elements of the point cloud decoder in FIG. 11, although not shown, include hardware, software, firmware, or a combination thereof including one or more processors or integrated circuits communicatively set with one or more memories included in the point cloud providing device. One or more processors perform any of the operations and / or functions of the elements of the point cloud decoder in FIG. 11 described above. Also, one or more processors operate or execute a software program and / or a set of instructions for performing the operations and / or functions of the elements of the point cloud decoder in FIG. 11.
[0160] FIG. 12 shows an example of a transmission device according to the embodiment.
[0161] The transmission device shown in FIG. 12 is an example of the transmission device 10000 in FIG. 1 (or the point cloud encoder in FIG. 4). The transmission device shown in FIG. 12 performs any of the operations and methods that are the same as or similar to the operations and encoding methods of the point cloud encoder described in FIGS. 1 to 9. The transmission device according to the embodiment includes a data input unit 12000, a quantization processing unit 12001, a voxelization processing unit 12002, an octree occupancy code generation unit 12003, a surface model processing unit 12004, an intra / inter coding processing unit 12005, an arithmetic coder 12006, a metadata processing unit 12007, a hue conversion processing unit 12008, a characteristic conversion processing unit (or attribute conversion processing unit) 12009, a prediction / lift / RAHT conversion processing unit 12010, an arithmetic coder 12011, and / or a transmission processing unit 12012.
[0162] The data input unit 12000 according to the embodiment receives or acquires point cloud data. The data input unit 12000 performs an operation and / or acquisition method that is the same as or similar to the operation and / or acquisition method of the point cloud video acquisition unit 10001 (or the acquisition process 20000 shown in FIG. 2).
[0163] The data input unit 12000, the quantization processing unit 12001, the voxelization processing unit 12002, the octree occupancy code generation unit 12003, the surface model processing unit 12004, the intra / inter coding processing unit 12005, and the arithmetic coder 12006 perform geometry encoding. Since the geometry encoding according to the embodiment is the same as or similar to the geometry encoding described in FIGS. 1 to 9, specific description is omitted.
[0164] The quantization processing unit 12001 according to the embodiment quantizes geometry (for example, the position value of a point, or the position value). The operation and / or quantization of the quantization processing unit 12001 is the same as or similar to the operation and / or quantization of the quantization unit 40001 shown in FIG. 4. The specific description is as described in FIGS. 1 to 9.
[0165] The voxelization processing unit 12002 according to the embodiment voxelizes the position values of the quantized points. The voxelization processing unit 120002 performs operations and / or processes identical or similar to the operations and / or processes of the quantization unit 40001 and / or the voxelization process shown in FIG. 4. The specific description is as described in FIGS. 1 to 9.
[0166] The octree occupancy code generation unit 12003 according to the embodiment performs octree coding on the positions of the voxelized points based on the octree structure. The octree occupancy code generation unit 12003 generates occupancy codes. The octree occupancy code generation unit 12003 performs operations and / or methods identical or similar to the operations and / or methods of the point cloud encoder (or octree analysis unit 40002) described in FIGS. 4 and 6. The specific description is as described in FIGS. 1 to 9.
[0167] The surface model processing unit 12004 according to the embodiment performs trisoup geometry encoding that reconstructs the positions of points in a specific region (or node) based on a surface model onto a voxel basis. The surface model processing unit 12004 performs operations and / or methods identical or similar to the operations and / or methods of the point cloud encoder (for example, surface approximation analysis unit 40003) shown in FIG. 4. The specific description is as described in FIGS. 1 to 9.
[0168] The intra / inter coding processing unit 12005 according to the embodiment performs intra / inter coding on the point cloud data. The intra / inter coding processing unit 12005 performs coding identical or similar to the intra / inter coding described in FIG. 7. The specific description is as described in FIG. 7. In the embodiment, the intra / inter coding processing unit 12005 is included in the arithmetic coder 12006.
[0169] The arithmetic coder 12006 according to the embodiment entropy - encodes the octree and / or the approximated octree of the point cloud data. For example, the encoding method includes the arithmetic encoding method. The arithmetic coder 12006 performs operations and / or methods that are the same as or similar to the operations and / or methods of the arithmetic encoder 40004.
[0170] The metadata processing unit 12007 according to the embodiment processes metadata related to the point cloud data, such as set values, and provides them to necessary processing procedures such as geometric encoding and / or texture encoding. Also, the metadata processing unit 12007 according to the embodiment generates and / or processes signaling information related to geometric encoding and / or texture encoding. The signaling information according to the embodiment is encoded separately from geometric encoding and / or texture encoding. Also, the signaling information according to the embodiment may be interleaved.
[0171] The hue conversion processing unit 12008, the texture conversion processing unit 12009, the prediction / lift / RAHT conversion processing unit 12010, and the arithmetic coder 12011 perform texture encoding. Since the texture encoding according to the embodiment is the same as or similar to the texture encoding described in FIGS. 1 to 9, specific descriptions are omitted.
[0172] The hue conversion processing unit 12008 according to the embodiment performs hue conversion coding for converting the hue value included in the texture. The hue conversion processing unit 12008 performs hue conversion coding based on the reconstructed geometry. The description of the reconstructed geometry is as described in FIGS. 1 to 9. Also, it performs operations and / or methods that are the same as or similar to the operations and / or methods of the color conversion unit 40006 described in FIG. 4. Specific descriptions are omitted.
[0173] The characteristic conversion processing unit 12009 according to the embodiment performs characteristic conversion that converts characteristics based on positions where geometry encoding is not performed and / or the reconstructed geometry. The characteristic conversion processing unit 12009 performs operations and / or methods that are the same as or similar to the operations and / or methods of the characteristic conversion unit 40007 described in FIG. 4. A specific description is omitted. The prediction / lift / RAHT conversion processing unit 12010 according to the embodiment codes the converted characteristics by any one or a combination of RAHT coding, prediction conversion coding, and lift conversion coding. The prediction / lift / RAHT conversion processing unit 12010 performs any of the operations that are the same as or similar to the operations of the RAHT conversion unit 40008, the LOD generation unit 40009, and the lift conversion unit 40010 described in FIG. 4. Also, since the descriptions of prediction conversion coding, lift conversion coding, and RAHT conversion coding are as described in FIGS. 1 to 9, a specific description is omitted.
[0174] The arithmetic coder 12011 according to the embodiment codes the coded characteristics based on arithmetic coding. The arithmetic coder 12011 performs operations and / or methods that are the same as or similar to the operations and / or methods of the arithmetic encoder 400012.
[0175] The transmission processing unit 12012 according to the embodiment transmits each bitstream including the encoded geometry and / or the encoded trait and the metadata information, or configures the encoded geometry and / or the encoded trait and the metadata information into one bitstream and transmits it. When the encoded geometry and / or the encoded trait and the metadata information according to the embodiment are configured into one bitstream, the bitstream includes one or more sub-bitstreams. The bitstream according to the embodiment includes signaling information including SPS (Sequence Parameter Set) for sequence-level signaling, GPS (Geometry Parameter Set) for signaling of geometry information coding, APS (Attribute Parameter Set) for signaling of trait information coding, TPS (Tile Parameter Set) for tile-level signaling, and slice data. The slice data includes information regarding one or more slices. One slice according to the embodiment is one geometry bitstream (Geom0 0 ) and one or more trait bitstreams (Attr0 0 , Attr1 0 ).
[0176] A slice refers to a series of syntax elements indicating the whole or a part of the coded point cloud frame.
[0177] The TPS according to the embodiment includes information regarding each tile (e.g., coordinate value information of the bounding box and height / size information, etc.) for one or more tiles. The geometry bitstream includes a header and a payload. The header of the geometry bitstream according to the embodiment includes identification information (geom_parameter_set_id) of the parameter set included in the GPS, tile identifier (geom_tile_id), slice identifier (geom_slice_id), and information regarding the data included in the payload, etc. As described above, the metadata processing unit 12007 according to the embodiment can generate and / or process signaling information and transmit it to the transmission processing unit 12012. In the embodiment, the element performing geometry encoding and the element performing trait encoding can share mutual data / information as if they were dot-line processed. The transmission processing unit 12012 according to the embodiment performs the same or similar operations and / or transmission methods as the operations and / or transmission methods of the transmitter 10003. Since the specific description is as described in FIGS. 1 and 2, it is omitted.
[0178] FIG. 13 shows an example of a receiving device according to the embodiment.
[0179] The receiving device shown in FIG. 13 is an example of the receiving device 10004 in FIG. 1 (or the point cloud decoder in FIGS. 10 and 11). The receiving device shown in FIG. 13 performs any one of the same or similar operations and methods as the operations and decoding methods of the point cloud decoder described in FIGS. 1 to 11.
[0180] The receiving device according to the embodiment includes a receiving unit 13000, a reception processing unit 13001, an arithmetic decoder 13002, an octree reconstruction processing unit 13003 for an occupancy code base, a surface model processing unit (triangle reconstruction, upsampling, voxelization) 13004, an inverse quantization processing unit 13005, a metadata parser 13006, an arithmetic decoder 13007, an inverse quantization processing unit 13008, a prediction / lift / RAHT inverse transformation processing unit 13009, a hue inverse transformation processing unit 13010 and / or a renderer 13011. Each component of the decoding according to the embodiment performs the reverse process of the component of the encoding according to the embodiment.
[0181] The receiving unit 13000 according to the embodiment receives point cloud data. The receiving unit 13000 performs operations and / or receiving methods that are the same as or similar to the operations and / or receiving methods of the receiver 10005 in FIG. 1. Specific descriptions are omitted.
[0182] The reception processing unit 13001 according to the embodiment obtains a geometry bit stream and / or a texture bit stream from the received data. The reception processing unit 13001 is included in the receiving unit 13000.
[0183] The arithmetic decoder 13002, the octree reconstruction processing unit 13003 for the occupancy code base, the surface model processing unit 13004 and the inverse quantization processing unit 13005 perform geometry decoding. Since the geometry decoding according to the embodiment is the same as or similar to the geometry decoding described in FIGS. 1 to 10, specific descriptions are omitted.
[0184] The arithmetic decoder 13002 according to the embodiment decodes the geometry bit stream based on arithmetic coding. The arithmetic decoder 13002 performs operations and / or coding that are the same as or similar to the operations and / or coding of the arithmetic decoder 11000.
[0185] The octree reconstruction processing unit 13003 of the occupancy code base according to the embodiment acquires the occupancy code from the decoded geometry bitstream (or from the decoded result, information regarding the secured geometry) and reconstructs the octree. The octree reconstruction processing unit 13003 of the occupancy code base performs operations and / or methods that are the same as or similar to the operations of the octree synthesis unit 11001 and / or the octree generation method. When trisoup geometry encoding is applied, the surface model processing unit 13004 according to the embodiment performs trisoup geometry decoding and related geometry reconstruction (for example, triangle reconstruction, upsampling, voxelization) based on the surface model method. The surface model processing unit 13004 performs operations that are the same as or similar to the operations of the surface approximation synthesis unit 11002 and / or the geometry reconstruction unit 11003.
[0186] The inverse quantization processing unit 13005 according to the embodiment inverse quantizes the decoded geometry.
[0187] The metadata parser 13006 according to the embodiment analyzes the metadata included in the received point cloud data, for example, set values and the like. The metadata parser 13006 transmits the metadata to geometry decoding and / or trait decoding. Since the specific description regarding the metadata is as described in FIG. 12, it is omitted.
[0188] The arithmetic decoder 13007, the inverse quantization processing unit 13008, the prediction / lift / RAHT inverse transformation processing unit 13009, and the hue inverse transformation processing unit 13010 perform trait decoding. Since the trait decoding is the same as or similar to the trait decoding described in FIGS. 1 to 10, the specific description is omitted.
[0189] The arithmetic decoder 13007 according to the embodiment decodes the trait bitstream into arithmetic coding. The arithmetic decoder 13007 decodes the trait bitstream based on the reconstructed geometry. The arithmetic decoder 13007 performs operations and / or coding that are the same as or similar to the operations and / or coding of the arithmetic decoder 11005.
[0190] The inverse quantization processing unit 13008 according to the embodiment inverse quantizes the decoded characteristic bit stream. The inverse quantization processing unit 13008 performs operations and / or methods that are the same as or similar to the operations and / or methods of the inverse quantization unit 11006.
[0191] The prediction / lift / RAHT inverse transform processing unit 13009 according to the embodiment processes the reconstructed geometry and the inverse quantized characteristics. The prediction / lift / RAHT inverse transform processing unit 13009 performs any of the operations and / or decoding that are the same as or similar to the operations and / or decoding of the RAHT transform unit 11007, the LOD generation unit 11008, and / or the inverse lift unit 11009. The hue inverse transform processing unit 13010 according to the embodiment performs inverse transform coding for inverse transforming the color values (or textures) included in the decoded characteristics. The hue inverse transform processing unit 13010 performs operations and / or inverse transform coding that are the same as or similar to the operations and / or inverse transform coding of the color inverse transform unit 11010. The renderer 13011 according to the embodiment renders point cloud data.
[0192] FIG. 14 is a diagram showing an example of a structure that can be linked to the method / apparatus for transmitting and receiving point cloud data according to the embodiment.
[0193] The structure of FIG. 14 shows a configuration in which any one of the server 1460, the robot 1410, the autonomous driving vehicle 1420, the XR device 1430, the smartphone 1440, the home appliance 1450, and / or the HMD 1470 is connected to the cloud network 1410. Devices such as the robot 1410, the autonomous driving vehicle 1420, the XR device 1430, the smartphone 1440, or the home appliance 1450 are also called devices. Also, the XR device 1430 corresponds to or is linked to the point cloud data (PCC) device according to the embodiment.
[0194] The cloud network 1400 refers to a network that constitutes a part of the cloud computing infrastructure or exists within the cloud computing infrastructure. Here, the cloud network 1400 is configured using a 3G network, a 4G or LTE network, or a 5G network, etc.
[0195] The server 1460 is connected by the cloud network 1400 to any one of the robot 1410, the autonomous vehicle 1420, the XR device 1430, the smartphone 1440, the home appliance 1450, and / or the HMD 1470, and can assist with at least a part of the processing of the connected devices 1410 to 1470.
[0196] The HMD (Head-Mount Display) 1470 indicates any one of the types in which the XR device and / or the PCC device according to the embodiment is implemented. The device of the HMD type according to the embodiment includes a communications unit, a control unit, a memory unit, an I / O unit, a sensor unit, a power supply unit, etc.
[0197] Hereinafter, various embodiments of the devices 1410 to 1450 to which the above technology is applied will be described. Here, the devices 1410 to 1450 shown in FIG. 14 can be interlocked / connected to the point cloud data transceiver according to the above-described embodiment.
[0198] <PCC+XR>
[0199] The XR / PCC device 1430 is applied with PCC and / or XR (AR+VR) technology and can also be implemented in an HMD (Head-Mount Display), a HUD (Head-Up Display) provided in a vehicle, a TV, a mobile phone, a smartphone, a computer, a wearable device, a home appliance, a digital signboard, a vehicle, a stationary robot, a mobile robot, etc.
[0200] The XR / PCC device 1430 can obtain information about the surrounding space or real objects by analyzing 3D point cloud data or image data acquired by various sensors or from external devices to generate position data and characteristic data for 3D points, and output an XR object by rendering. For example, the XR / PCC device 1430 can output an XR object including additional information about the recognized object corresponding to the recognized object.
[0201] <PCC+XR+Mobile Phone>
[0202] The XR / PCC device 1430 is implemented in a mobile phone 1440 and the like by applying the PCC technology.
[0203] The mobile phone 1440 decrypts and displays point cloud content based on the PCC technology.
[0204] <PCC+Autonomous Driving+XR>
[0205] The autonomous driving vehicle 1420 is implemented in a mobile robot, a vehicle, a drone, etc. by applying the PCC technology and the XR technology.
[0206] The autonomous driving vehicle 1420 to which the XR / PCC technology is applied means an autonomous driving vehicle equipped with means for providing an XR video, an autonomous driving vehicle that is a target of control / interaction within the XR video, etc. In particular, the autonomous driving vehicle 1420 that is a target of control / interaction within the XR video is distinguished from the XR device 1430 and interlocked with each other.
[0207] The autonomous driving vehicle 1420 equipped with means for providing an XR / PCC video obtains sensor information from sensors including a camera, and outputs an XR / PCC video generated based on the obtained sensor information. For example, the autonomous driving vehicle 1420 can provide an XR / PCC object corresponding to a real object or an object within the screen to the passengers by outputting an XR / PCC video with an HUD.
[0208] At this time, when the XR / PCC object is output to the HUD, at least a part of the XR / PCC object is output so as to overlap the actual object that the rider's line of sight is directed to. On the contrary, when the XR / PCC object is output to a display provided in the autonomous vehicle, at least a part of the XR / PCC object is output so as to overlap the object within the screen. For example, the autonomous vehicle 1220 can output XR / PCC objects corresponding to objects such as a road, other vehicles, traffic lights, traffic signs, motorcycles, pedestrians, buildings, and the like.
[0209] The VR (Virtual Reality) technology, AR (Augmented Reality) technology, MR (Mixed Reality) technology, and / or PCC (Point Cloud Compression) technology according to the embodiments are applicable to various devices.
[0210] That is, the VR technology is a display technology that provides only CG images of real objects and backgrounds. On the contrary, the AR technology is a technology that shows virtual CG images together on the video of actual things. Also, the MR technology is similar to the above AR technology in that it shows a virtual object mixed in the real world. However, in the AR technology, the distinction between the real object and the virtual object composed of the CG image is clear, and the virtual object is used in a form that complements the real object, while in the MR technology, it is distinguished from the AR technology in that the virtual object and the real object are regarded as having the same nature. More specifically, for example, the hologram service is an example where the above MR technology is applied.
[0211] However, recently, rather than clearly distinguishing between VR, AR, and MR technologies, they are also called XR (extended Reality) technologies. Therefore, the embodiments of the present invention are applicable to any of VR, AR, MR, and XR technologies. Such technologies are applied with encoding / decoding of PCC, V-PCC, and G-PCC technology bases.
[0212] The PCC method / device according to the embodiment can be applied to a vehicle that provides an autonomous driving service.
[0213] The vehicle providing the autonomous driving service is communicably connected to the PCC device, either wirelessly or via a wired connection.
[0214] The point cloud data transmission method / device according to the embodiment is interpreted as terms such as the transmission device 10000 in FIG. 1, the point cloud video encoder 10002, the transmitter 10003, the acquisition - encoding - transmission 20000 - 20001 - 20002 in FIG. 2, the encoder in FIG. 4, the transmission device in FIG. 12, the device in FIG. 14, and the encoder in FIG. 21.
[0215] The point cloud data reception method / device according to the embodiment is interpreted as terms such as the reception device 10004 in FIG. 1, the receiver 10005, the point cloud video decoder 10006, the transmission - decoding - rendering 20002 - 20003 - 20004 in FIG. 2, the decoders in FIGS. 10 and 11, the reception device in FIG. 13, the device in FIG. 14, and the decoder in FIG. 22.
[0216] Also, the point cloud data transmission / reception method / device according to the embodiment is also simply referred to as the method / device according to the embodiment.
[0217] In the embodiment, the geometric data, geometric information, position information, etc. that make up the point cloud data are interpreted in the same meaning. The characteristic data, characteristic information, attribute information, etc. that make up the point cloud data are also interpreted in the same meaning.
[0218] The point cloud data transmission / reception method / device according to the embodiment provides an extended predictive tree construction scheme for low - latency 3D map point cloud geometry information compression (A method to build predictive geometry tree for low - latency geometry coding of 3D map point cloud).
[0219] The point cloud data transmission / reception method / apparatus according to the embodiment assists in constructing a predictive tree for efficient geometric compression of G-PCC (Geometry-based Point Cloud Compression) when integrating point cloud frames captured by a LiDAR (Light Detection and Ranging) device into one point cloud content. For example, it includes an origin selection method and a signaling method, a laser angle-based alignment method for generating a predictive tree, and / or a rapid predictive tree construction method, etc.
[0220] The embodiment relates to a solution for improving the compression efficiency of G-PCC for three-dimensional point cloud data compression. Hereinafter, the encoder and the coder are referred to as the encoder, and the decoder and the decoder are referred to as the decoder.
[0221] A point cloud consists of a set of points, and each point has geometric information and characteristic information. The geometric information is three-dimensional position (XYZ) information, and the characteristic information is hue (RGB, YUV, etc.) and / or reflectance value.
[0222] In the G-PCC encoding process, the point cloud is divided into tiles by regions, and each tile is divided into slices for parallel processing. The geometry is compressed in each slice unit, and the attribute information is compressed based on the reconstructed geometry (decoded geometry) with the position information changed by the compression.
[0223] In the G-PCC decoding process, the encoded geometry bitstream and attribute bitstream of the slice unit are transmitted to decode the geometry, and the attribute information is decoded based on the geometry reconstructed in the decoding process.
[0224] Compression techniques based on an octree, a predictive tree, or a trisoup are used for geometric information compression.
[0225] The point cloud data transmission / reception method / apparatus according to the embodiment performs a geometric compression technique based on a predictive tree for increasing the geometric compression efficiency of 3D map content captured by a lidar device.
[0226] For capturing point cloud content, a laser pulse is irradiated and the time for the reflected light to return is measured to measure the position coordinates of the reflector. Depth information can be obtained by a lidar device using such a laser system. The point cloud content generated by the lidar device may consist of a plurality of frames, or a plurality of frames may be integrated into one content.
[0227] 3D map point cloud content means data generated by capturing a plurality of frames with a lidar device and integrating them into one content. In this case, since the data of each frame captured at different positions of the center position of the lidar device are mixed, the angular characteristics shown by the data captured by the lidar device, that is, the angle The rule between points when changed to JPEG2025107332000013.jpg989 is hidden, and thus the application of the angle mode is not more efficient than the compression based on the Cartesian coordinate system.
[0228] Therefore, the geometric compression method based on a predictive tree applied to 3D map point cloud content cannot increase the compression efficiency using the angle mode, and a solution for increasing the compression efficiency from the regularity of the points in the content is required.
[0229] Since the prediction tree predicts the position of the current point based on the vector of the parent node, the residual value between the predicted point and the current point becomes small depending on whether a regular parent node is well selected, and the size of the bitstream can be reduced.
[0230] In an embodiment, it supports a prediction tree construction method for assisting in efficient geometric compression of a prediction tree base of 3D map data captured by a lidar device and integrated into one content.
[0231] The prediction tree construction according to the embodiment is performed by the geometry encoder of the PCC encoder and is restored by the geometry decoding process of the PCC decoder.
[0232] FIG. 15 shows additional attribute data included in the point cloud data according to the embodiment.
[0233] Point cloud data compressed and restored by the transmission device 10000 in FIG. 1, the point cloud video encoder 10002, the transmitter 10003, the acquisition-encoding-transmission 20000-20001-20002 in FIG. 2, the encoder in FIG. 4, the transmission device in FIG. 12, the device in FIG. 14, the encoder in FIG. 21, the reception device 10004 in FIG. 1, the receiver 10005, the point cloud video decoder 10006, the transmission-decoding-rendering 20002-20003-20004 in FIG. 2, the decoders in FIGS. 10 and 11, the reception device in FIG. 13, the device in FIG. 14, and the decoder in FIG. 22 have attributes as shown in FIG. 15.
[0234] When there is a laser angle value according to the embodiment, a method for selecting the origin position:
[0235] Referring to FIG. 15, the 3D map data captured by a LiDAR-equipped device and integrated into one piece of content has additional attribute data such as time, laser angle, normal position (nx, ny, nz), etc., in addition to the position (x, y, z) and attribute (red, green, blue, reflectance) values.
[0236] FIG. 16 shows an example of the origin position for point cloud data according to an embodiment.
[0237] The transmission device 10000 in FIG. 1, the point cloud video encoder 10002, the transmitter 10003, the acquisition-encoding-transmission 20000-20001-20002 in FIG. 2, the encoder in FIG. 4, the transmission device in FIG. 12, the device in FIG. 14, the encoder in FIG. 21, the reception device 10004 in FIG. 1, the receiver 10005, the point cloud video decoder 10006, the transmission-decoding-rendering 20002-20003-20004 in FIG. 2, the decoders in FIGS. 10 and 11, the reception device in FIG. 13, the device in FIG. 14, and the decoder in FIG. 22, etc., set the origin of the point cloud data and the point cloud data in the form of a 3D map for compression / decompression.
[0238] When performing geometry compression based on a prediction tree on a 3D map point cloud, the left, bottom, and front positions of the bounding box of the slice are set as the position of the origin 16000. The position of the origin affects the point alignment process for generating the prediction tree, the aligned form affects the construction of the prediction tree, the prediction tree affects the predicted value, and thus affects the residual value with the predicted value and further affects the size of the bitstream.
[0239] That is, when the encoder and / or decoder according to the embodiment processes slice 0, one point of the bounding box corresponding to slice 0 is processed as the origin. The origin of the bounding box corresponding to slice 1 is at the left / bottom / front position of the bounding box.
[0240] FIG. 17 shows an example of the origin position according to the embodiment.
[0241] When the point cloud content has a laser angle value for each point, the method / apparatus according to the embodiment can calculate the origin position within the slice by the following process.
[0242] - A candidate angle value (origin_laser_angle) corresponding to the origin is input. For example, 90° is the candidate angle value.
[0243] - The position (origin_direction) of the point corresponding to the origin is input. For example, left is the position of the point.
[0244] The following process is performed for all points belonging to the slice.
[0245] 1. When the laser angle value of point p is the same as the set origin_laser_angle (=90),
[0246] A. If there is no origin position value, set p as the origin position value.
[0247] B. If there is an origin position value, compare the position value of p with the origin position value and set it as the origin position value according to origin_direction.
[0248] For example, when origin_direction is left, if origin.x > p.x, set p as the origin position value. This is because p is located to the left of the origin.
[0249] That is, in the method / apparatus according to the embodiment, the coordinates for the points can be converted to set the origin position. The orthogonal coordinate system is changed to a spherical coordinate system with the origin as the reference.
[0250] The coordinate transformation is used when applying the angular mode of 1) predictive geom coding, and 2) converts and aligns coordinate values for point alignment in the general mode and / or angular mode of predictive geom.
[0251] Referring to FIG. 17, with reference to the newly set origin, the (x, y, z) orthogonal coordinates are changed to azimuth, radius, elevation (laser ID). Since the origin moves on the road, the corresponding azimuth and radius can be found.
[0252] FIG. 17 shows an example where origin_laser_angle = 90° and origin_direction = left. When setting the origin at 90° and to the left for each slice, the point displayed at 17000 becomes the origin position of each slice.
[0253] Comparing with the case where the left / bottom / front of the bounding box is set as the origin 16000, 17001, it can be seen that the position of the origin set based on the additional attribute data changes.
[0254] Referring to the attributes of the point cloud data, when the point cloud data indicates roads, buildings, people, etc., the position 17002 where the set of points starts corresponds to the road area.
[0255] That is, it can be seen that the origin 16000, 17001 of the left / bottom / front of the bounding box is far from the road area 17002 that is the reference for point arrangement. Conversely, it can be seen that the position of the origin set based on the additional attribute data is set as the road start point 17002.
[0256] FIG. 18 shows an example of setting the origin position when there is no laser angle according to the embodiment.
[0257] In FIG. 18, when there is no laser angle value for each point in the point cloud content in FIG. 17, the origin position within the slice is calculated by the following process.
[0258] To determine the position of the bounding box of the point corresponding to the origin, a reference axis is input. For example, the x-axis becomes the reference axis.
[0259] To determine the position of the bounding box of the point corresponding to the origin, a second reference axis is input. For example, it becomes the y-axis.
[0260] To determine the position of the bounding box of the point corresponding to the origin, a vector range is input. For example, a range of -0.2 to -1 is set.
[0261] To determine the position of the bounding box of the point corresponding to the origin, set the position of the origin when it belongs to the vector range. For example, set the origin to the left / top / front.
[0262] To determine the position of the bounding box of the point corresponding to the origin, the position of the origin when it does not belong to the vector range is input. For example, set the origin to the left / bottom / front.
[0263] The following process is performed for all points belonging to the slice.
[0264] Search for point L that exists at the minimum value and point R that exists at the maximum value with respect to the reference axis of point p.
[0265] In diff, which is the normalized value of the R - L value, check whether it corresponds to the vector range with respect to the second reference axis. If it belongs to the range, set the specified position as the origin.
[0266] For example, if the reference axis is the x-axis and the second reference axis is the y-axis, when belonging to the vector range, set the left / top / front as the origin. If not belonging, set the left / bottom / front as the origin.
[0267] For example, in order to set the origin at slices 18000, 18003, 18006, while searching for points in a certain direction 18002, 18005, 18008 along the reference axis and vector range at points 18001, 18004, 18007, set the origin that fits.
[0268] Laser angle base alignment method according to an embodiment:
[0269] The method / apparatus according to the embodiment aligns points (Points[*]) according to Morton code, radius, azimuth, elevation, sensor ID criteria, or the captured chronological order, etc. before generating the prediction tree.
[0270] The alignment method according to the embodiment is set according to the characteristics of the content. For example, in the case of content having the form of spinning data captured by LiDAR equipment, a prediction tree can be generated more efficiently when aligned based on azimuth.
[0271] Align the points with Morton code. Or azimuth alignment is more efficient. When the left / bottom / front of the bounding box is the origin, there may be a problem that the angular difference of azimuth between points is large and the error becomes large.
[0272] Since the generation of the prediction tree proceeds in order based on the aligned points, the order of the aligned points affects the configuration of the prediction tree, the prediction tree affects the prediction value, and thus affects the residual value with the prediction value and affects the size of the bitstream.
[0273] FIG. 19 shows an example of laser angle base alignment according to an embodiment.
[0274] The transmission device 10000 in FIG. 1, the point cloud video encoder 10002, the transmitter 10003, the acquisition - encoding - transmission 20000 - 20001 - 20002 in FIG. 2, the encoder in FIG. 4, the transmission device in FIG. 12, the device in FIG. 14, the encoder in FIG. 21, the reception device 10004 in FIG. 1, the receiver 10005, the point cloud video decoder 10006, the transmission - decoding - rendering 20002 - 20003 - 20004 in FIG. 2, the decoder in FIGS. 10 - 11, the reception device in FIG. 13, the device in FIG. 14, the decoder in FIG. 22, etc. set the origin as shown in FIGS. 16 and 18 using the attributes in FIG. 15, and align the points as shown in FIG. 19.
[0275] When there is a laser angle value for each point of the point cloud content, the points may be aligned based on the laser angle, and the laser angles may be grouped and aligned. For example, the laser angles from 0 to 5° may be regarded as the same laser angle to align the order of the points. When the laser angles are the same or the laser angle groups are the same, align based on the radius, and when the radii are the same or the radius groups are the same, align based on the elevation.
[0276] The point cloud content is in the form of a road, and the points of the point cloud data are as shown in FIG. 19.
[0277] The method / device according to the embodiment sets a point with a laser angle of 90° and coordinates on the left side of the axis as the origin 19001. Starting from the origin, align the points based on the laser angle 19002. When the laser angle values (or the laser angle reference ranges) are the same in the alignment process based on the laser angle, align the points based on the radius 19003.
[0278] Instead of setting the left / bottom / front 19004 of the bounding box of the slice as the origin, in the method / apparatus according to the embodiment, a point 19005 where the laser angle is 90° and is on the left side of the axis may be set as the new origin. Also, points may be aligned based on the laser angle. Since the position of the origin is set at the starting point 19005 of the road, points for objects on the road can be aligned in the order of the laser angle along the road. When the laser angle values between points are the same, the points can be aligned based on the radius.
[0279] Due to the characteristics of the point cloud content of the object on the road, when points are aligned based on the laser angle, the order of the aligned points has a form aligned along the road, so errors in the process of encoding / decoding the points can be effectively reduced.
[0280] Method for constructing a rapid prediction tree according to the embodiment:
[0281] In the method / apparatus according to the embodiment, a prediction tree is generated while selecting the nearest prediction point as the parent node through the KD-Tree generation / search process. This process takes a considerable amount of execution time. In a scenario aiming for low-latency geometry compression, the KD-Tree-based prediction tree generation technique has an issue of execution time.
[0282] When the point cloud content has a laser angle value for each point, a method of quickly constructing a prediction tree without using a KD-Tree can be applied by selecting the origin position based on the laser angle and aligning based on the laser angle.
[0283] The prediction tree generation process according to the embodiment is as follows.
[0284] 1. Set the first point as the root node. Set the current point as the latest point of the current laser angle. Set the current point as the first point of the current laser angle.
[0285] 2. Perform the following process for all points on the slice.
[0286] 1) If the latest point of the laser angle of point p exists, set the latest point as the parent node on the prediction tree of the current point. Again, set the current point as the latest value of the current laser angle.
[0287] 2) If the latest point of the laser angle of point p does not exist, set the current point as the latest point of the current laser angle. Set the current point as the first point of the current laser angle.
[0288] 3. Perform the following process for all laser angles.
[0289] 1) Set the first point of the previous laser angle as the parent node of the first point of the current laser angle.
[0290] Figure 20 shows an example of a laser group and prediction tree generation according to an embodiment.
[0291] If there are a first laser angle group 20000 and a second laser group 20001, the second laser group 20001 is the current layer angle group, and the first laser angle group 20000 is the laser group processed before the second laser angle group 20001.
[0292] For example, in the case of the current laser angle group 20001, the first point 20002 is set as the root node. The current point 20002 is set as the latest point of the current laser angle. The current point 20002 is set as the first point of the current laser angle.
[0293] If the latest point 20002 of the laser angle of point p20003 exists, the latest point 20002 is set as the parent node on the prediction tree of the current point 20003. Again, the current point 20003 is set as the latest value of the current laser angle.
[0294] Since the latest point of the laser angle of the next point 20004 is point 20003, point 20003 becomes the parent of point 20004.
[0295] The first point 20005 of the previous laser angle 20000 is set as the parent node of the first point 20002 of the current laser angle 20001.
[0296] As shown in FIG. 20, the method / apparatus according to the embodiment can generate a rapid prediction tree using points aligned based on the laser angle.
[0297] In the embodiment, the latest point means the first point among the points included in the corresponding group and aligned. For example, as described above, for encoding, the points are aligned based on the laser angle (group by azimuth value or azimuth range).
[0298] Also, points are captured by a LiDAR, and the captured points have strong regularity based on the radius and / or azimuth value.
[0299] As shown in FIG. 20, the latest point within the laser angle group becomes the first point. Specifically, the range of group 20001 by the laser angle (azimuth) is 0 to 5°, including point 20002, and since point 20002 is the first point in the aligned order, it becomes the root node (point). In group 20000 corresponding to the laser angle of 0 to 5°, among the aligned points, since point 20005 is the first point, it becomes the root. Therefore, if there are groups and points according to a specific laser angle in this way, the parent / child relationship between the points within the group can be set, and the parent / child relationship between the groups can be set.
[0300] For example, when the acquisition unit (lidar) according to the embodiment rotates (rotation) to capture points, there is a difference in time for each azimuth angle, but when capturing in the flash type, since it is captured at once by sensors that are numerous in a specific area, there is no time difference. In the meaning of the previous laser angle group according to the embodiment, when grouped into angles of 0 to 5 and 5 to 10, the 0 to 5° group is the previous laser angle group based on the 5 to 10° group standard. Also, when the device according to the embodiment rotates to capture (spinning LiDAR, a general case), there is a time difference.
[0301] FIG. 21 shows a point cloud data transmission device according to an embodiment.
[0302] The transmission device 10000 in FIG. 1, the point cloud video encoder 10002, the transmitter 10003, the acquisition-encoding-transmission 20000-20001-20002 in FIG. 2, the encoder in FIG. 4, the transmission device in FIG. 12, the device in FIG. 14, the encoder in FIG. 21, etc. are point cloud data transmission devices according to corresponding embodiments. Each component corresponds to hardware, software, a processor, and / or a combination thereof.
[0303] The input of the symbolizer receives PCC data and outputs an encoded geometry information bitstream and an attribute information bitstream after encoding.
[0304] The data input unit receives geometry data and texture data. The data input unit receives parameter setting values related to encoding.
[0305] The coordinate system conversion unit sets the coordinate system related to the position of the points in the geometry data as a system adapted to encoding.
[0306] The geometry information conversion quantization processing unit converts and quantizes the geometry data.
[0307] The space division unit divides the point cloud data into a spatial structure adapted to encoding.
[0308] When the geometry information encoding unit has prediction-based coding as the geometry coding type, the prediction tree generation unit generates a prediction tree, and based on the prediction tree generated by the prediction determination unit, performs an RDO (Rate Distortion Optimization) process to select the optimal prediction mode. An optimal geometry prediction value can be generated according to the optimal prediction mode.
[0309] The geometry information encoding unit performs octree-based geometry coding by the octree generation unit or trisoup-based geometry coding by the trisoup generation unit.
[0310] The geometry position reconstruction unit restores the encoded geometry data and provides it for texture coding.
[0311] The residual value from the predicted value in the geometry information entropy encoding unit is entropy-coded to form a geometry information bitstream.
[0312] The detailed operation of the prediction tree generation unit is as follows.
[0313] When the prediction tree generation unit determines that the point has a laser angle, it receives the candidate angle value (origin_laser_angle) corresponding to the origin of the origin and the direction (origin_direction) of the point corresponding to the origin. Based on the received values, it selects the points to be used as the origin within the slice according to origin_laser_angle and origin_direction. The origin value is transmitted to the decoder through signal information.
[0314] The prediction tree generation unit receives a point alignment method and aligns the points according to the alignment method. The point alignment methods include Morton code, radius, azimuth, elevation, sensor ID criteria, laser angle, or the order of captured time, etc. In the case of the laser angle, the order of the points is determined based on the selected origin. When aligned by the laser angle, if they belong to the same laser angle or the same laser angle group, the point order is determined based on the radius or the same radius. If the radius values are the same, the point order is determined based on the elevation. The applied alignment method is transmitted to the decoder through signal information.
[0315] A method for generating a prediction tree is input to the prediction tree generation unit, and the prediction tree is generated according to the input method. The tree generation methods include the Fast prediction tree generation scheme based on the aligned order, the prediction tree generation scheme based on the distance criterion, the prediction tree generation scheme based on the angle (angular), etc. It can be selected according to the content characteristics and service type. The applied prediction tree generation method is transmitted to the decoder through signal information.
[0316] A maximum distance value is input to the prediction tree generation unit. When using the prediction point list, search for adjacent prediction points for parent node selection, and register them as child nodes only when the distance to the searched points is smaller than the maximum distance value. The maximum distance value is either input or automatically set by content analysis.
[0317] The geometry encoder uses the additional attribute data in FIG. 15 via, for example, a prediction tree generation unit, selects an origin position as shown in FIGS. 17 and 19, aligns points based on the laser angle, and generates a fast prediction tree from the points based on the laser angle group. A prediction coding is performed by quickly setting the parent-child relationship from the fast prediction tree. To encode the current geometry data at present, prediction geometry data for the current geometry data is calculated by the fast prediction tree. Residual data between the current geometry data (original) and the prediction geometry data is generated, and a geometry bit stream including the residual data is generated.
[0318] The attribute information encoding unit encodes using the restored geometry data of the characteristic data. Information related to the attribute information encoding can be transmitted to the decoder as signaling information.
[0319] FIG. 21 shows a transmission method / apparatus (point cloud data transmission method / apparatus) according to an embodiment, and the configuration (encoding process) of a point cloud data encoder.
[0320] The prediction geometry coding in FIG. 21 is an alternative to the octree-based method. The prediction coding technique according to the embodiment supports low latency and provides low-complexity decoding.
[0321] The prediction structure is applied to, for example, content corresponding to Category 3. A prediction structure for point cloud data is generated to generate a prediction tree. The points of the point cloud data correspond to the vertices of the tree. Each vertex can be predicted from above (parent) within the tree. The prediction geometric coding according to the embodiment performs prediction geometric coding using the tree structure. A tree structure having parent / child relationships between points is generated. The prediction modes include No prediction, Delta prediction (i.e., p0), Linear prediction (i.e., 2p0 - p1), Parallelogram predictor (i.e., 2p + p1 - p2), etc. Here, p0, p1, and p2 represent the parent, grandparent, and great-grandparent points of the current point. The prediction mode is selected based on the RDO method. The mode corresponding to the case where the residual according to the prediction mode is the smallest is selected, and the used prediction mode (predictor) is transmitted by signaling information.
[0322] FIG. 22 shows a point cloud data receiving device according to an embodiment.
[0323] The receiving device 10004 in FIG. 1, the receiver 10005, the point cloud video decoder 10006, the transmitting - decoding - rendering 20002 - 20003 - 20004 in FIG. 2, the decoders in FIGS. 10 and 11, the receiving device in FIG. 13, the device in FIG. 14, the decoder in FIG. 22, etc. are point cloud data receiving devices according to embodiments. Each component corresponds to hardware, software, a processor, and / or their combination.
[0324] The receiving operation in FIG. 22 corresponds to the transmitting operation in FIG. 21 or performs the reverse process of the transmitting operation.
[0325] The geometric information entropy decoding unit entropy - decodes the geometric data.
[0326] The octree reconstruction unit reconstructs the geometric data based on the octree when octree - based coding is applied to the geometric data.
[0327] The detailed operation of the prediction tree reconstruction unit is as follows.
[0328] The prediction tree reconstruction unit restores the prediction tree generation method, the origin position value, and the point alignment method transmitted, and thereby reconstructs the prediction tree for use in decoding the predicted value of the geometry.
[0329] The geometry decoder grasps the position of the origin by the prediction tree reconstruction unit, grasps the point alignment method, and when fast prediction tree generation is applied, predicts the geometry data by the fast prediction tree and restores the geometry data in combination with the received residual geometry data.
[0330] The geometry position reconstruction unit reconstructs the position of the geometry data and applies it to the feature decoder.
[0331] The geometry information prediction unit generates predicted data of the geometry data.
[0332] When the geometry information inverse quantization processing unit is quantized on the transmission side, it inversely applies quantization to the geometry data based on the quantization parameter.
[0333] When the coordinate system related to the geometry data is converted on the transmission side, the coordinate system inverse conversion unit inversely converts the coordinate system.
[0334] The attribute information decoding unit entropy decodes the residual data of the feature data from the bit stream including the feature data through the attribute residual information entropy decoding unit.
[0335] Based on the attribute decoding method, the attribute information decoding unit decodes the feature data.
[0336] When the residual attribute information inverse quantization processing unit is quantized on the transmission side, it inversely quantizes the residual attribute information based on the quantization parameter. The decoder restores the feature data according to the encoding method on the transmission side.
[0337] In addition, the point cloud data receiving method / apparatus according to the embodiment receives a bitstream in the tree order generated by the encoder and does not perform the point alignment (coordinate transformation) process on the transmission side. The receiving method / apparatus restores the prediction tree from the bitstream in the receiving order. The origin information is used in the process of restoring the prediction tree, and the finally restored position can be transformed into xyz coordinates through the coordinate transformation process.
[0338] FIG. 23 shows a bitstream including point cloud data and parameter information according to the embodiment.
[0339] The point cloud data transmission apparatus according to the embodiment such as FIG. 21 generates a bitstream such as FIG. 23, and the point cloud data receiving apparatus according to the embodiment such as FIG. 22 receives a bitstream such as FIG. 23 and decodes the point cloud data based on the parameter information.
[0340] To add / perform the embodiment, relevant information can be signaled. The signaling information according to the embodiment is used at the transmission end or the reception end, etc. The signaling information according to the embodiment is generated and transmitted by the metadata processing unit (also referred to as a metadata generator, etc.) of the transmission apparatus according to the embodiment and received and acquired by the metadata parser of the reception apparatus. Each operation of the reception apparatus according to the embodiment is performed based on the signaling information. The configuration of the encoded point cloud is as shown in FIG. 23.
[0341] The meanings of the abbreviations are as follows. Each abbreviation may be referred to by other terms within the same meaning range. SPS: Sequence Parameter Set, GPS: Geometry Parameter Set, APS: Attribute Parameter Set, TPS: Tile Parameter Set, Geom: Geometry bitstream = geometry slice header + geometry slice data, Attr: Attribute bitstream = attribute blick header + attribute brick data.
[0342] The prediction tree generation related option information can be added to the SPS or GPS and signaled.
[0343] The prediction tree generation related option information can be added to the TPS or the geometry header for each slice and signaled.
[0344] Tiles or slices are provided so that the point cloud can be processed by region.
[0345] When dividing by region, different neighboring point set generation options can be set for each region to provide a solution with low complexity but slightly reduced result reliability, or conversely, a selection solution with high complexity but high reliability. This setting varies depending on the processing capacity of the receiver.
[0346] Therefore, when the point cloud is divided into tiles, different options can be applied to each tile. When the point cloud is divided into slices, different options can be applied to each slice.
[0347] Figure 24 shows the sequence parameter set according to the embodiment.
[0348] FIG. 24 is a sequence parameter set included in the bitstream of FIG. 23.
[0349] The method / apparatus according to the embodiment includes information related to the prediction tree generation according to the embodiment in the sequence parameter set to provide efficient signaling.
[0350] Prediction geometry tree sorting type (pred_geom_tree_sorting_type): Indicates the sorting method applied when generating a prediction geometry tree in the corresponding sequence. For example, it indicates the sorting method according to each integer value: 0 = no sorting, 1 = sorting in Morton code order, 2 = sorting in radius order, 3 = sorting in azimuth order, 4 = sorting in elevation order, 5 = sorting in sensor ID order, 6 = sorting in captured time order, 7 = sorting in laser angle order
[0351] Prediction geometry tree generation method (pred_geom_tree_build_method): Indicates the method for generating a prediction geometry tree in the corresponding sequence. For example, 0 = Fast prediction tree generation scheme, 1 = distance-based prediction tree generation scheme, 2 = angle-based prediction tree generation scheme
[0352] Profile (profile_idc) indicates the profile that the bitstream conforms to as specified in Appendix A. The bitstream does not contain a profile_idc value that is not a value according to the embodiment. Other values of profile_idc are reserved for future use by ISO / IEC.
[0353] Profile compatibility flag (profile_compatibility_flags): The same profile_compatibility_flags as 1 indicates that the bitstream conforms to the profile represented by the same profile_idc as j.
[0354] The number of SPS characteristic sets (sps_num_attribute_sets) indicates the number of attributes coded in the bitstream. The value of sps_num_attribute_sets ranges from 0 to 63.
[0355] The characteristic dimension (attribute_dimension[i]) indicates the number of components of the i-th attribute.
[0356] The characteristic instance identifier (attribute_instance_id[i]) indicates the instance ID for the i-th attribute.
[0357] Figure 25 shows a geometry parameter set according to an embodiment.
[0358] Figure 25 is the geometry parameter set included in the bitstream of Figure 23.
[0359] The method / apparatus according to the embodiment includes information related to the prediction tree generation according to the embodiment in the geometry parameter set to provide efficient signaling.
[0360] Predicted geometry tree sorting type (pred_geom_tree_sorting_type): Indicates the sorting method to be applied when generating a predicted geometry tree for the corresponding sequence. For example, 0 = no sorting, 1 = sorting in Morton code order, 2 = sorting in radius order, 3 = sorting in azimuth order, 4 = sorting in elevation order, 5 = sorting in sensor ID order, 6 = sorting in captured time order, 7 = sorting in laser angle order
[0361] Predicted geometry tree generation method (pred_geom_tree_build_method): Indicates the method for generating a predicted geometry tree for the corresponding sequence. For example, 0 = Fast prediction tree generation scheme, 1 = distance-based prediction tree generation scheme, 2 = angle-based prediction tree generation scheme
[0362] GPS Geometry Parameter Set ID (gps_geom_parameter_set_id): Provides an identifier for GPS so that it can be referenced by other syntax elements. The value of gps_seq_parameter_set_id ranges from 0 to 15.
[0363] GPS Sequence Parameter Set ID (gps_seq_parameter_set_id): Indicates the sps_seq_parameter_set_id value for the active SPS. The value of gps_seq_parameter_set_id ranges from 0 to 15.
[0364] Figure 26 shows a tile parameter set according to an embodiment.
[0365] Figure 26 is the tile parameter set included in the bitstream of Figure 23.
[0366] The method / apparatus according to the embodiment includes information related to the prediction tree generation according to the embodiment in the tile parameter set to provide efficient signaling.
[0367] Prediction Geometry Tree Sorting Type (pred_geom_tree_sorting_type): Indicates the sorting method to be applied when generating the prediction geometry tree for the corresponding tile. For example, 0 = no sorting, 1 = sort in Morton code order, 2 = sort in radius order, 3 = sort in azimuth order, 4 = sort in elevation order, 5 = sort in sensor ID order, 6 = sort in captured time order, 7 = sort in laser angle order.
[0368] Prediction Geometry Tree Building Method (pred_geom_tree_build_method): Indicates the method for generating the prediction geometry tree for the corresponding tile. For example, 0 = Fast prediction tree generation scheme, 1 = distance-based prediction tree generation scheme, 2 = angle-based prediction tree generation scheme.
[0369] GPS Geometry Parameter Set ID (gps_geom_parameter_set_id): Provides an identifier for GPS so that it can be referenced by other syntax elements. The value of gps_seq_parameter_set_id ranges from 0 to 15.
[0370] GPS Sequence Parameter Set ID (gps_seq_parameter_set_id): Indicates the value of sps_seq_parameter_set_id for the active SPS. The value of gps_seq_parameter_set_id ranges from 0 to 15.
[0371] The number of tiles (num_tiles) indicates the number of tiles signaled for the bitstream. If not present, num_tiles is assumed to be 0.
[0372] Tile Bounding Box Offset X (tile_bounding_box_offset_x[i]) indicates the x offset of the i-th tile in Cartesian coordinates. If not present, the value of tile_bounding_box_offset_x[0] is assumed to be the same as sps_bounding_box_offset_x.
[0373] Tile Bounding Box Offset Y (tile_bounding_box_offset_y[i]) indicates the y offset of the i-th tile in Cartesian coordinates. If not present, the value of tile_bounding_box_offset_y[0] is assumed to be the same as sps_bounding_box_offset_y.
[0374] Tile Bounding Box Offset Z (tile_bounding_box_offset_z[i]) indicates the z offset of the i-th tile in Cartesian coordinates. If not present, the value of tile_bounding_box_offset_z[0] is assumed to be the same as SPS_bounding_box_offset_z.
[0375] FIG. 27 shows a geometry slice header according to an embodiment.
[0376] FIG. 27 is the geometry slice header included in the bitstream of FIG. 23.
[0377] The method / apparatus according to the embodiment provides efficient signaling by including information related to prediction tree generation according to the embodiment in the geometry slice header.
[0378] Prediction origin (pred_origin[i]): Indicates the position value of the origin applied in the corresponding slice.
[0379] Prediction geometry tree sorting type (pred_geom_tree_sorting_type): Indicates the sorting method applied when generating the prediction geometry tree in the corresponding slice. For example, 0 = no sorting, 1 = sorting in Morton code order, 2 = sorting in radius order, 3 = sorting in azimuth order, 4 = sorting in elevation order, 5 = sorting in sensor ID order, 6 = sorting in the order of captured time, 7 = sorting in laser angle order
[0380] Prediction geometry tree generation method (pred_geom_tree_build_method): Indicates the method for generating the prediction geometry tree in the corresponding slice. For example, 0 = Fast prediction tree generation scheme, 1 = distance-based prediction tree generation scheme, 2 = angle-based prediction tree generation scheme
[0381] The GSH geometry parameter set ID (gsh_geometry_parameter_set_id) indicates the gps_geom_parameter_set_id value of the active GPS.
[0382] The GSH tile identifier (gsh_tile_id) indicates the value of the tile ID referenced by the GSH. The value of gsh_tile_id ranges from 0 to XX.
[0383] The GSH slice ID (gsh_slice_id) identifies the slice header that is referenced by other syntax elements. The value of gsh_slice_id ranges from 0 to XX.
[0384] FIG. 28 shows a point cloud data transmission method according to an embodiment.
[0385] Point cloud data transmission devices such as the transmission device 10000 in FIG. 1, the point cloud video encoder 10002, the transmitter 10003, the acquisition - encoding - transmission 20000 - 20001 - 20002 in FIG. 2, the encoder in FIG. 4, the transmission device in FIG. 12, the device in FIG. 14, and the encoder in FIG. 21 encode and transmit point cloud data in the following stages.
[0386] S2800 Stage of encoding point cloud data
[0387] The point cloud data transmission method according to the embodiment includes a stage of encoding point cloud data. The encoding stage according to the embodiment includes operations such as the transmission device 10000 in FIG. 1, the point cloud video acquisition 10001, the point cloud video encoder 10002, the acquisition - encoding 20000 - 20001 in FIG. 2, the encoder in FIG. 4, the transmission device in FIG. 12, the XR device 1430 in FIG. 14, the origin position selection, point alignment, prediction tree generation according to FIGS. 15 and 20, the encoder in FIG. 21, and the bitstream and parameter generation according to FIGS. 23 and 27.
[0388] S2810 Stage of transmitting a bitstream including point cloud data
[0389] The point cloud data transmission method according to the embodiment further includes a step of transmitting a bitstream including the point cloud data. The transmission operations according to the embodiment include operations such as the transmission of the transmission device 10000, the transmitter 10003 in FIG. 1, the transmission 20002 in FIG. 2, the geometric bitstream and the texture bitstream transmission in FIGS. 4, 12, and 14, and the transmission of the bitstream (FIGS. 23 and 27) including the point cloud data encoded according to FIGS. 15 and 20.
[0390] FIG. 29 shows the point cloud data reception method according to the embodiment.
[0391] The point cloud data reception devices such as the reception device 10004, the receiver 10005, the point cloud video decoder 10006 in FIG. 1, the transmission-decoding-rendering 20002-20003-20004 in FIG. 2, the decoders in FIGS. 10 and 11, the reception device in FIG. 13, the device in FIG. 14, and the decoder in FIG. 22 receive and decode the point cloud data in the following steps. The receiving side process is the reverse process of the transmitting side process.
[0392] S2900 Step of receiving a bitstream including point cloud data
[0393] The point cloud data reception method according to the embodiment includes a step of receiving a bitstream including the point cloud data. The receiving steps according to the embodiment include operations such as the reception by the reception device 10004, the receiver 10005 in FIG. 1, the reception by the transmission 20002 in FIG. 2, the reception of the geometric bitstream and the texture bitstream in FIGS. 10 and 11, the reception of the point cloud data encoded according to the reception device in FIG. 13, the XR device 1430 in FIG. 14, FIGS. 15 and 20, etc., the reception of the bitstream by the decoder in FIG. 22, and the reception of the bitstreams in FIGS. 23 and 27.
[0394] S2910 Step of decoding the point cloud data
[0395] The point cloud data receiving method according to the embodiment further includes a step of decoding the point cloud data. The decoding operation according to the embodiment includes operations such as the point cloud video decoder 10006 and renderer 10007 in FIG. 1, the decode-renderer-feedback 20003-20005 in FIG. 2, the decoding in FIGS. 10 and 11, the reception / decoding in FIG. 13, the restoration of the point cloud data encoded by FIGS. 15 and 20, etc., the decoder in FIG. 22, and the restoration of the geometry data and texture data included in the parameter-based bitstream in FIGS. 23 and 37.
[0396] Referring to FIGS. 17 and 18, the method / apparatus according to the embodiment re-orders the geometry data (point positions) for geometry prediction tree coding.
[0397] In the process of capturing points with a laser sensor according to the embodiment, since the laser sensor rotates at a certain angle, the laser angle attribute is possessed by the points (see FIG. 15). The method for using such a laser angle is applied to the determination of the origin for point re-alignment.
[0398] Origin determination according to the embodiment:
[0399] Since the points have a laser angle and the laser angle is used to determine the origin. The origin is selected and determined by the method / apparatus according to the embodiment according to the following settings.
[0400] As shown in FIGS. 17 and 18, the center laser angle may be 90°. It may also be the leftmost point.
[0401] FIGS. 17 and 18 show the selected origins 18001, 08004, 18007 for each slice.
[0402] Re-ordering method based on laser angle
[0403] Referring to FIG. 19, it is used to re-align points instead of the azimuth for which the laser angle is calculated. Points are aligned according to the laser angle value (laser angle range). They may be arranged based on the color which is an attribute of the points. For example, when re-aligning points by laser angle, a result is obtained where the points are aligned in an exemplary order such as yellow points, blue points, orange points, green points.
[0404] Referring to FIG. 1, it includes a stage of encoding point cloud data and a stage of transmitting a bit stream including the point cloud data.
[0405] Referring to FIG. 15, in relation to the laser angle, the stage of encoding point cloud data includes a stage of encoding the geometric data of the point cloud data, and the geometric data is encoded based on the laser angle with respect to the point cloud data.
[0406] Referring to FIGS. 17 and 18, in relation to the origin position and point alignment, in the method / apparatus according to the embodiment, the geometric data of the point cloud data has a laser angle, the laser angle is 90°, and the point of the leftmost coordinate is selected as the origin of the geometric data.
[0407] Referring to FIG. 19, in relation to the alignment based on the laser angle, the geometric data is aligned based on the laser angle.
[0408] Referring to FIG. 20, in the method according to the embodiment related to the prediction tree generation, a prediction tree is generated with a point having the latest laser angle based on the laser angle as the parent, and the root node of the second laser group including a plurality of points is set as the parent node of the root node of the first laser group including a plurality of points. The second laser group has a smaller laser angle value than the first laser group. The latest refers to a point at the first position in the aligned state, or a point having a small laser angle value.
[0409] Referring to FIG. 21, the step of encoding the point cloud data related to the geometry encoding includes the step of encoding the geometry data of the point cloud data. The step of encoding the geometry data includes converting the coordinate system of the geometry data to set an origin based on the laser angle, aligning the geometry data based on the origin, generating a prediction tree based on the aligned geometry data, generating a predicted value of the point cloud data based on the prediction tree, and generating a residual value from the predicted value to generate a geometry bitstream.
[0410] According to the PCC encoding method, PCC decoding method, and signaling method of the embodiment, the following effects can be obtained.
[0411] In the scenario where the lidar equipment captures and stores one frame at a time, the angle mode can be applied. However, when the lidar equipment captures multiple frames and integrates them into one content for 3D map data generation, since data with different central positions of the lidar equipment are mixed, the angular characteristics shown by the data captured by the lidar equipment, that is, the rule between points when changing to the angle (r, Φ, i) is hidden, and thus the application of the angle mode is not more efficient than the compression based on the Cartesian coordinate system.
[0412]
Number
[0413]
Number
[0414]
Number
[0415] Therefore, in order to improve the compression efficiency by using the characteristics captured by a LiDAR (Light Detection and Ranging) device, a solution is also needed to improve the compression efficiency from the regularity of the points in the 3D map data as well as in the content.
[0416] This embodiment supports a method for selecting an origin, an alignment method, and a method for quickly constructing a prediction tree for a prediction tree structure for efficiently geometrically compressing 3D map data captured by a LiDAR device and integrated into one content.
[0417] Thereby, the embodiment can improve the geometric compression efficiency of an encoder / decoder of G-PCC (Geometry-based Point Cloud Compression) for 3D point cloud data compression and provide a point cloud content stream.
[0418] The PCC encoder and / or PCC decoder according to the embodiment provides an efficient prediction tree generation solution, and by considering the influence degree between prediction points, provides an effect of increasing the geometric compression coding / decoding efficiency.
[0419] Therefore, the transmission method / apparatus according to the embodiment can efficiently compress point cloud data, transmit the data, and transmit the signaling information therefor, so that the receiving method / apparatus according to the embodiment can also efficiently decode / restore the point cloud data.
[0420] The embodiment is described from the perspectives of a method and / or an apparatus, and the description of the method and the description of the apparatus can be complementarily applied to each other.
[0421] For the sake of convenience of explanation, each figure has been described separately, but it is also possible to design to embody a new embodiment by combining the embodiments described in each figure. Also, depending on the needs of an ordinary technician, designing a computer-readable recording medium on which a program for executing the previously described embodiments is recorded also falls within the scope of the rights of the embodiments. The apparatus and method according to the embodiments, as described above, are not limited to the application of the described configurations and methods of the embodiments, and the embodiments can also be configured by selectively combining all or part of each embodiment in various deformable ways. Although the preferred embodiments of the embodiments have been shown and described, the embodiments are not limited to the specific embodiments described above, and without departing from the gist of the embodiments claimed in the claims, various modifications can be made by those having ordinary knowledge in the technical field to which the invention pertains, and such modifications should not be individually understood from the technical idea and prospect of the embodiments.
[0422] The various components of the device according to the embodiments are configured by hardware, software, firmware, or a combination thereof. The various components of the embodiments are implemented by one chip, for example, one hardware circuit. In the embodiments, the components according to the embodiments are each implemented by an individual chip. In the embodiments, any of the components of the device according to the embodiments is composed of one or more processors capable of executing one or more programs, and one or more programs include instructions for causing or executing any one or more of the operations / methods according to the embodiments. The executable instructions for performing the method / operation of the device according to the embodiments are stored in a non-transitory CRM or other computer program product configured to be executed by one or more processors, or can be stored in a transitory CRM or other computer program product configured to be executed by one or more processors. Also, the memory according to the embodiments is used as a concept that includes not only volatile memory (such as RAM, etc.), but also non-volatile memory, flash memory, PROM, etc. Also included is being embodied in the form of a carrier wave such as transmission via the Internet. Also, the recording medium read by the processor can be distributed in a computer system connected by a network, and the code read by the processor can be stored and executed in a distributed manner.
[0423] In this specification, " / " and "," are interpreted as "and / or". For example, "A / B" is interpreted as "A and / or B", and "A, B" is interpreted as "A and / or B". Further, "A / B / C" means "any one of A, B, and / or C". Also, "A, B, C" also means "any one of A, B, and / or C". Further, in this document, "or" is interpreted as "and / or". For example, "A or B" means 1) only "A", 2) only "B", or 3) "A and B". In other words, "or" in this specification means "additionally or alternatively".
[0424] Terms such as first and second are used to describe various components of the embodiments. However, the various components according to the embodiments should not be limited in interpretation by the above terms. Such terms are only used to distinguish one component from another. For example, the first user input signal can be referred to as the second user input signal. Similarly, the second user input signal can be referred to as the first user input signal. The use of such terms does not depart from the scope of the various embodiments. Although both the first user input signal and the second user input signal are user input signals, they do not mean the same user input signal unless clearly indicated in the context.
[0425] The terms used for explaining the embodiments are used for explaining specific embodiments and do not limit the embodiments. As used in the description of the embodiments and the claims, unless clearly stated in the context, the singular includes the plural. The expression "and / or" is used in the sense of including all possible combinations between terms. "Including" describes the existence of features, numbers, steps, elements and / or components, and does not mean excluding further features, numbers, steps, elements and / or components. Conditional expressions such as "when ~" and "when ~" used for explaining the embodiments are not interpreted as limited only in optional cases. It is intended that when specific conditions are met, related operations are performed corresponding to the specific conditions or related definitions are interpreted.
[0426] Also, the operations according to the embodiments described in this specification are performed by a transmission / reception device including a memory and / or a processor according to the embodiments. The memory stores a program for processing / controlling the operations according to the embodiments, and the processor controls the various operations described in this specification. The processor is also referred to as a controller, etc. The operations of the embodiments are performed by firmware, software and / or a combination thereof, and the firmware, software and / or a combination thereof are stored in the processor or stored in the memory.
[0427] On one hand, the operations according to the above-described embodiments are performed by a transmission device and / or a reception device according to the embodiments. The transmission / reception device includes a transmission / reception unit that transmits and receives media data, a memory that stores instructions (program codes, algorithms, flowcharts, and / or data) for the processes according to the embodiments, and a processor that controls the operations of the transmission / reception device.
[0428] The processor is also referred to as a controller, for example, and corresponds to hardware, software, and / or a combination thereof. The operations according to the above-described embodiments are performed by the processor. The processor is embodied as an encoder / decoder or the like for the operations of the above-described embodiments.
[0429] As described above, the related content regarding the best mode for implementing the embodiments is explained.
Industrial Applicability
[0430] As described above, the embodiments can be applied in whole or in part to a point cloud data transmission / reception device and system.
[0431] A person skilled in the art can make various changes and modifications to the embodiments within the scope of the embodiments.
[0432] The embodiments include changes / modifications, and the changes / modifications are within the scope of the claims and those identical thereto.
Claims
1. Encoding the geometric data of point cloud data, the step including generating a prediction tree based on the azimuth for the geometric data, encoding the attribute data of the point cloud data, transmitting a bitstream including the point cloud data, and the bitstream includes type information related to the prediction tree, a method for transmitting point cloud data.
2. The method according to claim 1, wherein the geometric data is encoded based on the laser angle for the point cloud data.
3. The geometric data of the point cloud data includes a laser angle, and the point is selected as the origin based on the laser angle of the point being 90 degrees and the position of the point being on the left, the method according to claim 1.
4. The method according to claim 3, wherein the geometric data is aligned based on the laser angle.
5. The method further includes: generating a prediction tree including as a parent a point having a laser angle based on the laser angle; and setting a root node of a second laser group including points as a parent node of a root node of a first laser group including points, wherein the laser angle of the second laser group is smaller than the laser angle of the first laser group, the method according to claim 4.
6. The step of encoding the geometric data includes: setting an origin based on the laser angle by converting the coordinates of the geometric data; aligning the geometric data based on the origin; generating a prediction tree based on the aligned geometric data; generating a predicted value based on the prediction tree; generating a residual from the predicted value; and generating a geometric bitstream, wherein the bitstream includes information for indicating the selected origin, the method according to claim 4.
7. An encoder, encoding the geometric data of point cloud data, and encoding the geometric data includes generating a prediction tree based on the azimuth for the geometric data, An encoder configured to encode attribute data of the point cloud data; A transmitter configured to transmit a bitstream including the point cloud data, comprising: The bitstream includes type information related to the prediction tree, and is a device for transmitting point cloud data. **Claim 8** The geometry data is encoded based on a laser angle with respect to the point cloud data. The geometry data of the point cloud data includes a laser angle. Based on the laser angle of the point being 90 degrees and the position of the point being on the left, the point is selected as the origin. The geometry data is aligned based on the laser angle. The device Generates a prediction tree including a point having a laser angle based on the laser angle as a parent. Is further configured to set a root node of a second laser group including points as a parent node of a root node of a first laser group including points. The laser angle of the second laser group is smaller than the laser angle of the first laser group. The device according to claim 7. **Claim 9** The encoder Sets an origin based on a laser angle by converting coordinates of the geometry data. Aligns the geometry data based on the origin. Generates a prediction tree based on the aligned geometry data. Is further configured to generate a predicted value based on the prediction tree, generate a residual from the predicted value, and generate a geometry bitstream. The bitstream includes information for indicating the selected origin. The device according to claim 8. **Claim 10** Decoding the geometry data of the point cloud data in the bitstream; Decoding the attribute data of the point cloud data, including: The step of decoding the geometry data includes generating a prediction tree based on an azimuth with respect to the geometry data. The bitstream includes type information related to the prediction tree, and is a method for receiving point cloud data. **Claim 11** Based on the laser angle of the point being 90 degrees and the position of the point being on the left, the point is selected as the origin, The method according to claim 10, wherein the bitstream includes information for indicating the selected origin.
12. The geometry data is aligned based on the laser angle, The method is, Generating a prediction tree including a point having a laser angle based on the laser angle as a parent, Setting a root node of a second laser group including points as a parent node of a root node of a first laser group including points, further comprising, The method according to claim 11, wherein the laser angle of the second laser group is smaller than the laser angle of the first laser group.
13. The step of decrypting the geometry data is, Setting an origin based on the laser angle by converting the coordinates of the geometry data, Aligning the geometry data based on the origin, Generating a prediction tree based on the aligned geometry data, Generating a predicted value based on the prediction tree, Adding the predicted value and the residual, Reconstructing the geometry data, the method according to claim 12.
14. A receiving unit configured to receive a bitstream including point cloud data, A decoder configured to decrypt the geometry data of the point cloud data and decrypt the attribute data of the point cloud data, Decrypting the geometry data includes generating a prediction tree based on the azimuth for the geometry data, The apparatus for receiving point cloud data, wherein the bitstream includes type information related to the prediction tree.
15. The geometry data is decrypted based on the laser angle for the point cloud data, The apparatus according to claim 14, wherein the bitstream includes information for indicating a selected origin.
Citation Information
Patent Citations
A method and apparatus for encoding / decoding the colors of a colored point cloud whose geometry is represented by an octree-based structure
US20200143568A1
Method and apparatus for point cloud compression
US20200394822A1
Coding of laser angles for angular and azimuthal modes in geometry-based point cloud compression
US20210326734A1
Method and apparatus for interframe point cloud attribute coding
WO2020197966A1
High level syntax for laser rotation in geometry point cloud compression (g-PCC)
WO2022076124A1