Point cloud data transmitting device, point cloud data transmitting method, point cloud data receiving device, and point cloud data receiving method.
By extracting and encoding the geometric information and feature information of point cloud data, and using appropriate prediction modes and encoding methods, the challenges of latency, complexity and compression performance in point cloud data transmission are solved, and efficient and low-latency point cloud data transmission is achieved.
Patent Information
- Application Number
- JP2022520572
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-10-03
- Filing Date
- 2020-10-05
- Publication Date
- 2025-05-08
- Estimated Expiration
- 2040-10-05
AI Technical Summary
The prior art is difficult to efficiently process and transmit large amounts of point cloud data, especially in solving the challenges of latency, encoding/decoding complexity, and compression performance.
By extracting geometric information and feature information of point cloud data, appropriate prediction modes and coding methods are adopted, including generating prediction candidates, calculating scores, setting prediction modes and encoding residual feature values, to improve the transmission efficiency of point cloud data.
High-quality point cloud data transmission is realized, reducing latency and encoding/decoding complexity, while improving the compression performance and transmission efficiency of point cloud data.
Smart Images

Figure 0007673057000035 
Figure 0007673057000036 
Figure 0007673057000037
Abstract
Description
[Technical field]
[0001] The embodiments relate to a method and apparatus for processing Point Cloud Content. [Background technology]
[0002] Point cloud content is content represented by a point cloud, which is a collection of points belonging to a coordinate system that represents a three-dimensional space. Point cloud content can represent three-dimensional media and is used to provide various services such as VR (Virtual Reality), AR (Augmented Reality), MR (Mixed Reality, XR (Extended Reality)) and autonomous driving services. However, to represent point cloud content, tens of thousands to hundreds of thousands of point data are required. Therefore, a method for efficiently processing a huge amount of point data is required. Summary of the Invention [Problem to be solved by the invention]
[0003] The technical objective of the embodiments is to provide a point cloud data transmitting device, transmitting method, point cloud data receiving device and receiving method for efficiently transmitting and receiving point clouds in order to solve the above-mentioned problems.
[0004] A technical problem according to the embodiments is to provide a point cloud data transmitting device and transmitting method, and a point cloud data receiving device and receiving method for solving latency and encoding / decoding complexity.
[0005] A technical objective of the present invention is to provide a point cloud data transmitting device and method, and a point cloud data receiving device and method, which improve the compression performance of point clouds by improving an encoding technique of attribute information of geometry-based point cloud compression (G-PCC).
[0006] A technical objective of the embodiment is to provide a point cloud data transmitting device and transmitting method, and a point cloud data receiving device and receiving method for improving compression efficiency while supporting parallel processing of G-PCC characteristic information.
[0007] The technical objective of the embodiment is to provide a point cloud data transmitting device, transmitting method, and point cloud data receiving device and receiving method that reduce the size of the feature bitstream and improve the efficiency of feature compression by proposing a method for selecting a suitable predictor candidate from among predictor candidates when encoding G-PCC feature information.
[0008] However, the scope of the invention is not limited to the above-mentioned technical problems, and may be extended to other technical problems that a person skilled in the art can derive based on all the contents described herein. [Means for solving the problem]
[0009] To achieve the above objectives and other advantages, a point cloud data transmission method according to an embodiment includes the steps of acquiring point cloud data, encoding geometry information including positions of points of the point cloud data, encoding attribute information including attribute values of the points of the point cloud data based on the geometry information, and transmitting the encoded geometry information, the encoded attribute information, and signaling information.
[0010] In one embodiment, the step of encoding the feature information includes the steps of generating predictor candidates based on nearest neighbors of the point to be encoded, obtaining a score for each predictor candidate, setting a prediction mode corresponding to the predictor candidate having the lowest score as a prediction mode for the point, and encoding the prediction mode of the point and residual feature values obtained based on the prediction mode of the point.
[0011] In one embodiment, the maximum difference between the feature values of the nearest neighboring points is equal to or greater than a threshold value.
[0012] In one embodiment, the step of encoding the characteristic information includes setting one of prediction modes 1 to 3 as a prediction mode for a point, and in prediction mode 1, the predicted characteristic value of the point is determined based on the characteristic value of the nearest adjacent point among the nearest adjacent points, in prediction mode 2, the predicted characteristic value of the point is determined based on the characteristic value of the second nearest adjacent point among the nearest adjacent points, and in prediction mode 3, the predicted characteristic value of the point is determined based on the characteristic value of the third nearest adjacent point among the nearest adjacent points.
[0013] In one embodiment, the step of encoding the feature information includes obtaining a score for each predictor candidate based on the feature value of each residual obtained by applying each of prediction modes 1 to 3.
[0014] The step of encoding the feature information includes a step of setting a predetermined prediction mode as a prediction mode of a point when a maximum difference value between feature values of nearest neighboring points of the point is less than a threshold value, and a step of encoding residual feature values obtained based on the set prediction mode.
[0015] In one embodiment, the predetermined prediction mode is prediction mode 0, in which the predicted feature value of a point is determined by a distance-based weighted average of its nearest neighboring points.
[0016] A point cloud data transmitting device according to an embodiment includes an acquisition unit for acquiring point cloud data, a geometry encoder for encoding geometry information including positions of points of the point cloud data, a feature encoder for encoding feature information including feature values of points of the point cloud data based on the geometry information, and a transmission unit for transmitting the encoded geometry information, the encoded feature information, and signaling information.
[0017] In one embodiment, the feature encoder generates predictor candidates based on the nearest neighbors of the point to be encoded, obtains a score for each predictor candidate, sets the prediction mode corresponding to the predictor candidate with the lowest score as the prediction mode of the point, and encodes the prediction mode of the point and the residual feature value obtained based on the prediction modes.
[0018] In one embodiment, the maximum difference between the feature values of a point's nearest neighbors is equal to or greater than the threshold value.
[0019] In one embodiment, the feature encoder sets one of prediction modes 1 to 3 as the prediction mode for a point, and in prediction mode 1, the predicted feature value of the point is determined based on the feature value of the nearest adjacent point among the nearest adjacent points, in prediction mode 2, the predicted feature value of the point is determined based on the feature value of the second nearest adjacent point among the nearest adjacent points, and in prediction mode 3, the predicted feature value of the point is determined based on the feature value of the third nearest adjacent point among the nearest adjacent points.
[0020] In one embodiment, the feature encoder obtains a score for each predictor candidate based on the feature values of the residuals obtained by applying prediction mode 1 to prediction mode 3, respectively.
[0021] In one embodiment, the feature encoder sets a predetermined prediction mode as the prediction mode of a point when the maximum difference value between the feature values of the nearest neighboring points of the point is smaller than a threshold value, and encodes the residual feature values obtained based on the set prediction mode.
[0022] In one embodiment, the predetermined prediction mode is prediction mode 0, in which the predicted feature value of a point is determined by a distance-based weighted average of its nearest neighboring points.
[0023] A point cloud data receiving device according to an embodiment includes a receiving unit that receives geometry information, attribute information and signaling information, a geometry decoder that decodes the geometry information based on the signaling information to restore the positions of points, a attribute decoder that decodes the attribute information based on the signaling information and the geometry information to restore the attribute values of points, and a renderer that renders the restored point cloud data based on the positions and attribute values of points.
[0024] In one embodiment, the feature decoder obtains a prediction mode of the point being decoded, obtains a predicted feature value of the point based on the prediction mode of the point, and restores the feature value of the point based on the predicted feature value of the point and the residual feature value of the point included in the feature information.
[0025] In one embodiment, the feature decoder derives a prediction mode for a point from the feature information when the maximum difference value between the feature values of the point's nearest neighbors is equal to or greater than a threshold value, and the obtained prediction mode is one of prediction modes 1 to 3.
[0026] In one embodiment, in prediction mode 1, the predicted feature value of a point is determined based on the feature value of its nearest neighbor among the nearest neighboring points, in prediction mode 2, the predicted feature value of a point is determined based on the feature value of its second nearest neighboring point among the nearest neighboring points, and in prediction mode 3, the predicted feature value of a point is determined based on the feature value of its third nearest neighboring point among the nearest neighboring points.
[0027] In one embodiment, if the maximum difference value between the feature values of a point's nearest neighbors is less than a threshold value, the prediction mode of the point is the predetermined prediction mode.
[0028] In one embodiment, the predetermined prediction mode is prediction mode 0, in which the predicted feature value of a point is determined by a distance-based weighted average of its nearest neighboring points. Effect of the Invention
[0029] The point cloud data transmitting method and transmitting device, point cloud data receiving method and receiving device according to the embodiments provide a high-quality point cloud service.
[0030] The point cloud data transmitting method, transmitting device, point cloud data receiving method, and receiving device according to the embodiments achieve various video codec methods.
[0031] The point cloud data transmitting method and transmitting device, point cloud data receiving method and receiving device according to the embodiments provide general-purpose point cloud content such as an autonomous driving service.
[0032] The point cloud data transmitting method, transmitting device, point cloud data receiving method, and receiving device according to the embodiments provide improved parallel processing and scalability by performing spatially adaptive partitioning of point cloud data for independent encoding and decoding of the point cloud data.
[0033] The point cloud data transmitting method, transmitting device, point cloud data receiving method, and receiving device according to the embodiments can improve the performance of encoding and decoding of point cloud by spatially dividing point cloud data into tile and / or slice units and encoding and decoding the data, and by signaling the data necessary for this purpose.
[0034] The point cloud data transmitting method, transmitting device, point cloud data receiving method, and receiving device according to the embodiments can reduce the size of the bitstream of the characteristic information and improve the compression efficiency of the characteristic information by changing the method of selecting the most suitable predictor from among predictor candidates when encoding characteristic information.
[0035] According to an embodiment of the point cloud data transmitting method, transmitting device, point cloud data receiving method, and receiving device, when the maximum difference value between the feature values of adjacent points registered in the predictor of the corresponding point is equal to or greater than a predetermined threshold value, a prediction mode that calculates a predicted feature value by a weighted average is not set as the prediction mode of the corresponding point, thereby reducing the size of the feature bitstream including the residual feature values, thereby improving the compression efficiency of the features. [Brief description of the drawings]
[0036] The accompanying drawings are provided for a clearer understanding of the embodiments and together with the description relating to the embodiments illustrate the embodiments.
[0037]
Figure 1
[0038]
Figure 2
[0039]
Figure 3
[0040]
Figure 4
[0041]
Figure 5
[0042]
Figure 6
[0043]
Figure 7
[0044]
Figure 8
[0045]
Figure 9
[0046]
Figure 10
[0047]
Figure 11
[0048]
Figure 12
[0049]
Figure 13
[0050]
Figure 14
[0051]
Figure 15
[0052]
Figure 16
[0053]
Figure 17
[0054]
Figure 18
[0055]
Figure 19
[0056]
Figure 20
[0057]
Figure 21
[0058]
Figure 22
[0059]
Figure 23
[0060]
Figure 24
[0061]
Figure 25
[0062]
Figure 26
[0063]
Figure 27
[0064]
Figure 28
[0065]
Figure 29
[0066]
Figure 30
[0067]
Figure 31
[0068]
Figure 32
[0069]
Figure 33
[0070]
Figure 34
[0071]
Figure 35
[0072]
Figure 36
[0073] Hereinafter, the embodiments described in this specification will be described in detail with reference to the accompanying drawings, and the same or similar components will be given the same reference numerals regardless of the drawing numerals, and duplicated explanations will be omitted. The following embodiments are for the purpose of embodying the present invention, and do not limit or restrict the scope of the present invention. Anything that a person skilled in the art to which the present invention pertains can easily infer from the detailed description and embodiments of the present invention is deemed to fall within the scope of the present invention.
[0074] The detailed description of this specification should not be construed as limiting in all respects, but should be considered as illustrative. The scope of the present invention should be determined based on a reasonable interpretation of the appended claims, and all modifications within the equivalent scope of the present invention are included in the scope of the present invention.
[0075] Preferred embodiments will be described in detail with reference to the accompanying drawings. The following detailed description with reference to the accompanying drawings is intended to describe preferred embodiments rather than showing only embodiments that can be implemented by the embodiments. The following detailed description includes details to provide a thorough understanding of the embodiments. However, it is apparent to those skilled in the art that the embodiments can be implemented without such details. Most of the terms used in the embodiments are common terms widely used in the relevant field, but some are arbitrarily selected by the applicant, and their meanings will be described in detail below as necessary. Therefore, the embodiments should be understood based on the intended meaning of the terms, rather than the simple names and meanings of the terms. In addition, the following drawings and detailed description should not be interpreted as being limited to the specifically described embodiments, but should be interpreted as including equivalents or alternatives to the embodiments described in the drawings and detailed description.
[0076] FIG. 1 is a diagram illustrating an example of a point cloud content providing system according to an embodiment.
[0077] The point cloud content providing system shown in Fig. 1 includes a transmission device 10000 and a reception device 10004. The transmission device 10000 and the reception device 10004 are capable of wired or wireless communication to transmit and receive point cloud data.
[0078] The transmitting device 10000 according to the embodiment secures, processes, and transmits a point cloud video (or point cloud content). In the embodiment, the transmitting device 10000 includes a fixed station, a base transceiver system (BTS), a network, an AI (Artificial Intelligence) device and / or system, a robot, an AR / VR / XR device and / or server, etc. In the embodiment, the transmitting device 10000 includes a device that communicates with a base station and / or other wireless devices using a wireless connection technology (e.g., 5G NR (New RAT), LTE (Long Term Evolution)), a robot, a vehicle, an AR / VR / XR device, a mobile device, a home appliance, an IoT (Internet of Things) device, an AI device / server, etc.
[0079] The transmitting device 10000 according to the embodiment includes a Point Cloud Video Acquisition unit 10001, a Point Cloud Video Encoder 10002, and / or a Transmitter (or a communication module) 10003.
[0080] The point cloud video acquisition unit 10001 according to the embodiment acquires a point cloud video through a process such as capture, synthesis, or generation. The point cloud video is a point cloud content represented by a point cloud, which is a collection of points located in a three-dimensional space, and is also called point cloud video data. The point cloud video according to the embodiment includes one or more frames. One frame represents a still image / picture. Therefore, the point cloud video includes a point cloud image / frame / picture, and is called any of a point cloud image, a frame, and a picture.
[0081] The point cloud video encoder 10002 according to the embodiment encodes the secured point cloud video data. The point cloud video encoder 10002 encodes the point cloud video data based on point cloud compression coding. The point cloud compression coding according to the embodiment includes Geometry-based Point Cloud Compression (G-PCC) coding and / or Video based Point Cloud Compression (V-PCC) coding or next-generation coding. Note that the point cloud compression coding according to the embodiment is not limited to the above-mentioned embodiment. The point cloud video encoder 10002 can output a bitstream including encoded point cloud video data. The bitstream includes not only encoded point cloud video data but also signaling information related to the encoding of the point cloud video data.
[0082] The transmitter 10003 according to the embodiment transmits a bitstream including encoded point cloud video data. The bitstream according to the embodiment is encapsulated into a file or a segment (e.g., a streaming segment) and transmitted via various networks such as a broadcast network and / or a broadband network. Although not shown, the transmitter 10000 includes an encapsulation unit (or an encapsulation module) that performs an encapsulation operation. In the embodiment, the encapsulation unit is included in the transmitter 10003. In the embodiment, the file or segment is transmitted to the receiving device 10004 via a network or stored in a digital storage medium (e.g., USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.). The transmitter 10003 according to the embodiment can perform wired or wireless communication with the receiving device 10004 (or a receiver 10005) via a network such as 4G, 5G, or 6G. The transmitter 10003 can also perform a required data processing operation via a network system (e.g., a communication network system such as 4G, 5G, or 6G). The transmitting device 10000 can also transmit encapsulated data by an on-demand method.
[0083] The receiving device 10004 according to the embodiment includes a receiver 10005, a point cloud video decoder 10006, and / or a renderer 10007. In the embodiment, the receiving device 10004 includes a device, a robot, a vehicle, an AR / VR / XR device, a mobile device, a home appliance, an IoT (Internet of Things) device, an AI device / server, etc., that communicates with a base station and / or other wireless devices using a wireless connection technology (e.g., 5G NR (New RAT), LTE (Long Term Evolution)).
[0084] The receiver 10005 according to the embodiment receives a bitstream including point cloud video data or a file / segment in which the bitstream is encapsulated from a network or storage medium. The receiver 10005 performs a data processing operation required by a network system (e.g., a communication network system such as 4G, 5G, 6G, etc.). The receiver 10005 according to the embodiment decapsulates the received file / segment and outputs a bitstream. In the embodiment, the receiver 10005 also includes a decapsulation unit (or a decapsulation module) for performing the decapsulation operation. The decapsulation unit is also embodied as an element (or component) separate from the receiver 10005.
[0085] The point cloud video decoder 10006 decodes a bitstream including point cloud video data. The point cloud video decoder 10006 can decode the point cloud video data in the manner in which it was encoded (e.g., the reverse process of the operation of the point cloud video encoder 10002). Thus, the point cloud video decoder 10006 can decode the point cloud video data by performing point cloud reconstruction coding, which is the reverse process of point cloud compression. The point cloud reconstruction coding includes G-PCC coding.
[0086] The renderer 10007 renders the decoded point cloud video data. The renderer 10007 renders not only the point cloud video data but also audio data to output point cloud content. In an embodiment, the renderer 10007 includes a display for displaying the point cloud content. In an embodiment, the display is not included in the renderer 10007 and is embodied by a separate device or component.
[0087] In the drawing, the dotted arrow indicates a transmission path of feedback information obtained by the receiving device 10004. The feedback information is information for reflecting interaction with a user consuming point cloud content, and includes user information (e.g., head orientation information, viewport information, etc.). In particular, when the point cloud content is for a service requiring interaction with a user (e.g., an autonomous driving service, etc.), the feedback information may be transmitted to a content transmitting side (e.g., the transmitting device 10000) and / or a service provider. In an embodiment, the feedback information may be used not only by the transmitting device 10000 but also by the receiving device 10004, or may not be provided.
[0088] According to an embodiment, the head orientation information is information regarding the position, direction, angle, movement, etc. of the user's head. According to an embodiment, the receiving device 10004 calculates viewport information based on the head orientation information. The viewport information is information regarding the area of the point cloud video that the user is looking at. The viewpoint is the point at which the user is looking at the point cloud video, and means the center of the viewport area. That is, the viewport is an area centered on the viewpoint, and the size, shape, etc. of the area are determined by the FOV (Field Of View). Therefore, the receiving device 10004 can extract viewport information based on the vertical or horizontal FOV supported by the device in addition to the head orientation information. In addition, the receiving device 10004 performs gaze analysis, etc. to confirm the user's point cloud consumption manner, the point cloud video area that the user is gazing at, the gaze time, etc. In an embodiment, the receiving device 10004 can transmit feedback information including the result of the gaze analysis to the transmitting device 10000. According to an embodiment, the feedback information is obtained during the rendering and / or display process. In some embodiments, the feedback information is obtained by one or more sensors included in the receiving device 10004. In some embodiments, the feedback information is obtained by the renderer 10007 or another external element (or device, component, etc.).
[0089] The dotted lines shown in FIG. 1 indicate the transmission process of the feedback information secured by the renderer 10007. The point cloud content providing system processes (encodes / decodes) point cloud data based on the feedback information. Thus, the point cloud video data decoder 10006 can perform a decoding operation based on the feedback information. Also, the receiving device 10004 can transmit the feedback information to the transmitting device 10000. The transmitting device 10000 (or the point cloud video data encoder 10002) can perform an encoding operation based on the feedback information. Thus, the point cloud content providing system does not process (encodes / decodes) all point cloud data, but efficiently processes necessary data (e.g., point cloud data corresponding to the user's head position) based on the feedback information to provide point cloud content to the user.
[0090] In an embodiment, the sending device 10000 may be referred to as an encoder, a sending device, a transmitter, etc., and the receiving device 10004 may be referred to as a decoder, a receiving device, a receiver, etc.
[0091] The point cloud data processed (processed in a series of steps of acquisition / encoding / transmission / decoding / rendering) in the point cloud content providing system of FIG. 1 according to the embodiment is also called point cloud content data or point cloud video data. In the embodiment, the point cloud content data can be used as a concept including metadata or signaling information related to the point cloud data.
[0092] The elements of the point cloud content providing system illustrated in FIG. 1 may be implemented in hardware, software, a processor, and / or a combination thereof.
[0093] FIG. 2 is a block diagram illustrating operations for providing point cloud content according to an embodiment.
[0094] Figure 2 is a block diagram showing the operation of the point cloud content providing system described in Figure 1. As described above, the point cloud content providing system processes point cloud data based on point cloud compression coding (eg, G-PCC).
[0095] In a point cloud content providing system (e.g., a point cloud transmitting device 10000 or a point cloud video acquiring unit 10001) according to an embodiment, a point cloud video is acquired (20000). The point cloud video is represented by a point cloud belonging to a coordinate system representing a three-dimensional space. The point cloud video according to an embodiment includes a Ply (Polygon File format or the Stanford Triangle format) file. When the point cloud video has one or more frames, the acquired point cloud video includes one or more Ply files. A Ply file includes point cloud data such as the geometry and / or attributes of a point. The geometry includes the position of the point. The position of each point is represented by parameters (e.g., values of each of the X-axis, Y-axis, and Z-axis) indicating a three-dimensional coordinate system (e.g., a coordinate system consisting of XYZ axes). The attributes include the attributes of the point (e.g., texture information of each point, hue (YCbCr or RGB), reflectance (r), transparency, etc.). One point has one or more attributes (or attributes). For example, a point may have one attribute of hue, or it may have two attributes of hue and reflectance.
[0096] In embodiments, geometry may also be referred to as position, geometry information, geometry data, etc., and features may also be referred to as features, feature information, feature data, etc.
[0097] In addition, a point cloud content providing system (eg, a point cloud transmitting device 10000 or a point cloud video acquiring unit 10001) can obtain point cloud data from information related to the point cloud video acquiring process (eg, depth information, color information, etc.).
[0098] A point cloud content providing system (e.g., a transmitting device 10000 or a point cloud video encoder 10002) according to an embodiment encodes point cloud data (20001). The point cloud content providing system encodes point cloud data based on point cloud compression coding. As described above, point cloud data includes geometry and attributes of points. Thus, the point cloud content providing system can perform geometry encoding to encode the geometry and output a geometry bitstream. The point cloud content providing system can perform attribute encoding to encode the attributes and output an attribute bitstream. In an embodiment, the point cloud content providing system can perform attribute encoding based on the geometry encoding. The geometry bitstream and the attribute bitstream according to the embodiment are multiplexed and output as one bitstream. The bitstream according to the embodiment further includes signaling information related to the geometry encoding and the attribute encoding.
[0099] A point cloud content providing system (e.g., transmitting device 10000 or transmitter 10003) according to an embodiment transmits encoded point cloud data (20002). As described in FIG. 1, the encoded point cloud data is represented by a geometry bitstream and a feature bitstream. The encoded point cloud data is transmitted in the form of a bitstream together with signaling information related to the encoding of the point cloud data (e.g., signaling information related to geometry encoding and feature encoding). The point cloud content providing system also encapsulates the bitstream for transmitting the encoded point cloud data and transmits it in the form of a file or segment.
[0100] A point cloud content providing system (e.g., receiving device 10004 or receiver 10005) according to an embodiment receives a bitstream including encoded point cloud data, and the point cloud content providing system (e.g., receiving device 10004 or receiver 10005) can demultiplex the bitstream.
[0101] The point cloud content providing system (e.g., the receiving device 10004 or the point cloud video decoder 10005) decodes the encoded point cloud data (e.g., geometry bitstream, attribute bitstream) transmitted in a bitstream. The point cloud content providing system (e.g., the receiving device 10004 or the point cloud video decoder 10005) decodes the point cloud video data based on signaling information related to the encoding of the point cloud video data included in the bitstream. The point cloud content providing system (e.g., the receiving device 10004 or the point cloud video decoder 10005) decodes the geometry bitstream to restore the position (geometry) of the point. The point cloud content providing system decodes the attribute bitstream based on the restored geometry to restore the attribute of the point. The point cloud content providing system (e.g., the receiving device 10004 or the point cloud video decoder 10005) restores the point cloud video based on the position according to the restored geometry and the decoded attribute.
[0102] A point cloud content providing system (e.g., receiving device 10004 or renderer 10007) according to an embodiment renders the decoded point cloud data (20004). The point cloud content providing system (e.g., receiving device 10004 or renderer 10007) renders the geometry and characteristics decoded during the decoding process using various rendering methods. Points of the point cloud content are rendered as a fixed point with a certain thickness, a cube with a predetermined minimum size centered on the position of the fixed point, or a circle centered on the position of the fixed point. All or a part of the region of the rendered point cloud content is provided to a user via a display (e.g., a VR / AR display, a general display, etc.).
[0103] The point cloud content providing system according to the embodiment (e.g., receiving device 10004) can obtain feedback information (20005). The point cloud content providing system encodes and / or decodes point cloud data based on the feedback information. The feedback information and the operation of the point cloud content providing system according to the embodiment are the same as the feedback information and the operation described in FIG. 1, so a detailed description will be omitted.
[0104] FIG. 3 illustrates an example of a point cloud video capture process according to an embodiment.
[0105] FIG. 3 illustrates an example of a point cloud video capture process of the point cloud content providing system described in FIG. 1 and FIG.
[0106] Point cloud content includes point cloud videos (images and / or videos) showing objects and / or environments located in various 3D spaces (e.g., a 3D space showing a real environment, a 3D space showing a virtual environment, etc.). Thus, the point cloud content providing system according to the embodiment captures point cloud videos using one or more cameras (e.g., an infrared camera capable of obtaining depth information, an RGB camera capable of extracting color information corresponding to the depth information, etc.), projectors (e.g., an infrared pattern projector for obtaining depth information, etc.), LiDAR, etc. to generate point cloud content. The point cloud content providing system according to the embodiment extracts a geometric form composed of points in a 3D space from the depth information, and extracts characteristics of each point from the color information to obtain point cloud data. The images and / or videos according to the embodiment are captured based on either an inward-facing approach or an outward-facing approach.
[0107] The inward-looking method is shown on the left side of Figure 3. The inward-looking method is a method in which one or more cameras (or camera sensors) positioned around a central object capture the central object. The inward-looking method is used to generate point cloud content that provides the user with a 360-degree image of the core object (e.g., VR / AR content that provides the user with a 360-degree image of an object (e.g., a core object such as a character, player, item, actor, etc.)).
[0108] The outward-looking approach is shown on the right side of Figure 3. In the outward-looking approach, one or more cameras (or camera sensors) positioned around a central object capture the environment of the central object that is not the central object. The outward-looking approach is used to generate point cloud content to provide the surrounding environment from a user's perspective (e.g., content showing the external environment provided to a user of an autonomous vehicle).
[0109] As shown in FIG. 3, the point cloud content is generated based on the capture operation of one or more cameras. In this case, since the coordinate systems of the respective cameras are different, the point cloud content providing system performs calibration of one or more cameras to set a global coordinate system before the capture operation. The point cloud content providing system generates the point cloud content by combining the image and / or video captured by the above-mentioned capture method with an arbitrary image and / or video. When generating point cloud content representing a virtual space, the point cloud content providing system does not perform the capture operation described in FIG. 3. The point cloud content providing system according to the embodiment can also perform post-processing on the captured image and / or video. That is, the point cloud content providing system can remove undesired areas (e.g., background) or can recognize a space where the captured images and / or videos are connected and fill a spatial hole if one exists.
[0110] The point cloud content providing system can also generate one point cloud content by performing coordinate system transformation on the points of the point cloud video acquired from each camera. The point cloud content providing system performs coordinate system transformation of the points based on the position coordinates of each camera. In this way, the point cloud content providing system can generate content showing a wide range or generate point cloud content with a high point density.
[0111] FIG. 4 is a diagram illustrating an example of a point cloud video encoder according to an embodiment.
[0112] 4 shows an example of the point cloud video encoder 10002 of FIG. 1. The point cloud video encoder performs encoding by reconstructing point cloud data (e.g., point positions and / or characteristics) to adjust the quality of the point cloud content (e.g., lossless, lossy, near-lossless) according to the network condition or application. If the overall size of the point cloud content is large (e.g., point cloud content of 60 Gbps for 30 fps), the point cloud content providing system cannot stream the corresponding content in real time. Therefore, the point cloud content providing system can reconstruct the point cloud content based on the maximum target bit rate to provide it according to the network environment.
[0113] As shown in Figures 1 and 2, a point cloud video encoder can perform geometry encoding and feature encoding, where the geometry encoding is performed before the feature encoding.
[0114] The point cloud video encoder according to the embodiment includes a Transformation Coordinates unit 40000, a Quantization Unit 40001, an Octree Analysis unit 40002, a Surface Approximation Analysis unit 40003, an Arithmetic Encoder 40004, a Geometry Reconstruction unit 40005, a Color Transformation unit 40006, an Attribute Transformation unit 40007, a Region Adaptive Hierachical Transform (RAHT) unit 40008, an LOD Generation unit 40009, a Lifting Transformation unit 40010, a Coefficient Quantization unit 40011 and / or an Arithmetic Encoder 40012.
[0115] The coordinate system conversion unit 40000, the quantization unit 40001, the octree analysis unit 40002, the surface approximation analysis unit 40003, the arithmetic encoder 40004, and the geometry reconstruction unit 40005 can perform geometry encoding. The geometry encoding according to the embodiment includes octree geometry coding, direct coding, trisoup geometry encoding, and entropy coding. The direct coding and trisoup geometry encoding are applied selectively or in combination. Note that the geometry encoding is not limited to the above examples.
[0116] As shown in the figure, the coordinate system conversion unit 40000 according to the embodiment receives a position and converts it into a coordinate system. For example, the position is converted into position information in a three-dimensional space (e.g., a three-dimensional space expressed in an XYZ coordinate system). The position information in the three-dimensional space according to the embodiment is also referred to as geometry information.
[0117] The quantizer 40001 according to the embodiment quantizes the geometry. For example, the quantizer 40001 quantizes the points based on the minimum position value of all points (for example, the minimum value on each axis for the X-axis, Y-axis, and Z-axis). The quantizer 40001 performs a quantization operation by multiplying the difference between the minimum position value and the position value of each point by a predetermined quantization scale value, and then rounding down or up to find the closest integer value. Therefore, one or more points can have the same quantized position (or position value). The quantizer 40001 according to the embodiment performs voxelization based on the quantized position to reconstruct the quantized points. The smallest unit including 2D image / video information is a pixel, and the points of the point cloud content (or 3D point cloud video) according to the embodiment are included in one or more voxels. A voxel is a combination of the words volume and pixel, and refers to a three-dimensional cubic space that is generated when a three-dimensional space is divided into units (unit=1.0) based on an axis (e.g., X-axis, Y-axis, Z-axis) that represents the three-dimensional space. The quantization unit 40001 can match a group of points in the three-dimensional space with a voxel. In an embodiment, one voxel can include only one point. In an embodiment, one voxel includes one or more points. In addition, to represent one voxel with one point, the position of the center point of the corresponding voxel can be set based on the position of one or more points included in the voxel. In this case, the characteristics of all positions included in one voxel are combined and assigned to the corresponding voxel.
[0118] The octree analysis unit 40002 according to the embodiment performs octree geometry coding (or octree coding) to represent the voxels in an octree structure, which represents points matched to the voxels based on an octet structure.
[0119] The surface approximation analysis unit 40003 according to the embodiment analyzes and approximates the octree. The octree analysis and approximation according to the embodiment is a process of analyzing a region containing a large number of points to voxelize it in order to efficiently provide an octree and voxelization.
[0120] The arithmetic encoder 40004 according to the embodiment entropy encodes the octree and / or the approximated octree. For example, the encoding method includes an arithmetic encoding method. As a result of the encoding, a geometry bitstream is generated.
[0121] The color transform unit 40006, the feature transform unit 40007, the RAHT transform unit 40008, the LOD generating unit 40009, the lift transform unit 40010, the coefficient quantization unit 40011 and / or the arithmetic encoder 40012 perform feature coding. As described above, one point has one or more features. The feature coding according to the embodiment is equally applied to the features of one point. However, when one feature (e.g., hue) includes one or more elements, an independent feature coding is applied to each element. The feature coding according to the embodiment includes color transformation coding, feature transformation coding, RAHT (Region Adaptive Hierarchial Transform) coding, prediction transform (Interpolarization-based hierarchical nearest-neighbour prediction-Prediction Transform) coding, and lift transform (interpolation-based hierarchical nearest-neighbour prediction with an update / lifting step (Lifting Transform)) coding. Depending on the point cloud content, the above-mentioned RAHT coding, predictive transform coding, and lift transform coding may be selectively used, or a combination of one or more coding may be used. Also, the feature coding according to the embodiment is not limited to the above examples.
[0122] The color converter 40006 according to the embodiment performs color conversion coding to convert the color values (or textures) included in the characteristics. For example, the color converter 40006 converts the format of the hue information (e.g., converts from RGB to YCbCr). The operation of the color converter 40006 according to the embodiment is optionally applied depending on the color values included in the characteristics.
[0123] The geometry reconstruction unit 40005 according to the embodiment reconstructs (restores) the octree and / or the approximated octree. The geometry reconstruction unit 40005 reconstructs the octree / voxel based on the result of analyzing the distribution of points. The reconstructed octree / voxel is also called the reconstructed geometry (or restored geometry).
[0124] The feature converter 40007 according to the embodiment performs feature conversion to convert features based on a position where geometry encoding has not been performed and / or a reconstructed geometry. As described above, features depend on geometry, so the feature converter 40007 can convert features based on reconstructed geometry information. For example, the feature converter 40007 can convert features of a point included in a voxel based on a position value of the point. As described above, when the position of the center point of a voxel is set based on the positions of one or more points included in the voxel, the feature converter 40007 converts features of one or more points. When trisoup geometry encoding is performed, the feature converter 40007 can convert features based on the trisoup geometry encoding.
[0125] The attribute conversion unit 40007 performs attribute conversion by calculating the average value of the attributes or attribute values (e.g., the hue or reflectance of each point) of adjacent points within a certain position / radius from the position (or position value) of the center point of each voxel. When calculating the average value, the attribute conversion unit 40007 applies a weighting value according to the distance from the center point to each point. Thus, each voxel has a position and a calculated attribute (or attribute value).
[0126] The feature conversion unit 40007 searches for adjacent points within a specific position / radius from the position of the center point of each voxel based on the KD tree or Moulton code. The KD tree supports a data structure that manages points based on their position so that a Nearest Neighbor Search (NNS) can be performed quickly using a binary search tree. The Moulton code is generated by mixing bits and expressing the coordinate values (e.g., (x, y, z)) indicating the three-dimensional position of all points. For example, if the coordinate value indicating the position of a point is (5, 9, 1), the bit value of the coordinate value is (0101, 1001, 0001). If the bit values are mixed according to the bit index in the order of z, y, and x, the result is 010001000111. This value is expressed as 1095 in decimal. In other words, the Moulton code value of a point with coordinate values (5, 9, 1) is 1095. The feature converter 40007 aligns points based on the Moulton code value and performs nearest neighbor search (NNS) through a depth-first traversal process. If nearest neighbor search (NNS) is required in other conversion processes for feature coding after the feature conversion operation, a KD tree or Moulton code is used.
[0127] As shown, the transformed attributes are input to a RAHT transformer 40008 and / or an LOD generator 40009 .
[0128] The RAHT converter 40008 according to the embodiment performs RAHT coding to predict feature information based on the reconstructed geometry information. For example, the RAHT converter 40008 can predict feature information of a node at a higher level of the octree based on feature information associated with a node at a lower level of the octree.
[0129] According to the embodiment, the LOD generating unit 40009 generates LOD (Level of Detail). According to the embodiment, the LOD indicates a degree of detail of the point cloud content, and it is indicated that the smaller the LOD value, the lower the detail of the point cloud content, and the larger the LOD value, the higher the detail of the point cloud content. Points can be classified according to the LOD.
[0130] The lift transform unit 40010 according to the embodiment performs lift transform coding, which transforms the characteristics of the point cloud based on weights. As described above, the lift transform coding is selectively applied.
[0131] The coefficient quantization unit 40011 according to the embodiment quantizes the feature-coded feature based on the coefficients.
[0132] The arithmetic encoder 40012 according to the embodiment encodes the quantized feature based on arithmetic coding.
[0133] The elements of the point cloud video encoder of FIG. 4 may be implemented in hardware, software, firmware, or a combination thereof, including one or more processors or integrated circuits (not shown) configured to communicate with one or more memories included in the point cloud providing apparatus. The one or more processors may perform any one of the operations and / or functions of the elements of the point cloud video encoder of FIG. 4 described above. The one or more processors may also operate or execute a software program and / or set of instructions to perform the operations and / or functions of the elements of the point cloud video encoder of FIG. 4. According to an embodiment, the one or more memories may include high speed random access memory or may include non-volatile memory (e.g., one or more magnetic disk storage devices, flash memory devices, or other non-volatile solid-state memory devices).
[0134] FIG. 5 is a diagram showing an example of a voxel according to the embodiment.
[0135] FIG. 5 shows a voxel located in a three-dimensional space represented by a coordinate system consisting of three axes, the X-axis, the Y-axis, and the Z-axis. As shown in FIG. 4, a point cloud video encoder (e.g., the quantizer 40001, etc.) performs voxelization. A voxel means a three-dimensional cubic space that is generated when a three-dimensional space is divided into units (unit=1.0) based on the axes (e.g., the X-axis, the Y-axis, and the Z-axis) that represent the three-dimensional space. FIG. 5 shows two extreme points (0,0,0) and (2 d , 2 d , 2 d ) is generated by an octree structure that recursively subdivides a bounding box (cubical axis-aligned bounding box) defined by . A voxel contains at least one point. The spatial coordinates of a voxel can be estimated from its positional relationship with a voxel group. As mentioned above, a voxel has characteristics (such as hue or reflectance) just like a pixel in a 2D image / video. A detailed description of voxels is omitted here as it has been described in FIG. 4.
[0136] FIG. 6 is a diagram illustrating an example of an octree and occupancy code according to an embodiment.
[0137] As shown in Figures 1 to 4, the point cloud content providing system (point cloud video encoder 10002) or the octree analysis unit 40002 of the point cloud video encoder performs octree geometry coding (or octree coding) based on an octree structure to efficiently manage the area and / or position of voxels.
[0138] The upper part of FIG. 6 shows an octree structure. The three-dimensional space of the point cloud content according to the embodiment is represented by the axes of a coordinate system (e.g., X-axis, Y-axis, and Z-axis). The octree structure is represented by two extreme points (0,0,0) and (2 d , 2 d , 2 d ), where 2d is set to the value that constitutes the smallest bounding box that encloses all the points in the point cloud content (or point cloud video). The value of d is determined by the following equation (1): int n , y int n , z int n ) indicates the quantized position (or position value) of the point.
[0139]
number
[0140] As shown in the upper center of Figure 6, the entire 3D space is divided into eight spaces through division. Each divided space is represented as a cube with six sides. As shown in the upper right of Figure 6, each of the eight spaces is again divided by the axes of the coordinate system (e.g., X-axis, Y-axis, Z-axis). Therefore, each space is again divided into eight smaller spaces. The divided smaller spaces are also represented as cubes with six sides. This division method is applied until the leaf nodes of the octree become voxels.
[0141] The lower part of FIG. 6 shows an occupancy code of an octree. The occupancy code of an octree is generated to indicate whether each of the eight divided spaces generated by dividing one space contains at least one point. Therefore, one occupancy code is expressed by eight child nodes. Each child node indicates the occupancy of the divided space, and has a 1-bit value. Therefore, the occupancy code is expressed by an 8-bit code. That is, if the space corresponding to a child node contains at least one point, the corresponding node has a value of 1. If the space corresponding to a node does not contain a point (empty), the corresponding node has a value of 0. The occupancy code shown in FIG. 6 is 00100001, which indicates that the spaces corresponding to the third and eighth child nodes of the eight child nodes each contain at least one point. As shown in the figure, the third and eighth child nodes each have eight child nodes, and each child node is expressed by an 8-bit occupancy code. The figure shows that the occupied code of the third child node is 10000111, and the occupied code of the eighth child node is 01001111. A point cloud video encoder according to an embodiment (e.g., the arithmetic encoder 40004) can entropy code the occupied code. In addition, to improve compression efficiency, the point cloud video encoder can intra / inter code the occupied code. A receiving device according to an embodiment (e.g., the receiving device 10004 or the point cloud video decoder 10006) reconstructs an octree based on the occupied code.
[0142] A point cloud video encoder (e.g., the octree analyzer 40002) according to an embodiment performs voxelization and octree coding to store the positions of points. However, since points in a 3D space are not always uniformly distributed, there may be certain areas where there are not many points. Therefore, it is inefficient to perform voxelization on the entire 3D space. For example, if there are almost no points in a certain area, there is no need to perform voxelization on that area.
[0143] Therefore, the point cloud video encoder according to the embodiment does not perform voxelization for the above-mentioned specific region (or nodes other than the leaf nodes of the octree), but performs direct coding, which directly codes the positions of points included in the specific region. The coordinates of the direct coding points according to the embodiment are called a direct coding mode (DCM). The point cloud video encoder according to the embodiment can also perform trisoup geometry encoding, which reconstructs the positions of points in a specific region (or node) based on voxels based on a surface model. Trisoup geometry encoding is a geometry encoding that represents an object as a series of triangle meshes. Therefore, the point cloud video decoder can generate a point cloud from the mesh surface. The direct coding and trisoup geometry encoding according to the embodiment are selectively performed. The direct coding and trisoup geometry encoding according to the embodiment can also be combined with octree geometry coding (or octree coding).
[0144] In order to perform direct coding, a direct mode use option for applying direct coding must be activated, and the node to which direct coding is applied must not be a leaf node, and points below a threshold must exist within a specific node. Also, the total number of points to be subject to direct coding must not exceed a predetermined threshold. If the above conditions are met, a point cloud video encoder (e.g., the computation encoder 40004) according to an embodiment can entropy code the positions (or position values) of the points.
[0145] The point cloud video encoder according to the embodiment (e.g., the surface approximation analysis unit 40003) can determine a specific level of the octree (if the level is smaller than the depth d of the octree), and from that level, perform trisoup geometry encoding, which uses a surface model to reconstruct the positions of points in the node area based on voxels (trisoup mode). The point cloud video encoder according to the embodiment can specify the level to which the trisoup geometry encoding is applied. For example, if the specified level is equal to the depth of the octree, the point cloud video encoder does not operate in the trisoup mode. That is, the point cloud video encoder according to the embodiment can operate in the trisoup mode only when the specified level is smaller than the depth value of the octree. The 3D cubic area of the node at the specified level according to the embodiment is called a block. One block includes one or more voxels. A block or a voxel may also correspond to a brick. Within each block, the geometry is expressed as a surface. The surface according to the embodiment can intersect each edge of the block at most once.
[0146] Since a block has 12 edges, there are at least 12 intersections within a block. Each intersection is called a vertex. A vertex along an edge is detected if there is at least one occupied voxel adjacent to the edge among all blocks that share the edge. In this embodiment, an occupied voxel refers to a voxel that contains a point. The position of a vertex detected along an edge is the average position along the edge of all voxels adjacent to the edge among all blocks that share the edge.
[0147] When a vertex is detected, the point cloud video encoder according to the embodiment may entropy code the edge start point (x, y, z), the edge direction vector (Δx, Δy, Δz), and the vertex position value (relative position value within the edge). When trisoup geometry coding is applied, the point cloud video encoder according to the embodiment (e.g., the geometry reconstruction unit 40005) may perform triangle reconstruction, up-sampling, and voxelization processes to generate a restored geometry.
[0148] Vertices located on the edges of blocks determine a surface that passes through the blocks. In this embodiment, the surface is a non-planar polygon. In the triangulation process, the surface represented by triangles is reconstructed based on the start points of the edges, the edge direction vectors, and the vertex position values. The triangulation process is as shown in Equation 2 below. (1) Calculate the centroid value of each vertex, (2) subtract the centroid value from each vertex value, and (3) square the result and add up all the values.
[0149]
number
[0150] Then, the minimum value of the added values is found, and a projection process is performed along the axis where the minimum value is. For example, if the x element is the minimum, each vertex is projected onto the x-axis based on the center of the block, and then onto the (y,z) plane. If the value obtained by projecting onto the (y,z) plane is (ai,bi), the θ value is found using atan2(bi,ai) and the vertices are aligned based on the θ value. Table 1 below shows the combination of vertices to generate a triangle depending on the number of vertices. The vertices are aligned in order from 1 to n. Table 1 below shows that for four vertices, two triangles are formed by the combination of vertices. The first triangle is formed from the 1st, 2nd, and 3rd vertices of the aligned vertices, and the second triangle is formed from the 3rd, 4th, and 1st vertices of the aligned vertices.
[0151] Table 1.Triangles formed from vertices ordered 1
[0152] [Table 1]
[0153] The upsampling process is performed to add intermediate points along the edges of triangles for voxelization. The additional points are generated based on an upsampling factor and the width of the block. The additional points are called refined vertices. The point cloud video encoder according to the embodiment can voxelize the refined vertices. The point cloud video encoder can also perform feature coding based on the voxelized positions (or position values).
[0154] FIG. 7 is a diagram illustrating an example of an adjacent node pattern according to the embodiment.
[0155] In order to increase the compression efficiency of the point cloud video, the point cloud video encoder according to the embodiment performs entropy coding based on context adaptive arithmetic coding.
[0156] As described in FIG. 1 to FIG. 6, the point cloud content providing system or the point cloud video encoder 10002 of FIG. 2 or the point cloud video encoder or the arithmetic encoder 40004 of FIG. 4 can immediately perform entropy coding on the occupancy code. In addition, the point cloud content providing system or the point cloud video encoder can perform entropy coding (intra coding) based on the occupancy code of the current node and the occupancy rate of the adjacent node, or can perform entropy coding (inter coding) based on the occupancy code of the previous frame. A frame according to the embodiment means a collection of point cloud videos generated at the same time. The compression efficiency of intra coding / inter coding according to the embodiment varies depending on the number of adjacent nodes to be referenced. Although the complexity increases as the number of bits increases, the compression efficiency can be improved by leaning to one side. For example, if there is a 3-bit context, coding is performed in 8 ways, which is the cube of 2. The part to be coded separately affects the complexity of the implementation. Therefore, it is necessary to match the compression efficiency and the appropriate level of the complexity.
[0157] FIG. 7 illustrates a process of determining an occupancy pattern based on the occupancy of adjacent nodes. The point cloud video encoder according to the embodiment obtains a neighbor pattern value by determining the occupancy of adjacent nodes of each node of an occupancy tree. The neighbor pattern is used to infer the occupancy pattern of the corresponding node. The left side of FIG. 7 illustrates a cube corresponding to the node (the cube located in the middle) and six cubes (neighboring nodes) that share at least one side with the corresponding cube. The illustrated nodes are nodes at the same depth. The illustrated numbers indicate the weights (1, 2, 4, 8, 16, 32, etc.) associated with each of the six nodes. Each weight is assigned in order according to the position of the neighboring node.
[0158] The right side of FIG. 7 shows the adjacent node pattern value. The adjacent node pattern value is the sum of values multiplied by the weight values of occupied adjacent nodes (adjacent nodes having points). Therefore, the adjacent node pattern value has a value from 0 to 63. An adjacent node pattern value of 0 means that there is no node (occupied node) having points among the adjacent nodes of the corresponding node. An adjacent node pattern value of 63 means that all adjacent nodes are occupied nodes. As shown in the figure, adjacent nodes to which weight values 1, 2, 4, and 8 are assigned are occupied nodes, so the adjacent node pattern value is 15, which is a value obtained by adding 1, 2, 4, and 8. The point cloud video encoder can perform coding according to the adjacent node pattern value (e.g., if the adjacent node pattern value is 63, 64 coding is performed). In an embodiment, the point cloud video encoder can reduce coding complexity by changing the adjacent node pattern value (e.g., based on a table that changes 64 to 10 or 6).
[0159] FIG. 8 is a diagram illustrating an example of a point configuration for each LOD according to the embodiment.
[0160] As described in Figures 1 to 7, before feature coding is performed, the coded geometry is reconstructed (restored). When direct coding is applied, the operation of geometry reconstruction includes changing the placement of the direct coded points (e.g., placing the direct coded points in front of the point cloud data). When trisoup geometry coding is applied, the process of geometry reconstruction includes the processes of triangulation, upsampling, and voxelization. Since features are dependent on geometry, feature coding is performed based on the reconstructed geometry.
[0161] A point cloud video encoder (e.g., the LOD generator 40009) can reorganize or group points by LOD. FIG. 8 shows point cloud content corresponding to LOD. The leftmost part of FIG. 8 shows original point cloud content. The second from the left in FIG. 8 shows the distribution of points with the lowest LOD, and the rightmost part of FIG. 8 shows the distribution of points with the highest LOD. That is, the points with the lowest LOD have a sparse distribution, and the points with the highest LOD have a fine distribution. That is, as the LOD increases along the direction of the arrow shown at the bottom of FIG. 8, the interval (or distance) between points becomes shorter.
[0162] FIG. 9 is a diagram illustrating an example of a point configuration for each LOD according to the embodiment.
[0163] As described in FIG. 1 to FIG. 8, a point cloud content providing system or a point cloud video encoder (e.g., point cloud video encoder 10002 in FIG. 2, point cloud video encoder or LOD generator 40009 in FIG. 4) generates LOD. The LOD is generated by realigning points with a set of refinement levels according to a set LOD distance value (or a set of Euclidean Distances). The LOD generation process is performed not only in a point cloud video encoder but also in a point cloud video decoder.
[0164] The upper part of Figure 9 shows an example of points (P0 to P9) of point cloud content distributed in 3D space. The original order in Figure 9 shows the order of points P0 to P9 before LOD generation. The LOD based order in Figure 9 shows the order of points after LOD generation. Points are re-ordered for each LOD. Also, higher LODs include points that belong to lower LODs. As shown in Figure 9, LOD0 includes P0, P5, P4, and P2. LOD1 includes points of LOD0, P1, P6, and P3. LOD2 includes points of LOD0, points of LOD1, and P9, P8, and P7.
[0165] As described in FIG. 4, the point cloud video encoder according to the embodiment may selectively or in combination perform LOD-based predictive transform coding, lift transform coding, and RAHT transform coding.
[0166] The point cloud video encoder according to the embodiment performs LOD-based predictive conversion coding to generate a predictor for each point and set a prediction characteristic (or prediction characteristic value) for each point. That is, N predictors are generated for N points. The predictor according to the embodiment can calculate a weight (=1 / distance) based on the LOD value of each point, index information for adjacent points within a distance set for each LOD, and a distance value to the adjacent point.
[0167] The predicted feature (or feature value) according to the embodiment is set as an average value of the feature (or feature value, e.g., hue, reflectance, etc.) of the adjacent point set in the predictor of each point multiplied by the weight (or weight value) calculated based on the distance to each adjacent point. The point cloud video encoder (e.g., the coefficient quantizer 40011) according to the embodiment may quantize and inverse quantize the residual value (also called residual feature, residual feature value, feature prediction residual value, prediction error feature value, etc.) of the corresponding point obtained by subtracting the corresponding predicted feature (feature value) from the feature (i.e., original feature value) of the corresponding point. The quantization process performed on the residual feature value at the transmitter is as shown in Table 2. Also, the inverse quantization process performed on the residual feature value quantized as shown in Table 2 at the receiver is as shown in Table 3.
[0168] [Table 2]
[0169] [Table 3]
[0170] The point cloud video encoder (e.g., the computation encoder 40012) according to the embodiment performs entropy coding of the quantized and dequantized residual feature values as described above if there are adjacent points in the predictor of each point. If there are no adjacent points in the predictor of each point, the point cloud video encoder (e.g., the computation encoder 40012) according to the embodiment does not perform the above process and performs entropy coding of the feature of the corresponding point.
[0171] The point cloud video encoder (e.g., lift transform unit 40010) according to the embodiment generates a predictor for each point, sets the calculated LOD in the predictor, registers adjacent points, and sets weights according to the distance to the adjacent points to perform lift transform coding. The lift transform coding according to the embodiment is similar to the above-mentioned point transform coding, but differs in that weights are cumulatively applied to feature values. The process of cumulatively applying weights to feature values according to the embodiment is as follows.
[0172] 1) Create an array QW (QuantizationWeight) to store the weight value of each point. The initial value of all elements of QW is 1.0. The QW value of the predictor index of the adjacent node registered in the predictor is multiplied by the predictor weight value of the current point and added.
[0173] 2) Lift prediction process: To calculate the predicted attribute value, the attribute value of the point is multiplied by the weighting value and subtracted from the existing attribute value.
[0174] 3) Create temporary arrays called updateweight and update, and initialize the temporary arrays to 0.
[0175] 4) The calculated weights for all predictors are multiplied by the weights stored in the QW corresponding to the predictor index, and the calculated weights are accumulated and summed in the update weight array as the index of the adjacent node. The update array is accumulated and summed by multiplying the attribute values of the index of the adjacent node by the calculated weights.
[0176] 5) Lift update process: For every predictor, divide the feature value in the update array by the weight value in the update weight array for the predictor index, and then add the divided value back to the existing feature value.
[0177] 6) For all predictors, the feature values updated in the lift update process are further multiplied by the weights updated in the lift prediction process (stored in the QW) to calculate predicted feature values. A point cloud video encoder (e.g., coefficient quantizer 40011) according to the embodiment quantizes the predicted feature values. Also, a point cloud video encoder (e.g., arithmetic encoder 40012) entropy codes the quantized feature values.
[0178] The point cloud video encoder according to the embodiment (e.g., the RAHT converter 40008) performs RAHT transform coding, which predicts the characteristics of higher level nodes using characteristics associated with lower level nodes of the octree. RAHT transform coding is an example of characteristic intra coding using octree backward scan. The point cloud video encoder according to the embodiment scans the entire region from the voxel and repeats the merging process up to the root node while combining the voxels into larger blocks at each step. The merging process according to the embodiment is performed only for occupied nodes. The merging process is not performed for empty nodes, but is performed for the node immediately above the empty node.
[0179] The following equation 3 shows the RAHT transformation matrix. JPEG0007673057000006.jpg11151 shows the average quality value of voxels at level l. JPEG0007673057000007.jpg11151 is JPEG0007673057000008.jpg10151 and Calculated from JPEG0007673057000009.jpg10152. JPEG0007673057000010.jpg10151 and The weighting for JPEG0007673057000011.jpg10152 is JPEG0007673057000012.jpg12153 and The image is JPEG0007673057000013.jpg13152.
[0180]
number
[0181] JPEG0007673057000015.jpg10152 is the low-pass value that will be used in the next higher level merging process. JPEG0007673057000016.jpg10152 are high-pass coefficients, and the high-pass coefficients at each step are quantized and entropy coded (e.g., the encoding of the computational encoder 40012). The weights are The root node is calculated by the last JPEG0007673057000018.jpg11152 and JPEG0007673057000019.jpg10153 is generated as shown in the following equation 4.
[0182]
number
[0183] The gDC values are also quantized and entropy coded like the high-pass coefficients.
[0184] FIG. 10 is a diagram illustrating an example of a point cloud video decoder according to an embodiment.
[0185] The point cloud video decoder illustrated in FIG. 10 is an example of the point cloud video decoder 10006 illustrated in FIG. 1, and performs the same or similar operations as the point cloud video decoder 10006 described in FIG. 1. As illustrated, the point cloud video decoder receives a geometry bitstream and an attribute bitstream included in one or more bitstreams. The point cloud video decoder includes a geometry decoder and an attribute decoder. The geometry decoder performs geometry decoding on the geometry bitstream to output a decoded geometry. The attribute decoder performs attribute decoding on the attribute bitstream based on the decoded geometry to output decoded attributes. The decoded geometry and decoded attributes are used to restore the point cloud content (decoded point cloud).
[0186] FIG. 11 illustrates an example of a point cloud video decoder according to an embodiment.
[0187] The point cloud video decoder shown in FIG. 11 is an example of the point cloud video decoder described in FIG. 10, and performs a decoding operation that is the reverse process of the encoding operation of the point cloud video encoder described in FIGS. 1 to 9.
[0188] As explained in Figure 1 and Figure 10, the point cloud video decoder performs geometry decoding and feature decoding, where geometry decoding is performed before feature decoding.
[0189] The point cloud video decoder according to the embodiment includes an arithmetic decoder (11000), an octree synthesis unit (11001), a surface approximation synthesis unit (11002), a geometry reconstruction unit (11003), a coordinates inverse transformation unit (11004), an arithmetic decoder (11005), an inverse quantization unit (11006), a RAHT transform unit 11007, an LOD generation unit (11008), an inverse lifting unit (11009), and / or a color inverse transformation unit (11010).
[0190] The arithmetic decoder 11000, the octree synthesis unit 11001, the surface approximation synthesis unit 11002, the geometry reconstruction unit 11003, and the coordinate system inverse transformation unit 11004 perform geometry decoding. The geometry decoding according to the embodiment includes direct decoding and trisoup geometry decoding. The direct decoding and trisoup geometry decoding are selectively applied. Furthermore, the geometry decoding is not limited to the above examples, and is performed by the reverse process of the geometry encoding described in FIG. 1 to FIG. 9.
[0191] The arithmetic decoder 11000 according to the embodiment decodes the received geometry bitstream based on arithmetic coding. The operation of the arithmetic decoder 11000 corresponds to the inverse process of the arithmetic encoder 40004.
[0192] The octree synthesis unit 11001 according to the embodiment obtains an exclusive code from the decoded geometry bitstream (or from the decoded result, information on the secured geometry) to generate an octree. The exclusive code is described in detail in FIG. 1 to FIG. 9.
[0193] The surface approximation synthesis unit 11002 according to the embodiment synthesizes a surface based on the decoded geometry and / or the generated octree if trisoup geometry encoding is applied.
[0194] The geometry reconstruction unit 11003 according to the embodiment regenerates geometry based on the surface and / or the decoded geometry. As described in FIG. 1 to FIG. 9, direct coding and trisoup geometry coding are selectively applied. Therefore, the geometry reconstruction unit 11003 directly brings and adds position information of the point to which the direct coding is applied. Also, when the trisoup geometry coding is applied, the geometry reconstruction unit 11003 restores the geometry by performing the reconstruction operation of the geometry reconstruction unit 40005, for example, triangulation, upsampling, and voxelization operations. The detailed contents are omitted since they are the same as those described in FIG. 6. The restored geometry includes a point cloud picture or frame that does not include features.
[0195] A coordinate system inverse transformation unit 11004 according to the embodiment transforms the coordinate system based on the reconstructed geometry to obtain the position of the point.
[0196] The arithmetic decoder 11005, the inverse quantization unit 11006, the RAHT transform unit 11007, the LOD generating unit 11008, the inverse lift unit 11009 and / or the color inverse transform unit 11010 perform the feature decoding described in FIG. 10. The feature decoding according to the embodiment includes RAHT (Region Adaptive Hierarchial Transform) decoding, Interpolarization-based hierarchical nearest-neighbour prediction-Prediction Transform decoding, and Lift Transform (interpolation-based hierarchical nearest-neighbour prediction with an update / lifting step (Lifting Transform)) decoding. The above three decodings are used selectively, or one or more combinations of decodings are used. In addition, the feature decoding according to the embodiment is not limited to the above examples.
[0197] An arithmetic decoder 11005 according to an embodiment decodes the attribute bitstream into arithmetic coding.
[0198] The inverse quantization unit 11006 according to the embodiment inverse quantizes the decoded feature bitstream or the information on the feature acquired as a result of decoding, and outputs the inverse quantized feature (or feature value). The inverse quantization is selectively applied based on the feature coding of the point cloud video encoder.
[0199] In an embodiment, the RAHT transform unit 11007, the LOD generator 11008 and / or the inverse lift unit 11009 process the reconstructed geometry and dequantized features. As described above, the RAHT transform unit 11007, the LOD generator 11008 and / or the inverse lift unit 11009 selectively perform corresponding decoding operations according to the encoding of the point cloud video encoder.
[0200] The color inverse converter 11010 according to the embodiment performs inverse transform coding to inversely transform color values (or textures) included in the decoded features. The operation of the color inverse converter 11010 is selectively performed based on the operation of the color converter 40006 of the point cloud video encoder.
[0201] The elements of the point cloud video decoder of Figure 11 may be implemented in hardware, software, firmware, or a combination thereof, including one or more processors or integrated circuits (not shown) configured to communicate with one or more memories included in the point cloud providing device. The one or more processors perform any of the operations and / or functions of the elements of the point cloud video decoder of Figure 11 described above. The one or more processors also operate or execute a software program and / or set of instructions to perform the operations and / or functions of the elements of the point cloud video decoder of Figure 11.
[0202] FIG. 12 illustrates an example of a transmitting device according to an embodiment.
[0203] The transmitting device shown in FIG. 12 is an example of the transmitting device 10000 in FIG. 1 (or the point cloud video encoder in FIG. 4). The transmitting device shown in FIG. 12 performs any of the same or similar operations and methods as the operations and encoding methods of the point cloud video encoder described in FIG. 1 to FIG. 9. The transmitting device according to the embodiment includes a data input unit 12000, a quantization processing unit 12001, a voxelization processing unit 12002, an octree occupation code generating unit 12003, a surface model processing unit 12004, an intra / inter coding processing unit 12005, an arithmetic coder 12006, a metadata processing unit 12007, a hue conversion processing unit 12008, a characteristic conversion processing unit (or attribute conversion processing unit) 12009, a prediction / lift / RAHT conversion processing unit 12010, an arithmetic coder 12011 and / or a transmission processing unit 12012.
[0204] According to the embodiment, the data input unit 12000 receives or acquires point cloud data. The data input unit 12000 performs an operation and / or an acquisition method that is the same as or similar to the operation and / or acquisition method of the point cloud video acquisition unit 10001 (or the acquisition process 20000 shown in FIG. 2).
[0205] Geometry encoding is performed by a data input unit 12000, a quantization processing unit 12001, a voxelization processing unit 12002, an octree occupation code generation unit 12003, a surface model processing unit 12004, an intra / inter coding processing unit 12005, and an arithmetic coder 12006. The geometry encoding according to the embodiment is the same as or similar to the geometry encoding described in Figures 1 to 9, so a detailed description will be omitted.
[0206] The quantization unit 12001 according to the embodiment quantizes geometry (e.g., position values of points). The operation and / or quantization of the quantization unit 12001 is the same as or similar to the operation and / or quantization of the quantization unit 40001 shown in Fig. 4. The specific description is as described in Figs. 1 to 9.
[0207] The voxelization processing unit 12002 according to the embodiment voxels the position values of the quantized points. The voxelization processing unit 120002 performs operations and / or processes that are the same as or similar to the operations and / or voxelization processes of the quantization unit 40001 shown in Fig. 4. The details are as described in Figs. 1 to 9.
[0208] The octree occupation code generator 12003 according to the embodiment performs octree coding on the positions of the voxelized points based on the octree structure. The octree occupation code generator 12003 generates an occupation code. The octree occupation code generator 12003 performs operations and / or methods that are the same as or similar to the operations and / or methods of the point cloud encoder (or the octree analysis unit 40002) described in Figures 4 and 6. The details are as described in Figures 1 to 9.
[0209] The surface model processing unit 12004 according to the embodiment performs trisoup geometry encoding to reconstruct the positions of points within a specific region (or node) on a voxel basis based on a surface model. The surface model processing unit 12004 performs operations and / or methods that are the same as or similar to those of the point cloud video encoder (e.g., the surface approximation analysis unit 40003) shown in Fig. 4. The details are as described in Figs. 1 to 9.
[0210] The intra / inter coding unit 12005 according to the embodiment performs intra / inter coding of the point cloud data. The intra / inter coding unit 12005 performs coding that is the same as or similar to the intra / inter coding described in FIG. 7. The details are as described in FIG. 7. In the embodiment, the intra / inter coding unit 12005 is included in the arithmetic coder 12006.
[0211] According to an embodiment, the arithmetic coder 12006 entropy codes the octree and / or approximated octree of the point cloud data. For example, the coding scheme includes an arithmetic coding method. The arithmetic coder 12006 performs operations and / or methods that are the same as or similar to those of the arithmetic encoder 40004.
[0212] The metadata processing unit 12007 according to the embodiment processes metadata related to point cloud data, such as setting values, and provides the metadata to necessary processing such as geometry encoding and / or feature encoding. The metadata processing unit 12007 according to the embodiment also generates and / or processes signaling information related to geometry encoding and / or feature encoding. The signaling information according to the embodiment is encoded and processed separately from the geometry encoding and / or feature encoding. The signaling information according to the embodiment may also be interleaved.
[0213] The hue conversion processor 12008, the feature conversion processor 12009, the prediction / lift / RAHT conversion processor 12010, and the arithmetic coder 12011 perform feature coding. The feature coding according to the embodiment is the same as or similar to the feature coding described in Figs. 1 to 9, so a detailed description thereof will be omitted.
[0214] The hue conversion processing unit 12008 according to the embodiment performs hue conversion coding to convert a hue value included in a feature. The hue conversion processing unit 12008 performs hue conversion coding based on the reconstructed geometry. The reconstructed geometry is as described in FIG. 1 to FIG. 9. The hue conversion processing unit 12008 performs operations and / or methods that are the same as or similar to the operations and / or methods of the color conversion unit 40006 described in FIG. 4. Detailed description will be omitted.
[0215] The feature conversion processing unit 12009 according to the embodiment performs feature conversion to convert features based on positions where geometry coding has not been performed and / or reconstructed geometry. The feature conversion processing unit 12009 performs the same or similar operation and / or method as the feature conversion unit 40007 described in FIG. 4. Detailed description will be omitted. The prediction / lift / RAHT conversion processing unit 12010 according to the embodiment codes the converted features by one or a combination of RAHT coding, predictive conversion coding, and lift conversion coding. The prediction / lift / RAHT conversion processing unit 12010 performs one of the same or similar operations as the RAHT conversion unit 40008, the LOD generating unit 40009, and the lift conversion unit 40010 described in FIG. 4. In addition, the predictive conversion coding, lift conversion coding, and RAHT conversion coding are the same as those described in FIG. 1 to FIG. 9, so detailed description will be omitted.
[0216] The arithmetic coder 12011 according to the embodiment encodes the coded attribute based on arithmetic coding. The arithmetic coder 12011 performs operations and / or methods that are the same as or similar to the operations and / or methods of the arithmetic encoder 40012.
[0217] The transmission processing unit 12012 according to the embodiment transmits each bitstream including the coded geometry and / or coded attribute and metadata information, or transmits the coded geometry and / or coded attribute and metadata information in one bitstream. When the coded geometry and / or coded attribute and metadata information according to the embodiment is configured in one bitstream, the bitstream includes one or more sub-bitstreams. The bitstream according to the embodiment includes signaling information including a Sequence Parameter Set (SPS) for sequence level signaling, a Geometry Parameter Set (GPS) for signaling geometry information coding, an Attribute Parameter Set (APS) for signaling attribute information coding, and a Tile Parameter Set (TPS) for tile level signaling, and slice data. The slice data includes information about one or more slices. One slice according to the embodiment includes one geometry bitstream (Geometry Parameter Set (GPS) for signaling geometry information coding, an Attribute Parameter Set (APS) for signaling attribute information coding, and a Tile Parameter Set (TPS) for tile level signaling. 0 ) and one or more attribute bit streams (Attr0 0 , Attr1 0 ).
[0218] According to an embodiment, a TPS includes information about one or more tiles (eg, bounding box coordinate information and height / size information, etc.).
[0219] The geometry bitstream includes a header and a payload. The header of the geometry bitstream according to the embodiment includes identification information of a parameter set included in the GPS (geom_parameter_set_id), a tile identifier (geom_tile_id), a slice identifier (geom_slice_id), and information about data included in the payload. As described above, the metadata processing unit 12007 according to the embodiment can generate and / or process signaling information and transmit it to the transmission processing unit 12012. In the embodiment, the element performing geometry coding and the element performing attribute coding can share data / information with each other as shown by the dotted line. The transmission processing unit 12012 according to the embodiment performs an operation and / or a transmission method that is the same as or similar to the operation and / or transmission method of the transmitter 10003. A detailed description is omitted as it is the same as that described in FIG. 1 and FIG. 2.
[0220] FIG. 13 illustrates an example of a receiving device according to the embodiment.
[0221] The receiving device shown in Figure 13 is an example of the receiving device 10004 in Figure 1 (or the point cloud video decoder in Figures 10 and 11). The receiving device shown in Figure 13 performs any of the same or similar operations and methods as the operations and decoding methods of the point cloud video decoder described in Figures 1 to 11.
[0222] The receiving device according to the embodiment includes a receiving unit 13000, a receiving processing unit 13001, an arithmetic decoder 13002, an occupancy code based octree reconstruction processing unit 13003, a surface model processing unit (triangle reconstruction, upsampling, voxelization) 13004, an inverse quantization processing unit 13005, a metadata analysis 13006, an arithmetic decoder 13007, an inverse quantization processing unit 13008, a prediction / lift / RAHT inverse transform processing unit 13009, a hue inverse transform processing unit 13010, and / or a renderer 13011. Each component of the decoding according to the embodiment performs the inverse process of the component of the encoding according to the embodiment.
[0223] The receiving unit 13000 according to the embodiment receives point cloud data. The receiving unit 13000 performs an operation and / or a receiving method that is the same as or similar to the operation and / or the receiving method of the receiver 10005 of Fig. 1. A detailed description will be omitted.
[0224] The receiving processor 13001 according to the embodiment obtains a geometry bit stream and / or an attribute bit stream from the received data. The receiving processor 13001 is included in the receiving unit 13000.
[0225] Geometry decoding is performed by an arithmetic decoder 13002, an exclusive code based octree reconstruction processor 13003, a surface model processor 13004, and an inverse quantization processor 13005. The geometry decoding according to the embodiment is the same as or similar to the geometry decoding described in Figures 1 to 10, so a detailed description will be omitted.
[0226] The computation decoder 13002 according to the embodiment decodes the geometry bitstream based on computation coding. The computation decoder 13002 performs operations and / or coding that are the same as or similar to those of the computation decoder 11000.
[0227] The exclusive code based octree reconstruction processor 13003 according to the embodiment obtains an exclusive code from the decoded geometry bitstream (or from the decoded result, information on the secured geometry) and reconstructs an octree. The exclusive code based octree reconstruction processor 13003 performs an operation and / or method identical or similar to the operation of the octree synthesis unit 11001 and / or an octree generation method. The surface model processor 13004 according to the embodiment performs trisoup geometry decoding and associated geometry reconstruction (e.g., triangle reconstruction, upsampling, voxelization) based on a surface model scheme when trisoup geometry encoding is applied. The surface model processor 13004 performs an operation identical or similar to the operation of the surface approximation synthesis unit 11002 and / or the geometry reconstruction unit 11003.
[0228] The inverse quantization unit 13005 according to the embodiment inverse quantizes the decoded geometry.
[0229] According to an embodiment, the metadata analysis 13006 analyzes metadata, such as setting values, included in the received point cloud data. The metadata analysis 13006 transmits the metadata to the geometry decoder and / or the feature decoder. A detailed description of the metadata is omitted here as it has been described in FIG. 12.
[0230] The arithmetic decoder 13007, the inverse quantization processor 13008, the prediction / lift / RAHT inverse transform processor 13009, and the hue inverse transform processor 13010 perform feature decoding. The feature decoding is the same as or similar to the feature decoding described in FIG. 1 to FIG. 10, so a detailed description thereof will be omitted.
[0231] The computation decoder 13007 according to the embodiment decodes the attribute bitstream into computation coding. The computation decoder 13007 performs decoding of the attribute bitstream based on the reconstructed geometry. The computation decoder 13007 performs operations and / or coding that are the same or similar to those of the computation decoder 11005.
[0232] The inverse quantization unit 13008 according to the embodiment inversely quantizes the decoded characteristic bit stream. The inverse quantization unit 13008 performs the same or similar operation and / or method as the inverse quantization unit 11006.
[0233] The prediction / lift / RAHT inverse transform processing unit 13009 according to the embodiment processes the reconstructed geometry and the inverse quantized features. The prediction / lift / RAHT inverse transform processing unit 13009 performs any of the same or similar operations and / or decoding as those of the RAHT transform unit 11007, the LOD generating unit 11008, and / or the inverse lift unit 11009. The hue inverse transform processing unit 13010 according to the embodiment performs inverse transform coding for inversely transforming color values (or textures) included in the decoded features. The hue inverse transform processing unit 13010 performs the same or similar operations and / or inverse transform coding as those of the color inverse transform unit 11010. The renderer 13011 according to the embodiment renders the point cloud data.
[0234] FIG. 14 illustrates an example of a structure that can be linked to a point cloud data transmitting / receiving method / apparatus according to an embodiment.
[0235] 14 shows a configuration in which any of a server 17600, a robot 17100, an autonomous vehicle 17200, an XR device 17300, a smartphone 17400, a home appliance 17500, and / or an HMD (Head-Mounted Display) 17700 is connected to a cloud network 17100. The robot 17100, the autonomous vehicle 17200, the XR device 17300, the smartphone 17400, or the home appliance 17500 may also be referred to as an apparatus. The XR device 17300 may correspond to a point cloud data (PCC) device according to an embodiment or may be linked to a PCC device.
[0236] The cloud network 17000 refers to a network that constitutes a part of a cloud computing infrastructure or exists within a cloud computing infrastructure. Here, the cloud network 17000 is configured using a 3G network, a 4G or LTE network, a 5G network, or the like.
[0237] The server 17600 can be connected to any of the robot 17100, the autonomous vehicle 17200, the XR device 17300, the smartphone 17400, the home appliance 17500, and / or the HMD 17700 via a cloud network 17000 and can assist with at least a portion of the processing of the connected devices 17100-17700.
[0238] The HMD (Head-Mount Display) 17700 represents any of the types of XR devices and / or PCC devices according to the embodiment. The HMD type device according to the embodiment includes a communications unit, a control unit, a memory unit, an I / O unit, a sensor unit, and a power supply unit.
[0239] Hereinafter, various embodiments of the devices 17100 to 17500 to which the above-mentioned technology is applied will be described. Here, the devices 17100 to 17500 shown in Fig. 14 can be linked / coupled to the point cloud data transmitting / receiving device according to the above-mentioned embodiment.
[0240] <PCC+XR>
[0241] The XR / PCC device 17300 may be embodied in a Head-Mount Display (HMD), a Head-Up Display (HUD) installed in a vehicle, a TV, a mobile phone, a smartphone, a computer, a wearable device, a home appliance, a digital sign, a vehicle, a fixed robot, a mobile robot, etc. by applying PCC and / or XR (AR+VR) technology.
[0242] The XR / PCC device 17300 may obtain information about a surrounding space or a real object by analyzing 3D point cloud data or image data obtained by various sensors or from an external device to generate position data and attribute data for 3D points, and may render and output an XR object to be output. For example, the XR / PCC device 17300 may output an XR object including additional information about the recognized object in correspondence with the recognized object.
[0243] <PCC+Self-propelled+XR>
[0244] The autonomous vehicle 17200 is realized as a mobile robot, vehicle, unmanned aerial vehicle, etc. by applying PCC technology and XR technology.
[0245] The autonomous vehicle 17200 to which the XR / PCC technology is applied refers to an autonomous vehicle equipped with a means for providing an XR image, an autonomous vehicle that is the target of control / interaction within the XR image, etc. In particular, the autonomous vehicle 17200 that is the target of control / interaction within the XR image can be separated from the XR device 17300 and linked to each other.
[0246] The autonomous vehicle 17200 equipped with a means for providing an XR / PCC image obtains sensor information from a sensor including a camera, and outputs an XR / PCC image generated based on the obtained sensor information. For example, the autonomous vehicle 17200 is equipped with a HUD and outputs an XR / PCC image, thereby providing the passenger with an XR / PCC object corresponding to a real object or an object on a screen.
[0247] At this time, when the XR / PCC object is output to the HUD, at least a part of the XR / PCC object is output to overlap with an actual object toward which the passenger's gaze is directed. On the other hand, when the XR / PCC object is output to a display provided in the autonomous vehicle, at least a part of the XR / PCC object is output to overlap with an object in the screen. For example, the autonomous vehicle 1220 may output XR / PCC objects corresponding to objects such as a road, another vehicle, a traffic light, a traffic sign, a motorcycle, a pedestrian, a building, etc.
[0248] The VR (Virtual Reality) technology, the AR (Augmented Reality) technology, the MR (Mixed Reality) technology, and / or the PCC (Point Cloud Compression) technology according to the embodiment can be applied to various devices.
[0249] That is, VR technology is a display technology that provides real objects and backgrounds only in CG images. On the other hand, AR technology is a technology that displays virtual CG images on top of images of real things. Also, MR technology is similar to the AR technology in that it mixes virtual objects into the real world. However, AR technology clearly distinguishes between real objects and virtual objects made of CG images, and uses virtual objects to complement real objects, while MR technology is different from AR technology in that virtual objects and real objects are considered to have the same characteristics. More specifically, for example, the MR technology is applied to hologram services.
[0250] However, recently, rather than clearly distinguishing between VR, AR, and MR technologies, they are also referred to as XR (extended reality) technology. Therefore, the embodiments of the present invention can be applied to any of VR, AR, MR, and XR technologies. These technologies are applied with encoding / decoding based on PCC, V-PCC, and G-PCC technologies.
[0251] The PCC method / apparatus according to the embodiment can be applied to vehicles providing autonomous driving services.
[0252] Vehicles providing autonomous driving services are connected to the PCC device via wired or wireless communication.
[0253] When the point cloud data (PCC) transceiver according to the embodiment is connected to a vehicle for wired or wireless communication, it can receive / process AR / VR / PCC service related content data that can be provided along with an autonomous driving service and transmit it to the vehicle. Also, when the point cloud data transceiver is installed in a vehicle, the point cloud transceiver can receive / process AR / VR / PCC service related content data according to a user input signal inputted to a user interface device and provide it to a user. The vehicle or user interface device according to the embodiment receives a user input signal. The user input signal according to the embodiment includes a signal instructing an autonomous driving service.
[0254] Meanwhile, the transmitting point cloud video encoder further performs a spatial division process of spatially dividing (division or partitioning) the point cloud data into one or more 3D blocks before encoding the point cloud data. That is, the transmitting device spatially divides the point cloud data into a plurality of regions so that the encoding and transmission operations of the transmitting device and the decoding and rendering operations of the receiving device can be performed in real time and at the same time with low delay processing. The transmitting device also provides an effect of enabling random access and parallel encoding in the 3D space occupied by the point cloud data by encoding and decoding the spatially divided regions (or blocks) independently or non-independently. The transmitting device and the receiving device also perform encoding and decoding independently or non-independently in units of the spatially divided regions (or blocks) to prevent repeated errors during the encoding and decoding process.
[0255] FIG. 15 is a diagram illustrating yet another example of a point cloud transmitting device according to an embodiment, which includes a space dividing unit.
[0256] The point cloud transmitting device according to the embodiment includes a data input unit 51001, a coordinate system conversion unit 51002, a quantization processing unit 51003, a spatial division unit 51004, a signaling processing unit 51005, a geometry encoder 51006, a feature encoder 51007, and a transmission processing unit 51008. According to the embodiment, the coordinate system conversion unit 51002, the quantization processing unit 51003, the spatial division unit 51004, the geometry encoder 51006, and the feature encoder 51007 are referred to as a point cloud video encoder.
[0257] A data input unit 51001 performs part or all of the operations of the point cloud video acquisition unit 10001 in Fig. 1, or performs part or all of the operations of the data input unit 12000 in Fig. 12. A coordinate system conversion unit 51002 performs part or all of the operations of the coordinate system conversion unit 40000 in Fig. 4. Furthermore, a quantization processing unit 51003 performs part or all of the operations of the quantization unit 40001 in Fig. 4, or performs part or all of the operations of the quantization processing unit 12001 in Fig. 12.
[0258] The spatial division unit 51004 spatially divides the point cloud data quantized and output by the quantization processing unit 51003 into one or more 3D blocks based on a bounding box and / or a sub-bounding box. At this time, the 3D block means a tile group, a tile, a slice, a coding unit (CU), a prediction unit (PU), or a transform unit (TU). In addition, in one embodiment, the signaling information for spatial division is entropy coded by the signaling processing unit 51005 and then transmitted by the transmission processing unit 51008 in the form of a bitstream.
[0259] Figures 16(a) to 16(c) show an example of dividing a bounding box into one or more tiles. As shown in Figure 16(a), a point cloud object corresponding to point cloud data is shown in the form of a box based on a coordinate system, which is called a bounding box. In other words, a bounding box means a hexahedron that can contain all the point cloud points.
[0260] Figures 16(b) and 16(c) show an example in which the bounding box of Figure 16(a) is divided into tile 1 (tile1#) and tile 2 (tile2#), and tile 2 (tile2#) is further divided into slice 1 (slice1#) and slice 2 (slice2#).
[0261] In one embodiment, the point cloud content may be one or more people, such as an actor, or one or more things, but may also be a map for autonomous driving, a map for in-real-world navigation of a robot, etc., in a larger scope. In this case, the point cloud content is a huge amount of data that is geographically linked. Since the point cloud content cannot be encoded / decoded all at once, tile partitioning is performed before compressing the point cloud content. For example, the number 101 in a building is divided into one tile and the other number 102 is divided into another tile. The divided tiles are further divided (or divided) into slices to support rapid encoding / decoding by applying parallelism. This is called slice partitioning (or division).
[0262] That is, a tile refers to a portion of a three-dimensional space occupied by point cloud data according to an embodiment (e.g., a rectangular cube). A tile according to an embodiment includes one or more slices. A tile according to an embodiment is divided (distributed) into one or more slices, so that a point cloud video encoder can encode point cloud data in parallel.
[0263] A slice means a unit of data (or bitstream) that can be independently encoded in a point cloud video encoder according to an embodiment, and / or a unit of data (or bitstream) that can be independently decoded in a point cloud video decoder according to an embodiment. A slice according to an embodiment means a set of data in a three-dimensional space occupied by point cloud data, or a set of partial data of point cloud data. A slice means a region of points or a set of points included in a tile according to an embodiment. A tile according to an embodiment is divided into one or more slices based on the number of points included in one tile. For example, one tile means a set of points divided by the number of points. A tile according to an embodiment may be divided into one or more slices based on the number of points, and some data may be split or merged during the division process. That is, a slice is a unit that is independently coded within a corresponding tile. The tiles spatially divided in this way are further divided into one or more slices for fast and efficient processing.
[0264] The point cloud video encoder according to the embodiment encodes the point cloud data in slice units or in tile units including one or more slices, and may perform quantization and / or transformation differently for each tile or slice.
[0265] The positions of one or more 3D blocks (e.g., slices) spatially divided by the spatial division unit 51004 are output to a geometry encoder 51006, and feature information (or features) is output to a feature encoder 51007. The positions are position information of points included in the divided units (boxes or blocks or tiles or tile groups or slices) and are referred to as geometry information.
[0266] The geometry encoder 51006 constructs an octree based on the positions output by the spatial division unit 51004, encodes (i.e., compresses) it, and outputs a geometry bitstream. The geometry encoder 51006 also reconstructs the octree and / or an approximated octree, and outputs it to the feature encoder 51007. The reconstructed octree is also called a reconstructed geometry (or a restored geometry).
[0267] The feature encoder 51007 encodes (ie, compresses) the feature output by the spatial divider 51004 based on the reconstructed geometry output by the geometry encoder 51006, and outputs a feature bitstream.
[0268] FIG. 17 is a detailed block diagram showing another example of the geometry encoder 51006 and the feature encoder 51007 according to an embodiment.
[0269] The voxelization processing unit 53001, octree generation unit 53002, geometry information prediction unit 53003 and arithmetic coder 53004 of the geometry encoder 51006 in Figure 17 perform some or all of the operations of the octree analysis unit 40002, surface approximation analysis unit 40003, arithmetic encoder 40004 and geometry reconstruction unit 40005 in Figure 4, or perform some or all of the operations of the voxelization processing unit 12002, octree occupation code generation unit 12003, surface model processing unit 12004, intra / inter coding processing unit 12005 and arithmetic coder 12006 in Figure 12.
[0270] The feature encoder 51007 in FIG. 17 includes a hue conversion processing unit 53005, a feature conversion processing unit 53006, an LOD construction unit 53007, an adjacent point set construction unit 53008, a feature information prediction unit 53009, a residual feature information quantization processing unit 53010 and an arithmetic coder 53011.
[0271] In one embodiment, a quantization processing unit is further provided between the spatial division unit 51004 and the voxelization processing unit 53001. The quantization processing unit quantizes the positions of one or more 3D blocks (e.g., slices) spatially divided by the spatial division unit 51004. In this case, the quantization unit performs part or all of the operations of the quantization unit 40001 in Fig. 4, or part or all of the operations of the quantization processing unit 12001 in Fig. 12. When a quantization processing unit is further provided between the spatial division unit 51004 and the voxelization processing unit 53001, the quantization processing unit 51003 in Fig. 15 may or may not be omitted.
[0272] The voxelization processing unit 53001 according to the embodiment performs voxelization based on the positions of one or more spatially divided 3D blocks (e.g., slices) or quantized positions. Voxelization refers to the smallest unit that represents position information in 3D space. Points of the point cloud content (or 3D point cloud video) according to the embodiment may be included in one or more voxels. Depending on the embodiment, one voxel includes one or more points. In one embodiment, if quantization is performed before performing voxelization, multiple points may belong to one voxel.
[0273] In this specification, when two or more points are included in one voxel, these two or more points are called duplicated points. That is, in the geometry encoding process, duplicated points can be generated by geometry quantization and voxelization.
[0274] The voxelization processing unit 53001 according to this embodiment does not merge overlapping points belonging to one voxel, and outputs them as they are to the octree generating unit 53002, or merges the overlapping points into one point and outputs the merged points to the octree generating unit 53002.
[0275] The octree generating unit 53002 according to this embodiment generates an octree based on the voxels output by the voxelization processing unit 53001 .
[0276] The geometry information prediction unit 53003 according to the embodiment predicts and compresses geometry information based on the octree generated by the octree generation unit 53002, and outputs the prediction and compression to the arithmetic coding unit 53004. The geometry information prediction unit 53003 also reconstructs geometry based on the position changed by the compression, and outputs the reconstructed (or decoded) geometry to the LOD construction unit 53007 of the feature encoder 51007. The reconstruction of geometry information may be performed by a device or component separate from the geometry information prediction unit 53003. In another embodiment, the reconstructed geometry is also provided to the feature conversion processing unit 53006 of the feature encoder 51007.
[0277] The hue conversion processing unit 53005 of the feature encoder 51007 corresponds to the color conversion unit 40006 of FIG. 4 or the hue conversion processing unit 12008 of FIG. 12. The hue conversion processing unit 53005 according to the embodiment performs hue conversion coding to convert the hue value (or texture) included in the feature provided by the data input unit 51001 and / or the space division unit 51004. For example, the hue conversion processing unit 53005 can convert the format of the hue information (e.g., convert from RGB to YCbCr). The operation of the hue conversion processing unit 53005 according to the embodiment is applied optionally according to the hue value included in the feature. In another embodiment, the hue conversion processing unit 53005 performs hue conversion coding based on the reconstructed geometry. For a detailed description of the geometry reconstruction, please refer to the description of FIG. 1 to FIG. 9.
[0278] The feature transformation processor 53006 according to the embodiment performs feature transformation to transform features based on the position and / or reconstructed geometry where no geometry coding is performed.
[0279] The characteristic conversion processing unit 53006 is also called a hue readjustment unit (recoloring).
[0280] The operation of the feature conversion unit 53006 according to the embodiment is applied selectively (optionally) depending on whether or not duplicated points are merged. In one embodiment, the determination of whether or not to merge duplicated points is performed in the voxelization unit 53001 or the octree generation unit 53002 of the geometry encoder 51006.
[0281] In this specification, when points belonging to one voxel are merged into one point by the voxelization processing unit 53001 or the octree generation unit 53002, the attribute conversion processing unit 53006 performs attribute conversion in one embodiment.
[0282] The attribute conversion processing unit 53006 performs operations and / or methods that are the same as or similar to the operations and / or methods of the attribute conversion unit 40007 in FIG. 4 or the attribute conversion processing unit 12009 in FIG.
[0283] According to an embodiment, the geometry information reconstructed by the geometry information prediction unit 53003 and the feature information output by the feature conversion processing unit 53006 are provided to the LOD construction unit 53007 for feature compression.
[0284] According to an embodiment, the feature information output from the feature conversion processing unit 53006 is compressed based on the reconstructed geometry information by combining one or more of a RAHT coding technique, an LOD-based predictive transformation coding technique, and a lift transformation coding technique.
[0285] In this specification, one embodiment is to perform feature compression by combining one or two of the LOD-based predictive transform coding technique and the lift transform coding technique. Therefore, the RAHT coding technique will not be described. For the description of the RAHT transform coding, please refer to the description of FIGS. 1 to 9.
[0286] The LOD construction unit 53007 according to the embodiment generates LOD (Level of Detail).
[0287] LOD is a measure of the detail of the point cloud content, with smaller LOD values indicating less detail in the point cloud content and larger LOD values indicating more detail in the point cloud content. Points are classified according to their LOD.
[0288] In one embodiment, the predictive transform coding technique and the lift transform coding technique can group points according to LOD.
[0289] This is called the LOD generation process, and groups with different LODs are called LOD l Here, l indicates the LOD and is an integer starting from 0. LOD0 is the set consisting of points with the largest distance between them, and the larger l is, the lower the LOD is. l The distance between points belonging to becomes smaller.
[0290] In the embodiment, the adjacent point set construction unit 53008 performs LOD in the LOD construction unit 53007. l Once the set is generated, the LOD l Based on the set, X (>0) nearest neighbors are searched for in groups with equal or smaller LOD (i.e., large distance between nodes) and registered in the predictor as a set of neighboring points. X is the maximum number that can be set as a neighboring point, and is input by a user parameter or signaled in the signaling information via the signaling processing unit 51005 (for example, lifting_num_pred_nearest_neighbours field signaled to the APS).
[0291] For example, in FIG. 9, adjacent points of P3 belonging to LOD1 are searched from LOD0 and LOD1. For example, if the maximum number (X) set as adjacent points is 3, the three adjacent nodes closest to P3 are P2, P4, and P6. These nodes are registered as an adjacent point set in the predictor of P3. Among these, in one embodiment, adjacent node P4 is closest to P3 based on distance, followed by P6, and then P2. Here, X=3 is an example to help those skilled in the art understand, and the value of X can be changed.
[0292] As mentioned above, every point in the point cloud data can have its own predictor.
[0293] According to an embodiment, the feature information predictor 53009 predicts feature values from adjacent points registered in the predictor. The predictor has a set of registered adjacent points and registers a weight value (1 / 2 distance) based on the distance value from each adjacent point. For example, the predictor of the P3 node has a set of adjacent points (P2, P4, P6) and calculates a weight value based on the distance value from each adjacent point. According to an embodiment, the weight value of each adjacent point is The result is JPEG0007673057000021.jpg11152.
[0294] According to an embodiment, when the adjacent point set of the predictor is set, the adjacent point set construction unit 53008 or the characteristic information prediction unit 53009 can normalize the weighted value of each adjacent point by combining the weighted values of the adjacent points.
[0295] For example, add up the weights of all the neighbors in the neighbor set of node P3 (total_weight = JPEG0007673057000022.jpg9152+ JPEG0007673057000023.jpg18151), and then divide that value again by the weighted value of each neighboring point ( JPEG0007673057000024.jpg9153 / total_weight, JPEG0007673057000025.jpg10151 / total_weight, JPEG0007673057000026.jpg10152 / total_weight), normalizing the weight value of each adjacent point.
[0296] Thereafter, the attribute information predictor 53009 can predict the attribute value via a predictor.
[0297] According to an embodiment, the average of the values obtained by multiplying the characteristics (e.g., hue, reflectance, etc.) of adjacent points registered in the predictor by a weighting value (or a normalized weighting value) is set as the predicted result (i.e., predicted characteristic value), or the characteristic of a specific point is set as the predicted result (i.e., predicted characteristic value). According to an embodiment, the predicted characteristic value is also referred to as predicted characteristic information. In addition, the residual characteristic value (or residual characteristic information, or residual) can be obtained by subtracting the predicted characteristic value of the corresponding point (which is referred to as the predicted characteristic value or predicted characteristic information) from the characteristic value of the corresponding point (i.e., the original characteristic value).
[0298] According to an embodiment, various prediction modes (or predictor indexes) are applied to pre-calculate compressed results, and then the prediction mode (i.e., predictor index) that generates the smallest bitstream can be selected from among these.
[0299] Next, the process of selecting a prediction mode will be described in detail.
[0300] In this specification, prediction mode is used synonymously with predictor index (Preindex), and is also broadly referred to as prediction method.
[0301] In this specification, it is assumed as one embodiment that the process of searching for the optimal prediction mode for each point and setting the searched prediction mode in the predictor for the corresponding point is performed by the characteristic information prediction unit 53009.
[0302] According to the embodiment, a prediction mode in which a predicted feature value is calculated by weighted average (i.e., average value of the features of adjacent points set in the predictor of each point multiplied by weighted values calculated based on the distance to each adjacent point) is called prediction mode 0. Also, a prediction mode in which the feature value of the first adjacent point is the predicted feature value is called prediction mode 1, a prediction mode in which the feature value of the second adjacent point is the predicted feature value is called prediction mode 2, and a prediction mode in which the feature value of the third adjacent point is the predicted feature value is called prediction mode 3. In other words, a prediction mode (or predictor index) value of 0 indicates that the feature value is predicted by the weighted average, 1 indicates that the feature value is predicted by the first adjacent node (i.e., adjacent point), 2 indicates that the feature value is predicted by the second adjacent node, and 3 indicates that the feature value is predicted by the third adjacent node.
[0303] According to the embodiment, the residual feature value for prediction mode 0, the residual feature value for prediction mode 1, the residual feature value for prediction mode 2, and the residual feature value for prediction mode 3 are obtained, and each score (score or double score) can be calculated based on each residual feature value. Also, the prediction mode having the lowest calculated score is selected and set as the prediction mode for the corresponding point. The score calculation method will be described in detail later.
[0304] For example, if the score calculated based on the residual feature value when the feature of the second adjacent point is used as the predicted feature value is the lowest, the predictor index value is 2, i.e., prediction mode 2 is selected as the prediction mode for the corresponding point.
[0305] As above, among the neighbors P2, P4, and P6 of P3, assume that neighbor P4 is closest to point P3 based on distance, followed by neighbor P6, which is then closest to neighbor P2. Then, the first neighbor is P4, the second neighbor is P6, and the third neighbor is P2.
[0306] This is just one example to aid in understanding this specification, and this specification is not limited to the above example, as the values assigned to the prediction modes, the number of candidate prediction modes, and the prediction methods can be added, deleted, and modified by those skilled in the art.
[0307] According to an embodiment, the process of searching for an optimal prediction mode among a plurality of prediction modes and setting it as the prediction mode of the corresponding point is performed when a predetermined condition is satisfied. Therefore, if the predetermined condition is not satisfied, the process of searching for an optimal prediction mode is not performed, and a fixed prediction mode, for example, prediction mode 0 that calculates a prediction characteristic value by a weighted average, is set as the prediction mode of the corresponding point. In one embodiment, this process is performed for each point.
[0308] According to the embodiment, a case where a certain condition is satisfied for a particular point is when the difference value of the characteristic elements (e.g., R, G, B) between adjacent points registered in the predictor of the corresponding point is equal to or greater than a predetermined threshold value (e.g., lifting_adaptive_prediction_threshold), or when the sum of the largest element value of the difference value of the characteristic elements (e.g., R, G, B) between adjacent points registered in the predictor of the corresponding point is equal to or greater than a predetermined threshold value. For example, assume that the P3 point is the corresponding point, and P2, P4, and P6 points are registered as adjacent points of P3. Furthermore, it is assumed that when the respective difference values of R, G, and B between the P2 point and the P4 point are obtained, the respective difference values of R, G, and B between the P2 point and the P6 point are obtained, and the respective difference values of R, G, and B between the P4 point and the P6 point are obtained, the R difference value is largest between the P2 point and the P4 point, the G difference value is largest between the P4 point and the P6 point, and the B difference value is largest between the P2 point and the P6 point. It is also assumed that the R difference value between the P2 point and the P4 point is largest among the largest R difference value (i.e., between P2 and P4), the largest G difference value (i.e., between P4 and P6), and the largest B difference value (i.e., between P2 and P6).
[0309] In this assumption, if the value of the R difference between the P2 point and the P4 point is equal to or greater than a predetermined threshold value, or if the sum of the R difference value between the P2 point and the P4 point, the G difference value between the P4 point and the P6 point, and the B difference value between the P2 point and the P6 point is equal to or greater than a predetermined threshold value, a process of searching for an optimal prediction mode among a plurality of candidate prediction modes can be performed. Also, a prediction mode (e.g., predIndex) can be signaled only if the value of the R difference between the P2 point and the P4 point is equal to or greater than a predetermined threshold value, or if the sum of the R difference value between the P2 point and the P4 point, the G difference value between the P4 point and the P6 point, and the B difference value between the P2 point and the P6 point is equal to or greater than a predetermined threshold value.
[0310] According to another embodiment in which a certain condition is satisfied for a particular point, the maximum difference between the attribute values (e.g., reflectance) of adjacent points registered in the predictor of the corresponding point may be equal to or greater than a predetermined threshold (e.g., lifting_adaptive_prediction_threshold). For example, when the reflectance difference value between points P2 and P4, the reflectance difference value between points P2 and P6, and the reflectance difference value between points P4 and P6 are obtained, it is assumed that the reflectance difference value between points P2 and P4 is the largest.
[0311] Under this assumption, if the reflectance difference value between the P2 point and the P4 point is equal to or greater than a predetermined threshold, a process of searching for an optimal prediction mode among a plurality of candidate prediction modes can be performed, and a prediction mode (e.g., predIndex) can be signaled only if the reflectance difference value between the P2 point and the P4 point is equal to or greater than a predetermined threshold.
[0312] According to an embodiment, the prediction mode (e.g., predIndex) can be signaled in the feature slice data. In another embodiment, if the prediction mode is not signaled, the transmitting side calculates the predicted feature value based on a default prediction mode (e.g., prediction mode 0), calculates the residual feature value based on the difference between the original feature value and the predicted feature value, and transmits the residual feature value. The receiving side calculates the predicted feature value based on a default prediction mode (e.g., prediction mode 0), and restores the feature value together with the received residual feature value.
[0313] According to an embodiment, the threshold value is either directly input or signaled in the signaling information by the signaling processing unit 51005 (eg, lifting_adaptive_prediction_threshold field signaled in the APS).
[0314] According to an embodiment, as described above, if a certain condition is satisfied for a particular point, a predictor candidate may be generated, which is also called a prediction mode or a predictor index.
[0315] According to an embodiment, the predictor candidates include prediction mode 1 to prediction mode 3. According to an embodiment, prediction mode 0 may or may not be included in the predictor candidates. According to an embodiment, at least one prediction mode not mentioned above may further be included in the predictor candidates.
[0316] 18 is a diagram showing an example of a method for selecting a prediction mode for feature coding according to an embodiment of the present invention, which is performed by the feature information predictor 53009 as an embodiment.
[0317] FIG. 18 shows an embodiment of searching for the optimal prediction mode using prediction modes 0 through 3 for the i-th point of a feature.
[0318] That is, the maximum difference between the characteristic values of adjacent points registered in the predictor of the corresponding point is calculated (step 55001), and it is determined whether the maximum difference calculated in step 55001 is equal to or greater than a predetermined threshold value (step 55002).
[0319] In step 55002 according to an embodiment, if it is determined that the maximum difference value is less than a predetermined threshold value, the prediction mode of the corresponding point is set as a predetermined base prediction mode (step 55003). In one embodiment, the predetermined base prediction mode is prediction mode 0 (i.e., feature value prediction based on weighted value average). The predetermined base prediction mode may be a prediction mode other than prediction mode 0. According to an embodiment, when the prediction mode of the corresponding point is set as a predetermined base prediction mode, the prediction mode of the corresponding point is not signaled. In this case, it is assumed that the predetermined base prediction mode is already known by the transmitting / receiving side.
[0320] If, in step 55002 according to an embodiment, it is determined that the maximum difference value is equal to or greater than a predefined threshold, a candidate predictor is generated (step 55004).
[0321] In one embodiment, the number of predictor candidates generated in operation 55004 is greater than 1. According to the embodiment, the predictor candidates generated in operation 55004 correspond to prediction modes 0 to 3.
[0322] In step 55004, when predictor candidates are generated, a process of selecting an optimal predictor by applying a rate-distortion optimization (RDO) procedure is performed.
[0323] To this end, when four predictor candidates corresponding to prediction modes 0 to 3 are generated in step 55004, a score or double score of each predictor is calculated (step 55005).
[0324] According to an embodiment, one of the four predictor candidates is set as a base predictor, and the scores of the base predictor and the remaining predictors are calculated. In one embodiment, the predictor corresponding to prediction mode 0 is set as the base predictor.
[0325] The score of the base predictor is calculated as follows:
[0326] [Number 5] double score=attrResidualQuant+kAttrPredLambdaR *(quant.stepSize()>>kFixedPointAttributeShift)
[0327] The scores of the remaining predictors are calculated as follows:
[0328] [Number 6] double score=attrResidualQuant+idxBits*kAttrPredLambdaR *(quant.stepSize()>>kFixedPointAttributeShift)
[0329] Here, idxBits = i + (i == lifting_max_num_direct_predictors - 1?1:2).
[0330] In Equation 5 and Equation 6, attrResidualQuant is a quantized value of the residual attribute value of the attribute, and kAttrPredLambdaR is a constant, which is 0.01 in one embodiment. Also, i indicates which adjacent point (i.e., node) it is (0, 1, 2, ...), and quant.stepSize represents the feature quantization size. Also, kFixedPointAttributeShift is a constant, which is 8 in one embodiment, and in this case, quant.stepSize() >> kFixedPointAttributeShift means shifting the feature quantization size to the right by 8 bits. lifting_max_num_direct_predictors is the maximum number of predictors used for prediction. In the embodiment, the value of lifting_max_num_direct_predictors is set to 3 as a default, and can be input by a user parameter or signaled by the signaling processing unit 51005 in the signaling information.
[0331] The difference between Equation 5 and Equation 6 is that idxBits is not applied in the score calculation of the base predictor.
[0332] Among the scores of the four predictor candidates obtained by applying Equations 5 and 6, a predictor candidate having the lowest score is found (step 55006). A prediction mode corresponding to the predictor candidate having the lowest score is set as the prediction mode of the corresponding point (step 55007). That is, one of prediction modes 0 to 3 is set as the prediction mode of the corresponding point (step 55008). For example, if the score obtained by applying the first adjacent point is the smallest, prediction mode 1 (i.e., predictor index 1) is set as the prediction mode of the corresponding point.
[0333] According to an embodiment, the prediction mode (eg, predIndex) of the corresponding point set in step 55007 is signaled in the feature slice data.
[0334] When a predictor candidate (i.e., a prediction mode) is selected from among multiple predictor candidates (i.e., multiple candidate prediction modes) by applying Numbers 5 and 6, idxBits is not applied to the score calculation of the base predictor corresponding to prediction mode 0, so there is a high probability that the base predictor corresponding to prediction mode 0 will be selected.
[0335] In this case, if the maximum difference between the feature values of adjacent points registered in the predictor of the corresponding point is equal to or larger than a predetermined threshold value, it means that the feature difference between the adjacent points is already large. However, if prediction mode 0 is selected as the prediction mode of the corresponding point, the predicted feature value is calculated by the average value of the values multiplied by the weights of the features of the adjacent points (i.e., the average value of the values multiplied by the weights calculated based on the distances to each adjacent point and the features of the adjacent points set in the predictor of each point). Therefore, the size of the bit stream of the residual feature value, which is the difference between the original feature value and the predicted feature value, becomes larger than the size calculated from each registered adjacent point.
[0336] In this manner, a predictor candidate selected from a plurality of predictor candidates influences a predicted feature value, which in turn influences a residual feature value.
[0337] In another embodiment of this specification, in order to improve the performance efficiency, the optimal predictor candidate may be selected from the remaining predictor candidates excluding the predictor corresponding to prediction mode 0.
[0338] For example, if the maximum difference between the attribute values of adjacent points registered in the predictor of the corresponding point is equal to or greater than a predetermined threshold value, the prediction mode with the lowest score among three predictor candidates, i.e., prediction mode 1 to prediction mode 3 (i.e., prediction mode 0 is excluded), is set as the prediction mode of the corresponding point, and if the maximum difference value is smaller than a predetermined threshold value, prediction mode 0 is set as the prediction mode of the corresponding point. Prediction mode 0 is a prediction mode that calculates a predicted attribute value by weighted average (i.e., the average value of the values obtained by multiplying the attributes of adjacent points set in the predictor of each point by weighted values calculated based on the distance to each adjacent point).
[0339] 19 is a diagram illustrating another embodiment of a method for selecting a prediction mode for feature coding according to an embodiment of the present invention, which is performed by the feature information predictor 53009 as an embodiment.
[0340] FIG. 19 shows an embodiment of searching for the optimal prediction mode using prediction modes 1 through 3 for the i-th point of a feature.
[0341] That is, the maximum difference between the characteristic values of adjacent points registered in the predictor of the corresponding point is calculated (step 57001), and it is determined whether the maximum difference calculated in step 57001 is equal to or greater than a predetermined threshold value (step 57002).
[0342] In step 57002 according to an embodiment, if it is determined that the maximum difference value is less than a predetermined threshold value, the prediction mode of the corresponding point is set as a predetermined base prediction mode (step 57003). In one embodiment, the predetermined base prediction mode is prediction mode 0 (i.e., feature value prediction based on weighted value average). The predetermined base prediction mode may be a prediction mode other than prediction mode 0. According to an embodiment, when the prediction mode of the corresponding point is set as a predetermined base prediction mode, the prediction mode of the corresponding point is not signaled. In this case, it is assumed that the predetermined base prediction mode is already known by the transmitting / receiving side.
[0343] If, in step 57002 according to the embodiment, it is determined that the maximum difference value is equal to or greater than a predefined threshold, a predictor candidate is generated (step 57004).
[0344] In one embodiment, the number of predictor candidates generated in step 57004 is greater than 1. According to this embodiment, the predictor candidates generated in step 57004 are predictor candidates corresponding to prediction mode 1 to prediction mode 3. That is, a predictor corresponding to prediction mode 0 is not included in the predictor candidates. In other words, the predictor candidates correspond to prediction modes in which a feature value of an adjacent point registered to a corresponding point is set as a prediction feature value.
[0345] When predictor candidates are generated in step 57004, a process of selecting an optimal predictor is performed by applying a rate-distortion optimization (RDO) procedure.
[0346] To this end, when three predictor candidates corresponding to prediction modes 1 to 3 are generated in step 57004, a score or double score of each predictor is calculated (step 57005).
[0347] The score of each predictor is calculated as follows:
[0348] [Number 7] double score=attrResidualQuant+kAttrPredLambdaR *(quant.stepSize()>>kFixedPointAttributeShift)
[0349] In Equation 7, attrResidualQuant is a quantized value of the residual attribute value of the attribute, kAttrPredLambdaR is a constant, and in one embodiment, it is 0.01. In addition, quant.stepSize indicates the feature quantization size, and kFixedPointAttributeShift is a constant, and in one embodiment, it is 8. In this case, quant.stepSize()>>kFixedPointAttributeShift means to shift the feature quantization size to the right by 8 bits.
[0350] According to an embodiment, in Equation 7, any of kAttrPredLambdaR, quant.stepSize(), and kFixedPointAttributeShift may be omitted.
[0351] In another embodiment, Equation 6 can be applied to calculate the score of each predictor.
[0352] In this specification, as an embodiment, a predictor candidate having the lowest score among the scores of the three predictor candidates obtained by applying Equation 7 is found (step 57006). A prediction mode corresponding to the predictor candidate having the lowest score is set as the prediction mode of the corresponding point (step 57007). That is, one of prediction modes 1 to 3 is set as the prediction mode of the corresponding point (step 57008). For example, if the score obtained by applying the first adjacent point is the smallest, prediction mode 1 (i.e., predictor index 1) is set as the prediction mode of the corresponding point.
[0353] According to an embodiment, the prediction mode (e.g., predIndex) of the corresponding point set in step 57007 can be signaled in the feature slice data.
[0354] In this way, if the difference value of the feature elements between adjacent points registered to the corresponding point is equal to or greater than the limit value, the size of the residual feature value of the corresponding point can be reduced by not using prediction mode 0, which calculates the predicted feature value by weighted average, as the prediction mode of the corresponding point.
[0355] In another embodiment, this specification provides that when three predictor candidates corresponding to prediction modes 1 to 3 are generated in step 57004, one of the three predictor candidates is set as a base predictor, and the base predictor is applied with equation (5), and the remaining two predictors are applied with equation (6) to calculate scores. In addition, a prediction mode corresponding to a predictor candidate having the lowest score among the scores of the three predictor candidates is set as a prediction mode of the corresponding point, and the set prediction mode (e.g., predIndex) is signaled to the feature slice data.
[0356] According to the embodiment, a similarity value between the attribute value of the corresponding point and the adjacent points registered to the corresponding point is calculated, and the adjacent point having the most similar attribute value is selected from the adjacent points registered to the corresponding point, and the predictor candidate that selects the corresponding adjacent point is set as the basic predictor.
[0357] To verify attribute similarity according to the embodiment, Euclidean Color Distance, Correlated Color Temperature, or the distance metric CIE94 defined by the Commission on Illumination (CIE) can be selectively used.
[0358] The Euclidean color distance according to the embodiment is calculated by the following equation (8).
[0359]
number
[0360] According to the embodiment, the CCT (correlated color temperature) is calculated by converting the RGB values into CIE (XYZ) values and normalizing them to chromatic values as shown in the following Equation 9.
[0361]
number
[0362] In this embodiment, CIE94 is calculated by the following formula (10).
[0363]
number
[0364] In this manner, the adjacent point having the most similar feature value between the feature value of the corresponding point and the adjacent points registered to the corresponding point is selected, and the predictor candidate corresponding to the corresponding adjacent point is set as the base predictor.
[0365] According to an embodiment, optional information related to predictor selection used to verify attribute similarity is signaled in the signaling information. The optional information related to predictor selection according to an embodiment includes a similar attribute search method (attribute_similarity_check_method_type) and a base predictor selection method (attribute_base_predictor selection_type). The signaling information according to an embodiment is any one of a sequence parameter set, a feature parameter set, a tile parameter set, and a feature slice header.
[0366] In this way, when the difference value of the feature element between the adjacent points registered to the corresponding point is equal to or greater than the limit value, the size of the residual feature value of the corresponding point can be reduced by setting the predictor candidate corresponding to the adjacent point having the most similar feature value between the feature value of the corresponding point and the adjacent point registered to the corresponding point as the basic predictor.
[0367] Through the above process, the prediction mode set for each point and the residual feature value in the set prediction mode are output to the residual feature information quantization processor 53010 .
[0368] According to the embodiment, the residual feature information quantization unit 53010 applies zero run length coding to the input residual feature values.
[0369] In one embodiment, this specification uses quantization and zero run length coding on the residual feature values.
[0370] The arithmetic coder 53011 according to the embodiment applies arithmetic coding to the prediction mode and the residual feature value output from the residual feature information quantization processor 53010, and outputs a feature bitstream.
[0371] The geometry bit stream compressed and output by the geometry encoder 51006 and the feature bit stream compressed and output by the feature encoder 51007 are output to the transmission processing unit 51008.
[0372] The transmission processing unit 51008 according to the embodiment performs an operation and / or a transmission method that is the same as or similar to the operation and / or the transmission method of the transmission processing unit 12012 in Fig. 12, and performs an operation and / or a transmission method that is the same as or similar to the operation and / or the transmission method of the transceiver 10003 in Fig. 1. For a detailed description, refer to the description of Fig. 1 or Fig. 12, and a detailed description will be omitted here.
[0373] In the embodiment, the transmission processing unit 51008 transmits the geometry bitstream output from the geometry encoder 51006, the feature bitstream output from the feature encoder 51007, and the signaling bitstream output from the signaling processing unit 51005 individually, or multiplexes them into a single bitstream and transmits them.
[0374] The transmission processing unit 51008 according to the embodiment can also encapsulate the bitstream into a file or a segment (eg, a streaming segment) and then transmit it via various networks such as a broadcast network and / or a broadband network.
[0375] The signaling processing unit 51005 according to the embodiment generates and / or processes signaling information and outputs it in the form of a bit stream to the transmission processing unit 51008. The signaling information generated and / or processed by the signaling processing unit 51005 is provided to the geometry encoder 51006, the attribute encoder 51007 and / or the transmission processing unit 51008 for geometry encoding, attribute encoding and transmission processing, or the signaling information generated by the geometry encoder 51006, the attribute encoder 51007 and / or the transmission processing unit 51008 is provided to the signaling processing unit 51005.
[0376] In this specification, the signaling information is signaled and transmitted in units of parameter sets (SPS: sequence parameter set, GPS: geometry parameter set, APS: attribute parameter set, TPS: tile parameter set, etc.). It can also be signaled and transmitted in units of coding units of each image such as slices or tiles. In this specification, the signaling information includes metadata (e.g., setting values, etc.) related to point cloud data, and is provided to the geometry encoder 51006, the attribute encoder 51007 and / or the transmission processing unit 51008 for geometry encoding, attribute encoding and transmission processing. Depending on the application, the signaling information is also defined at the system end such as file format, dynamic adaptive streaming over HTTP (DASH), MPEG media transport (MMT), or at the wired interface end such as High Definition Multimedia Interface (HDMI), DisplayPort, Video Electronics Standards Association (VESA), CTA, etc.
[0377] In order for the method / apparatus according to the embodiment to add / perform the operation of the embodiment, relevant information can be signaled. The signaling information according to the embodiment can be used in the transmitting device and / or the receiving device.
[0378] In this specification, the maximum number of predictors used for feature prediction (lifting_max_num_direct_predictors), threshold information for enabling feature adaptive prediction (lifting_adaptive_prediction_threshold), and option information related to predictor selection, such as a base predictor selection method (attribute_base_predictor selection_type), a similar feature search method (attribute_similarity_check_method_type), etc. are signaled in one of a sequence parameter set, a feature parameter set, a tile parameter set, and a feature slice header. In addition, predictor index information (predIndex) indicating a prediction mode corresponding to a predictor candidate selected from a plurality of predictor candidates is signaled in feature slice data.
[0379] On the other hand, the point cloud video decoder of the receiving device also uses the LOD, just like the point cloud video encoder of the transmitting device. l Generate a set and LOD l The method performs the same or similar process of searching for nearest neighbor points based on the set, registering them as a neighbor point set in the predictor, calculating weights based on distance values from each neighbor point, and normalizing the weights. The method also decodes the received prediction mode and predicts the feature value of the corresponding point according to the decoded prediction mode. The method also decodes the received residual feature value and restores the feature value of the corresponding point by adding the predicted feature value.
[0380] FIG. 20 is a diagram illustrating yet another example of a point cloud receiving device according to an embodiment.
[0381] The point cloud receiving device according to the embodiment includes a receiving processor 61001, a signaling processor 61002, a geometry decoder 61003, a feature decoder 61004, and a post-processor 61005. According to the embodiment, the geometry decoder 61003 and the feature decoder 61004 are also referred to as a point cloud video decoder. According to the embodiment, the point cloud video decoder is also referred to as a PCC decoder, a PCC decoder, a point cloud decoder, a point cloud decoder, etc.
[0382] The receiving processor 61001 according to the embodiment receives one bitstream, or receives a geometry bitstream, an attribute bitstream, and a signaling bitstream. When a file and / or a segment is received, the receiving processor 61001 according to the embodiment decapsulates the received file and / or segment and outputs the file and / or segment as a bitstream.
[0383] In this embodiment, when a bitstream is received (or decapsulated), the receiving processing unit 61001 demultiplexes the geometry bitstream, attribute bitstream, and / or signaling bitstream from the bitstream, and outputs the demultiplexed signaling bitstream to the signaling processing unit 61002, the geometry bitstream to the geometry decoder 61003, and the attribute bitstream to the attribute decoder 61004.
[0384] In the embodiment, when the geometry bitstream, attribute bitstream, and / or signaling bitstream are received (or decapsulated), the receiving processing unit 61001 transmits the signaling bitstream to the signaling processing unit 61002, the geometry bitstream to the geometry decoder 61003, and the attribute bitstream to the attribute decoder 61004.
[0385] The signaling processor 61002 parses and processes signaling information, such as information included in SPS, GPS, APS, TPS, metadata, etc., from the input signaling bitstream, and provides the information to the geometry decoder 61003, the attribute decoder 61004, and the post processor 61005. As another example, the signaling information included in the geometry slice header and / or the attribute slice header may also be pre-parsed by the signaling processor 61002 before decoding the corresponding slice data. That is, when the point cloud data is divided into tiles and / or slices as shown in FIG. 16 at the transmitting side, the TPS includes the number of slices included in each tile, so that the point cloud video decoder according to the embodiment can check the number of slices and quickly parse information for parallel decoding.
[0386] Therefore, the point cloud video decoder according to the present specification can quickly parse a bitstream including point cloud data by receiving an SPS with a reduced amount of data. The receiving device can maximize the efficiency of decoding by decoding the corresponding tile as soon as the tile is received and decoding each slice based on the GPS and APS included in each tile.
[0387] That is, the geometry decoder 61003 performs the reverse process of the geometry encoder 51006 in Fig. 15 on the compressed geometry bitstream based on signaling information (e.g., geometry-related parameters) to restore the geometry. The geometry restored (or reconstructed) by the geometry decoder 61003 is provided to the feature decoder 61004. The feature decoder 61004 performs the reverse process of the feature encoder 51007 in Fig. 15 on the compressed feature bitstream based on signaling information (e.g., feature-related parameters) and the reconstructed geometry to restore the feature. According to an embodiment, when the point cloud data is divided into tiles and / or slices at the transmitting side as shown in Fig. 16, the geometry decoder 61003 and the feature decoder 61004 perform geometry decoding and feature decoding in tile and / or slice units.
[0388] FIG. 21 is a detailed block diagram showing another example of the geometry decoder 61003 and the feature decoder 61004 according to an embodiment.
[0389] The arithmetic decoder 63001, octree reconstruction unit 63002, geometry information prediction unit 63003, inverse quantization processing unit 63004, and coordinate system inverse transformation unit 63005 included in the geometry decoder 61003 of Fig. 21 perform part or all of the operations of the arithmetic decoder 11000, octree synthesis unit 11001, surface approximation synthesis unit 11002, geometry reconstruction unit 11003, and coordinate system inverse transformation unit 11004 of Fig. 11, or perform part or all of the operations of the arithmetic decoder 13002, exclusive code based octree reconstruction processing unit 13003, surface model processing unit 13004, and inverse quantization processing unit 13005 of Fig. 13. The position restored by the geometry decoder 61003 is output to a post-process 61005.
[0390] According to an embodiment, when information on the maximum number of predictors used for feature prediction (lifting_max_num_direct_predictors), threshold information for enabling adaptive prediction of the feature (lifting_adaptive_prediction_threshold), and option information related to predictor selection, such as the base predictor selection method (attribute_base_predictor selection_type), similar feature search method (attribute_similarity_check_method_type), etc., is signaled in any of the sequence parameter set (SPS), feature parameter set (APS), tile parameter set (TPS), and feature slice header, these can be obtained by the signaling processing unit 61002 and provided to the feature decoder 61004, or can be obtained directly by the feature decoder 61004.
[0391] The feature decoder 61004 according to the embodiment includes a calculation decoder 63006, an LOD construction unit 63007, an adjacent point set construction unit 63008, a feature information prediction unit 63009, a residual feature information inverse quantization processing unit 63010 and a hue inverse transformation processing unit 63011.
[0392] The operation decoder 63006 according to the embodiment performs operation decoding of the input feature bit stream. The operation decoder 63006 performs decoding of the feature bit stream based on the reconstructed geometry. The operation decoder 63006 performs operations and / or decoding that are the same as or similar to the operations and / or decoding of the operation decoder 11005 in FIG. 11 or the operation decoder 13007 in FIG. 13.
[0393] According to an embodiment, the characteristic bitstream output from the arithmetic decoder 63006 is decoded based on the reconstructed geometry information using any one or a combination of two of the RAHT decoding, the LOD-based predictive transform decoding technique, and the lift transform decoding technique.
[0394] In this specification, since the transmitting device performs feature compression by combining one or two of the LOD-based predictive transform coding technique and the lift transform coding technique as an embodiment, the receiving device also performs feature decoding by combining one or two of the LOD-based predictive transform decoding technique and the lift transform decoding technique as an embodiment. Therefore, the description of the RAHT decoding technique in the receiving device is omitted.
[0395] According to one embodiment, the feature bitstream arithmetically decoded by the arithmetic decoder 63006 is provided to the LOD composition unit 63007. According to an embodiment, the feature bitstream provided from the arithmetic decoder 63006 to the LOD composition unit 63007 includes a prediction mode and a residual feature value.
[0396] The LOD construction unit 63007 according to the embodiment generates an LOD in the same or similar manner as the LOD construction unit 53007 of the transmitting device, and outputs the LOD to the adjacent point set construction unit 63008.
[0397] According to the embodiment, the LOD construction unit 63007 divides the points into LODs and groups them. At this time, the groups having different LODs are called LODs. l Here, l indicates the LOD and is an integer starting from 0. LOD0 is the set consisting of points with the largest distance between them, and the larger l is, the lower the LOD is. l The distance between points belonging to decreases.
[0398] According to an embodiment, the prediction modes and residual feature values coded at the transmitter are present for each LOD or only for the leaf nodes.
[0399] In one embodiment, the LOD configuration unit 63007 l When the set is generated, the adjacent point set constructor 63008 calculates the LOD lBased on the set, X (>0) nearest neighbors are searched from groups with equal or smaller LOD (i.e., large distance between nodes) and registered as a neighbor point set in the predictor. Here, X is the maximum number of neighbor points that can be set (e.g., lifting_num_pred_nearest_neighbours), which is input by a user parameter or received in signaling information such as SPS, APS, GPS, TPS, geometry slice header, and attribute slice header.
[0400] In another embodiment, the neighboring point set configuration unit 63008 selects neighboring points for each point based on signaling information such as SPS, APS, GPS, TPS, geometry slice header, and attribute slice header. To this end, the neighboring point set configuration unit 63008 is provided with the corresponding information from the signaling processing unit 61002.
[0401] For example, in FIG. 9, points P2, P4, and P6 may be selected as adjacent points of point P3 (i.e., node) belonging to LOD1 and registered as an adjacent point set in the predictor of P3.
[0402] According to an embodiment, the feature information predictor 63009 performs a process of predicting a feature value of a particular point based on a prediction mode of the particular point. In one embodiment, the feature prediction process is performed for all points or at least a part of points of the reconstructed geometry.
[0403] The prediction mode of a particular point according to the embodiment is any one of prediction modes 0 to 3.
[0404] According to an embodiment, in the transmitting side feature encoder, if the maximum difference between the feature values of adjacent points registered in the predictor of the corresponding point is smaller than a predetermined threshold, the feature encoder sets prediction mode 0 as the prediction mode of the corresponding point, and if the maximum difference is equal to or larger than the predetermined threshold, the feature encoder applies the RDO method to a number of candidate prediction modes and sets one of them as the prediction mode of the corresponding point. In one embodiment, this process is performed for each point.
[0405] According to an embodiment, in the transmitting side feature encoder, if the maximum difference between feature values of adjacent points registered in the predictor of the corresponding point is equal to or greater than a predetermined threshold value, the predictor corresponding to prediction mode 0 is set as a base predictor, and the base predictor applies Equation 5, and the predictor corresponding to the adjacent point of the corresponding point applies Equation 6 to calculate a score. Also, the prediction mode corresponding to the predictor with the lowest score is set as the prediction mode of the corresponding point. The prediction mode of the corresponding point is set to any one of prediction mode 0 to prediction mode 3.
[0406] According to the embodiment, in the feature encoder of the transmitting side, when the value of the maximum difference between the feature values of the adjacent points registered in the predictor of the corresponding point is equal to or larger than a predetermined threshold value, a similar feature value between the feature value of the corresponding point and the adjacent points registered in the corresponding point is obtained, and a predictor of the adjacent point having the most similar feature value among the adjacent points registered in the corresponding point is set as a base predictor. In addition, the base predictor applies Equation 5, and the predictors corresponding to the remaining adjacent points excluding the adjacent point set as the base predictor apply Equation 6 to calculate scores, and then a prediction mode corresponding to the predictor with the lowest score is set as a prediction mode of the corresponding point. That is, a predictor corresponding to prediction mode 0 is not used to set a prediction mode of the corresponding point. In other words, the prediction mode of the corresponding point is set to one of prediction modes 1 to 3.
[0407] According to the embodiment, in the transmitting side feature encoder, if the value of the maximum difference between the feature values of adjacent points registered in the predictor of the corresponding point is equal to or greater than a predetermined threshold value, the feature encoder calculates a score by applying Equation (7) to the predictor corresponding to the adjacent point of the corresponding point, and sets the prediction mode corresponding to the predictor with the lowest score as the prediction mode of the corresponding point. That is, the predictor corresponding to prediction mode 0 is not used to set the prediction mode of the corresponding point. In other words, the prediction mode of the corresponding point is set to one of prediction modes 1 to 3.
[0408] According to the embodiment, the prediction mode (predIndex) of the corresponding point selected by applying the RDO method is signaled to the feature slice data, so that the prediction mode of the corresponding point is obtained from the feature slice data.
[0409] According to an embodiment, the characteristic information predicting unit 63009 may predict the characteristic value of each point based on the prediction mode of each point set as described above.
[0410] For example, assuming that the prediction mode of point P3 is 0, the average of the values obtained by multiplying the characteristics of adjacent points P2, P4, and P6 registered in the predictor of P3 by weighted values (or normalized weighted values) is calculated, and the average value is determined as the predicted characteristic value of point P3.
[0411] As another example, assuming that the prediction mode of the point P3 is 1, the feature value of the adjacent point P4 registered in the predictor of P3 is determined as the predicted feature value of the point P3.
[0412] As yet another example, assuming that the prediction mode of the P3 point is 2, the feature value of the adjacent point P6 registered in the predictor of P3 is determined as the predicted feature value of the P3 point.
[0413] As yet another example, assuming that the prediction mode of the P3 point is 3, the feature value of the adjacent point P2 registered in the predictor of P3 is determined as the predicted feature value of the P3 point.
[0414] When the characteristic information prediction unit 63009 obtains a predicted characteristic value of the corresponding point based on the prediction mode of the corresponding point, the residual characteristic information inverse quantization processing unit 63010 adds the predicted characteristic value of the corresponding point predicted by the characteristic information prediction unit 63009 to the residual characteristic value of the received corresponding point to restore the characteristic value of the corresponding point, and then performs inverse quantization, which is the reverse of the quantization process of the transmitting device.
[0415] In one embodiment, when zero run length coding is applied to the residual feature value of the point at the transmitting side, the residual feature information inverse quantization processing unit 63010 performs zero run length decoding on the residual feature value of the point and then performs inverse quantization.
[0416] The characteristic values restored by the residual characteristic information inverse quantization processing unit 63010 are output to a hue inverse conversion processing unit 63011 .
[0417] The hue inverse conversion processor 63011 performs inverse conversion coding to inversely convert the color values (or texture) included in the restored feature values, and outputs the feature to the post-processor 61005. The hue inverse conversion processor 63011 performs operations and / or inverse conversion coding that are the same as or similar to the operations and / or inverse conversion coding of the color inverse conversion unit 11010 of Fig. 11 or the hue inverse conversion processor 13010 of Fig. 13.
[0418] The post-processing unit 61005 matches the positions restored and output by the geometry decoder 61003 with the features restored and output by the feature decoder 61004 to reconstruct point cloud data. If the reconstructed point cloud data is in units of tiles and / or slices, the post-processing unit 61005 performs the reverse process of the spatial division of the transmitting side based on the signaling information. For example, if a bounding box as shown in FIG. 16(a) is divided into tiles and slices as shown in FIG. 16(b) and FIG. 16(c), the post-processing unit 61005 combines the tiles and / or slices based on the signaling information to reconstruct the bounding box as shown in FIG. 16(a).
[0419] FIG. 22 is a diagram illustrating an example of a bitstream structure of point cloud data for transmission / reception according to an embodiment.
[0420] When the geometry bitstream, the attribute bitstream, and the signaling bitstream according to the embodiment are configured as one bitstream, the bitstream includes one or more sub-bitstreams. The bitstream according to the embodiment includes a sequence parameter set (SPS) for sequence level signaling, a geometry parameter set (GPS) for geometry information coding signaling, one or more attribute parameter sets (APS0, APS1) for attribute information coding signaling, a tile parameter set (TPS) for tile level signaling, and one or more slices (slice 0 to slice N). That is, the bitstream of point cloud data according to the embodiment includes one or more tiles, and each tile is a slice group including one or more slices (slice 0 to slice n). The TPS according to the embodiment includes information about each tile for one or more tiles (e.g., bounding box coordinate value information and height / size information, etc.). Each slice includes one geometry bitstream (Geom0) and one or more attribute bitstreams (Attr0, Attr1). For example, the first slice (slice 0) contains one geometry bitstream (Geom0 0 ) and one or more attribute bit streams (Attr0 0 , Attr1 0 ).
[0421] The geometry bitstream in each slice (also called geometry slice) is composed of a geometry slice header (geom_slice_header) and geometry slice data (geom_slice_data). The geometry slice header (geom_slice_header) according to the embodiment includes identification information (geom_parameter_set_id) of a parameter set included in the geometry parameter set (GPS), a tile identifier (geom_tile_id), a slice identifier (geom_slice_id), and information on data included in the geometry slice data (geom_slice_data) (geomBoxOrigin, geom_box_log2_scale, geom_max_node_size_log2, geom_num_points), etc. geomBoxOrigin is geometry box origin information indicating the box origin of the corresponding geometry slice data, geom_box_log2_scale is information indicating the log scale of the corresponding geometry slice data, geom_max_node_size_log2 is information indicating the size of the root geometry octree node, and geom_num_points is information regarding the number of points of the corresponding geometry slice data. Geometry slice data (geom_slice_data) according to the embodiment includes geometry information (or geometry data) of the point cloud data in the corresponding slice.
[0422] Each feature bit stream (or feature slice) in each slice is composed of a feature slice header (attr_slice_header) and feature slice data (attr_slice_data). According to an embodiment, the feature slice header (attr_slice_header) contains information about the feature slice data, and the feature slice data contains feature information (or feature data or feature value) of the point cloud data in the slice. When there are multiple feature bit streams in one slice, each contains different feature information. For example, one feature bit stream contains feature information corresponding to hue, and another feature stream contains feature information corresponding to reflectance.
[0423] FIG. 23 is a diagram illustrating an example of a bit stream structure of point cloud data according to an embodiment.
[0424] FIG. 24 is a diagram illustrating inter-component connection relationships within a bit stream of point cloud data according to an embodiment.
[0425] The bit stream structure of the point cloud data shown in FIGS. 23 and 24 means the bit stream structure of the point cloud data shown in FIG.
[0426] In the embodiment, the SPS includes an identifier (seq_parameter_set_id) for identifying the corresponding SPS, the GPS includes an identifier (geom_parameter_set_id) for identifying the corresponding GPS and an identifier (seq_parameter_set_id) indicating the active SPS to which the corresponding GPS belongs, and the APS includes an identifier (attr_parameter_set_id) for identifying the corresponding APS and an identifier (seq_parameter_set_id) indicating the active SPS to which the corresponding APS belongs.
[0427] A geometry bitstream (also referred to as a geometry slice) according to an embodiment includes a geometry slice header and geometry slice data, where the geometry slice header includes an identifier (geom_parameter_set_id) of an active GPS referenced by the corresponding geometry slice. The geometry slice header further includes an identifier (geom_slice_id) for identifying the corresponding geometry slice and / or an identifier (geom_tile_id) for identifying the corresponding tile. The geometry slice data includes geometry information belonging to the corresponding slice.
[0428] According to an embodiment, a feature bitstream (or feature slice) includes a feature slice header and feature slice data, where the feature slice header includes an identifier (attr_parameter_set_id) of an active APS referenced by the feature slice and an identifier (geom_slice_id) for identifying a geometry slice associated with the feature slice. The feature slice data includes feature information belonging to the slice.
[0429] That is, a geometry slice references a GPS, which references an SPS, which lists the available attributes and assigns them an identifier to identify how they are decoded, and attribute slices are mapped to output attributes by their identifiers, and attribute slices themselves have dependencies on the previous (decoded) geometry slice and the APS, which references an SPS.
[0430] According to an embodiment, parameters required for encoding point cloud data are newly defined in a parameter set of the point cloud data and / or in a corresponding slice header. For example, they can be added to a feature parameter set (APS) when encoding feature information, and to a tile and / or slice header when encoding based on tiles.
[0431] As shown in Figures 22 to 24, the bitstream of point cloud data provides tiles or slices so that the point cloud data can be divided and processed by region. Each region of the bitstream according to the embodiment has a different importance. Therefore, when the point cloud data is divided into tiles, a different filter (encoding method) and a different filter unit can be provided for each tile. Also, when the point cloud data is divided into slices, a different filter and a different filter unit can be applied to each slice.
[0432] When a transmitting device and a receiving device according to an embodiment segment and compress point cloud data, the transmitting device and the receiving device can transmit and receive a bitstream with a high-level syntax structure for selective transmission of characteristic information within the segmented regions.
[0433] The transmitting device according to the embodiment can provide a scheme for applying different encoding operations according to importance and using a good quality encoding method for important areas by transmitting point cloud data according to the bitstream structures as shown in Figures 22 to 24. Also, the transmitting device can support efficient encoding and transmission according to the characteristics of point cloud data and provide characteristic values according to user requirements.
[0434] The receiving device according to the embodiment receives point cloud data according to the bitstream structure as shown in Fig. 22 to Fig. 24, and can apply different filtering (decoding) to each region (region divided into tiles or slices) instead of using a complex decoding (filtering) method on the entire point cloud data depending on the processing capacity of the receiving device. Therefore, it is possible to provide better image quality to the regions that are more important to the user and ensure appropriate latency in the system.
[0435] As described above, tiles or slices are provided to process point cloud data by dividing it into regions. When dividing point cloud data into regions, an option to generate different adjacent point sets for each region can be set to provide a low complexity but somewhat lower reliability option, or conversely, a high complexity but highly reliable option.
[0436] According to an embodiment, any of the SPS, APS, TPS, and attribute slice header for each slice includes option information related to predictor selection. According to an embodiment, the option information related to predictor selection includes a base predictor selection method (attribute_base_predictor selection_type) and a similarity attribute search method (attribute_similarity_check_method_type).
[0437] The term field, as used in the syntax of this specification below, has the same meaning as parameter or element.
[0438] 25 is a diagram showing an example of a syntax structure of a sequence parameter set (seq_parameter_set_rbsp()) (SPS) according to this specification. The SPS includes sequence information of a point cloud data bit stream, and in particular includes option information related to predictor selection.
[0439] The SPS according to the embodiment includes a profile_idc field, a profile_compatibility_flags field, a level_idc field, an sps_bounding_box_present_flag field, an sps_source_scale_factor field, an sps_seq_parameter_set_id field, an sps_num_attribute_sets field, and an sps_extension_present_flag field.
[0440] The profile_idc field indicates the profile that the bitstream conforms to.
[0441] A value of 1 in the profile_compatibility_flags field indicates that the bitstream conforms to the profile indicated by profile_idc.
[0442] The level_idc field indicates the level to which the bitstream conforms.
[0443] The sps_bounding_box_present_flag field indicates whether source bounding box information is signaled to the SPS. The source bounding box information includes source bounding box offset and size information. For example, a value of 1 in the sps_bounding_box_present_flag field indicates that source bounding box information is signaled to the SPS, and a value of 0 indicates that it is not signaled. The sps_source_scale_factor field indicates the scale factor of the source point cloud.
[0444] The sps_seq_parameter_set_id field provides an identifier for the SPS for reference by other syntax elements.
[0445] The sps_num_attribute_sets field indicates the number of coded attributes in the bitstream.
[0446] The sps_extension_present_flag field indicates whether the sps_extension_data syntax structure is present in the corresponding SPS syntax structure. For example, when the value of the sps_extension_present_flag field is 1, the sps_extension_data syntax structure is present in this SPS syntax structure, and when it is 0, it is not present (equal to 1 specifies that the sps_extension_data syntax structure is present in the SPS syntax structure. The sps_extension_present_flag field equal to 0 specifies that this syntax structure is not present. When not present, the value of the sps_extension_present_flag field is inferred to be equal to 0).
[0447] An SPS according to the embodiment further includes, when the value of the sps_bounding_box_present_flag field is 1, an sps_bounding_box_offset_x field, an sps_bounding_box_offset_y field, an sps_bounding_box_offset_z field, an sps_bounding_box_scale_factor field, an sps_bounding_box_size_width field, an sps_bounding_box_size_height field, and an sps_bounding_box_size_depth field.
[0448] The sps_bounding_box_offset_x field indicates the x-offset of the source bounding box in Cartesian coordinates. If there is no source bounding box x-offset, the value of the sps_bounding_box_offset_x field is 0.
[0449] The sps_bounding_box_offset_y field indicates the y offset of the source bounding box in Cartesian coordinates. If there is no source bounding box y offset, the value of the sps_bounding_box_offset_y field is 0.
[0450] The sps_bounding_box_offset_z field indicates the z-offset of the source bounding box in Cartesian coordinate system. If there is no source bounding box z-offset, the value of the sps_bounding_box_offset_z field is 0.
[0451] The sps_bounding_box_scale_factor field indicates the scale factor of the source bounding box in Cartesian coordinate system. If there is no source bounding box scale factor, the value of the sps_bounding_box_scale_factor field is 1.
[0452] The sps_bounding_box_size_width field indicates the width of the source bounding box in Cartesian coordinates. If the source bounding box width is not present, the value of the sps_bounding_box_size_width field is 1.
[0453] The sps_bounding_box_size_height field indicates the height of the source bounding box in Cartesian coordinates. If the source bounding box height is not present, the value of the sps_bounding_box_size_height field is 1.
[0454] The sps_bounding_box_size_depth field indicates the depth of the source bounding box in the Cartesian coordinate system. If the source bounding box depth is not present, the value of the sps_bounding_box_size_depth field is 1.
[0455] The SPS according to the embodiment includes a repeat statement that is repeated by the value of the sps_num_attribute_sets field. In this case, in one embodiment, i is initialized to 0, and increases by 1 each time the repeat statement is executed, and the repeat statement is repeated until the i value becomes the value of the sps_num_attribute_sets field. This repeat statement includes an attribute_dimension[i] field, an attribute_instance_id[i] field, an attribute_bitdepth[i] field, an attribute_cicp_colour_primaries[i] field, an attribute_cicp_transfer_characteristics[i] field, an attribute_cicp_matrix_coeffs[i] field, an attribute_cicp_video_full_range_flag[i] field, and a known_attribute_label_flag[i] field.
[0456] The attribute_dimension[i] field specifies the number of components of the i-th attribute.
[0457] The attribute_instance_id[i] field indicates the instance identifier of the i-th attribute.
[0458] The attribute_bitdepth[i] field specifies the bitdepth of the i-th attribute signal(s).
[0459] The Attribute_cicp_colour_primaries[i] field indicates the chromaticity coordinates of the colour attribute source primaries of the i-th attribute.
[0460] The attribute_cicp_transfer_characteristics[i] field indicates the reference opto-electronic transfer characteristic function of the colour attribute as a function of a source input linear optical intensity with a nominal real-valued range of 0 to 1 or indicates the inverse of the reference electro-optical transfer characteristic function as a function of an output linear optical intensity.
[0461] The attribute_cicp_matrix_coeffs[i] field describes the matrix coefficients used in deriving luma and chroma signals from the green, blue, and red, or Y, Z, and X primaries, of the i-th attribute.
[0462] The attribute_cicp_video_full_range_flag[i] field specifies indicates the black level and range of the luma and chroma signals as derived from E'Y, E'PB, and E'PR or E'R, E'G, and E'B real-valued component signals of the i-th attribute.
[0463] The known_attribute_label[i] field indicates whether the know_attribute_label field or the attribute_label_four_bytes field is signaled for the i-th attribute. For example, a value of 1 in the known_attribute_label_flag[i] field indicates that the know_attribute_label field is signaled for the i-th attribute, and a value of 1 in the known_attribute_label_flag[i] field indicates that the attribute_label_four_bytes field is signaled for the i-th attribute.
[0464] The known_attribute_label[i] field indicates the type of the attribute. For example, a value of 0 in the known_attribute_label[i] field indicates that the i-th attribute is color, a value of 1 in the known_attribute_label[i] field indicates that the i-th attribute is reflectance, and a value of 1 in the known_attribute_label[i] field indicates that the i-th attribute is frame index.
[0465] The attribute_label_four_bytes field indicates a known attribute type as a four-byte code.
[0466] In one embodiment, a value of 0 in the attribute_label_four_bytes field indicates color and a value of 1 indicates reflectance.
[0467] According to the embodiment, when the value of the SPS_extension_present_flag field is 1, the SPS may further include an SPS_extension_data_flag field.
[0468] The sps_extension_data_flag field can have any value.
[0469] The SPS according to the embodiment further includes option information related to predictor selection. According to the embodiment, the option information related to predictor selection includes a base predictor selection method (attribute_base_predictor selection_type) and a similarity attribute search method (attribute_similarity_check_method_type).
[0470] According to the embodiment, the option information related to predictor selection is included in a repeat statement that is repeated the same number of times as the value of the sps_num_attribute_sets field.
[0471] That is, the iteration statement can further include an attribute_base_predictor selection_type[i] field and an attribute_similarity_check_method_type[i] field.
[0472] The attribute_base_predictor selection_type[i] field specifies the base predictor selection method for the i-th attribute. For example, a value of 0 in the attribute_base_predictor selection_type[i] field indicates that the base predictor is selected based on the weighted average of neighbours, and a value of 1 indicates that the base predictor is selected based on attribute similarity.
[0473] The attribute_similarity_check_method_type[i] field indicates the attribute similarity search method when referring to the attribute similarity of the i-th attribute. For example, if the value of the attribute_similarity_check_method_type[i] field is 0, it indicates that Euclidean Color Distance as shown in Equation 8 is used, if the value is 1, it indicates that Correlated Color Temperature as shown in Equation 9 is used, and if the value is 2, it indicates that CIE94 as shown in Equation 10 is used.
[0474] In this way, the attribute_base_predictor selection_type[i] field and the attribute_similarity_check_method_type[i] field are signaled to the SPS.
[0475] 26 is a diagram illustrating an example syntax structure of a geometry_parameter_set() (GPS) according to this specification. According to an embodiment, a GPS contains information about how to encode geometry information of point cloud data included in one or more slices.
[0476] The GPS in the embodiment includes a gps_geom_parameter_set_id field, a gps_seq_parameter_set_id field, a gps_box_present_flag field, a unique_geometry_points_flag field, a neighborhood_context_restriction_flag field, an inferred_direct_coding_mode_enabled_flag field, a bitwise_occupancy_coding_flag field, an adjacent_child_contextualization_enabled_flag field, a log2_neighbour_avail_boundary field, a log2_intra_pred_max_node_size field, a log2_trisoup_node_size field, and a gps_extension_present_flag field.
[0477] The gps_geom_parameter_set_id field provides an identifier for the GPS for reference by other syntax elements.
[0478] The gps_seq_parameter_set_id field indicates the value of the seq_parameter_set_id field for the active SPS.
[0479] The gps_box_present_flag field indicates whether additional bounding box information is provided from a geometry slice header that references the current GPS. For example, a value of 1 in the gps_box_present_flag field indicates that additional bounding box information is provided in a geometry header that references the current GPS. Thus, when the value of the gps_box_present_flag field is 1, the GPS further includes a gps_gsh_box_log2_scale_present_flag field.
[0480] The gps_gsh_box_log2_scale_present_flag field indicates whether the gps_gsh_box_log2_scale field is signaled in each geometry slice header that references the current GPS. For example, a value of 1 in the gps_gsh_box_log2_scale_present_flag field indicates that the gps_gsh_box_log2_scale field is signaled in each geometry slice header that references the current GPS. As another example, a value of 0 in the gps_gsh_box_log2_scale_present_flag field indicates that the gps_gsh_box_log2_scale field is not signaled in each geometry slice header that references the current GPS, and a common scale for all slices is signaled in the gps_gsh_box_log2_scale field of the current GPS.
[0481] If the value of the gps_gsh_box_log2_scale_present_flag field is 0, the GPS further includes a gps_gsh_box_log2_scale field.
[0482] The gps_gsh_box_log2_scale field indicates the common scale factor of the bounding box origin for all slices that reference the current GPS.
[0483] The unique_geometry_points_flag field indicates whether all output points have unique positions. For example, when the value of the unique_geometry_points_flag field is 1, it indicates that all output points have unique positions. When the value of the unique_geometry_points_flag field is 0, it indicates that two or more output points may have the same positions (equal to 1 indicates that all output points have unique positions. unique_geometry_points_flag field equal to 0 indicates that the output points may have same positions).
[0484] The neighbour_context_restriction_flag field indicates which contexts octree occupancy coding uses. For example, a value of 0 in the neighbour_context_restriction_flag field indicates that octree occupancy coding uses contexts determined from six neighbouring parent nodes. A value of 1 in the neighbour_context_restriction_flag field indicates that octree occupancy coding uses contexts determined from sibling nodes only.
[0485] The inferred_direct_coding_mode_enabled_flag field indicates whether the direct_mode_flag field exists in the corresponding geometry node syntax. For example, if the value of the inferred_direct_coding_mode_enabled_flag field is 1, it indicates that the direct_mode_flag field exists in the corresponding geometry node syntax. For example, if the value of the inferred_direct_coding_mode_enabled_flag field is 0, it indicates that the direct_mode_flag field does not exist in the corresponding geometry node syntax.
[0486] The bitwise_occupancy_coding_flag field indicates whether the geometry node occupancy is coded using the bit contextualization of its syntax element occupancy map. For example, a value of 1 in the bitwise_occupancy_coding_flag field indicates that the geometry node occupancy is coded using the bit contextualization of its syntax element occupancy_map. For example, a value of 0 in the bitwise_occupancy_coding_flag field indicates that the geometry node occupancy is coded using the directory-coded syntax element occupancy_byte.
[0487] The adjacent_child_contextualization_enabled_flag field indicates whether adjacent children of neighboring octree nodes are used for bitwise occupancy contextualization. For example, a value of 1 in the adjacent_child_contextualization_enabled_flag field indicates that adjacent children of neighboring octree nodes are used for bitwise occupancy contextualization. For example, a value of 0 in the adjacent_child_contextualization_enabled_flag field indicates that children of neighboring octree nodes are not used for bit occupancy contextualization.
[0488] The log2_neighbour_avail_boundary field specifies the value of the variable NeighbAvailBoundary that is used in the decoding process as follows:
[0489] NeighbAvailBoundary = 2 log2_neighbour_avail_boundary
[0490] For example, if the value of the neighbour_context_restriction_flag field is 1, NeighbAvailabilityMask is set to 1. For example, if the value of the neighbour_context_restriction_flag field is 0, NeighbAvailabilityMask is set to 1 << log2_neighbour_avail_boundary.
[0491] The log2_intra_pred_max_node_size field specifies the octree nodesize eligible for occupancy intra prediction.
[0492] The log2_trisoup_node_size field specifies the variable TrisoupNodeSize as the size of the triangle nodes as follows.
[0493] TrisoupNodeSize = 1 << log2_trisoup_node_size
[0494] The gps_extension_present_flag field indicates whether the gps_extension_data syntax structure is present in the corresponding GPS syntax. For example, when the value of the gps_extension_present_flag field is 1, it indicates that the gps_extension_data syntax structure is present in the corresponding GPS syntax. For example, when the value of the gps_extension_present_flag field is 0, it indicates that the gps_extension_data syntax structure is not present in the corresponding GPS syntax.
[0495] In the embodiment, when the value of the gps_extension_present_flag field is 1, the GPS further includes a gps_extension_data_flag field.
[0496] The gps_extension_data_flag field may have any value. Its presence and value do not affect decoder conformance to profiles.
[0497] 27 is a diagram showing an example of a syntax structure of an attribute parameter set (attribute_parameter_set()) (APS) according to this specification. An APS according to an embodiment includes information on how to encode attribute information of point cloud data included in one or more slices, and in particular, shows an example including option information related to predictor selection.
[0498] An APS according to the embodiment includes an aps_attr_parameter_set_id field, an aps_seq_parameter_set_id field, an attr_coding_type field, an aps_attr_initial_qp field, an aps_attr_chroma_qp_offset field, an aps_Slice_qp_delta_present_flag field, and an aps_extension_present_flag field.
[0499] The aps_attr_parameter_set_id field indicates an identifier for the APS for reference by other syntax elements.
[0500] The aps_seq_parameter_set_id field indicates the value of sps_seq_parameter_set_id for an active SPS.
[0501] The attr_coding_type field indicates the coding type for the attribute.
[0502] In one embodiment, a value of 0 in the attr_coding_type field indicates predictive weight lifting, a value of 1 indicates RAHT, and a value of 2 indicates fix weight lifting.
[0503] The aps_attr_initial_qp field indicates the initial value of the variable slice quantization parameter (SliceQp) for each slice referring to the APS. The initial value of SliceQp is modified at the attribute slice segment layer when a non-zero value of Slice_qp_delta_luma or Slice_qp_delta_luma are decoded.
[0504] The aps_attr_chroma_qp_offset field specifies the offsets to the initial quantization parameter signalled by the syntax aps_attr_initial_qp.
[0505] The aps_slice_qp_delta_present_flag field indicates whether the ash_attr_qp_delta_luma and ash_attr_qp_delta_luma syntax elements are present in the corresponding attribute slice header (ASH). For example, a value of the aps_slice_qp_delta_present_flag field of 1 indicates that the ash_attr_qp_delta_luma and ash_attr_qp_delta_luma syntax elements are present in the corresponding attribute slice header (ASH). For example, a value of 0 in the aps_slice_qp_delta_present_flag field specifies that the ash_attr_qp_delta_luma and ash_attr_qp_delta_luma syntax elements are not present in the ASH.
[0506] In the embodiment, when the value of the attr_coding_type field is 0 or 2, i.e., the coding type is predictive weight lifting or fix weight lifting, the APS further includes lifting_num_pred_nearest_neighbours field, lifting_max_num_direct_predictors field, lifting_search_range field, lifting_lod_regular_sampling_enabled_flag field, and Lifting_num_detail_levels_minus1 field.
[0507] The lifting_num_pred_nearest_neighbours field indicates the maximum number of nearest neighbors used for prediction.
[0508] The lifting_max_num_direct_predictors field indicates the maximum number of predictors to be used for direct prediction. The value of the variable MaxNumPredictors that is used in the point cloud data decoding process according to the embodiment is expressed as follows:
[0509] MaxNumPredictors = lifting_max_num_direct_predictors field + 1
[0510] The lifting_lifting_search_range field specifies the search range used to determine nearest neighbours to be used for prediction and to build distance-based levels of detail.
[0511] The lifting_num_detail_levels_minus1 field specifies the number of levels of detail for the attribute coding.
[0512] The lifting_lod_regular_sampling_enabled_flag field indicates whether levels of detail (LOD) are built by a regular sampling strategy. For example, a value of 1 in the lifting_lod_regular_sampling_enabled_flag field indicates that LODs are built by a regular sampling strategy. For example, a value of 0 in the lifting_lod_regular_sampling_enabled_flag field indicates that a distance-based sampling strategy is used instead.
[0513] The APS according to the embodiment includes a repetition statement that is repeated by the value of the lifting_num_detail_levels_minus1 field. In this case, the index (idx) is initialized to 0 and increases by 1 each time the repetition statement is executed, and the repetition statement is repeated until the index (idx) becomes larger than the value of the lifting_num_detail_levels_minus1 field. If the value of the lifting_lod_decimation_enabled_flag field is true (e.g., 1), this repetition statement includes a lifting_sampling_period[idx] field, and if the value is false (e.g., 0), it includes a lifting_sampling_distance_squared[idx] field.
[0514] The lifting_sampling_period[idx] field specifies the sampling period for the level of detail idx.
[0515] The lifting_sampling_distance_squared[idx] field specifies the square of the sampling distance for the level of detail idx.
[0516] According to the embodiment, if the value of the attr_coding_type field is 0, that is, the coding type is predictive weight lifting, the APS further includes a lifting_adaptive_prediction_threshold field and a lifting_intra_lod_prediction_num_layerS field.
[0517] The lifting_adaptive_prediction_threshold field specifies the threshold to enable adaptive prediction.
[0518] The lifting_intra_lod_prediction_num_layers field specifies number of LOD layer where decoded points in the same LOD layer could be referred to generate prediction value of target point. For example, the lifting_intra_lod_prediction_num_layers field equal to num_detail_levels_minus1 plus 1 indicates that target point could refer decoded points in the same LOD layer for all LOD layers. For example, the lifting_intra_lod_prediction_num_layers field equal to 0 indicates that target point could not refer decoded points in the same LOD layer for any LOD layers.
[0519] The aps_extension_present_flag field indicates whether the aps_extension_data syntax structure is present in the corresponding APS syntax structure. For example, if the value of the aps_extension_present_flag field is 1, it indicates that the aps_extension_data syntax structure is present in the corresponding APS syntax structure. For example, if the value of the aps_extension_present_flag field is 0, it indicates that the aps_extension_data syntax structure is not present in the corresponding APS syntax structure.
[0520] According to the embodiment, when the value of the aps_extension_present_flag field is 1, the APS further includes an aps_extension_data_flag field.
[0521] The aps_extension_data_flag field may have any value. Its presence and value do not affect decoder conformance to profiles.
[0522] The APS according to the embodiment further includes option information related to predictor selection.
[0523] According to the embodiment, when the value of the attr_coding_type field is 0 or 2, i.e., the coding type is predictive weight lifting or fix weight lifting, the APS further includes an attribute_base_predictor selection_type field and an attribute_similarity_check_method_type field.
[0524] The attribute_base_predictor selection_type field indicates the base predictor selection method. For example, a value of 0 in the attribute_base_predictor selection_type field indicates that the base predictor is selected based on the weighted average of neighbours, and a value of 1 indicates that the base predictor is selected based on attribute similarity.
[0525] The attribute_similarity_check_method_type field indicates a search method for attribute similarity. For example, if the value of the attribute_similarity_check_method_type field is 0, it indicates that Euclidean Color Distance (Equation 8) is used, if the value is 1, it indicates that Correlated Color Temperature (Equation 9) is used, and if the value is 2, it indicates that CIE94 (Equation 10) is used.
[0526] In this way, the attribute_base_predictor selection_type field and the attribute_similarity_check_method_type field can be signaled to the APS.
[0527] FIG. 28 is a diagram showing an example of a syntax structure of a tile parameter set (tile_parameter_set()) (TPS) according to this specification. In some embodiments, the TPS (Tile Parameter Set) is also called a tile inventory. In this embodiment, the TPS includes information related to each tile for each tile, and in particular, shows an example including option information related to prediction predictor selection.
[0528] A TPS according to an embodiment includes a num_tiles field.
[0529] The num_tiles field indicates the number of tiles signaled for the bitstream. When not present, the value of the num_tiles field is inferred to be 0.
[0530] The TPS according to the embodiment includes a repeat statement that is repeated by the value of the num_tiles field. In this case, in one embodiment, i is initialized to 0 and increments by 1 each time the repeat statement is executed, and the repeat statement is repeated until the i value becomes the value of the num_tiles field. This repeat statement includes a tile_bounding_box_offset_x[i] field, a tile_bounding_box_offset_y[i] field, a tile_bounding_box_offset_z[i] field, a tile_bounding_box_size_width[i] field, a tile_bounding_box_size_height[i] field, and a tile_bounding_box_size_depth[i] field.
[0531] The tile_bounding_box_offset_x[i] field indicates the X offset of the i-th tile in the Cartesian coordinates.
[0532] The tile_bounding_box_offset_y[i] field gives the y offset of the ith tile in the Cartesian coordinate system.
[0533] The tile_bounding_box_offset_z[i] field indicates the z offset of the i-th tile in the Cartesian coordinate system.
[0534] The tile_bounding_box_size_width[i] field indicates the width of the i-th tile in the Cartesian coordinate system.
[0535] The tile_bounding_box_size_height[i] field indicates the height of the i-th tile in the Cartesian coordinate system.
[0536] The tile_bounding_box_size_depth[i] field indicates the depth of the i-th tile in the Cartesian coordinate system.
[0537] The TPS according to the embodiment further includes option information related to predictor selection.
[0538] According to the embodiment, the option information related to predictor selection is included in a repeat statement that is repeated as many times as the value of the num_tiles field described above.
[0539] That is, the iterative statement further includes an attribute_base_predictor selection_type[i] field and an attribute_similarity_check_method_type[i] field.
[0540] The attribute_base_predictor selection_type[i] field indicates the base predictor selection method for the i-th tile. For example, a value of 0 in the attribute_base_predictor selection_type[i] field indicates that the base predictor is selected based on the weighted average of neighbours, and a value of 1 indicates that the base predictor is selected based on attribute similarity.
[0541] The attribute_similarity_check_method_type[i] field indicates the attribute similarity search method in the i-th tile. For example, if the value of the attribute_similarity_check_method_type[i] field is 0, it indicates that Euclidean Color Distance (Equation 8) is used, if the value is 1, it indicates that Correlated Color Temperature (Equation 9) is used, and if the value is 2, it indicates that CIE94 (Equation 10) is used.
[0542] In this way, the attribute_base_predictor selection_type[i] field and the attribute_similarity_check_method_type[i] field can be signaled to the TPS.
[0543] FIG. 29 is a diagram showing an example of a syntax structure of a geometry slice bitstream according to this specification.
[0544] A geometry slice bitstream (geometry_slice_bitstream()) according to the embodiment includes a geometry slice header (geometry_slice_header()) and geometry slice data (geometry_slice_data()).
[0545] FIG. 30 is a diagram showing an example of the syntax structure of a geometry slice header (geometry_slice_header()) according to this specification.
[0546] A bitstream transmitted by a transmitting device (or received by a receiving device) according to an embodiment includes one or more slices. Each slice includes a geometry slice and an attribute slice. The geometry slice includes a Geometry Slice Header (GSH). The attribute slice includes an attribute slice Header (ASH).
[0547] The geometry slice header (geometry_slice_header()) according to the embodiment includes a gsh_geom_parameter_set_id field, a gsh_tile_id field, a gsh_slice_id field, a gsh_max_node_size_log2 field, a gsh_num_points field, and a byte_alignment() field.
[0548] In the embodiment, the geometry slice header (geometry_slice_header()) further includes a gsh_box_log2_scale field, a gsh_box_origin_x field, a gsh_box_origin_y field, and a gsh_box_origin_z field when the value of the gps_box_present_flag field included in the geometry parameter set (GPS) is true (e.g., 1) and the value of the gps_gsh_box_log2_scale_present_flag field is true (e.g., 1).
[0549] The gsh_geom_parameter_set_id field specifies the value of the gps_geom_parameter_set_id of the active gps.
[0550] The gsh_tile_id field indicates the identifier of the corresponding tile referenced by the corresponding Geometry Slice Header (GSH).
[0551] The gsh_slice_id indicates the identifier of the slice for reference by other syntax elements.
[0552] The gsh_box_log2_scale field indicates the scale factor of the bounding box origin for this slice.
[0553] The gsh_box_origin_x field indicates the x value of the bounding box origin scaled by the value of the gsh_box_log2_scale field.
[0554] The gsh_box_origin_y field indicates the y value of the bounding box origin scaled by the value of the gsh_box_log2_scale field.
[0555] The gsh_box_origin_z field indicates the z value of the bounding box origin scaled by the value of the gsh_box_log2_scale field.
[0556] The gsh_max_node_size_log2 field indicates the size of the root geometry octree node.
[0557] The gsh_points_number field indicates the number of coded points in the slice.
[0558] 31 is a diagram showing an example of a syntax structure of geometry slice data (geometry_slice_data()) according to this specification. Geometry slice data (geometry_slice_data()) according to the embodiment transmits a geometry bitstream belonging to a corresponding slice.
[0559] The geometry slice data (geometry_slice_data()) according to the embodiment includes a first repeat statement that is repeated by the value of MaxGeometryOctreeDepth. In this case, in one embodiment, depth is initialized to 0 and increases by 1 each time the repeat statement is executed, and the first repeat statement is repeated until depth becomes the value of MaxGeometryOctreeDepth. The first repeat statement includes a second repeat statement that is repeated by the value of NumNodesAtDepth. In this case, in one embodiment, nodeIdx is initialized to 0 and increases by 1 each time the repeat statement is executed, and the second repeat statement is repeated until nodeIdx becomes the value of NumNodesAtDepth. The second repeat statement includes xN=NodeX[depth][nodeIdx], yN=NodeY[depth][nodeIdx], zN=NodeZ[depth][nodeIdx], geometry_node(depth, nodeIdx, xN, yN, zN). MaxGeometryOctreeDepth indicates the maximum value of the geometry octree depth, and NumNodesAtDepth indicates the number of nodes to be decoded at the corresponding depth. The variables NodeX[depth][nodeIdx], NodeY[depth][nodeIdx], and NodeZ[depth][nodeIdx] indicate the x, y, and z coordinates of the nodeIdx-th node at a given depth in decoding order. geometry_node(depth, nodeIdx, xN, yN, zN) transmits the geometry bitstream of the corresponding node at the corresponding depth.
[0560] The geometry slice data (geometry_slice_data()) according to the embodiment further includes geometry_trisoup_data() if the value of the log2_trisoup_node_size field is greater than 0. In other words, if the size of the triangle node is greater than 0, a geometry bitstream encoded with trisoup geometry is transmitted by geometry_trisoup_data().
[0561] FIG. 32 is a diagram showing an example of a syntax structure of a feature slice bitstream() according to this specification.
[0562] The attribute slice bitstream (attribute_slice_bitstream()) according to the embodiment includes an attribute slice header (attribute_slice_header()) and attribute slice data (attribute_slice_data()).
[0563] 33 is a diagram showing an example of a syntax structure of an attribute slice header (attribute_slice_header()) according to this specification. The attribute slice header according to the embodiment includes signaling information for the corresponding attribute slice, and in particular, shows an example including option information related to predictor selection.
[0564] The attribute slice header (attribute_slice_header()) according to the embodiment includes an ash_attr_parameter_set_id field, an ash_attr_sps_attr_idx field, and an ash_attr_geom_slice_id field.
[0565] The attribute slice header (attribute_slice_header()) according to the embodiment further includes ash_qp_delta_luma and ash_qp_delta_chroma fields when the value of the aps_slice_qp_delta_present_flag field of the attribute parameter set (APS) is true (eg, 1).
[0566] The ash_attr_parameter_set_id field indicates the value of the aps_attr_parameter_set_id field of the currently active APS (for example, the aps_attr_parameter_set_id field included in the APS described in FIG. 27).
[0567] The ash_attr_sps_attr_idx field identifies an attribute set within the current active SPS. The value of the ash_attr_sps_attr_idx field ranges from 0 to the sps_num_attribute_sets field contained in the current active SPS.
[0568] The ash_attr_geom_slice_id field indicates the value of the gsh_slice_id field of the current geometry slice header.
[0569] The ash_qp_delta_luma field indicates the luma delta quantization parameter (qp) derived from the initial slice qp in the active attribute parameter set.
[0570] The ash_qp_delta_chroma field indicates the chroma delta quantization parameter (qp) derived from the initial slice qp in the active characteristic parameter set.
[0571] The attribute slice header (attribute_slice_header()) according to the embodiment further includes optional predictor selection related information, such as an attribute_base_predictor selection_type field and an attribute_similarity_check_method_type field.
[0572] The attribute_base_predictor selection_type field indicates the base predictor selection method. For example, a value of 0 in the attribute_base_predictor selection_type field indicates that the base predictor is selected based on the weighted average of neighbours, and a value of 1 indicates that the base predictor is selected based on attribute similarity.
[0573] The attribute_similarity_check_method_type field indicates the attribute similarity search method. For example, if the value of the attribute_similarity_check_method_type field is 0, it indicates that Euclidean Color Distance (Equation 8) is used, if the value is 1, it indicates that Correlated Color Temperature (Equation 9) is used, and if the value is 2, it indicates that CIE94 (Equation 10) is used.
[0574] In this way, the attribute_base_predictor selection_type field and the attribute_similarity_check_method_type field can be signaled in the attribute slice header (attribute_slice_header()).
[0575] 34 is a diagram showing an example of a syntax structure of attribute slice data (attribute_slice_data()) according to this specification. The attribute slice data (attribute_slice_data()) according to the embodiment transmits an attribute bitstream belonging to a corresponding slice.
[0576] In the attribute slice data (attribute_slice_data()) of Figure 34, dimension = attribute_dimension [ash_attr_sps_attr_idx] indicates the attribute dimension (attribute_dimension) of the attribute set identified by the ash_attr_sps_attr_idx field in the corresponding attribute slice header. The attribute dimension (attribute_dimension) means the number of components that make up the attribute. In the embodiment, the attribute indicates reflectance, hue, etc. Therefore, the number of components that the attribute has varies. For example, the attribute corresponding to hue has three hue components (e.g., RGB). Therefore, the attribute corresponding to reflectance is a mono-dimensional attribute, and the attribute corresponding to hue is a three-dimensional attribute.
[0577] Feature coding in accordance with the preferred embodiment is per-dimension feature coding.
[0578] For example, a feature corresponding to reflectance and a feature corresponding to hue may be feature-coded separately. Also, features according to the embodiment may be feature-coded together regardless of dimension. For example, a feature corresponding to reflectance and a feature corresponding to hue may be feature-coded together.
[0579] In FIG. 34, zerorun specifies the number of 0 prior to the residual.
[0580] Also in FIG. 34, i means the i-th point value of the attribute, and the attr_coding_type field and lifting_adaptive_prediction_threshold field are signaled to the APS in one embodiment.
[0581] Also, in FIG. 34, the variable MaxNumPredictors is a variable used in the process of decoding point cloud data, and is obtained as follows based on the lifting_adaptive_prediction_threshold field value signaled to the APS.
[0582] MaxNumPredictoRS = lifting_max_num_direct_predictors field + 1
[0583] Here, the lifting_max_num_direct_predictors field indicates the maximum number of predictors used for direct prediction.
[0584] In this embodiment, predIindex[i] specifies the predictor index (or prediction mode) to decode the i-th point value of the attribute. The value of predIndex[i] ranges from 0 to the lifting_max_num_direct_predictors field value.
[0585] The variable MaxPredDiff{i] according to the embodiment is calculated as follows:
[0586] minValue=maxValue= JPEG0007673057000030.jpg9152
[0587] for (j=0;j <k;j++){
[0588] minValue = Min(minValue, JPEG0007673057000031.jpg9152)
[0589] maxValue = Max(maxValue, JPEG0007673057000032.jpg9151)
[0590] }
[0591] MaxPredDiff[i]=maxValue-minValue;
[0592] where k i is the set of k-nearest neighbors of the current point i, Let JPEG0007673057000033.jpg10151 be their decoded / reconstructed attribute value (Let k i be the set of the k-nearest neighbors of the current point and let JPEG0007673057000034.jpg10151be their decoded / reconstructed attribute values). Also, the number of nearest neighbors, k i The number of nearest neighbours, k, ranges from 1 to the value of the lifting_num_pred_nearest_neighbours field. i lifting_num_pred_nearest_neighbours shall be in the range of 1 to lifting_num_pred_nearest_neighbours. According to an embodiment, the decoded / reconstructed attribute value of neighbours are derived according to the Predictive lifting decoding process.
[0593] The lifting_num_pred_nearest_neighbours field is signaled to the APS and indicates the maximum number of nearest neighbors to be used for prediction.
[0594] FIG. 35 is a flowchart of a method for transmitting point cloud data according to an embodiment.
[0595] A point cloud data transmission method according to an embodiment includes a step 71001 of encoding geometry contained in the point cloud data, a step 71002 of encoding features contained in the point cloud data based on an input and / or reconstructed geometry, and a step 71003 of transmitting a bitstream including the encoded geometry, the encoded features and signaling information.
[0596] Steps 71001, 71002 for encoding the geometry and features contained in the point cloud data perform some or all of the operations of the point cloud video encoder 10002 of Figure 1, the encoder 20001 of Figure 2, the point cloud video encoder of Figure 4, the point cloud video encoder of Figure 12, the point cloud encoding of Figure 14, the point cloud video encoder of Figure 15, and the geometry encoder and feature encoder of Figure 17.
[0597] In one embodiment, the prediction mode for a particular point set in step 71002 of encoding a feature is one of prediction modes 0 to 3.
[0598] According to the embodiment, if the maximum difference between the feature values of the adjacent points registered in the predictor of the corresponding point is smaller than a predetermined threshold, the prediction mode 0 is set as the prediction mode of the corresponding point, and if it is equal to or larger than the predetermined threshold, the RDO method is applied to a plurality of candidate prediction modes, and one of them is set as the prediction mode of the corresponding point. In one embodiment, this process is performed for each point.
[0599] According to an embodiment, if the maximum difference between the feature values of adjacent points registered in the predictor of the corresponding point is equal to or greater than a predetermined threshold value, a predictor corresponding to prediction mode 0 is set as a base predictor, and the base predictor applies Equation 5, and a predictor corresponding to an adjacent point of the corresponding point applies Equation 6 to calculate a score. Also, a prediction mode corresponding to a predictor with the lowest score is set as the prediction mode of the corresponding point. The prediction mode of the corresponding point is set to any one of prediction mode 0 to prediction mode 3.
[0600] According to the embodiment, if the maximum difference between the feature values of the adjacent points registered in the predictor of the corresponding point is equal to or larger than a predetermined threshold value, a similar feature value between the feature value of the corresponding point and the adjacent points registered in the corresponding point is obtained, and a predictor of the adjacent point having the most similar feature value among the adjacent points registered in the corresponding point is set as a base predictor. In addition, the base predictor applies Equation 5, and predictors corresponding to the remaining adjacent points excluding the adjacent point set as the base predictor apply Equation 6 to calculate scores, and then a prediction mode corresponding to a predictor with the lowest score is set as a prediction mode of the corresponding point. That is, a predictor corresponding to prediction mode 0 is not used to set a prediction mode of the corresponding point. In other words, the prediction mode of the corresponding point is set to one of prediction modes 1 to 3.
[0601] According to the embodiment, if the value of the maximum difference between the feature values of adjacent points registered in the predictor of the corresponding point is equal to or greater than a predetermined threshold value, a score is calculated by applying Equation 7 to the predictor corresponding to the adjacent point of the corresponding point, and the prediction mode corresponding to the predictor with the lowest score is set as the prediction mode of the corresponding point. That is, the predictor corresponding to prediction mode 0 is not used to set the prediction mode of the corresponding point. In other words, the prediction mode of the corresponding point is set to any one of prediction modes 1 to 3.
[0602] According to an embodiment, the prediction mode (predIndex) of the corresponding point selected by applying the RDO method is signaled in the feature slice data.
[0603] According to an embodiment, the step 71002 of encoding the feature may apply zero run length encoding to the residual feature value of each point.
[0604] According to an embodiment, steps 71001 and 71002 for encoding geometry and attributes may perform encoding in units of slices or tiles including one or more slices.
[0605] The step 71003 of transmitting a bitstream including the encoded geometry, encoded attributes and signaling information can also be performed in the transceiver 10003 of FIG. 1, the transmitting step 20002 of FIG. 2, the transmitting processing unit 12012 of FIG. 12 or the transmitting processing unit 51008 of FIG. 15.
[0606] FIG. 36 is a flow chart of a method for receiving point cloud data according to an embodiment.
[0607] A method for receiving point cloud data according to an embodiment includes step 81001 of receiving a bitstream including encoded geometry, encoded attributes, and signaling information, step 81002 of decoding the geometry based on the signaling information, step 81003 of decoding the attributes based on the decoded / reconstructed geometry and signaling information, and step 81004 of rendering reconstructed point cloud data based on the decoded geometry and decoded attributes.
[0608] The step 81001 of receiving a bitstream including encoded geometry, encoded attributes, and signaling information according to the embodiment can be performed by reception 10005 in FIG. 1, transmission 20002 or decoding 20003 in FIG. 2, reception unit 13000 or reception processing unit 13001 in FIG. 13, or reception processing unit 61001 in FIG. 20.
[0609] According to an embodiment, steps 81002, 81003 for decoding geometry and attributes perform decoding in units of slices or tiles including one or more slices.
[0610] The step 81002 of decoding geometry according to the embodiment performs some or all of the operations of the point cloud video decoder 10006 of Figure 1, the decoder 20003 of Figure 2, the point cloud video decoder of Figure 11, the point cloud video decoder of Figure 13, the geometry decoder of Figure 20, and the geometry decoder of Figure 21.
[0611] The step 81003 of decoding features according to the embodiment performs some or all of the operations of the point cloud video decoder 10006 of Figure 1, the decoder 20003 of Figure 2, the point cloud video decoder of Figure 11, the point cloud video decoder of Figure 13, the feature decoder of Figure 20 and the feature decoder of Figure 21.
[0612] According to an embodiment, the signaling information, e.g., any of the sequence parameter set, attribute parameter set, tile parameter set, and attribute slice header, includes information on the maximum number of predictors used for attribute prediction (lifting_max_num_direct_predictors), information on a threshold for enabling adaptive prediction of the attribute (lifting_adaptive_prediction_threshold), a base predictor selection method (attribute_base_predictor selection_type), and a similar attribute search method (attribute_similarity_check_method_type), etc.
[0613] According to an embodiment, the feature slice data includes predictor index information (predIndex) that indicates the prediction mode for a particular point.
[0614] According to an embodiment, in the step 81003 of decoding a feature, when a particular point is feature-decoded, if predictor index information of the corresponding point is not signaled in the corresponding feature slice data, the step predicts the feature value of the corresponding point based on prediction mode 0.
[0615] According to an embodiment, in the step 81003 of decoding a feature, when a particular point is feature-decoded, if predictor index information of the corresponding point is signaled in the corresponding feature slice data, the step predicts the feature value of the corresponding point based on the signaled prediction mode.
[0616] According to an embodiment, when a particular point is decoded, the step of decoding attributes 81003 adds the residual attribute value of the corresponding point received to the predicted attribute value to restore the attribute value of the corresponding point. The step of decoding attributes 81003 performs this process for each point to restore the attribute value of each point.
[0617] In some embodiments, if the residual attribute value is zero-run length coded, the step of decoding the attribute 81003 performs zero-run length decoding, which is the reverse process of the transmitting side, before recovering the attribute value.
[0618] The step 81004 of rendering the restored point cloud data based on the decoded geometry and the decoded attributes according to the embodiment may render the restored point cloud data in various rendering manners. For example, the points of the point cloud content are rendered as a vertex having a certain thickness, a cube having a certain minimum size centered on the position of the corresponding vertex, or a circle centered on the position of the vertex. All or a part of the region of the rendered point cloud content is provided to a user via a display (e.g., a VR / AR display, a general display, etc.).
[0619] The step 81004 of rendering the point cloud data according to the embodiment is performed by the renderer 10007 in FIG. 1 or the rendering 20004 in FIG. 2 or the renderer 13011 in FIG.
[0620] This specification provides that if the maximum difference between the feature values of adjacent points registered in the predictor of the corresponding point is equal to or greater than a predetermined threshold value, a prediction mode that calculates a predicted feature value by a weighted average in a predictor candidate is not set as the prediction mode of the corresponding point, thereby reducing the size of the bitstream including the residual feature values and thereby improving the compression efficiency of the features.
[0621] Each of the above-mentioned parts, modules or units is software, a processor or a hardware part that performs a continuous execution process stored in a memory (or a storage unit). Each step described in the above embodiments is performed by a processor, software or hardware part. Each of the modules / blocks / units described in the above embodiments operates as a processor, software or hardware. Also, the method presented in the embodiments is executed as code. This code is written in a processor-readable storage medium and is thus read by the processor provided by the device.
[0622] In addition, in the entire specification, when a part "includes" a certain element, this does not mean excluding other elements, but means further including other elements, unless otherwise specified. Furthermore, the term "part" or the like in the specification means a unit that processes at least one function or operation, which is realized by hardware, software, or a combination of hardware and software.
[0623] For convenience of explanation, each drawing is described separately, but it is possible to combine the embodiments described in each drawing to design a new embodiment. In addition, according to the needs of a person skilled in the art, it is also within the scope of the embodiment to design a computer readable recording medium on which a program for executing the previously described embodiment is recorded.
[0624] As described above, the apparatus and method according to the embodiments are not limited to the configurations and methods of the described embodiments, and the embodiments can be modified in various ways and can be configured by selectively combining all or part of each embodiment.
[0625] Although preferred embodiments have been shown and described, the embodiments are not limited to the specific embodiments described above, and various modifications may be made by a person having ordinary skill in the art to which the invention pertains without departing from the spirit of the embodiments claimed in the claims, and such modifications should not be understood individually from the technical ideas or prospects of the embodiments.
[0626] Various components of the apparatus according to the embodiment may be implemented by hardware, software, firmware, or a combination thereof. Various components of the embodiment may be implemented by one chip, for example, one hardware circuit. In the embodiment, the components according to the embodiment may be implemented by individual chips. In the embodiment, any of the components of the apparatus according to the embodiment may be implemented by one or more processors capable of executing one or more programs, and the one or more programs may include instructions for performing or causing any one or more of the operations / methods according to the embodiment to be performed. Executable instructions for performing the method / operation of the apparatus according to the embodiment may be stored in a non-transitory CRM or other computer program product configured to be executed by one or more processors, or may be stored in a transitory CRM or other computer program product configured to be executed by one or more processors. In addition, the memory according to the embodiment is used as a concept including not only volatile memory (e.g., RAM, etc.) but also non-volatile memory, flash memory, PROM, etc. It may also be implemented in the form of a carrier wave, such as transmission over the Internet. In addition, the recording medium read by the processor may be distributed to computer systems connected by a network, and the code read by the processor may be stored and executed in a distributed manner.
[0627] In this specification, " / " and "," are interpreted as "and / or". For example, "A / B" is interpreted as "A and / or B", and "A, B" is interpreted as "A and / or B". Furthermore, "A / B / C" means "any of A, B and / or C". Also, "A, B, C" means "any of A, B and / or C". Furthermore, in this document, "or" is interpreted as "and / or". For example, "A or B" means 1) only "A", 2) only "B", or 3) "A and B". In other words, in this specification, "or" means "additionally or alternatively".
[0628] Terms such as first and second are used to describe various components of the embodiments. However, the various components according to the embodiments should not be limited in interpretation by the above terms. Such terms are merely used to distinguish one component from another. For example, a first user input signal can be referred to as a second user input signal. Similarly, a second user input signal can be referred to as a first user input signal. Use of such terms does not depart from the scope of various embodiments. Although the first user input signal and the second user input signal are both user input signals, they do not mean the same user input signal unless the context clearly indicates otherwise.
[0629] The terms used to describe the embodiments are used to describe specific embodiments and are not intended to limit the embodiments. As used in the description of the embodiments and the claims, the singular includes the plural unless the context clearly dictates otherwise. The term "and / or" is used to include all possible combinations between terms. "Comprise" describes the presence of a feature, number, step, element, and / or component, and does not mean that additional features, numbers, steps, elements, and / or components are not included. Conditional expressions such as "if" and "when" used to describe the embodiments are only intended to be interpreted as selective, and are not interpreted as limiting. When a particular condition is met, a related operation is performed in response to the particular condition, or a related definition is intended to be interpreted. In addition, the operations according to the embodiments described in this specification are performed by a transceiver including a memory and / or a processor according to the embodiments. The memory stores a program for processing / controlling the operations according to the embodiments, and the processor controls various operations described in this specification. The processor is also referred to as a controller. In the embodiments, the operations are performed by firmware, software, and / or a combination thereof, and the firmware, software, and / or a combination thereof are stored in the processor or in the memory.
[0630] The best mode for carrying out the embodiment has been described above. [Industrial Applicability]
[0631] As mentioned above, the embodiments may be applied in whole or in part to devices and systems for transmitting and receiving point cloud data.
[0632] Those skilled in the art may modify or alter the embodiments in various ways within the scope of the embodiments.
[0633] The embodiments include modifications / variations, which modifications / variations are within the scope of the claims and their scope.
Claims
1. 1. A method for processing point cloud data in a transmitting device, comprising: acquiring the point cloud data; encoding geometry information including positions of points within the point cloud data; encoding feature information including feature values of points in the point cloud data based on the geometry information; transmitting the encoded geometry information, the encoded characteristic information and signaling information; The step of encoding the characteristic information comprises: generating predictor candidates based on a plurality of nearest neighboring points of a point to be encoded when a maximum difference value between feature values of the nearest neighboring points of the point to be encoded is equal to or greater than a threshold value, each predictor candidate corresponding to a respective prediction mode, and a predictor corresponding to prediction mode 0 is excluded from the predictor candidates, the prediction mode 0 being a mode based on a distance-based weighted average value of the nearest neighboring points; obtaining a score for each of the predictor candidates and setting a prediction mode corresponding to the predictor candidate having the lowest score, the score for each of the predictor candidates being obtained based on a respective residual feature value obtained by applying the input feature value of the point to a prediction feature value for the respective prediction mode; quantizing residual feature values of the points based on the set prediction mode; and encoding said quantized residual feature values of said points.
2. The method of claim 1 , wherein the quantized residual feature values of the points are entropy coded.
3. The set prediction mode is one of prediction mode 1, prediction mode 2, or prediction mode 3, the predicted feature value of the point for the prediction mode 1 is set as the feature value of a first nearest neighbor of the plurality of nearest neighbor points; the predicted feature value of the point for the prediction mode 2 is set as the feature value of a second nearest neighbor of the plurality of nearest neighbor points; The method of claim 1 , wherein the predicted feature value of the point for the prediction mode 3 is set as the feature value of a third nearest neighbor of the plurality of nearest neighbor points.
4. setting the prediction mode to 0 if the maximum difference value between the feature values of the nearest neighboring points is less than the threshold value; quantizing residual feature values of the points based on the set prediction mode; The method of claim 1 , further comprising the step of: encoding the quantized residual feature values of the points.
5. A transmitting device for processing point cloud data, comprising: an acquisition unit configured to acquire the point cloud data; a geometry encoder configured to encode geometry information including positions of points within the point cloud data; a feature encoder configured to encode feature information including feature values of points in the point cloud data based on the geometry information; a transmitter configured to transmit the encoded geometry information, the encoded characteristic information and signaling information; The feature encoder includes: generating predictor candidates based on a plurality of nearest neighboring points of a point to be encoded if a maximum difference value between the feature values of the nearest neighboring points of the point to be encoded is equal to or greater than a threshold value, each predictor candidate corresponding to a respective prediction mode, and a predictor corresponding to a prediction mode 0 is excluded from the predictor candidates, the prediction mode 0 being a mode based on a distance-based weighted average value of the nearest neighboring points; obtaining a score for each predictor candidate and setting a prediction mode corresponding to the predictor candidate having the lowest score, the score for each predictor candidate being obtained based on a respective residual feature value obtained by applying the input feature value of the point to a prediction feature value for the respective prediction mode; quantizing the residual feature value of the point based on the set prediction mode; A transmitting device for encoding the quantized residual quality values of the points.
6. 6. A transmitting device according to claim 5, wherein the quantized residual quality values of the points are entropy coded.
7. The set prediction mode is one of prediction mode 1, prediction mode 2, or prediction mode 3, the predicted feature value of the point for the prediction mode 1 is set as the feature value of a first nearest neighbor of the plurality of nearest neighbor points; the predicted feature value of the point for the prediction mode 2 is set as the feature value of a second nearest neighbor of the plurality of nearest neighbor points; The transmitting device of claim 5 , wherein the predicted quality value of the point for the prediction mode 3 is set as the quality value of a third nearest neighbor point of the plurality of nearest neighbor points.
8. The feature encoder includes: If the maximum difference value between the feature values of the nearest neighboring points is less than the threshold, set the prediction mode to 0; quantizing the residual feature value of the point based on the set prediction mode 0; 6. A transmitting device as claimed in claim 5, further comprising means for encoding a quality value of the quantized residual of the point.
9. 1. A receiving device for processing point cloud data, comprising: a receiver configured to receive geometry information, attribute information and signaling information; a geometry decoder configured to decode the geometry information based on the signaling information to recover a position of a point; a feature decoder configured to decode the feature information based on the signaling information and the geometry information to recover feature values of the points; a renderer configured to render the reconstructed point cloud data based on the point positions and the attribute values; The attribute decoder comprises: obtaining a prediction mode for a point to be decoded, and if a maximum difference value between feature values of a plurality of nearest neighboring points of the point is equal to or greater than a threshold value, a prediction mode 0 is excluded from the obtained prediction modes, the prediction mode 0 being a mode based on a distance-based weighted average value of the plurality of nearest neighboring points, and a predicted feature value of the point is set as a feature value of one of the plurality of nearest neighboring points according to the obtained prediction mode; performing inverse quantization on the residual feature values of said points in said feature information; A receiving device recovering a feature value of the point based on the predicted feature value of the point and the dequantized residual feature value of the point in the feature information.
10. the obtained prediction mode of the point is one of prediction mode 1, prediction mode 2, or prediction mode 3; the predicted feature value of the point for the prediction mode 1 is set as the feature value of a first nearest neighbor of the plurality of nearest neighbor points; the predicted feature value of the point for the prediction mode 2 is set as the feature value of a second nearest neighbor of the plurality of nearest neighbor points; 10. The receiving apparatus of claim 9, wherein the predicted quality value of the point for the prediction mode 3 is set as the quality value of a third nearest neighbor point of the plurality of nearest neighbor points.
11. if the maximum difference value between the feature values of the nearest neighboring points is less than the threshold, the obtained prediction mode of the point is the prediction mode 0; The receiving apparatus of claim 9 , wherein the predicted quality value of the point for the prediction mode 0 is set as the distance-based weighted average of the nearest neighbor points.
12. 1. A method for processing point cloud data at a receiving device, comprising: receiving geometry information, attribute information and signaling information; decoding the geometry information based on the signaling information to recover a position of the point; decoding the attribute information based on the signaling information and the geometry information to recover attribute values of the points; rendering the reconstructed point cloud data based on the point locations and attribute values; The step of decoding the characteristic information comprises: obtaining a prediction mode for a point to be decoded, where if a maximum difference value between feature values of a plurality of nearest neighboring points of the point is equal to or greater than a threshold value, a prediction mode 0 is excluded from the obtained prediction modes, the prediction mode 0 being a mode based on a distance-based weighted average value of the plurality of nearest neighboring points, and a predicted feature value of the point is set as one feature value of the plurality of nearest neighboring points according to the obtained prediction mode; performing inverse quantization on residual feature values of said points in said feature information; and recovering the feature value of the point based on the predicted feature value of the point and the dequantized residual feature value of the point in the feature information.
13. the obtained prediction mode of the point is one of prediction mode 1, prediction mode 2, or prediction mode 3; the predicted feature value of the point for the prediction mode 1 is set as the feature value of a first nearest neighbor of the plurality of nearest neighbor points; the predicted feature value of the point for the prediction mode 2 is set as the feature value of a second nearest neighbor of the plurality of nearest neighbor points; The method of claim 12 , wherein the predicted feature value of the point for the prediction mode 3 is set as the feature value of a third nearest neighbor of the plurality of nearest neighbor points.
14. if the maximum difference value between the feature values of the nearest neighboring points is less than the threshold, the obtained prediction mode of the point is the prediction mode 0; The method of claim 12 , wherein the predicted quality value of the point for the prediction mode 0 is set as the distance-based weighted average of the nearest neighbor points.
Citation Information
Patent Citations
Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device
WO2021002443A1