Point cloud data transmission device, transmission method, processing device and processing method
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-26
- Publication Date
- 2026-08-14
Smart Images

Figure CN115210765B_ABST
Abstract
Description
Technical Field
[0001] The implementation provides point cloud content to offer services including virtual reality (VR), augmented reality (AR), mixed reality (MR), and autonomous driving services. Background Technology
[0002] Point cloud content is content represented by point clouds, which are a set of points belonging to a coordinate system representing three-dimensional space. Point cloud content can represent 3D configured media and is used to provide various services such as virtual reality (VR), augmented reality (AR), mixed reality (MR), and autonomous driving services. However, tens of thousands to hundreds of thousands of point data points are required to represent point cloud content. Therefore, methods for efficiently processing large amounts of point data are needed. Summary of the Invention
[0003] Technical issues
[0004] The embodiments provide an apparatus and method for efficiently processing point cloud data. The embodiments provide a point cloud data processing method and apparatus for addressing latency and encoding / decoding complexity.
[0005] The technical scope of the implementation is not limited to the above-described technical objectives, but can be extended to other technical objectives that can be inferred by those skilled in the art based on the full content disclosed herein.
[0006] Technical solution
[0007] To achieve these and other advantages, and in accordance with the purposes of this disclosure, in some embodiments, a method for transmitting point cloud data may include: encoding point cloud data comprising geometry and attributes, and transmitting a bitstream comprising the encoded point cloud data. In some embodiments, the geometry represents the location of points in the point cloud data, and the attributes include at least one of the point's color and reflectivity.
[0008] In some embodiments, an apparatus for transmitting point cloud data may include: an encoder configured to encode point cloud data including geometry and attributes; and a transmitter configured to transmit a bit stream including the encoded point cloud data.
[0009] In some embodiments, a method for processing point cloud data may include receiving a bitstream comprising the point cloud data and decoding the point cloud data. In some embodiments, the point cloud data includes geometry and attributes, the geometry representing the location of points in the point cloud data, and the attributes including at least one of the color and reflectivity of the points.
[0010] In some embodiments, an apparatus for processing point cloud data may include: a receiver for receiving a bitstream comprising the point cloud data; and a decoder for decoding the point cloud data. In some embodiments, the point cloud data includes geometry and attributes, the geometry representing the location of points in the point cloud data, and the attributes including at least one of the color and reflectivity of the points.
[0011] Beneficial effects
[0012] The apparatus and method according to the embodiments effectively process point cloud data.
[0013] The apparatus and method according to the embodiments provide point cloud services with high quality.
[0014] The apparatus and method according to the embodiments provide point cloud content to provide general services including VR services and autonomous driving. Attached Figure Description
[0015] The accompanying drawings are included to provide a further understanding of this disclosure and are incorporated in and constitute a part of this application. The drawings illustrate embodiments of the disclosure and, together with the description, serve to illustrate the principles of the disclosure. For a better understanding of the various embodiments described below, reference should be made to the following description of embodiments in conjunction with the accompanying drawings. The same reference numerals will be used throughout the drawings to denote the same or similar parts. In the drawings:
[0016] Figure 1 An exemplary point cloud content delivery system according to an implementation is shown.
[0017] Figure 2 This is a block diagram illustrating the operation of providing point cloud content according to an embodiment.
[0018] Figure 3 An exemplary process for capturing point cloud video according to an embodiment is shown.
[0019] Figure 4 An exemplary point cloud encoder according to an implementation is shown.
[0020] Figure 5 An example of a voxel according to an embodiment is shown.
[0021] Figure 6 An example of an octree and occupancy code according to an implementation is shown.
[0022] Figure 7 An example of an adjacent node pattern according to an implementation method is shown.
[0023] Figure 8 Examples of point configurations in various Levels of Detail (LODs) according to the implementation method are shown.
[0024] Figure 9 Examples of point configurations in various Levels of Detail (LODs) according to the implementation method are shown.
[0025] Figure 10 A point cloud decoder according to an embodiment is shown.
[0026] Figure 11 A point cloud decoder according to an embodiment is shown.
[0027] Figure 12 A transmitting device according to an embodiment is shown.
[0028] Figure 13 A receiving device according to an embodiment is shown.
[0029] Figure 14 An exemplary structure is shown that can be operated in conjunction with a point cloud data transmission / reception method / apparatus according to an embodiment.
[0030] Figure 15 This is a flowchart illustrating the operation of a point cloud transmitting device according to an embodiment.
[0031] Figure 16 This is a block diagram of a point cloud transmitting device according to an embodiment.
[0032] Figure 17 The distribution of points according to the LOD level is shown.
[0033] Figure 18 An example of a neighboring point set search method is shown.
[0034] Figure 19 An example of point cloud content is shown.
[0035] Figure 20 An example of an octree-based LOD generation process is shown.
[0036] Figure 21 An example of the baseline distance is shown.
[0037] Figure 22 An example of the maximum nearest neighbor distance is shown.
[0038] Figure 23 An example of the maximum nearest neighbor distance is shown.
[0039] Figure 24 An example of neighbor set search is shown.
[0040] Figure 25 An example of a bitstream structure diagram is shown.
[0041] Figure 26 This is an example of signaling information according to the implementation method.
[0042] Figure 27 An example of signaling information according to an implementation method is shown.
[0043] Figure 28 An example of signaling information according to an implementation method is shown.
[0044] Figure 29 An example of signaling information according to an implementation method is shown.
[0045] Figure 30 This is a flowchart illustrating the operation of a point cloud receiving device according to an embodiment.
[0046] Figure 31 This is a flowchart illustrating a point cloud data transmission method according to an embodiment.
[0047] Figure 32 This is a flowchart illustrating a method for processing point cloud data according to an embodiment. Detailed Implementation
[0048] Preferred embodiments of the present disclosure will now be described in detail, examples of which are illustrated in the accompanying drawings. The detailed description given below with reference to the drawings is intended to illustrate exemplary embodiments of the present disclosure, and not to show only embodiments that can be implemented according to the present disclosure. The following detailed description includes specific details in order to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure can be practiced without these specific details.
[0049] Although most of the terms used in this disclosure are selected from commonly used terms in the art, some terms have been arbitrarily chosen by the applicant and their meanings are explained in detail in the following description as needed. Therefore, this disclosure should be understood based on the intended meaning of the terms rather than their simple names or meanings.
[0050] Figure 1 An exemplary point cloud content delivery system according to an implementation is shown.
[0051] Figure 1 The point cloud content providing system shown may include a transmitting device 10000 and a receiving device 10004. The transmitting device 10000 and the receiving device 10004 are capable of transmitting and receiving point cloud data via wired or wireless communication.
[0052] The point cloud data transmission device 10000 according to an embodiment can acquire and process point cloud video (or point cloud content) and transmit it. According to an embodiment, the transmission device 10000 may include a fixed station, a base transceiver system (BTS), a network, an artificial intelligence (AI) device and / or system, a robot, an AR / VR / XR device, and / or a server. According to an embodiment, the transmission device 10000 may include devices configured to communicate with base stations and / or other wireless devices using radio access technologies (e.g., 5G New RAT (NR), Long Term Evolution (LTE)), robots, vehicles, AR / VR / XR devices, portable devices, home appliances, Internet of Things (IoT) devices, and AI devices / servers.
[0053] The transmitting device 10000 according to the embodiment includes a point cloud video acquirer 10001, a point cloud video encoder 10002 and / or a transmitter (or communication module) 10003.
[0054] The point cloud video acquirer 10001 according to an embodiment acquires point cloud video through processing procedures such as capture, synthesis, or generation. Point cloud video is point cloud content represented by a point cloud, which is a set of points located in 3D space, and may be referred to as point cloud video data. The point cloud video according to an embodiment may include one or more frames. A frame represents a still image / scene. Therefore, point cloud video may include point cloud images / frames / scenes, and may be referred to as point cloud images, frames, or scenes.
[0055] The point cloud video encoder 10002 according to an embodiment encodes the acquired point cloud video data. The point cloud video encoder 10002 can encode the point cloud video data based on point cloud compression coding. The point cloud compression coding according to an embodiment may include geometry-based point cloud compression (G-PCC) coding and / or video-based point cloud compression (V-PCC) coding or next-generation coding. The point cloud compression coding according to an embodiment is not limited to the above embodiments. The point cloud video encoder 10002 can output a bitstream containing encoded point cloud video data. The bitstream may contain not only the encoded point cloud video data but also signaling information related to the encoding of the point cloud video data.
[0056] According to an embodiment, transmitter 10003 transmits a bitstream containing encoded point cloud video data. The bitstream, according to an embodiment, is encapsulated in a file or segment (e.g., a streaming segment) and transmitted via various networks such as broadcast networks and / or broadband networks. Although not shown in the figures, transmitting device 10000 may include a wrapper (or wrapper module) configured to perform encapsulation operations. According to an embodiment, the wrapper may be included in transmitter 10003. According to an embodiment, the file or segment may be transmitted via a network to receiving device 10004 or stored in a digital storage medium (e.g., USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.). Transmitter 10003 according to an embodiment is capable of wired / wireless communication with receiving device 10004 (or receiver 10005) via networks such as 4G, 5G, and 6G. Additionally, the transmitter may perform necessary data processing operations depending on the network system (e.g., a 4G, 5G, or 6G communication network system). Transmitting device 10000 can transmit encapsulated data on demand.
[0057] The receiving device 10004 according to an embodiment includes a receiver 10005, a point cloud video decoder 10006, and / or a renderer 10007. According to an embodiment, the receiving device 10004 may include devices, robots, vehicles, AR / VR / XR devices, portable devices, home appliances, Internet of Things (IoT) devices, and AI devices / servers configured to communicate with base stations and / or other wireless devices using radio access technologies (e.g., 5G New RAT (NR), Long Term Evolution (LTE)).
[0058] According to an embodiment, receiver 10005 receives a bitstream containing point cloud video data or a file / segment encapsulated with a bitstream from a network or storage medium. Receiver 10005 may perform necessary data processing according to the network system (e.g., a communication network system such as 4G, 5G, 6G, etc.). According to an embodiment, receiver 10005 may decapsulate the received file / segment and output a bitstream. According to an embodiment, receiver 10005 may include a decapsulator (or decapsulator module) configured to perform a decapsulation operation. The decapsulator may be implemented as a separate element (or component) from receiver 10005.
[0059] The point cloud video decoder 10006 decodes the bitstream containing point cloud video data. The point cloud video decoder 10006 can decode the point cloud video data according to the method in which the point cloud video data is encoded (e.g., the inverse process of the operation of the point cloud video encoder 10002). Therefore, the point cloud video decoder 10006 can decode the point cloud video data by performing point cloud decompression encoding (the inverse process of point cloud compression). Point cloud decompression encoding includes G-PCC encoding.
[0060] Renderer 10007 renders decoded point cloud video data. Renderer 10007 can output point cloud content by rendering not only the point cloud video data but also the audio data. According to one embodiment, renderer 10007 may include a display configured to display the point cloud content. According to another embodiment, the display may be implemented as a separate device or component rather than included in renderer 10007.
[0061] The arrows indicated by dashed lines in the diagram represent the transmission paths of the feedback information acquired by the receiving device 10004. The feedback information reflects the interactivity of the user consuming the point cloud content and includes information about the user (e.g., head orientation information, viewport information, etc.). Specifically, when the point cloud content is for a service requiring user interaction (e.g., autonomous driving services, etc.), the feedback information may be provided to the content sender (e.g., the sending device 10000) and / or the service provider. Depending on the implementation, the feedback information may be used in both the receiving device 10004 and the sending device 10000, or it may not be provided.
[0062] According to the embodiment, head orientation information is information about the user's head position, orientation, angle, movement, etc. The receiving device 10004 according to the embodiment can calculate viewport information based on the head orientation information. The viewport information can be information about the area of the point cloud video that the user is viewing. The viewpoint is the point through which the user views the point cloud video, and can refer to the center point of the viewport area. That is, the viewport is the area centered on the viewpoint, and the size and shape of the area can be determined by the field of view (FOV). Therefore, in addition to head orientation information, the receiving device 10004 can also extract viewport information based on the vertical or horizontal FOV supported by the device. Furthermore, the receiving device 10004 performs gaze analysis, etc., to examine the way the user consumes the point cloud, the area the user gazes at in the point cloud video, the gaze duration, etc. According to the embodiment, the receiving device 10004 can send feedback information including the gaze analysis results to the transmitting device 10000. The feedback information according to the embodiment can be acquired during rendering and / or display. The feedback information according to the embodiment can be acquired by one or more sensors included in the receiving device 10004. According to the implementation method, the feedback information can be obtained by the renderer 10007 or by a separate external component (or device, component, etc.). Figure 1The dashed lines in the diagram represent the process of sending feedback information obtained by the renderer 10007. The point cloud content providing system can process (encode / decode) point cloud data based on the feedback information. Therefore, the point cloud video decoder 10006 can perform decoding operations based on the feedback information. The receiving device 10004 can send the feedback information to the transmitting device 10000. The transmitting device 10000 (or the point cloud video data encoder 10002) can perform encoding operations based on the feedback information. Therefore, the point cloud content providing system can effectively process necessary data (e.g., point cloud data corresponding to the user's head position) based on the feedback information instead of processing (encoding / decoding) the entire point cloud data, and provide the point cloud content to the user.
[0063] According to the implementation, the transmitting device 10000 may be referred to as an encoder, transmitting device, transmitter, etc., and the receiving device 10004 may be referred to as a decoder, receiving device, receiver, etc.
[0064] According to the implementation method Figure 1 Point cloud data processed in a point cloud content provision system (through a series of processes including acquisition, encoding, transmission, decoding, and rendering) can be referred to as point cloud content data or point cloud video data. Depending on the implementation, point cloud content data can be used as a concept encompassing metadata or signaling information related to point cloud data.
[0065] Figure 1 The components of the point cloud content provided by the system can be implemented by hardware, software, processors, and / or combinations thereof.
[0066] = This is a block diagram illustrating the operation of providing point cloud content according to an embodiment.
[0067] Figure 2 The block diagram shows Figure 2 The operation of the point cloud content providing system described herein. As mentioned above, the point cloud content providing system can process point cloud data based on point cloud compression encoding (e.g., G-PCC).
[0068] A point cloud content providing system according to an embodiment (e.g., point cloud sending device 10000 or point cloud video acquirer 10001) can acquire point cloud video (20000). The point cloud video is represented by a point cloud belonging to a coordinate system used to represent 3D space. The point cloud video according to an embodiment may include Ply (Polygon file format or Stanford Triangle format) files. When the point cloud video has one or more frames, the acquired point cloud video may include one or more Ply files. The Ply file contains point cloud data such as point geometry and / or attributes. Geometry includes the position of the points. The position of each point may be represented by parameters (e.g., values of the X, Y, and Z axes) representing a three-dimensional coordinate system (e.g., a coordinate system consisting of X, Y, and Z axes). Attributes include the attributes of the points (e.g., information about the texture, color (YCbCr or RGB), reflectivity r, transparency, etc., of each point). A point has one or more attributes. For example, a point may have a color attribute or two attributes: color and reflectivity. According to the implementation, geometry can be referred to as location, geometric information, geometric data, etc., and attributes can be referred to as attributes, attribute information, attribute data, etc. A point cloud content providing system (e.g., point cloud sending device 10000 or point cloud video acquirer 10001) can acquire point cloud data from information related to the point cloud video acquisition process (e.g., depth information, color information, etc.).
[0069] A point cloud content providing system (e.g., a transmitting device 10000 or a point cloud video encoder 10002) according to an embodiment can encode point cloud data (20001). The point cloud content providing system can encode point cloud data based on point cloud compression encoding. As described above, point cloud data can include the geometry and attributes of points. Therefore, the point cloud content providing system can perform geometry encoding to encode the geometry and output a geometry bitstream. The point cloud content providing system can perform attribute encoding to encode the attributes and output an attribute bitstream. According to an embodiment, the point cloud content providing system can perform attribute encoding based on geometry encoding. The geometry bitstream and attribute bitstream according to an embodiment can be multiplexed and output as a single bitstream. The bitstream according to an embodiment may also contain signaling information related to geometry encoding and attribute encoding.
[0070] A point cloud content providing system according to an embodiment (e.g., transmitting device 10000 or transmitter 10003) can transmit encoded point cloud data (20002). For example... Figure 1 As shown, encoded point cloud data can be represented by geometric bitstreams and attribute bitstreams. Additionally, the encoded point cloud data can be transmitted as a bitstream along with signaling information related to the encoding of the point cloud data (e.g., signaling information related to geometric encoding and attribute encoding). The point cloud content providing system can encapsulate the bitstream carrying the encoded point cloud data and transmit it as a file or fragment.
[0071] The point cloud content providing system (e.g., receiving device 10004 or receiver 10005) according to the embodiment can receive a bitstream containing encoded point cloud data. Additionally, the point cloud content providing system (e.g., receiving device 10004 or receiver 10005) can demultiplex the bitstream.
[0072] A point cloud content providing system (e.g., receiving device 10004 or point cloud video decoder 10005) can decode encoded point cloud data (e.g., geometric bitstream, attribute bitstream) transmitted in a bitstream. The point cloud content providing system (e.g., receiving device 10004 or point cloud video decoder 10005) can decode point cloud video data based on signaling information related to the encoding of the point cloud video data contained in the bitstream. The point cloud content providing system (e.g., receiving device 10004 or point cloud video decoder 10005) can decode the geometric bitstream to reconstruct the location (geometry) of the points. The point cloud content providing system can reconstruct the attributes of the points by decoding the attribute bitstream based on the reconstructed geometry. The point cloud content providing system (e.g., receiving device 10004 or point cloud video decoder 10005) can reconstruct point cloud video based on location according to the reconstructed geometry and the decoded attributes.
[0073] A point cloud content providing system (e.g., receiving device 10004 or renderer 10007) according to an embodiment can render decoded point cloud data (20004). The point cloud content providing system (e.g., receiving device 10004 or renderer 10007) can use various rendering methods to render the geometry and attributes decoded through the decoding process. Points in the point cloud content can be rendered as vertices with a specific thickness, cubes with a specific minimum size centered at the corresponding vertex position, or circles centered at the corresponding vertex position. All or part of the rendered point cloud content is provided to the user through a display (e.g., a VR / AR display, a general display, etc.).
[0074] The point cloud content providing system (e.g., receiving device 10004) according to the embodiment can acquire feedback information (20005). The point cloud content providing system can encode and / or decode point cloud data based on the feedback information. The feedback information and operation / reference of the point cloud content providing system according to the embodiment... Figure 1 The feedback information and operation described are the same, so their detailed description is omitted.
[0075] Figure 1 An exemplary process for capturing point cloud video according to an implementation method is shown.
[0076] Figure 3 Show reference Figure 3 The described point cloud content provides an exemplary point cloud video capture process for the system.
[0077] Point cloud content includes point cloud videos (images and / or videos) representing objects and / or environments located in various 3D spaces (e.g., 3D spaces representing real environments, 3D spaces representing virtual environments, etc.). Therefore, the point cloud content providing system according to embodiments can use one or more cameras (e.g., infrared cameras capable of acquiring depth information, RGB cameras capable of extracting color information corresponding to the depth information, etc.), projectors (e.g., infrared pattern projectors acquiring depth information), LiDAR, etc., to capture point cloud videos. The point cloud content providing system according to embodiments can extract geometry composed of points in 3D space from the depth information and extract attributes of each point from the color information to obtain point cloud data. Images and / or videos according to embodiments can be captured based on at least one of inward-oriented and outward-oriented techniques.
[0078] Figures 1 to 2 The left side illustrates inward-facing technology. Inward-facing technology refers to the technique of using one or more cameras (or camera sensors) positioned around a central object to capture an image of that object. Inward-facing technology can be used to generate point cloud content that provides users with 360-degree images of key objects (e.g., VR / AR content that provides users with 360-degree images of objects such as characters, players, objects, or actors).
[0079] Figure 3 The right side illustrates outward-facing techniques. Outward-facing techniques refer to techniques that use one or more cameras (or camera sensors) positioned around a central object to capture images of the environment surrounding the central object, rather than the central object itself. Outward-facing techniques can be used to generate point cloud content that provides a representation of the surrounding environment from the user's perspective (e.g., content representing the external environment that can be provided to users of self-driving vehicles).
[0080] As shown in the figure, point cloud content can be generated based on the capture operations of one or more cameras. In this case, the coordinate systems may differ between the cameras, so the point cloud content providing system can calibrate one or more cameras to set the global coordinate system before the capture operation. Alternatively, the point cloud content providing system can generate point cloud content by compositing arbitrary images and / or videos with images and / or videos captured using the aforementioned capture techniques. The point cloud content providing system may not perform the following when generating point cloud content representing virtual space: Figure 3 The capture operations described herein. The point cloud content providing system according to an embodiment can perform post-processing on the captured images and / or videos. In other words, the point cloud content providing system can remove unwanted areas (e.g., background), identify spaces to which the captured images and / or videos are connected, and perform a space-filling operation when spatial holes exist.
[0081] A point cloud content delivery system can generate point cloud content by performing coordinate transformations on points in point cloud video captured from various cameras. The system can perform these transformations based on the position coordinates of each camera. Therefore, the system can generate content representing a wide range or point cloud content with high point density.
[0082] Figure 3 An exemplary point cloud encoder according to an implementation is shown.
[0083] Figure 4 Show Figure 4 An example of a point cloud video encoder 10002. The point cloud encoder reconstructs and encodes point cloud data (e.g., point locations and / or attributes) to adjust the quality of the point cloud content (e.g., lossless, lossy, or near-lossless) based on network conditions or applications. When the total size of the point cloud content is large (e.g., 60Gbps of point cloud content for 30fps), the point cloud content providing system may not be able to stream the content in real time. Therefore, the point cloud content providing system can reconstruct the point cloud content based on a maximum target bitrate to provide the point cloud content according to network conditions, etc.
[0084] For reference Figure 1 The point cloud encoder described herein can perform geometric encoding and attribute encoding. Geometric encoding is performed before attribute encoding.
[0085] The point cloud encoder according to the implementation includes a coordinate transformer (transform coordinates) 40000, a quantizer (quantize and remove points (voxarization)) 40001, an octree analyzer (analyze octrees) 40002, a surface approximation analyzer (analyze surface approximations) 40003, an arithmetic encoder (arithmetic encoding) 40004, a geometry reconstructor (reconstruct geometry) 40005, a color transformer (transform colors) 40006, an attribute transformer (transform attributes) 40007, a RAHT transformer (RAHT) 40008, an LOD generator (generate LODs) 40009, a lift transformer (lift) 40010, a coefficient quantizer (quantize coefficients) 40011, and / or an arithmetic encoder (arithmetic encoding) 40012.
[0086] The coordinate transformer 40000, quantizer 40001, octree analyzer 40002, surface approximation analyzer 40003, arithmetic encoder 40004, and geometric reconstructor 40005 are capable of performing geometric encoding. Geometric encoding according to the implementation may include octree geometric encoding, direct encoding, triplet geometric encoding, and entropy encoding. Direct encoding and triplet geometric encoding are applied selectively or in combination. Geometric encoding is not limited to the examples described above.
[0087] As shown in the figure, the coordinate transformer 40000 according to the embodiment receives the position and transforms it into coordinates. For example, the position can be transformed into position information in three-dimensional space (e.g., three-dimensional space represented by the XYZ coordinate system). The position information in three-dimensional space according to the embodiment can be referred to as geometric information.
[0088] The quantizer 40001 according to the embodiment quantizes geometry. For example, the quantizer 40001 may quantize points based on the minimum position value of all points (e.g., the minimum value on each of the X, Y, and Z axes). The quantizer 40001 performs a quantization operation: multiplying the difference between the minimum position value and the position value of each point by a preset quantization scale value, and then finding the nearest integer value by rounding the value obtained by multiplication. Thus, one or more points may have the same quantized position (or position value). The quantizer 40001 according to the embodiment performs voxelization based on the quantized position to reconstruct the quantized points. As in the case of pixels (the smallest unit containing 2D image / video information), points in the point cloud content (or 3D point cloud video) according to the embodiment may be included in one or more voxels. As a combination of volume and pixel, the term voxel refers to the 3D cubic space generated when 3D space is divided into units (unit = 1.0) based on axes representing 3D space (e.g., X-axis, Y-axis, and Z-axis). The quantizer 40001 allows a group of points in 3D space to be matched with voxels. In one embodiment, a voxel may include only one point. In another embodiment, a voxel may include one or more points. To represent a voxel as a point, the location of the voxel's center can be set based on the locations of one or more points included in the voxel. In this case, attributes of all locations included in a voxel can be combined and assigned to the voxel.
[0089] The octree analyzer 40002 according to the implementation performs octree geometric encoding (or octree coding) to represent voxels in an octree structure. The octree structure represents points based on the matching of octree structures with voxels.
[0090] The surface approximation analyzer 40003 according to the embodiment can analyze and approximate an octree. The octree analysis and approximation according to the embodiment is a process of analyzing a region containing multiple points to efficiently provide an octree and voxelization.
[0091] The arithmetic encoder 40004 according to the implementation performs entropy encoding on octrees and / or approximate octrees. For example, the encoding scheme includes arithmetic encoding. As a result of the encoding, a geometric bitstream is generated.
[0092] The attribute encoding is performed by a color transformer 40006, an attribute transformer 40007, a RAHT transformer 40008, an LOD generator 40009, a boosting transformer 40010, a coefficient quantizer 40011, and / or an arithmetic encoder 40012. As described above, a point may have one or more attributes. The attribute encoding according to the embodiments is also applied to the attributes possessed by a point. However, when an attribute (e.g., color) comprises one or more elements, attribute encoding is applied independently to each element. The attribute encoding according to the embodiments includes color transformation encoding, attribute transformation encoding, region adaptive hierarchical transformation (RAHT) encoding, interpolation-based hierarchical nearest neighbor prediction (prediction transformation) encoding, and interpolation-based hierarchical nearest neighbor prediction (boosting transformation) encoding with update / boosting steps. Depending on the point cloud content, the above-described RAHT encoding, prediction transformation encoding, and boosting transformation encoding may be used selectively, or a combination of one or more encoding schemes may be used. The attribute encoding according to the embodiments is not limited to the examples described above.
[0093] The color converter 40006 according to the embodiment performs color transformation encoding that transforms the color values (or textures) included in the attributes. For example, the color converter 40006 can transform the format of color information (e.g., from RGB to YCbCr). Optionally, the operation of the color converter 40006 according to the embodiment can be applied based on the color values included in the attributes.
[0094] The geometry reconstructor 40005, according to the implementation method, reconstructs (decompresses) octrees and / or approximate octrees. The geometry reconstructor 40005 reconstructs the octree / voxel based on the results of analyzing the point distribution. The reconstructed octree / voxel can be referred to as the reconstructed geometry (recovered geometry).
[0095] The attribute transformer 40007 according to the embodiment performs attribute transformation to transform attributes based on reconstructed geometry and / or positions without performing geometric encoding. As described above, since attributes depend on geometry, the attribute transformer 40007 can transform attributes based on reconstructed geometric information. For example, based on the position value of a point included in a voxel, the attribute transformer 40007 can transform the attributes of the point at that position. As described above, when the center position of a voxel is set based on the positions of one or more points included in the voxel, the attribute transformer 40007 transforms the attributes of one or more points. When performing triadic geometric encoding, the attribute transformer 40007 can transform attributes based on the triadic geometric encoding.
[0096] The attribute transformer 40007 performs attribute transformation by calculating the average of the attributes or attribute values (e.g., the color or reflectivity of each point) of neighboring points within a specific location / radius from the center (or position value) of each voxel. The attribute transformer 40007 can apply weights based on the distance from the center to each point when calculating the average. Therefore, each voxel has a position and a calculated attribute (or attribute value).
[0097] The attribute transformer 40007 can search for neighboring points within a specific location / radius of the center of each voxel based on a KD-tree or Morton code. A KD-tree is a binary search tree and supports a data structure that allows points to be managed based on location, enabling fast nearest neighbor search (NNS). Morton codes are generated by representing the coordinates (e.g., (x, y, z)) of the 3D location of all points as bit values and mixing the bits. For example, when the coordinates representing the point location are (5, 9, 1), the bit values are (0101, 1001, 0001). Mixing the bit values according to the bit index in the order of z, y, and x produces 010001000111. This value is represented as the decimal number 1095. That is, the Morton code value for the point with coordinates (5, 9, 1) is 1095. The attribute transformer 40007 can sort the points based on their Morton code values and perform NNS using a depth-first traversal process. After an attribute transformation operation, use a KD tree or Morton code when an NNS is needed in another transformation process used for attribute encoding.
[0098] As shown in the figure, the transformation properties are input to the RAHT transformer 40008 and / or the LOD generator 40009.
[0099] According to the implementation, the RAHT transformer 40008 performs RAHT encoding for predicting attribute information based on the reconstructed geometric information. For example, the RAHT transformer 40008 can predict the attribute information of higher-level nodes in an octree based on the attribute information associated with lower-level nodes in the octree.
[0100] The LOD generator 40009 according to the embodiment generates a Level of Detail (LOD) to perform predictive transform coding. The LOD according to the embodiment represents the level of detail of the point cloud content. As the LOD value decreases, it indicates a deterioration in the detail of the point cloud content. As the LOD value increases, it indicates an enhancement in the detail of the point cloud content. Points can be classified according to LOD.
[0101] The lift transformer 40010 according to the implementation performs lift transform coding to transform point cloud attributes based on weights. As described above, lift transform coding may optionally be applied.
[0102] According to the implementation method, the coefficient quantizer 40011 quantizes the attribute encoding based on the coefficient.
[0103] According to the implementation method, the arithmetic encoder 40012 encodes the quantized attributes based on arithmetic encoding.
[0104] Although not shown in the figure, Figures 1 to 2 The elements of the point cloud encoder can be implemented by hardware, software, firmware, or a combination thereof, including one or more processors or integrated circuits configured to communicate with one or more memories included in the point cloud providing device. One or more processors can perform the above-described... Figure 4 At least one of the operation and / or functions of the elements of the point cloud encoder. Additionally, one or more processors are operable or perform operations for executing... Figure 4 The software program and / or instruction set for the operation and / or function of the elements of the point cloud encoder. One or more memories according to the embodiment may include high-speed random access memory, or may include non-volatile memory (e.g., one or more disk storage devices, flash memory devices or other non-volatile solid-state memory devices).
[0105] Figure 4 An example of a voxel according to an embodiment is shown.
[0106] Figure 5 This shows voxels in 3D space represented by a coordinate system consisting of three axes (X, Y, and Z). (See reference...) Figure 5 The point cloud encoder (e.g., quantizer 40001) can perform voxelization. A voxel is a 3D cubic space generated when the 3D space is divided into cells (unit = 1.0) based on axes representing the 3D space (e.g., X-axis, Y-axis, and Z-axis). Figure 4 An example of voxels generated via an octree structure is shown, where a cubic axis-aligned bounding box defined by two poles (0,0,0) and (2d,2d,2d) is recursively subdivided. A voxel comprises at least one point. The spatial coordinates of a voxel can be estimated from its positional relationship to a group of voxels. As mentioned above, voxels possess properties similar to pixels in a 2D image / video (e.g., color or reflectivity). Details of the voxels and references are shown. Figure 5 The descriptions are the same, so their descriptions are omitted.
[0107] Figure 4 An example of an octree and occupancy code according to an implementation is shown.
[0108] For reference Figure 6As described, the point cloud content delivery system (point cloud video encoder 10002) or point cloud encoder (e.g., octree analyzer 40002) performs octree geometric encoding (or octree encoding) based on an octree structure to efficiently manage the regions and / or locations of voxels.
[0109] Figures 1 to 4 The upper part shows an octree structure. The 3D space of the point cloud content according to the embodiment is represented by the axes of a coordinate system (e.g., the X, Y, and Z axes). This is achieved by using two poles (0,0,0) and (2... d ,2 d ,2 d An octree structure is created by recursively subdividing the bounding box with the cubic axis aligned to the bounding box. Here, 2 d This can be set to the value of the minimum bounding box that constitutes all points surrounding the point cloud content (or point cloud video). Here, d represents the depth of the octree. The value of d is determined in the following formula. In the following formula, (x int n ,y int n ,z int n ) indicates the position (or position value) of the quantized point.
[0110] d=Ceil(Log2(Max(x_n∧int, y_n∧int, z_n∧int, n=1,...,N)+1))
[0111] like Figure 6 As shown in the upper center, the entire 3D space can be divided into eight spaces. Each divided space is represented by a cube with six faces. (See diagram below.) Figure 6 As shown in the upper right, the eight spaces are further subdivided based on coordinate system axes (e.g., the X, Y, and Z axes). Thus, each space is divided into eight smaller spaces. These smaller spaces are also represented by cubes with six faces. This partitioning scheme is applied until the leaf nodes of the octree become voxels.
[0112] Figure 6 The lower part shows the octree occupancy code. The occupancy code generates the octree to indicate whether each of the eight partitions generated by dividing a space contains at least one point. Therefore, a single occupancy code is represented by eight child nodes. Each child node represents the occupancy of a partitioned space, and each child node has a 1-bit value. Therefore, the occupancy code is represented as an 8-bit code. That is, when the space corresponding to a child node contains at least one point, the node is assigned a value of 1. When the space corresponding to a child node does not contain a point (the space is empty), the node is assigned a value of 0. Since... Figure 6The occupancy code shown is 00100001, indicating that the spaces corresponding to the third and eighth child nodes among the eight child nodes each contain at least one point. As shown, each of the third and eighth child nodes has eight child nodes, and each child node is represented by an 8-bit occupancy code. The attached figure shows the occupancy code for the third child node as 10000111, and the occupancy code for the eighth child node as 01001111. A point cloud encoder (e.g., an arithmetic encoder 40004) according to an embodiment can perform entropy encoding on the occupancy code. To increase compression efficiency, the point cloud encoder can perform intra / inter-frame encoding on the occupancy code. A receiving device (e.g., receiving device 10004 or point cloud video decoder 10006) according to an embodiment reconstructs the octree based on the occupancy code.
[0113] A point cloud encoder according to an implementation method (e.g., Figure 6 A point cloud encoder or octree analyzer (40002) can perform voxelization and octree encoding to store point locations. However, points are not always uniformly distributed in 3D space, so there may be specific regions with fewer points. Therefore, performing voxelization over the entire 3D space is inefficient. For example, when a specific region contains very few points, voxelization is not necessary in that specific region.
[0114] Therefore, for the aforementioned specific region (or nodes other than the leaf nodes of the octree), the point cloud encoder according to the embodiment can skip voxelization and perform direct encoding to directly encode the point positions included in the specific region. The coordinates of the directly encoded points according to the embodiment are called the Direct Encoding Mode (DCM). The point cloud encoder according to the embodiment can also perform triadic geometry encoding based on the surface model, which reconstructs the point positions in the specific region (or node) based on voxels. Triadic geometry encoding is a geometric encoding that represents an object as a series of triangular meshes. Therefore, the point cloud decoder can generate a point cloud from the mesh surface. Direct encoding and triadic geometry encoding according to the embodiment can be performed selectively. In addition, direct encoding and triadic geometry encoding according to the embodiment can be performed in combination with octree geometry encoding (or octree encoding).
[0115] To perform direct encoding, the option to apply direct encoding using direct mode should be enabled. The node to which direct encoding is applied must not be a leaf node, and there should be fewer than a threshold number of points within that node. Furthermore, the total number of points to which direct encoding is applied should not exceed a preset threshold. When the above conditions are met, the point cloud encoder (or arithmetic encoder 40004) according to the implementation method can perform entropy encoding on the point locations (or location values).
[0116] A point cloud encoder according to an embodiment (e.g., a surface approximation analyzer 40003) can determine a specific level of the octree (a level less than the depth d of the octree) and can start using a surface model from that level to perform triadic geometry encoding to reconstruct the point positions in the node region based on voxels (triadic mode). The point cloud encoder according to an embodiment can specify the level to which triadic geometry encoding is to be applied. For example, when the specific level is equal to the depth of the octree, the point cloud encoder does not operate in triadic mode. In other words, the point cloud encoder according to an embodiment can operate in triadic mode only when the specified level is less than the depth value of the octree. The 3D cubic region of a node at a specified level according to an embodiment is called a block. A block may include one or more voxels. A block or voxel may correspond to a cube. Geometry is represented as surfaces within each block. A surface according to an embodiment may intersect each edge of a block at most once.
[0117] A block has 12 edges, therefore a block contains at least 12 intersections. Each intersection is called a vertex. Vertices along an edge are detected when there is at least one occupied voxel adjacent to the edge in all blocks sharing the edge. An occupied voxel, according to the implementation, refers to a voxel containing a point. The vertex position detected along an edge is the average position of the edges of all voxels adjacent to the edge in all blocks sharing the edge.
[0118] Once a vertex is detected, the point cloud encoder according to the implementation can perform entropy encoding on the edge's origin (x, y, z), the edge's direction vector (Δx, Δy, Δz), and the vertex position value (relative position value within the edge). When applying triad geometry encoding, the point cloud encoder according to the implementation (e.g., geometry reconstructor 40005) can generate the restored geometry (reconstructed geometry) by performing triangle reconstruction, upsampling, and voxelization processes.
[0119] Vertices located at the edges of a block determine the surface passing through the block. The surface, according to the implementation, is a non-planar polygon. During triangle reconstruction, the surface represented by triangles is reconstructed based on the origin of the edges, the direction vectors of the edges, and the position values of the vertices. The triangle reconstruction process is performed as follows: i) calculating the centroid value of each vertex, ii) subtracting the centroid value from each vertex value, and iii) estimating the sum of squares of the values obtained through the subtraction.
[0120] ① ② ③
[0121] Estimate the minimum value of the sum and perform a projection process based on the axis with the minimum value. For example, when element x is minimum, each vertex is projected onto the x-axis relative to the center of the block, and the projection is onto the (y,z) plane. When the value obtained by the projection onto the (y,z) plane is (ai,bi), the value of θ is estimated by atan2(bi,ai), and the vertices are sorted based on the value of θ. The following shows the vertex combinations for creating triangles based on the number of vertices. Vertices are sorted from 1 to n. The following shows that for 4 vertices, two triangles can be constructed based on the vertex combinations. The first triangle can be composed of vertices 1, 2, and 3 from the sorted vertices, and the second triangle can be composed of vertices 3, 4, and 1 from the sorted vertices.
[0122] Table 2-1. Triangles formed from vertices sorted by 1, ..., n
[0123] n triangles
[0124] 3(1,2,3)
[0125] 4(1,2,3),(3,4,1)
[0126] 5(1,2,3),(3,4,5),(5,1,3)
[0127] 6(1,2,3),(3,4,5),(5,6,1),(1,3,5)
[0128] 7(1,2,3),(3,4,5),(5,6,7),(7,1,3),(3,5,7)
[0129] 8(1,2,3),(3,4,5),(5,6,7),(7,8,1),(1,3,5),(5,7,1)
[0130] 9(1,2,3),(3,4,5),(5,6,7),(7,8,9),(9,1,3),(3,5,7),(7,9,3)
[0131] 10(1,2,3),(3,4,5),(5,6,7),(7,8,9),(9,10,1),(1,3,5),(5,7,9),(9,1,5)
[0132] 11(1,2,3),(3,4,5),(5,6,7),(7,8,9),(9,10,11),(11,1,3),(3,5,7),(7,9,11),(11,3,7)
[0133] 12(1,2,3),(3,4,5),(5,6,7),(7,8,9),(9,10,11),(11,12,1),(1,3,5),(5,7,9),(9,11,1),(1,5,9)
[0134] An upsampling process is performed to add points along the edges of the triangle at the center, and voxelization is then performed. The added points are generated based on the upsampling factor and the width of the block. These added points are called refined vertices. A point cloud encoder according to an implementation can voxelize the refined vertices. Additionally, the point cloud encoder can perform attribute encoding based on the voxelized positions (or position values).
[0135] Figure 4 An example of an adjacent node pattern according to an implementation method is shown.
[0136] To increase the compression efficiency of point cloud videos, the point cloud encoder according to the implementation method can perform entropy coding based on context-adaptive arithmetic coding.
[0137] For reference Figure 7 As described, a point cloud content providing system or point cloud encoder (e.g., point cloud video encoder 10002, Figures The point cloud encoder or arithmetic encoder (40004) can immediately perform entropy coding on the occupancy code. Alternatively, the point cloud content providing system or point cloud encoder can perform entropy coding (intra-frame coding) based on the occupancy code of the current node and the occupancy of neighboring nodes, or entropy coding (inter-frame coding) based on the occupancy code of a previous frame. According to the embodiment, a frame represents a collection of simultaneously generated point cloud videos. The compression efficiency of the intra-frame coding / inter-frame coding according to the embodiment can depend on the number of neighboring nodes referenced. As the number of bits increases, the computation becomes more complex, but the coding can be biased to one side, which can increase compression efficiency. For example, when given a 3-bit context, 2... 3 There are 8 methods to perform the encoding. The division of the encoding affects the implementation complexity. Therefore, it is necessary to achieve an appropriate level of compression efficiency and complexity.
[0138] This illustrates the process of obtaining an occupancy pattern based on the occupancy of neighboring nodes. A point cloud encoder according to an embodiment determines the occupancy of neighboring nodes of each node in an octree and obtains the value of the neighbor pattern. The neighboring node pattern is used to infer the occupancy pattern of a node. The left side of the diagram shows the cube corresponding to the node (the cube in the middle) and six cubes (adjacent nodes) that share at least one face with this cube. The nodes shown in the diagram are nodes at the same depth. The numbers shown in the diagram represent the weights associated with the six nodes (1, 2, 4, 8, 16, and 32). Weights are assigned sequentially based on the position of adjacent nodes.
[0139] The right side shows the neighboring node pattern values. The neighboring node pattern value is the sum of values multiplied by the weights of the occupying neighboring nodes (neighboring nodes with points). Therefore, the neighboring node pattern values range from 0 to 63. When the neighboring node pattern value is 0, it indicates that none of the node's neighboring nodes have points (no occupying node). When the neighboring node pattern value is 63, it indicates that all neighboring nodes are occupying nodes. As shown, since the neighboring nodes assigned weights 1, 2, 4, and 8 are occupying nodes, the neighboring node pattern value is 15 (the sum of 1, 2, 4, and 8). The point cloud encoder can perform encoding based on the neighboring node pattern values (e.g., when the neighboring node pattern value is 63, 64 types of encoding can be performed). According to implementations, the point cloud encoder can reduce encoding complexity by changing the neighboring node pattern values (e.g., based on a table that changes 64 to 10 or 6).
[0140] Examples of point configurations in various Levels of Detail (LODs) according to the implementation method are shown.
[0141] For reference The description describes the geometric reconstruction (decompression) of the encoded geometry before performing attribute encoding. When direct encoding is applied, the geometric reconstruction operation may include changing the placement of the directly encoded points (e.g., placing the directly encoded points in front of the point cloud data). When triadic geometric encoding is applied, the geometric reconstruction process is performed through triangle reconstruction, upsampling, and voxelization. Since attributes depend on geometry, attribute encoding is performed based on the reconstructed geometry.
[0142] A point cloud encoder (e.g., LOD generator 40009) can classify (reorganize) points according to LOD. The figure shows the point cloud content corresponding to LOD. The leftmost view in the figure represents the original point cloud content. The second view from the left in the figure represents the point distribution in the lowest LOD, and the rightmost view represents the point distribution in the highest LOD. That is, points are sparsely distributed in the lowest LOD and densely distributed in the highest LOD. In other words, as the LOD increases in the direction indicated by the arrow at the bottom of the figure, the space (or distance) between points narrows.
[0143] An example of point configuration for each LOD is shown according to an implementation method.
[0144] For reference As described, a point cloud content providing system or point cloud encoder (e.g., point cloud video encoder 10002, A point cloud encoder or LOD generator (40009) can generate LODs. LODs are generated by reorganizing points into a set of refined levels based on a set of LOD distance values (or a set of Euclidean distances). The LOD generation process is performed not only by the point cloud encoder but also by the point cloud decoder.
[0145] The upper part shows examples of points (P0 to P9) of point cloud content distributed in 3D space. In this context, the original order represents the order of points P0 to P9 before LOD generation. In this context, LOD-based order represents the order in which points are generated according to their LOD. Points are reorganized by LOD. Additionally, higher LODs include points belonging to lower LODs. For example... As shown, LOD 0 contains P0, P5, P4, and P2. LOD1 contains the points of LOD 0, P1, P6, and P3. LOD2 contains the points of LOD 0, the points of LOD1, P9, P8, and P7.
[0146] For reference As described, the point cloud encoder according to the implementation can selectively or in combination perform predictive transform coding, lifting transform coding, and RAHT transform coding.
[0147] The point cloud encoder according to the embodiment can generate predictors for points to perform predictive transform coding for setting the predicted attributes (or predicted attribute values) of each point. That is, N predictors can be generated for N points. The predictors according to the embodiment can calculate weights (= 1 / distance) based on the LOD value of each point, index information of neighboring points existing within a set distance of each LOD, and the distance to the neighboring points.
[0148] According to the implementation, the predicted attribute (or attribute value) is set as the average of the values obtained by multiplying the attributes (or attribute values) of neighboring points (e.g., color, reflectivity, etc.) set in the predictor of each point by a weight (or weight value) calculated based on the distance to each neighboring point. The point cloud encoder (e.g., coefficient quantizer 40011) according to the implementation can quantize and inverse quantize the residuals (which may be referred to as residual attributes, residual attribute values, or attribute prediction residuals) obtained by subtracting the predicted attribute (attribute value) from the attributes (attribute values) of each point. The quantization process is configured as shown in the table below.
[0149] Table. Pseudocode for Attribute Prediction Residual Quantization
[0150] int PCCQuantization(int value,int quantStep){
[0151] if(value>=0){
[0152] return floor(value / quantStep+1.0 / 3.0);
[0153] }else{
[0154] return-floor(-value / quantStep+1.0 / 3.0);
[0155] }
[0156] }
[0157] Table. Pseudocode for Inverse Quantization of Attribute Prediction Residuals
[0158] int PCCInverseQuantization(int value,int quantStep){
[0159] if(quantStep==0){
[0160] return value;
[0161] }else{
[0162] return value * quantStep;
[0163] }
[0164] }
[0165] When the predictors of each point have neighboring points, the point cloud encoder (e.g., arithmetic encoder 40012) according to the embodiment can perform entropy encoding on the residual values of quantization and inverse quantization as described above. When the predictors of each point do not have neighboring points, the point cloud encoder (e.g., arithmetic encoder 40012) according to the embodiment can perform entropy encoding on the attributes of the corresponding points without performing the above operations.
[0166] The point cloud encoder (e.g., lift transformer 40010) according to the embodiment can generate predictors for each point, set the calculated LOD and register neighboring points in the predictors, and set weights based on the distance to the neighboring points to perform lift transform coding. The lift transform coding according to the embodiment is similar to the prediction transform coding described above, but the difference is that the weights are applied cumulatively to the attribute values. The process of applying weights cumulatively to the attribute values according to the embodiment is configured as follows.
[0167] 1) Create an array quantized weights (QW) to store the weight values of each point. The initial value of all elements of QW is 1.0. Multiply the QW value of the predictor index of the neighboring nodes registered in the predictor by the weight of the current point's predictor, and add the values obtained by multiplication.
[0168] 2) Improve the prediction process: Subtract the value obtained by multiplying the attribute value of the point by the weight from the existing attribute value to calculate the predicted attribute value.
[0169] 3) Create temporary arrays called updateweight and update, and initialize the temporary arrays to zero.
[0170] 4) The weights calculated by multiplying the weights computed for all predictors by the weights stored in the QW corresponding to the predictor index are summed with the updateweight array and used as the index of the neighboring node. The values obtained by multiplying the attribute values of the neighboring node indices by the calculated weights are summed with the update array.
[0171] 5) Improve the update process: Divide the attribute values of the update array of all predictors by the weight values of the updateweight array of the predictor index, and add the existing attribute values to the values obtained by division.
[0172] 6) For all predictors, the predicted attribute is calculated by multiplying the attribute value updated through the boosting update process by the weight updated through the boosting prediction process (stored in QW). The predicted attribute value is quantized by a point cloud encoder (e.g., coefficient quantizer 40011) according to the implementation. In addition, the point cloud encoder (e.g., arithmetic encoder 40012) performs entropy encoding on the quantized attribute value.
[0173] A point cloud encoder according to an embodiment (e.g., RAHT transformer 40008) can perform RAHT transform coding, where attributes associated with lower-level nodes in an octree are used to predict attributes of higher-level nodes. RAHT transform coding is an example of intra-frame attribute coding performed by scanning backward through an octree. The point cloud encoder according to an embodiment scans the entire region starting from voxels and repeats a merging process of merging voxels into larger blocks at each step until the root node is reached. The merging process according to the embodiment is performed only on occupied nodes. The merging process is not performed on empty nodes. The merging process is performed on the node directly above an empty node.
[0174] The following equation represents the RAHT transformation matrix. In this equation, This represents the average attribute value of the voxels at level l. Based on and To calculate. and The weight is and
[0175]
[0176] here, It is a low-pass value and is used in the next higher level of merging. This represents the high-pass coefficient. The high-pass coefficients at each step are quantized and subjected to entropy encoding (e.g., encoded by an arithmetic encoder 400012). The weights are calculated as follows: pass and Create the root node as follows.
[0177]
[0178] A point cloud decoder according to an embodiment is shown.
[0179] The point cloud decoder shown is The example of the point cloud video decoder 10006 described in [the document], and it can be executed with [the following]. The operation of the point cloud video decoder 10006 shown is the same or similar. As shown, the point cloud decoder can receive a geometry bitstream and an attribute bitstream contained in one or more bitstreams. The point cloud decoder includes a geometry decoder and an attribute decoder. The geometry decoder performs geometry decoding on the geometry bitstream and outputs the decoded geometry. The attribute decoder performs attribute decoding based on the decoded geometry and the attribute bitstream and outputs the decoded attributes. The decoded geometry and decoded attributes are used to reconstruct the point cloud content (the decoded point cloud).
[0180] A point cloud decoder according to an embodiment is shown.
[0181] The point cloud decoder shown is The example shown is a point cloud decoder that can perform decoding operations. The reverse processing of the encoding operation of the point cloud encoder shown.
[0182] For reference and As described, the point cloud decoder can perform geometry decoding and attribute decoding. Geometry decoding is performed before attribute decoding.
[0183] The point cloud decoder according to the implementation includes an arithmetic decoder (arithmetic decoding) 11000, an octree synthesizer (synthesized octree) 11001, a surface approximation synthesizer (synthesized surface approximation) 11002, a geometry reconstructor (reconstructed geometry) 11003, an inverse coordinate transformer (inverse coordinate transformation) 11004, an arithmetic decoder (arithmetic decoding) 11005, an inverse quantizer (inverse quantization) 11006, a RAHT transformer 11007, an LOD generator (generated LOD) 11008, an inverse lifter (inverse lift) 11009, and / or a color inverse transformer (inverse color transformation) 11010.
[0184] An arithmetic decoder 11000, an octree synthesizer 11001, a surface approximation synthesizer 11002, a geometry reconstructor 11003, and a coordinate inverse transformer 11004 can perform geometric decoding. Geometric decoding according to the embodiment may include direct encoding and triplet geometric decoding. Direct encoding and triplet geometric decoding are selectively applied. Geometric decoding is not limited to the examples described above and is provided as a reference. The inverse processing of the described geometric encoding is performed.
[0185] The arithmetic decoder 11000 according to the embodiment decodes the received geometric bitstream based on arithmetic coding. The operation of the arithmetic decoder 11000 corresponds to the inverse processing of the arithmetic encoder 40004.
[0186] The octree synthesizer 11001 according to the embodiment can generate an octree by obtaining occupancy codes (or information about the geometry obtained as a decoding result) from the decoded geometry bitstream. The occupancy codes are as follows:
[0187] to Please describe that configuration in detail.
[0188] When applying triplet geometry encoding, the surface approximation synthesizer 11002 according to the implementation can synthesize the surface based on the decoded geometry and / or the generated octree.
[0189] According to the embodiments, the geometry reconstructor 11003 can regenerate geometry based on surface and / or decoded geometry. See also... As described, direct encoding and triadic geometric encoding are selectively applied. Therefore, the geometry reconstructor 11003 directly imports and sums the positional information of points for which direct encoding has been applied. When triadic geometric encoding is applied, the geometry reconstructor 11003 can reconstruct the geometry by performing the reconstruction operations (e.g., triangle reconstruction, upsampling, and voxelization) of the geometry reconstructor 40005. Details and references The descriptions are the same for all of them, so their descriptions are omitted. The reconstructed geometry may include point cloud images or frames that do not contain attributes.
[0190] According to the implementation, the inverse coordinate transformer 11004 can obtain the point position based on the reconstructed geometric transformation coordinates.
[0191] The arithmetic decoder 11005, inverse quantizer 11006, RAHT transformer 11007, LOD generator 11008, inverse booster 11009, and / or color inverse transformer 11010 are executable references. The attribute decoding described herein includes Region Adaptive Hierarchical Transformation (RAHT) decoding, interpolation-based hierarchical nearest neighbor prediction (prediction transformation) decoding, and interpolation-based hierarchical nearest neighbor prediction (lifting transformation) decoding with update / lifting steps. These three decoding schemes may be used selectively, or a combination of one or more decoding schemes may be used. The attribute decoding according to the embodiments is not limited to the examples described above.
[0192] According to the embodiment, the arithmetic decoder 11005 decodes the attribute bitstream by arithmetic encoding.
[0193] The inverse quantizer 11006 according to the implementation dequantizes information about the decoded attribute bitstream or the attributes obtained as a decoding result, and outputs the dequantized attributes (or attribute values). Inverse quantization can be selectively applied based on the attribute encoding of the point cloud encoder.
[0194] According to the implementation, the RAHT transformer 11007, LOD generator 11008, and / or inverse lifter 11009 can handle the reconstructed geometry and inverse quantization attributes. As described above, the RAHT transformer 11007, LOD generator 11008, and / or inverse lifter 11009 can selectively perform decoding operations corresponding to the encoding of the point cloud encoder.
[0195] According to the implementation, the color inverse transformer 11010 performs inverse transformation encoding to inversely transform the color values (or textures) included in the decoded attributes. The operation of the color inverse transformer 11010 can be selectively performed based on the operation of the color transformer 40006 of the point cloud encoder.
[0196] Although not shown in the figure, The elements of the point cloud decoder can be implemented by hardware, software, firmware, or a combination thereof, including one or more processors or integrated circuits configured to communicate with one or more memories included in the point cloud providing device. One or more processors can perform the above-described... The point cloud decoder's components include at least one or more operations and / or functions. Additionally, one or more processors are operable or perform operations for executing... The software program and / or instruction set for the operation and / or function of the elements of the point cloud decoder.
[0197] A transmitting device according to an embodiment is shown.
[0198] The transmitting device shown is The transmitting device 10000 (or Example of a point cloud encoder. The transmitting device shown can perform the same operation as the reference. The described point cloud encoder includes one or more of the same or similar operations and methods. The transmitting apparatus according to the embodiment may include a data input unit 12000, a quantization processor 12001, a voxelization processor 12002, an octree occupancy code generator 12003, a surface model processor 12004, an intra / inter-frame coding processor 12005, an arithmetic encoder 12006, a metadata processor 12007, a color transformation processor 12008, an attribute transformation processor 12009, a prediction / boosting / RAHT transformation processor 12010, an arithmetic encoder 12011, and / or a transmission processor 12012.
[0199] According to the embodiment, the data input unit 12000 receives or acquires point cloud data. The data input unit 12000 can perform operations and / or acquisition methods similar to those of the point cloud video acquirer 10001 (or refer to...). The described acquisition process (20000) is the same as or similar to the operation and / or acquisition method.
[0200] The data input unit 12000, quantization processor 12001, voxelization processor 12002, octree occupancy code generator 12003, surface model processor 12004, intra / inter-frame coding processor 12005, and arithmetic encoder 12006 perform geometric coding. Geometric coding according to the implementation method and reference... The geometric codes described are the same or similar, so their detailed descriptions are omitted.
[0201] The quantization processor 12001 according to the embodiment quantizes geometry (e.g., point position values). The operation of the quantization processor 12001 and / or the quantization with reference... The operation and / or quantization of the described quantizer 40001 are the same or similar. Details and references The descriptions are the same.
[0202] According to the embodiment, the voxelization processor 12002 voxels the quantized position values of points. The voxelization processor 120002 can execute and reference... The operation and / or voxelization process of the described quantizer 40001 is the same as or similar to the operation and / or process. Details and references The descriptions are the same.
[0203] According to the implementation method, the octree occupancy code generator 12003 performs octree encoding based on the voxelized positions of points in the octree structure. The octree occupancy code generator 12003 can generate occupancy codes. The octree occupancy code generator 12003 can execute and reference... and The operations and / or methods described are the same as or similar to those of the point cloud encoder (or octree analyzer 40002). Details and references The descriptions are the same.
[0204] According to the implementation, the surface model processor 12004 can perform triadic geometry encoding based on a surface model to reconstruct point positions in a specific region (or node) based on voxels. The surface model processor 12004 can perform operations related to reference... The operations and / or methods described are the same as or similar to those of the point cloud encoder (e.g., surface approximation analyzer 40003). Details and references are available. The descriptions are the same.
[0205] The intra / inter-frame coding processor 12005 according to the embodiment can perform intra / inter-frame coding on point cloud data. The intra / inter-frame coding processor 12005 can perform operations similar to those described above. The described intra / inter-frame coding is the same or similar. Details and references The descriptions are the same. According to an implementation, the intra / inter-frame coding processor 12005 may be included in the arithmetic encoder 12006.
[0206] The arithmetic encoder 12006 according to the embodiment performs entropy encoding on octrees and / or approximate octrees of point cloud data. For example, the encoding scheme includes arithmetic encoding. The arithmetic encoder 12006 performs the same or similar operations and / or methods as the arithmetic encoder 40004.
[0207] The metadata processor 12007 according to an embodiment processes metadata (e.g., set values) about point cloud data and provides it to necessary processing procedures such as geometric encoding and / or attribute encoding. Additionally, the metadata processor 12007 according to an embodiment can generate and / or process signaling information related to geometric encoding and / or attribute encoding. The signaling information according to an embodiment can be encoded separately from the geometric encoding and / or attribute encoding. The signaling information according to an embodiment can be interleaved.
[0208] Color transformation processor 12008, attribute transformation processor 12009, prediction / boosting / RAHT transformation processor 12010, and arithmetic encoder 12011 perform attribute encoding. Attribute encoding according to the implementation method and reference... The attribute codes described are the same or similar, so their detailed descriptions are omitted.
[0209] According to the implementation, the color transformation processor 12008 performs color transformation encoding to transform color values included in attributes. The color transformation processor 12008 can perform color transformation encoding based on reconstructed geometry. The reconstructed geometry and reference... The description is the same. Furthermore, its execution is the same as the reference. The operation and / or methods of the described color converter 40006 are the same as or similar to those described. Detailed descriptions are omitted.
[0210] According to the implementation, the attribute transformation processor 12009 performs attribute transformation to transform attributes based on reconstructed geometry and / or locations where geometric encoding is not performed. The attribute transformation processor 12009 performs transformations with reference to... The described attribute transformer 40007 operates and / or uses the same or similar operations and / or methods. Detailed descriptions are omitted. The prediction / boosting / RAHT transformation processor 12010 according to the embodiment can encode the transformed attributes using any one or a combination of RAHT encoding, prediction transformation encoding, and boosting transformation encoding. The prediction / boosting / RAHT transformation processor 12010 performs and references... The RAHT transformer 40008, LOD generator 40009, and lift transformer 40010 described herein operate at least one of the same or similar operations. Furthermore, the predictive transform coding, lift transform coding, and RAHT transform coding are similar to those of the reference transformer. The descriptions are the same, so their detailed descriptions are omitted.
[0211] The arithmetic encoder 12011 according to the embodiment can encode the attributes of the encoding based on arithmetic encoding. The arithmetic encoder 12011 performs the same or similar operations and / or methods as the arithmetic encoder 400012.
[0212] According to an embodiment, the transmission processor 12012 can transmit individual bitstreams containing encoded geometric and / or encoded attribute and metadata information, or transmit a single bitstream configured with encoded geometric and / or encoded attribute and metadata information. When the encoded geometric and / or encoded attribute and metadata information according to an embodiment is configured as a single bitstream, the bitstream may include one or more sub-bitstreams. The bitstream according to an embodiment may include signaling information and slice data. The signaling information includes a sequence parameter set (SPS) for sequence-level signaling, a geometric parameter set (GPS) for geometric information encoded signaling, an attribute parameter set (APS) for attribute information encoded signaling, and a tile parameter set (TPS) for tile-level signaling. The slice data may include information about one or more slices. A slice according to an embodiment may include a geometric bitstream Geom00 and one or more attribute bitstreams Attr00 and Attr10.
[0213] A slice is a series of syntactic elements that represent all or part of a encoded point cloud frame.
[0214] The TPS according to an embodiment may include information about each tile in one or more tiles (e.g., coordinate information and height / size information about the bounding box). The geometric bitstream may include a header and a payload. The header of the geometric bitstream according to an embodiment may include a geom_parameter_set_id, a geom_tile_id, and a geom_slice_id included in the GPS, as well as information about the data included in the payload. As described above, the metadata processor 12007 according to an embodiment may generate and / or process signaling information and transmit it to the transmission processor 12012. According to an embodiment, the element performing geometry encoding and the element performing attribute encoding may share data / information with each other, as indicated by the dashed lines. The transmission processor 12012 according to an embodiment may perform the same or similar operations and / or transmission methods as the transmitter 10003. Details and References and The descriptions are the same, so their descriptions are omitted.
[0215] A receiving device according to an embodiment is shown.
[0216] The receiving device shown is The receiving device 10004 (or and Example of a point cloud decoder. The receiving device shown can perform the same operation as the reference. The same or similar one or more operations and methods described in the point cloud decoder.
[0217] The receiving apparatus according to the embodiment includes a receiver 13000, a receiving processor 13001, an arithmetic decoder 13002, an octree reconstruction processor based on occupancy codes 13003, a surface model processor (triangle reconstruction, upsampling, voxelization) 13004, an inverse quantization processor 13005, a metadata parser 13006, an arithmetic decoder 13007, an inverse quantization processor 13008, a prediction / boost / RAHT inverse transform processor 13009, a color inverse transform processor 13010, and / or a renderer 13011. Each decoding element according to the embodiment can perform the inverse processing of the operation of the corresponding encoding element according to the embodiment.
[0218] Receiver 13000, according to an embodiment, receives point cloud data. Receiver 13000 can perform operations related to... The operation and / or receiving method of the receiver 10005 are the same as or similar to those of the receiver. Detailed description omitted.
[0219] According to an embodiment, the receiving processor 13001 can acquire a geometric bitstream and / or an attribute bitstream from the received data. The receiving processor 13001 may be included in the receiver 13000.
[0220] Arithmetic decoder 13002, octet-based octree reconstruction processor 13003, surface model processor 13004, and inverse quantization processor 13005 are capable of performing geometric decoding. Geometric decoding and reference according to the implementation method. The described geometric decodings are the same or similar, so their detailed descriptions are omitted.
[0221] The arithmetic decoder 13002 according to the embodiment can decode a geometric bitstream based on arithmetic coding. The arithmetic decoder 13002 performs the same or similar operations and / or encodings as the arithmetic decoder 11000.
[0222] According to an embodiment, the octree reconstruction processor 13003 based on occupancy codes can reconstruct an octree by obtaining occupancy codes from the decoded geometric bitstream (or information about the geometry obtained as a decoding result). The octree reconstruction processor 13003 based on occupancy codes performs the same or similar operations and / or methods as the octree synthesizer 11001 and / or the octree generation method. When applying triad geometry encoding, the surface model processor 13004 according to an embodiment can perform triad geometry decoding and related geometric reconstruction (e.g., triangle reconstruction, upsampling, voxelization) based on surface model methods. The surface model processor 13004 performs the same or similar operations as the surface approximation synthesizer 11002 and / or the geometry reconstructor 11003.
[0223] The geometry of reversible quantization decoding according to the inverse quantization processor 13005 of the embodiment.
[0224] The metadata parser 13006 according to the implementation can parse metadata (e.g., settings) contained in received point cloud data. The metadata parser 13006 can pass the metadata to geometry decoding and / or attribute decoding. Metadata and reference The metadata described is the same, so its detailed description is omitted.
[0225] The arithmetic decoder 13007, inverse quantization processor 13008, prediction / boost / RAHT inverse transform processor 13009, and color inverse transform processor 13010 perform attribute decoding. Attribute decoding and reference The properties described are decoded the same or similarly, so their detailed descriptions are omitted.
[0226] The arithmetic decoder 13007 according to the embodiment can decode the attribute bitstream via arithmetic coding. The arithmetic decoder 13007 can decode the attribute bitstream based on the reconstructed geometry. The arithmetic decoder 13007 performs the same or similar operations and / or encodings as the arithmetic decoder 11005.
[0227] The inverse quantization processor 13008 according to the embodiment can reversibly quantize and decode the attribute bitstream. The inverse quantization processor 13008 performs the same or similar operations and / or methods as the inverse quantizer 11006 and / or the inverse quantization method.
[0228] According to an embodiment, the prediction / boosting / RAHT inverse transform processor 13009 can process reconstructed geometry and inverse quantization attributes. The prediction / boosting / RAHT inverse transform processor 13009 performs one or more operations and / or decodings that are the same as or similar to those of the RAHT transformer 11007, LOD generator 11008, and / or inverse booster 11009. According to an embodiment, the color inverse transform processor 13010 performs inverse transform encoding to inverse transform color values (or textures) included in the decoded attributes. The color inverse transform processor 13010 performs operations and / or inverse transform encodings that are the same as or similar to those of the color inverse transformer 11010. According to an embodiment, the renderer 13011 can render point cloud data.
[0229] An exemplary structure is shown that can be operated in conjunction with a point cloud data transmission / reception method / apparatus according to an embodiment.
[0230] The structure represents a configuration in which at least one of the following components—server 1460, robot 1410, autonomous vehicle 1420, XR device 1430, smartphone 1440, home appliance 1450, and / or head-mounted display (HMD) 1470—is connected to cloud network 1400. Robot 1410, autonomous vehicle 1420, XR device 1430, smartphone 1440, or home appliance 1450 are referred to as devices. Furthermore, XR device 1430 may correspond to a point cloud data (PCC) device according to an embodiment or be operatively connected to a PCC device.
[0231] Cloud network 1400 can refer to a network that forms part of or exists within a cloud computing infrastructure. Here, cloud network 1400 can be configured using a 3G network, a 4G or Long Term Evolution (LTE) network, or a 5G network.
[0232] Server 1460 can be connected via cloud network 1400 to at least one of robot 1410, self-driving vehicle 1420, XR device 1430, smartphone 1440, home appliance 1450 and / or HMD 1470, and can assist at least a portion of the processing of connected devices 1410 to 1470.
[0233] HMD 1470 represents one of the implementation types of an XR device and / or PCC device according to an embodiment. An HMD-type device according to an embodiment includes a communication unit, a control unit, a memory, an I / O unit, a sensor unit, and a power supply unit.
[0234] Hereinafter, various embodiments of the apparatus 1410 to 1450 that apply the above-described technology will be described. The devices 1410 to 1450 shown can be operably connected / coupled to the point cloud data sending and receiving device according to the above-described embodiments.
[0235] <PCC+XR>
[0236] The XR / PCC device 1430 can adopt PCC technology and / or XR (AR+VR) technology, and can be implemented as an HMD, a head-up display (HUD) set in a vehicle, a television, a mobile phone, a smartphone, a computer, a wearable device, a household appliance, a digital sign, a vehicle, a stationary robot, or a mobile robot.
[0237] The XR / PCC device 1430 can analyze 3D point cloud data or image data obtained through various sensors or from external devices and generate position data and attribute data regarding 3D points. Thus, the XR / PCC device 1430 can obtain information about the surrounding space or real objects, and render and output XR objects. For example, the XR / PCC device 1430 can match an XR object including auxiliary information about the identified object with the identified object and output the matched XR object.
[0238]
[0239] The XR / PCC device 1430 can be implemented as a mobile phone (smartphone) 1440 by applying PCC technology.
[0240] The mobile phone 1440 can decode and display point cloud content based on PCC technology.
[0241] <PCC+Autonomous driving+XR>
[0242] The autonomous driving vehicle 1420 can be implemented as a mobile robot, a vehicle, an unmanned aerial vehicle, etc. by applying PCC technology and XR technology.
[0243] The autonomous driving vehicle 1420 applied with XR / PCC technology can represent an autonomous driving vehicle provided with means for providing an XR image, or an autonomous driving vehicle as a control / interaction target in an XR image. Specifically, as a control / interaction target in an XR image, the autonomous driving vehicle 1420 can be distinguished from the XR device 1430 and can be operably connected thereto.
[0244] The autonomous driving vehicle 1420 having means for providing an XR / PCC image can obtain sensor information from sensors including a camera, and output the generated XR / PCC image based on the obtained sensor information. For example, the autonomous driving vehicle 1420 can have a HUD and output an XR / PCC image thereto, thereby providing an XR / PCC object corresponding to a real object or an object presented on the screen to passengers.
[0245] When an XR / PCC object is output to a HUD, at least a portion of the XR / PCC object can be output to overlap with the actual object being pointed at by the passenger's eyes. Conversely, when an XR / PCC object is output to a display installed within the autonomous vehicle, at least a portion of the XR / PCC object can be output to overlap with objects on the screen. For example, the autonomous vehicle 1220 can output XR / PCC objects corresponding to objects such as roads, other vehicles, traffic lights, traffic signs, two-wheeled vehicles, pedestrians, and buildings.
[0246] Virtual reality (VR), augmented reality (AR), mixed reality (MR), and / or point cloud compression (PCC) technologies according to the implementation methods are applicable to various devices.
[0247] In other words, VR technology is a display technology that only provides CG images of real-world objects, backgrounds, etc. On the other hand, AR technology refers to the technology of displaying virtually created CG images on top of images of real objects. MR technology is similar to AR technology in that the virtual objects to be displayed are mixed and combined with the real world. However, MR technology differs from AR technology in that AR technology clearly distinguishes between real objects and virtual objects created as CG images and uses virtual objects as supplementary objects to real objects, while MR technology treats virtual objects as objects with the same characteristics as real objects. More specifically, an example of MR technology application is holographic services.
[0248] Recently, VR, AR, and MR technologies have often been referred to as Extended Display (XR) technologies rather than being clearly distinguished from each other. Therefore, embodiments of this disclosure are applicable to any of VR, AR, MR, and XR technologies. Encoding / decoding based on PCC, V-PCC, and G-PCC technologies are applicable to such technologies.
[0249] The PCC method / apparatus according to the embodiments can be applied to vehicles that provide self-driving services.
[0250] Vehicles providing autonomous driving services connect to the PCC device for wired / wireless communication.
[0251] When the point cloud data (PCC) transmitting / receiving device according to the embodiment is connected to a vehicle for wired / wireless communication, the device can receive / process content data related to AR / VR / PCC services (which may be provided together with autonomous driving services) and transmit it to the vehicle. When the PCC transmitting / receiving device is installed in the vehicle, it can receive / process content data related to AR / VR / PCC services based on user input signals input through a user interface device and provide it to the user. The vehicle or user interface device according to the embodiment can receive user input signals. User input signals according to the embodiment may include signals indicating autonomous driving services.
[0252] This is a flowchart illustrating the operation of a point cloud transmitting device according to an embodiment.
[0253] As referenced above The described point cloud transmitting device (e.g., The transmitting device Point cloud encoder and The transmitting device receives input point cloud data and processes the input data (1510). The input data includes geometry and / or attributes. According to an embodiment, geometry is information representing the position of a point in the point cloud, and attributes are information corresponding to at least one of color and / or reflectivity of each point. According to an embodiment, geometry can be used as a concept relating to the position of a single point or the positions of all points. According to an embodiment, attributes can be used as a concept relating not only to the color (e.g., R value of RGB) or reflectivity of a single point, but also to the color and / or reflectivity of multiple points. The geometry according to an embodiment can be represented by parameters of at least one coordinate system, such as a Cartesian coordinate system, a cylindrical coordinate system, or a spherical coordinate system. Therefore, the point cloud transmitting device (e.g., referring to...) receives input point cloud data and processes the input data (1510). The described coordinate transformer 40000 can transform received geometric information into information in a coordinate system, so as to represent the positions of each point indicated by the input geometric information as positions in 3D space. The coordinate system according to the embodiment is not limited to the example described above. Furthermore, point cloud transmitting devices (e.g., refer to...) The described quantizer 40001 can quantize geometric information represented in a coordinate system and generate transformed quantized geometric information. For example, a point cloud transmitting device can apply one or more transformations, such as position transformation and / or rotation transformation, to the position of a point indicated by geometric information, and perform quantization by dividing the transformed geometric information by a quantization value. The quantization value according to an embodiment can vary based on the distance between the encoding unit (e.g., a tile, slice, etc.) and the origin of the coordinate system or the angle with a reference direction. According to an embodiment, the quantization value can be a preset value. When quantization is performed, one or more points may have the same position. These points are called duplicate points. That is, one or more points may have the same quantized position but different attributes. The point cloud transmitting device can remove duplicates (remove duplicate points) by combining duplicate points into one point, and send information about multiple attributes associated with the combined point to an attribute transmission module for attribute encoding. The point cloud transmitting device (e.g., The quantizer (40001) can perform quantization and / or duplicate point removal, followed by voxelization to assign attributes to the remaining points. Voxelization is the process of grouping points into voxels. See the detailed process and references... The description is identical, so its description will be skipped. Furthermore, the point cloud transmitting device can transform the properties of an attribute (e.g., color). For example, when the attribute is color, the point cloud transmitting device can transform the color space of that attribute (from RGB to YCbCr or vice versa). The individual sub-components of the attribute (e.g., luminance and chrominance) are processed independently.
[0254] According to an implementation method, the point cloud transmitting device can divide point cloud data into tiles and slices in 3D space to store point information about the point cloud data. A tile is a bounding box (e.g., refer to...). The bounding box (described) is a (3D) cuboid. The bounding box may include one or more tiles. One tile may completely or partially overlap with another tile. One tile may include one or more slices. A slice is a set of points and is represented as a series of syntactic elements representing all or part of the encoded point cloud data. A slice may or may not have dependencies on other slices. A point cloud transmission apparatus according to an embodiment may include a spatial segmenter configured to segment the point cloud data.
[0255] According to the implementation method, the point cloud transmitting device performs geometric encoding (or coding) (1520) on the geometry. The point cloud transmitting device can perform reference... The described coordinate transformer 40000, quantizer 40001, octree analyzer 40002, surface approximation analyzer 40003, arithmetic encoder 40004, and geometry reconstructor (reconstructed geometry) 40005 are described as operating at least one of the following: The operation of the described data input unit 12000, quantization processor 12001, voxelization processor 12002, octree occupancy code generator 12003, surface model processor 12004, intra / inter-frame coding processor 12005, arithmetic encoder 12006, and metadata processor 12007 is at least one of the following operations.
[0256] According to the implementation method, geometric encoding may include, but is not limited to, at least one of octree-based octree geometric encoding and ternary geometric encoding. Geometric encoding and reference. The description is identical, so its description is skipped. Octree geometry encoding generates octrees representing voxels. Octrees represent points matching voxels based on the octree structure. Voxels and octrees are compared with references. and The descriptions are the same, so I will skip their detailed descriptions.
[0257] The point cloud transmitting device predicts geometric information (1521). The point cloud transmitting device can perform quantization on the encoded geometry and calculate predicted values (or predicted geometric information) based on the quantization values of adjacent coding units. Furthermore, the point cloud transmitting device can generate reconstructed geometry based on the predicted geometric information and residual information reconstructed from the quantized geometry. The reconstructed geometry is provided for attribute encoding through additional processing such as filtering.
[0258] Point cloud transmitting device (e.g., The arithmetic encoder 40004 performs geometric entropy coding (1522) on the quantized geometry. According to the implementation, the entropy coding may include exponential Golomb coding, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC).
[0259] Point cloud transmitting device (e.g., The attribute transformer 40007 performs attribute transfer (recoloring) for attribute encoding (1530). The point cloud transmitting apparatus according to the embodiment can receive reconstructed geometry, geometry and attributes as input, and determine the attribute values that minimize attribute distortion.
[0260] The point cloud transmitting apparatus performs attribute encoding (1531). According to an embodiment, the point cloud processing apparatus selectively uses RAHT encoding, predictive transform encoding, and lifting transform encoding, or a combination of one or more encoding schemes, to perform attribute encoding (or attribute coding) based on the point cloud content. For example, RAHT encoding and lifting transform encoding can be used for lossy encoding that significantly compresses the point cloud content data. Predictive transform encoding can be used for lossless encoding. The point cloud transmitting apparatus according to the embodiment can perform reference... The described operation includes at least one of the following: color transformer 40006, attribute transformer 40007, RAHT transformer 40008, LOD generator 40009, boost transformer 40010, coefficient quantizer 40011, and / or arithmetic encoder 40012. Furthermore, the point cloud transmitting device can perform the reference... At least one of the operations of the described color transformation processor 12008, attribute transformation processor (or attribute transformation processor) 12009, prediction / lifting / RAHT transformation processor 12010, and arithmetic encoder 12011.
[0261] The point cloud transmitter performs attribute quantization (1532). The point cloud transmitter can receive the transformed residual attribute of the encoded attribute and generate the transformed and quantized residual attribute based on the quantized value.
[0262] Point cloud transmitting device (e.g., The arithmetic encoder 12011 performs attribute entropy encoding (1533). The point cloud transmitting device can receive the transformed quantized residual attribute information, entropy encode it, and output the attribute information bit stream. The entropy encoding according to the implementation may include any one or more of exponential Columbus, CAVLC, and CABAC, but is not limited thereto.
[0263] This is a block diagram of a point cloud transmitting device according to an embodiment.
[0264] Point cloud transmitting device (e.g., reference) The described point cloud transmitting device performs attribute encoding. According to the implementation method, attribute encoding is generated based on the Level of Detail (LOD). See reference... and As described, a Level of Detail (LOD) can be generated by reorganizing points distributed in 3D space into a set of refinement levels. According to an embodiment, an LOD may include one or more points distributed at regular intervals. As mentioned above, the LOD according to an embodiment is the level of detail of the point cloud content. Therefore, as the level of LOD (or LOD value) indication decreases, the detail of the point cloud content deteriorates. As the level of LOD indication increases, the detail of the point cloud content is enhanced. According to an embodiment, a point cloud encoder (e.g., Point cloud encoders and point cloud decoders (e.g., The point cloud encoder and decoder can generate a Level of Distributed Object (LOD) to increase the attribute compression ratio. This is because points with similar attributes are likely to be near the target point, so the residual value between the predicted attribute obtained from neighboring points with similar attributes and the attribute of the target point is likely to be close to zero. Therefore, the point cloud encoder and decoder can generate an LOD to select appropriate neighboring points that can be used to predict attributes.
[0265] As shown in the figure, the point cloud transmission device includes a LOD configurator 1610 and a neighboring point set configurator 1620. The LOD configurator 1610 can perform the same or similar operations as the LOD generator 40009. That is, as shown in the figure, the LOD configurator 1610 can receive attributes and reconstructed geometry, and configure one or more LODs based on the received attributes and reconstructed geometry.
[0266] According to the implementation, the LOD configurator 1610 can use one or more methods to configure the LOD. As described above, the point cloud decoder (e.g., referring to...) and The described point cloud decoder should also generate a Level of Detail (LOD). Therefore, information related to one or more LOD configuration methods (or LOD generation methods) according to the embodiments is included in the bitstream generated based on geometric encoding and attribute encoding. The LOD output of the LOD configurator 1610 is sent to at least one of a prediction transformer / inverse transformer and a boost transformer / inverse transformer.
[0267] The LOD configurator 1610 can generate LODs based on one or more methods, which are capable of generating LODs. l The set maintains a specific interval between points while reducing the computational complexity of the point spacing.
[0268] According to the implementation, the LOD configurator 1610 can configure the LOD based on the Morton code of the points. As described above, the Morton code is generated by representing the coordinate values (e.g., (x, y, z)) indicating the 3D positions of all points as bit values and mixing these bits.
[0269] According to an implementation, the LOD configurator 1610 can generate Morton codes for each point based on the reconstructed geometry, and can organize the points in ascending order based on the Morton codes. The order in which the points are organized in ascending order of Morton codes can be referred to as the Morton order. The LOD configurator 1610 can configure the LOD by performing sampling on the points organized in the Morton order. The LOD configurator can perform sampling in various ways. For example, according to an implementation, the LOD configurator can sequentially select points based on the gaps in the sampling rate, according to the Morton order of the points included in the regions corresponding to the node. That is, the LOD configurator 1610 selects the point first organized according to the Morton order (point 0), and points separated from the first point by the same number of samples at the same sampling rate (e.g., the 5th point separated from the first point when the sampling rate is 5). According to an implementation, the LOD configurator can select the point with the Morton code value closest to the Morton code value of the center point of the node, based on the Morton code value of the center point and the Morton code values of the adjacent points.
[0270] For reference The point cloud encoder (or attribute information predictor) according to the embodiment generates predictors for points and performs predictive transformation encoding to set the predicted attributes (or predicted attribute values) of each point. That is, it can generate N predictors for N points.
[0271] When each LOD (or LOD set) is generated, the adjacent point set configurator 1620 can search or retrieve LODs. l A set of points has one or more neighboring points. The number of one or more neighboring points can be represented by X, where X is an integer greater than 0. According to the implementation, neighboring points are those in 3D space that are adjacent to the LOD (Level of Dimension). l The nearest neighbor (NN) of the point in the set, and included in the same LOD as the target LOD (e.g., LOD 1). l It is included in, or is included in a set of LODs at a lower level than the target LOD (e.g., LOD 10 ... l-1 LOD l-2 …, LOD0). The neighbor set configurator 1620 can register one or more searched neighbor points as a neighbor set in the predictor. According to an implementation, the number of neighbor points can be set to a maximum number of neighbor points based on an input signal from the user, or the number of neighbor points can be preset to a specific value according to the neighbor search method.
[0272] According to the implementation method, the adjacent point set configurator 1620 can search LOD0 and LOD1 to find points belonging to... The points adjacent to point P3 in LOD1 are shown. As shown, LOD0 includes P0, P5, P4, and P2. LOD2 includes the points of LOD0, the points of LOD1, and P9, P8, and P7. When the number of neighboring points X is 3, the neighboring point set configurator 1620... In the 3D space shown at the top, points belonging to LOD0 or LOD1 are searched to find the three nearest neighbors of P3. That is, the neighbor set configurator 1620 can search for P6, which belongs to LOD1 (same LOD level), and P2 and P4, which belong to LOD0 (lower LOD level), as neighbors of P3. In the 3D space, P7 is a point close to P3, but is not searched as a neighbor because it is at a higher LOD level. The neighbor set configurator 1620 can register the searched neighbor points P2, P4, and P6 in the predictor as the neighbor set of P3. The method for generating the neighbor set according to the embodiment is not limited to this example. Furthermore, information about the method for generating the neighbor set according to the embodiment (hereinafter, neighbor set generation information) is included in the bitstream containing the encoded point cloud video data and is sent to the receiving device (e.g., The receiving device 10004 or and (point cloud decoder).
[0273] As described above, each point can have a predictor. According to the embodiment, the point cloud encoder can apply the predictor to encode the attribute values of the corresponding points and generate predicted attributes (or predicted attribute values). According to the embodiment, after generating the LOD, a predictor is generated based on the searched neighboring points. The predictor is used to predict the attributes of the target point. The attribute values of the points and the residual values of the attributes predicted by the point's predictor are sent to the point cloud receiving device via a bitstream.
[0274] The distribution of points according to the LOD level is shown.
[0275] The leftmost portion of the diagram, 1700, represents the content corresponding to the highest level of LOD. Arrow 1710 indicates the LOD level increasing from the lowest LOD level. As the LOD level decreases, the pixel spacing increases (density decreases). As the LOD level increases, the pixel spacing decreases (density increases).
[0276] An example illustrating the neighbor set search method.
[0277] It is determined by the neighboring point set configurator (e.g., An example of a method for searching neighboring point sets using the neighboring point set configurator 1620. The arrows shown in the figure indicate the Morton order according to the implementation.
[0278] As described above, the point cloud encoder according to the embodiment can present the position values of points (e.g., coordinate values (x, y, z) representing a coordinate system in 3D space) as bit values and generate Morton codes by mixing the bits. Points are organized in ascending order based on the size of the Morton codes. Therefore, points located before other points organized in Morton order have the smallest Morton codes. To provide points belonging to the LOD... l The set of points Px generates a set of neighboring points. According to the implementation method, the neighboring point set configurator is used to align points belonging to LOD0 to LOD0. l-1 The set of points (reserved list) and belonging to LOD l A neighbor set search is performed on points in the set that precede point Px in the Morton order (i.e., points whose Morton codes are less than or equal to Px). According to an implementation, when a Level of Optimization (LOD) exists, the neighbor set configurator can determine the neighbor search range based on the predicted position of the target point (e.g., Px) in the Morton order.
[0279] The neighbor point set configurator can search among points preceding point Px in Morton order for a point with a Morton code closest to that of point Px. According to an embodiment, the searched point may be referred to as the center point. The method for searching the center point is not limited to this example. According to an embodiment, the neighbor point configurator can increase the compression ratio based on attribute encoding by moving the search range according to the method for searching the center point. According to an embodiment, information about the method for searching the center point is included in the neighbor point set generation information and is transmitted to a receiving device (e.g., via a bitstream containing the aforementioned encoded point cloud video data). The receiving device 10004 or and (Point cloud decoder). According to the implementation, when a LOD exists, the neighbor point set configurator can determine the neighbor point search range based on the predicted target point (e.g., Px) position in the Morton order.
[0280] The neighbor set configurator can set the neighbor point search range based on the searched center point. The neighbor point search range can include one or more points located before and after the center point in Morton order. According to an implementation, information about the neighbor point search range can be included in neighbor point set generation information and transmitted to a receiving device via a bitstream containing encoded point cloud video data. The neighbor set configurator compares the distances between points within the neighbor point search range determined around the center point and point Px, and registers as many as the maximum number of neighbor points set (e.g., 3) the points closest to point Px as neighbors.
[0281] Because the registered neighboring points are close to the target point (e.g., Px), they are determined to have high attribute relevance. However, depending on the characteristics of the point cloud content, there may be cases where the distances between the registered neighboring points are not all adjacent to the target point (e.g., Px) in terms of distance.
[0282] An example of point cloud content is shown.
[0283] The left side of the figure shows point cloud content 1900 with high-density points, and the right side shows point cloud content 1910 with low-density points. The cloud content exhibits different characteristics depending on the object, capture method, and device used.
[0284] For example, point cloud content 1900 with high-density points is generated by capturing a relatively narrow area of an object using a 3D scanner or similar device, thus exhibiting high correlation between points. Point cloud content 1910 with low-density points is generated by capturing a large area at low density using a LiDAR device or similar device, thus exhibiting relatively low correlation between points.
[0285] Therefore, when point cloud content 1900 has a high density, the efficiency of attribute encoding will not degrade even when registered neighboring points are separated from the target point (e.g., point Px) because the attribute correlation between points is high. When point cloud content 1910 has a low density, the efficiency of attribute encoding will not decrease only when the distance between registered neighboring points and the target point is within the range that guarantees the attribute correlation between points.
[0286] Therefore, to ensure optimal attribute compression efficiency independent of the characteristics of the point cloud content, the point cloud encoder according to the embodiment performs a neighbor search method, wherein the maximum nearest neighbor distance between neighboring points and the target point is set (calculated) taking into account the characteristics of the point cloud content to select and register neighboring points. The maximum nearest neighbor distance can vary depending on the LOD generation method, or it can be set by the user. According to the embodiment, information about the maximum nearest neighbor distance, such as information about the maximum nearest neighbor distance calculation method and related parameters, is transmitted via a bitstream. Therefore, the point cloud receiving device can obtain information about the maximum nearest neighbor distance from the bitstream and search for neighboring points based on the obtained information to perform attribute decoding.
[0287] An example of an octree-based LOD generation process is shown.
[0288] The LOD 2000 according to the implementation may include points grouped based on distances between points. A point cloud encoder reorganizes the points based on an octree 2010 to generate one or more LODs. An iterative generation algorithm suitable for octree decoding, based on the position or order of the points (e.g., Morton code order, etc.), is applied to the grouped points. In each sequential iteration, one or more refinement levels R0, R1…, Ri belonging to a single LOD (e.g., LODi) are generated. That is, the level of the LOD is a combination of multiple refinement levels. The octree has one or more depths. (See reference...) In an octree, the upper node is called the root node, the lower node is called the leaf node, and the depth increases in the direction from the root node to the leaf node. Each depth of the octree can correspond to one or more Levels of Depth (LOD). For example, the root node corresponds to LOD0, while the leaf nodes correspond to the maximum level LOD N. The depth of point S in the octree 2010 shown in the figure corresponds to the level LOD Nx.
[0289] According to the implementation method, an approximate nearest neighbor search method is used to generate a Level of Detail (LOD) in the octree from the lowest point to the highest point. That is, in the direction from the lowest point to the highest point, a LOD is generated at a level lower than or equal to the current LOD (e.g., LOD 1). l Level of LOD (e.g., LOD) l-1 The point cloud processing apparatus (or transmitting apparatus) according to the embodiment can search for the nearest neighbor of the corresponding point in the current LOD. It can calculate the maximum nearest neighbor distance and select points whose distance from the corresponding point is less than or equal to the maximum nearest neighbor distance as neighboring points. According to the embodiment, the maximum nearest neighbor distance is determined based on a reference distance, which is the diagonal distance of the block corresponding to the upper node of the octree node, and the distance between the octree node and the LOD to which the point belongs (e.g., LOD 1). l Corresponding to.
[0290] According to the implementation method, the reference distance of each LOD is represented as follows.
[0291] or
[0292]
[0293] Here, the parameter LOD represents the level of each LOD.
[0294] An example of the baseline distance is shown.
[0295] For reference The reference distance for each LOD represents the diagonal distance of the block to which each LOD belongs.
[0296] The left side of the diagram shows block 2100 corresponding to LOD0, and the right side shows block 2110 corresponding to LOD1. In an octree, a parent node (or upper node) has 8 child nodes. Therefore, a 3D block corresponding to a parent node can include 8 3D blocks corresponding to child nodes (or lower nodes). Therefore, block 2110 corresponding to LOD1 can include blocks belonging to the same parent node (e.g., corresponding to...). The 8 blocks of the LOD0 node shown.
[0297] According to reference The formula described, the baseline distance of LOD0 is The baseline distance for LOD1 is
[0298] An example of the maximum nearest neighbor distance is shown.
[0299] According to an embodiment, a point cloud transmission apparatus can set the octree node range as a neighboring point search range based on the characteristics of the point cloud content. According to an embodiment, the neighboring point search range is calculated as the number of neighboring nodes (e.g., parent nodes) of the octree node corresponding to the LOD level to which the corresponding point belongs. The neighboring point search range can be represented as NN_Range. NN_Range represents the number of octree nodes surrounding the current point. That is, NN_Range indicates the number of one or more octree nodes (or parent octree nodes) corresponding to the depth of the octree, which is equal to or less than the depth of the octree corresponding to the LOD level to which the current point belongs. According to an embodiment, the neighboring point search range can include a maximum range and a minimum range. According to an embodiment, based on reference... The baseline distance and the search range for neighboring points are described as follows: determine (calculate) the maximum nearest neighbor distance.
[0300]
[0301] The maximum nearest neighbor distance calculation method according to the implementation method can be called the maximum nearest neighbor distance calculation of adjacent points based on octree.
[0302] The left side shows the maximum nearest neighbor distance when the value of NN_Range is 1. Examples, and The right side shows the maximum nearest neighbor distance when the value of NN_Range is 3. Example. Depending on the implementation, the value of NN_Range can be set to any value independent of the octree node range. Therefore, the maximum nearest neighbor distance is expressed as:
[0303]
[0304] Here, LOD is a parameter indicating the LOD level to which the predicted target point belongs. Depending on the implementation, NN_Range can be set by calculating the distances between the target point (e.g., Px) and its neighbors. For example, the distance calculation method can use the L2 norm. Alternatively, the L1 norm (L1 = (nΣi|xi|) = (|x1| + |x2| + |x3| + ... + |xn|) can be used to reduce the computational load, or L2 norm can be used. 2 and L1 2 The value of NN_Range is adjusted according to the distance calculation method described above. Alternatively, NN_Range can be set based on the characteristics of the point cloud content. The point cloud transmitting device selects points (candidate neighbor points) that are at a distance equal to or less than the maximum nearest neighbor distance from the current point as neighbor points. Furthermore, the point cloud transmitting device sends information about NN_Range to the point cloud receiving device via a bitstream. Therefore, the point cloud receiving device (or point cloud decoder) obtains the information about NN_Range, calculates the maximum nearest neighbor distance, and generates a set of neighboring points based on the maximum nearest neighbor distance to perform attribute decoding.
[0305] An example of the maximum nearest neighbor distance is shown.
[0306] According to the implementation method, Levels of Depth (LODs) can be generated based on the distances between points. The distances used to select points for each LOD can be set by the user or calculated. The distances used to select points for each LOD are represented as follows.
[0307] dist2 L , where L represents the level of LOD.
[0308] The distance used to select points for each LOD is set as the baseline distance for searching neighboring points.
[0309] The baseline distances between adjacent points belonging to each LOD are shown. The distance between points P0, P2, P4, and P5 belonging to LOD0 is dist2. LOD0 The distance between points P7, P8, and P9, which belong to LOD2, is dist. LOD2 As per the above reference. The search range for neighboring points can be represented as NN_Range. The maximum nearest neighbor distance, according to the implementation method, is determined based on the baseline distance and the search range as follows.
[0310] Maximum nearest neighbor distance = dist2 L ×NN RANGE ,or
[0311] L2 maximum nearest neighbor distance = dist2L 2 ×NN RANGE
[0312] The maximum nearest neighbor distance calculation method according to the implementation method can be called the distance-based maximum nearest neighbor distance calculation method.
[0313] According to the implementation, NN_Range can be set by calculating the distance between the target point (e.g., Px) and its neighboring points. For example, the distance calculation method can use the L2 norm. Alternatively, the L1 norm (L1 = (nΣi|xi|) = (|x1| + |x2| + |x3| + ... + |xn|) can be used to reduce the computational load, or L2 norm can be used. 2 and L1 2 .
[0314] According to the implementation method, a Level of Detail (LOD) can be generated based on sampling (sampling). That is, the point cloud transmitting device can organize points based on Morton code values, and then register points in the current LOD that do not correspond to the k-th point according to the organization order. Unregistered points are used to generate other LODs different from the current LOD. LODs can have the same or different k values. The k value for each LOD is represented as k L The reference distance used for neighbor search according to the implementation method is set as the average point of each LOD and the kth consecutive point. L The average distance between points k. This can be applied to the k-th point. L The average value of all or some of the points is calculated. The reference distance according to the implementation method is given below.
[0315] Baseline distance = average_dist2norm(MC_points[(i+1)*k)) L ]-MC_points[i*k L ]), where MC_points represents points organized based on Morton code.
[0316] As referred above The search range for neighboring points is denoted as NN_Range. The maximum nearest neighbor distance, according to the implementation method, is determined based on the aforementioned baseline distance and search range, and is given below.
[0317] Maximum nearest neighbor distance = average_dist2norm(MC_points[(i+1)*k)) L ]-MC_points[i*k L ])x NN RANGE ,or
[0318] L2-based maximum nearest neighbor distance = average_distnorm(MC_points[(i+1)*k)) L ]-MC_points[i*k L ])xNN RANGE 2 .
[0319] Here, average_dist2norm and average_distnorm represent operations that calculate the average distance using the point's position difference vector as input. The maximum nearest neighbor distance calculation method according to the implementation can be referred to as sampled maximum nearest neighbor distance calculation.
[0320] According to the implementation method, NN_Range can be set by calculating the distance between the target point (Px) and its neighboring points. For example, the distance calculation method can use the L2 norm. Alternatively, the L1 norm (L1 = (nΣi|xi|) = (|x1| + |x2| + |x3| + ... + |xn|) can be used to reduce the computational load, or L2 norm can be used. 2 and L1 2 .
[0321] According to the implementation method, the reference distance for neighbor search can be calculated as the average difference of the Morton codes between sampled points (a method for calculating the maximum nearest neighbor distance of neighboring points by calculating the average difference of the Morton codes for each LOD), or it can be calculated as the average distance between points calculated for the currently configured LOD, independent of distance and sampling (a method for calculating the maximum nearest neighbor distance of neighboring points by calculating the average distance difference for each LOD). The maximum nearest neighbor distance can be set according to user input and can be sent to the point cloud receiving device via a bitstream.
[0322] An example of neighbor set search is shown.
[0323] Similar to This shows the configurator of adjacent point sets (e.g., The neighbor set configurator 1610) is based on the reference set. The process of performing a neighbor set search using the maximum nearest neighbor distance is described. The arrows in the diagram indicate the Morton order according to the implementation method. This is to provide information for points belonging to the LOD. l The set of points Px generates a set of neighboring points. According to the implementation method, the neighboring point set configurator is used to align points belonging to LOD0 to LOD0. l-1 The set of points (reserved list) and belonging to LOD lThe neighbor set configurator performs a neighbor set search on points preceding point Px in Morton order (i.e., points with Morton codes less than or equal to Px). The neighbor set configurator searches for points among those preceding point Px in Morton order that have a Morton code closest to that of point Px. Depending on the implementation, the searched point may be referred to as the center point. The neighbor set configurator compares multiple distances between point Px and multiple points within the neighbor set search range to the left and right of the searched center point. The neighbor set configurator determines the maximum nearest neighbor distance at the LOD to which Px belongs (e.g., referring to...). The point closest to Px among the points within the maximum nearest neighbor distance described is registered as the neighboring point set.
[0324] According to the implementation method, when searching for neighboring points among points within a search range, the neighboring point set configurator divides the search range into one or more blocks to reduce search time. The method for searching for neighboring points for a set of points corresponding to a block is configured as follows.
[0325] The points in the reserved list are divided into K groups. The point cloud transmitting device calculates the position and size of the bounding boxes surrounding the points in each group.
[0326] The point cloud transmitting device searches for the point with the closest Morton code value to point Px in a reserved list of points. The found point is represented by Pi.
[0327] It can calculate the distance between Px and multiple points in the group to which the searched point Pi belongs, and can register points that are close to each other as neighboring points.
[0328] When the number of registered neighboring points is less than the maximum number of neighboring points, the distance between the corresponding point and the points in the adjacent group can be calculated.
[0329] When the number of registered neighboring points exceeds the maximum number of neighboring points, the distance between the bounding box of the next neighboring group and point Px is calculated. The calculated distance is compared to the longest distance among the registered neighboring points and the longest distance between them. If the calculated distance is greater than the longest distance, points in that group are not used as neighboring point candidates. If the calculated distance is less than the longest distance, the distances between multiple points in that group and Px can be calculated, and points with close proximity can be registered as neighbors (neighboring point update). The neighboring point search according to this implementation is performed only within the search range (e.g., 128).
[0330] As described above, the neighbor set configurator searches for neighboring points around Pi, saving time in neighbor point search. Furthermore, when the number of neighboring points found is the same as the maximum number of neighboring points, the neighbor point registration process can be omitted. Additionally, since the neighbor set configurator searches for neighboring points based on bounding boxes, the accuracy of neighbor point search can be improved.
[0331] As mentioned above, the maximum nearest neighbor distance can be applied when calculating the distance between points or after searching all neighboring points. The resulting set of neighboring points can vary depending on the step in applying the maximum nearest neighbor distance, and can affect the bitstream size and PSNR.
[0332] Therefore, the point cloud sending apparatus (or neighboring point set generator) according to the embodiment can calculate the distance between Px and the points based on the characteristics of the point cloud content and / or service, and register points with a distance less than or equal to the maximum nearest neighbor distance as neighboring points. Furthermore, the point cloud sending apparatus can register the maximum neighboring points of Px based on the characteristics of the point cloud content and / or service, determine whether the distance between the registered points and point Px is shorter than or longer than the maximum nearest neighbor distance, i.e., whether the distance is within the maximum nearest neighbor distance, and remove points not within the maximum nearest neighbor distance from the neighboring points. The maximum nearest neighbor distance application step (or location) according to the embodiment is included with respect to the reference... The information described regarding the maximum nearest neighbor distance, such as the calculation method and related parameters, is transmitted via a bitstream. Therefore, the point cloud receiving device can obtain the maximum nearest neighbor distance information from the bitstream, search for neighboring points based on the obtained information, and perform attribute decoding.
[0333] An example of a bitstream structure diagram is shown.
[0334] Point cloud processing device (e.g., reference) , and The described transmitting device can transmit encoded point cloud data in the form of a bit stream. A bit stream is a sequence of bits that form a representation of point cloud data (or point cloud frames).
[0335] Point cloud data (or point cloud frames) can be divided into tiles and slices.
[0336] Point cloud data can be segmented into multiple slices and encoded in a bitstream. A slice is a set of points and is represented as a series of syntactic elements representing all or part of the encoded point cloud data. A slice may or may not have dependencies on other slices. Furthermore, a slice may include a geometric data unit and may have one or more attribute data units or zero attribute data units. Since attribute encoding is performed based on geometry encoding as described above, attribute data units are based on geometric data units within the same slice. That is, a point cloud data receiving device (e.g., receiving device 10004 or point cloud video decoder 10006) can process attribute data based on the decoded geometric data. Therefore, geometric data units must appear before the relevant attribute data units within a slice. Data units within a slice must be consecutive, and the order of slices is not specified.
[0337] A piece is a bounding box (e.g., see reference). A bounding box (described as a 3D cuboid). The bounding box may include one or more tiles. A tile may completely or partially overlap with another tile. A tile may include one or more slices.
[0338] Therefore, the point cloud data transmission device can provide high-quality point cloud content by processing data corresponding to tiles according to their importance. That is, the point cloud data transmission device according to the embodiment can use point cloud compression encoding to process data corresponding to areas important to the user with better compression efficiency and appropriate latency.
[0339] According to the implementation method, the bitstream contains signaling information and multiple slices (slice 0, ..., slice n). As shown in the figure, the signaling information precedes the slices in the bitstream. Therefore, the point cloud data receiving device can first obtain the signaling information and process the multiple slices sequentially or selectively based on the signaling information. As shown in the figure, slice 0 includes a geometric data unit (Geom0). 0 ) and two attribute data units (Attr0) 0 and Attr1 0 ).
[0340] Furthermore, geometric data units precede attribute data units within the same slice. Therefore, the point cloud data receiving device first processes (decodes) the geometric data units (or geometric data) and then processes the attribute data units (or attribute data) based on the processed geometric data. The signaling information according to the implementation may be referred to as signaling data, metadata, etc., but is not limited thereto.
[0341] According to the implementation, the signaling information includes a Sequence Parameter Set (SPS), a Geometric Parameter Set (GPS), and one or more Attribute Parameter Sets (APS). The SPS encodes information about the entire sequence, such as a profile or hierarchy, and may include comprehensive information about the entire sequence (sequence hierarchy), such as screen resolution and video format. The GPS is information about geometric encoding applied to the geometry included in the sequence (bitstream). The GPS may include information about an octree (e.g., reference...). The description includes information about the octree and its depth. APSs are information about the attribute encoding applied to the attributes included in the sequence (bitstream). As shown, the bitstream contains one or more APSs (e.g., APS0, APS1… in the figure) based on the identifiers used to identify the attributes.
[0342] According to the implementation, the signaling information may further include TPS. TPS is information about the piece and may include information about the piece identifier, piece size, etc. According to the implementation, the signaling information is applied to the corresponding bitstream as information about the sequence, i.e., the bitstream level. Furthermore, the signaling information has a syntactic structure including syntactic elements and descriptors describing the syntactic elements. Pseudocode for describing the syntax can be used. Additionally, the point cloud receiving device can sequentially parse and process the syntactic elements appearing in the syntax.
[0343] Although not shown in the figure, the geometric data unit and attribute data unit each include a geometric header and an attribute header, respectively. The geometric header and attribute header are signaling information applied at the corresponding slice level and have the syntactic structure described above.
[0344] The geometry header includes information (or signaling information) for processing the corresponding geometric data unit. Therefore, the geometry header appears first within the geometric data unit. The point cloud receiving device can process the geometric data unit by first parsing the geometry header. The geometry header is associated with GPS, which contains information about the entire geometry. Therefore, the geometry header contains information specifying the gps_geom_parameter_set_id included in the GPS. Furthermore, the geometry header contains slice information (e.g., tile_id) related to the slice to which the geometric data unit belongs, as well as a slice identifier.
[0345] The attribute header contains information (or signaling information) for processing the corresponding attribute data unit. Therefore, the attribute header appears first within the attribute data unit. The point cloud receiving device can process the attribute data unit by first parsing the attribute header. The attribute header is associated with the APS, which contains information about all attributes. Therefore, the attribute header contains information specifying the aps_attr_parameter_set_id included in the APS. As mentioned above, attribute decoding is based on geometry decoding. Therefore, the attribute header contains information specifying the slice identifier contained in the geometry header to determine the geometric data unit associated with the attribute data unit.
[0346] When the point cloud data transmission device is used by the application When the maximum nearest neighbor distance (MRD) described in the document generates a set of neighboring points and performs attribute encoding, the signaling information in the bitstream may include information related to the MRD. According to an implementation, the MRD-related information may be included in the sequence-level signaling information (e.g., SPS, APS, etc.) or in the slice-level (e.g., attribute header).
[0347] This is an example of signaling information according to the implementation method.
[0348] Indicates reference The syntactic structure of SPS is described, and its relationship with the reference is shown. The information related to the maximum nearest neighbor distance described is included in the example of the SPS at the sequence level.
[0349] The syntax of the SPS according to the implementation method includes the following syntactic elements.
[0350] `profile_idc`: Indicates the profile applied to the bitstream. This profile specifies the constraints applied to the bitstream to define its capabilities for bitstream decoding. Each profile, as a subset of algorithmic features and constraints, is supported by all decoders that conform to that profile. It is intended for decoding and can be defined according to standards.
[0351] profile_compatibility_flag: Indicates whether the bitstream conforms to a specific profile or another profile used for decoding.
[0352] sps_num_attribute_sets: Indicates the number of encoded attributes in the bitstream. The value of sps_num_attribute_sets should be in the range of 0 to 63.
[0353] The for statement following sps_num_attribute_sets includes elements indicating information about the individual attributes, which is the same number indicated by sps_num_attribute_sets. In this diagram, i represents the individual attribute (or set of attributes), and the value of i is greater than or equal to 0 and less than the number indicated by sps_num_attribute_sets.
[0354] `attribute_dimension_minus1[i]`: Incrementing by 1 specifies the number of components for the i-th attribute. When the attribute is color, it corresponds to a three-dimensional signal representing the light characteristics of the target point. For example, this attribute can be represented as a signal with three RGB (red, green, blue) components. This attribute can also be represented as a signal with three YUV components: luminance (brightness) and two chromaticity (saturation). When the attribute is reflectance, it corresponds to a one-dimensional signal representing the ratio of light reflection intensity at the target point.
[0355] `attribute_instance_id[i]`: Specifies the instance ID used for the i-th attribute. The value of `attribute_instance_id` can be used to distinguish attributes with the same attribute label. For example, it is useful for point clouds that have multiple colors when viewed from different angles.
[0356] As mentioned above, the SPS syntax contains information about the maximum nearest neighbor distance. The following elements represent information about the reference... Information describing the maximum nearest neighbor distance.
[0357] nn_base_distance_calculation_method_type: Indicates the type of base distance calculation method applied in the attribute encoding of the bitstream (sequence). The method for calculating the maximum nearest neighbor distance is specified according to the value of nn_base_distance_calculation_method_type as follows:
[0358] 0: Use the input baseline (maximum) distance;
[0359] 1: Calculation of the maximum nearest neighbor distance based on an octree;
[0360] 2: Calculation of the maximum nearest neighbor distance based on distance;
[0361] 3: Calculation of the maximum nearest neighbor distance based on the sampled neighboring points;
[0362] 4: Calculate the maximum nearest neighbor distance by calculating the average difference of the Morton codes for each LOD; and
[0363] 5: Calculate the maximum nearest neighbor distance of adjacent points by calculating the average distance difference for each LOD.
[0364] When the value of nn_base_distance_calculation_method_type is 0, the SPS syntax includes the following elements.
[0365] nn_base_distance: Represents the base distance value used when calculating the maximum nearest neighbor distance.
[0366] nearest_neighbour_max_range[i]: Indicates the maximum range of neighboring points when encoding attributes of the bitstream (e.g., refer to...). The maximum range of the neighboring point search range described is NN_range.
[0367] nearest_neighbour_min_range[i]: Indicates the minimum range of neighboring points when encoding the attributes of the bitstream (e.g., refer to...). The minimum range of the neighboring point search range described is NN_range.
[0368] NN_range_filtering_location_type: Indicates the application method of maximum nearest neighbor distance when encoding the attributes of the bitstream. The method is specified according to the value of NN_range_filtering_location_type as follows:
[0369] 0: Calculate the distance between each point and apply the maximum nearest neighbor distance when registering adjacent points; and
[0370] 1: Register all neighboring points and remove only those points that exceed the maximum neighbor distance among the registered points.
[0371] An example of signaling information according to an implementation method is shown.
[0372] Indicates reference The syntactic structure of the APS is described, and its relationship with the reference is shown. The information related to the maximum nearest neighbor distance described is included in the example of the APS at the sequence level.
[0373] According to the implementation method, the syntax of APS includes the following syntactic elements.
[0374] `aps_attr_parameter_set_id`: Provides an identifier for the APS for reference by other syntax elements. The value of `aps_attr_parameter_set_id` should be in the range of 0 to 15 (inclusive). One or more attribute data units are contained in the bitstream (e.g., reference...). The described bitstream contains attribute data units (APS), and each attribute data unit includes an attribute header. The attribute header includes a field (e.g., ash_attr_parameter_set_id) that has the same value as aps_attr_parameter_set_id. The point cloud receiving apparatus according to the embodiment parses the APS and processes attribute data units referencing the same aps_attr_parameter_set_id based on the parsed APS and attribute header.
[0375] aps_seq_parameter_set_id: Specifies the value of sps_seq_parameter_set_id for the active SPS. The value of aps_seq_parameter_set_id should be in the range of 0 to 15 (inclusive).
[0376] `attr_coding_type`: Indicates the encoding type used for an attribute for a given value of `attr_coding_type`. Attribute encoding (coding) means attribute coding (encoding). As mentioned above, attribute coding uses at least one of RAHT encoding, predictive transform coding, and lifting transform coding, and `attr_coding_type` indicates any of these three encoding types. Therefore, the value of `attr_coding_type` in the bitstream is equal to any one of 0, 1, or 2. Other values of `attr_coding_type` may be used later by ISO / IEC. Therefore, the point cloud receiving apparatus according to the embodiment ignores `attr_coding_type` values other than 0, 1, and 2. When `attr_coding_type` equals 0, the attribute encoding type is predictive transform coding. When `attr_coding_type` equals 1, the attribute encoding type is RAHT encoding. When `attr_coding_type` equals 2, the attribute encoding type is lifting transform coding. The value of `attr_coding_type` and the encoding type indicated by that value can be changed. For example, when the value is 0, the attribute encoding type can be RAHT encoding.
[0377] The following are the syntactic elements included in the APS when the attribute encoding is a lifting transform encoding or a predictive transform encoding.
[0378] When the attribute encoding is a lifting transform encoding, the APS includes syntactic elements related to the maximum nearest neighbor distance mentioned above.
[0379] different_NN_range_in_tile_flag: Indicates whether to use different maximum / minimum ranges of adjacent points for each tile.
[0380] different_NN_range_per_lod_flag: Indicates whether to use different maximum / minimum ranges for adjacent points for each LOD.
[0381] When `different_NN_range_per_lod_flag` indicates that the maximum / minimum ranges of neighboring points are not used differently for each LOD, the ASP may include at least one or more of the following syntax elements: `nearest_neighbour_max_range`, `nearest_neighbour_min_range`, `nn_base_distance_calculation_method_type`, `nn_base_distance`, and `NN_range_filtering_type`. The descriptions of each element are as follows... The description is the same as the previous one, so it is skipped.
[0382] When `different_NN_range_per_lod_flag` indicates different maximum / minimum ranges of adjacent points for the corresponding LOD, the APS according to the implementation includes information for each LOD. The following statements include elements indicating information for each LOD that is the same as the indicated number `num_detail_levels_minus1`. `idx` shown in the figure represents each LOD, and the value of `idx` is greater than or equal to 0 and less than or equal to the number indicated by `num_detail_levels_minus1`.
[0383] According to the implementation, the ASP includes at least one or more of the elements nearest_neighbour_max_range[idx], nearest_neighbour_min_range[idx], nn_base_distance_calculation_method_type[idx], nn_base_distance[idx], and NN_range_filtering_location_type[idx] for each LOD. The descriptions of each element are... The description is the same as the previous one, so it is skipped.
[0384] An example of signaling information according to an implementation method is shown.
[0385] Indicates reference The syntactic structure of TPS is described, and its comparison with the reference is shown. The information related to the maximum nearest neighbor distance described is included in the example of TPS at the sequence level.
[0386] The syntax of TPS according to the implementation method includes the following syntax elements.
[0387] `num_tiles` indicates the number of tiles signaled for the bitstream. When no tiles are signaled for the bitstream, this value is inferred to be 0. The `for` statement following `num_tiles` has signaling parameters for each tile. In this statement, `i` is a parameter representing each tile and has a value greater than or equal to 0 and less than the value indicated by `num_tiles`.
[0388] tile_bounding_box_offset_x[i] represents the x-offset of the i-th tile in the Cartesian coordinate system. When this parameter is not present, the value of tile_bounding_box_offset_x[0] is regarded as the value of sps_bounding_box_offset_x included in SPS.
[0389] tile_bounding_box_offset_y[i] represents the y-offset of the i-th tile in the Cartesian coordinate system. When this parameter is not present, the value of tile_bounding_box_offset_y[0] is regarded as the value of sps_bounding_box_offset_y included in SPS.
[0390] When reference The `different_NN_range_intile_flag` in the described ASP indicates the different maximum / minimum ranges of adjacent points used for each tile (`different_NN_range_in_tile_flag == true`). The TPS syntax, according to the implementation, includes at least one or more of the elements `nearest_neighbour_max_range[i]`, `nearest_neighbour_min_range[i]`, `nn_base_distance_calculation_method_type[i]`, `nn_base_distance[i]`, and `NN_range_filtering_location_type[i]`. The descriptions of each element are... The description is the same as the previous one, so it is skipped.
[0391] The TPS according to this embodiment also includes the following elements.
[0392] different_NN_range_in_slice_flag[i]: Indicates whether the maximum / minimum range of different neighboring points is used for the slice in piece i.
[0393] When the maximum / minimum range of different neighboring points is used for a slice in piece i (different_NN_range_in_slice_flag[i] == true), the TPS syntax includes the following elements.
[0394] nearest_neighbor_offset_range_in_slice_flag[i]: Indicates whether the maximum / minimum range of neighboring points defined in the slice is indicated by the range offset in the maximum / minimum range of neighboring points defined in the tile or by the absolute value.
[0395] An example of signaling information according to an implementation method is shown.
[0396] Indicates reference The syntactic structure of the attribute header is described, and its references are shown. The information describing the maximum nearest neighbor distance is included in the example of the attribute header at the slice level. Although two syntaxes are shown separately in the figure for simplicity, these two syntaxes constitute a single attribute header syntax.
[0397] The syntax of the attribute header according to the implementation method includes the following syntax elements.
[0398] ash_attr_parameter_set_id has the same value as aps_attr_parameter_set_id for the active SPS.
[0399] ash_attr_sps_attr_idx specifies the order of attribute sets in the active SPS. The value of ash_attr_sps_attr_idx falls within the range from 0 to the value of sps_num_attribute_set included in the active SPS.
[0400] ash_attr_geom_slice_id indicates the value of the slice ID (e.g., gsh_slice_id) included in the geometry header.
[0401] When slices use different maximum / minimum ranges of adjacent points (see reference) The TPS description of different_NN_range_in_slice_flag == true, the attribute header syntax includes the following elements.
[0402] different_NN_range_per_lod_flag: Indicates whether to use different maximum / minimum ranges of adjacent points for the corresponding LOD.
[0403] When the different maximum / minimum ranges of adjacent points are not used for each LOD, the attribute header syntax includes the following elements.
[0404] When nearest_neighbour_offset_range_in_slice_flag (e.g., included in reference) When the nearest_neighbour_offset_range_in_slice_flag in the TPS description indicates the maximum / minimum range of neighboring points defined in the slice, indicated by absolute values, the attribute header syntax includes nearest_neighbour_absolute_max_range and nearest_neighbour_absolute_min_range as elements.
[0405] nearest_neighbour_absolute_max_range: Indicates the maximum range of neighboring points.
[0406] `nearest_neighbour_absolute_min_range`: Indicates the minimum range of neighboring points. The attribute header syntax includes at least one or more of `nn_base_distance_calculation_method_type`, `nn_base_distance`, or `NN_range_filtering_location_type`. The descriptions of each element are... The description is the same as the previous one, so it is skipped.
[0407] When nearest_neighbour_offset_range_in_slice_flag (e.g., included in reference) The attribute header syntax, which describes the `nearest_neighbour_offset_range_in_slice_flag` in TPS, indicates that the maximum / minimum range of neighboring points defined in the slice is indicated by the range offset in the maximum / minimum range of neighboring points defined in the tile. The elements include `nearest_neighbour_max_range_offset` and `nearest_neighbour_min_range_offset`.
[0408] nearest_neighbour_max_range_offset: Indicates the maximum range offset of a slice's neighboring points. The baseline is the maximum range of neighboring points of the tile to which the slice belongs.
[0409] nearest_neighbour_min_range_offset: Indicates the minimum range offset of a slice's neighboring points. The baseline is the minimum range of neighboring points of the tile to which the slice belongs.
[0410] The attribute header syntax includes at least one or more of the elements nn_base_distance_calculation_method_type, nn_base_distance, and NN_range_filtering_location_type. The descriptions of each element are as follows: The description is the same as the previous one, so it is skipped.
[0411] When different maximum / minimum ranges of adjacent points are used for each LOD, the attribute header syntax also includes information about the maximum nearest neighbor distance for each LOD.
[0412] When nearest_neighbour_offset_range_in_slice_flag (e.g., included in reference) When the `nearest_neighbour_offset_range_in_slice_flag` in the TPS description indicates the maximum / minimum range of neighboring points defined in the slice, indicated by absolute values, the attribute header syntax includes the following elements. The following statements have signaling parameters for each LOD. `idx` is a parameter representing each LOD and has a value greater than or equal to 0 and less than the value indicated by `num_detail_level_munus1`. The attribute header syntax includes at least one or more of the elements `nearest_neighbour_absolute_max_range[idx]`, `nearest_neighbour_absolute_min_range[idx]`, `nn_base_distance_calculation_method_type[idx]`, `nn_base_distance[idx]`, and `NN_range_filtering_location_type[idx]` for each LOD. The descriptions of the individual elements are the same as above and are therefore skipped.
[0413] When nearest_neighbour_offset_range_in_slice_flag (e.g., included in reference) When the `nearest_neighbour_offset_range_in_slice_flag` in the TPS description indicates that the maximum / minimum range of neighboring points defined in the slice is indicated by the range offset in the maximum / minimum range of neighboring points defined in the tile, the attribute header syntax includes the following elements. The following statements have signaling parameters for each LOD. `idx` is a parameter representing each LOD and has a value greater than or equal to 0 and less than the value indicated by `num_detail_level_munus1`. The attribute header syntax includes at least one or more of the elements `nearest_neighbour_max_range_offset[idx]`, `nearest_neighbour_min_range_offset[idx]`, `nn_base_distance_calculation_method_type[idx]`, `nn_base_distance[idx]`, and `NN_range_filtering_location_type[idx]` for each LOD. These elements are the same as those described above, therefore their description is skipped.
[0414] This is a flowchart illustrating the operation of a point cloud receiving device according to an embodiment.
[0415] The operation of the point cloud receiving device corresponds to the reference The operation of the point cloud transmitting device is described.
[0416] A point cloud receiving device according to an embodiment (e.g., The receiving device 10004 and Point cloud decoder, and The receiving device receives the point cloud bit stream (e.g., refer to...). The described bitstream (3000). See reference... The bitstream contains signaling information, encoded geometry, and encoded attributes.
[0417] The point cloud receiving device obtains the geometric bit stream and attribute bit stream (3010, 3020) from the received bit stream. The point cloud receiving device obtains the geometric bit stream (or geometric data unit) and attribute bit stream (attribute data unit) based on the signaling information in the bit stream.
[0418] The point cloud receiving device performs entropy decoding (3011) on the acquired geometric bit stream. The point cloud receiving device can perform entropy decoding, which is based on... The inverse process of entropy coding 1522 is described. As mentioned above, entropy coding operations may include exponential Golomb, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC), and entropy decoding operations may include exponential Golomb, CAVLC, and CABAC based on the entropy coding operations.
[0419] The point cloud receiving device performs geometric decoding (3012) on the entropy-decoded geometry. According to embodiments, geometric decoding may include octree geometric decoding and ternary geometric decoding, but is not limited thereto. The point cloud receiving device performs reference... The described operation includes at least one of the following: an arithmetic decoder (arithmetic decoder) 12000, an octree synthesizer (synthesized octree) 12001, a surface approximation synthesizer (synthesized surface approximation) 12002, a geometry reconstructor (reconstructed geometry) 12003, and an inverse coordinate transformer (inverse coordinate transform) 12004. Furthermore, the point cloud receiving device performs a reference... The described operation comprises at least one of the following: an arithmetic decoder 13002, an octet-based octree reconstruction processor 13003, a surface model processor (triangle reconstruction, upsampling, voxelization) 13004, and an inverse quantization processor 13005. The point cloud receiving device outputs the reconstructed geometric information as the result of geometric decoding. The reconstructed geometric information is used for attribute decoding.
[0420] The point cloud receiving device processes the geometry after geometry decoding (3013). For example, as referenced... As described above, when the point cloud transmitting device performs coordinate transformation on the geometry, the point cloud receiving device performs inverse transformation on the coordinates of the geometry to output the geometry.
[0421] The point cloud receiving device according to the embodiment performs entropy decoding (3021) on the obtained attribute bit stream.
[0422] The point cloud receiving device performs attribute decoding (3022) on the attributes obtained from entropy decoding. According to an embodiment, attribute decoding may include at least one or a combination of RAHT coding, predictive transform coding, or lifting transform coding. The point cloud receiving device may generate a Level of Detail (LOD) based on the reconstructed geometric information (octree). Additionally, the point cloud receiving device acquires information about a reference... The signaling information describing the maximum nearest neighbor distance is used to generate a system based on the acquired signaling information. The method for calculating the maximum nearest neighbor distance described in [the text] sets up neighboring points. The methods for generating the neighboring point set and calculating the maximum nearest neighbor distance are described in [the text]. The description is the same.
[0423] The point cloud receiver processes the decoded attributes (3023). For example, when the point cloud transmitter performs a color transformation on the attribute, the point cloud receiver performs an inverse color transformation on the attribute and outputs the attribute.
[0424] This is a flowchart illustrating a point cloud data transmission method according to an embodiment.
[0425] Flowchart 3100 indicates that it is used for reference The described point cloud data transmission device (e.g., refer to...) , and The described method for transmitting point cloud data (using a transmitting device or point cloud encoder). The flowchart is intended to illustrate the point cloud data transmission method, and the order of operations in the point cloud data transmission method is not limited to the order shown in the figure.
[0426] The point cloud data transmitting device encodes point cloud data, including geometry and attributes (3110). Geometry is information indicating the location of points in the point cloud data, and attributes include at least one of the point's color and reflectivity. Attributes are encoded based on one or more Levels of Detail (LODs) generated by reorganizing the points, and one or more neighboring points around the points belonging to each LOD are selected based on the maximum nearest neighbor distance. Maximum nearest neighbor distance and reference... The description is identical, therefore its description is skipped. The point cloud data transmission device is for the reference... The geometry and attributes are encoded. The point cloud data transmitting device (or attribute encoder) can generate one or more LODs based on the octree of the encoded geometry and perform predictive transformations based on one or more selected neighboring points. (See reference...) The distance between each neighboring point and the predicted transformation target point is less than or equal to the maximum nearest neighbor distance. A geometry-based octree generates one or more Levels of Detail (LODs). According to an implementation, the octree has a reference... The level of each LOD corresponds to one or more depths in the octree. Therefore, each LOD level corresponds to a depth or one or more depths in the octree. Furthermore, as referenced... The octree mentioned above comprises one or more octree nodes. The maximum nearest neighbor distance is represented as... Here, LOD is a parameter indicating the LOD level of the predicted transformation target point, and NN_range is a reference... The description of the neighbor search range. NN_range represents the number of one or more octree nodes surrounding the point. The description of the neighbor search range is related to... The description is identical to that described above, therefore it is skipped. The point cloud data transmitting device transmits a bit stream containing encoded point cloud data. The bit stream (e.g., see reference...) The described bitstream contains information related to the maximum nearest neighbor distance (e.g., reference). The SPS and APS syntax is described. According to the implementation, information related to the maximum nearest neighbor distance includes `nearest_neighbour_max_range`. That is, as described in reference... As described, information related to the maximum nearest neighbor distance can be sent at the sequence level or the slice level.
[0427] This is a flowchart illustrating a method for processing point cloud data according to an embodiment.
[0428] Flowchart 3200 shows a reference The point cloud data processing method of the described point cloud data processing apparatus (e.g., receiving device 10004 or point cloud video decoder 10006). The flowchart is intended to illustrate the point cloud data processing method, and the order of operations in the point cloud data processing method is not limited to the order shown in the figure.
[0429] Point cloud data processing device (e.g., receiver, The receiver receives a bitstream (3210) containing point cloud data including geometry and attributes. Geometry is information indicating the location of points in the point cloud data, and attributes include at least one of the point's color and reflectivity.
[0430] Point cloud data processing device (e.g., The decoder decodes the point cloud data (3220). The attributes are decoded based on one or more Levels of Detail (LODs) generated by reorganizing the points, and one or more neighboring points around the points belonging to each LOD are selected based on the maximum nearest neighbor distance. Maximum nearest neighbor distance and reference... The description is identical, therefore its description is skipped. Point cloud data receiving device (e.g., A geometry decoder decodes the geometry included in the point cloud data. A point cloud data receiving device (e.g., The attribute decoder decodes the attributes. The point cloud data receiving device generates one or more Levels of Dependencies (LODs) based on the octree of the decoded geometry, and performs a prediction transformation on points included in each LOD based on one or more selected neighbor pairs. The distance between each neighbor point and the target point of the prediction transformation is less than or equal to the maximum nearest neighbor distance. One or more LODs are generated based on the geometry-based octree. According to the implementation, the octree has a reference... The level of each LOD corresponds to one or more depths in the octree. Therefore, each LOD level corresponds to a depth or one or more depths in the octree. Furthermore, as referenced... The octree described herein comprises one or more octree nodes. (See reference...) The maximum nearest neighbor distance is calculated based on the search range of adjacent points. The maximum nearest neighbor distance is expressed as...
[0431] Here, LOD is a parameter indicating the LOD level of the predicted transformation target point, and NN_range is a reference... The description of the neighbor search range. NN_range represents the number of one or more octree nodes surrounding the point. The description of the neighbor search range is related to... The description is the same as the previous one, so it is skipped.
[0432] Bitstream (e.g., reference) The described bitstream contains information related to the maximum nearest neighbor distance (e.g., reference). The SPS and APS syntax is described. According to the implementation, information related to the maximum nearest neighbor distance includes `nearest_neighbour_max_range`. That is, as described in reference... As described, information related to the maximum nearest neighbor distance can be transmitted at the sequence level or the slice level. Therefore, the point cloud data receiving device can perform attribute decoding based on the information related to the maximum nearest neighbor distance contained in the bitstream.
[0433] Reference This describes point cloud data processing according to an embodiment. Components of the apparatus may include one or more processors coupled to a memory. It can be implemented using hardware, software, firmware, or a combination thereof. Components of the apparatus according to an embodiment are chips, for example, they can be implemented as hardware circuits. Furthermore, each component of the point cloud data processing apparatus according to an embodiment can be implemented as a separate chip. Additionally, at least one or more components of the point cloud data processing apparatus according to an embodiment are capable of executing one or more programs. It may consist of one or more processors, and one or more programs are part of the operation / method of the point cloud data processing apparatus described with reference to the accompanying drawings, to perform or execute any one or more operations. It may contain instructions.
[0434] Although the accompanying drawings have been described separately for simplicity, new embodiments can be designed by combining the embodiments shown in the various drawings. Designing a computer-readable recording medium having a program recorded thereon for performing the above embodiments, as needed by those skilled in the art, also falls within the scope of the appended claims and their equivalents. The apparatus and methods according to the embodiments are not limited to the configurations and methods of the above embodiments. Various modifications can be made to the embodiments by selectively combining all or some of the embodiments. Although preferred embodiments have been described with reference to the accompanying drawings, those skilled in the art will understand that various modifications and variations can be made to the embodiments without departing from the spirit or scope of this disclosure as described in the appended claims. These modifications should not be understood in isolation from the technical concept or viewpoint of the embodiments.
[0435] The implementation has been described from the perspective of methods and / or apparatus, and the descriptions of methods and apparatus can be applied to complement each other.
[0436] Various elements of the device according to the embodiments can be implemented by hardware, software, firmware, or a combination thereof. Various elements in the embodiments can be implemented by a single chip, such as a single hardware circuit. According to the embodiments, the components according to the embodiments can be implemented as separate chips. According to the embodiments, at least one or more of the components of the device according to the embodiments can include one or more processors capable of executing one or more programs. One or more programs can execute any one or more operations / methods according to the embodiments, or include instructions for performing the operations / methods. Executable instructions for performing the methods / operations of the device according to the embodiments can be stored in a non-temporary CRM or other computer program product configured to be executed by one or more processors, or can be stored in a temporary CRM or other computer program product configured to be executed by one or more processors. Furthermore, the memory according to the embodiments can be used as a concept that covers not only volatile memory (e.g., RAM) but also non-volatile memory, flash memory, and PROM. Furthermore, it can also be implemented in the form of a carrier wave, such as transmission via the Internet. Furthermore, a processor-readable recording medium can be distributed to a computer system connected via a network, such that processor-readable code can be and executed.
[0437] In this specification, the terms “ / ” and “,” should be interpreted as meaning “and / or”. For example, the expression “A / B” can mean “A and / or B”. Furthermore, “A, B” can mean “A and / or B”. Furthermore, “A / B / C” can mean “at least one of A, B, and / or C”. Furthermore, “A / B / C” can mean “at least one of A, B, and / or C”. Additionally, in this specification, the term “or” should be interpreted as meaning “and / or”. For example, the expression “A or B” can mean 1) only A, 2) only B, or 3) both A and B. In other words, the term “or” as used herein should be interpreted as meaning “additionally or alternatively”.
[0438] Terms such as "first" and "second" can be used to describe various elements of the embodiments. However, the various components according to the embodiments should not be limited by the terms used above. These terms are used only to distinguish one element from another. For example, a first user input signal may be referred to as a second user input signal. Similarly, a second user input signal may be referred to as a first user input signal. The use of these terms should be interpreted without departing from the scope of the various embodiments. Both a first user input signal and a second user input signal are user input signals, but do not necessarily mean the same user input signal, unless the context clearly specifies otherwise.
[0439] The terminology used to describe embodiments is for the purpose of describing a particular embodiment and is not intended to limit the embodiments. As used in the description of embodiments and claims, the singular forms “a,” “the,” and “the” include plural indicators unless the context clearly indicates otherwise. The expression “and / or” is used to include all possible combinations of terms. Terms such as “comprising” or “having” are intended to indicate the presence of drawings, numbers, steps, elements, and / or components, and should be understood not to exclude the possibility of the additional presence of drawings, numbers, steps, elements, and / or components. As used herein, conditional expressions such as “if” and “when” are not limited to optional cases and are intended to be interpreted as performing the relevant operation or interpreting the relevant definition according to the specific condition when a particular condition is met.
[0440] Invention Model
[0441] As described above, the relevant content has been described in the best manner for implementing the embodiments.
[0442] Industrial practicality
[0443] It will be apparent to those skilled in the art that various changes and modifications can be made to this invention without departing from its spirit or scope. Therefore, this invention is intended to cover modifications and variations thereof, provided they fall within the scope of the appended claims and their equivalents.
Claims
1. A method for transmitting point cloud data, the method comprising the following steps: Encoding point cloud data that includes geometry and attributes, wherein the geometry represents the position of points in the point cloud data, and the attributes include at least one of the color and reflectivity of the points; as well as Send a bitstream containing encoded point cloud data. The attributes are encoded based on one or more levels of detail (LODs). Specifically, for points with LOD (Level of Detail), the attribute is predicted based on one or more neighboring points. Among them, points that discard one or more neighboring points based on distance, The distance mentioned here is derived based on coefficient three, the LOD, and the range related to the search for neighboring points. Where, if one or more neighboring points of a point based on the LOD exceed the specified distance, the one or more neighboring points of that point are discarded. Wherein, the distance is related to the block size of each LOD, and The bit stream includes information related to the range associated with the neighbor point search.
2. The method according to claim 1, wherein, The steps for encoding the point cloud data include: Encode the geometry; and The attribute is encoded.
3. The method according to claim 2, wherein, The steps for encoding the attribute include: One or more levels of detail (LOD) are generated by rearranging the points; and Predictive transformations are performed by selecting one or more neighboring points of a point within each Level of Dimension (LOD). Wherein, the neighbor distance between the point and each of the adjacent points in the one or more adjacent points is less than or equal to the distance.
4. The method according to claim 3, wherein, The level of one or more LODs corresponds to the depth of an octree of encoded geometry, the octree comprising octree nodes. The distance is represented as , The LOD refers to the LOD level of the point. NN RANGE It is the range related to the search of neighboring points, and NN RANGE This represents the number of one or more octree nodes surrounding the point.
5. An apparatus for transmitting point cloud data, the apparatus comprising: An encoder configured to encode point cloud data including geometry and attributes, the geometry representing the positions of points in the point cloud data, and the attributes including at least one of the color and reflectivity of the points; and A transmitter configured to transmit a bitstream comprising encoded point cloud data. The attributes are encoded based on one or more levels of detail (LODs). Specifically, for points with LOD (Level of Detail), the attribute is predicted based on one or more neighboring points, and Among them, points that discard one or more neighboring points based on distance, The distance mentioned here is derived based on coefficient three, the LOD, and the range related to the search for neighboring points. Where, if one or more neighboring points of a point based on the LOD exceed the specified distance, the one or more neighboring points of that point are discarded. Wherein, the distance is related to the block size of each LOD, and The bit stream includes information related to the range associated with the neighbor point search.
6. The apparatus according to claim 5, in, The encoder includes: A geometry encoder, configured to encode the geometry. An attribute encoder, configured to encode the attribute.
7. The apparatus according to claim 6, in, The attribute encoder is configured as follows: One or more levels of detail (LOD) are generated by rearranging the points; and Predictive transformations are performed by selecting one or more neighboring points of a point within each Level of Dimension (LOD). Wherein, the neighbor distance between the point and each of the adjacent points in the one or more adjacent points is less than or equal to the distance.
8. The apparatus according to claim 7, The level of one or more LODs corresponds to the depth of an octree of encoded geometry, the octree comprising octree nodes. The distance is represented as , The LOD refers to the LOD level of the point. NN RANGE It is the range related to the search of neighboring points, and NN RANGE This represents the number of one or more octree nodes surrounding the point.
9. A method for processing point cloud data, the method comprising the following steps: Receive a bitstream including the point cloud data, the point cloud data including geometry and attributes, the geometry representing the position of points in the point cloud data, and the attributes including at least one of the color and reflectivity of the points; as well as Decode the point cloud data. Specifically, the attribute is decoded based on one or more levels of detail (LODs). Specifically, for points with LOD (Level of Detail), the attribute is predicted based on one or more neighboring points. Among them, points that discard one or more neighboring points based on distance, The distance mentioned here is derived based on coefficient three, the LOD, and the range related to the search for neighboring points. Where, if one or more neighboring points of a point based on the LOD exceed the specified distance, the one or more neighboring points of that point are discarded. Wherein, the distance is related to the block size of each LOD, and The bit stream includes information related to the range associated with the neighbor point search.
10. The method according to claim 9, wherein, The steps for decoding the point cloud data include: Decode the geometry; and Decode the attribute.
11. The method according to claim 10, wherein, The steps for decoding the attribute include: One or more levels of detail (LOD) are generated by rearranging the points; and Predictive transformations are performed by selecting one or more neighboring points of a point within each Level of Dimension (LOD). Wherein, the neighbor distance between the point and each of the adjacent points in the one or more adjacent points is less than or equal to the distance.
12. The method according to claim 11, wherein, The level of one or more LODs corresponds to the depth of an octree of decoded geometry, the octree containing octree nodes, and the distance is calculated based on information related to the distance at which a point is discarded if it exceeds the specified distance, the information related to the distance at which a point is discarded if it exceeds the specified distance representing the number of one or more octree nodes around the point.
13. An apparatus for processing point cloud data, the apparatus comprising: A receiver is configured to receive a bitstream comprising the point cloud data, the point cloud data comprising geometry and attributes, the geometry representing the position of points in the point cloud data, and the attributes comprising at least one of the color and reflectivity of the points. as well as A decoder, used to decode the point cloud data. Specifically, the attribute is decoded based on one or more levels of detail (LODs). Specifically, for points with LOD (Level of Detail), the attribute is predicted based on one or more neighboring points, and Among them, points that discard one or more neighboring points based on distance, The distance mentioned here is derived based on coefficient three, the LOD, and the range related to the search for neighboring points. Where, if one or more neighboring points of a point based on the LOD exceed the specified distance, the one or more neighboring points of that point are discarded. Wherein, the distance is related to the block size of each LOD, and The bit stream includes information related to the range associated with the neighbor point search.
14. The apparatus according to claim 13, wherein, The decoder includes: A geometry decoder, wherein the geometry decoder is used to decode the geometry; and An attribute decoder is used to decode the attribute.
15. The apparatus according to claim 14, wherein, The attribute decoder is also configured to: One or more levels of detail (LOD) are generated by rearranging the points; and Predictive transformations are performed by selecting one or more neighboring points of a point within each Level of Dimension (LOD). Wherein, the neighbor distance between the point and each of the adjacent points in the one or more adjacent points is less than or equal to the distance.
16. The apparatus according to claim 15, wherein, The level of one or more LODs corresponds to the depth of an octree of decoded geometry, the octree containing octree nodes, and the distance is calculated based on information related to the distance at which a point is discarded if it exceeds the specified distance, the information related to the distance at which a point is discarded if it exceeds the specified distance representing the number of one or more octree nodes around the point.
Citation Information
Patent Citations
Point Cloud Compression
US20190080483A1