Point cloud data sending device, sending method, processing device and processing method
By encoding and bitstream processing of point cloud data, the problem of low point cloud data processing efficiency in the prior art is solved, an efficient and simplified encoding/decoding process is realized, and high-quality point cloud services are provided.
Patent Information
- Application Number
- CN202180008152.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-03-23
- Filing Date
- 2021-01-05
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2041-01-05
AI Technical Summary
The prior art is difficult to efficiently process large amounts of point cloud data, resulting in long wait times and high encoding/decoding complexity.
By encoding point cloud data, including encoding geometric structures and attributes, and converting the encoded data into a bitstream for transmission. The specific steps include encoding the geometric structure, representing it using an octree structure, encoding the attributes, and encoding based on the encoded geometric structure.
It realizes efficient processing of point cloud data, reduces waiting time, and simplifies the encoding/decoding process, providing high-quality point cloud services.
Smart Images

Figure CN114930853B_ABST
Abstract
Description
Technical Field
[0001] Embodiments provide for providing point cloud content to provide various services such as VR (Virtual Reality), augmented reality (AR), mixed reality (MR), and autonomous driving services to a user. Background Art
[0002] Point cloud content is content represented by a point cloud, which is a set of points belonging to a coordinate system representing a three-dimensional space. Point cloud content can represent three-dimensionally configured media and is used to provide various services such as virtual reality (VR), augmented reality (AR), mixed reality (MR), and autonomous driving services. However, tens of thousands to hundreds of thousands of point data are required to represent point cloud content. Therefore, a method for efficiently processing a large amount of point data is needed. Summary of the Invention
[0003] Technical Problem
[0004] Embodiments provide an apparatus and method for efficiently processing point cloud data. Embodiments provide a method and apparatus for processing point cloud data to solve latency and encoding / decoding complexity.
[0005] The technical scope of the embodiments is not limited to the above-mentioned technical objectives and may be extended to other technical objectives that those skilled in the art can infer based on all the content disclosed herein.
[0006] Technical Solution
[0007] Therefore, in order to efficiently process point cloud data, a method for transmitting point cloud data according to some embodiments may include the following steps: encoding point cloud data including a geometric structure and attributes; and transmitting a bitstream including the encoded point cloud data. The geometric structure according to some embodiments represents the positions of the points of the point cloud data, and the attributes according to some embodiments include at least one of the color and reflectivity of the points. The step of encoding the point cloud data includes encoding the geometric structure and encoding the attributes based on a complete or partial octree of the encoded geometric structure.
[0008] An apparatus for transmitting point cloud data according to some embodiments may include: an encoder configured to encode point cloud data including a geometric structure and attributes; and a transmitter configured to transmit a bitstream including the encoded point cloud data. The geometric structure according to some embodiments represents the positions of the points of the point cloud data, and the attributes according to some embodiments include at least one of the color and reflectivity of the points. The encoder according to some embodiments includes a geometric structure encoder that encodes the geometric structure and an attribute encoder that encodes the attributes based on a complete or partial octree of the encoded geometric structure.
[0009] A method for processing point cloud data according to some embodiments may include the following steps: receiving a bitstream including point cloud data; and decoding the point cloud data based on signaling information included in the bitstream. The step of decoding the point cloud data includes decoding the geometric structure included in the point cloud data and decoding attributes including at least one of the color and reflectivity of points based on a complete or partial octree of the decoded geometric structure. The geometric structure according to some embodiments represents the positions of the points of the point cloud data.
[0010] An apparatus for processing point cloud data according to some embodiments may include: a receiver that receives a bitstream including point cloud data; and a decoder that decodes the point cloud data based on signaling information included in the bitstream. The decoder includes: a geometric structure decoder that decodes the geometric structure included in the point cloud data; and an attribute decoder that decodes attributes including at least one of the color and reflectivity of points based on a complete or partial octree of the decoded geometric structure. The geometric structure according to some embodiments represents the positions of the points of the point cloud data.
[0011] Advantageous Effects
[0012] The apparatus / method according to some embodiments efficiently processes point cloud data.
[0013] The apparatus / method according to some embodiments provides a high-quality point cloud service.
[0014] The apparatus / method according to some embodiments provides point cloud content for providing general services such as VR services, autonomous driving services, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings, which are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this application, illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure. For a better understanding of the various embodiments described below, reference should be made to the description of the following embodiments in conjunction with the accompanying drawings. The same reference numerals will be used throughout the drawings to refer to the same or similar components.
[0016] Figure 1 An exemplary point cloud content providing system according to an embodiment is shown;
[0017] Figure 2 is a block diagram illustrating a point cloud content providing operation according to an embodiment;
[0018] Figure 3 An exemplary process of capturing a point cloud video according to an embodiment is illustrated;
[0019] Figure 4Illustrates an exemplary point cloud encoder according to an embodiment;
[0020] Figure 5 Shows an example of a voxel according to an embodiment;
[0021] Figure 6 Shows an example of an octree and occupancy code according to an embodiment;
[0022] Figure 7 Shows an example of a neighboring node pattern according to an embodiment;
[0023] Figure 8 Illustrates an example of a point configuration in each LOD according to an embodiment;
[0024] Figure 9 Illustrates an example of a point configuration in each LOD according to an embodiment;
[0025] Figure 10 Illustrates a point cloud decoder according to an embodiment;
[0026] Figure 11 Illustrates a point cloud decoder according to an embodiment;
[0027] Figure 12 Illustrates a transmitting device according to an embodiment;
[0028] Figure 13 Illustrates a receiving device according to an embodiment;
[0029] Figure 14 Illustrates an exemplary structure operable in conjunction with a point cloud data transmission / reception method / device according to an embodiment;
[0030] Figure 15 Is a flowchart illustrating an example of point cloud encoding;
[0031] Figure 16 Illustrates an example of a point and its neighboring points;
[0032] Figure 17 Shows a diagram of an exemplary bitstream structure;
[0033] Figure 18 Shows an example of signaling information according to an embodiment;
[0034] Figure 19 Illustrates an example of signaling information according to an embodiment;
[0035] Figure 20 Illustrates a method for encoding relevant weights according to an embodiment;
[0036] Figure 21Illustrates an example of spatial scalability decoding;
[0037] Figure 22 Shows an exemplary improved quantization weight derivation process;
[0038] Figure 23 Is a flowchart illustrating point cloud encoding according to an embodiment;
[0039] Figure 24 Is a flowchart illustrating a method of transmitting point cloud data according to an embodiment;
[0040] Figure 25 Is a flowchart illustrating a method of processing point cloud data according to an embodiment. Detailed Description
[0041] Now, reference will be made in detail to the preferred embodiments of the present disclosure, examples of which are illustrated in the accompanying drawings. The following detailed description made with reference to the accompanying drawings is intended to explain the exemplary embodiments of the present disclosure, and not to show the only embodiments that can be implemented according to the present disclosure. The following detailed description includes specific details to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure can be practiced without these specific details.
[0042] Although most of the terms used in the present disclosure have been selected from commonly used general terms widely used in the art, some terms have been arbitrarily selected by the applicant, and their meanings will be explained in detail as needed in the following description. Therefore, the present disclosure should be understood based on the intended meaning of the terms rather than their simple names or meanings.
[0043] Figure 1 Shows an exemplary point cloud content providing system according to an embodiment.
[0044] Figure 1 The point cloud content providing system illustrated in may include a transmitting device 10000 and a receiving device 10004. The transmitting device 10000 and the receiving device 10004 are capable of wired or wireless communication to transmit and receive point cloud data.
[0045] The point cloud data transmitting device 10000 according to an embodiment can obtain and process a point cloud video (or point cloud content), and transmit the point cloud video (or point cloud content). According to an embodiment, the transmitting device 10000 can include a fixed station, a base transceiver system (BTS), a network, an artificial intelligence (AI) device and / or system, a robot, an AR / VR / XR device and / or a server. According to an embodiment, the transmitting device 10000 can include a device, a robot, a vehicle, an AR / VR / XR device, a portable device, a home appliance, an Internet of Things (IoT) device, and an AI device / server configured to perform communication with a base station and / or other wireless devices using a radio access technology (e.g., 5G new radio access technology (NR), long term evolution (LTE)).
[0046] The transmitting device 10000 according to an embodiment includes a point cloud video acquirer 10001, a point cloud video encoder 10002, and / or a transmitter (or communication module) 10003.
[0047] The point cloud video acquirer 10001 according to an embodiment acquires a point cloud video through a processing process such as capturing, synthesizing, or generating. The point cloud video is point cloud content represented as a set of points in a 3D space, and can be referred to as point cloud video data. The point cloud video according to an embodiment can include one or more frames. A frame represents a still image / picture. Therefore, the point cloud video can include point cloud images / frames / pictures, and can be referred to as a point cloud image, a frame, or a picture.
[0048] The point cloud video encoder 10002 according to an embodiment encodes the acquired point cloud video data. The point cloud video encoder 10002 can encode the point cloud video data based on point cloud compression coding. The point cloud compression coding according to an embodiment can include geometry-based point cloud compression (G-PCC) coding and / or video-based point cloud compression (V-PCC) coding or next-generation coding. The point cloud compression coding according to an embodiment is not limited to the above embodiment. The point cloud video encoder 10002 can output a bitstream including the encoded point cloud video data. The bitstream can include not only the encoded point cloud video data, but also signaling information related to the encoding of the point cloud video data.
[0049] The transmitter 10003 according to an embodiment transmits a bitstream including encoded point cloud video data. The bitstream according to an embodiment is encapsulated in a file or segment (e.g., a streaming segment) and transmitted through various networks such as a broadcast network and / or a broadband network. Although not shown in the figure, the transmitting device 10000 may include an encapsulator (or encapsulation module) configured to perform an encapsulation operation. According to an embodiment, the encapsulator may be included in the transmitter 10003. According to an embodiment, the file or segment may be transmitted to the receiving device 10004 through a network or stored in a digital storage medium (e.g., USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.). The transmitter 10003 according to an embodiment is capable of performing wired / wireless communication with the receiving device 10004 (or the receiver 10007) through networks such as 4G, 5G, 6G, etc. Additionally, the transmitter may perform necessary data processing operations according to the network system (e.g., 4G, 5G, or 6G communication network system). The transmitting device 10000 may transmit the encapsulated data in an on-demand manner.
[0050] The receiving device 10004 according to an embodiment includes a receiver 10007, a point cloud video decoder 10006, and / or a renderer 10005. According to an embodiment, the receiving device 10004 may include devices, robots, vehicles, AR / VR / XR devices, portable devices, household appliances, Internet of Things (IoT) devices, and AI devices / servers configured to communicate with a base station and / or other wireless devices using radio access technologies (e.g., 5G new radio access technology (NR), long-term evolution (LTE)).
[0051] The receiver 10007 according to an embodiment receives a bitstream including point cloud video data or a file / segment in which the bitstream is encapsulated from a network or a storage medium. The receiver 10007 may perform necessary data processing according to the network system (e.g., a communication network system such as 4G, 5G, 6G, etc.). The receiver 10007 according to an embodiment may decompose the received file / segment and output a bitstream. According to an embodiment, the receiver 10007 may include a decompressor (or decompression module) configured to perform a decompression operation. The decompressor may be implemented as an element (or component) separate from the receiver 10007.
[0052] The point cloud video decoder 10006 decodes a bitstream including point cloud video data. The point cloud video decoder 10006 may decode the point cloud video data according to the method of encoding the point cloud video data (e.g., in the reverse process of the operation of the point cloud video encoder 10002). Therefore, the point cloud video decoder 10006 may decode the point cloud video data by performing point cloud decompression encoding, which is the reverse process of point cloud compression. The point cloud decompression encoding includes G-PCC encoding.
[0053] The renderer 10005 renders the decoded point cloud video data. The renderer 10005 may output point cloud content by rendering not only the point cloud video data but also audio data. According to an embodiment, the renderer 10005 may include a display configured to display the point cloud content. According to an embodiment, the display may be implemented as a separate device or component instead of being included in the renderer 10005.
[0054] The arrow indicated by the dashed line in the figure represents the transmission path of the feedback information acquired by the receiving device 10004. The feedback information is information for reflecting the interactivity with the user consuming the point cloud content and includes information about the user (e.g., head orientation information, viewport information, etc.). In particular, when the point cloud content is content of a service that requires interaction with the user (e.g., an autonomous driving service, etc.), the feedback information may be provided to the content sender (e.g., the transmitting device 10000) and / or the service provider. According to an embodiment, the feedback information may be used in the receiving device 10004 and the transmitting device 10000, or may not be provided.
[0055] The head orientation information according to an embodiment is information regarding the user's head position, orientation, angle, movement, etc. The receiving device 10004 according to an embodiment can calculate viewport information based on the head orientation information. The viewport information can be information regarding the area of the point cloud video that the user is viewing. The viewing point is the point through which the user is viewing the point cloud video, and can refer to the center point of the viewport area. That is, the viewport is an area centered on the viewing point, and the size and shape of this area can be determined by the field of view (FOV). Therefore, in addition to the head orientation information, the receiving device 10004 can also extract viewport information based on the vertical or horizontal FOV supported by the device. In addition, the receiving device 10004 performs gaze analysis, etc., to examine the way the user consumes the point cloud, the area in the point cloud video where the user gazes, the gaze time, etc. According to an embodiment, the receiving device 10004 can send feedback information including the gaze analysis result to the sending device 10000. The feedback information according to an embodiment can be obtained during the rendering and / or display processing. The feedback information according to an embodiment can be obtained by one or more sensors included in the receiving device 10004. According to an embodiment, the feedback information can be obtained by the renderer 10005 or a separate external element (or device, component, etc.). Figure 1 The dashed line in represents the process of sending the feedback information obtained by the renderer 10005. The point cloud content providing system can process (encode / decode) the point cloud data based on the feedback information. Therefore, the point cloud video decoder 10006 can perform a decoding operation based on the feedback information. The receiving device 10004 can send the feedback information to the sending device 10000. The sending device 10000 (or the point cloud video encoder 10002) can perform an encoding operation based on the feedback information. Therefore, the point cloud content providing system can efficiently process the necessary data (e.g., the point cloud data corresponding to the user's head position) based on the feedback information instead of processing (encoding / decoding) the entire point cloud data, and provide the point cloud content to the user.
[0056] According to an embodiment, the sending device 10000 can be referred to as an encoder, a sending device, a transmitter, etc., and the receiving device 10004 can be referred to as a decoder, a receiving device, a receiver, etc.
[0057] In the Figure 1 point cloud content providing system according to an embodiment, the point cloud data processed (through a series of processes of acquisition / encoding / sending / decoding / rendering) can be referred to as point cloud content data or point cloud video data. According to an embodiment, the point cloud content data can be used as a concept covering metadata or signaling information related to the point cloud data.
[0058] Figure 1 The elements of the point cloud content providing system illustrated in can be implemented by hardware, software, a processor, and / or a combination thereof.
[0059] Figure 2 is a block diagram illustrating an operation of providing point cloud content according to an embodiment.
[0060] Figure 2 The block diagram of Figure 1 illustrates the operation of the point cloud content providing system described in. As described above, the point cloud content providing system can process point cloud data based on point cloud compression coding (e.g., G-PCC).
[0061] A point cloud content providing system according to an embodiment (e.g., the point cloud transmitting device 10000 or the point cloud video acquirer 10001) can acquire a point cloud video (20000). The point cloud video is represented by point clouds belonging to a coordinate system for representing a 3D space. The point cloud video according to an embodiment may include a Ply (Polygon File Format or Stanford Triangle Format) file. When the point cloud video has one or more frames, the acquired point cloud video may include one or more Ply files. The Ply file contains point cloud data such as point geometry and / or attributes. The geometry includes the position of the points. The position of each point can be represented by parameters (e.g., values of the X, Y, and Z axes) representing a three-dimensional coordinate system (e.g., a coordinate system composed of the X, Y, and Z axes). The attributes include attributes of the points (e.g., information about texture, color (YCbCr or RGB), reflectance r, transparency, etc. for each point). A point has one or more attributes. For example, a point may have one attribute, i.e., color, or two attributes, i.e., color and reflectance. According to an embodiment, the geometry may be referred to as position, geometry information, geometry data, etc., and the attributes may be referred to as attributes, attribute information, attribute data, etc. The point cloud content providing system (e.g., the point cloud transmitting device 10000 or the point cloud video acquirer 10001) can obtain point cloud data from information (e.g., depth information, color information, etc.) related to the acquisition process of the point cloud video.
[0062] A point cloud content providing system according to an embodiment (e.g., the transmitting device 10000 or the point cloud video encoder 10002) can encode the point cloud data (20001). The point cloud content providing system can encode the point cloud data based on point cloud compression coding. As described above, the point cloud data may include the geometry and attributes of the points. Therefore, the point cloud content providing system can perform geometry encoding for encoding the geometry and output a geometry bitstream. The point cloud content providing system can perform attribute encoding for encoding the attributes and output an attribute bitstream. According to an embodiment, the point cloud content providing system can perform attribute encoding based on geometry encoding. The geometry bitstream and the attribute bitstream according to an embodiment can be multiplexed and output as one bitstream. The bitstream according to an embodiment may also include signaling information related to geometry encoding and attribute encoding.
[0063] A point cloud content providing system according to an embodiment (e.g., the transmitting device 10000 or the transmitter 10003) may transmit encoded point cloud data (20002). As Figure 1 illustrated, the encoded point cloud data may be represented by a geometry bitstream and an attribute bitstream. Additionally, the encoded point cloud data may be transmitted in the form of a bitstream together with signaling information related to the encoding of the point cloud data (e.g., signaling information related to geometry encoding and attribute encoding). The point cloud content providing system may encapsulate the bitstream carrying the encoded point cloud data and transmit the bitstream in the form of a file or a segment.
[0064] A point cloud content providing system according to an embodiment (e.g., the receiving device 10004 or the receiver 10007) may receive a bitstream containing the encoded point cloud data. Additionally, the point cloud content providing system (e.g., the receiving device 10004 or the receiver 10007) may demultiplex the bitstream.
[0065] A point cloud content providing system (e.g., the receiving device 10004 or the point cloud video decoder 10006) may decode the encoded point cloud data (e.g., the geometry bitstream, the attribute bitstream) transmitted in the bitstream. The point cloud content providing system (e.g., the receiving device 10004 or the point cloud video decoder 10006) may decode the point cloud video data based on the signaling information related to the encoding of the point cloud video data contained in the bitstream. The point cloud content providing system (e.g., the receiving device 10004 or the point cloud video decoder 10006) may decode the geometry bitstream to reconstruct the positions (geometry) of the points. The point cloud content providing system may reconstruct the attributes of the points by decoding the attribute bitstream based on the reconstructed geometry. The point cloud content providing system (e.g., the receiving device 10004 or the point cloud video decoder 10006) may reconstruct the point cloud video based on the reconstructed geometry and the decoded attributes based on the positions.
[0066] A point cloud content providing system according to an embodiment (e.g., the receiving device 10004 or the renderer 10005) may render the decoded point cloud data (20004). The point cloud content providing system (e.g., the receiving device 10004 or the renderer 10005) may use various rendering methods to render the geometry and attributes decoded through the decoding process. The points in the point cloud content may be rendered as vertices with a certain thickness, cubes with a specific minimum size centered at the corresponding vertex positions, or circles centered at the corresponding vertex positions. All or part of the rendered point cloud content is provided to the user through a display (e.g., a VR / AR display, a common display, etc.).
[0067] A point cloud content providing system according to an embodiment (e.g., the receiving device 10004) may obtain feedback information (20005). The point cloud content providing system may encode / decode point cloud data based on the feedback information. The feedback information and operations of the point cloud content providing system according to an embodiment are the same as those described in the reference Figure 1 and thus a detailed description thereof is omitted.
[0068] Figure 3 Illustrates an exemplary process of capturing a point cloud video according to an embodiment.
[0069] Figure 3 Illustrates the Figures 1 to 2 exemplary point cloud video capture process of the point cloud content providing system described in the reference.
[0070] The point cloud content includes a point cloud video (image and / or video) representing objects and / or environments located in various 3D spaces (e.g., a 3D space representing a real environment, a 3D space representing a virtual environment, etc.). Therefore, a point cloud content providing system according to an embodiment may use one or more cameras (e.g., an infrared camera capable of obtaining depth information, an RGB camera capable of extracting color information corresponding to the depth information, etc.), a projector (e.g., an infrared mode projector for obtaining depth information), LiDAR, etc. to capture a point cloud video. A point cloud content providing system according to an embodiment may extract the shape of the geometric structure formed by points in the 3D space from the depth information, and extract the attributes of each point from the color information to obtain point cloud data. Images and / or videos according to an embodiment may be captured based on at least one of an inward-facing technique and an outward-facing technique.
[0071] Figure 3 The left part of
[0072] Figure 3 illustrates the inward-facing technique. The inward-facing technique refers to a technique of capturing an image of a central object using one or more cameras (or camera sensors) disposed around the central object. The inward-facing technique may be used to generate point cloud content that provides a 360-degree image of a key object to the user (e.g., VR / AR content that provides a 360-degree image of an object (e.g., a key object such as a character, a player, an object, or an actor) to the user).
[0073] As shown in the figure, point cloud content can be generated based on the capture operations of one or more cameras. In this case, the coordinate system is different for each camera. Therefore, the point cloud content providing system can calibrate one or more cameras before the capture operation to set a global coordinate system. Additionally, the point cloud content providing system can generate point cloud content by synthesizing any image and / or video with the images and / or videos captured by the above capture techniques. The point cloud content providing system cannot perform the capture operations described in Figure 3 when it generates point cloud content representing a virtual space. The point cloud content providing system according to an embodiment can perform post-processing on the captured images and / or videos. In other words, the point cloud content providing system can remove unwanted regions (e.g., the background), identify the space to which the captured images and / or videos are connected, and perform an operation to fill spatial holes when there are spatial holes.
[0074] The point cloud content providing system can generate a piece of point cloud content by performing coordinate transformation on the points of the point cloud videos obtained from each camera. The point cloud content providing system can perform coordinate transformation on the points based on the coordinates of each camera position. Therefore, the point cloud content providing system can generate point cloud content representing a wide range of content, or can generate point cloud content with a high density of points.
[0075] Figure 4 An exemplary point cloud encoder according to an embodiment is illustrated.
[0076] Figure 4 Shows Figure 1 an example of the point cloud video encoder 10002. The point cloud encoder reconstructs and encodes the point cloud data (e.g., the positions and / or attributes of the points) to adjust the quality of the point cloud content (e.g., losslessly, lossily, or near-losslessly) according to the network conditions or the application. When the overall size of the point cloud content is large (e.g., for 30fps, a point cloud content of 60Gbps is given), the point cloud content providing system may not be able to stream the content in real time. Therefore, the point cloud content providing system can reconstruct the point cloud content based on the maximum target bit rate according to the network environment, etc., to provide the point cloud content.
[0077] As described with reference to Figures 1 to 2 the point cloud encoder can perform geometric structure encoding and attribute encoding. The geometric structure encoding is performed before the attribute encoding.
[0078] The point cloud encoder according to the embodiment includes a coordinate transformer (transform coordinates) 40000, a quantizer (quantize and remove points (voxelize)) 40001, an octree analyzer (analyze octree) 40002, a surface approximation analyzer (analyze surface approximation) 40003, an arithmetic encoder (arithmetic coding) 40004, a geometric structure reconstructor (reconstruct geometric structure) 40005, a color transformer (transform color) 40006, an attribute transformer (transform attribute) 40007, a RAHT transformer (RAHT) 40008, a LOD generator (generate LOD) 40009, a lifting transformer (lifting) 40010, a coefficient quantizer (quantize coefficients) 40011, and / or an arithmetic encoder (arithmetic coding) 40012.
[0079] The coordinate transformer 40000, the quantizer 40001, the octree analyzer 40002, the surface approximation analyzer 40003, the arithmetic encoder 40004, and the geometric structure reconstructor 40005 may perform geometric structure coding. The geometric structure coding according to the embodiment may include octree geometric structure coding, direct coding, trisoup geometric structure coding (trisoup geometry encoding), and entropy coding. The direct coding and the trisoup geometric structure coding are selectively or combinatorially applied. The geometric structure coding is not limited to the above examples.
[0080] As shown in the figure, the coordinate transformer 40000 according to the embodiment receives a position and transforms it into coordinates. For example, the position may be transformed into position information in a three-dimensional space (e.g., a three-dimensional space represented by an XYZ coordinate system). The position information in the three-dimensional space according to the embodiment may be referred to as geometric structure information.
[0081] The quantizer 40001 according to an embodiment quantizes the geometric structure. For example, the quantizer 40001 may quantize points based on the minimum position values of all points (e.g., the minimum values on each of the X, Y, and Z axes). The quantizer 40001 performs the following quantization operation: multiplying the difference between the minimum position value and the position value of each point by a preset quantization scaling value, and then finding the closest integer value by rounding the value obtained by the multiplication. Thus, one or more points may have the same quantized position (or position value). The quantizer 40001 according to an embodiment performs voxelization based on the quantized position to reconstruct the quantized points. As in the case of pixels which are the smallest units containing 2D image / video information, the points of the point cloud content (or 3D point cloud video) according to an embodiment may be included in one or more voxels. The term voxel, which is a composite of volume and pixel, refers to a 3D cubic space generated when a 3D space is divided into units (unit = 1.0) based on axes representing the 3D space (e.g., the X axis, the Y axis, and the Z axis). The quantizer 40001 may match multiple sets of points in the 3D space with voxels. According to an embodiment, one voxel may include only one point. According to an embodiment, one voxel may include one or more points. To represent a voxel as one point, the position of the center of the voxel may be set based on the positions of one or more points included in the voxel. In this case, the attributes of all the positions included in one voxel may be combined and assigned to the voxel.
[0082] The octree analyzer 40002 according to an embodiment performs octree geometric structure encoding (or octree encoding) to represent the voxels in an octree structure. The octree structure represents the points matched with the octree structure.
[0083] The surface approximation analyzer 40003 according to an embodiment may analyze and approximate the octree. The octree analysis and approximation according to an embodiment is a process of analyzing a region containing multiple points to efficiently provide the octree and voxelization.
[0084] The arithmetic encoder 40004 according to an embodiment performs entropy encoding on the octree and / or the approximated octree. For example, the encoding scheme includes arithmetic encoding. As a result of the encoding, a geometric structure bitstream is generated.
[0085] The color transformer 40006, the attribute transformer 40007, the RAHT transformer 40008, the LOD generator 40009, the lifting transformer 40010, the coefficient quantizer 40011, and / or the arithmetic coder 40012 perform attribute encoding. As described above, a point may have one or more attributes. The attribute encoding according to an embodiment is equally applied to the attributes that a point has. However, when an attribute (e.g., color) includes one or more elements, the attribute encoding is independently applied to each element. The attribute encoding according to an embodiment includes color transform encoding, attribute transform encoding, region adaptive hierarchical transform (RAHT) encoding, interpolation-based hierarchical nearest neighbor prediction (prediction transform) encoding, and interpolation-based hierarchical nearest neighbor prediction with an update / lifting step (lifting transform) encoding. Depending on the point cloud content, the above-described RAHT encoding, prediction transform encoding, and lifting transform encoding may be selectively used, or a combination of one or more encoding schemes may be used. The attribute encoding according to an embodiment is not limited to the above examples.
[0086] The color transformer 40006 according to an embodiment performs color transform encoding of the color value (or texture) included in the transform attributes. For example, the color transformer 40006 may transform the format of the color information (e.g., from RGB to YCbCr). The operation of the color transformer 40006 according to an embodiment may be selectively applied according to the color value included in the attributes.
[0087] The geometry reconstructor 40005 according to an embodiment reconstructs (decompresses) an octree and / or an approximate octree. The geometry reconstructor 40005 reconstructs an octree / voxel based on the result of analyzing the point distribution. The reconstructed octree / voxel may be referred to as a reconstructed geometry (restored geometry).
[0088] The attribute transformer 40007 according to an embodiment performs an attribute transform to transform an attribute based on the reconstructed geometry and / or a position where geometry encoding has not been performed. As described above, since an attribute depends on the geometry, the attribute transformer 40007 may transform an attribute based on the reconstructed geometry information. For example, based on the position value of the points included in a voxel, the attribute transformer 40007 may transform the attributes of the points at that position. As described above, when the position of the voxel center is set based on the position of one or more points included in the voxel, the attribute transformer 40007 transforms the attributes of the one or more points. When performing trisoup geometry encoding, the attribute transformer 40007 may transform an attribute based on the trisoup geometry encoding.
[0089] The attribute transformer 40007 can perform an attribute transformation by calculating the average value of the attributes or attribute values (e.g., the color or reflectivity of each point) of the neighboring points within a specific position / radius from the center position (or position value) of each voxel. The attribute transformer 40007 can apply weights according to the distance from the center to each point when calculating the average value. Thus, each voxel has a position and a calculated attribute (or attribute value).
[0090] The attribute transformer 40007 can search for neighboring points within a specific position / radius from the center position of each voxel based on a K-D tree or a Morton code. A K-D tree is a binary search tree and supports a data structure that can manage points based on their positions so that a nearest neighbor search (NNS) can be performed quickly. The Morton code is generated by representing the coordinates (e.g., (x, y, z)) of the 3D positions of all points as bit values and mixing the bits. For example, when the coordinates representing the position of a point are (5, 9, 1), the bit values of the coordinates are (0101, 1001, 0001). Mixing the bit values in the order of z, y, and x according to the bit index produces 010001000111. This value is represented as a decimal number 1095. That is, the Morton code value of the point with coordinates (5, 9, 1) is 1095. The attribute transformer 40007 can sort the points based on the Morton code values and perform NNS through depth-first traversal processing. When NNS is required in another transformation process for attribute encoding after the attribute transformation operation, a K-D tree or a Morton code is used.
[0091] As shown in the figure, the transformed attributes are input to the RAHT transformer 40008 and / or the LOD generator 40009.
[0092] The RAHT transformer 40008 according to an embodiment performs RAHT encoding for predicting attribute information based on the reconstructed geometric structure information. For example, the RAHT transformer 40008 can predict the attribute information of the nodes in the higher levels of the octree based on the attribute information associated with the nodes in the lower levels of the octree.
[0093] The LOD generator 40009 according to an embodiment generates a level of detail (LOD) to perform predictive transform encoding. The LOD according to an embodiment is the level of detail of the point cloud content. As the LOD value decreases, it indicates a decrease in the level of detail of the point cloud content. As the LOD value increases, it indicates an increase in the level of detail of the point cloud content. The points can be classified according to the LOD.
[0094] The lifting transformer 40010 according to an embodiment performs lifting transform encoding for transforming the point cloud attributes based on weights. As described above, the lifting transform encoding can be optionally applied.
[0095] The coefficient quantizer 40011 according to the embodiment quantizes the attribute after the attribute encoding based on the coefficient.
[0096] The arithmetic encoder 40012 according to the embodiment encodes the quantized attributes based on arithmetic coding.
[0097] Although not shown in this figure, Figure 4 The elements of the point cloud encoder may be implemented by hardware, software, firmware, or a combination thereof including one or more processors or integrated circuits configured to communicate with one or more memories included in the point cloud providing device. The one or more processors may perform the above Figure 4 At least one of the operations and / or functions of the elements of the point cloud encoder. In addition, one or more processors can operate or execute a set of software programs and / or instructions to perform Figure 4 The operation and / or functionality of the elements of the point cloud encoder. One or more memories according to an embodiment may include high-speed random access memory, or include non-volatile memory (e.g., one or more disk storage devices, flash memory devices, or other non-volatile solid-state storage devices).
[0098] Figure 5 An example of a voxel according to an embodiment is shown.
[0099] Figure 5 1 shows a voxel in a 3D space represented by a coordinate system consisting of three axes, namely, the X axis, the Y axis, and the Z axis. Figure 4 As described, the point cloud encoder (eg, quantizer 40001) may perform voxelization. A voxel refers to a 3D cubic space generated when a 3D space is divided into units (unit=1.0) based on axes representing the 3D space (eg, X-axis, Y-axis, and Z-axis). Figure 5 An example of a voxel generated by an octree structure is shown, in which the octree consists of two poles (0, 0, 0) and (2 d ,2 d ,2 d ) is recursively subdivided. A voxel consists of at least one point. The spatial coordinates of a voxel can be estimated based on the positional relationship with the voxel group. As mentioned above, a voxel has properties like pixels of a 2D image / video (such as color or reflectivity). The details of the voxel are similar to those of the reference Figure 4 The details described are the same, so their description is omitted.
[0100] Figure 6 Examples of octrees and occupancy codes are shown according to an embodiment.
[0101] As reference Figures 1 to 4As described, a point cloud content providing system (point cloud video encoder 10002) or a point cloud encoder (e.g., octree analyzer 40002) performs octree geometry encoding (or octree encoding) based on an octree structure to efficiently manage regions and / or positions of voxels.
[0102] Figure 6 The upper part shows an octree structure. The 3D space of the point cloud content according to an embodiment is represented by axes of a coordinate system (e.g., X-axis, Y-axis, and Z-axis). The octree structure is created by recursively subdividing a cube-axis-aligned bounding box defined by two extreme points (0, 0, 0) and (2 d , 2 d , 2 d ). Here, 2 d can be set to a value that constitutes a minimum bounding box surrounding all points of the point cloud content (or point cloud video). Here, d represents the depth of the octree. The value of d is determined by the following formula. In the following formula, (x int n , y int n , z int n ) represents the position (or position value) of a quantized point.
[0103]
[0104] As Figure 6 shown in the middle of the upper part, the entire 3D space can be divided into eight spaces according to partitioning. Each divided space is represented by a cube having six faces. As Figure 6 shown in the upper right of, each of the eight spaces is further divided based on the axes of the coordinate system (e.g., X-axis, Y-axis, and Z-axis). Thus, each space is divided into eight smaller spaces. The divided smaller spaces are also represented by cubes having six faces. This partitioning scheme is applied until the leaf nodes of the octree become voxels.
[0105] Figure 6 The lower part shows an octree occupancy code. An octree occupancy code is generated to indicate whether each of the eight divided spaces resulting from dividing a space contains at least one point. Thus, a single occupancy code is represented by eight child nodes. Each child node represents the occupancy of the divided space, and the child node has a value in units of 1 bit. Thus, the occupancy code is represented as an 8-bit code. That is, when at least one point is included in the space corresponding to the child node, the node is assigned the value 1. When no point is included in the space corresponding to the child node (the space is empty), the node is assigned the value 0. Since Figure 6The occupancy code shown is 00100001, so it indicates that the spaces corresponding to the third and eighth child nodes among the eight child nodes each contain at least one point. As shown in the figure, each of the third and eighth child nodes has 8 child nodes, and the child nodes are represented by 8-bit occupancy codes. The occupancy code of the third child node shown in the figure is 10000111, and the occupancy code of the eighth child node is 01001111. A point cloud encoder according to an embodiment (e.g., arithmetic encoder 40004) may perform entropy encoding on the occupancy code. To improve compression efficiency, the point cloud encoder may perform intra / inter-frame encoding on the occupancy code. A receiving device according to an embodiment (e.g., receiving device 10004 or point cloud video decoder 10006) reconstructs the octree based on the occupancy code.
[0106] A point cloud encoder according to an embodiment (e.g., Figure 4 the point cloud encoder or octree analyzer 40002) may perform voxelization and octree encoding to store the positions of the points. However, the points are not always evenly distributed in the 3D space, so there will be specific regions where there are fewer points. Therefore, it is inefficient to perform voxelization on the entire 3D space. For example, when a specific region contains fewer points, voxelization does not need to be performed in the specific region.
[0107] Therefore, for the above specific regions (or nodes other than the leaf nodes of the octree), a point cloud encoder according to an embodiment may skip voxelization and perform direct encoding to directly encode the positions of the points included in the specific regions. The coordinates of the directly encoded points according to an embodiment are referred to as the direct coding mode (DCM). A point cloud encoder according to an embodiment may also perform trisoup geometry encoding based on a surface model, thereby reconstructing the positions of the points in the specific region (or node) based on voxels. Trisoup geometry encoding is a geometry encoding that represents an object as a series of triangular meshes. Therefore, the point cloud decoder can generate a point cloud from the mesh surface. The direct encoding and trisoup geometry encoding according to an embodiment may be selectively performed. In addition, the direct encoding and trisoup geometry encoding according to an embodiment may be performed in combination with octree geometry encoding (or octree encoding).
[0108] To perform direct encoding, the option to use the direct mode to apply direct encoding should be enabled. The node to which direct encoding will be applied is not a leaf node, and there should be fewer than a threshold number of points within the specific node. In addition, the total number of points to which direct encoding will be applied should not exceed a preset threshold. When the above conditions are met, a point cloud encoder according to an embodiment (or arithmetic encoder 40004) may perform entropy encoding on the position (or position value) of the points.
[0109] A point cloud encoder according to an embodiment (e.g., the surface approximation analyzer 40003) may determine a specific level of the octree (a level less than the depth d of the octree), and the surface model may be used starting from this level to perform trisoup geometry encoding to reconstruct the positions of points in the region of the node based on voxels (trisoup mode). The point cloud encoder according to an embodiment may specify the level at which trisoup geometry encoding will be applied. For example, when the specific level is equal to the depth of the octree, the point cloud encoder does not operate in the trisoup mode. In other words, the point cloud encoder according to an embodiment may operate in the trisoup mode only when the specified level is less than the depth value of the octree. The 3D cubic region of the node at the specified level according to an embodiment is called a block. A block may include one or more voxels. The block or voxel may correspond to a brick. The geometry is represented as a surface within each block. The surface according to an embodiment may intersect each edge of the block at most once.
[0110] A block has 12 edges, so there are at least 12 intersection points in a block. Each intersection point is called a vertex (or apex point). A vertex existing along an edge is detected when there is at least one occupied voxel adjacent to that edge among all the blocks sharing the edge. The occupied voxel according to an embodiment refers to a voxel containing a point. The position of the vertex detected along the edge is the average position of the edges of all the voxels adjacent to that edge among all the blocks sharing the edge.
[0111] Once the vertex is detected, the point cloud encoder according to an embodiment may perform entropy encoding on the starting point (x, y, z) of the edge, the direction vector (Δx, Δy, Δz) of the edge, and the voxel position value (relative position value within the edge). When applying trisoup geometry encoding, the point cloud encoder according to an embodiment (e.g., the geometry reconstructor 40005) may generate a restored geometry (reconstructed geometry) by performing triangle reconstruction, upsampling, and voxelization processing.
[0112] The vertex at the edge of the block determines the surface passing through the block. The surface according to an embodiment is a non-planar polygon. In the triangle reconstruction process, the surface represented by a triangle is reconstructed based on the starting point of the edge, the direction vector of the edge, and the position value of the vertex. The triangle reconstruction process is performed through the following steps: i) calculating the centroid value of each vertex, ii) subtracting the centroid value from each vertex value, and iii) estimating the sum of the squares of the values obtained by the subtraction.
[0113] i) ii) iii)
[0114] Estimate the minimum value of the sum and perform projection processing based on the axis with the minimum value. For example, when the element x is the minimum value, each vertex is projected onto the x-axis with respect to the center of the block and projected onto the (y, z) plane. When the values obtained by projecting onto the (y, z) plane are (ai, bi), estimate the value of θ by atan2(bi, ai), and sort the vertices according to the value of θ. The following table shows the vertex combinations for creating triangles according to the number of vertices. The vertices are sorted from 1 to n. The following table shows that for four vertices, two triangles can be constructed according to the vertex combinations. The first triangle can be composed of vertices 1, 2, and 3 among the sorted vertices, and the second triangle can be composed of vertices 3, 4, and 1 among the sorted vertices.
[0115] Table 2-1 Triangles formed by vertices sorted as 1, …, n n triangles
[0116]
[0117] Perform upsampling processing to add points in the middle along the sides of the triangles and perform voxelization. The added points are generated based on the upsampling factor and the width of the block. The added points are called refined vertices. The point cloud encoder according to the embodiment can voxelize the refined vertices. In addition, the point cloud encoder can perform attribute encoding based on the voxelization position (or position value).
[0118] Figure 7 An example of the neighboring node pattern according to the embodiment is shown.
[0119] In order to improve the compression efficiency of the point cloud video, the point cloud encoder according to the embodiment can perform entropy encoding based on context-adaptive arithmetic coding.
[0120] As referred to Figures 1 to 6 described, the point cloud content providing system or the point cloud encoder (e.g., the point cloud video encoder 10002, the point cloud encoder or Figure 4 the arithmetic encoder 40004) can immediately perform entropy encoding on the occupancy code. In addition, the point cloud content providing system or the point cloud encoder can perform entropy encoding (intra-frame encoding) based on the occupancy code of the current node and the occupancy of the neighboring nodes, or perform entropy encoding (inter-frame encoding) based on the occupancy code of the previous frame. The frame according to the embodiment represents a set of point cloud videos generated simultaneously. The compression efficiency of the intra-frame encoding / inter-frame encoding according to the embodiment can depend on the number of neighboring nodes being referred to. When the number of bits increases, the operation becomes complex, but the encoding may be biased to one side, thereby increasing the compression efficiency. For example, when given 3-bit context, 2 3 = 8 methods are required to perform encoding. The parts divided for encoding affect the complexity of the implementation. Therefore, an appropriate level of compression efficiency and complexity must be satisfied.
[0121] Figure 7 Illustrates a process of obtaining an occupancy pattern based on the occupancy of neighboring nodes. The point cloud encoder according to an embodiment determines the occupancy of neighboring nodes of each node of the octree and obtains the value of the neighboring pattern. The occupancy pattern of the node is inferred using the neighboring node pattern. Figure 7 The left part of shows a cube corresponding to a node (the cube in the middle) and six cubes (neighboring nodes) that share at least one face with the cube. The nodes shown in the figure are nodes at the same depth. The numbers shown in the figure respectively represent the weights (1, 2, 4, 8, 16, and 32) associated with the six nodes. The weights are assigned sequentially according to the positions of the neighboring nodes.
[0122] Figure 7 The right part of shows the neighboring node pattern values. The neighboring node pattern value is the sum of the values obtained by multiplying the weights of the occupied neighboring nodes (neighboring nodes with points). Therefore, the neighboring node pattern value is from 0 to 63. When the neighboring node pattern value is 0, this indicates that there are no nodes without points (unoccupied nodes) among the neighboring nodes of the node. When the neighboring node pattern value is 63, this indicates that all neighboring nodes are occupied nodes. As shown in the figure, since the neighboring nodes assigned weights 1, 2, 4, and 8 are occupied nodes, the neighboring node pattern value is 15, which is the sum of 1, 2, 4, and 8. The point cloud encoder can perform encoding according to the neighboring node pattern value (for example, when the neighboring node pattern value is 63, 64 kinds of encoding can be performed). According to an embodiment, the point cloud encoder can reduce the encoding complexity by changing the neighboring node pattern value (for example, based on a table through which 64 is changed to 10 or 6).
[0123] Figure 8 Illustrates an example of the point configuration in each LOD according to an embodiment.
[0124] As referred to Figures 1 to 7 described, before performing attribute encoding, the encoded geometry is reconstructed (decompressed). When direct encoding is applied, the geometry reconstruction operation may include changing the placement of the points after direct encoding (for example, placing the points after direct encoding at the front of the point cloud data). When trisoup geometry encoding is applied, the geometry reconstruction process is performed through triangle reconstruction, upsampling, and voxelization. Since the attributes depend on the geometry, the attribute encoding is performed based on the reconstructed geometry.
[0125] A point cloud encoder (e.g., the LOD generator 40009) can classify (reorganize) points by LOD. This figure shows the point cloud content corresponding to the LOD. The leftmost picture in this figure represents the original point cloud content. The second picture from the left in this figure represents the distribution of points in the lowest LOD, and the rightmost picture in this figure represents the distribution of points in the highest LOD. That is, the points in the lowest LOD are sparsely distributed, and the points in the highest LOD are densely distributed. That is, as the LOD rises in the direction indicated by the arrow at the bottom of this figure, the space (or distance) between points narrows.
[0126] Figure 9 An example of the point configuration for each LOD according to an embodiment is illustrated.
[0127] As referred to Figures 1 to 8 described, a point cloud content providing system or a point cloud encoder (e.g., the point cloud video encoder 10002, Figure 4 the point cloud encoder or the LOD generator 40009) can generate LOD. LOD is generated by reorganizing points into a set of refinement levels according to a set of LOD distance values (or a set of Euclidean distances). The LOD generation process is performed not only by the point cloud encoder but also by the point cloud decoder.
[0128] Figure 9 The upper part of Figure 9 shows examples of points (P0 to P9) in the point cloud content distributed in 3D space. In Figure 9 the original order represents the order of points P0 to P9 before generating the LOD. In Figure 9 the LOD-based order represents the order of points generated according to the LOD. Points are reorganized by the LOD. Additionally, the high LOD contains the points belonging to the lower LOD. As shown in Figure 9 LOD0 contains P0, P5, P4, and P2. LOD1 contains the points of LOD0, P1, P6, and P3. LOD2 contains the points of LOD0, the points of LOD1, P9, P8, and P7.
[0129] As referred to Figure 4 described, the point cloud encoder according to an embodiment can selectively or combinatorially perform predictive transform coding, lifting transform coding, and RAHT transform coding.
[0130] The point cloud encoder according to an embodiment can generate predictors for points to perform predictive transform coding to set the prediction attributes (or prediction attribute values) for each point. That is, N predictors can be generated for N points. The predictor according to an embodiment can calculate weights (= 1 / distance) based on the LOD value of each point, the indexed information about neighboring points existing within the distance set for each LOD, and the distance to the neighboring points.
[0131] The predicted attribute (or attribute value) according to the embodiment is set to the average value of the values obtained by multiplying the attributes (or attribute values) of the neighboring points set in the predictor of each point (e.g., color, reflectance, etc.) by the weights (or weight values) calculated based on the distances from the respective neighboring points. A point cloud encoder according to the embodiment (e.g., coefficient quantizer 40011) may quantize and inverse-quantize the residuals (which may be referred to as residual attributes, residual attribute values, or attribute prediction residuals) obtained by subtracting the predicted attribute (attribute value) from the attribute (attribute value) of each point. The quantization process is configured as shown in the following table.
[0132] Table: Attribute Prediction Residual Quantization Pseudocode
[0133]
[0134] Table: Attribute Prediction Residual Inverse Quantization Pseudocode
[0135]
[0136] When the predictor of each point has neighboring points, a point cloud encoder according to the embodiment (e.g., arithmetic encoder 40012) may perform entropy coding on the residual attribute values after quantization and inverse quantization as described above. When the predictor of each point has no neighboring points, a point cloud encoder according to the embodiment (e.g., arithmetic encoder 40012) may perform entropy coding on the attributes of the corresponding points without performing the above operations.
[0137] A point cloud encoder according to the embodiment (e.g., lifting transformer 40010) may generate the predictor of each point, set the calculated LOD, register the neighboring points in the predictor, and set weights according to the distances from the neighboring points to perform lifting transform coding. The lifting transform coding according to the embodiment is similar to the above-described prediction transform coding, but the difference is that the weights are applied to the attribute values in an accumulative manner. The process of applying weights to the attribute values in an accumulative manner according to the embodiment is configured as follows.
[0138] 1) Create an array quantization weight (QW) for storing the weight values of each point. The initial value of all elements of QW is 1.0. Multiply the QW value of the predictor index of the neighboring nodes registered in the predictor by the weight of the predictor of the current point, and add the values obtained by the multiplication.
[0139] 2) Lifting prediction process: Subtract the value obtained by multiplying the attribute value of the point by the weight from the existing attribute value to calculate the predicted attribute value.
[0140] 3) Create a temporary array called updateweight, and update and initialize this temporary array to zero.
[0141] 4) The weights calculated by multiplying the weights calculated for all predictors by the weights corresponding to the predictor indices stored in QW are cumulatively added to the updated weight array as the indices of neighboring nodes. The values obtained by multiplying the attribute values of the indices of neighboring nodes by the calculated weights are cumulatively added to the updated array.
[0142] 5) Boosting update process: Divide the attribute values of the updated array for all predictors by the weight values of the updated weight array of the predictor indices, and add the existing attribute values to the values obtained by the division.
[0143] 6) Calculate the predicted attributes by multiplying the attribute values updated by the boosting update process for all predictors by the weights updated by the boosting prediction process (stored in QW). Quantize the predicted attribute values according to the point cloud encoder according to the embodiment (e.g., coefficient quantizer 40011). Additionally, the point cloud encoder (e.g., arithmetic encoder 40012) performs entropy encoding on the quantized attribute values.
[0144] The point cloud encoder according to the embodiment (e.g., RAHT transformer 40008) can perform RAHT transform encoding in which the attributes of higher-level nodes are predicted using the attributes associated with the nodes at lower levels in the octree. RAHT transform encoding is an example of intra-frame encoding of attributes performed by reverse scanning of the octree. The point cloud encoder according to the embodiment scans the entire region of the voxels and repeats the merging process of merging the voxels into larger blocks at each step until reaching the root node. The merging process according to the embodiment is only performed on occupied nodes. The merging process is not performed on empty nodes. The merging process is performed on the upper nodes directly above the empty nodes.
[0145] The following formula represents the RAHT transform matrix. In this formula, represents the average attribute value of the voxels at level l. can be based on and to calculate. and The weights of and
[0146]
[0147] Here, is the low-pass value and is used in the merging process at the next higher level. represents the high-pass coefficient. The high-pass coefficients at each step are quantized and subjected to entropy encoding (e.g., encoded by arithmetic encoder 400012). The weights are calculated as as follows by and to create a root node.
[0148]
[0149] Figure 10 An example of a point cloud decoder according to an embodiment is illustrated.
[0150] Figure 10 The point cloud decoder illustrated in Figure 1 is an example of the point cloud video decoder 10006 described in Figure 1 and can perform operations the same as or similar to those of the point cloud video decoder 10006 illustrated in
[0151] Figure 11 An example of a point cloud decoder according to an embodiment is illustrated.
[0152] Figure 11 The point cloud decoder illustrated in Figure 10 is an example of the point cloud decoder illustrated in Figures 1 to 9 and can perform a decoding operation that is an inverse process of the encoding operation of the point cloud encoder illustrated in
[0153] As described with reference to Figure 1 and Figure 10 the point cloud decoder can perform geometric structure decoding and attribute decoding. Geometric structure decoding is performed before attribute decoding.
[0154] A point cloud decoder according to an embodiment includes an arithmetic decoder (arithmetic decoding) 11000, an octree synthesizer (synthesize octree) 11001, a surface approximation synthesizer (synthesize surface approximation) 11002, a geometric structure reconstructor (reconstruct geometric structure) 11003, a coordinate inverse transformer (inverse transform coordinates) 11004, an arithmetic decoder (arithmetic decoding) 11005, an inverse quantization unit (inverse quantization) 11006, a RAHT transformer 11007, a LOD generator (generate LOD) 11008, an inverse lifter (inverse lift) 11009, and / or a color inverse transformer (inverse transform color) 11010.
[0155] The arithmetic decoder 110000, octree synthesizer 11001, surface approximation synthesizer 11002, geometric structure reconstructor 11003, and coordinate inverse transformer 11004 can perform geometric structure decoding. Geometric structure decoding according to an embodiment may include direct decoding and trisoup geometric structure decoding. Direct decoding and trisoup geometric structure decoding are selectively applied. Geometric structure decoding is not limited to the above examples and is provided for reference. Figures 1 to 9 It is performed as the inverse process of the geometric structure encoding described above.
[0156] According to an embodiment, the arithmetic decoder 11000 decodes the received geometric structure bitstream based on arithmetic coding. The operation of the arithmetic decoder 11000 corresponds to the inverse process of the arithmetic encoder 40004.
[0157] According to an embodiment, the octree synthesizer 11001 can generate an octree by obtaining occupancy codes from the decoded geometric structure bitstream (or information about the geometric structure obtained as a decoding result). As described in the reference Figures 1 to 9 The occupancy codes are configured in detail.
[0158] When trisoup geometric structure encoding is applied, according to an embodiment, the surface approximation synthesizer 11002 can synthesize a surface based on the decoded geometric structure and / or the generated octree.
[0159] According to an embodiment, the geometric structure reconstructor 11003 can regenerate a geometric structure based on the surface and / or the decoded geometric structure. As described in the reference Figures 1 to 9 Direct encoding and trisoup geometric structure encoding are selectively applied. Therefore, the geometric structure reconstructor 11003 directly imports and adds the position information of the points to which direct encoding is applied. When trisoup geometric structure encoding is applied, the geometric structure reconstructor 11003 can reconstruct the geometric structure by performing the reconstruction operations of the geometric structure reconstructor 40005 (e.g., triangle reconstruction, upsampling, and voxelization). The details are the same as those described in the reference Figure 6 and are thus omitted here. The reconstructed geometric structure may include a point cloud picture or frame that does not contain attributes.
[0160] According to an embodiment, the coordinate inverse transformer 11004 can obtain the positions of points by transforming coordinates based on the reconstructed geometric structure.
[0161] The arithmetic decoder 11005, inverse quantizer 11006, RAHT transformer 11007, LOD generator 11008, inverse lifter 11009, and / or color inverse transformer 11010 can perform the reference Figure 10Described attribute decoding. Attribute decoding according to an embodiment includes region adaptive hierarchical transform (RAHT) decoding, interpolation-based hierarchical nearest neighbor prediction (prediction transform) decoding, and interpolation-based hierarchical nearest neighbor prediction with an update / lifting step (lifting transform) decoding. The above three decoding schemes can be selectively used, or a combination of one or more decoding schemes can be used. Attribute decoding according to an embodiment is not limited to the above examples.
[0162] An arithmetic decoder 11005 according to an embodiment decodes an attribute bitstream by arithmetic decoding.
[0163] An inverse quantizer 11006 according to an embodiment inverse quantizes information about the decoded attribute bitstream or attribute obtained as a decoding result, and outputs the inverse quantized attribute (or attribute value). Inverse quantization can be selectively applied based on the attribute encoding of the point cloud encoder.
[0164] According to an embodiment, a RAHT transformer 11007, a LOD generator 11008, and / or an inverse lifter 11009 can process the reconstructed geometry and the inverse quantized attributes. As described above, the RAHT transformer 11007, the LOD generator 11008, and / or the inverse lifter 11009 can selectively perform decoding operations corresponding to the encoding of the point cloud encoder.
[0165] A color inverse transformer 11010 according to an embodiment performs inverse transform decoding to inverse transform the color value (or texture) included in the decoded attribute. The operation of the color inverse transformer 11010 can be selectively performed based on the operation of the color transformer 40006 of the point cloud encoder.
[0166] Although not shown in this figure, Figure 11 the elements of the point cloud decoder can be implemented by hardware, software, firmware, or a combination thereof including one or more processors or integrated circuits configured to communicate with one or more memories included in the point cloud providing device. The one or more processors can perform at least one or more of the operations and / or functions of the elements of the point cloud video decoder described above. Additionally, the one or more processors can operate or execute a set of software programs and / or instructions to perform the operations and / or functions of the elements of the point cloud decoder. Figure 11 the elements of the point cloud video decoder can be implemented by hardware, software, firmware, or a combination thereof including one or more processors or integrated circuits configured to communicate with one or more memories included in the point cloud providing device. The one or more processors can perform at least one or more of the operations and / or functions of the elements of the point cloud video decoder described above. Additionally, the one or more processors can operate or execute a set of software programs and / or instructions to perform the operations and / or functions of the elements of the point cloud decoder. Figure 11 the elements of the point cloud decoder can be implemented by hardware, software, firmware, or a combination thereof including one or more processors or integrated circuits configured to communicate with one or more memories included in the point cloud providing device. The one or more processors can perform at least one or more of the operations and / or functions of the elements of the point cloud video decoder described above. Additionally, the one or more processors can operate or execute a set of software programs and / or instructions to perform the operations and / or functions of the elements of the point cloud decoder.
[0167] Figure 12 An exemplary transmitting device according to an embodiment is illustrated.
[0168] Figure 12 The transmitting device shown in Figure 1 is the transmitting device 10000 of Figure 4Example of a point cloud encoder). Figure 12 The transmission device illustrated in can perform one or more of the operations and methods that are the same as or similar to the operations and methods of the point cloud encoder described with reference to Figures 1 to 9 According to an embodiment, the transmission device may include a data input unit 12000, a quantization processor 12001, a voxelization processor 12002, an octree occupancy code generator 12003, a surface model processor 12004, an intra / inter-frame encoding processor 12005, an arithmetic encoder 12006, a metadata processor 12007, a color transformation processor 12008, an attribute transformation processor 12009, a prediction / lifting / RAHT transformation processor 12010, an arithmetic encoder 12011, and / or a transmission processor 12012.
[0169] The data input unit 12000 according to an embodiment receives or acquires point cloud data. The data input unit 12000 may perform operations and / or acquisition methods that are the same as or similar to the operations and / or acquisition method of the point cloud video acquirer 10001 (or the acquisition process 20000 described with reference to Figure 2 ).
[0170] The data input unit 12000, the quantization processor 12001, the voxelization processor 12002, the octree occupancy code generator 12003, the surface model processor 12004, the intra / inter-frame encoding processor 12005, and the arithmetic encoder 12006 perform geometric structure encoding. The geometric structure encoding according to an embodiment is the same as or similar to the geometric structure encoding described with reference to Figures 1 to 9 and thus a detailed description thereof is omitted.
[0171] The quantization processor 12001 according to an embodiment quantizes geometric structures (e.g., position values of points). The operations and / or quantization of the quantization processor 12001 are the same as or similar to the operations and / or quantization of the quantizer 40001 described with reference to Figure 4 . The details are the same as the details described with reference to Figures 1 to 9 .
[0172] The voxelization processor 12002 according to an embodiment voxelizes the quantized position values of points. The voxelization processor 120002 may perform operations and / or processes that are the same as or similar to the operations and / or voxelization process of the quantizer 40001 described with reference to Figure 4 . The details are the same as the details described with reference to Figures 1 to 9 .
[0173] According to an embodiment, the octree occupancy code generator 12003 performs octree encoding on the voxelized positions of points based on an octree structure. The octree occupancy code generator 12003 may generate an occupancy code. The octree occupancy code generator 12003 may perform operations and / or methods that are the same as or similar to those of the point cloud encoder (or octree analyzer 40002) described with reference to Figure 4 and Figure 6 The details are the same as those described with reference to Figures 1 to 9 described.
[0174] According to an embodiment, the surface model processor 12004 may perform trisoup geometry encoding based on a surface model to reconstruct the positions of points in a specific region (or node) based on voxels. The surface model processor 12004 may perform operations and / or methods that are the same as or similar to those of the point cloud encoder (e.g., surface approximation analyzer 40003) described with reference to Figure 4 The details are the same as those described with reference to Figures 1 to 9 described.
[0175] According to an embodiment, the intra / inter-frame encoding processor 12005 may perform intra / inter-frame encoding on point cloud data. The intra / inter-frame encoding processor 12005 may perform encoding that is the same as or similar to the intra / inter-frame encoding described with reference to Figure 7 The details are the same as those described with reference to Figure 7 described. According to an embodiment, the intra / inter-frame encoding processor 12005 may be included in the arithmetic encoder 12006.
[0176] According to an embodiment, the arithmetic encoder 12006 performs entropy encoding on the octree and / or approximate octree of point cloud data. For example, the encoding scheme includes arithmetic coding. The arithmetic encoder 12006 performs operations and / or methods that are the same as or similar to those of the arithmetic encoder 40004.
[0177] According to an embodiment, the metadata processor 12007 processes metadata about point cloud data (e.g., set values) and provides it to necessary processing procedures such as geometry encoding and / or attribute encoding. Additionally, according to an embodiment, the metadata processor 12007 may generate and / or process signaling information related to geometry encoding and / or attribute encoding. The signaling information according to an embodiment may be encoded separately from geometry encoding and / or attribute encoding. The signaling information according to an embodiment may be interleaved.
[0178] The color transformation processor 12008, the attribute transformation processor 12009, the prediction / lifting / RAHT transformation processor 12010, and the arithmetic encoder 12011 perform attribute encoding. The attribute encoding according to an embodiment is the same as that with reference toFigures 1 to 9 The described attribute encodings are the same or similar, so their detailed descriptions are omitted.
[0179] The color transformation processor 12008 according to the embodiment performs color transformation encoding to transform the color values included in the attributes. The color transformation processor 12008 may perform color transformation encoding based on the reconstructed geometric structure. The reconstructed geometric structure is the same as the reference Figures 1 to 9 described. Additionally, it performs operations and / or methods that are the same as or similar to the operations and / or methods of the color transformer 40006 described in the reference Figure 4 described. The detailed description thereof is omitted.
[0180] The attribute transformation processor 12009 according to the embodiment performs attribute transformation to transform the attributes based on the reconstructed geometric structure and / or the positions not encoded by the geometric structure. The attribute transformation processor 12009 performs operations and / or methods that are the same as or similar to the operations and / or methods of the attribute transformer 40007 described in the reference Figure 4 described. The detailed description thereof is omitted. The prediction / lifting / RAHT transformation processor 12010 according to the embodiment may encode the transformed attributes by any one or a combination of RAHT encoding, prediction transformation encoding, and lifting transformation encoding. The prediction / lifting / RAHT transformation processor 12010 performs at least one of the operations that are the same as or similar to the operations of the RAHT transformer 40008, the LOD generator 40009, and the lifting transformer 40010 described in the reference Figure 4 described. Additionally, the prediction transformation encoding, the lifting transformation encoding, and the RAHT transformation encoding are the same as those described in the reference Figures 1 to 9 described, so the detailed description thereof is omitted.
[0181] The arithmetic encoder 12011 according to the embodiment may encode the encoded attributes based on arithmetic coding. The arithmetic encoder 12011 performs operations and / or methods that are the same as or similar to the operations and / or methods of the arithmetic encoder 400012.
[0182] According to an embodiment, the transmitting processor 12012 may transmit each bitstream including the encoded geometry and / or the encoded attributes and metadata information, or transmit a single bitstream configured with the encoded geometry and / or the encoded attributes and metadata information. When the encoded geometry and / or the encoded attributes and metadata information according to an embodiment are configured as a single bitstream, the bitstream may include one or more sub-bitstreams. The bitstream according to an embodiment may include signaling information, which includes a sequence parameter set (SPS) for signaling at the sequence level, a geometry parameter set (GPS) for signaling for encoding geometry information, an attribute parameter set (APS) for signaling for encoding attribute information, and a tile parameter set (TPS) for signaling for tile level and slice data. The slice data may include information about one or more slices. A slice according to an embodiment may include a geometry bitstream Geom0 0 and one or more attribute bitstreams Attr0 0 and Attr1 0 .
[0183] A slice refers to a series of syntax elements that represent all or part of an encoded point cloud frame.
[0184] The TPS according to an embodiment may include information about each tile in one or more tiles (e.g., coordinate information and height / size information about the border). The geometry bitstream may include a header and a payload. The header of the geometry bitstream according to an embodiment may include a parameter set identifier (geom_parameter_set_id) included in the GPS, a tile identifier (geom_tile_id), and a slice identifier (geom_slice_id) included in the GPS, and information about the data included in the payload. As described above, the metadata processor 12007 according to an embodiment may generate and / or process the signaling information and send it to the transmitting processor 12012. According to an embodiment, the elements for performing geometry encoding and the elements for performing attribute encoding may share data / information with each other, as indicated by the dashed line. The transmitting processor 12012 according to an embodiment may perform operations and / or a transmission method that are the same as or similar to the operations and / or the transmission method of the transmitter 10003. The details are the same as the details described with reference to Figure 1 and Figure 2 and thus the description thereof is omitted.
[0185] Figure 13 Illustrates an exemplary receiving device according to an embodiment.
[0186] Figure 13 The receiving device illustrated in Figure 1 is the receiving device 10004 (orFigure 10 and Figure 11 an example of a point cloud decoder). Figure 13 The receiving device illustrated in Figures 1 to 11 can perform one or more of the operations and methods that are the same as or similar to the operations and methods of the point cloud decoder described in the reference
[0187] The receiving device according to an embodiment includes a receiver 13000, a receiving processor 13001, an arithmetic decoder 13002, an occupancy code-based octree reconstruction processor 13003, a surface model processor (triangle reconstruction, upsampling, voxelization) 13004, an inverse quantization processor 13005, a metadata parser 13006, an arithmetic decoder 13007, an inverse quantization processor 13008, a prediction / lifting / RAHT inverse transform processor 13009, a color inverse transform processor 13010, and / or a renderer 13011. Each element for decoding according to an embodiment can perform an inverse process of the operation of the corresponding element for encoding according to the embodiment.
[0188] The receiver 13000 according to an embodiment receives point cloud data. The receiver 13000 can perform operations and / or a receiving method that are the same as or similar to the operations and / or the receiving method of the receiver 10007 in Figure 1 . A detailed description thereof is omitted.
[0189] The receiving processor 13001 according to an embodiment can obtain a geometry bitstream and / or an attribute bitstream from the received data. The receiving processor 13001 can be included in the receiver 13000.
[0190] The arithmetic decoder 13002, the occupancy code-based octree reconstruction processor 13003, the surface model processor 13004, and the inverse quantization processor 1305 can perform geometry decoding. The geometry decoding according to an embodiment is the same as or similar to the geometry decoding described in the reference Figures 1 to 10 . Thus, a detailed description thereof is omitted.
[0191] The arithmetic decoder 13002 according to an embodiment can decode the geometry bitstream based on arithmetic coding. The arithmetic decoder 13002 performs operations and / or coding that are the same as or similar to the operations and / or coding of the arithmetic decoder 11000.
[0192] The occupancy code-based octree reconstruction processor 13003 according to an embodiment may reconstruct an octree by obtaining an occupancy code from a decoded geometric structure bitstream (or information about the geometric structure obtained as a decoding result). The occupancy code-based octree reconstruction processor 13003 performs operations and / or methods that are the same as or similar to those of the octree synthesizer 11001 and / or the octree generation method. When applying trisoup geometric structure encoding, the surface model processor 13004 according to an embodiment may perform trisoup geometric structure decoding and related geometric structure reconstruction (e.g., triangle reconstruction, upsampling, voxelization) based on the surface model method. The surface model processor 13004 performs operations that are the same as or similar to those of the surface approximation synthesizer 11002 and / or the geometric structure reconstructor 11003.
[0193] The inverse quantization processor 13005 according to an embodiment may perform inverse quantization on the decoded geometric structure.
[0194] The metadata parser 13006 according to an embodiment may parse metadata included in the received point cloud data, e.g., set values. The metadata parser 13006 may transmit the metadata for geometric structure decoding and / or attribute decoding. The metadata is the same as the metadata described in the reference Figure 12 and thus a detailed description thereof is omitted.
[0195] The arithmetic decoder 13007, the inverse quantization processor 13008, the prediction / lifting / RAHT inverse transform processor 13009, and the color inverse transform processor 13010 perform attribute decoding. The attribute decoding is the same as or similar to the attribute decoding described in the reference Figures 1 to 10 and thus a detailed description thereof is omitted.
[0196] The arithmetic decoder 13007 according to an embodiment may decode an attribute bitstream by arithmetic coding. The arithmetic decoder 13007 may decode the attribute bitstream based on the reconstructed geometric structure. The arithmetic decoder 13007 performs operations and / or coding that are the same as or similar to those of the arithmetic decoder 11005.
[0197] The inverse quantization processor 13008 according to an embodiment may perform inverse quantization on the decoded attribute bitstream. The inverse quantization processor 13008 performs operations and / or methods that are the same as or similar to those of the inverse quantizer 11006.
[0198] According to an embodiment, the prediction / lifting / RAHT inverse transform processor 13009 may process the reconstructed geometry and the inverse quantized attributes. The prediction / lifting / RAHT inverse transform processor 13009 performs one or more of the same or similar operations and / or decoding as those of the RAHT transform processor 11007, the LOD generator 11008, and / or the inverse lifter 11009. According to an embodiment, the color inverse transform processor 13010 performs inverse transform coding to inverse transform the color values (or textures) included in the decoded attributes. The color inverse transform processor 13010 performs the same or similar operations and / or inverse transform coding as those of the color inverse transform processor 11010. According to an embodiment, the renderer 13011 may render the point cloud data.
[0199] Figure 14 An exemplary structure operable in conjunction with a point cloud data transmission / reception method / apparatus according to an embodiment is illustrated.
[0200] Figure 14 The structure represents a configuration in which at least one of the server 17600, the robot 17100, the autonomous vehicle 17200, the XR device 17300, the smartphone 17400, the home appliance 17500, and / or the head-mounted display (HMD) 17700 is connected to the cloud network 17000. The robot 17100, the autonomous vehicle 17200, the XR device 17300, the smartphone 17400, or the home appliance 17500 is referred to as a device. Additionally, the XR device 17300 may correspond to a point cloud data (PCC) device according to an embodiment or may be operatively connected to a PCC device.
[0201] The cloud network 17000 may represent a part of or exist in a cloud computing infrastructure. Here, the cloud network 17000 may be configured using a 3G network, a 4G or Long-Term Evolution (LTE) network, or a 5G network.
[0202] The server 17600 may be connected to at least one of the robot 17100, the autonomous vehicle 17200, the XR device 17300, the smartphone 17400, the home appliance 17500, and / or the HMD 17700 via the cloud network 17000 and may assist in at least part of the processing of the connected devices 17100 to 17700.
[0203] The HMD 17700 represents one of the implementation types of an XR device and / or a PCC device according to an embodiment. The HMD-type device according to an embodiment includes a communication unit, a control unit, a memory, an I / O unit, a sensor unit, and a power unit.
[0204] In the following, various embodiments of apparatuses 17100 to 17500 to which the above technologies are applied will be described. According to the above embodiments, Figure 14 the apparatuses 17100 to 17500 illustrated in [reference] can be operably connected / linked to a point cloud data transmitting apparatus and a receiver.
[0205] <PCC+XR>
[0206] The XR / PCC apparatus 17300 can adopt PCC technology and / or XR (AR+VR) technology, and can be implemented as an HMD, a head-up display (HUD) provided in a vehicle, a television, a mobile phone, a smartphone, a computer, a wearable device, a household appliance, a digital signage, a vehicle, a stationary robot, or a mobile robot.
[0207] The XR / PCC apparatus 17300 can analyze 3D point cloud data or image data obtained through various sensors or from an external apparatus, and generate position data and attribute data regarding 3D points. Thereby, the XR / PCC apparatus 17300 can acquire information regarding the surrounding space or real objects, and render and output XR objects. For example, the XR / PCC apparatus 17300 can match an XR object including auxiliary information regarding a recognized object with the recognized object, and output the matched XR object.
[0208] <PCC + XR + Mobile Phone>
[0209] The XR / PCC apparatus 17300 can be implemented as a mobile phone (smartphone) 17400 by applying PCC technology.
[0210] The mobile phone 17400 can decode and display point cloud content based on PCC technology.
[0211] <PCC+Autopilot+XR>
[0212] The autonomous driving vehicle 17200 can be implemented as a mobile robot, a vehicle, a drone, etc. by applying PCC technology and XR technology.
[0213] The autonomous driving vehicle 17200 applying XR / PCC technology can represent an autonomous driving vehicle provided with an apparatus for providing an XR image or an autonomous driving vehicle that is a control / interaction target in an XR image. Specifically, the autonomous driving vehicle 17200 that is a control / interaction target in an XR image can be distinguished from the XR apparatus 17300, and can be operably connected to the XR apparatus 17300.
[0214] An autonomous vehicle 17200 having a device for providing XR / PCC images can acquire sensor information from sensors including cameras and output XR / PCC images generated based on the acquired sensor information. For example, the autonomous vehicle 17200 can have a HUD and output XR / PCC images thereto, thereby providing an XR / PCC object corresponding to a real object or an object existing on a screen to an occupant.
[0215] When an XR / PCC object is output to the HUD, at least a part of the XR object can be output to overlap with a real object at which an occupant's eyes are gazing. On the other hand, when an XR / PCC object is output to a display provided inside the autonomous vehicle, at least a part of the XR / PCC object can be output to overlap with an object on the screen. For example, the autonomous vehicle 17200 can output an XR / PCC object corresponding to objects such as a road, another vehicle, a traffic signal, a traffic sign, a two-wheeler, a pedestrian, and a building.
[0216] Virtual reality (VR) technology, augmented reality (AR) technology, mixed reality (MR) technology, and / or point cloud compression (PCC) technology according to an embodiment are applicable to various devices.
[0217] In other words, VR technology is a display technology that only provides CG images of real-world objects, backgrounds, etc. On the other hand, AR technology refers to a technology of showing a virtual created CG image on an image of a real object. The MR technology is similar to the above AR technology in that the virtual object to be shown is mixed and combined with the real world. However, the MR technology is different from the AR technology in that the AR technology clearly distinguishes between a real object and a virtual object created as a CG image and uses the virtual object as a supplementary object for the real object, while the MR technology regards the virtual object as an object having the same characteristics as the real object. More specifically, an example of the application of the MR technology is a hologram service.
[0218] Recently, VR, AR, and MR technologies are sometimes referred to as extended reality (XR) technologies without being clearly distinguished from each other. Therefore, embodiments of the present disclosure are applicable to any one of VR, AR, MR, and XR technologies. Encoding / decoding based on PCC, V-PCC, G-PCC technologies is applicable to such technologies.
[0219] The PCC method / device according to an embodiment can be applied to a vehicle providing an autonomous driving service.
[0220] A vehicle providing an autonomous driving service is connected to a PCC device to perform wired / wireless communication.
[0221] When the point cloud data (PCC) transmitting / receiving device according to an embodiment is connected to a vehicle for wired / wireless communication, the device may receive / process content data related to AR / VR / PCC services that may be provided together with an autonomous driving service and transmit it to the vehicle. In the case where the PCC transmitting / receiving device is installed on a vehicle, the PCC transmitting / receiving device may receive / process content data related to AR / VR / PCC services according to a user input signal input through a user interface device and provide it to the user. A vehicle or a user interface device according to an embodiment may receive a user input signal. A user input signal according to an embodiment may include a signal indicating an autonomous driving service.
[0222] As referred to Figures 1 to 14 described, the point cloud processing device according to an embodiment (e.g., Figure 1 , Figure 12 and Figure 14 the transmitting device or the point cloud encoder described therein) selectively uses RAHT coding, predictive transform coding, and lifting transform coding or a combination of one or more of the coding techniques according to the point cloud content to perform attribute coding. For example, RAHT coding and lifting transform coding may be used for lossy coding, which greatly compresses the point cloud content data. Predictive transform coding may be used for lossless coding.
[0223] As described above, the point cloud encoder according to an embodiment may generate a predictor for a point and perform predictive transform coding to set the prediction attribute (or prediction attribute value) for each point. According to an embodiment, predictive transform coding and lifting transform coding calculate the distance (or position) of each neighboring point based on the positions of points within the neighboring range of the point cloud (hereinafter referred to as target points). The calculated distance is used as a reference or reference weight to predict the attribute (e.g., color, reflectivity, etc.) of the target point or to update the attribute (or prediction attribute) of the target point when the distance to the neighboring point changes. The following formula represents the attribute of the target point predicted based on neighboring points.
[0224] [Equation 1]
[0225]
[0226] In the above formula, P represents the prediction attribute of the target point, and d1, d2, and d3 represent the distances to each of the three neighboring points of the target point. The corresponding distances are combined and used as a reference weight. C1, C2, and C3 (not shown in the formula) represent the attributes of the corresponding neighboring points. Shift is a parameter for adjusting the average energy or power magnitude of the point. The value of Shift is controlled by the hardware voltage operating range of the encoder or decoder.
[0227] According to an embodiment, considering the correlation between points (e.g., neighboring points), the point cloud encoder may change the above reference weights. Specifically, the point cloud encoder uses the correlation weights calculated by changing the above reference weights according to the correlation combination method to highlight the eigenvalue at the distance between the target point and the neighboring points. Therefore, the point cloud encoder can ensure a performance gain proportional to the correlation between points. Additionally, the point cloud encoder may generate a predictor considering the degree of correlation without changing the structures of the RAHT coding, predictive transform coding, and lifting transform coding described Figures 1 to 14 while considering the degree of correlation without changing the structures of the RAHT coding, predictive transform coding, and lifting transform coding described
[0228] Figure 15 is a flowchart illustrating an example of point cloud coding.
[0229] As described in the reference Figures 1 to 14 described, the point cloud transmitting device or the point cloud encoder (e.g., Figure 4 the point cloud encoder described in
[0230] receives an attribute (15100). The point cloud encoder (e.g., the LOD generator 4009) generates an LOD to perform predictive transform (15200). The point cloud encoder reorganizes the points into levels of detail to generate an LOD. Therefore, the larger the level value of the LOD, the more detailed the point cloud content. According to an embodiment, the LOD may include points grouped based on the distance between points. The point cloud encoder reorganizes the points based on an octree structure. An iterative generation algorithm applicable to octree decoding may be applied to the grouped points according to the position or order of the points (e.g., Morton code order, etc.). In each iteration sequence, one or more levels of detail R0, R1,..., Ri belonging to one LOD (e.g., LODi) are generated. That is, the level of the LOD is a combination of levels of detail.
[0231] Additionally, the point cloud encoder described in the reference Figures 1 to 14 supports spatially adaptive decoding. According to an embodiment, spatially adaptive decoding performs some or all of the geometric structure and / or attributes according to the decoding performance of the point cloud receiving device (e.g., Figure 1 the receiving device 10004 of Figure 10 and Figure 11 the point cloud decoder of Figure 13 and the receiving device of Figures 10 to 11 described, the point cloud decoder described in the reference Figure 13The described receiving device, etc.) performs adaptable attribute decoding. The LOD for supporting spatial adaptable decoding can be generated by searching for neighboring points through an approximate nearest neighbor search method for points from the lowest point to the highest point in the octree structure. The nearest neighbor points of the corresponding points in the current LOD (e.g., LODl) are searched from the LOD (e.g., LODl-1) at a level lower than the current LOD. The LODs at levels lower than the current LOD are a combination of the refinement levels of R0, R1, ..., and Rl-1.
[0232] Configure a specific LOD generation algorithm as follows. According to an embodiment, (P i ) i=1...N is referred to as the set of positions associated with the points of the point cloud. According to an embodiment, (M i ) i=1... N is the Morton code associated with the set of positions. The parameters D0 and ρ are defined as the initial sampling distance and the distance ratio between LODs, respectively. The distance ratio is always greater than 1 (ρ > 1).
[0233] According to an embodiment, the points are sorted in ascending order according to the Morton code values of the points. According to an embodiment, the parameter I represents an array of point indices sorted according to the above processing. The LOD generation algorithm is executed iteratively. In each iteration k, the points belonging to LODk are extracted, and predictors for the extracted points are generated starting from k equal to 0 until all points are assigned to an LOD. Below, a more detailed process is described.
[0234] The sampling distance D is initialized to the initial sampling distance D0. In the iteration k ranging from 0 to the number of LODs, L(k) is the set of indices of the points belonging to the k-th LOD, and O(k) is the set of points belonging to the LODs corresponding to levels higher than k. After L(k) and O(k) are initialized, the LOD assignment and residuals of the points are calculated iteratively and input sequentially. This process is repeated for all indices in the array I. Here, L(k) and O(k) can be calculated and used in the process of generating predictors associated with the points of L(k). According to an embodiment, R(k) is the set of points that need to be added to LOD(k-1) to obtain LOD(k) and is represented as follows.
[0235] R(k) = L(k) \ L(k-1), where "\" is the difference operator.
[0236] For each point i in R(k), an algorithm for finding h neighboring points of point i in O(k) and calculating the normal distance and the associated linear distance associated with point i is configured as follows. According to an embodiment, h, as a user-defined parameter, represents a constant for adjusting the maximum number of neighboring points used to predict point i.
[0237] The counter j is initialized to zero (j = 0).
[0238] For a point i in R(k), Mi represents the Morton code associated with the point i. Mj represents the Morton code associated with the j-th element in O(k).
[0239] When Mi is greater than or equal to Mj and j is less than the size of O(k) (M i ≥M j and j < SizeOf(O(k))), the counter j is incremented by 1 (j←j+1) and the distance between Mi and the point associated with the index in O(k) is measured. The point is within a specific search range [j - SR2, j + SR2], and h nearest neighbor points ((n1, n2,..., h )) and the normal distance between each neighboring point and i are tracked
[0240] In addition, based on the normal distance, the correlation squared distance between two nearest squared distances is calculated. The correlation squared distance can be expressed as follows.
[0241]
[0242] The calculation method is not limited to the above examples.
[0243] When the correlation squared distance between the target point and the last processed point is less than the threshold, the neighboring points of the last processed point are used for initial estimation and search. According to the embodiment, the threshold can be defined by the user. Points with a correlation squared distance greater than the threshold are excluded.
[0244] The LOD generation algorithm is also applied to a point cloud receiving device (for example, refer to Figure 10 、 Figure 11 and Figure 13 the described point cloud decoder and receiving device). Therefore, the point cloud receiving device generates the LOD based on the above LOD generation algorithm.
[0245] The point cloud encoder performs transform coding (15300). As described in reference Figures 1 to 14 the point cloud encoder according to the embodiment selectively uses RAHT coding, predictive transform coding, and lifting transform coding or a combination of one or more of the coding techniques according to the point cloud content.
[0246] According to an embodiment, the predictive transform coding includes prediction based on interpolation. Attributes associated with a point cloud are encoded and decoded in an order defined by a processing definition generated according to the LOD. In each operation, only the points that have been encoded or decoded are considered for prediction. The attribute of a point is predicted based on a weighted average of the attributes (or attribute values) of the neighboring points of the point. However, the points in the neighboring point group may be distributed near or far from the point. Therefore, when a greater weight is assigned to the densely distributed points compared to the points that are less densely distributed or far from the point, the actual correlation between these points can be reflected, and thus the predicted attribute can be calculated more accurately. Accordingly, the attribute (or attribute value) of a point according to an embodiment can be predicted based on the distance to the nearest neighboring points of the point and interpolation-based prediction (or interpolation prediction transform processing) using weights.
[0247] According to an embodiment, (a i ) i∈0...k-1 represents an attribute (or attribute value). represents a set of the k nearest neighboring points of a point. is the j-th decoded and reconstructed attribute. The following equation represents the process of calculating the correlation distance using a cyclic correlation shift matrix.
[0248] [Equation 2]
[0249]
[0250] In this equation, represents the distance between a point and its neighboring points. is the correlation distance (or referred to as the correlation value) calculated by applying a matrix. The weighted average predicted attribute (correlation weight) calculated based on the correlation distance calculated in the above equation is represented as follows
[0251] [Equation 3]
[0252]
[0253] According to an embodiment, the lifting transform coding uses an update operator to calculate the predicted attributes of each point. The lifting transform coding calculates the predicted values and residual values of the points belonging to the highest level of LOD (e.g., LODn). The lifting transform coding can calculate the weights of the points. The update operator of the points can calculate the updated attribute values based on the calculated weights and residual values. The calculated updated attribute values are used to calculate the predicted attributes of the points in the next LOD (e.g., LODn-1 which is one level lower than LODn). One level of LOD includes the points included in other LODs of higher levels. That is, since the points included in the lower-level LODs are more frequently used for prediction, the lifting transform coding based on LOD has a greater impact on the points belonging to the lower-level LODs. Therefore, the update operator can perform the update operation based on the weights updated by adding the weights of neighboring points to the weights of the corresponding points.
[0254] According to an embodiment, the lifting transform coding can use the updated weights reflecting the correlation between neighboring points. The update operator performs the update operation based on the updated weights. Hereinafter, the process of updating the weights based on the correlation between neighboring points will be described. The updated weights reflecting the correlation between neighboring points can be referred to as correlation weights.
[0255] w(P) is the weight associated with point p. The following recursive operation is used to calculate w(P).
[0256] For all points, the value of w(P) is defined as 1.
[0257] Traverse the points in the reverse order of the order defined in the LOD structure.
[0258] For each point Q(i, j) belonging to LOD(j), the weights of the neighboring points of the point are updated. The following represents the update process.
[0259] w(P)←w(P)+w(C[Q(i, j), j]α(P, Q(i, j))
[0260] Here, C[(Q(i, j), j] represents the correlation squared distance between the point Q(i, j) in the j-th set of the nearest neighboring points. Therefore, the correlation weight w(P) according to the embodiment is updated considering the correlation squared distance between neighboring points. The update operator updates the attribute values based on the correlation weights and the prediction residuals.
[0261] According to an embodiment, the update process can be executed by program instructions stored in one or more memories included in the point cloud transmitting device and the receiving device. According to an embodiment, the program instructions can be executed by the point cloud encoder and / or decoder (or processor), and cause the point cloud encoder and / or decoder to update the attribute values.
[0262] According to an embodiment, the point cloud encoder performs quantization (15400). Since the quantization is the same as the quantization described in the reference, a detailed description thereof will be skipped. The weights of the correlations between points as described above can also be applied to the quantization. Figures 1 to 14 According to an embodiment, the point cloud encoder performs arithmetic coding (15500). Since the arithmetic coding is the same as the arithmetic coding described in the reference, a detailed description thereof will be skipped. The weights of the correlations between points as described above can also be applied to the arithmetic coding.
[0263] According to an embodiment, the point cloud encoder performs arithmetic coding (15500). Since the arithmetic coding is the same as the arithmetic coding described in the reference, a detailed description thereof will be skipped. The weights of the correlations between points as described above can also be applied to the arithmetic coding. Figures 1 to 14 According to an embodiment, the point cloud encoder performs arithmetic coding (15500). Since the arithmetic coding is the same as the arithmetic coding described in the reference, a detailed description thereof will be skipped. The weights of the correlations between points as described above can also be applied to the arithmetic coding.
[0264] Figure 16 An example of a point and its neighboring points is illustrated.
[0265] Figure 16 Shown as a reference Figure 15 described target point P of the predictive transform coding and four neighboring points C1, C2, C3 and C4. As described in the reference Figure 15 described, the point cloud encoder performs predictive transform coding (e.g., interpolation-based prediction). The attribute (or attribute value) of point p can be a weighted average of the attributes of C1, C2, C3, and C4, which are the nearest neighboring points of point p. In this figure, δ j represents the distance between point p and neighboring point C1, and δ j+1 represents the distance between point p and neighboring point C2. δ j+2 represents the distance between point p and neighboring point C3, and δ j+3 represents the distance between point p and neighboring point C4. As Figure 16 indicated by the dashed line in, compared with neighboring point C1, neighboring points C2, C3, and C4 are relatively densely arranged.
[0266] Therefore, according to the embodiment, the predictive transform coding will Figure 16 illustrated in the distance between point p and neighboring points in is multiplied by the cyclic correlation shift matrix shown in Equation 2 to calculate a correlation distance reflecting the correlation between neighboring points (e.g., ). Therefore, the predictive transform coding calculates the weighted average attribute (Equation 3) considering the correlation according to the density of neighboring points.
[0267] Figure 17 An exemplary bitstream structure diagram is shown.
[0268] A point cloud processing device (e.g., a transmitting device described in the reference Figure 1 , Figure 12 and Figure 14 described) can transmit the encoded point cloud data in the form of a bitstream. The bitstream is a series of bits that form a representation of the point cloud data (or point cloud frame).
[0269] Point cloud data (or a point cloud frame) can be segmented into tiles and slices.
[0270] Point cloud data can be segmented into multiple slices and encoded in a bitstream. A slice is a set of points and is represented as a series of syntax elements representing all or part of the encoded point cloud data. A slice may or may not be dependent on other slices. Additionally, a slice may include a geometric structure data unit and may include one or more attribute data units, or may not include an attribute data unit. As described above, attribute encoding is performed based on geometric structure encoding. Thus, an attribute data unit is based on the geometric structure data unit in the same slice. That is, a point cloud data receiving device (e.g., receiving device 10004 or point cloud video decoder 10006) can process attribute data based on the decoded geometric structure data. Therefore, the geometric structure data unit must be before the associated attribute data unit in a slice. The data units in a slice must be consecutive, and no order of the slices is specified.
[0271] A tile is a (three-dimensional) rectangle in the shape of a parallelepiped within a bounding box (e.g., the bounding box described with reference to Figure 5 ). The bounding box may contain one or more tiles. A tile may completely or partially overlap with another tile. A tile may include one or more slices.
[0272] Thus, a point cloud data transmitting device can process data corresponding to a tile according to importance and provide high-quality point cloud content. That is, a point cloud data transmitting device according to an embodiment can perform point cloud compression encoding with better compression efficiency and appropriate latency on data corresponding to an area important to a user.
[0273] According to an embodiment, the bitstream contains signaling information and multiple slices (slice 0,..., slice n). As shown in the figure, before the signaling information slice in the bitstream. Thus, a point cloud data receiving device can first obtain the signaling information and sequentially or selectively process the multiple slices based on the signaling information. As shown in the figure, slice 0 includes a geometric structure data unit (Geom0 0 ) and two attribute data units (Attr0 0 and Attr1 0 ). Additionally, the geometric structure data unit is before the attribute data units in the same slice. Thus, a point cloud data receiving device processes (decodes) the geometric structure data unit (or geometric structure data) and then processes the attribute data unit (or attribute data) based on the processed geometric structure data. According to an embodiment, the signaling information may be referred to as signaling data, metadata, etc., and is not limited to this example.
[0274] According to an embodiment, the signaling information includes a Sequence Parameter Set (SPS), a Geometry Parameter Set (GPS), and one or more Attribute Parameter Sets (APS). The SPS encodes information about the entire sequence such as profiles and levels, and may include comprehensive information about the entire sequence (sequence level) such as picture resolution and video format. The GPS is information about the geometric structure coding applied to the geometric structure included in the sequence (bitstream). The GPS may include information about the octree (e.g., the octree described in Figure 6 the octree described in Figure 6 ) and information about the octree depth. The APS is information about the attribute coding applied to the attributes included in the sequence (bitstream). As shown in the figure, according to the identifier for identifying the attribute, the bitstream includes one or more APSs (e.g., APS0, APS1,... shown in the figure).
[0275] According to an embodiment, the signaling information may further include TPS. The TPS is information about tiles and may include information about identifiers, tile sizes, etc. According to an embodiment, the signaling information is information at the sequence level (i.e., bitstream level) and is applied to the corresponding bitstream. Additionally, the signaling information has a syntax structure including syntax elements and descriptors for describing the syntax elements. Pseudo-code for describing the syntax may be used. Additionally, the point cloud receiving device may sequentially parse and process the syntax elements in the syntax.
[0276] Although not shown in the figure, according to an embodiment, the geometric structure data unit and the attribute data unit respectively include a geometric structure header and an attribute header. According to an embodiment, the geometric structure header and the attribute header are signaling information applied at the corresponding slice level and have the above syntax structure.
[0277] According to an embodiment, the geometric structure header contains information (or signaling information) for processing the corresponding geometric structure data unit. Therefore, the geometric structure header first appears in the geometric structure data unit. The point cloud receiving device may first parse the geometric structure header to process the geometric structure data unit. The geometric structure header is related to the GPS that contains information about the entire geometric structure. Therefore, the geometric structure header contains information specifying the gps_geom_parameter_set_id included in the GPS. The geometric structure header also contains tile information (e.g., tile_id) and a slice identifier related to the slice to which the geometric structure data unit belongs.
[0278] According to an embodiment, the attribute header contains information (or signaling information) for processing the corresponding attribute data unit. Thus, the attribute header first appears in the attribute data unit. The point cloud receiving device may first parse the attribute header to process the attribute data unit. The attribute header is associated with the APS that contains information about all attributes. Thus, the attribute header contains information specifying the aps_attr_parameter_set_id included in the APS. As described above, attribute decoding is based on geometric structure decoding. Thus, to determine the geometric structure data unit associated with the attribute data unit, the attribute header contains information specifying the slice identifier included in the geometric structure header.
[0279] When the point cloud data processing device performs attribute encoding based on the correlation weights described in the reference Figures 15 to 16 the signaling information in the bitstream may include information about the correlation weights. According to an embodiment, the information about the correlation weights may be included in the signaling information at the sequence level (e.g., SPS, APS, etc.) or included in the slice level (e.g., the attribute header).
[0280] Figure 18 An example of the signaling information according to an embodiment is shown.
[0281] Figure 18 Shown is the reference Figure 17 describes the syntax structure of the SPS, and illustrates an example in which the information about the correlation weights described in the reference Figure 17 is included in the SPS at the sequence level.
[0282] The syntax of the SPS includes the following syntax elements.
[0283] The profile_compatibility_flags indicate whether the bitstream conforms to a specific profile for decoding or conforms to another profile. The profile specifies the constraints imposed on the bitstream to specify the ability to decode the bitstream. Each profile is a subset of algorithm features and constraints and is supported by all decoders that follow the profile. This is used for decoding and can be defined according to standards.
[0284] The level_idc indicates the level applied to the bitstream. This level is used in all profiles. Generally, the level corresponds to a specific decoder processing load and memory capacity.
[0285] The sps_bounding_box_present_flag indicates whether there is information about the bounding box in the SPS. A sps_bounding_box_present_flag equal to 1 indicates the existence of information about the bounding box. A sps_bounding_box_present_flag equal to 0 indicates that the information about the bounding box is undefined.
[0286] The following is the information about the bounding box included in the SPS when the sps_bounding_box_present_flag is equal to 1.
[0287] The sps_bounding_box_offset_x indicates the quantized x-axis offset of the source bounding box in the Cartesian coordinate system including the x, y, and z axes.
[0288] The sps_bounding_box_offset_y indicates the quantized y-axis offset of the source bounding box in the Cartesian coordinate system including the x, y, and z axes.
[0289] The sps_bounding_box_offset_z indicates the quantized z-axis offset of the source bounding box in the Cartesian coordinate system including the x, y, and z axes.
[0290] The sps_bounding_box_scale_factor specifies the scaling factor used to indicate the size of the source bounding box.
[0291] The sps_bounding_box_size_width indicates the width of the source bounding box in the Cartesian coordinate system including the x, y, and z axes.
[0292] The sps_bounding_box_size_height indicates the height of the source bounding box in the Cartesian coordinate system including the x, y, and z axes.
[0293] The sps_bounding_box_size_depth indicates the depth of the source bounding box in the Cartesian coordinate system including the x, y, and z axes.
[0294] The SPS syntax also includes the following elements.
[0295] The sps_source_scale_factor indicates the scaling factor of the source point cloud data.
[0296] The sps_seq_parameter_set_id is the identifier of the SPS for reference by other syntax elements (e.g., the seq_parameter_set_id in the GPS).
[0297] The sps_num_attribute_sets indicates the number of encoded attributes in the bitstream. The value of sps_num_attribute_sets is in the range of 0 to 64.
[0298] The following for statement includes as many elements indicating information about each attribute as the number indicated by sps_num_attribute_sets. In the figure, i represents each attribute (or attribute set), and the value of i is greater than or equal to 0 and less than the number indicated by sps_num_attribute_sets.
[0299] The attribute_dimension_minus1[i] indicates a value that is 1 less than the number of components of the i-th attribute. When the attribute is color, the attribute corresponds to a three-dimensional signal representing the light characteristics of the target point. For example, the attribute can be signaled as three components of RGB (red, green, blue). The attribute can be signaled as three components of YUV, namely luminance (illuminance) and two chrominances (saturation). When the attribute is reflectance, the attribute corresponds to a one-dimensional signal representing the intensity ratio of the light reflectance of the target point.
[0300] The attribute_instance_id[i] indicates the instance id of the i-th attribute. The attribute_instance_id is used to distinguish the same attribute label and attribute.
[0301] The attribute_bitdepth_minus1[i] indicates a value that is 1 less than the bit depth of the first component of the i-th attribute signal. Adding 1 to the value specifies the bit depth of the first component.
[0302] The attribute_cicp_colour_primaries[i] indicates the chromaticity coordinates of the color attribute source primaries of the i-th attribute.
[0303] The attribute_cicp_transfer_characteristics[i] either indicates the reference photoelectric transfer characteristic function of the color attribute as a function of the source input linear optical intensity Lc with a nominal true value range of 0 to 1, or indicates the reciprocal of the reference electro-optical transfer characteristic function as a function of the output linear light intensity Lo with a nominal true value range of 0 to 1.
[0304] The attribute_cicp_matrix_coeffs[i] indicates the matrix coefficients used to derive the luminance and chrominance signals from the RBG or YXZ primaries.
[0305] The attribute_cicp_video_full_range_flag[i] indicates the range and black level of the luminance and chrominance signals derived from the true value components of E′Y, E′PB, and E′PR or E′R, E′G, and E′B.
[0306] As described above, the SPS syntax contains information about the relevant weights. The following elements indicate information about the relevant weights with respect to the reference Figure 15 and Figure 16 described relevant weights.
[0307] The attribute_correlated_weight_flag indicates whether the calculated distance should be correlated. An attribute_correlated_weight_flag equal to 1 indicates that the calculated distance should be correlated (i.e., using the correlated weights). An attribute_correlated_weight_flag equal to 0 indicates that the calculated distance is not correlated. The default value is inferred to be 0.
[0308] The attribute_correlated_weight_method indicates the method for calculating the correlated weights when the attribute_correlated_weight_flag is equal to 1. For example, the methods for calculating the correlated weights include those described by Equations 2 to 4 with respect to the reference Figure 15 described. Accordingly, the point cloud receiving device calculates the correlated weights according to the method indicated by the attribute_correlated_weight_method and performs interpolation-based prediction and lifting transform coding.
[0309] The information about the correlated weights according to the embodiment is not limited to the above examples. Accordingly, the information about the correlated weights may also include information about the correlated weight variables and information about the number of neighboring nodes in the correlated point set.
[0310] According to an embodiment, the syntax of the SPS includes the following syntax elements.
[0311] The known_attribute_label_flag[i], known_attribute_label[i], and attribute_label_fourbytes[i] are used together to identify the data type carried in the i-th attribute. The known_attribute_label_flag[i] indicates whether the attribute is identified by the value of the known_attibute_label[i] or by another object identifier attribute_label_fourbytes[i].
[0312] The sps_extension_flag indicates whether sps_extension_data_flag exists in the SPS. A sps_extension_flag equal to 0 indicates that the sps_extension_data_flag syntax element does not exist in the SPS syntax structure. The value 1 of the sps_extension_flag is reserved for future use. The decoder may ignore all sps_extension_data_flag syntax elements following a sps_extension_flag equal to 1.
[0313] The sps_extension_data_flag indicates whether there is data reserved for future use and can have any value.
[0314] The SPS syntax is not limited to the above examples. For signaling efficiency, it may also include additional elements or may exclude some of the elements shown in the figure. Some elements may be signaled by signaling information other than the SPS (e.g., APS, attribute headers, etc.) or by attribute data units.
[0315] Figure 19 An example of signaling information according to an embodiment is illustrated.
[0316] Figure 19 is a reference Figure 17 to the syntax structure of the APS described and illustrates an example in which information about the relevant weights described with reference to Figure 17 is included in the APS at the sequence level.
[0317] The syntax of the APS includes the following syntax elements.
[0318] The aps_attr_parameter_set_id indicates the identifier of the APS for reference by other syntax elements. The value of the aps_attr_parameter_set_id is in the range of 0 to 15. One or more attribute data units are included in the bitstream (e.g., the bitstream described with reference to Figure 17 ), and each attribute data unit includes an attribute header. The attribute header includes a field having the same value as the aps_attr_parameter_set_id (e.g., ash_attr_parameter_set_id). The point cloud receiving device according to the embodiment parses the APS and processes the attribute data units referring to the same aps_attr_parameter_set_id based on the parsed APS and the attribute header.
[0319] The aps_seq_parameter_set_id specifies the value of the sps_seq_parameter_set_id that activates the SPS. The value of the aps_seq_parameter_set_id is in the range of 0 to 15.
[0320] The attr_coding_type indicates the attribute coding type for a given value of attr_coding_type. Attribute coding means attribute encoding. As described above, attribute encoding uses at least one of RAHT coding, predictive transform coding, and lifting transform coding, and the attr_coding_type indicates any one of the above-mentioned three coding types. Therefore, the value of the attr_coding_type is equal to any one of 0, 1, or 2 in the bitstream. Other values of the attr_coding_type may be used by ISO / IEC subsequently. Therefore, the point cloud receiving device according to the embodiment ignores the attr_coding_type having a value other than 0, 1, and 2. When the attr_coding_type is equal to 0, the attribute coding type is predictive transform coding. When the attr_coding_type is equal to 1, the attribute coding type is RAHT coding. When the attr_coding_type is equal to 2, the attribute coding type is lifting transform coding. The value of the attr_coding_type may vary and is not limited to this example. For example, an attr_coding_type equal to 0 indicates that the attribute coding type is RAHT coding, an attr_coding_type equal to 1 indicates that the attribute coding type is the LOD for predictive transform coding, and an attr_coding_type equal to 2 indicates that the attribute coding type is the LOD for lifting transform coding.
[0321] The aps_attr_initial_qp indicates the initial value of the variable SliceQp for each slice with reference to the current APS.
[0322] The aps_attr_chroma_qp_offset specifies the offset applied to the initial quantization parameter signaled by the aps_attr_initial_qp.
[0323] The aps_slice_qp_delta_present_flag indicates whether there is a component QP offset indicated by the ash_attr_qp_offset in the header of the attribute data unit.
[0324] As described above, the APS syntax contains information about the reference Figure 15and Figure 16 information on the relevant weights described. As shown in the figure, the APS syntax includes the attribute_correlated_weight_flag and the attribute_correlated_weight method. Each element is the same as the element described with reference to Figure 18 and thus the description thereof is skipped.
[0325] The information on the relevant weights according to the embodiment is not limited to the above examples. Thus, the information on the relevant weights may also include information on the relevant weight variables and information on the number of neighboring nodes in the relevant point set.
[0326] When the value of attr_coding_type indicates lifting transform coding, the following syntax elements exist in the APS.
[0327] lifting_num_pred_nearest_neighbours specifies the maximum number of nearest neighbours to be used for prediction.
[0328] lifting_max_num_direct_predictors indicates the maximum number of predictors to be used for direct prediction.
[0329] lifting_search_range specifies the search range for determining the nearest neighbours to be used for prediction and for establishing distance-based LOD.
[0330] lifting_lod_regular_sampling_enabled_flag indicates the sampling strategy for establishing LOD. A lifting_lod_regular_sampling_enabled_flag equal to 1 indicates using a regular sampling strategy to establish LOD. A lifting_lod_regular_sampling_enabled_flag equal to 0 indicates using a distance-based sampling strategy to establish LOD.
[0331] lifting_num_detail_levels_minus1 indicates the number of LODs for attribute coding. The value of lifting_num_detail_levels_minus1 is greater than or equal to 0.
[0332] The following for loop includes as many elements indicating information about each LOD as the number indicated by lifting_num_detail_levels_minus1. In this figure, idx indicates each LOD. The value of idx is greater than or equal to 0 and less than the number indicated by lifting_num_detail_levels_minus1.
[0333] When the value of lifting_lod_regular_sampling_enabled_flag is 1, lifting_sampling_period[idx] is included. When the value of lifting_lod_regular_sampling_enabled_flag is 0, lifting_sampling_distance_squared[idx] is included.
[0334] lifting_sampling_period[idx] specifies the sampling period for LOD idx.
[0335] lifting_sampling_distance_squared[idx] specifies the scaling factor used to derive the square of the sampling distance for LOD idx.
[0336] When attr_coding_type indicates that the attribute coding is predictive transform coding, the APS includes the following syntax elements.
[0337] lifting_adaptive_prediction_threshold indicates the threshold used to enable adaptive prediction.
[0338] lifting_intra_lod_prediction_num_layers specifies the number of LOD layers that can refer to decoded points in the same LoD layer to generate the predicted value of the target point.
[0339] According to an embodiment, the syntax of the APS includes the following syntax elements.
[0340] The aps_extension_flag indicates whether the aps_extension_data_flag exists in APS. An aps_extension_flag equal to 0 indicates that the aps_extension_data_flag syntax element does not exist in the APS syntax structure. The value 1 of the aps_extension_flag is reserved for future use. The decoder may ignore all aps_extension_data_flag syntax elements following an aps_extension_flag equal to 1.
[0341] The aps_extension_data_flag indicates whether there is data for future use and can have any value.
[0342] The APS syntax is not limited to the above examples. For signaling efficiency, it may also include additional elements or may exclude some of the elements shown in the figure. Some elements may be signaled through signaling information other than APS (e.g., attribute headers, etc.) or through attribute data units.
[0343] As described in the reference Figures 15 to 19 When there are two or more neighboring points, the correlation weight can be calculated based on the correlation between these points. As a method for calculating the correlation weight, the correlation weight can be calculated without changing the existing attribute coding algorithm, thus ensuring the flexibility of the system design. Additionally, by using the correlation weight to replace the weighting constant used in the prediction / lifting transform, the encoding and decoding performance is improved. The method for calculating the correlation weight according to the embodiment and the correlation weight can be applied to all functions that require prediction algorithms such as quantization.
[0344] The method for calculating the correlation degree (e.g., Equation 2) may or may not include the distance between the corresponding points. Additionally, the method for calculating the correlation degree may include operations of addition, multiplication, and division using the matrix described in the reference Figure 2 The method for calculating the correlation degree may also include operations of removing the correlation value or using the correlation value assignment and multiplying by each constant. Each constant may include not only integers but also complex numbers and may have a fixed value or a variable value.
[0345] The correlation weight is the sum of the weights based on the correlation degree and corresponds to the average value or variance (e.g., Equation 3). The weights not combined with the correlation may be used as the average value or variance.
[0346] Figure 20 Illustrates a method for encoding the correlation weight according to an embodiment.
[0347] Figure 20Instructions are shown that represent methods of calculating relevant weights (e.g., Equation 2 and Equation 3) in various ways when there are three points.
[0348] In this figure, weighted_sum represents the approximate sum (sum) of each point and relevance, and the method of calculating relevance can vary according to the embodiment. w0, w1, and w2 represent the relevant weights reflecting the relevance calculated for the corresponding points.
[0349] The first box represents the process of calculating relevance by multiplying the relevant values by any constants α, β, and γ. The second box represents the process of calculating relevance based solely on the distance between points. The third box represents the process of calculating relevance based on the square of the point distance. The fourth box represents the process of calculating relevance based on the sum of the point distances. Each box represents the relevant weights w0, w1, and w2 for the corresponding points, which reflect the relevance calculated according to the process of calculating relevance. The relevant weights can vary according to the calculation method and the type of relevance. Additionally, the method of calculating relevance is not limited to the above examples.
[0350] Reference Figures 1 to 20 The described point cloud processing device supports spatial adaptive decoding. Spatial adaptive decoding is performed on all or part of the geometry and / or attributes according to the decoding performance of the point cloud receiving device (e.g., Figure 1 the receiving device 10004 of Figure 10 and Figure 11 the point cloud decoder of Figure 13 and the receiving device) to provide decoding of point cloud content at various resolutions. According to an embodiment, the part of the geometry and attributes is referred to as partial geometry and partial attributes. The adaptive decoding applied to the geometry according to an embodiment is referred to as adaptive geometry decoding or geometry adaptive decoding. The adaptive decoding applied to the attributes according to an embodiment is referred to as adaptive attribute decoding or attribute adaptive decoding. As described in reference Figures 1 to 17 The points of the point cloud content are distributed in 3D space and the distributed points are represented in an octree structure (e.g., the octree described in reference Figure 6 ). The octree structure is an octal tree structure in which the depth increases from the upper node to the lower node. According to an embodiment, the depth is referred to as the level and / or layer.
[0351] The point cloud processing device (or geometric structure encoder) performs geometric structure encoding based on the octree structure. Additionally, the point cloud processing device (or attribute encoder) generates the LOD and performs attribute encoding (e.g., RAHT transform, prediction transform, lifting transform, etc.) based on the octree structure. Since the LOD is generated based on the octree structure, the octree structure is regarded as dividing the grouping of points and organizing the number of points of the geometric structure and attributes. The levels of the LOD can correspond to the depth of the octree. Since the LOD (or octree depth) must be large enough to represent the original quality, spatial scalability is very useful when the source point cloud is densely arranged even in a local area. Through spatial scalability, the point cloud receiving device (or decoder) can provide low-resolution point cloud content such as a thumbnail with low decoder complexity and / or small bandwidth. When spatial scalability decoding is supported, the point cloud processing device sends information for spatial scalability decoding by referring to Figure 17 the signaling information (e.g., SPS, APS, attribute headers, etc.) contained in the described bitstream.
[0352] The point cloud receiving device obtains the information for spatial scalability decoding through the signaling information contained in the bitstream. The point cloud receiving device performs geometric structure decoding on all or part of the geometric structures corresponding to a specific depth (or level) from the upper nodes to the lower nodes of the octree structure. As described above, attribute decoding is based on geometric structure decoding. Therefore, the point cloud receiving device can generate the LOD based on the decoded geometric structure (or decoded octree structure) and perform attribute decoding on all and / or part of the attributes (e.g., RAHT transform, prediction transform, lifting transform, etc.).
[0353] Figure 21 An example of spatial scalability decoding is illustrated.
[0354] The arrow 1800 shown in the figure indicates the direction in which the level of the LOD increases.
[0355] As referred to Figures 1 to 14 described, the point cloud processing device generates the LOD based on the octree structure. The LOD is designed to manage the attributes of points with an octree structure, and an increase in the LOD value indicates an increase in the detail of the point content. The LOD can correspond to one or more depths of the octree structure. The highest node of the octree structure corresponds to the lowest depth or the first depth and is called the root. The lowest node of the octree structure corresponds to the highest depth or the last depth and is called the leaf. The depth of the octree structure increases in the direction from the root to the leaf, which is the same as the direction indicated by the arrow.
[0356] According to an embodiment, the point cloud decoder performs decoding 1811 for providing full-resolution point cloud content or decoding 1812 for providing low-resolution point cloud content according to its performance. The point cloud decoder provides full-resolution point cloud content by decoding 1811 the geometric structure bitstream 1811-1 and the attribute bitstream 1812-1 corresponding to the entire octree structure. The point cloud decoder provides low-resolution point cloud content by decoding 1812 the partial geometric structure bitstream 1812-1 and the partial attribute bitstream 1812-2 corresponding to a specific depth of the octree structure. Figure 21 An example of the lifting transform as the attribute decoding is illustrated, but the embodiment is not limited to this example.
[0357] As described above, the signaling information (e.g., SPS, APS, attribute header, etc.) in the bitstream (e.g., Figure 17 the bitstream) may include adaptability information (e.g., scalable_lifting_enabled_flag or lifting_scalability_enabled_flag) related to the spatial adaptability decoding (or lifting transform) at the sequence level or slice level. As described above, the attribute decoding is performed based on the decoded geometric octree structure. The information related to the spatial adaptability decoding (or lifting transform) indicates whether the entire octree structure or a partial octree structure is required to decode partial attributes.
[0358] The point cloud receiving device obtains the signaling information of the bitstream and performs adaptable attribute decoding based on the entire octree structure or the partial octree structure as the result of the geometric structure decoding according to the information related to the spatial adaptability decoding.
[0359] As described above, the LOD is generated based on the octree structure (e.g., Figure 15 operation 1520 in). Therefore, the levels of the LOD are generated based on the depth of the octree. When the octree structure changes, the LOD structure also changes. The point cloud encoder (e.g., the point cloud encoder described with reference to Figure 15 performs the lifting transform based on the LOD (e.g., Figure 15 operation 1530 in). As described above, the point cloud processing device performs the lifting transform coding from the highest level of the LOD to the lowest level of the LOD. According to an embodiment, the lifting transform coding uses an update operator to calculate the predicted attributes of each point. The update operator for the corresponding point can calculate the updated attribute value based on the calculated weight and residual value. According to an embodiment, the point cloud processing device determines (or calculates or derives) the quantization weight according to the quantization weight derivation process and performs quantization based on the determined quantization weight.
[0360] The point cloud receiving device (or point cloud decoder) according to the embodiment may perform inverse quantization by using an update operator to restore the attribute value and calculating the quantization weight in the same manner as the point cloud processing device.
[0361] The density of the points belonging to one level of LOD is different from the density of the points belonging to another level of LOD. For example, the density of the points belonging to one level of LOD is lower than the density of the points belonging to a lower level of LOD. Therefore, according to the embodiment, the quantization weight is derived from the sum of distances in a higher level of LOD. However, when performing spatially adaptable decoding, the point cloud receiving device cannot accurately calculate the quantization weight because it does not have information on the low LOD. Therefore, the quantization weight is fixed by the number of points of the LOD. The following shows the process of calculating the quantization weight for each LOD to support attribute decoding (e.g., lifting transform decoding) performed based on a partial octree structure.
[0362] For i = 0 to
[0363]
[0364] Here, i is a parameter indicating the level of each LOD, and the value of i is greater than or equal to 0 and less than the number of LODs (LODcount). pointCount is the number of points belonging to the corresponding LOD, and predictorCount represents the number of predictors of the points belonging to an LOD lower than the corresponding LOD. predictorCount[i] represents the number of predictors of the points belonging to the corresponding LOD. As shown in this formula, the weight is calculated based on the number of attributes and a fixed constant (e.g., kFixedPointweightShift).
[0365] The above formula is used to calculate the quantization weight based on the LOD in which the points are densely arranged rather than the information of the points having predicted values for each LOD. Therefore, when the points are evenly distributed in one or more LODs, or when the gap of the LOD index is large and the distribution of the points is irregular, the performance of the decoder deteriorates according to the quantization weight.
[0366] Therefore, the point cloud transmitting device and the point cloud receiving device according to the embodiment perform an improved quantization weight derivation process to obtain a mathematical optimization when performing a lifting transform coding (e.g., a lifting transform coding performed based on a partial octree structure). The improved quantization weight derivation process can calculate an improved quantization weight that can be changed without applying a fixed constant. Therefore, the improved quantization value derivation process minimizes the change in the quantization weight according to the point cloud system without changing the fixed constant for each system. In addition, since the quantization weight corresponds to a mathematically optimized value, a higher performance gain is ensured compared to the existing quantization weight. In addition, since the improved quantization weight derivation process does not require operations on each point, the complexity of the point cloud receiver can be reduced.
[0367] According to an embodiment, the quantization weight derivation process can be performed by program instructions stored in one or more memories included in the point cloud transmitting device and the receiving device. The program instructions are executed by the point cloud encoder and / or decoder (or processor) and cause the point cloud encoder and / or decoder to calculate / deduce the quantization weight.
[0368] According to an embodiment, the lifting transform is represented as a linear function representing the sum of the quantization weight and the attribute value. The following equation is a linear function representing the lifting transform and is determined to satisfy the maximum benefit with the total resource allocation value. The resource according to the embodiment refers to the sum of the products of the weight of each point and the attribute value (the voltage representing the attribute value).
[0369] [Equation 4]
[0370]
[0371] The parameter j is the index of each point, greater than or equal to 0 and less than or equal to N. The parameter N represents the total number of points in the expected point cloud. That is, N corresponds to the total number of points to be transmitted by the point cloud transmitting device. The parameter w j is the quantization weight (or improved quantization weight) and is determined (or calculated or deduced) by the improved quantization weight derivation process. The parameter a j represents the attribute value of each point. SumAttribute represents the sum of the quantization weight and the attribute value and takes the form of multiplication and addition of N linear functions.
[0372] The point cloud receiving device can obtain the brightness and color values of the corresponding point through the value of SumAttribute. The condition for optimizing the function representing the above lifting transform is represented by the following equation.
[0373] [Equation 5]
[0374]
[0375] Experience
[0376] w j ≥ 1, where j = 1, ..., N
[0377] As shown in the above equation, the function representing the lifting transformation can be optimized by applying constraints to the parameter w j of. According to an embodiment, w j is greater than or equal to 1. Additionally, w j (where j is greater than or equal to 0 and less than or equal to N) the sum (or total weight) of the N values is less than or equal to the total number of predicted points (TotalPredictedCount). This is intended to prevent the quantization weights from increasing infinitely and to prevent overflow problems from occurring in the memory required for the calculation. Additionally, the equation for driving the lifting transformation consists of a linear form and an integer space, maintaining the convexity of the function.
[0378] Any defined function fj constituting the above equation consists of multiplication and addition of integers and decimals. Therefore, the function fj is represented as a convex function, and the function fj is a concave function according to the necessary and sufficient conditions. When the values of w1, w2, .., and w N are greater than or equal to 1, the functions f1 of w1, f2 of w1, ... and the function fN are also convex functions, so the combination of two or more of the convex functions takes the form of a convex function. According to an embodiment, the weight (or weight constant) w j can be modified or changed and has an optimized value. Additionally, when the function fj is a combination or clustering of convex functions, g, which is the clustering function among the fi, i.e., the clustering function, is also configured as a convex function. Therefore, the optimized value of the clustering of the weights grouped by LOD can be obtained. According to an embodiment, the function g has a maximum value and a minimum value.
[0379] As described above, the quantization weights are determined (calculated or derived) by TotalPredictedCount. When the point cloud transmitting device and the point cloud receiving device according to an embodiment transmit / receive information about the total number of points (e.g., TotalCount), the quantization weights have a global optimized value determined based on the total number of points. When the point cloud receiving device predicts TotalPredictedCount, the quantization weights have a local optimized value determined based on TotalPredictedCount. The above equation uses various methods such as Karush - Kuhn - Tucker (KKT) conditions, geometric / non - geometric structure techniques, and group - based power constraints. Since there is a Lagrange multiplier (λ ∈ R) in the real space, inequality constraints, and w representing the optimized value are defined *≥0, so the point cloud transmitting device and receiving device according to the embodiment can perform an improved quantization weight derivation process by generating conditions for calculating the quantization weight (e.g., constraint change).
[0380] Figure 22 An exemplary improved quantization weight derivation process is shown.
[0381] Figure 22 It is a flowchart illustrating an improved quantization weight derivation process. The flowchart includes one or more operations. These operations can be executed simultaneously or sequentially.
[0382] The improved quantization weight derivation process defines an initial weight (or initial quantization weight) (22100). According to the embodiment, the initial weight is determined based on the average transceiver power. That is, considering the data memory size, computational complexity, etc., the initial weight can be determined based on the average energy value or total energy value at the 32-bit and 64-bit levels. Additionally, when all point cloud data is normalized, the initial weight is determined to be 1.
[0383] The improved quantization weight derivation process reorders / counts the number of predicted points for each LOD (e.g., TotalPredictedCount as described above) (22200). This point represents a geometric structure point or an attribute point. The number of predicted points for each LOD can be obtained from the decoded geometric structure. The number of predicted points for each LOD is stored as a parameter. The improved quantization weight derivation process can store the number of actual points (e.g., TotalCount as described above) as a parameter without counting the number of predicted points.
[0384] The improved quantization weight derivation process calculates the total constraint (22300). Apply the constraint to the weight or the sum of weights (e.g., w in Equation 8). j sum). According to the embodiment, the constraint can include but is not limited to the number of points of the total cumulative LOD for each LOD, the number of points of some LODs, the number of points belonging to the subset, and the number of points grouped according to the LOD.
[0385] The improved quantization weight derivation process calculates the weight assignment determined based on the calculated constraint (22400). That is, the improved quantization weight derivation process allocates resources based on the constraint according to each LoD level and the points in the LoD. As described above, the resource is the attribute multiplied by the sum of the weights of each point. Therefore, the resources are allocated proportionally according to the LoD level and the points in the LoD.
[0386] That is, according to the weight based on the constraint, the total number of points is allocated to each LoD level.
[0387] Improve the quantization weight derivation process to update the weights (22500). The updated weights (or updated quantization weights) can be defined as the assigned weights, or can be generated by accumulating the assigned weights to the initial weights, or by modifying and combining some existing updated weights. The finally updated weights have optimized values.
[0388] In the following, the process of calculating the weight assignment determined based on the constraints calculated based on the reference Figure 22 will be described.
[0389] As described above, the improved quantization weight derivation process is based on the total number of prediction points (e.g., TotalPredictedCount above). Since the total number of prediction points is limited, the improved quantization weight derivation process can derive the same optimized value or a similar optimized value. According to an embodiment, the optimized value is calculated based on the ratio of the resources (the number of accumulated points of the LOD) to the total resources (the total number of points).
[0390] The process of calculating the number of accumulated points for each LOD is executed by program instructions stored in one or more memories included in the point cloud sending device and the receiving device. According to an embodiment, the program instructions are executed by the point cloud encoder and / or decoder (or processor), and cause the point cloud encoder and / or decoder to calculate the number of accumulated points for each LOD. The process of calculating the i-th optimized value (improved quantization weight) is represented as follows.
[0391] For i = 0 < LoDCount {
[0392] OptimolWeight[i] = numberOfPointsPerLOD[LodCount - 1] / numberOfPointsPerLOD[i];
[0393] }
[0394] OptimalWeight[i] represents the optimized value for the i-th LOD. numberOfPointsPerLOD[LodCount - 1] is the number of all points up to the LOD level corresponding to the value one less than the value of LoDCount. That is, numberOfPointsPerLOD[LodCount - 1] represents the sum of the number of points belonging to each LOD from the LOD at the level where i equals 0 to the LOD at the level where i equals LodCount - 1 (for example, the number of points of LOD 0 equal to 1, the number of points of LOD 1 equal to 7,..., and the sum of the number of points of LOD LodCount - 1 equal to XX). According to an embodiment, numberOfPointsPerLOD[LodCount - 1] can be stored in the form of an index or the like, and can be obtained from the maximum index stored.
[0395] As described above, the i-th optimized value is generated based on the constraint and Ratio[i]. According to an embodiment, Ratio[i] can change according to the irregularity of the total energy that appears in the actual environment such as the point cloud noise error. For example, numberOfpointsperLOD[i - 1] can be used instead of numberOfpointsperLOD[i], or a part of numberOfpointsperLOD[i] can be grouped and regarded as a variable. According to an embodiment, the grouped variables can include sequential variables such as i + 1 and i + 2 or non-sequential variables such as i, i + 4, and i + 6. The constraint numberOfPointsPerLOD[LodCount - 1] can be represented as numberOfPointsPerLOD[k], where k is i - 1, i - 2,..., 0 or i + 1, i + 2. In addition, a part of numberOfPointsPerLOD[k] can be grouped and regarded as a variable. For example, according to an embodiment, the grouped variables include sequential variables such as i + 1 and i + 2 or non-sequential variables such as k, k + 4, and k + 6. A specific constant can be subtracted from, added to, or combined with NumberOfPointsPerLOD[i] and NumberOfPointsPerLOD[k] represented by the variables i and k. For example, when two constants α and β are given for i and k, numberOfPointsPerLOD[i](+ or -)α or numberOfPointsPerLOD[i](* or / )α can be obtained, and numberOfPointsPerLOD[k](+ or can be -)β or numberOfPointsPerLOD[k](* or / )β can be obtained.
[0396] Different improved quantization weight derivation processes can be applied to points at corresponding LOD levels with the same attribute. For example, when i = 1, the optimized value can be inferred by numberOfPointPerLOD[LoDCount - 1] / numberOfPointsPerLoD[i]. When i > 1, the optimized value can be inferred by (numberOfPointPerLOD[LoDCount - 1] - numberOfPointsPerLoD[i]) / numberOfPointsPerLoD[i].
[0397] Different improved quantization weight derivation processes can be applied to points according to the LOD levels of different attributes. Therefore, the quantization weights for the LOD levels given when the attribute is reflectance and the quantization weights for the LOD levels given when the attribute is color have different optimized values.
[0398] According to the LOD levels of corresponding sub-components (such as luminance, chrominance, etc.) in the same attribute, different improved quantization weight derivation processes can be applied to points. Therefore, when the attribute is color (YCbCr), the quantization weights of the sub-component Y (luminance) and the quantization weights of the sub-components Cb and Cr (chrominance) have different optimized values.
[0399] Figure 23 It is a flowchart illustrating point cloud encoding according to an embodiment.
[0400] Reference Figures 1 to 22 The point cloud transmission device or point cloud encoder described can perform attribute encoding to support the spatial adaptability decoding described in Reference Figure 21 According to an embodiment, the attribute encoding includes at least one of RAHT encoding, predictive transform encoding, or lifting transform encoding. The lifting transform encoding performed to support spatial adaptability decoding performs the improved quantization weight calculation process described in Reference Figures 21 to 22
[0401] As Figure 23 shown, the point cloud encoder receives an input of an attribute (23100). According to an embodiment, the attribute includes color and reflectance.
[0402] The point cloud encoder according to an embodiment performs attribute encoding (23200, 23210) according to the attribute type. As shown in the figure, the point cloud encoder independently performs encoding in the case where the attribute is color and in the case where the attribute is reflectance. The two attribute encoding operations can be performed simultaneously or sequentially. The attribute encoding performs a lifting transform based on LOD (for example, Figure 15 operation 1530). As described above, the point cloud processing device performs lifting transform encoding from the highest LOD level to the lowest LOD level.
[0403] As described above, point cloud encoding can support spatially adaptive decoding (23300, 23310). According to an embodiment, spatially adaptive decoding may perform decoding on some or all of the geometric structures and / or attributes according to the decoding performance of a point cloud receiving device (e.g., Figure 1 the receiving device 10004 of Figure 10 and Figure 11 the point cloud decoder of Figure 13 and the receiving device) to provide point cloud content of various resolutions. Through spatial adaptability, the point cloud receiving device (or decoder) can provide low-resolution point cloud content such as a thumbnail with low decoder complexity and / or small bandwidth. When spatially adaptive decoding is supported, the point cloud processing device transmits information for spatially adaptive decoding by referring to Figure 17 the signaling information (e.g., SPS, APS, attribute header, etc.) included in the bitstream described in Figure 17 the bitstream of
[0404] As described above, the signaling information (e.g., SPS, APS, attribute header, etc.) in the bitstream (e.g., Figures 21 to 23 the bitstream of Figures 21 to 22 may include adaptability information related to spatially adaptive decoding (or lifting transform) at the sequence level or slice level (adaptability information indicating whether decoding of attributes can be performed based on a spatial octree structure) (e.g., scalable_lifting_enabled_flag or lifting_scalability_enabled_flag). As described above, attribute decoding is performed based on the decoded geometric octree structure. The information related to spatially adaptive decoding (or lifting transform) indicates whether the entire octree structure or a partial octree structure is required for partial attribute decoding. Therefore, the point cloud receiving device obtains the signaling information of the bitstream and performs adaptable attribute decoding based on the entire octree structure or partial octree structure as the result of geometric structure decoding according to the information related to spatially adaptive decoding.
[0405] The point cloud encoder compresses color attributes and reflectance attributes (23500, 23510).
[0406] It can be executed by program instructions stored in one or more memories included in the point cloud sending device and the receiving device. Figure 23 The point cloud encoding shown in. According to an embodiment, the program instructions are executed by a point cloud encoder and / or decoder (or processor), and cause the point cloud encoder and / or decoder to perform point cloud encoding and / or decoding.
[0407] Figure 24 It is a flowchart illustrating a method of sending point cloud data according to an embodiment.
[0408] Figure 24 The flowchart 2400 of illustrates a method of sending point cloud data by referring to Figures 1 to 23 The point cloud data sending device described, for example, referring to Figure 1 , Figure 12 And Figure 14 The sending device or point cloud encoder described. The point cloud data sending device encodes the point cloud data including geometric structures and attributes (2410). The geometric structure is information indicating the positions of points in the point cloud data, and the attributes include at least one of the color and reflectivity of the points. The point cloud data sending device encodes the geometric structure. As described above with reference to Figures 1 to 23 The attribute encoding depends on the geometric structure encoding. Therefore, the point cloud data sending device encodes the attributes based on the entire or part of the octree structure of the encoded geometric structure, as described with reference to Figure 21 The attributes are encoded based on the quantization weights of the points included in each level of detail (LOD) among one or more levels of detail. The quantization weights are determined based on the number of points and the number of points belonging to the level l represented by the LOD. The details are the same as those described with reference to Figures 20 to 23 And the description thereof will be skipped.
[0409] The quantization weights according to an embodiment are represented as follows.
[0410]
[0411] Here, Represents the quantization weight of the LOD with level l, "total number of points" represents the number of points, "number of points in LOD l " represents the cumulative number of points belonging to the level l represented by the LOD, and "number of points in R i " represents the number of points belonging only to the LOD with level i. The value of the number of points in LOD l Is equal to the sum of the values of the number of points in R i , where the range of i is from 0 to l. The quantization weights according to an embodiment are the same as those with reference to Figures 21 to 22The processing for calculating the i-th optimized value (improved quantization weight) or the calculated quantization weight is the same, and thus the detailed description thereof will be skipped.
[0412] As described in the reference Figures 15 to 23 The point cloud data transmitting apparatus generates one or more LODs by reordering points, performs a lifting transform coding on attributes based on the one or more LODs, and quantizes the attributes after the lifting transform coding based on quantization weights. The quantization weights according to an embodiment are used to decode attributes encoded based on a geometric structure-based partial octree structure. That is, as described in the reference Figures 1 to 23 The point cloud transmitting apparatus supports spatially adaptive decoding. The spatially adaptive decoding performs all or part of the geometric structure and / or attributes according to the decoding performance of a point cloud receiving apparatus (e.g., Figure 1 the receiving apparatus 10004 of Figure 10 and Figure 11 the point cloud decoder of Figure 13 and the receiving apparatus) to provide decoding of point cloud content at various resolutions.
[0413] The point cloud data transmitting apparatus transmits a bitstream (e.g., the bitstream described in the reference Figure 17 ) including the encoded point cloud data (2420).
[0414] Therefore, the bitstream according to an embodiment (e.g., Figure 17 ) includes adaptability information (e.g., scalable_lifting_enabled_flag, lifting_scalability_enabled_flag) indicating whether attributes encoded based on a partial octree structure can be decoded. As described in the reference Figures 20 to 23 The information for spatially adaptive decoding is transmitted through signaling information (e.g., SPS, APS, attribute headers, etc.) included in the bitstream described in the reference Figure 17 As described above, the signaling information (e.g., SPS, APS, attribute headers, etc.) in the bitstream (e.g., the bitstream of Figure 17 ) may include signaling information indicating whether attributes decoded based on a spatial octree structure at a sequence level or a slice level can be decoded. The receiver can obtain such information and perform spatially adaptive decoding. Since the operation of the point cloud data transmitting apparatus is the same as the operation described in the reference Figures 1 to 23 the detailed description thereof will be skipped.
[0415] Figure 25 is a flowchart illustrating a method of processing point cloud data according to an embodiment.
[0416] Figure 25 The flowchart 2500 ofFigures 1 to 23 Method for a described point cloud data receiving device (e.g., receiving device 10004 or point cloud video decoder 10006) to process point cloud data.
[0417] A point cloud data receiving device (e.g., receiving device 10004, Figure 13 receiver, etc.) receives a bitstream (2510) containing point cloud data. The bitstream according to an embodiment contains signaling information (e.g., SPS, APS, attribute header, etc.) necessary for decoding the point cloud data. As referenced Figures 20 to 23 described, information for spatially adaptable decoding is sent through the signaling information (e.g., SPS, APS, attribute header, etc.) contained in the referenced Figure 17 described bitstream. As described above, the signaling information (e.g., SPS, APS, attribute header, etc.) in the bitstream may include information related to spatially adaptable decoding (or lifting transform) at the sequence level or slice level.
[0418] A point cloud data receiving device (e.g., Figure 10 decoder) decodes the point cloud data (2520) based on the signaling information. A point cloud data receiving device (e.g., Figure 10 geometric structure decoder) decodes the geometric structure included in the point cloud data. According to an embodiment, the geometric structure is information indicating the positions of the points of the point cloud data. A point cloud data receiving device (e.g., Figure 10 attribute decoder) decodes an attribute including at least one of the color and reflectance of the points based on the entire or partial octree structure of the decoded geometric structure. The attribute is decoded based on the quantization weights (or improved quantization weights) of the points included in each level of detail (LOD) among one or more LODs. The quantization weights are determined based on the number of points and the number of points belonging to level l represented by the LOD.
[0419] The signaling information necessary for decoding the point cloud data further includes adaptability information (e.g., scalable_lifting_enabled_flag or lifting_scalability_enabled_flag) indicating whether an attribute can be decoded based on a partial octree structure. When the adaptability information indicates that an attribute can be decoded based on a partial octree structure, quantization weights are determined for each point included in each LOD from the 0th LOD to the last LOD. The quantization weight calculation process is the same as the quantization weight calculation process described in the referenced Figures 21 to 23 description, so the detailed description thereof is skipped.
[0420] The point cloud data receiving device according to the embodiment generates one or more LODs by reordering points, performs lifting transform decoding on attributes based on the one or more LODs, and inverse quantizes the attributes after lifting transform decoding based on quantization weights. As described above with reference to Figure 22 The quantization weights are determined by executing program instructions stored in a memory included in the point cloud receiving device for calculating the quantization weights. The point cloud data processing operations of the point cloud data receiving device are the same as those described with reference to Figures 1 to 23 and thus a detailed description thereof will be skipped.
[0421] According to Figures 1 to 25 The elements of the point cloud data processing device according to the embodiment described in can be implemented by hardware, software, firmware, or a combination thereof including one or more processors combined with a memory. The elements in the embodiment can be implemented by a single chip, for example, a single hardware circuit. According to an embodiment, the components according to the embodiment can be implemented separately as individual chips. Additionally, at least one or more components of the device according to the embodiment can include one or more processors capable of executing one or more programs. The one or more programs can execute Figures 1 to 25 any one or more of the operations / methods of the point cloud data processing device described, or include instructions for performing the same operations / methods.
[0422] Although the drawings have been separately described for simplicity, new embodiments can be designed by combining the embodiments illustrated in the corresponding figures. The design of a computer-readable recording medium on which a program for executing the above embodiments is recorded, which is required by those skilled in the art, also falls within the scope of the appended claims and their equivalents. The device and method according to the embodiment can be not limited by the configurations and methods of the above embodiments. Various modifications can be made to the embodiments by selectively combining all or part of the embodiments. Although the preferred embodiments have been described with reference to the drawings, those skilled in the art will appreciate that various modifications and variations can be made to the embodiments without departing from the spirit or scope of the present disclosure described in the appended claims. Such modifications will not be understood independently of the technical concepts or viewpoints of the embodiments.
[0423] The descriptions of the device and method according to the embodiment can be applied to complement each other. For example, the point cloud data sending method according to the embodiment can be executed by the point cloud data sending device according to the embodiment or components included in the point cloud data sending device. Additionally, the method of receiving point cloud data according to the embodiment can be executed by the device for receiving point cloud data according to the embodiment or components included in the device for receiving point cloud data according to the embodiment.
[0424] The various elements of the apparatus according to the embodiments may be implemented by hardware, software, firmware, or a combination thereof. The various elements in the embodiments may be implemented by a single chip (e.g., a single hardware circuit). According to an embodiment, the components according to the embodiment may be implemented separately as separate chips. According to an embodiment, at least one or more components of the apparatus according to the embodiment may include one or more processors capable of executing one or more programs. The one or more programs may execute any one or more of the operations / methods according to the embodiment, or include instructions for executing them. The executable instructions for executing the method / operation of the apparatus according to the embodiment may be stored in a non-transitory CRM or other computer program products configured to be executed by one or more processors, or may be stored in a transitory CRM or other computer program products configured to be executed by one or more processors. Additionally, the memory according to the embodiment may be used to cover not only volatile memory (e.g., RAM) but also non-volatile memory, flash memory, and PROM. Additionally, it may also be implemented in the form of a carrier wave such as being transmitted via the Internet. Additionally, the processor-readable recording medium may be distributed among computer systems connected via a network such that the processor-readable code may be stored and executed in a distributed manner.
[0425] In this specification, the terms " / " and "," should be construed as indicating "and / or". For example, the expression "A / B" may mean "A and / or B". Additionally, "A, B" may mean "A and / or B". Additionally, "A / B / C" may mean "at least one of A, B, and / or C". Additionally, "A / B / C" may mean "at least one of A, B, and / or C". Additionally, in this specification, the term "or" should be construed as indicating "and / or". For example, the expression "A or B" may mean 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" used in this document should be construed as indicating "additionally or alternatively".
[0426] Terms such as first and second may be used to describe the various elements of the embodiments. However, the various components according to the embodiment should not be limited by the above terms. These terms are only used to distinguish one element from another. For example, the first user input signal may be referred to as the second user input signal. Similarly, the second user input signal may be referred to as the first user input signal. The use of these terms should be construed as not departing from the scope of the various embodiments. Both the first user input signal and the second user input signal are user input signals, but do not mean the same user input signal unless the context clearly indicates otherwise.
[0427] The terms used to describe the embodiments are used for the purpose of describing particular embodiments and are not intended to limit the embodiments. As used in the description of the embodiments and the claims, the singular forms "a", "an" and "the" include plural referents unless the context clearly dictates otherwise. The phrase "and / or" is used to include all possible combinations of terms. Terms such as "comprising" or "having" are intended to indicate the components of diagrams, numbers, steps, elements and / or parts, and should be understood as not precluding the possibility of the additional presence of diagrams, numbers, steps, elements and / or components. As used herein, conditional expressions such as "if" and "when" are not limited to alternative cases and are intended to be interpreted as performing the relevant operations or interpreting the relevant definitions according to the specific conditions when the specific conditions are met.
[0428] The mode of the present disclosure
[0429] As described above, the relevant content has been described in the best mode of implementing the embodiments.
[0430] Industrial applicability
[0431] Those skilled in the art will appreciate that various changes or modifications can be made to the embodiments within the scope of the embodiments. Therefore, the embodiments are intended to cover the modified forms and variations of the present disclosure, provided that they fall within the scope of the appended claims and their equivalents.
Claims
1. A method for sending point cloud data, the method comprising the following steps: Encode the point cloud data including geometric structures and attributes, where the geometric structures represent the positions of the points in the point cloud data, and the attributes include at least one of the color and reflectivity of the points. Wherein, the step of encoding the point cloud data includes: Encode the geometric structures; Encode the attributes based on a partial octree of the geometric structures by generating one or more levels of detail (LOD); Wherein, the attributes are encoded by applying quantization weights to the attributes; Wherein, the quantization weights are derived based on a portion of the number of points associated with the level of detail and the number of points in the level of detail; Transmit a bitstream including the encoded point cloud data.
2. The method according to claim 1, wherein, Encode the attributes based on a partial tree of the geometric structures; Wherein, the step of encoding the attributes includes: Generate one or more LOD by reorganizing the points; Perform lifting transform encoding on the attributes based on the one or more LOD; and Quantize the transformed attributes based on the weights.
3. The method according to claim 2, wherein, The weights are used to decode the encoded attributes based on the partial octree.
4. The method according to claim 3, wherein, The bitstream includes signaling information indicating whether to decode the encoded attributes based on the partial octree.
5. An apparatus for sending point cloud data, the apparatus comprising: An encoder configured to encode the point cloud data including geometric structures and attributes, wherein the geometric structures represent the positions of the points in the point cloud data, and the attributes include at least one of the color and reflectivity of the points. Wherein, the encoder includes: A geometric structure encoder configured to encode the geometric structures; and An attribute encoder configured to encode the attributes based on a partial octree of the geometric structures by generating one or more levels of detail (LOD); Wherein, the attributes are encoded by applying quantization weights to the attributes; Wherein, the quantization weights are derived based on a portion of the number of points associated with the level of detail and the number of points in the level of detail; and A transmitter configured to transmit a bitstream including the encoded point cloud data.
6. The apparatus according to claim 5, wherein, Encode the attributes based on a partial tree of the geometric structures; Wherein, the attribute encoder is configured to: Generate one or more LOD by reorganizing the points; Perform lifting transform encoding on the attributes based on the one or more LOD; Perform quantization processing on the transformed attributes.
7. A method for processing point cloud data, the method comprising the following steps: Receive a bitstream including the point cloud data, the bitstream including signaling information; And Decode the point cloud data based on the signaling information, Wherein, the step of decoding the point cloud data includes: Decode the geometric structures in the point cloud data, where the geometric structures represent the positions of the points in the point cloud data; and Decode the attributes in the point cloud data based on a partial octree of the geometric structures by generating one or more levels of detail (LOD), where the attributes include at least one of the color and reflectivity of the points. Among them, quantization weights are applied to the attributes, and the quantization weights are derived based on a part of the number of points related to the level of detail and the number of points in the level of detail.
8. The method according to claim 7, wherein, The signaling information includes adaptability information indicating whether lifting transform decoding of the attributes based on the partial octree is allowed.
9. The method according to claim 8, wherein, Decode the attributes based on the partial tree of the geometric structure, wherein the step of decoding the attributes includes: Generate one or more LODs by reorganizing the points; Perform lifting transform decoding of the attributes based on the one or more LODs; and Perform inverse quantization processing on the decoded attributes.
10. The method according to claim 9, wherein, Calculate the weights for each point in each LOD from the 0th LOD to the last LOD.
11. An apparatus for processing point cloud data, the apparatus comprising: A receiver for receiving a bitstream including the point cloud data, wherein the bitstream includes signaling information; and A decoder for decoding the point cloud data based on the signaling information, wherein the decoder includes: A geometric structure decoder for decoding the geometric structure included in the point cloud data, the geometric structure representing the positions of the points of the point cloud data; and An attribute decoder for decoding attributes including at least one of the color and reflectivity of the points based on a partial octree of the geometric structure by generating one or more levels of detail LODs, wherein the attributes are decoded based on weights related to one of the one or more levels of detail LODs, wherein quantization weights are applied to the attributes, and the quantization weights are derived based on a part of the number of points related to the level of detail and the number of points in the level of detail.
12. The apparatus according to claim 11, wherein, The signaling information includes adaptability information indicating whether lifting transform decoding of the attributes based on the partial octree is allowed, and the weights are calculated for each point in each LOD from the 0th LOD to the last LOD.
13. The apparatus according to claim 12, wherein, Decode the attributes based on the partial tree of the geometric structure, wherein the attribute decoder is configured to: Generate one or more LODs by reorganizing the points; Perform lifting transform decoding of the attributes based on the one or more LODs; and Perform inverse quantization processing on the decoded attributes.