Point cloud data transmission apparatus, point cloud data transmission method, point cloud data reception apparatus, and point cloud data reception method
Patent Information
- Application Number
- PCT/KR2026/004741
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-25
- Filing Date
- 2026-03-25
- Publication Date
- 2026-10-01
Smart Images

Figure KR2026004741_01102026_PF_FP_ABST
Abstract
Description
Point cloud data transmission device, point cloud data transmission method, point cloud data reception device and point cloud data reception method
[0001] The embodiments relate to a method and apparatus for processing point cloud content.
[0002] Point cloud content is content represented as a point cloud, which is a set of points belonging to a coordinate system that represents a three-dimensional space or volume. Point cloud content can represent three-dimensional media and is used to provide various services such as VR (Virtual Reality), AR (Augmented Reality), MR (Mixed Reality), XR (Extended Reality), and autonomous driving services. However, representing point cloud content requires tens of thousands to hundreds of thousands of point data points. Therefore, a method is required to efficiently process a vast amount of point data.
[0003] In other words, there is a problem in that a large amount of throughput is required to transmit and receive point cloud data. Therefore, encoding for compression and decoding for decompression are performed during the process of transmitting and receiving point cloud data; however, due to the large size of the point cloud data, the computations are complex and time-consuming.
[0004] The technical problem according to the embodiments is to provide an apparatus and method for efficiently transmitting and receiving point clouds in order to solve the aforementioned problems, etc.
[0005] The technical problem according to the embodiments is to provide an apparatus and method for solving latency and encoding / decoding complexity.
[0006] The technical problem according to the embodiments is to provide an apparatus and method for efficiently performing partial encoding and decoding.
[0007] However, the scope of rights of the embodiments is not limited to the technical problems described above, and may be extended to other technical problems that a person skilled in the art can infer based on the entire content described.
[0008] To achieve the above-described purpose and other advantages, the decoding method according to the embodiments may include the step of decoding geometry data of point cloud data in a bitstream and the step of decoding attribute data of the point cloud data.
[0009] According to embodiments, the geometry data is divided and included in subgroups of a layer group structure, at least one of the subgroups is a parent subgroup, the parent subgroup includes at least one child subgroup, the subgroup bounding box of each subgroup provides the boundary of the nodes included in the subgroup, and the at least one child subgroup may exist within the subgroup boundary of the parent subgroup.
[0010] According to embodiments, the step of decoding the geometry data can derive the minimum point location and the maximum point location of the subgroup bounding box of each child subgroup based on the minimum node size of the partial occupancy tree of the parent subgroup.
[0011] According to embodiments, the step of decoding the geometry data may select at least one subgroup based on the overlap between the bounding box of each subgroup and the region of interest, which is set based on the minimum point position and the maximum point position of the subgroup bounding box of each derived child subgroup, and decode the geometry data associated with the selected at least one subgroup.
[0012] According to embodiments, each child subgroup is identified by a pair of a layer group index and a subgroup index, and the bitstream further includes signaling information, and the signaling information includes origin information and size information of the subgroup bounding box of the subgroup identified by the pair of the layer group index and the subgroup index, and the origin information and size information of the subgroup bounding box can be expressed in units of the minimum node size of the partial occupancy tree of the parent subgroup.
[0013] According to embodiments, a decoding device includes a memory and at least one processor connected to the memory, and the at least one processor may be configured to decode geometry data of point cloud data within a bitstream and to decode attribute data of the point cloud data.
[0014] According to embodiments, the geometry data is divided and included in subgroups of a layer group structure, at least one of the subgroups is a parent subgroup, the parent subgroup includes at least one child subgroup, the subgroup bounding box of each subgroup provides the boundary of the nodes included in the subgroup, and the at least one child subgroup may exist within the subgroup boundary of the parent subgroup.
[0015] According to embodiments, the at least one processor can derive the minimum point location and the maximum point location of the subgroup bounding box of each child subgroup based on the minimum node size of the partial occupancy tree of the parent subgroup.
[0016] According to embodiments, the at least one processor can select at least one subgroup based on the overlap between the bounding box of each subgroup and the region of interest, which is set based on the minimum point position and the maximum point position of the subgroup bounding box of each derived child subgroup, and can decode geometry data associated with the selected at least one subgroup.
[0017] According to embodiments, each child subgroup is identified by a pair of a layer group index and a subgroup index, and the bitstream further includes signaling information, and the signaling information includes origin information and size information of the subgroup bounding box of the subgroup identified by the pair of the layer group index and the subgroup index, and the origin information and size information of the subgroup bounding box can be expressed in units of the minimum node size of the partial occupancy tree of the parent subgroup.
[0018] According to embodiments, the encoding method may include the step of encoding geometry data of point cloud data and the step of encoding attribute data of the point cloud data.
[0019] According to embodiments, the geometry data is divided and included in subgroups of a layer group structure, at least one of the subgroups is a parent subgroup, the parent subgroup includes at least one child subgroup, the subgroup bounding box of each subgroup provides the boundary of the nodes included in the subgroup, and the at least one child subgroup may exist within the subgroup boundary of the parent subgroup.
[0020] According to embodiments, each child subgroup is identified by a pair of a layer group index and a subgroup index, and a bitstream including the encoded geometry data and the encoded attribute data further includes signaling information, the signaling information includes origin information and size information of the subgroup bounding box of the subgroup identified by the pair of the layer group index and the subgroup index, and the origin information and size information of the subgroup bounding box can be expressed in units of the minimum node size of the partial occupancy tree of the parent subgroup.
[0021] According to embodiments, an encoding device may include a memory and at least one processor connected to the memory, and the at least one processor may be configured to encode geometry data of point cloud data and encode attribute data of point cloud data.
[0022] According to embodiments, a computer-readable storage medium can store a bitstream generated by the encoding method.
[0023] According to embodiments, the transmission method may include the step of obtaining a bitstream for point cloud data, the bitstream being generated based on the step of encoding geometry data of the point cloud data and the step of encoding attribute data of the point cloud data; and the step of transmitting data including the bitstream.
[0024] The device and method according to the embodiments can provide a high-quality point cloud service.
[0025] The device and method according to the embodiments can achieve various video codec methods.
[0026] The device and method according to the embodiments can provide general-purpose point cloud content, such as autonomous driving services.
[0027] The apparatus and method according to the embodiments can provide improved parallel processing and scalability by performing spatial adaptive partitioning of point cloud data for independent encoding and decoding of point cloud data.
[0028] The apparatus and method according to the embodiments can improve the encoding and decoding performance of a point cloud by dividing the point cloud data into tile and / or slice units to perform encoding and decoding, and by signaling the data necessary for this purpose.
[0029] The device and method according to the embodiments can divide and transmit compressed data according to certain criteria for point cloud data. In addition, when using layered coding, the compressed data can be divided and sent according to the layer. Therefore, the storage and transmission efficiency of the transmitting device can be increased.
[0030] The device and method according to the embodiments can efficiently manage a context buffer when decoding a bitstream in a layer group structure composed of a plurality of data units.
[0031] The apparatus and method according to the embodiments can efficiently manage context buffers during partial decoding by managing context buffers for each data unit on a layer group basis in a layer group structure composed of a plurality of data units.
[0032] The device and method according to the embodiments can efficiently manage the context buffer during partial decoding by managing a list internally within the decoder without additional signaling and managing the context buffer by determining whether decoding related to the region of interest is complete.
[0033] The apparatus and method according to the embodiments can reduce unnecessary decoding processing by selecting and decoding data units only when the ROI and the subgroup area overlap (loosely overlapped) or when actual points or nodes exist within the overlapping area (strictly overlapped).
[0034] The device and method according to the embodiments can minimize memory usage by releasing the context buffer of a subgroup early when there are no points or nodes within the subgroup or when there are no nodes / points in the overlapping area with the ROI.
[0035] The apparatus and method according to the embodiments can prevent a discrepancy between the number of subsequent subgroups signaled in the bitstream and the actual encoding structure and improve decoding stability by correcting the number of subsequent subgroups when a skipped FGS exists.
[0036] The apparatus and method according to the embodiments represent a subgroup bounding box by quantizing it based on the minimum node size of the parent subgroup, and restore it at the leaf node level of the occupancy tree during decoding in the decoder, so that subgroup bounding boxes belonging to different layer groups are represented with the same precision, and the hierarchical relationship between parent and child subgroups can be stably preserved.
[0037] The apparatus and method according to the embodiments derive the minimum point location and the maximum point location of the subgroup bounding box based on the minimum node size in the partial occupancy tree of the parent subgroup, thereby preventing child nodes generated from one parent node from being separated into different subgroups.
[0038] Drawings are included to further understand the embodiments, and the drawings illustrate the embodiments along with descriptions related to the embodiments.
[0039] FIG. 1 shows an example of a point cloud content provision system according to embodiments.
[0040] FIG. 2 is a block diagram illustrating a point cloud content provision operation according to embodiments.
[0041] FIG. 3 shows an example of a point cloud encoder according to embodiments.
[0042] FIG. 4 shows examples of octree and occupancy codes according to embodiments.
[0043] Figure 5 shows an example of a point configuration by LOD according to embodiments.
[0044] Figure 6 shows an example of a point configuration by LOD according to embodiments.
[0045] FIG. 7 shows an example of a point cloud decoder according to embodiments.
[0046] FIG. 8 is an example of a transmission device according to embodiments.
[0047] FIG. 9 is an example of a receiving device according to embodiments.
[0048] FIG. 10 shows an example of a structure that can be linked with a point cloud data transmission / reception method / device according to embodiments.
[0049] FIGS. 11(a) and FIGS. 11(b) show a single slice and divided slice-based geometry tree structure according to embodiments.
[0050] FIG. 12 shows the layer group structure of a geometry coding tree and the aligned layer group structure of an attribute coding tree according to embodiments.
[0051] FIG. 13 shows a layer group and subgroup structure according to embodiments.
[0052] FIG. 14 illustrates an example of context reference between groups according to embodiments.
[0053] FIG. 15 shows another example of context reference between groups according to embodiments.
[0054] FIG. 16 (a) to (c) illustrates an example of a context buffer management method according to embodiments.
[0055] FIG. 17 (a) to (c) shows other examples of context buffer management methods according to embodiments.
[0056] FIG. 18 (a) to (c) shows another example of a context buffer management method according to the embodiments.
[0057] FIG. 19 (a) to (c) shows another example of a context buffer management method according to the embodiments.
[0058] FIG. 20 (a) to (c) shows another example of a context buffer management method according to the embodiments.
[0059] FIG. 21 shows examples of the coding and / or decoding order of point cloud data according to embodiments.
[0060] FIG. 22 shows the structure of a bitstream containing point cloud data according to embodiments.
[0061] FIG. 23 shows an example of the syntax structure of a sequence parameter set (SPS) according to the embodiments.
[0062] FIG. 24 shows an example of the syntax structure of a dependent geometry data unit header according to embodiments.
[0063] FIG. 25 shows an example of the syntax structure of a dependent attribute data unit header according to embodiments.
[0064] FIGS. 26a and 26b show an example of the syntax structure of a layer group structure inventory (LGSI) according to the embodiments.
[0065] FIG. 27 shows a point cloud data transmission device / method according to embodiments.
[0066] FIG. 28 shows a point cloud data receiving device / method according to embodiments.
[0067] FIG. 29 illustrates a method for receiving point cloud data according to embodiments.
[0068] FIG. 30 illustrates a layer group-based point cloud data encoding method according to embodiments.
[0069] FIG. 31 illustrates a layer group-based point cloud data decoding method according to embodiments.
[0070] FIG. 32 (a) to (c) shows another example of a context buffer management method according to the embodiments.
[0071] FIG. 33 (a) to (c) shows another example of a context buffer management method according to the embodiments.
[0072] FIG. 34 (a) to (c) shows another example of a context buffer management method according to the embodiments.
[0073] FIG. 35 shows an example of the syntax structure of a geometry data unit header according to embodiments.
[0074] FIG. 36 shows another example of the syntax structure of a dependent geometry data unit header according to embodiments.
[0075] FIG. 37 shows an example of the syntax structure of an attribute data unit header according to embodiments.
[0076] FIG. 38 shows another example of the syntax structure of a dependent attribute data unit header according to embodiments.
[0077] FIG. 39 is a drawing showing an example of a layer group structure considering partial decoding according to embodiments.
[0078] FIG. 40 (a) to (d) shows another example of a context buffer management method according to the embodiments.
[0079] FIG. 41 shows another example of the syntax structure of a sequence parameter set according to the embodiments.
[0080] FIG. 42 shows another example of the syntax structure of a geometry data unit header according to embodiments.
[0081] FIG. 43 shows another example of the syntax structure of a dependent geometry data unit header according to embodiments.
[0082] FIG. 44 shows another example of the syntax structure of an attribute data unit header according to the embodiments.
[0083] FIG. 45 shows another example of the syntax structure of a dependent attribute data unit header according to the embodiments.
[0084] FIG. 46 (a) to (c) shows another example of a context buffer management method according to the embodiments.
[0085] FIG. 47 (a) to (c) shows another example of a context buffer management method according to embodiments.
[0086] FIG. 48 shows another example of the syntax structure of a sequence parameter set according to the embodiments.
[0087] FIG. 49 shows another example of the syntax structure of a geometry data unit header according to embodiments.
[0088] FIG. 50 shows another example of the syntax structure of a dependent geometry data unit header according to embodiments.
[0089] FIG. 51 (a) to (c) shows another example of a context buffer management method according to the embodiments.
[0090] FIGS. 52(a) and FIGS. 52(b) show another example of a context buffer management method according to embodiments.
[0091] FIGS. 53(a) and FIGS. 53(b) show another example of a context buffer management method according to embodiments.
[0092] FIGS. 54(a) and FIGS. 54(b) are drawings illustrating another example of a method for decoding point cloud data according to embodiments.
[0093] FIGS. 55a and FIGS. 55b are drawings showing examples of code implementations for performing context buffer management according to embodiments.
[0094] FIGS. 56a to 56c are drawings showing examples of code implementations for performing context buffer management of attributes according to embodiments.
[0095] FIGS. 57a and FIGS. 57b are drawings showing examples of code implementations for releasing a context state stored in a context buffer according to embodiments.
[0096] FIGS. 58a and FIGS. 58b are drawings showing examples of code implementation for releasing a parent node stored for decoding according to embodiments.
[0097] FIGS. 59a and 59b are drawings showing examples of code implementation for generating a list of sub-regions for an area overlapping with an ROI according to embodiments.
[0098] FIGS. 60(a) and FIGS. 60(b) are flowcharts showing examples of encoding operations according to embodiments.
[0099] FIGS. 61(a) and FIGS. 61(b) are drawings showing examples of ROI bounding box control methods according to embodiments.
[0100] FIGS. 62(a) and FIGS. 62(b) are drawings showing examples of ROI bounding box control methods according to embodiments.
[0101] FIGS. 63(a) and FIGS. 63(b) are drawings showing examples of ROI bounding box control methods according to embodiments.
[0102] FIGS. 64(a) and FIGS. 64(b) are drawings showing examples of ROI bounding box control methods according to embodiments.
[0103] FIG. 65 shows an example of the syntax structure of a dependent geometry data unit header according to embodiments.
[0104] FIG. 66 is a diagram showing an example of a node distribution by layer group according to embodiments.
[0105] FIGS. 67(a) and FIGS. 67(b) are drawings illustrating an example of subgroup unit node alignment according to embodiments.
[0106] FIGS. 68(a) and FIGS. 68(b) are drawings showing examples of subgroup boundary settings according to embodiments.
[0107] FIGS. 69a to 69d are tables showing the results of comparing coding performance and time complexity according to embodiments.
[0108] FIGS. 70a and FIGS. 70b are tables showing geometry bitstream reduction effects according to embodiments.
[0109] FIG. 71 is a diagram illustrating an example of compressing and servicing the geometry and attributes of point cloud data according to embodiments.
[0110] FIG. 72 is a diagram showing another example of compressing and servicing the geometry and attributes of point cloud data according to embodiments.
[0111] FIG. 73 is a diagram showing the operation of the transmitting and receiving end when transmitting point cloud data composed of layers according to the embodiments.
[0112] FIG. 74 shows a point cloud data transmission / reception device / method according to embodiments.
[0113] FIG. 75 shows a flowchart of a point cloud data encoding method according to embodiments.
[0114] FIG. 76 shows a flowchart of a point cloud data decoding method according to embodiments.
[0115] Preferred embodiments of the embodiments are described in detail, and examples thereof are shown in the accompanying drawings. The following detailed description, with reference to the accompanying drawings, is intended to describe preferred embodiments of the embodiments rather than merely embodiments that may be implemented according to the embodiments. The following detailed description includes details to provide a thorough understanding of the embodiments. However, it is obvious to those skilled in the art that the embodiments may be practiced without these details.
[0116] Most terms used in the embodiments are selected from those commonly used in the field, but some terms are chosen at the applicant's discretion, and their meanings are described in detail in the following description as necessary. Accordingly, the embodiments should be understood based on the intended meaning of the terms, rather than their mere names or meanings.
[0117] FIG. 1 shows an example of a point cloud content provision system according to embodiments.
[0118] The point cloud content providing system illustrated in FIG. 1 may include a transmission device (10000) and a reception device (10004). The transmission device (10000) and the reception device (10004) can communicate via wired or wireless means to transmit and receive point cloud data.
[0119] A transmission device (10000) according to embodiments can acquire, process, and transmit point cloud video (or point cloud content). According to embodiments, the transmission device (10000) may include a fixed station, a base transceiver system (BTS), a network, an Artificial Intelligence (AI) device and / or system, a robot, an AR / VR / XR device and / or server, etc. Additionally, according to embodiments, the transmission device (10000) may include a device that communicates with a base station and / or other wireless devices using wireless access technology (e.g., 5G NR (New RAT), LTE (Long Term Evolution)), a robot, a vehicle, an AR / VR / XR device, a mobile device, a home appliance, an Internet of Things (IoT) device, an AI device / server, etc.
[0120] A transmission device (10000) according to embodiments includes a point cloud video acquisition unit (10001), a point cloud video encoder (10002), and / or a transmitter (or communication module), 10003.
[0121] A point cloud video acquisition unit (10001) according to the embodiments acquires a point cloud video through processing steps such as capture, synthesis, or generation. The point cloud video is a point cloud content represented as a point cloud, which is a set of points located in a three-dimensional space, and may be referred to as point cloud video data, etc. The point cloud video according to the embodiments may include one or more frames. A frame represents a still image / picture. Accordingly, the point cloud video may include a point cloud image / frame / picture and may be referred to as any one of a point cloud image, a frame, and a picture.
[0122] A point cloud video encoder (10002) according to the embodiments encodes the obtained point cloud video data. The point cloud video encoder (10002) can encode the point cloud video data based on point cloud compression coding. The point cloud compression coding according to the embodiments may include Geometry-based Point Cloud Compression (G-PCC) coding and / or Video-based Point Cloud Compression (V-PCC) coding or next-generation coding. Furthermore, the point cloud compression coding according to the embodiments is not limited to the embodiments described above. The point cloud video encoder (10002) can output a bitstream containing the encoded point cloud video data. The bitstream may include not only the encoded point cloud video data but also signaling information related to the encoding of the point cloud video data.
[0123] A transmitter (10003) according to the embodiments transmits a bitstream containing encoded point cloud video data. The bitstream according to the embodiments is encapsulated into a file or segment (e.g., a streaming segment) and transmitted through various networks such as a broadcast network and / or a broadband network. Although not illustrated in the drawings, the transmission device (10000) may include an encapsulation unit (or encapsulation module) that performs an encapsulation operation. Additionally, according to the embodiments, the encapsulation unit may be included in the transmitter (10003). According to the embodiments, the file or segment may be transmitted to a receiving device (10004) via a network or stored on a digital storage medium (e.g., USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.). The transmitter (10003) according to the embodiments can communicate wired or wirelessly with the receiving device (10004) (or receiver (10005)) via a network such as 4G, 5G, or 6G. Additionally, the transmitter (10003) can perform necessary data processing operations according to a network system (e.g., a communication network system such as 4G, 5G, 6G, etc.). Additionally, the transmission device (10000) can transmit encapsulated data according to an on-demand method.
[0124] A receiving device (10004) according to embodiments includes a receiver (10005), a point cloud video decoder (10006), and / or a renderer (10007). According to embodiments, the receiving device (10004) may include a device, robot, vehicle, AR / VR / XR device, mobile device, home appliance, IoT (Internet of Thing) device, AI device / server, etc., that communicates with a base station and / or other wireless device using wireless access technology (e.g., 5G NR (New RAT), LTE (Long Term Evolution)).
[0125] A receiver (10005) according to the embodiments receives a bitstream containing point cloud video data or a file / segment containing the bitstream from a network or a storage medium. The receiver (10005) can perform necessary data processing operations according to a network system (e.g., a communication network system such as 4G, 5G, 6G, etc.). The receiver (10005) according to the embodiments can output a bitstream by decapsulating the received file / segment. Additionally, according to the embodiments, the receiver (10005) may include a decapsulation unit (or decapsulation module) for performing a decapsulation operation. Additionally, the decapsulation unit may be implemented as an element (or component) separate from the receiver (10005).
[0126] A point cloud video decoder (10006) decodes a bitstream containing point cloud video data. The point cloud video decoder (10006) can decode the point cloud video data according to the way the point cloud video data is encoded (e.g., the reverse process of the operation of a point cloud video encoder (10002)). Accordingly, the point cloud video decoder (10006) can decode the point cloud video data by performing point cloud decompression coding, which is the reverse process of point cloud compression. Point cloud decompression coding includes G-PCC coding.
[0127] The renderer (10007) renders the decoded point cloud video data. In one embodiment, the renderer (10007) can render the decoded point cloud video data according to a viewport, etc. The renderer (10007) can render not only the point cloud video data but also audio data to output point cloud content. According to embodiments, the renderer (10007) may include a display for displaying the point cloud content. According to embodiments, the display may not be included in the renderer (10007) but may be implemented as a separate device or component.
[0128] The arrows indicated by dotted lines in the drawing represent the transmission path of feedback information obtained from the receiving device (10004). The feedback information is information intended to reflect interaction with a user consuming point cloud content and includes user information (e.g., head orientation information, viewport information, etc.). In particular, if the point cloud content is content for a service requiring interaction with a user (e.g., autonomous driving service, etc.), the feedback information may be transmitted to the content transmitting side (e.g., the transmitting device (10000)) and / or the service provider. Depending on the embodiments, the feedback information may be used in the receiving device (10004) as well as the transmitting device (10000), or it may not be provided.
[0129] Head orientation information according to the embodiments may refer to information regarding the user's head position, direction, angle, movement, etc. The receiving device (10004) according to the embodiments may calculate viewport information based on head orientation information. Viewport information is information about the area of the point cloud video that the user is looking at (i.e., the area the user is currently looking at). That is, viewport information is information about the area the user is currently looking at within the point cloud video. In other words, the viewport or viewport area may refer to the area the user is looking at in the point cloud video. And the viewpoint is the point the user is looking at in the point cloud video, and may refer to the exact center point of the viewport area. That is, the viewport is an area centered on the viewpoint, and the size and shape occupied by that area may be determined by the FOV (Field Of View). Therefore, the receiving device (10004) may extract viewport information based on the vertical or horizontal FOV supported by the device in addition to head orientation information. Additionally, the receiving device (10004) can perform gaze analysis, etc. based on head orientation information and / or viewport information to determine the user's point cloud video consumption method, the point cloud video area the user gazes at, the gaze time, etc. According to embodiments, the receiving device (10004) can transmit feedback information including the gaze analysis results to the transmitting device (10000). According to embodiments, a device such as a VR / XR / AR / MR display can extract a viewport area based on the user's head position / direction, the vertical or horizontal FOV supported by the device, etc. According to embodiments, the head orientation information and viewport information may be referred to as feedback information, signaling information, or metadata.
[0130] Feedback information according to the embodiments may be obtained during the rendering and / or display process. Feedback information according to the embodiments may be obtained by one or more sensors included in the receiving device (10004). Additionally, according to the embodiments, feedback information may be obtained by the renderer (10007) or a separate external element (or device, component, etc.). The dotted line in FIG. 1 indicates the process of transmitting feedback information obtained from the renderer (10007). The feedback information may be transmitted to the transmitting side, as well as consumed at the receiving side. That is, the point cloud content providing system may process point cloud data (encoding / decoding / rendering) based on the feedback information. For example, the point cloud video decoder (10006) and the renderer (10007) may use the feedback information, namely head orientation information and / or viewport information, to preferentially decode and render only the point cloud video for the area currently being viewed by the user.
[0131] Additionally, the receiving device (10004) can transmit feedback information to the transmitting device (10000). The transmitting device (10000) (or the point cloud video encoder (10002)) can perform an encoding operation based on the feedback information. Thus, the point cloud content providing system can efficiently process necessary data (e.g., point cloud data corresponding to the user's head position) based on the feedback information without processing (encoding / decoding) all point cloud data, and provide point cloud content to the user.
[0132] According to embodiments, the transmission device (10000) may be referred to as an encoder, transmission device, transmitter, transmission system, etc., and the receiving device (10004) may be referred to as a decoder, receiving device, receiver, receiving system, etc.
[0133] Point cloud data processed in the point cloud content providing system of FIG. 1 according to embodiments (processed through a series of processes of acquisition / encoding / transmission / decoding / rendering) may be referred to as point cloud content data or point cloud video data. According to embodiments, point cloud content data may be used as a concept including metadata or signaling information related to point cloud data.
[0134] The elements of the point cloud content delivery system illustrated in FIG. 1 can be implemented using hardware, software, processors, and / or combinations thereof.
[0135] FIG. 2 is a block diagram illustrating a point cloud content provision operation according to embodiments.
[0136] The block diagram of FIG. 2 illustrates the operation of the point cloud content provision system described in FIG. 1. As described above, the point cloud content provision system can process point cloud data based on point cloud compression coding (e.g., G-PCC).
[0137] A point cloud content providing system according to the embodiments (e.g., a point cloud transmission device (10000) or a point cloud video acquisition unit (10001)) can acquire a point cloud video (20000). The point cloud video is represented as a point cloud belonging to a coordinate system representing a three-dimensional space. The point cloud video according to the embodiments may include a Ply (Polygon File format or the Stanford Triangle format) file. If the point cloud video has one or more frames, the acquired point cloud video may include one or more Ply files. The Ply file contains point cloud data such as the geometry and / or attributes of the points. The geometry includes the positions of the points. The position of each point may be represented by parameters (e.g., values of the X-axis, Y-axis, and Z-axis, respectively) representing a three-dimensional coordinate system (e.g., a coordinate system consisting of XYZ axes). Attributes include attributes of points (e.g., texture information, color (YCbCr or RGB), reflectance (r), transparency, etc.) of each point. A point has one or more attributes (or properties). For example, a point may have one attribute which is color, or two attributes which are color and reflectance. According to embodiments, geometry may be referred to as positions, geometry information, geometry data, etc., and attributes may be referred to as attributes, attribute information, attribute data, etc. Additionally, a point cloud content providing system (e.g., a point cloud transmission device (10000) or a point cloud video acquisition unit (10001)) may obtain point cloud data from information related to the acquisition process of point cloud video (e.g., depth information, color information, etc.).
[0138] A point cloud content providing system (e.g., a transmission device (10000) or a point cloud video encoder (10002)) according to embodiments can encode point cloud data (20001). The point cloud content providing system can encode point cloud data based on point cloud compression coding. As described above, point cloud data may include geometry and attributes of points. Accordingly, the point cloud content providing system can output a geometry bitstream by performing geometry encoding to encode geometry. The point cloud content providing system can output an attribute bitstream by performing attribute encoding to encode attributes. According to embodiments, the point cloud content providing system can perform attribute encoding based on geometry encoding. The geometry bitstream and attribute bitstream according to embodiments can be multiplexed and output as a single bitstream. The bitstream according to the embodiments may further include signaling information related to geometry encoding and attribute encoding.
[0139] A point cloud content providing system according to embodiments (e.g., a transmission device (10000) or a transmitter (10003)) can transmit encoded point cloud data (20002). As described in FIG. 1, the encoded point cloud data can be represented as a geometry bitstream and an attribute bitstream. Additionally, the encoded point cloud data can be transmitted in the form of a bitstream along with signaling information related to the encoding of the point cloud data (e.g., signaling information related to geometry encoding and attribute encoding). Additionally, the point cloud content providing system can encapsulate the bitstream transmitting the encoded point cloud data and transmit it in the form of a file or segment.
[0140] A point cloud content providing system according to embodiments (e.g., a receiving device (10004) or a receiver (10005)) can receive a bitstream containing encoded point cloud data. Additionally, the point cloud content providing system (e.g., a receiving device (10004) or a receiver (10005)) can demultiplex the bitstream.
[0141] A point cloud content providing system (e.g., a receiving device (10004) or a point cloud video decoder (10005)) can decode encoded point cloud data (e.g., a geometry bitstream, an attribute bitstream) transmitted as a bitstream. A point cloud content providing system (e.g., a receiving device (10004) or a point cloud video decoder (10005)) can decode point cloud video data based on signaling information related to the encoding of point cloud video data included in the bitstream. A point cloud content providing system (e.g., a receiving device (10004) or a point cloud video decoder (10005)) can decode the geometry bitstream to restore the positions (geometry) of the points. A point cloud content providing system can decode the attribute bitstream based on the restored geometry to restore the attributes of the points. A point cloud content delivery system (e.g., a receiving device (10004) or a point cloud video decoder (10005)) can restore a point cloud video based on positions according to the restored geometry and decoded attributes.
[0142] A point cloud content providing system according to embodiments (e.g., a receiving device (10004) or a renderer (10007)) can render decoded point cloud data (20004). The point cloud content providing system (e.g., a receiving device (10004) or a renderer (10007)) can render geometry and attributes decoded through a decoding process according to various rendering methods. Points of the point cloud content may be rendered as vertices having a certain thickness, cubes having a specific minimum size with the vertex location as the center, or circles with the vertex location as the center, etc. All or part of the rendered point cloud content is provided to a user through a display (e.g., a VR / AR display, a general display, etc.).
[0143] A point cloud content providing system (e.g., a receiving device (10004)) according to the embodiments can obtain feedback information (20005). The point cloud content providing system can encode and / or decode point cloud data based on the feedback information. Since the feedback information and the operation of the point cloud content providing system according to the embodiments are the same as the feedback information and operation described in FIG. 1, a detailed description is omitted.
[0144] FIG. 3 shows an example of a point cloud encoder according to embodiments.
[0145] FIG. 3 shows an example of the point cloud video encoder (10002) of FIG. 1. The point cloud encoder reconstructs point cloud data (e.g., positions and / or attributes of points) and performs encoding operations to adjust the quality of point cloud content (e.g., lossless, lossy, near-lossless) according to network conditions or applications. If the total size of the point cloud content is large (e.g., point cloud content of 60 Gbps in the case of 30 fps), the point cloud content delivery system may not be able to stream the content in real time. Therefore, the point cloud content delivery system may reconstruct the point cloud content based on a maximum target bitrate to provide it according to the network environment.
[0146] As described in FIGS. 1 and 2, the point cloud encoder can perform geometry encoding and attribute encoding. Geometry encoding is performed before attribute encoding.
[0147] The point cloud encoder according to the embodiments comprises a coordinate system transformation unit (Transformation Coordinates, 30000), a quantization unit (Quantize and Remove Points (Voxelize), 30001), an octree analysis unit (Analyze Octree, 30002), a surface approximation analysis unit (Analyze Surface Approximation, 30003), an arithmetic encoder (Arithmetic Encode, 30004), a geometry reconstruction unit (Reconstruct Geometry, 30005), a color transformation unit (Transform Colors, 30006), an attribute transformation unit (Transfer Attributes, 30007), a RAHT transformation unit (30008), an LOD generation unit (Generated LOD, 30009), a lifting transformation unit (Lifting) (30010), and a coefficient quantization unit (Quantize Coefficients, 30011). It includes an and / or arithmetic encoder (30012). In the point cloud encoder of FIG. 3, the coordinate system transformation unit (30000), quantization unit (30001), octree analysis unit (30002), surface approximation analysis unit (30003), arithmetic encoder (30004), and geometry reconstruction unit (30005) can be grouped and referred to as a geometry encoder. Also, the color conversion unit (30006), attribute conversion unit (30007), RAHT conversion unit (30008), LOD generation unit (30009), lifting conversion unit (30010), coefficient quantization unit (30011), and / or arithmetic encoder (30012) can be grouped and referred to as an attribute encoder.
[0148] The coordinate system transformation unit (30000), quantization unit (30001), octree analysis unit (30002), surface approximation analysis unit (30003), arismetic encoder (30004), and geometry reconstruction unit (30005) can perform geometry encoding. Geometry encoding according to the embodiments may include octree geometry coding, direct coding, trisoup geometry encoding, and entropy encoding. Direct coding and trisoup geometry encoding are applied optionally or in combination. Additionally, geometry encoding is not limited to the above examples.
[0149] As illustrated in the drawings, the coordinate system conversion unit (30000) according to the embodiments receives positions and converts them into a coordinate system. For example, the positions can be converted into position information in a three-dimensional space (e.g., a three-dimensional space expressed in an XYZ coordinate system). The position information in the three-dimensional space according to the embodiments may be referred to as geometry information.
[0150] The quantization unit (30001) according to the embodiments quantizes the geometry. For example, the quantization unit (30001) can quantize points based on the minimum position values of all points (e.g., minimum values on each axis for the X-axis, Y-axis, and Z-axis). The quantization unit (30001) performs a quantization operation to find the nearest integer value by multiplying the difference between the minimum position value and the position value of each point by a preset quantization scale value and then performing rounding down or rounding up. Thus, one or more points may have the same quantized position (or position value). The quantization unit (30001) according to the embodiments performs voxelization based on the quantized positions to reconstruct the quantized points. Just as the minimum unit containing 2D image / video information is a pixel, the points of the point cloud content (or 3D point cloud video) according to the embodiments may be contained in one or more voxels. A voxel is a combination of volume and pixel, and refers to a three-dimensional cubic space that is generated when a three-dimensional space is divided into units (unit=1.0) based on axes representing the three-dimensional space (e.g., X-axis, Y-axis, Z-axis). The quantization unit (40001) can match groups of points in the three-dimensional space to voxels. According to embodiments, a single voxel may contain only one point. According to embodiments, a single voxel may contain one or more points. In addition, to represent a single voxel as a single point, the position of the center of the voxel can be set based on the positions of one or more points included in the voxel. In this case, the attributes of all positions included in the voxel can be combined and assigned to the voxel.
[0151] The octree analysis unit (30002) according to the embodiments performs octree geometry coding (or octree coding) to represent the voxels in an octree structure. The octree structure represents points matched to the voxels based on an octree structure.
[0152] The surface approximation analysis unit (30003) according to the embodiments can analyze and approximate an octree. The octree analysis and approximation according to the embodiments is a process of analyzing to voxelize an area containing multiple points in order to efficiently provide octree and voxelization.
[0153] An arithmetic encoder (30004) according to the embodiments entropy-encodes an octree and / or an approximated octree. For example, the encoding method includes an arithmetic encoding method. As a result of the encoding, a geometry bitstream is generated.
[0154] The color conversion unit (30006), attribute conversion unit (30007), RAHT conversion unit (30008), LOD generation unit (30009), lifting conversion unit (30010), coefficient quantization unit (30011) and / or arismetic encoder (30012) perform attribute encoding. As described above, a point may have one or more attributes. The attribute encoding according to the embodiments is applied equally to the attributes of a point. However, if a single attribute (e.g., color) includes one or more elements, independent attribute encoding is applied to each element. The attribute encoding according to the embodiments may include color conversion coding, attribute conversion coding, Region Adaptive Hierarchial Transform (RAHT) coding, prediction transformation (Interpolaration-based hierarchical nearest-neighbour prediction-Prediction Transform) coding, and lifting transformation (interpolation-based hierarchical nearest-neighbour prediction with an update / lifting step (Lifting Transform)) coding. Depending on the point cloud content, the above-described RAHT coding, prediction transformation coding, and lifting transformation coding may be used optionally, or a combination of one or more of the codings may be used. Furthermore, the attribute encoding according to the embodiments is not limited to the examples described above.
[0155] The color conversion unit (30006) according to the embodiments performs color conversion coding that converts color values (or textures) included in attributes. For example, the color conversion unit (30006) can convert the format of color information (e.g., convert from RGB to YCbCr). The operation of the color conversion unit (30006) according to the embodiments may be applied optionally depending on the color values included in attributes.
[0156] The geometry reconstruction unit (30005) according to the embodiments reconstructs (decompresses) an octree and / or an approximated octree. The geometry reconstruction unit (30005) reconstructs an octree / voxel based on the results of analyzing the distribution of points. The reconstructed octree / voxel may be referred to as the reconstructed geometry (or restored geometry).
[0157] The attribute transformation unit (30007) according to the embodiments performs attribute transformation that transforms attributes based on positions where geometry encoding has not been performed and / or reconstructed geometry. As described above, since attributes are dependent on geometry, the attribute transformation unit (30007) can transform attributes based on reconstructed geometry information. For example, the attribute transformation unit (30007) can transform the attributes of a point at a position based on the position value of a point included in a voxel. As described above, when the position of the center point of a voxel is set based on the positions of one or more points included in a voxel, the attribute transformation unit (30007) transforms the attributes of one or more points. When trisoop geometry encoding is performed, the attribute conversion unit (30007) can convert attributes based on the trisoop geometry encoding.
[0158] The attribute transformation unit (30007) can perform attribute transformation by calculating the average value of attributes or attribute values (e.g., the color or reflectance of each point) of neighboring points within a specific location / radius from the position (or position value) of the center point of each voxel. The attribute transformation unit (30007) can apply a weight based on the distance from the center point to each point when calculating the average value. Thus, each voxel has a position and a calculated attribute (or attribute value).
[0159] The attribute conversion unit (30007) can search for neighbor points within a specific location / radius from the position of the center point of each voxel based on a KD tree or a Morton code. A KD tree is a binary search tree that supports a data structure capable of managing points based on their positions to enable rapid Nearest Neighbor Search (NNS). A Morton code is generated by representing the coordinate values (e.g., (x, y, z)) representing the 3D positions of all points as bit values and mixing the bits. For example, if the coordinate values representing the position of a point are (5, 9, 1), the bit values of the coordinate values are (0101, 1001, 0001). When the bit values are mixed according to the bit indices in the order of z, y, and x, it becomes 010001000111. When this value is represented in decimal, it becomes 1095. That is, the Molton code value of the point with coordinates (5, 9, 1) is 1095. The attribute transformation unit (30007) sorts the points based on the Molton code value and can perform shortest neighbor search (NNS) through a depth-first traversal process. After the attribute transformation operation, if shortest neighbor search (NNS) is required in other transformation processes for attribute coding, a KD tree or Molton code is utilized.
[0160] As shown in the drawing, the converted attributes are input to the RAHT conversion unit (30008) and / or LOD generation unit (30009).
[0161] The RAHT transformation unit (30008) according to the embodiments performs RAHT coding to predict attribute information based on reconstructed geometry information. For example, the RAHT transformation unit (30008) can predict attribute information of a node at an upper level of the octree based on attribute information associated with a node at a lower level of the octree.
[0162] The LOD generation unit (30009) according to the embodiments generates a Level of Detail (LOD) to perform predictive transformation coding. The LOD according to the embodiments represents the degree of detail of the point cloud content, and indicates that the smaller the LOD value, the lower the detail of the point cloud content, and the larger the LOD value, the higher the detail of the point cloud content. Points can be classified according to the LOD.
[0163] The lifting transformation unit (30010) according to the embodiments performs lifting transformation coding that transforms the attributes of the point cloud based on weights. As described above, the lifting transformation coding may be applied optionally.
[0164] The coefficient quantization unit (30011) according to the embodiments quantizes attribute-coded attributes based on coefficients.
[0165] The arismetic encoder (30012) according to the embodiments encodes quantized attributes based on arismetic coding.
[0166] The elements of the point cloud encoder of FIG. 3 may be implemented in hardware, software, firmware, or a combination thereof, comprising one or more processors or integrated circuits configured to communicate with one or more memories included in the point cloud providing device, although not illustrated in the drawing. One or more processors may perform at least one of the operations and / or functions of the elements of the point cloud encoder of FIG. 3 described above. Additionally, one or more processors may operate or execute a set of software programs and / or instructions for performing the operations and / or functions of the elements of the point cloud encoder of FIG. 3. One or more memories according to the embodiments may include high-speed random access memory and may include non-volatile memory (e.g., one or more magnetic disk storage devices, flash memory devices, or other non-volatile solid-state memory devices).
[0167] FIG. 4 shows examples of octree and occupancy codes according to embodiments.
[0168] As described in FIGS. 1 to 3, a point cloud content providing system (point cloud video encoder (10002)) or a point cloud encoder (e.g., an octree analysis unit (30002)) performs octree geometry coding based on an octree structure (or octree coding) to efficiently manage the area and / or position of a voxel.
[0169] The top of FIG. 4 shows an octree structure. The three-dimensional space of the point cloud content according to the embodiments is represented by the axes of the coordinate system (e.g., X-axis, Y-axis, Z-axis). The octree structure has two poles (0,0,0) and (2 d , 2 d , 2 d It is generated by recursively subdividing the bounding box (cubical axis-aligned bounding box) defined by ). 2d can be set to the value that constitutes the smallest bounding box enclosing all points of the point cloud content (or point cloud video). d represents the depth of the octree. The value of d is determined according to the following equation. In the equation below, (x int n , y int n , z int n ) represents the positions (or position values) of quantized points.
[0170]
[0171] As illustrated in the middle of the top of Fig. 4, the entire three-dimensional space can be divided into eight spaces according to the division. Each divided space is represented as a cube having six faces. As illustrated in the right of the top of Fig. 4, each of the eight spaces is further divided based on the axes of the coordinate system (e.g., X-axis, Y-axis, Z-axis). Thus, each space is again divided into eight smaller spaces. The divided smaller spaces are also represented as cubes having six faces. This division method is applied until the leaf nodes of the octree become voxels.
[0172] The bottom of Fig. 4 shows the occupancy code of an octree. The occupancy code of an octree is generated to indicate whether each of the eight partitioned spaces resulting from the partitioning of a single space contains at least one point. Therefore, one occupancy code is represented by eight child nodes. Each child node represents the occupancy of the partitioned space, and the child node has a value of 1 bit. Thus, the occupancy code is represented as an 8-bit code. That is, if the space corresponding to the child node contains at least one point, the node has a value of 1. If the space corresponding to the child node does not contain a point (empty), the node has a value of 0. Since the occupancy code shown in Fig. 4 is 00100001, it indicates that the spaces corresponding to the 3rd and 8th child nodes among the eight child nodes each contain at least one point. As illustrated in the drawing, the 3rd child node and the 8th child node each have 8 child nodes, and each child node is represented by an 8-bit Occupancy code. The drawing indicates that the Occupancy code of the 3rd child node is 10000111 and the Occupancy code of the 8th child node is 01001111. A point cloud encoder according to the embodiments (e.g., an arismetic encoder (30004)) can entropy-encode the Occupancy code. Additionally, to increase compression efficiency, the point cloud encoder can intra- / inter-encode the Occupancy code. A receiving device according to the embodiments (e.g., a receiving device (10004) or a point cloud video decoder (10006)) reconstructs the octree based on the Occupancy code.
[0173] A point cloud encoder according to the embodiments (e.g., the point cloud encoder of FIG. 3, or the octree analysis unit (30002)) can perform voxelization and octree coding to store the positions of the points. However, since points in a three-dimensional space are not always evenly distributed, there may be specific areas where few points exist. Therefore, performing voxelization on the entire three-dimensional space is inefficient. For example, if there are almost no points in a specific area, there is no need to perform voxelization up to that area.
[0174] Accordingly, the point cloud encoder according to the embodiments can perform direct coding, which directly codes the positions of points included in a specific region (or nodes excluding leaf nodes of an octree) without performing voxelization on the aforementioned specific region. The coordinates of the points directly coded according to the embodiments are referred to as the Direct Coding Mode (DCM). Additionally, the point cloud encoder according to the embodiments can perform trisoup geometry encoding, which reconstructs the positions of points within a specific region (or node) based on voxels using a surface model. Trisoup geometry encoding is a geometry encoding that represents an object as a series of triangle meshes. Therefore, the point cloud decoder can generate a point cloud from the mesh surface. Direct coding and trisoup geometry encoding according to the embodiments may be performed optionally. In addition, direct coding and trisoop geometry encoding according to the embodiments can be performed in combination with octree geometry coding (or octree coding).
[0175] To perform direct coding, the option to use direct mode for applying direct coding must be enabled, the node to which direct coding is to be applied must not be a leaf node, and there must be points within a specific node that are below a threshold. In addition, the total number of points subject to direct coding must not exceed a preset threshold. If the above conditions are satisfied, the point cloud encoder (or arismetic encoder (30004)) according to the embodiments can entropy-code the positions (or position values) of the points.
[0176] A point cloud encoder according to the embodiments (e.g., a surface approximation analysis unit (30003)) can determine a specific level of an octree (where the level is smaller than the depth d of the octree) and, starting from that level, perform trisoop geometry encoding to reconstruct the position of points within a node region based on voxels using a surface model (trisoop mode). The point cloud encoder according to the embodiments can specify the level to which trisoop geometry encoding is applied. For example, if the specified level is equal to the depth of the octree, the point cloud encoder does not operate in trisoop mode. That is, the point cloud encoder according to the embodiments can operate in trisoop mode only when the specified level is smaller than the depth value of the octree. A three-dimensional cubic region of nodes at a specified level according to the embodiments is referred to as a block. A block may contain one or more voxels. A block or a voxel may correspond to a brick. Within each block, geometry is represented as a surface. A surface according to the embodiments may intersect each edge of the block at most once.
[0177] Since one block has 12 edges, there are at least 12 intersection points within one block. Each intersection point is referred to as a vertex. A vertex along an edge is detected if there is at least one occupied voxel adjacent to that edge among all blocks sharing that edge. An occupied voxel according to the embodiments means a voxel containing a point. The position of a vertex detected along an edge is the average position along the edge of all voxels adjacent to that edge among all blocks sharing that edge.
[0178] When a vertex is detected, the point cloud encoder according to the embodiments includes the starting point (x, y, z) of the edge and the direction vector of the edge ( x, y, z), vertex position values (relative position values within the edge) can be entropy-coded. When trisoop geometry encoding is applied, a point cloud encoder according to the embodiments (e.g., geometry reconstruction unit (30005)) can generate restored geometry (reconstructed geometry) by performing triangle reconstruction, up-sampling, and voxelization processes.
[0179] The vertices located on the edges of the block determine the surface passing through the block. The surface according to the embodiments is a non-planar polygon. The triangle reconstruction process reconstructs the surface represented by triangles based on the edge start point, the edge direction vector, and the vertex position value. The triangle reconstruction process is as follows: ① calculate the centroid value of each vertex, ② subtract the centroid value from each vertex value, ③ square the result, and add all the result together.
[0180]
[0181] Then, the minimum value of the sum is calculated, and a projection process is performed along the axis where the minimum value is located. For example, if the x-element is at its minimum, each vertex is projected along the x-axis relative to the center of the block and onto the (y, z) plane. If the value obtained from projecting onto the (y, z) plane is (ai, bi), the θ value is calculated using atan2(bi, ai), and the vertices are aligned based on the θ value. The table below shows the combinations of vertices to generate triangles depending on the number of vertices. The vertices are sorted in order from 1 to n. Table 1 below indicates that for four vertices, two triangles can be formed depending on the combination of vertices. The first triangle is composed of the 1st, 2nd, and 3rd vertices among the aligned vertices, and the second triangle can be composed of the 3rd, 4th, and 1st vertices among the aligned vertices.
[0182] [Table 1] Triangles formed from vertices ordered 1,… , nnTriangles3(1,2,3)4(1,2,3), (3,4,1)5(1,2,3), (3,4,5), (5,1,3)6(1,2,3), (3,4,5), (5,6,1), (1,3,5)7(1,2,3), (3,4,5), (5,6,7), (7,1,3), (3,5,7)8(1,2,3), (3,4,5), (5,6,7), (7,8,1), (1,3,5), (5,7,1)9(1,2,3), (3,4,5), (5,6,7), (7,8,9), (9,1,3), (3,5,7), (7,9,3)10(1,2,3), (3,4,5), (5,6,7), (7,8,9), (9,10,1), (1,3,5), (5,7,9), (9,1,5)11(1,2,3), (3,4,5), (5,6,7), (7,8,9), (9,10,11), (11,1,3), (3,5,7), (7,9,11), (11,3,7)12(1,2,3), (3,4,5), (5,6,7), (7,8,9), (9,10,11), (11,12,1), (1,3,5), (5,7,9), (9,11,1), (1,5,9)
[0183] The upsampling process is performed to voxelize by adding intermediate points along the edges of the triangle. Additional points are generated based on the upsampling factor value and the width of the block. The additional points are referred to as refined vertices. A point cloud encoder according to the embodiments can voxelize the refined vertices. Additionally, the point cloud encoder can perform attribute encoding based on the voxelized positions (or position values).
[0184] Figure 5 shows an example of a point configuration by LOD according to embodiments.
[0185] As described in FIGS. 1 to 4, the encoded geometry is reconstructed (decompressed) before attribute encoding is performed. When direct coding is applied, the geometry reconstruction operation may include changing the arrangement of the direct-coded points (e.g., placing the direct-coded points at the front of the point cloud data). When trisoop geometry encoding is applied, the geometry reconstruction process involves triangle reconstruction, upsampling, and voxelization. Since attributes depend on geometry, attribute encoding is performed based on the reconstructed geometry.
[0186] A point cloud encoder (e.g., an LOD generation unit (30009)) can reorganize points by LOD. The drawing shows point cloud content corresponding to the LOD. The left side of the drawing shows the original point cloud content. The second figure from the left of the drawing shows the distribution of points of the lowest LOD, and the rightmost figure of the drawing shows the distribution of points of the highest LOD. That is, the points of the lowest LOD are sparsely distributed, while the points of the highest LOD are densely distributed. In other words, according to the direction of the arrow indicated at the bottom of the drawing, as the LOD increases, the spacing (or distance) between points becomes shorter.
[0187] Figure 6 shows an example of a point configuration by LOD according to embodiments.
[0188] As described in FIGS. 1 to 5, a point cloud content providing system or a point cloud encoder (e.g., a point cloud video encoder (10002), the point cloud encoder of FIG. 3, or an LOD generation unit (30009)) can generate an LOD. The LOD is generated by reorganizing points into a set of refinement levels according to a set LOD distance value (or a set of Euclidean distances). The LOD generation process is performed in a point cloud decoder as well as a point cloud encoder.
[0189] The top of Fig. 6 shows examples of points (P0 to P9) of point cloud content distributed in three-dimensional space. The Original Order in Fig. 6 represents the order of points P0 to P9 prior to LOD generation. The LOD-based Order in Fig. 6 represents the order of points following LOD generation. Points are rearranged by LOD. Additionally, higher LODs include points belonging to lower LODs. As illustrated in Fig. 6, LOD0 includes P0, P5, P4, and P2. LOD1 includes the points of LOD0 and P1, P6, and P3. LOD2 includes the points of LOD0, the points of LOD1, and P9, P8, and P7.
[0190] As described in FIG. 3, the point cloud encoder according to the embodiments can perform predictive transform coding, lifting transform coding, and RAHT transform coding selectively or in combination.
[0191] The point cloud encoder according to the embodiments can generate predictors for points and perform predictive transformation coding to set the predicted attribute (or predicted attribute value) of each point. That is, N predictors can be generated for N points. The predictor according to the embodiments can calculate a weight (=1 / distance) value based on the LOD value of each point, indexing information for neighboring points within a set distance per LOD, and the distance value to the neighboring points.
[0192] According to the embodiments, the predicted attribute (or attribute value) is set as the average value of the values obtained by multiplying the attributes (or attribute values, e.g., color, reflectance, etc.) of neighboring points set in the predictor of each point by a weight (or weight value) calculated based on the distance to each neighboring point. The point cloud encoder according to the embodiments (e.g., coefficient quantization unit (30011)) can quantize and inverse quantize the residual values (which may be referred to as residual attributes, residual attribute values, attribute prediction residual values, etc.) obtained by subtracting the predicted attribute (attribute value) from the attribute (attribute value) of each point. The quantization process is as shown in Tables 2 and 3 below.
[0193] int PCCQuantization(int value, int quantStep) {if( value >=0) {return floor(value / quantStep + 1.0 / 3.0);} else {return -floor(-value / quantStep + 1.0 / 3.0);}}
[0194] int PCCInverseQuantization(int value, int quantStep) {if( quantStep ==0) {return value;} else {return value * quantStep;}}
[0195] A point cloud encoder according to the embodiments (e.g., an arismetic encoder (30012)) can entropy-code the quantized and inversely quantized residual values as described above when there are neighboring points in the predictor of each point. A point cloud encoder according to the embodiments (e.g., an arismetic encoder (30012)) can entropy-code the attributes of the corresponding point without performing the process described above when there are no neighboring points in the predictor of each point.
[0196] A point cloud encoder according to the embodiments (e.g., a lifting transformation unit (30010)) can perform lifting transformation coding by generating a predictor for each point, setting the LOD calculated in the predictor, registering neighboring points, and setting weights based on the distance to neighboring points. The lifting transformation coding according to the embodiments is similar to the prediction transformation coding described above, but differs in that weights are cumulatively applied to attribute values. The process of cumulatively applying weights to attribute values according to the embodiments is as follows.
[0197] 1) Create an array QW (QuantizationWight) to store the weight values of each point. The initial value of all elements in QW is 1.0. Add the value obtained by multiplying the current point's predictor weight by the QW value of the predictor index of the neighboring node registered in the predictor.
[0198] 2) Lift prediction process: To calculate the predicted attribute value, the value obtained by multiplying the point's attribute value by a weight is subtracted from the existing attribute value.
[0199] 3) Create temporary arrays named updateweight and update, and initialize the temporary arrays to 0.
[0200] 4) For all predictors, the calculated weight is additionally multiplied by the weight stored in the QW corresponding to the predictor index, and the resulting weight is accumulated in the update weight array with the neighbor node index. In the update array, the value obtained by multiplying the attribute value of the neighbor node index by the calculated weight is accumulated.
[0201] 5) Lift update process: For all predictors, the attribute value of the update array is divided by the weight value of the update weight array at the predictor index, and the original attribute value is added back to the divided value.
[0202] 6) For all predictors, the predicted attribute value is calculated by additionally multiplying the attribute value updated through the lift update process by the weight (stored in QW) updated through the lift prediction process. A point cloud encoder according to the embodiments (e.g., coefficient quantizer (30011)) quantizes the predicted attribute value. Additionally, a point cloud encoder (e.g., arismetic encoder (30012)) entropies the quantized attribute value.
[0203] A point cloud encoder according to the embodiments (e.g., a RAHT transform unit (30008)) can perform RAHT transform coding to predict attributes of upper-level nodes using attributes associated with nodes at lower levels of the octree. RAHT transform coding is an example of attribute intra-coding through octree backward scanning. A point cloud encoder according to the embodiments scans from voxels to the entire region and repeats the merging process up to the root node, merging voxels into larger blocks at each step. The merging process according to the embodiments is performed only on occupied nodes. The merging process is not performed on empty nodes, and the merging process is performed on the node immediately above the empty node.
[0204] The following equation represents the RAHT transformation matrix. g lx,y,z represents the average attribute value of the voxels at level l. g lx,y,z can be calculated from gl+1 2x,y,z and gl+1 2x+1,y,z. g l 2x,y,z and the weights of gl 2x+1,y,z are w1=w l 2x,y,z and w2=wl 2x+1,y,z.
[0205]
[0206] g l-1 x,y,z is a low-pass value used in the merging process at the next higher level. h l-1 x,y,z ε₀ are high-pass coefficients, and the high-pass coefficients at each step are quantized and entropy-coded (e.g., encoding of an arismetic encoder (30012)). The weights are w l-1 x,y,z = w l 2x,y,z + wl is calculated as 2x+1,y,z. The root node is the last g 1 0,0,0 and g 1 0,0,1 It is generated as follows through.
[0207]
[0208] The gDC value is also quantized and entropy-coded, just like the high-pass coefficient.
[0209] FIG. 7 shows an example of a point cloud decoder according to embodiments.
[0210] The point cloud decoder illustrated in FIG. 7 is an example of a point cloud decoder and can perform a decoding operation, which is the reverse process of the encoding operation of the point cloud encoder described in FIG. 1 to 6.
[0211] As described in Fig. 1, the point cloud decoder can perform geometry decoding and attribute decoding. Geometry decoding is performed before attribute decoding.
[0212] A point cloud decoder according to the embodiments comprises an arithmetic decoder (7000), a synthesize octree (7001), a synthesize surface approximation (7002), a reconstruct geometry (7003), an inverse transform coordinates (7004), an arithmetic decoder (7005), an inverse quantize (7006), a RAHT transform (7007), an LOD generater (7008), an inverse lifting (7009), and / or an inverse transform colors (7010).
[0213] An arismetic decoder (7000), an octree composite unit (7001), a surface offset composite unit (7002), a geometry reconstruction unit (7003), and a coordinate system inverse transformation unit (7004) can perform geometry decoding. Geometry decoding according to the embodiments may include direct coding and trisoup geometry decoding. Direct coding and trisoup geometry decoding are applied optionally. Additionally, geometry decoding is not limited to the above examples and is performed as the reverse process of geometry encoding described in FIGS. 1 through 6.
[0214] The arismetic decoder (7000) according to the embodiments decodes the received geometry bitstream based on arismetic coding. The operation of the arismetic decoder (7000) corresponds to the reverse process of the arismetic encoder (30004).
[0215] The octree synthesis unit (7001) according to the embodiments can generate an octree by obtaining an Occupancy code from a decoded geometry bitstream (or information regarding the geometry obtained as a result of decoding). A specific description of the Occupancy code is as described in FIGS. 1 to 6.
[0216] The surface off-relation synthesis unit (7002) according to the embodiments can synthesize a surface based on the decoded geometry and / or the generated octree when trisoop geometry encoding is applied.
[0217] The geometry reconstruction unit (7003) according to the embodiments can regenerate geometry based on a surface and / or decoded geometry. As described in FIGS. 1 through 6, direct coding and trisoop geometry encoding are applied optionally. Accordingly, the geometry reconstruction unit (7003) directly retrieves and adds position information of points to which direct coding has been applied. In addition, when trisoop geometry encoding is applied, the geometry reconstruction unit (7003) can restore geometry by performing reconstruction operations of the geometry reconstruction unit (30005), such as triangle reconstruction, up-sampling, and voxelization operations. Specific details are omitted as they are the same as those described in FIG. 4. The restored geometry may include a point cloud picture or frame that does not contain attributes.
[0218] The coordinate system inverse transformation unit (7004) according to the embodiments can obtain the positions of the points by transforming the coordinate system based on the restored geometry.
[0219] An arismetic decoder (7005), an inverse quantization unit (7006), a RAHT transformation unit (7007), an LOD generation unit (7008), an inverse lifting unit (7009), and / or a color inverse transformation unit (7010) may perform attribute decoding. Attribute decoding according to the embodiments may include RAHT (Region Adaptive Hierarchial Transform) decoding, prediction transformation (Interpolaration-based hierarchical nearest-neighbour prediction-Prediction Transform) decoding, and lifting transformation (interpolation-based hierarchical nearest-neighbour prediction with an update / lifting step (Lifting Transform)) decoding. The three decodings described above may be used optionally, or a combination of one or more decodings may be used. Furthermore, attribute decoding according to the embodiments is not limited to the examples described above.
[0220] The arismetic decoder (7005) according to the embodiments decodes the attribute bitstream into arismetic coding.
[0221] The inverse quantization unit (7006) according to the embodiments inverse quantizes information about the decoded attribute bitstream or the attribute obtained as a result of decoding and outputs the inverse quantized attributes (or attribute values). Inverse quantization may be optionally applied based on the attribute encoding of the point cloud encoder.
[0222] According to embodiments, the RAHT transformation unit (7007), LOD generation unit (7008), and / or inverse lifting unit (7009) can process the reconstructed geometry and inverse quantized attributes. As described above, the RAHT transformation unit (7007), LOD generation unit (7008), and / or inverse lifting unit (7009) can optionally perform a corresponding decoding operation according to the encoding of the point cloud encoder.
[0223] The color inverse conversion unit (7010) according to the embodiments performs inverse conversion coding to inversely convert the color value (or texture) included in the decoded attributes. The operation of the color inverse conversion unit (7010) may be selectively performed based on the operation of the color conversion unit (30006) of the point cloud encoder.
[0224] The elements of the point cloud decoder of FIG. 7 may be implemented in hardware, software, firmware, or a combination thereof, comprising one or more processors or integrated circuits configured to communicate with one or more memories included in a point cloud providing device, although not illustrated in the drawing. One or more processors may perform at least one of the operations and / or functions of the elements of the point cloud decoder of FIG. 7 described above. Additionally, one or more processors may operate or execute a set of software programs and / or instructions for performing the operations and / or functions of the elements of the point cloud decoder of FIG. 7.
[0225] FIG. 8 is an example of a transmission device according to embodiments.
[0226] The transmission device illustrated in FIG. 8 is an example of the transmission device (10000) of FIG. 1 (or the point cloud encoder of FIG. 3). The transmission device illustrated in FIG. 8 can perform at least one of the same or similar operations and methods as the operations and encoding methods of the point cloud encoder described in FIG. 1 to 6. A transmission device according to embodiments may include a data input unit (8000), a quantization processing unit (8001), a voxelization processing unit (8002), an octree occupancy code generation unit (8003), a surface model processing unit (8004), an intra / inter coding processing unit (8005), an arithmetic coder (8006), a metadata processing unit (8007), a color conversion processing unit (8008), an attribute conversion processing unit (or attribute conversion processing unit) (8009), a prediction / lifting / RAHT conversion processing unit (8010), an arithmetic coder (8011) and / or a transmission processing unit (8012).
[0227] The data input unit (8000) according to the embodiments receives or acquires point cloud data. The data input unit (8000) may perform an operation and / or acquisition method identical or similar to the operation and / or acquisition method of the point cloud video acquisition unit (10001) (or the acquisition process (20000) described in FIG. 2).
[0228] The data input unit (8000), quantization processing unit (8001), voxelization processing unit (8002), octree occupancy code generation unit (8003), surface model processing unit (8004), intra / inter coding processing unit (8005), and arithmetic coder (8006) perform geometry encoding. Since the geometry encoding according to the embodiments is identical or similar to the geometry encoding described in FIGS. 1 to 6, a detailed description is omitted.
[0229] The quantization processing unit (8001) according to the embodiments quantizes geometry (e.g., location values of points, or position values). The operation and / or quantization of the quantization processing unit (8001) is the same or similar to the operation and / or quantization of the quantization unit (30001) described in FIG. 3. The specific description is the same as that described in FIG. 1 through 6.
[0230] The voxelization processing unit (8002) according to the embodiments voxelizes the position values of the quantized points. The voxelization processing unit (80002) may perform the same or similar operation and / or process as the operation and / or voxelization process of the quantization unit (30001) described in FIG. 3. The specific description is the same as that described in FIG. 1 to 6.
[0231] The octree occupancy code generation unit (8003) according to the embodiments performs octree coding on the positions of voxelized points based on an octree structure. The octree occupancy code generation unit (8003) can generate an occupancy code. The octree occupancy code generation unit (8003) can perform operations and / or methods identical or similar to the operations and / or methods of the point cloud encoder (or octree analysis unit (30002)) described in FIGS. 3 and 4. The specific description is the same as that described in FIGS. 1 through 6.
[0232] The surface model processing unit (8004) according to the embodiments can perform trisup geometry encoding that reconstructs the positions of points within a specific region (or node) based on a voxel based on a surface model. The surface model processing unit (8004) can perform operations and / or methods identical or similar to the operations and / or methods of the point cloud encoder (e.g., surface approximation analysis unit (30003)) described in FIG. 3. The specific description is the same as that described in FIG. 1 through 6.
[0233] According to the embodiments, the intra / inter coding processing unit (8005) can intra / inter code point cloud data. The intra / inter coding processing unit (8005) can perform coding identical or similar to intra / inter coding. According to the embodiments, the intra / inter coding processing unit (8005) may be included in an arismetic coder (8006).
[0234] An arithmetic coder (8006) according to the embodiments entropy-encodes an octree and / or an approximated octree of point cloud data. For example, the encoding method includes an arithmetic encoding method. The arithmetic coder (8006) performs the same or similar operation and / or method as the arithmetic encoder (30004).
[0235] A metadata processing unit (8007) according to the embodiments processes metadata regarding point cloud data, such as setting values, and provides it to necessary processing processes such as geometry encoding and / or attribute encoding. Additionally, a metadata processing unit (8007) according to the embodiments may generate and / or process signaling information related to geometry encoding and / or attribute encoding. The signaling information according to the embodiments may be encoded separately from geometry encoding and / or attribute encoding. Additionally, the signaling information according to the embodiments may be interleaved.
[0236] The color conversion processing unit (8008), attribute conversion processing unit (8009), prediction / lifting / RAHT conversion processing unit (8010), and arithmetic coder (8011) perform attribute encoding. Since the attribute encoding according to the embodiments is identical or similar to the attribute encoding described in FIGS. 1 to 6, a detailed description is omitted.
[0237] The color conversion processing unit (8008) according to the embodiments performs color conversion coding that converts color values included in attributes. The color conversion processing unit (8008) may perform color conversion coding based on reconstructed geometry. The description of the reconstructed geometry is the same as that described in FIGS. 1 through 6. In addition, it performs the same or similar operation and / or method as the operation and / or method of the color conversion unit (30006) described in FIG. 3. A detailed description is omitted.
[0238] The attribute transformation processing unit (8009) according to the embodiments performs attribute transformation that transforms attributes based on positions where geometry encoding has not been performed and / or reconstructed geometry. The attribute transformation processing unit (8009) performs operations and / or methods identical or similar to the operations and / or methods of the attribute transformation unit (30007) described in FIG. 3. A detailed description is omitted. The prediction / lifting / RAHT transformation processing unit (8010) according to the embodiments may code the transformed attributes by RAHT coding, prediction transformation coding, and lifting transformation coding, or a combination thereof. The prediction / lifting / RAHT transformation processing unit (8010) performs at least one of operations identical or similar to the operations of the RAHT transformation unit (30008), LOD generation unit (30009), and lifting transformation unit (30010) described in FIG. 3. In addition, the descriptions of predictive transformation coding, lifting transformation coding, and RAHT transformation coding are the same as those described in Figures 1 to 6, so a detailed description is omitted.
[0239] The arismetic coder (8011) according to the embodiments can encode coded attributes based on arismetic coding. The arismetic coder (8011) performs the same or similar operation and / or method as the operation and / or method of the arismetic encoder (300012).
[0240] A transmission processing unit (8012) according to embodiments may transmit each bitstream containing encoded geometry and / or encoded attributes and metadata information, or may transmit the encoded geometry and / or encoded attributes and metadata information by configuring them into a single bitstream. When the encoded geometry and / or encoded attributes and metadata information according to embodiments is configured into a single bitstream, the bitstream may include one or more sub-bitstreams. The bitstream according to embodiments may include signaling information and slice data, including SPS (Sequence Parameter Set) for sequence-level signaling, GPS (Geometry Parameter Set) for signaling of geometry information coding, APS (Attribute Parameter Set) for signaling of attribute information coding, and TPS (Tile Parameter Set) for tile-level signaling. The slice data may include information regarding one or more slices. One slice according to embodiments is one geometry bitstream (Geom0 0 ) and one or more attribute bitstreams (Attr0 0 , Attr1 0 It may include ).
[0241] A slice refers to a series of syntax elements representing all or part of a coded point cloud frame.
[0242] According to the embodiments, the TPS may include information regarding each tile (e.g., coordinate value information of a bounding box and height / size information, etc.) for one or more tiles. The geometry bitstream may include a header and a payload. The header of the geometry bitstream according to the embodiments may include identification information of a parameter set included in the GPS (geom_parameter_set_id), a tile identifier (geom_tile_id), a slice identifier (geom_slice_id), and information regarding data included in the payload, etc. As described above, the metadata processing unit (8007) according to the embodiments may generate and / or process signaling information and transmit it to the transmission processing unit (8012). According to the embodiments, the elements performing geometry encoding and the elements performing attribute encoding may share data / information with each other as indicated by the dotted lines. The transmission processing unit (8012) according to the embodiments may perform an operation and / or transmission method identical or similar to the operation and / or transmission method of the transmitter (10003). A detailed explanation is omitted as it is the same as that described in FIGS. 1 and 2.
[0243] FIG. 9 is an example of a receiving device according to embodiments.
[0244] The receiving device illustrated in FIG. 9 is an example of the receiving device (10004) of FIG. 1. The receiving device illustrated in FIG. 9 can perform at least one of the same or similar operations and methods as the operations and decoding methods of the point cloud decoder described in FIG. 1 to FIG. 8.
[0245] A receiving device according to the embodiments may include a receiving unit (9000), a receiving processing unit (9001), an arithmetic decoder (9002), an occupancy code-based octree reconstruction processing unit (9003), a surface model processing unit (triangle reconstruction, up-sampling, voxelization) (9004), an inverse quantization processing unit (9005), a metadata parser (9006), an arithmetic decoder (9007), an inverse quantization processing unit (9008), a prediction / lifting / RAHT inverse transformation processing unit (9009), a color inverse transformation processing unit (9010), and / or a renderer (9011). Each component of the decoding according to the embodiments may perform the inverse process of the components of the encoding according to the embodiments.
[0246] A receiver (9000) according to the embodiments receives point cloud data. The receiver (9000) may perform an operation and / or a receiving method identical or similar to the operation and / or receiving method of the receiver (10005) of FIG. 1. A detailed description is omitted.
[0247] A receiving processing unit (9001) according to the embodiments can obtain a geometry bitstream and / or an attribute bitstream from the received data. The receiving processing unit (9001) may be included in the receiving unit (9000).
[0248] The arismetic decoder (9002), the occupancy code-based octree reconstruction processing unit (9003), the surface model processing unit (9004), and the inverse quantization processing unit (9005) can perform geometry decoding. Since the geometry decoding according to the embodiments is identical or similar to the geometry decoding described in at least one of FIGS. 1 to 8, a detailed description is omitted.
[0249] The arismetic decoder (9002) according to the embodiments can decode a geometry bitstream based on arismetic coding. The arismetic decoder (9002) performs the same or similar operation and / or coding as the operation and / or coding of the arismetic decoder (7000).
[0250] According to the embodiments, the Occupancy code-based octree reconstruction processing unit (9003) can reconstruct an octree by obtaining an Occupancy code from a decoded geometry bitstream (or information regarding geometry obtained as a result of decoding). The Occupancy code-based octree reconstruction processing unit (9003) performs the same or similar operations and / or methods as the octree synthesis unit (7001) and / or octree generation method. According to the embodiments, the surface model processing unit (9004) can perform trisup geometry decoding and related geometry reconstruction (e.g., triangle reconstruction, up-sampling, voxelization) based on the surface model method when trisup geometry encoding is applied. The surface model processing unit (9004) performs the same or similar operations as the surface offset synthesis unit (7002) and / or geometry reconstruction unit (7003).
[0251] The inverse quantization processing unit (9005) according to the embodiments can inverse quantize the decoded geometry.
[0252] A metadata parser (9006) according to the embodiments can parse metadata included in the received point cloud data, such as setting values, etc. The metadata parser (9006) can pass the metadata to geometry decoding and / or attribute decoding. A specific description of the metadata is omitted as it is the same as described in FIG. 8.
[0253] The arismetic decoder (9007), inverse quantization processing unit (9008), prediction / lifting / RAHT inverse transformation processing unit (9009), and color inverse transformation processing unit (9010) perform attribute decoding. Since attribute decoding is identical or similar to the attribute decoding described in at least one of FIGS. 1 to 8, a detailed description is omitted.
[0254] The arismetic decoder (9007) according to the embodiments can decode an attribute bitstream into arismetic coding. The arismetic decoder (9007) can perform decoding of the attribute bitstream based on reconstructed geometry. The arismetic decoder (9007) performs the same or similar operation and / or coding as the operation and / or coding of the arismetic decoder (7005).
[0255] The inverse quantization processing unit (9008) according to the embodiments can inverse quantize the decoded attribute bitstream. The inverse quantization processing unit (9008) performs the same or similar operation and / or method as the operation and / or inverse quantization method of the inverse quantization unit (7006).
[0256] The prediction / lifting / RAHT inverse transformation processing unit (9009) according to the embodiments can process the reconstructed geometry and inverse quantized attributes. The prediction / lifting / RAHT inverse transformation processing unit (9009) performs at least one of the same or similar operations and / or decodings as the operations and / or decodings of the RAHT transformation unit (7007), LOD generation unit (7008), and / or inverse lifting unit (7009) of FIG. 7. The color inverse transformation processing unit (9010) according to the embodiments performs inverse transformation coding to inversely transform the color values (or textures) included in the decoded attributes. The color inverse transformation processing unit (9010) performs the same or similar operations and / or inverse transformation coding as the operations and / or inverse transformation coding of the color inverse transformation unit (7010) of FIG. 7. A renderer (9011) according to the embodiments can render point cloud data.
[0257] FIG. 10 shows an example of a structure capable of interoperability with a point cloud data transmission / reception method / device according to embodiments.
[0258] The structure of FIG. 10 represents a configuration in which at least one of a server (1060), a robot (1010), an autonomous vehicle (1020), an XR device (1030), a smartphone (1040), a home appliance (1050) and / or an HMD (1070) is connected to a cloud network (1010). The robot (1010), the autonomous vehicle (1020), the XR device (1030), the smartphone (1040), or the home appliance (1050) are referred to as devices. Additionally, the XR device (1030) may correspond to a point cloud data (PCC) device according to the embodiments or may be linked with a PCC device.
[0259] The cloud network (1000) may refer to a network that constitutes part of the cloud computing infrastructure or exists within the cloud computing infrastructure. Here, the cloud network (1000) may be configured using a 3G network, a 4G or LTE (Long Term Evolution) network, or a 5G network, etc.
[0260] The server (1060) is connected to at least one of a robot (1010), an autonomous vehicle (1020), an XR device (1030), a smartphone (1040), a home appliance (1050) and / or an HMD (1070) via a cloud network (1000) and can assist in at least some of the processing of the connected devices (1010 to 1070).
[0261] The HMD (Head-Mount Display) (1070) represents one of the types in which an XR device and / or PCC device according to the embodiments may be implemented. A device of the HMD type according to the embodiments includes a communication unit, a control unit, a memory unit, an I / O unit, a sensor unit, and a power supply unit, etc.
[0262] Hereinafter, various embodiments of the device (1010 to 1050) to which the above-described technology is applied will be described. Here, the device (1010 to 1050) illustrated in FIG. 10 may be linked / coupled with a point cloud data transmission / reception device according to the above-described embodiments.
[0263] <PCC+XR>
[0264] The XR / PCC device (1030) may be implemented as a Head-Mount Display (HMD), a Head-Up Display (HUD) equipped in a vehicle, a television, a mobile phone, a smartphone, a computer, a wearable device, a home appliance, digital signage, a vehicle, a stationary robot, or a mobile robot by applying PCC and / or XR (AR+VR) technology.
[0265] The XR / PCC device (1030) can obtain information about surrounding space or real objects by analyzing 3D point cloud data or image data obtained through various sensors or from an external device to generate position data and attribute data for 3D points, and can render and output an XR object to be output. For example, the XR / PCC device (1030) can output an XR object containing additional information about a recognized object by associating it with the recognized object.
[0266] <PCC+XR+모바일폰>
[0267] The XR / PCC device (1030) can be implemented as a mobile phone (1040) or the like by applying PCC technology.
[0268] The mobile phone (1040) can decode and display point cloud content based on PCC technology.
[0269] <PCC+자율주행+XR>
[0270] The autonomous vehicle (1020) can be implemented as a mobile robot, vehicle, unmanned aerial vehicle, etc. by applying PCC technology and XR technology.
[0271] An autonomous vehicle (1020) equipped with XR / PCC technology may refer to an autonomous vehicle equipped with means for providing XR images, or an autonomous vehicle that is the subject of control / interaction within the XR images. In particular, the autonomous vehicle (1020) that is the subject of control / interaction within the XR images is distinguished from the XR device (1030) and can be interconnected with it.
[0272] An autonomous vehicle (1020) equipped with means for providing XR / PCC images can acquire sensor information from sensors including cameras and output XR / PCC images generated based on the acquired sensor information. For example, the autonomous vehicle (1020) can provide an XR / PCC object corresponding to a real object or an object in the screen to the occupant by providing an XR / PCC object by outputting an XR / PCC image with a HUD.
[0273] At this time, when the XR / PCC object is displayed on the HUD, at least a portion of the XR / PCC object may be displayed so as to overlap with the actual object to which the occupant's gaze is directed. On the other hand, when the XR / PCC object is displayed on a display provided inside the autonomous vehicle, at least a portion of the XR / PCC object may be displayed so as to overlap with an object on the screen. For example, the autonomous vehicle (1220) may display XR / PCC objects corresponding to objects such as lanes, other vehicles, traffic lights, traffic signs, motorcycles, pedestrians, buildings, etc.
[0274] VR (Virtual Reality) technology, AR (Augmented Reality) technology, MR (Mixed Reality) technology and / or PCC (Point Cloud Compression) technology according to the embodiments can be applied to various devices.
[0275] In other words, VR technology is a display technology that provides real-world objects or backgrounds solely as CG images. On the other hand, AR technology refers to a technology that displays virtual CG images alongside images of real objects. Furthermore, MR technology is similar to the aforementioned AR technology in that it mixes and combines virtual objects with the real world. However, it is distinguished from AR technology in that while AR technology maintains a clear distinction between real-world objects and virtual objects created from CG images, and uses virtual objects to complement real-world objects, MR technology regards virtual objects as having the same nature as real-world objects. To give a more specific example, the aforementioned MR technology is applied in hologram services.
[0276] However, recently, rather than clearly distinguishing between VR, AR, and MR technologies, they are also referred to as XR (extended Reality) technology. Accordingly, the embodiments of the present disclosure are applicable to all VR, AR, MR, and XR technologies. These technologies may utilize encoding / decoding based on PCC, V-PCC, and G-PCC technologies.
[0277] The PCC method / device according to the embodiments can be applied to a vehicle providing autonomous driving services.
[0278] Vehicles providing autonomous driving services are connected to PCC devices to enable wired / wireless communication.
[0279] When a point cloud data (PCC) transmitting / receiving device according to the embodiments is connected to a vehicle for wired / wireless communication, it can receive and process content data related to AR / VR / PCC services that can be provided along with an autonomous driving service, and transmit it to the vehicle. Additionally, when the point cloud data transmitting / receiving device is mounted on a vehicle, the point cloud transmitting / receiving device can receive and process content data related to AR / VR / PCC services according to a user input signal received through a user interface device, and provide it to the user. A vehicle or a user interface device according to the embodiments can receive a user input signal. The user input signal according to the embodiments may include a signal indicating an autonomous driving service.
[0280] The point cloud data transmission method / device according to the embodiments is interpreted as a term referring to the transmission device (10000) of FIG. 1, the point cloud video encoder (10002), the transmitter (10003), the acquisition-encoding-transmission (20000-20001-20002) of FIG. 2, the point cloud video encoder of FIG. 3, the transmission device of FIG. 8, the device of FIG. 10, the transmission device of FIG. 27, the encoding method of FIG. 30, or the encoding method of FIG. 65, etc.
[0281] The point cloud data receiving method / device according to the embodiments is interpreted as a term referring to the receiving device (10004), receiver (10005), point cloud video decoder (10006) of FIG. 1, the transmission-decoding-rendering (20002-20003-20004) of FIG. 2, the point cloud video decoder of FIG. 7, the receiving device of FIG. 9, the device of FIG. 10, the receiving device of FIG. 28, the decoding method of FIG. 29, the decoding method of FIG. 31, the decoding method of FIG. 54, or the decoding method of FIG. 66, etc.
[0282] In addition, the point cloud data transmission / reception method / device according to the embodiments may be referred to simply as the method / device according to the embodiments.
[0283] According to the embodiments, geometry data, geometry information, location information, etc. constituting the point cloud data are interpreted as having the same meaning. Attribute data, attribute information, attribute information, etc. constituting the point cloud data are interpreted as having the same meaning.
[0284] The method / device according to the embodiments can process point cloud data with consideration of scalable transmission.
[0285] The method / device according to the embodiments describes a method for efficiently supporting selective decoding of a portion of data when such decoding is required due to receiver performance or transmission speed when transmitting / receiving point cloud data. In particular, this document proposes a technique to enhance the efficiency of scalable coding, wherein the encoder on the transmitting side can selectively transmit information required by the decoder on the receiving side regarding already compressed data, and the decoder can decode it. In this case, the coding unit may be a tree level, LOD, layer group, subgroup, data unit, slice, fine segmentation slice (FGS), etc.
[0286] In particular, this document proposes a method to enhance the efficiency of scalable coding among point cloud data compression methods. Here, scalable coding is a technology capable of gradually changing the resolution of data according to conditions such as the request / processing speed / performance / transmission bandwidth of the receiver, enabling the transmitter to efficiently transmit compressed data and the receiver to decode the compressed data. To this end, packing for effectively transmitting point cloud data configured based on layers may be applied in conjunction with the technology of this disclosure.
[0287] Referring to the point cloud data transmission / reception device (or may be abbreviated as encoder / decoder) according to the embodiments illustrated in FIGS. 3 and 7, the point cloud data consists of a set of points, and each point consists of geometry information (or referred to as geometry or geometry data) and attribute information (or referred to as attribute or attribute data). The geometry information is the three-dimensional position information (xyz) of each point. That is, the position of each point is expressed by parameters on a coordinate system representing three-dimensional space (for example, parameters of the three axes representing space, namely the X-axis, Y-axis, and Z-axis (x,y,z)). And, the attribute information refers to the color (RGB, YUV, etc.), reflectance, normal vectors, transparency, etc. of the point. Point Cloud Compression (PCC) performs octree-based compression to efficiently compress distribution characteristics that are non-uniformly distributed in three-dimensional space, and compresses attribute information based on this. The point cloud video encoder and point cloud video decoder illustrated in FIGS. 3 and 7 can process operation(s) according to the embodiments through each component.
[0288] According to the embodiments, the transmitting device compresses the geometry information (e.g., location) and attribute information (e.g., color / brightness / reflectivity, etc.) of the point cloud data and transmits them to the receiving device. At this time, the point cloud data can be configured according to an octree structure with layers or according to the Level of Detail (LoD) based on the level of detail, and based on this, scalable point cloud data coding and representation are possible. At this time, it is possible to decode or represent only a part of the point cloud data depending on the performance or transmission speed of the receiving device, but currently there is no method to remove unnecessary data in advance.
[0289] The present disclosure distinguishes between scalable transmission and scalable decoding depending on the purpose. According to the embodiments, scalable transmission can be used for the purpose of selecting information up to a specific layer without passing through a decoder at a transmitting or receiving device. According to the embodiments, scalable decoding can be used for the purpose of selecting a specific layer during coding. That is, scalable transmission supports the selection of necessary information without passing through a decoder in a compressed state (i.e., at the bitstream stage), thereby enabling the identification of a specific layer at the transmitting or receiving device. On the other hand, scalable decoding can be used in cases such as scalable representation by supporting encoding / decoding only up to the parts required during the encoding / decoding process.
[0290] In this case, the layer configuration for scalable transmission and the layer configuration for scalable decoding may differ. For example, the lower three octree layers including the leaf node may constitute a single layer from the perspective of scalable transmission, but from the perspective of scalable decoding, scalable decoding may be possible for the leaf node layer, leaf node layer-1, and leaf node layer-2 respectively, if they include all layer information.
[0291] FIGS. 11(a) and FIGS. 11(b) show a geometry tree structure based on a single slice and a segmented slice according to embodiments.
[0292] The method / device according to the embodiments may configure a slice for transmitting point cloud data as shown in FIG. 11(a) and FIG. 11(b).
[0293] FIGS. 11(a) and FIGS. 11(b) show geometry tree structures included in different slice structures. According to G-PCC technology, the entire coded bitstream can be included in a single slice. Furthermore, for multiple slices, each slice can include a sub-bitstream. The order of the slices can be the same as the order of the sub-bitstreams. Bitstreams are accumulated in a width-first order of the geometry tree, and each slice can be matched with a group of tree layers ( FIGS. 11(a), FIGS. 11(b)). The divided slices can inherit the layering structure of the G-PCC bitstream.
[0294] Just as the upper layers of a geometry tree do not affect the lower layers, following slices may not affect previous slices.
[0295] The segmented slices according to the embodiments are efficient in terms of error robustness, effective transmission, and supporting region of interest.
[0296] 1) Error resilience
[0297] Compared to a single-slice structure, split slices can be more robust to errors. If a slice contains the entire bitstream of a frame, data loss can affect the entire frame data. On the other hand, if the bitstream is split into multiple slices, even if a portion of the slice is lost, the slices that are unaffected by the loss can still be decoded.
[0298] 2) Scalable transmission
[0299] Consider a case where multiple decoders with different capacities can be supported. If the coded data is in a single slice, the LOD of the coded point cloud can be determined prior to encoding. Therefore, multiple pre-encoded bitstreams of point cloud data with different resolutions can be transmitted independently. This can be inefficient in terms of large bandwidth or storage space.
[0300] When a PCC bitstream is generated and contained within divided slices, a single bitstream can support different levels of decoders. From the decoder's perspective, the receiver can select target layers and pass the partially selected bitstream to the decoder. Similarly, by using a single PCC bitstream instead of partitioning the entire bitstream, partial PCC bitstreams can be efficiently generated at the transmitter's side.
[0301] 3) Region-based spatial scalability
[0302] In terms of G-PCC requirements, region-based spatial scalability can be defined as follows. The compressed bitstream can be configured to have one or more layers. A specific region of interest may have additional layers and high density, and the layers can be predicted from the lower layers.
[0303] To support this requirement, it is necessary to support different detailed representations of regions. For example, in VR / AR applications, it is desirable to represent distant objects with low accuracy and nearby objects with high accuracy. Alternatively, the decoder can increase the resolution of the region of interest upon request. This can be implemented by using scalable structures of G-PCC, such as geometry octrees and scalable attribute coding schemes. Based on the current slice structure containing the entire geometry or attributes, decoders must access the entire bitstream. This can lead to bandwidth, memory, and decoder inefficiencies. On the other hand, if the bitstream is segmented into multiple slices, and each slice contains sub-bitstreams according to scalable layers, the decoder according to the embodiments can select a slice as needed before efficiently parsing the bitstream.
[0304] FIG. 12 shows a layer group structure of a geometry tree and an attribute layer group structure aligned with the geometry tree according to embodiments.
[0305] That is, the left side of FIG. 12 shows a layer group structure of a geometry tree according to the embodiments, and the right side of FIG. 12 shows an attribute layer group structure aligned with the geometry tree on the left side of FIG. 12.
[0306] The method / device according to the embodiments can generate a slice layer group using a layer structure or tree structure of point cloud data as shown in FIG. 12.
[0307] The method / device according to the embodiments may apply segmentation of geometry and attribute bitstreams contained in different slices. Additionally, the coding tree structure of each slice and geometry and attribute coding contained in partial tree information in terms of tree depth may be used.
[0308] Referring to the left side of Fig. 12, an example of a geometry tree structure and a proposed slice segment is shown.
[0309] For example, there are eight layers within the octree (i.e., layers 0 through 7), and five slices can be used to contain sub-bitstreams of one or more layers. A group represents a group of geometry tree layers. For example, Group 1 consists of layers 0 through 4, Group 2 includes layer 5, and Group 3 includes layers 6 and 7. Additionally, a group can be divided into three sub-groups. Parent and child pairs exist in each sub-group. Groups 3-1 through 3-3 are sub-groups of Group 3. When scalable attribute coding is used, the tree structure is identical to the geometry tree structure. The same octree-slice mapping can be used to create attribute slice segments (right side of Fig. 12).
[0310] Layer group: Represents a group of layer structure units that occur in G-PCC coding, such as octree layers, LoD layers, etc.
[0311] Sub-group: For a single layer group, it can be represented as a set of adjacent nodes based on location information. Alternatively, groupings can be formed based on the lowest layer within the layer group (which may refer to the layer closest to the root direction; for example, layer 6 in the case of group 3), or groupings of adjacent nodes can be formed according to Morton code order, distance-based adjacent node groupings, or coding order. Additionally, nodes in a parent-child relationship can be defined to exist within a single sub-group.
[0312] When defining subgroups, boundaries occur in the middle of the layer, and regarding whether to maintain continuity at the boundary, you can indicate whether to use entropy continuously, such as sps_entropy_continuation_enabled_flag, gsh_entropy_continuation_flag, and maintain continuity with the previous slice by providing ref_slice_id.
[0313] The tree structure of the geometry according to the embodiments may be an octree structure, and the attribute layer structure or attribute tree structure according to the embodiments may include a Level of Detail (LOD) structure. That is, the tree structure for point cloud data includes layers corresponding to depth or level, and the layers may be grouped.
[0314] The method / device according to the embodiments (e.g., the octree analysis unit (30002) or LOD generation unit (30009) of FIG. 3, the octree synthesis unit (7001) or LOD generation unit (7008) of FIG. 7) can generate an octree structure of geometry or generate an LOD tree structure of attributes. Additionally, point cloud data can be grouped based on layers of a tree structure as shown in FIG. 11 and FIG. 12.
[0315] Referring to FIG. 12, a plurality of layers are grouped to form a first to third group. A single group may be further divided to form subgroups. FIG. 12 illustrates that the third group is divided into three subgroups.
[0316] The method / device according to the embodiments can generate geometry-based slices and attribute-based slice layers.
[0317] An attribute coding layer can have a different structure from a geometry coding tree.
[0318] To efficiently use the layering structure of G-PCC, it is possible to provide segmenting slices that are paired with the geometry and attribute layering structure.
[0319] For a geometry slice segment, each slice segment may contain data coded from a layer group. Here, a layer group is defined as a group of consecutive tree layers, and the start and end depths of the tree layers may be specific numbers within the tree depth, and the start is smaller than the end.
[0320] For an attribute slice segment, each slice segment contains data coded from a layer group, where the layers may be tree depths or LODs according to an attribute coding scheme.
[0321] The order of coded data within slice segments can be the same as the order of coded data within a single slice.
[0322] As sets of parameters included in the bitstream, the following can be provided.
[0323] FIG. 13 shows a layer group and subgroup structure according to embodiments.
[0324] Referring to Fig. 13, point cloud data and bitstreams can be represented separated by bounding boxes.
[0325] Referring to FIG. 13, a subgroup structure and a bounding box corresponding to the subgroup are illustrated. Layer group 2 is divided into two subgroups (group2-1, group2-2) and included in different slices, and layer group 3 is divided into four subgroups (group3-1, group3-2, group3-3, group3-4) and included in different slices. Given the slices of the layer groups and subgroups along with bounding box information, spatial access can be performed by 1) comparing the bounding box of each slice with the ROI, 2) selecting the slice in which the subgroup bounding box overlaps with the ROI, and 3) decoding the selected slice.
[0326] When the ROI is considered in region 3-3, slices 1, 3, and 6 are selected as the subgroup bounding boxes of layer group 1, subgroup 2-2, and 3-3 that cover the ROI area. For effective spatial access, it is assumed that there is no dependency between subgroups of the same layer group. For live streaming or low-latency use cases, time efficiency can be improved by performing selection and decoding when each slice segment is received.
[0327] The method / device according to the embodiments may represent data as a tree (22000) composed of layers (which may be referred to as depths, levels, etc.) when encoding geometry and / or attributes. Point cloud data corresponding to each layer (depth / level) may be grouped into layer groups (or groups). Four layers may be grouped to form layer group 1 (22001). Layer group 2 (22002) may be further divided (segmented) into two subgroups, and layer group 3 (22003) may be further divided (segmented) into four subgroups. Each subgroup may be composed of each slice to generate a bitstream.
[0328] A receiving device according to the embodiments receives a bitstream, selects a specific slice from the bitstream, and can decode a bounding box corresponding to a subgroup included in the selected slice. For example, if slice 1 is selected, a bounding box (22004) corresponding to layer group 1 can be decoded. Layer group 1 may be data corresponding to the largest area. When additionally displaying a detailed area for layer group 1, the method / device according to the embodiments may select slice 3 and / or slice 6 to hierarchically partially access the bounding boxes (point cloud data) of subgroup 2-2 and / or subgroup 3-3 for the detailed area included in the area of layer group 1.
[0329] Encoding and decoding of point cloud data using the layer group and subgroup of FIG. 13 may be performed by at least one of the transmitting / receiving device of FIG. 1, the encoding and decoding of FIG. 2, the transmitting device / method of FIG. 3, the receiving device / method of FIG. 7, the transmitting / receiving device / method of FIG. 8 and FIG. 9, the devices of FIG. 10, the encoding device of FIG. 27, the encoding method of FIG. 30, the decoding device of FIG. 28, FIG. 29, FIG. 31, or the decoding method of FIG. 54.
[0330] As such, the present disclosure is configured by dividing a geometry bitstream or attribute bitstream into slices in the form of layer-group / subgroup units, and can efficiently compress and restore geometry information or attribute information in slice units. In this case, the continuity of context references can be used as a method to reduce loss of coding efficiency.
[0331] In other words, the context table used during the coding process of a single slice can be used when coding another slice. A context table is based on the correlations between nodes within a geometry tree and can be used to improve coding efficiency. In this case, as a method to consider local correlations between layer groups, a context reference relationship can be established only when the subgroup bounding box of the referenced slice includes or is identical to the subgroup bounding box of the referencing slice. That is, coding efficiency can be further enhanced by using the context table of a slice that has a parent subgroup-child subgroup relationship or an ancestor subgroup-child subgroup relationship. Alternatively, as a method to reduce the burden on the buffer storing the context table, the context table of the first slice can be used in the coding of the following slice.
[0332] The following describes an example applying the context reference continuity among the previously presented criteria. To ensure the continuity of context reference information, the layer-group index and subgroup index of each node are determined, and the reference context reference corresponding to each index is used as the initial value of the context table at the start of each subgroup. Although nodes may exist regardless of the subgroup order during the encoding process, the continuous use of the context table within the subgroup can be guaranteed through the process of storing and retrieving the encoder context state in the buffer.
[0333] This process consists of two stages. When the depth changes, it determines whether the layer group changes and performs different actions accordingly.
[0334] The relationship between parent nodes and child nodes included in a tree structure (e.g., an octree or LOD) according to the embodiments can be expressed as a connection relationship between a node at an upper level and a node at a lower level. The transmitting / receiving device / method according to the embodiments can generate a tree structure and group point cloud data into a plurality of groups based on the layers of the tree structure. The groups may be layer groups, subgroups, etc., and may correspond to slices according to the embodiments.
[0335] The transmitting / receiving device / method according to the embodiments can perform encoding / decoding by loading a context (or context information) stored in a buffer based on a reference layer group, reference subgroup, or reference slice when encoding / decoding point cloud data belonging to a layer group, subgroup, or slice. At this time, the reference layer group, reference subgroup, or reference slice may include nodes corresponding to the parents of the nodes belonging to the layer group, subgroup, or slice to be encoded / decoded.
[0336] In other words, the point cloud data transmission / reception device / method according to the embodiments encodes / decodes a second group based on context information for a first group and stores the context information of the second group in a buffer. The stored context information of the second group may be used to encode / decode a third group. At this time, the first to third groups may be groups grouped based on arbitrary layers. Additionally, the first group may be a group corresponding to the parent layer of the second group.
[0337] Additionally, the point cloud data transmission / reception device / method according to the embodiments may store the context of a parent layer group, parent subgroup, or parent slice in a buffer and use the stored context when encoding / decoding a child layer group, child subgroup, or child slice. The parent group or slice is located at a higher level in the tree structure than the child group or slice, and a node belonging to the parent group (layer group, subgroup) / slice may correspond to a parent-child relationship with a node belonging to the child group / slice.
[0338] FIG. 14 illustrates an example of context reference between layer groups according to embodiments. In particular, FIG. 14 is an example of a fixed context reference, where the current subgroup refers to the parent subgroup.
[0339] Referring to FIG. 14, FGS (fine-granularity slice) may represent a subgroup or slice, and the subgroup or slice may be a group of point cloud data according to the embodiments.
[0340] In FIG. 14, FGS 1(1,0) represents subgroup 0 of layer group 1, and FGS 2(1,1) represents subgroup 1 of layer group 1. FGS N+1(2,0) represents subgroup 0 of layer group 2, and FGS N+2(2,1) represents subgroup 1 of layer group 2.
[0341] In FIG. 14, Save states indicates encoding or decoding the corresponding subgroup and saving context information, and Refer context indicates referring to the saved context information to encode or decode the corresponding subgroup.
[0342] Accordingly, in FIG. 14, it is exemplified that FGS 1 refers to the context information of FGS 0 (23001), and FGS N+1 (23003) refers to the context information of FGS 1 (23002). As shown in FIG. 14, a subgroup (or slice) belonging to layer group 2 may refer to the context of a subgroup (or slice) belonging to layer group 1. Also, a subgroup belonging to layer group 1 may be a parent subgroup of a subgroup belonging to layer group 2.
[0343] Context inheritance
[0344] FIG. 14 illustrates a context reference structure of layer group slicing according to embodiments. In the figure, FGS may correspond to subgroups according to embodiments, and subgroups in the same row are considered to be in the same layer group. Subgroups may correspond to slices. In the figure, an arrow pointing from one slice to another indicates a context reference relationship between two slices. The context reference of the current slice may be one of the slices decoded prior to the current slice.
[0345] Considering an example of spatial random access, parent subgroup referencing can be a good choice to ensure independence between subgroups. However, as the number of subgroups or layer groups increases, the number of context buffers may increase. If the number of context buffers is considered in terms of the number of referenced slices, the number of context buffers can be the sum of all subgroups excluding those belonging to the first and last layer groups. This can be formulated as follows, where N represents the number.
[0346] N (context buffer) =1+N subgroup ×(N (layer-group) -2)
[0347] FIG. 15 illustrates an example of context reference between groups according to embodiments. In particular, FIG. 15 is an example of a flexible context reference, in which all subgroups refer to the context information of the root subgroup (i.e., the root slice).
[0348] In FIG. 15, FGS 1 (24002) to FGS N, FGS N+1 (24003) to FGS 2N+1 refer to the context information of FGS 0 (24001). That is, all subgroups (or slices) excluding FGS 0 (24001) refer to FGS 0 (24001).
[0349] One method for mitigating the size of the context buffer according to the embodiments may be to reduce the number of subgroups referenced by following slices. An extreme case of this method is referencing the root slice as shown in FIG. 15. This approach based on the number of context buffers mentioned above results in a single context buffer because the number of referenced subgroup slices is zero.
[0350] N (context buffer) =1
[0351] That is, all slices refer to the first slice. Compared to the case of Fig. 14, the number of stored context states is reduced to one, and all dependent slices refer to the first slice.
[0352] FIG. 16 (a) to (c) illustrates examples of context buffer management according to embodiments. In particular, it is an example when referencing a parent subgroup.
[0353] Referring to FIG. 16 (a), a transceiver device according to the embodiments processes FGS 0 (0,0) (25001) and stores context information (25002) for the corresponding group (FGS 0) in a context buffer.
[0354] Referring to FIG. 16(b), the transceiver according to the embodiments can load context information (25004) of FGS 0 (0,0) stored in a context buffer to process (encode or decode) FGS 2 (1,1) (25003) (load states), process FGS 2 (1,1), and save context information of FGS 2 (1,1) to the context buffer (save states). At this time, the loading process of context information can be performed by referring to the ref_layer_group_id and ref_subgroup_id parameters.
[0355] Referring to FIG. 16 (c), the transceiver according to the embodiments may load context information (context states (1,1)) (25006) of FGS 2 (1,1) to process (encode or decode) FGS N+2 (2,1) (25005). Since FGS N+2 (2,1) belongs to the last layer group, the context information is not stored in the context buffer.
[0356] In this way, for the parent subgroup reference method, changes in the context buffer are described in FIGS. 16 (a) through (c). When the bitstream of a slice is decoded, the context state is stored in the context buffer as in FIG. 16 (a). When context inheritance is used, the context state of the following slices is initialized by one of the stored context states of the previous slices indicated by ref_layer_group_id and ref_subgroup_id (Fig. 16 (b)). Therefore, the context state of the slice belonging to the last layer group is initialized by the stored context state of the previous slice as in FIG. 16 (c). However, the receiving device (or decoder) according to the embodiments may decide not to store the context of FGS N+1 to 2N if there is a restriction that it does not reference a subgroup belonging to the same layer group. Based on the previous information, the smart decoder can save the context buffer.
[0357] FIG. 17 (a) to (c) illustrates examples of context buffer management according to embodiments. In particular, examples are provided when referring to the root layer group.
[0358] Referring to FIG. 17(a), a transceiver according to the embodiments processes FGS 0(0,0)(26001) and stores context information (26002) for the corresponding group (FGS 0) in a context buffer.
[0359] Referring to FIG. 17(b), the transceiver according to the embodiments can load context information (26004) of FGS 0 stored in a context buffer to process (encode or decode) FGS 2(1,1)(26003) (load states), process FGS 2(1,1), and save context information of FGS 2(1,1) to the context buffer (save states). At this time, the loading process of context information can be performed by referring to the ref_layer_group_id and ref_subgroup_id parameters.
[0360] Referring to FIG. 17(c), the transceiver according to the embodiments may load context information (context states (0,0)) (26006) of FGS 0 (0,0) (26001) to process (encode or decode) FGS N+2 (2,1) (26005). Since FGS N+2 (2,1) (26005) belongs to the last layer group, the context information is not stored in the context buffer.
[0361] When flexible context references are allowed, there exist context states that are not used by following slices. For example, consider a root layer group reference where all context states of dependent slices are initialized by the context states stored from the first slice. Decoders, unaware of the overall reference structure, may use the current context of the following slices. However, as expected, the context states stored from (1,0) to (1,N-1) are not used by any slice. In this example, inefficiency arises from a lack of information on the decoder's side.
[0362] Figures 18 (a) to (c) show examples of context buffer management according to embodiments.
[0363] Referring to FIG. 18 (a), a transceiver device according to the embodiments processes FGS 0 (0,0) (27001) and stores context information (27002) for the corresponding group (FGS 0) in a context buffer.
[0364] Referring to FIG. 18(b), the transmitting / receiving device / method according to the embodiments may retrieve context information (context states (0,0)) (27004) of FGS 0 (0,0) (27001) from the context buffer to process (encode or decode) FGS 2 (1,1) (27003). Then, the context information of FGS 2 (1,1) may be stored in the context buffer based on context reference indication flag (context_reference_indication_flag) information. That is, depending on the context_reference_indication_flag information, the context information of FGS 2 (1,1) may or may not be stored in the buffer. The context_reference_indication_flag information may be generated and transmitted by the transmitting device / method according to the embodiments and may be used by the receiving device / method according to the embodiments.
[0365] Referring to FIG. 18 (c), the transmitting / receiving device / method according to the embodiments can retrieve context information (context states (0,0)) (27006) of FGS 0 (0,0) (27001) from the context buffer to process (encode or decode) FGS N+2 (2,1) (27005).
[0366] Since the context buffer always stores only the context information of FGS 0, the transmitting and receiving device according to the embodiments can efficiently use the memory of the buffer.
[0367] A transceiver device according to the embodiments proposes a new signal to assist a decoder in increasing context buffer management efficiency by indicating whether the current context of the followed slices is in use and determining whether to save the context.
[0368] A followed slice according to the embodiments may represent a slice that is referenced by other slices. A following slice according to the embodiments may represent a slice that references other slices.
[0369] To improve context buffer management from the decoder perspective, context reference information is proposed in terms of the current slice. Specifically, context state reference indication by the followed slice is proposed.
[0370] FIG. 18(b) illustrates a proposed signal buffer management case. Compared to previous cases without other information, storing the context state in the context buffer is determined by the context_reference_indication_flag information. If the context reference indicator is on, the decoder stores the current context state in the context buffer to allow the context state to be used by the followed slices. On the other hand, if the context reference indicator is off, the decoder does not store the current context state in the context buffer to save context memory. Comparing FIG. 16(c) and FIG. 17(c), the context buffer of the root layer reference case has only one context compared to the context buffer of the parent subgroup reference case. The size of the saved context memory is N units, where N represents the number of subgroups.
[0371] Meanwhile, in order to prevent the reduction in coding efficiency caused by dividing slices, the receiver may use context, phiBuffer, Planar context, and buffer continuity between slices, and the receiver may also determine whether to store context, phiBuffer, and Planar context based on whether to reuse context.
[0372] if (_dep_gbh.context_reuse_flag) {
[0373] _refIdxToSavedArrayIdx[curLayerGroup][_dep_gbh.subgroup_id] = _ctxtMemSaved.size();
[0374] _ctxtMemSaved.push_back(cur_ctxtMem);
[0375]
[0376] if(_gps->geom_angular_mode_enabled_flag)
[0377] _phiBufferSaved.push_back(cur_phiBuffer);
[0378] if(_gps->geom_planar_mode_enabled_flag)
[0379] _planarSaved.push_back(cur_planar);
[0380] int idx = _refIdxToSavedArrayIdx[curLayerGroup][_dep_gbh.subgroup_id];
[0381] }
[0382] Figures 19 (a) to (c) illustrate a context buffer release method according to embodiments.
[0383] The method according to the embodiments includes a method for determining the decoder context release time.
[0384] When indicating whether to reuse a context, it allows for efficient use of context buffer memory by specifying whether to store the context used in a particular slice in the context buffer, thereby saving only the context necessary for coding subsequent slices. However, in this case, since the context remains in the context buffer until coding for that frame is finished, the buffer can become burdened as the number of stored contexts increases. To use the context buffer more efficiently, a method can be used to delete (i.e., release) stored contexts from the context buffer when they are no longer in use.
[0385] 1) List-based context memory management method
[0386] FIGS. 19 (a) to (c) illustrates a method for managing a list of slices / subgroups that use context information stored in a buffer as a method for effectively managing context memory (i.e., a buffer). FIGS. 19 (a) to (c) illustrates a method for managing context memory based on a list.
[0387] FIG. 19 (a): The context reference indication flag (context_reference_indication_flag) can indicate whether the context used to code the current slice is saved. That is, if the context reference indication flag (context_reference_indication_flag) is 1, it indicates that the context of the current slice / subgroup can be used in the slice / subgroup passed thereafter, and to this end, the context of the current slice / subgroup can be saved in the context buffer. To use the context later, the index of the context can be made to have the same index as the subgroup index. In this case, if information about the slice / subgroup using the current context is provided as a list, the target list can be saved and compared with the list of cases actually used. The list of slices / subgroups using the context can be examined in advance by the encoder and passed, or the list can be generated by the decoder by estimating the context reference relationship based on a predetermined layer group structure. For example, if fixed by referencing the context of a parent subgroup, a list of subgroups that are in a child relationship with the current subgroup can be examined, and in this embodiment, this can be called List A.
[0388] Fig. 19 (b): When a new slice / subgroup is passed, the context used to decode the slice / subgroup can be specified based on the context reference ID. In the following example, the context corresponding to (0,0) is used, and (1,1), which is the index of the current slice / subgroup, can be added to the list of used subgroups. During the coding process, the list of slices / subgroups that use a specific context is called List B. If the context reference indication flag (context_reference_indication_flag) is 1, it means that the context of the current slice / subgroup will be used later, so it can be stored in the context buffer, and List B (1,1) of the context buffer (1,1) can be initialized.
[0389] (c) of FIG. 19: Memory can be managed by releasing (i.e., freeing) the context buffer when List A and List B are the same for a specific context buffer. As shown below, when the context reference id is (1,1), coding can be done using the context state (1,1) in the context buffer. At this time, the index (2,1) of the current slice / subgroup can be added to List B. Since List B and List A are identical at (2,1), it can be seen that the context state (1,1) will not be used in the future, and in this case, the memory of the context buffer can be effectively managed by releasing the memory that stored the context state (1,1).
[0390] The embodiments include a method for specifying an index of the context based on a layer group index and a subgroup index. If a unique slice index is assigned to each slice, the context buffer list can be managed based on the slice index.
[0391] FIG. 20 illustrates a context buffer release method according to embodiments.
[0392] The embodiments include a method for managing context memory based on the number of references.
[0393] The context buffer can be used efficiently by managing context states that are no longer in use in real time based on the number of times each context state is used. The embodiments can manage the context buffer through a target number and a counter for each context state in the context buffer.
[0394] FIG. 20 (a): The target number represents the number of times the corresponding context state is used, and the counter can update the number of times the context state is used. That is, the counter value increases by 1 each time the corresponding context state is used. FIG. 20 (a) represents the case of the first slice / subgroup, and the context state (0,0) can be stored in the context buffer when the context_reference_indication_flag is 1 or when layer-group slicing is used. Additionally, N, the number of times the context state is used, can be stored in the target number; by using the case of using the parent subgroup's context state as an example, this can be set to the same value as N, the number of subgroups belonging to layer-group 1. Alternatively, the list of child subgroups of FGS 0 is FGS 1 (1,0), FGS 2 (1,1), … , it can be obtained as FGS N (1, N-1), and the number of elements N in the list can be specified as the target number.
[0395] FIG. 20 (b): When coding a new slice / subgroup, the context reference of the slice / subgroup can be found in the context buffer via the context reference ID. The context state (0,0) is used, and the counter can be increased by 1. Since the context state (0,0) was already used to code FGS 1 (1,0), the counter can have a value of 2. Additionally, the context state of FGS 2 (1,1) can be stored in the context buffer because if the context_reference_indication_flag is 1, it means that it can be used as a context reference in the following slice / subgroup. At this time, the number of child slices / subgroups using the context state (1,1) can be stored in the target number.
[0396] Fig. 20 (c): Memory usage of the context buffer can be managed through memory release for context states where the target number and counter managed in the context buffer are the same. The context state (1,1) corresponding to the context reference ID of FGS N+2(2,1) can be used, and the counter (1,1) can be incremented by 1. In this case, the target number and the counter value of the context state (1,1) become the same, which means that the context state (1,1) is no longer used. Therefore, even if the context state (1,1) is deleted from the context buffer, it does not affect the coding of subsequent slices / subgroups. Memory usage of the context buffer can be minimized by deleting context states that are no longer used from the context buffer.
[0397] The following shows an example of decoder code implementation.
[0398] if (_dep_gbh.context_reference_indication_flag) {
[0399] _refIdxToSavedArrayIdx[curLayerGroup][_dep_gbh.subgroup_id] = _ctxtMemSaved.size();
[0400] _ctxtMemSaved.push_back(cur_ctxtMem); _numSubsequentSubgroups.push_back(_dep_gbh.numSubsequentSubgroups);
[0401] if (_gps->geom_angular_mode_enabled_flag)
[0402] _phiBufferSaved.push_back(cur_phiBuffer);
[0403] if (_gps->geom_planar_mode_enabled_flag)
[0404] _planarSaved.push_back(cur_planar);
[0405] }
[0406] _numSubsequentSubgroups[refArrayIdx]--;
[0407] if (_numSubsequentSubgroups[refArrayIdx] == 0) {
[0408] _ctxtMemSaved[refArrayIdx].resetMap();
[0409] _ctxtMemSaved[refArrayIdx].reset();
[0410] }
[0411] The target number for each context state can be passed after the encoder examines the number of subsequent subgroups. If necessary, by additionally passing a list of slices / subgroups used as context references, the accuracy of the number of subsequent subgroups can be verified, and the decision to delete the context state can be made.
[0412] FIG. 21 illustrates a context memory management method according to embodiments.
[0413] The embodiments include a method for managing context memory based on a data unit coding structure.
[0414] Slices / subgroups can be passed based on a specific order. Representative methods for this include breadth-first search and depth-first search. Breadth-first search is a method that codes subgroups belonging to the same layer group first, and then codes subgroups belonging to child layer groups. In contrast, depth-first search is a method that reaches the subgroup corresponding to the maximum depth first, and then codes the children belonging to the same parent first.
[0415] If each node is considered as an FGS slice index, it can be assumed that Slice 0 belongs to Layer Group 0, Slices 1 and 2 to Layer Group 1, and Slices 3, 4, 5, and 6 to Layer Group 2. Additionally, slices connected by solid lines can be assumed to represent slice pairs belonging to a parent-child relationship. When coding based on Breadth-First Search, the order can be coded as 0, 1, 2, 3, 4, 5, 6, and when coding based on Depth-First Search, the order is 0, 1, 3, 4, 2, 5, 6.
[0416] When managing context buffer memory based on slice order, it may operate as follows. In this case, it can be assumed that the parent context state is used as the reference context. If coding is based on breadth-first search, the context state of the parent layer group is no longer used when the coding of the subgroup belonging to each layer group is finished. That is, context state 0 is used when coding slices 1 and 2, but context state 1 is no longer used when slice 2 is coded. In this case, the context state of the parent subgroup (the context state of slice 0) can be removed from the context buffer at the time the layer group switches (slice 2).
[0417] If coded based on depth-first search, the context state can be removed when coding for a child subgroup ends or when the layer group transitions (transition from bottom to root). After coding for slices 3 and 4 ends, the process moves on to slice 2. At this point, slices 3 and 4 belong to layer group 2, and slice 2 belongs to layer group 1. Slices 3 and 4 are coded using the context state of slice 1, and since context state 1 is no longer used, it can be removed from the context buffer.
[0418] In this case, context memory can be used efficiently without additional information such as the list of subsequent subgroups or the number of subsequent subgroups.
[0419] FIG. 22 shows a bitstream containing point cloud data according to embodiments.
[0420] An encoder according to the embodiments can encode point cloud data and generate related parameter information (i.e., signaling information) to generate a bitstream. A decoder according to the embodiments can receive the bitstream and decode the point cloud data based on the parameter information (i.e., signaling information).
[0421] Information regarding separated slices can be defined in parameter sets and SEI messages as follows. It can be defined in the Sequence Parameter Set (SPS), Geometry Parameter Set (GPS), Attribute Parameter Set (APS), Geometry Slice Header (GSH), and Attribute Slice Header (ASH). Depending on the application or system, it can be defined in corresponding locations or separate locations to allow for different application scopes and methods. In other words, the signaling can have different meanings depending on where it is transmitted. If defined in the SPS, it may apply uniformly to the entire sequence; if defined in the GPS, it may indicate use for location restoration; if defined in the APS, it may indicate application for attribute restoration; if defined in the TPS (Tile Parameter Set), it may indicate that the signaling is applied only to points within a tile; and if transmitted at the slice level, it may indicate that the signaling is applied only to that specific slice. Additionally, depending on the application or system, it can be defined in corresponding locations or separate locations to allow for different application scopes and methods. In addition, if the syntax element (or field) defined below can be applied to multiple point cloud data streams as well as the current point cloud data stream, it can be conveyed through a higher-level concept parameter set, etc.
[0422] Each abbreviation means the following. Each abbreviation may be referred to by other terms within the scope of equivalent meaning: SPS: Sequence Parameter Set, GPS: Geometry Parameter Set, APS: Attribute Parameter Set, TPS: Tile Parameter Set, Geom: Geometry bitstream = geometry slice header + geometry slice data, Attr: Attribute bitstream = attribute blick header + attribute brick data.
[0423] The embodiments may generate the information independently of the coding technique or in conjunction with the coding method. A set of tile parameters may be defined to support locally different scalability. Alternatively, a bitstream may be selected at the system level by defining a Network Abstract Layer (NAL) unit and transmitting relevant information that allows selecting a layer, such as a layer ID (layer_id).
[0424] The parameters (which may be referred to in various ways, such as metadata, signaling information, etc.) according to the embodiments described below may be generated during the process of the transmitter according to the embodiments described below, and may be transmitted to the receiver according to the embodiments and used in the reconstruction process.
[0425] For example, parameters according to the embodiments may be generated in the metadata processing unit (or metadata generator) of the transmitting device according to the embodiments described below, and may be obtained in the metadata parser of the receiving device according to the embodiments.
[0426] FIG. 23 shows an example of a syntax structure of a sequence parameter set (SPS) of a bitstream according to embodiments.
[0427] FIG. 24 shows an example of the syntax structure of a dependent geometry data unit header of a bitstream according to embodiments.
[0428] FIG. 25 shows an example of the syntax structure of a dependent attribute data unit header of a bitstream according to embodiments.
[0429] FIGS. 26a and 26b show an example of the syntax structure of a layer group structure inventory (LGSI) according to the embodiments.
[0430] The definitions of each syntax element in FIGS. 23 through 26a and FIG. 26b are as follows:
[0431] If the layer_group_enabled_flag is equal to 1, it indicates that the geometry and / or attribute bitstreams of a frame or tile are contained in multiple slices corresponding to a coding layer group or its subgroups. If the layer_group_enabled_flag is equal to 0, it indicates that the geometry bitstreams of a frame or tile are contained in a single slice.
[0432] The layer_group_slice_order_type indicates the ordering type of the layer group slices. If layer_group_slice_order_type is 0, it indicates a breadth-first search order of the layer group slices. If layer_group_slice_order_type is 1, it indicates a depth-first search order of the layer group slices. A layer_group_slice_order_type of 2 indicates that no ordering type is specified.
[0433] If the context_reference_indication_flag is 1, it indicates that the context state of the current dependent slice is inherited by one or more followed dependent slices. If context_reference_indication_flag is 0, it indicates that the context state of the current dependent slice is not inherited by followed dependent slices.
[0434] Decoders can manage the context buffer using context_reference_indication_flag. If context_reference_indication_flag is 1, the context state of the current dependent slice is stored in the context buffer at the end of decoding. If context_reference_indication_flag is 0, the context state of the current dependent slice is not stored in the context buffer.
[0435] The number of subsequent data units (num_subsequent_data_units) represents the number of subsequent dependent data units that use the context state of the current data unit.
[0436] If the subsequent_data_unit_list_present_flag is 1, it indicates that a list of subsequent data units exists. If next_data_unit_list_present_flag is 0, it indicates that a list of subsequent data units does not exist.
[0437] The number of layer groups (number_of_layer_groups) indicates the number of layer groups in the list of subsequent data units.
[0438] The subsequent layer group id (subsequent_layer_group_id) represents the layer group index of the subsequent data unit.
[0439] The number of subgroups indicates the number of subgroups in the layer group of the list of subsequent data units.
[0440] The subsequent subgroup ID (subsequent_subgroup_id) represents the subgroup index within the layer group of the subsequent data unit.
[0441] Layer group structure inventory syntax
[0442] The sequence parameter set id (lgsi_seq_parameter_set_id) represents the sps_seq_parameter_set_id value. An lgsi_seq_parameter_set_id of 0 is a requirement for bitstream conformity.
[0443] lgsi_frame_ctr_lsb_bits represents the length of the lgsi_frame_ctr_lsb syntax element (or field) in bits.
[0444] lgsi_frame_ctr_lsb represents the least significant bit of lgsi_frame_ctr_lsb_bits of a FrameCtr for which the group structure inventory is valid. The layer group structure inventory remains valid until it is replaced by another layer group structure inventory.
[0445] Count of slices (lgsi_num_slice_ids_minus1): This value plus 1 represents the number of slices in the layer group structure inventory.
[0446] slice_id (gi_slice_id) represents the slice ID of the sid-th slice within the layer group structure inventory. It is a requirement of bitstream conformance that all values of lgsi_slice_id must be unique within the layer group structure inventory.
[0447] Layer group count(lgsi_num_layer_groups_minus1)+ 1 represents the number of layer groups.
[0448] Layer group id (lgsi_layer_group_id) represents the layer group indicator. The range of lgsi_layer_group_id is from 0 to lgsi_num_layer_groups_minus1.
[0449] The number of layers (lgsi_num_layers_minus1) + 1 represents the number of coded layers in the slice of the i-th layer group of the sid-th slice. The total number of coded layers required to decode the n-th layer group is equal to the sum of lgsi_num_layers_minus1[sid][i] + 1 for i from 0 to n.
[0450] The number of subgroups (lgsi_num_subgroups_minus1) + 1 represents the number of subgroups in the i-th layer group of the sid-th slice.
[0451] Subgroup ID (lgsi_subgroup_id) represents the layer group indicator. The range of lgsi_subgroup_id is from 0 to lgsi_num_subgroups_minus1.
[0452] The parent subgroup id (lgsi_parent_subgroup_id) represents the indicator of a subgroup within the layer group pointed to by lgsi_subgroup_id. The range of lgsi_parent_subgroup_id is from 0 to gi_num_subgroups_minus1 in the layer group pointed to by lgsi_subgroup_id.
[0453] The subgroup bounding box origin (lgsi_subgroup_bbox_origin) and subgroup bounding box size (lgsi_subgroup_bbox_size) represent the bounding box of the current subgroup.
[0454] The subgroup bounding box origin (lgsi_subgroup_bbox_origin) represents the origin of the subgroup bounding box of the subgroup pointed to by lgsi_subgroup_id among the layer groups pointed to by lgsi_layer_group_id.
[0455] lgsi_subgroup_bbox_size represents the size of the subgroup bounding box of the subgroup indicated by lgsi_subgroup_id among the layer groups indicated by lgsi_layer_group_id.
[0456] lgsi_origin_bits_minus + 1 represents the length of the lgsi_origin_xyz syntax element in bits.
[0457] The origin location (lgsi_origin_xyz) indicates the origin of all partitions. The value of lgsi_origin_xyz[ k ] is equal to sps_bounding_box_offset[ k ].
[0458] lgsi_origin_log2_scale represents a scaling factor for scaling the components of lgsi_origin_xyz. The value of lgsi_origin_log2_scale is the same as sps_bounding_box_offset_log2_scale.
[0459] FIG. 27 illustrates a point cloud data transmission device according to embodiments. The elements of the transmission device illustrated in FIG. 27 may be implemented in hardware, software, processors connected to memory, and / or combinations thereof. That is, the elements of the transmission device of FIG. 27 may be implemented in hardware, software, firmware, or a combination thereof, including one or more processors or integrated circuits configured to communicate with one or more memories, although not illustrated in the drawing. One or more processors may perform at least one of the operations and / or functions of the elements of the transmission device of FIG. 27 described above. Additionally, one or more processors may operate or execute a set of software programs and / or instructions for performing the operations and / or functions of the elements of the transmission device of FIG. 27. The execution order of each block in FIG. 27 may be changed, some blocks may be omitted, and some blocks may be newly added.
[0460] Referring to FIG. 27, an example of a detailed functional configuration for encoding / transmitting point cloud data is shown. When point cloud data is input, the encoder can encode geometry information (geometry data: e.g., XYZ coordinates, phi-theta coordinates, etc.) and attribute information (attribute data: e.g., color, reflectance, intensity, grayscale, opacity, medium, material, glossiness, etc.). The compressed data is divided into units for transmission, and can be packed into units suitable for selecting necessary information in bitstream units according to layering structure information through a sub-bitstream generator (40010).
[0461] According to the embodiments, when different types of bitstreams are included in a single slice, the encoder can separate the generated bitstream (AEC bitstream or DC bitstream) according to the purpose. Then, each slice or adjacent information can be included in a single slice according to the layer group information. At this time, through the metadata generator (40006), information such as layer-group information, layer information included in the layer-group, number of nodes, layer depth information, number of nodes included in the sub-group, bitstream type, bitstream_offset, bitstream_length, and bitstream direction can be transmitted according to each slice ID.
[0462] When point cloud data is input to a transmitting device according to the embodiments, the geometry encoder (40002) encodes position information (geometry data: e.g., XYZ coordinates, phi-theta coordinates, etc.), and the attribute encoder (40004) encodes attribute information (attribute data: e.g., color, reflectance, intensity, grayscale, opacity, medium, material, glossiness, etc.).
[0463] The compressed (encoded) data is divided into units for transmission, and can be packed into units suitable for selecting necessary information in bitstream units according to layering structure information through the sub-bitstream generation unit (40010).
[0464] According to the embodiments, an octree-coded geometry bitstream is input to an octree-coded geometry bitstream segmentation (40011), and a direct-coded geometry bitstream is input to a direct-coded geometry bitstream segmentation (40012).
[0465] The octree-coded geometry bitstream segmentation unit (40011) performs the process of dividing the octree-coded geometry bitstream into one or more groups and / or subgroups based on information about the segmented (separated) slices generated in the layer-group structure generation unit (40014) and / or information related to direct coding.
[0466] Additionally, the direct-coded geometry bitstream segmentation unit (40012) performs the process of dividing the direct-coded geometry bitstream into one or more groups and / or sub-groups based on information about the segmented (separated) slices generated in the layer-group structure generation unit (40014) and / or information related to direct coding.
[0467] The output of the octree-coded geometry bitstream segmentation unit (40011) and the output of the direct-coded geometry bitstream segmentation unit (40012) are input to the geometry bitstream bonding unit (40013).
[0468] The geometry bitstream bonding unit (40013) performs a geometry bitstream bonding process based on information about the segmented (separated) slices generated by the layer-group structure generation unit (40014) and / or information related to direct coding, and outputs sub-bitstreams in layers to the segmented slice generation unit (40016). For example, the geometry bitstream bonding unit (40013) performs a process of joining an AEC bitstream and a DC bitstream within a single slice. Final slices are created in the geometry bitstream bonding unit (40013).
[0469] The coded attribute bitstream segmentation unit (40015) performs the process of dividing the coded attribute bitstream into one or more groups and / or subgroups based on information about the segmented (separated) slices generated in the layer-group structure generation unit (40014) and / or information related to direct coding. One or more groups and / or subgroups of attribute information may be linked with one or more groups and / or subgroups for geometry information, or may be generated independently.
[0470] The segmented slice generation unit (40016) receives the geometry bitstream bonding unit (40013) and / or the coded attribute bitstream segmentation unit (40015) based on information about the segmented (separated) slice generated by the metadata generation unit (40006) and / or information related to direct coding, and performs the process of segmenting one slice into multiple slices. Each sub-bitstream is transmitted through each slice segment. At this time, the AEC bitstream and the DC bitstream may be transmitted through one slice or through different slices.
[0471] The multiplexer (40008) multiplexes the output of the segmented slice generation unit (40016) and the output of the metadata generation unit (40006) layer by layer and outputs them to the transmitter (40009).
[0472] When different types of bitstreams (e.g., AEC bitstream and DC bitstream) are included in a single slice, the geometry encoder (40002) can separate the generated bitstreams (e.g., AEC bitstream and DC bitstream) according to the purpose. Then, each slice or adjacent information can be included in a single slice according to information about the segmented (separated) slices and / or information related to direct coding (i.e., layer-group information) generated by the layer-group structure generation unit (40014) and / or metadata generation unit (40006). According to embodiments, information about the segmented (separated) slices and / or information related to direct coding (e.g., information such as bitstream type, bitstream_offset, bitstream_length, bitstream direction, etc., along with layer-group information, layer information included in the layer-group, number of nodes, layer depth information, and number of nodes included in the sub-group) according to each slice id) can be transmitted through the metadata generation unit (40006). Information about segmented (separated) slices and / or information related to direct coding (e.g., layer-group information according to each slice ID, layer information included in the layer-group, number of nodes, layer depth information, number of nodes included in the sub-group, along with information such as bitstream type, bitstream_offset, bitstream_length, bitstream direction, etc.) may be signaled in SPS, APS, GPS, geometry data unit headers, attribute data unit headers, or SEI messages, etc.
[0473] FIG. 28 shows a point cloud data receiving device according to embodiments.
[0474] The receiving method of FIG. 28 may follow the reverse process of the transmitting device of FIG. 27. The elements of the receiving device illustrated in FIG. 28 may be implemented in hardware, software, processors connected to memory, and / or combinations thereof. That is, the elements of the receiving device of FIG. 28 may be implemented in hardware, software, firmware, or a combination thereof, including one or more processors or integrated circuits configured to communicate with one or more memories, although not illustrated in the drawing. One or more processors may perform at least one of the operations and / or functions of the elements of the receiving device of FIG. 28 described above. Additionally, one or more processors may operate or execute a set of software programs and / or instructions for performing the operations and / or functions of the elements of the receiving device of FIG. 28. The execution order of each block in FIG. 28 may be changed, some blocks may be omitted, and some blocks may be newly added.
[0475] FIG. 28 is an example of a detailed functional configuration for receiving / decoding point cloud data (i.e., PCC data). When a bitstream is input, the receiving device according to the embodiments can process the bitstream for location information and the bitstream for attribute information by distinguishing them. At this time, the sub-bitstream classifier (41010) can transmit the bitstream to an appropriate decoder based on the information in the bitstream header. Alternatively, the layer required by the receiver can be selected during this process. Depending on the characteristics of the data, the classified bitstream can be restored into geometry data and attribute data in the geometry decoder (41006) and attribute decoder (41008), respectively, and then converted into a format for final output in the renderer (41009).
[0476] When different types of geometry bitstreams are included, each bitstream can be decoded separately through a bitstream splitter (41014) as shown below. In an embodiment of the present invention, an octree coding-based arithmetic entropy-coded bitstream and a direct-coded bitstream can be distinguished and processed in a geometry decoder (41006). At this time, separation can be performed based on information regarding bitstream type, bitstream_offset, bitstream_length, and bitstream direction. For the separated bitstreams, a process of attaching (or connecting) bitstream segments of the same type can be performed by a bitstream segment connector (41016). This can be included as a process for processing bitstreams separated by layer-groups into a continuous bitstream, and bitstreams can be sorted in order based on layer-group information. If the bitstream is capable of parallel processing, it can be processed in the decoder without a concatenation process.
[0477] The receiver (41002) can receive the bitstream.
[0478] The demultiplexer (41004) can output point cloud data and metadata (signaling information) included in the bitstream.
[0479] The sub-bitstream classifier (41010) can select a slice, split the bitstream, and concatenate the bitstream segments of the octree-coded geometry bitstream and the direct-coded geometry bitstream.
[0480] The metadata parser (41005) can provide information about slices and / or layer groups.
[0481] The slice selector (41012) can select one or more slices included in the bitstream.
[0482] The bitstream splitter (41014) can split the geometry bitstream. The geometry data can be encoded based on an octree and / or coded directly.
[0483] A bitstream segment concatenation (41016) can concatenate octree-coded geometry bitstreams and direct-coded geometry bitstreams according to their encoding types. For layer group-based geometry bitstreams, it can concatenate bitstream segments containing multiple groups / subgroups related to the decoding area.
[0484] The geometry decoder (41006) can decode the geometry bitstream and output geometry data.
[0485] The attribute decoder (41008) can decode attribute data included in the selected slice.
[0486] The renderer (41009) can render point cloud data based on geometry data and / or attributable data.
[0487] FIG. 29 illustrates a method for receiving point cloud data according to embodiments.
[0488] FIG. 29 illustrates the operation of the sub-bitstream classifier (41010) shown in FIG. 28 in more detail.
[0489] A receiving device receives data in slice units, and a metadata parser transmits parameter set information such as SPS, GPS, APS, and TPS (e.g., information about segmented (separated) slices and / or information related to direct coding). Based on the transmitted information, it can be determined whether scalability is possible. If scalability is possible, a slice structure for scalable transmission is identified as shown in FIG. 29 (42011). First, a geometry slice structure can be identified based on information such as num_scalable_layers, scalable_layer_id, tree_depth_start, tree_depth_end, node_size, num_nodes, num_slices_in_scalable_layer, and slice_id transmitted via GPS.
[0490] If the value of the aligned_slice_structure_enabled_flag field is 1 (42017), the attribute slice structure can be identified in the same way (for example, if geometry is octree-based, attributes are encoded based on scalable LoD or scalable RAHT, and geometry / attribute slice pairs generated through the same slice partitioning have the same number of nodes for the same octree layer).
[0491] If the structure is identical, the range of geometry slice IDs is determined based on the target scalable layer, the range of attribute slice IDs is determined through slice_id_offset, and geometry / attribute slices are selected according to the determined range (42012-42014, 42018, 42019).
[0492] If aligned_slice_structure_enabled_flag = 0, the attribute slice structure is identified separately based on information such as num_scalable_layers, scalable_layer_id, tree_depth_start, tree_depth_end, node_size, num_nodes, num_slices_in_scalable_layer, and slice_id transmitted through the APS, and the range of attribute slice id required according to the scalable purpose can be limited, and based on this, the required slice can be selected through each slice id before reconstruction (42020-42021, 42019). The geometry / attribute slice selected through the above process is transmitted as an input to a receiving device according to the embodiments.
[0493] In the above description, the decoding process based on the slice structure was explained based on scalable transmission or the receiver's scalable selection, but it can also be used for non-scalable processes by selecting the entire slice and omitting the ranging geom / attr slice id process when scalable_transmission_enabled_flag is 0. In this case, information about the preceding slice (e.g., a slice belonging to a higher layer or a slice specified via ref_slice_id) can be used through slice structure information transmitted via parameter sets such as SPS, GPS, APS, TPS, etc. (e.g., information about segmented (separated) slices and / or information related to direct coding).
[0494] In cases where different types of geometry bitstreams exist, all slices included within the range for different bitstreams can be selected during the slice selection process. If different types of bitstreams are included within a single slice, each bitstream can be separated based on offset and length information, and the separated bitstreams can be rearranged according to the layer group order for decoding.
[0495] FIG. 30 illustrates a layer group-based point cloud data encoding method according to embodiments.
[0496] An encoder according to the embodiments includes the flowchart of FIG. 30. When point cloud data is input, a layer group structure is configured and relevant parameters are obtained. Based on the layer group structure (or already by external input), reference relationships between subgroups are established. Encoding can be performed based on the layer group structure and the reference structure. For each subgroup / slice, whether it is used as a reference is checked, and if it is used, context_reference_indication_flag can be set to 1. If it is not used, context_reference_indication_flag can be set to 0. Whether it is used as a reference can be determined after encoding and can be obtained directly through the reference structure. If it is used as a reference, the number of times the corresponding slice / subgroup is used as a reference can be signaled as the number of subsequent data units (num_subsequent_data_units). Additionally, if a specific list of slices / subgroups used as references is passed, the subsequent_subgroup_list_present_flag can be set to 1 and the subsequent data unit list can be passed as the layer-group index and subgroup index. The parameters required for decoding can be included in the data unit header, and the encoded compressed bitstream can be included in the data unit to generate the bitstream for each slice, and this process can be performed for every data unit / slice / subgroup.
[0497] FIG. 31 illustrates a layer group-based point cloud data decoding method according to embodiments.
[0498] The decoding method of Fig. 31 can follow the reverse process of the encoding method of Fig. 30.
[0499] The decoder can prepare for decoding by analyzing the data unit header for each slice. At this time, the context state used for decoding can be retrieved from the context buffer using the reference layer group ID (ref_layer_group_id) and reference subgroup ID (ref_subgroup_id). After initializing the decoder based on the retrieved context state, decoding is performed. For new context states generated during the decoding process, whether to store them can be determined by the context_reference_indication_flag included in the data unit header. If used for decoding a subsequent slice / subgroup / data unit, context_reference_indication_flag is signaled as 1. In this case, to manage the memory of the corresponding context state, the target number of the subsequent data unit can be set to the number of subsequent data units (num_subsequent_data_units). Additionally, if subsequent_subgroup_list_present_flag = 1, the layer group index and subgroup index of the subsequent data unit can be stored in the list of subsequent data units.
[0500] For a used context state, the context state counter and the list of used data units can be updated. In this case, if the target number of subsequent data units of the context state is the same as the counter, or if the number of subsequent data units in the given list (list_given) is the same as the updated list of used data units (lilst_updated), the context state can be released from memory in the counter buffer.
[0501] The context buffer management method can be applied equally to encoders as well as decoders. It can also be applied not only to slices based on layer-group slicing but also to cases where context buffers are referenced within a regular frame or between frame slices.
[0502] FIG. 32 (a) to (c) illustrates a context buffer management method according to embodiments.
[0503] In Fine Granularity Slicing (FGS), context inheritance between slices is used to mitigate coding loss caused by discontinuities between adjacent nodes or coding layers. However, as the number of slices or groups of layers increases, the number of context states in memory (or referred to as the context buffer) increases. To help decoders manage the context buffer, an indication of future use of the current context by subsequent slices is signaled.
[0504] Referring to FIGS. 32 (a) through (c), the context buffer management method of the current layer group slicing method is illustrated. When the bitstream of a slice is decoded and the context reference indication flag (context_reference_indication_flag) is enabled, the context state of the decoder output is stored in the context buffer as shown in FIG. 32 (a). When context inheritance is used, the context state of the following slice can be initialized to one of the stored context states of the previous subgroups indicated by the reference layer group ID (ref_layer_group_id) and the reference subgroup ID (ref_subgroup_id) (see FIG. 32 (b)). In FIG. 32 (c), the context state of the slice belonging to the last layer group is initialized by the stored context state of the parent slice. However, since the context_reference_indication_flag is disabled, the output context state is not stored in the context buffer.
[0505] Using the context_reference_indication_flag allows you to reduce the overall size of the context buffer by selecting context states known to be used by following slices. However, there is a limitation in that decoders cannot know the release time of each stored context. Therefore, all context states must be stored in the context buffer until the decoding of all subgroups is complete.
[0506] Referring to FIG. 32 (a), the first slice, FGS 0 (0,0), represents subgroup 0 of layer group 0 of the point cloud data. When encoding (or decoding) starting from FGS 0 (0,0), the context state (0, 0) for FGS 0 is stored in the context buffer for subsequent FGSs.
[0507] Referring to FIG. 32(b), the third slice, FGS 2(1,1), can be encoded (or decoded) sequentially. FGS 2(1,1) represents subgroup 1 of layer group 1 of point cloud data. FGS 2(1,1) may be a subgroup belonging to (or dependent on) FGS 0(0,0) (parent-child relationship). Therefore, FGS 2(1,1) can be efficiently encoded (or decoded) by referring to the context state (0, 0) for FGS 0(0,0) stored in the context buffer. Then, the context state (1, 1) is stored in the context buffer for the subsequent slice (FGS).
[0508] The method according to the embodiments can solve this problem as follows through a mechanism for releasing the context state.
[0509] FIG. 33 (a) to (c) illustrates a context buffer management method according to embodiments.
[0510] The method according to the embodiments includes a method for indicating the number of subgroups referencing the current subgroup so that the decoder can know the timing for releasing the stored context state.
[0511] FIGS. 33 (a) through (c) illustrate a method for releasing the context buffer using the proposed signal. Compared to FIGS. 32 (a) through (c), the context buffer has two additional columns to indicate the number of subsequent (i.e., following) subgroups referencing the current subgroup, and to calculate the number of subgroups for which the context state has already been used in subgroup decoding. When the context_reference_indication_flag is enabled, the number of subsequent (i.e., following) subgroups (num_subsequent_subgroups) is signaled, and the corresponding number is stored along with the context state as in FIG. 33 (a). As in FIG. 33 (b), three context states (i.e., (0,0), (1,0), (1,1)) are stored in the context buffer known to be referenced N times or 1 time. For example, when there are N subsequent (i.e., following) subgroups referencing the context state of FGS 0(0, 0), and the context state of FGS 0(0, 0) is referenced and encoded (or decoded) in FGS 2(1, 1), the context state of FGS 0(0, 0) is referenced by FGS 1(1, 0) and FGS 2(1, 1) respectively (i.e., twice), so the value of the counter for the context state (0, 0) becomes 2. As shown in FIG. 33(c), the context state of FGS 2(1, 1) is released after being used in FGS N+2(2, 1). The context states (0, 0) and (1, 0) are released beforehand because there are no longer any subsequent subgroups referencing them. Releasing the context state means removing the context state from the context buffer (or memory).
[0512] FIG. 34 (a) to (c) illustrates a context buffer management method according to embodiments.
[0513] When different orders are used for FGS (which may simply be referred to as slices, sub-slices, etc.), the proposed method can be effective in the same way. Referring to FIGS. 34 (a) through (c), the method according to the embodiments is also applicable to depth-first ordering. For comparison with the breadth-first ordering case of FIGS. 33 (a) through (c), the names of each slice are the same, and the passing order is changed from FGS 0, FGS 1, FGS 2, FGS 3, … to FGS 0, FGS 1, FGS N+1, FGS 2. As shown in FIG. 34 (b), when the decoding of FGS2 is complete, the output context state (1, 1) is stored in the context buffer and the number of subsequent subgroups (e.g., 1) is stored. As shown in Fig. 34c), when the next slice, FGS N+2, is decoded, the context state is initialized by the context state (1, 1), and the corresponding counter in the context buffer is incremented by 1. Since the stored value of the count in the counter is the same as the number of subsequent subgroups, the context state (1, 1) can be released from this point onward. In this example, the maximum number of context states in the context buffer is 2, which is the number of layer groups minus 1.
[0514] At this time, the size of the context buffer can be anticipated in advance at the receiver. For example, there are N layer groups, and assuming the size of the subgroup within the nth layer group is S[n], a case can be considered where the parent subgroup (or upper subgroup) is referenced.
[0515] In this case, the number of context states to be stored in memory (or called the context buffer) can be estimated as follows. That is, since the maximum number N-1 is not referenced in subsequent steps, the addition process is performed only up to N-2.
[0516]
[0517] As an extreme opposite case, consider the scenario where the initial slice / subgroup is referenced; in this case, the number of context states to be stored in memory can be estimated as follows.
[0518] (number of subgroup in the root layer-group)=1
[0519] For example, the number of subgroups within the root layer group can be 1.
[0520] Referencing a parent subgroup may be the case that uses the most context state, while referencing the first slice / subgroup (e.g., root subgroup) may be the case that uses the least context state. Therefore, when using layer group slicing with different reference relationships, one may consider that it lies between the maximum and minimum values.
[0521] When using the method according to the embodiments to dynamically release memory, space may be required to store fewer context states than is required. For example, if FGS generated by layer group slicing are generated and / or delivered based on breadth-first order and reference only parent subgroups, context states belonging to the parent layer group may be released when coding for a single layer group is finished. Thus, the number of context states to be stored in memory can be estimated as follows.
[0522] MAX(S[0],S[1],…,S[N-2])
[0523] If FGS generated by layer group slicing is generated / transmitted in depth-first order, the context memory of the parent subgroup can be released (deleted) when the coding of the child subgroup is finished by the method according to the embodiments. In this case, only the context of the parent subgroup with remaining children needs to be stored, and since leaf layer groups are excluded, the number of context states to be stored in memory can be estimated as follows.
[0524] N-1
[0525] If it is necessary to estimate the memory size for storing the context buffer (or context state) at the receiver, information related to this (slice coding order type, depth first / breadth first, number of layer-groups, number of subgroups in each layer-group, context reference method, parent reference / root reference, etc.) can be provided, and the number of context states can be estimated as in the method presented above.
[0526] In the following, based on the context memory management mechanism according to the embodiments, the number of subsequent data units (num_subsequent_data_units), which is new signaling information, is signaled to the geometry data unit header and the dependent geometry data unit header.
[0527] An encoder according to the embodiments encodes point cloud data and generates related signaling information. It then generates and transmits a bitstream containing the encoded point cloud data and parameter information. In the reverse process, a decoder according to the embodiments receives the bitstream, parses the parameter information contained in the bitstream, and decodes the point cloud data based on the parameter information. FIGS. 35 to 37 illustrate the syntax of the parameter information contained in the bitstream.
[0528] FIG. 35 shows an example of the syntax structure of a geometry data unit header according to embodiments.
[0529] FIG. 36 shows an example of the syntax structure of a dependent geometry data unit header according to embodiments.
[0530] FIG. 37 shows an example of the syntax structure of an attribute data unit header according to embodiments.
[0531] FIG. 38 shows an example of the syntax structure of a dependent attribute data unit header according to embodiments.
[0532] In FIGS. 35 through 38, the number of subsequent data units (num_subsequent_data_units) represents the number of subsequent dependent data units that reference the current data unit or dependent data unit. A data unit may be a slice. A slice may be an FGS for a subgroup within a layer group according to the embodiments.
[0533] The geometry parameter set ID (dgsh_geometry_parameter_set_id) is a geometry parameter set identifier. It may be information identifying a parameter set for a dependent geometry data unit.
[0534] The slice ID (dgsh_slice_id) is a slice identifier. It can be a slice identifier associated with a dependent geometry data unit.
[0535] The layer group ID (layer_group_id) is a layer group identifier. It may be the identifier of the layer group associated with the dependent geometry data unit.
[0536] The subgroup ID (subgroup_id) is a subgroup identifier. It can be the identifier of a subgroup associated with a dependent geometry data unit.
[0537] The subgroup bounding box location (subgroup_bbox_origin[i]) represents the origin location of the subgroup bounding box.
[0538] The subgroup bounding box size (subgroup_bbox_size[i]) represents the size of the subgroup bounding box.
[0539] The reference layer group ID (ref_layer_group_id) is a reference layer group identifier. It may be the identifier of the layer group referenced by the dependent geometry data unit.
[0540] The referenced subgroup ID (ref_subgroup_id) is the referenced subgroup identifier. It can be the identifier of the subgroup referenced by the layer group for the layer group identifier.
[0541] The context reference indication flag (context_reference_indication_flag) is a flag that indicates whether a context is being referenced.
[0542] The reference count (num_referenced) indicates the number of times it is referenced.
[0543] Next, we will explain how to control the context buffer (or memory) considering partial decoding.
[0544] The context memory control method described in FIGS. 33 and 34 above can efficiently manage the context buffer when receiving a bitstream composed of multiple slices and decoding all received slices. However, in the case of partial decoding (i.e., when only some slices are decoded when there is a region of interest or a resolution of interest), the reference count (num_subsequent_data_units) transmitted by the encoder may not be fully filled. In this disclosure, slice, subgroup, and data unit may be used interchangeably with the same meaning.
[0545] For example, considering the case of a layer group structure divided into three layer groups, the decoder of the receiving device may decide not to use the last layer group. In this case, if the number referenced in the second layer group is known, there is no need to store the context (or context state or context information) in the context buffer to decode the last layer group.
[0546] To this end, the present disclosure may enable context memory release in a decoder after a specified number of times by signaling the number of times a context is used (or referenced) in a specified unit (e.g., data unit). Additionally, the present disclosure may enable context memory release in a decoder after a specified number of times by signaling per layer group when signaling the number of times a context is used (i.e., referenced) in a specified unit (e.g., data unit). For example, if there are two or more layer groups referencing the current data unit, the number of subsequent data units (i.e., data units referencing the current data unit) may be signaled per layer group. In the present disclosure, a data unit may be a subgroup or a slice.
[0547] In this case, the number of times a context is used (i.e., referenced) for each data unit can have a relationship with the total number of times a context is used (i.e., referenced) as shown in the following formula: num_subsequent_data_units = i in subsequeny layer-groupsnum_sdu_per_layer_group[i]
[0548] In the above formula, subsequent layer-groups represents layer groups containing subgroups (i.e., data units) that reference the context of the current data unit, and num_sdu_per_layer_group represents the number of times subgroups belonging to each layer group reference the context of the current subgroup. That is, num_sdu_per_layer_group represents the number of data units (or subgroups) belonging to each layer group that reference the context of the current data unit of the current layer group.
[0549] If we assume that FGS1 (1,0) of layer-group #1 in Fig. 39 is the current data unit, then the subsequent layer-groups are layer-group #2 and layer-group #3. And, in layer-group #2, the number of data units referencing the context of the current data unit FGS1 (1,0) is 1. Also, in layer-group #3, the number of data units referencing the context of the current data unit FGS1 (1,0) is 1.
[0550] FIG. 39 is a diagram showing an example of a layer group structure considering partial decoding according to embodiments. That is, for partial decoding in terms of layer groups, a context reference structure such as FIG. 39 can be considered. More specifically, a subgroup belonging to layer-group #1 may refer to the context state of a subgroup of layer-group #0 (i.e., root subgroup), a subgroup belonging to layer-group #2 may refer to the context state of a parent subgroup belonging to layer-group #1, and a subgroup belonging to layer-group #3 may refer to the context state of a grandparent subgroup belonging to layer-group #1.
[0551] At this time, in a layer group structure divided into four layer groups as shown in Fig. 39, the decoder of the receiving device can perform partial decoding that does not use the last layer group (i.e., layer-group #3).
[0552] According to the embodiments, when a specific layer group, e.g., layer-group #3, is skipped and partial decoding is performed, the context buffer stores num_subsequent_subgroups (or num_subsequent_data_units) for each layer group and can determine whether the context reference count is satisfied for each subsequent layer group. In an example such as FIG. 39, since layer-group #3 is skipped, the counters in the context buffer do not use or need to store num_subsequent_subgroups (or num_subsequent_data_units) for layer-group #3.
[0553] FIG. 40 (a) to (d) illustrates a context buffer management method according to embodiments.
[0554] In FIG. 40 (a) to (d), the context buffer is managed per data unit. According to the embodiments, the context buffer may be divided into three storage areas per data unit, for example, a context state storage area (51010), a reference count information storage area (51020), and a counter information storage area (51030). At this time, the reference count information storage area (51020) and the counter information storage area (51030) may each be further divided into sub-storage areas corresponding to the number of layer groups within the layer group structure.
[0555] For example, the layer group structure of FIG. 39 includes four layer groups (i.e., layer-group #0-layer-group #3), so the reference count information storage area (51020) and the counter information storage area (51030) in the context buffer can each be divided into four sub-storage areas.
[0556] Referring to FIG. 40(a), the context state storage area (51010) of the context buffer stores the context state of FGS0 (0,0) (i.e., context state (0,0)) after FGS0 (0,0) is encoded or decoded. Then, the reference count information storage area (51020) stores the number of data units (i.e., subgroups) that reference the context state (0,0) in each layer group, i.e., four layer groups (layer-group #0-layer-group #3). At this time, since the context state (0,0) references only the N data units of layer-group #0 (i.e., FGS1 (1,0) ~ FGSN (1,N-1)) and is not referenced by other layer groups, the value {- | N | 0 | 0} is stored in the reference count information storage area (51020). Here, "-" means no reference. That is, for the context state (0,0), it indicates that there is no reference in layer-group #0, it is referenced N times in layer-group #1, and it is not referenced at all in layer-group #2 and layer-group #3. Also, in FIG. 40(a), since the context state (0,0) has not yet been referenced, the value {- | 0 | 0 | 0} is stored in the counter information storage area (51030).
[0557] That is, in FIG. 40 (a), when FGS 0 / subgroup (0,0) is decoded (FGS 0(0,0)), the context state storage area (51010) of the context buffer stores the context state (0,0), and the reference count information storage area (51020) can store num_subsequent_subgroups (corresponding to num_subsequent_data_units in the signaling information of FIG. 41 to 45) transmitted through the data unit header for each subsequent layer group. In this example, the value {- | N | 0 | 0} can be stored.
[0558] In other words, when context_reference_indication_flag is enabled, the number of subsequent (i.e., following) subgroups (num_subsequent_subgroups or num_subsequent_data_units) is signaled, and as shown in FIG. 40 (a), the context state storage area (51010) stores the context state (0,0), and the corresponding number for each layer group is stored in the reference count information storage area (51020). At this time, the counter information storage area (51030) is set to {- | 0 | 0 | 0}.
[0559] These rules apply equally to other data units.
[0560] Taking FGS1 (1,0) in Fig. 40 (b) as an example, the context state storage area (51010) of the context buffer stores the context state of FGS1 (1,0) (i.e., context state (1,0)) after FGS1 (1,0) is encoded or decoded. At this time, since the context state (1,0) is referenced by FGS N+1 (2,0) of layer-group #2 and FGS 2N+1 (3,0) of layer-group #3, the reference count information storage area (51020) stores the value {- | - | 1 | 1}. That is, for the context state (1,0), it indicates that there is no reference in layer-group #0 and layer-group #1, that it is referenced once in layer-group #2, and that it is referenced once in layer-group #3. And, in FIG. 40(b), since the context state (1,0) has not yet been referenced, the value {- | 0 | 0 | 0} is stored in the counter information storage area (51030). At this time, since the context state (0,0) has been referenced in FGS1 (1,0), the value {- | 1 | 0 | 0} is stored in the counter information storage area (51030) corresponding to FGS0 (0,0).
[0561] That is, in FIG. 40 (b), when FGS 1 / subgroup (1,0) is decoded (FGS 1(1,0)), since context_reference_indication_flag is 1, the context state storage area (51010) of the context buffer stores the context state (1,0), and the reference count information storage area (51020) can store num_subsequent_subgroups (corresponding to num_subsequent_data_units in the signaling information of FIG. 41 to 45) transmitted through the data unit header for each subsequent layer group. In this example, the value {- | - | 1 | 1} can be stored.
[0562] In the present disclosure, the loading process of context states (i.e., context information) can be performed by referring to the ref_layer_group_id and ref_subgroup_id parameters. In FIG. 40(b), context information (context states(0,0)) of FGS 0 (0,0) can be loaded to process (encode or decode) FGS 1 (1,0). Then, context states (0,0) are used to process FGS 1 (1,0), and the counter value is increased by 1. That is, the value {- | 1 | 0 | 0} is stored in the counter information storage area (51030) corresponding to FGS 0 (0,0).
[0563] If the data unit processing order is a breadth-first search method, then after FGS 1 (1,0) is processed, the next data unit of the same layer group (i.e., layer-group #1) (i.e., FGS 2 (1,1)) is processed. In contrast, if the data unit processing order is a depth-first search method, then after FGS 1 (1,0) is processed, the first data unit of a different layer group (i.e., layer-group #2) (i.e., FGS N+1 (2,0)) is processed. FIGS. 40(a) to 40(d) illustrate an example of a depth-first search method. This is an example embodiment, and the present disclosure may also be applied to a breadth-first search method.
[0564] Here, breadth-first search is a method that codes / decodes subgroups belonging to the same layer group and then codes / decodes subgroups belonging to the child layer group. In contrast, depth-first search is a method that reaches the subgroup corresponding to the maximum depth and then codes / decodes the children belonging to the same parent first.
[0565] The decoder can prepare for decoding by analyzing the data unit header for each data unit (i.e., subgroup or slice). At this time, the context state used for decoding can be retrieved from the context buffer using the reference layer group ID (ref_layer_group_id) and the reference subgroup ID (ref_subgroup_id). Based on the retrieved context state, the decoder initializes the context buffer for the current data unit and then performs decoding.
[0566] When decoding FGS N+1 / subgroup (2, 0) in Fig. 40 (b) (FGS N+1(2,0)), the encoder / decoder can initialize the context state of FGS N+1 / subgroup (2, 0) based on the context state (1,0) corresponding to ref_layer_group_id=0 and ref_subgroup_id=0. That is, the context state can be initialized based on the context state (1,0). At this time, since the context state (1, 0) is referenced in FGS N+1(2,0) of layer-group #2, the value of the counter information storage area (51030) is changed from { - | 0 | 0 | 0} to { - | 0 | 1 | 0}. And, considering the skip layer group, since no reference will be made in layer-group #3, the reference count for each layer-group stored in the context buffer matches the reference count of the counter, so the context state (1, 0) can be released (i.e., deleted). That is, since the value (=1) corresponding to layer-group #2 stored in the reference count information storage area (51020) of the context buffer corresponding to FGS (1,0) and the value (=1) corresponding to layer-group #2 stored in the counter information storage area (51030) are identical, the context state (1,0), reference count information, and counter information of the context buffer corresponding to FGS (1,0) are deleted from the context buffer. In this way, the context state of FGS 1 (1, 0) and related information (i.e., reference count information and count information) are released after being used in FGS N+1 (2, 0). Releasing the context state means deleting the context state from the context buffer (or memory).
[0567] Additionally, since it is assumed that layer-group #3 is skipped during decoding, FGS N+2 (2,1) belongs to the last layer group, and therefore the context state of FGS N+2 (2,1) is not stored in the context buffer. As another example, even if it is assumed that layer-group #3 is not skipped, there are no data units referencing layer-group #2 in FIG. 39, so the context state of FGS N+2 (2,1) is not stored in the context buffer. That is, if the current data unit is a data unit of the last layer group due to the skipped layer group, and / or there are no data units referencing the current data unit, the context state of the encoded or decoded current data unit is not stored in the context buffer.
[0568] As shown in FIG. 40 (c), when FGS 2 / subgroup (1,1) is decoded (FGS 2(1,1)), since context_reference_indication_flag is 1, the context state storage area (51010) of the context buffer stores the context state (1,1), and the reference count information storage area (51020) can store num_subsequent_subgroups (or num_subsequent_data_units) transmitted through the data unit header for each subsequent layer group. In this example, the value {- | - | 1 | 1} can be stored. For a detailed description of FIG. 40 (c), refer to FIG. 40 (b).
[0569] As shown in Fig. 40 (d), when decoding FGS 2N / subgroup (2, N-1) (FGS 2N(2, N-1)), the encoder / decoder can initialize the context state of FGS 2N / subgroup (2, N-1) based on the context state (1, N-1) corresponding to ref_layer_group_id=1 and ref_subgroup_id=N-1. That is, the context state of FGS 2N / subgroup (2, N-1) can be initialized based on the context state (1, N-1) identified by ref_layer_group_id and ref_subgroup_id.
[0570] At this time, since context state (1, N-1) is referenced in FGS 2N(2, N-1) of layer-group #2, the value of the counter information storage area (51030) is changed from { - | 0 | 0 | 0} to { - | 0 | 1 | 0}. And, considering the skip layer group, since no reference will be made in layer-group #3, the number of references for each layer-group stored in the context buffer matches the number of references in the counter, so context state (1, N-1) can be released (i.e. deleted). That is, since the value corresponding to layer-group #2 (=1) stored in the reference count information storage area (51020) of the context buffer corresponding to FGS (1,N-1) and the value corresponding to layer-group #2 (=1) stored in the counter information storage area (51030) are identical, the context state (1,N-1), reference count information, and counter information of the context buffer corresponding to FGS (1,N-1) are deleted from the context buffer. In this way, the context state and related information (i.e., reference count information and count information) of FGS N (1,N-1) are released after being used in FGS 2N (2,N-1).
[0571] Additionally, since it is assumed that layer-group #3 is skipped during decoding, FGS 2N (2,N-1) belongs to the last layer group, and therefore the context state of FGS 2N (2,N-1) is not stored in the context buffer. As another example, even if it is assumed that layer-group #3 is not skipped, there are no data units referencing layer-group #2 in FIG. 39, so the context state of FGS 2N (2,N-1) is not stored in the context buffer. That is, if the current data unit is a data unit of the last layer group due to the skipped layer group, and / or there are no data units referencing the current data unit, the context state of the encoded or decoded current data unit is not stored in the context buffer.
[0572] In the present disclosure, the group of layers for which decoding is skipped can be determined by the decoder depending on the application.
[0573] As explained so far, context memory (i.e., context buffer) can be early released in partial decoding situations.
[0574] Next, we will explain the signaling information required to release context memory (i.e., the context buffer) in partial decoding situations.
[0575] According to embodiments, the present disclosure may signal and transmit the entire context reference structure. In one embodiment, the encoder of the transmitting device of the present disclosure signals the entire context reference structure via SPS and transmits it to the decoder of the receiving device. In this case, actual relationships that allow the decoder to identify the reference structure, such as information about the layer group structure and reference subgroup identification information / number of subsequent subgroups (reference subgroup ID / number of subsequent subgroups), may be transmitted. Furthermore, the decoder may perform an early release of the context (or context state) in the case of partial decoding based on the entire identified reference structure.
[0576] According to embodiments, the present disclosure may signal subsequent subgroup identification information (i.e., subsequent subgroup ID) and transmit it directly to a decoder of a receiving device. In one embodiment, the encoder of the transmitting device of the present disclosure transmits the subsequent subgroup identification information (i.e., subsequent subgroup ID) directly to the decoder of the receiving device through a data unit header. In this case, a context reference structure is transmitted in units of data units. Then, the decoder checks whether a request is received from the corresponding subgroup ID in the buffer, and if all corresponding contexts are used in the subgroups used in partial decoding, an early release may be performed.
[0577] According to embodiments, the present disclosure may divide and transmit information on the number of subsequent subgroups. In one embodiment, the encoder of the transmitting device of the present disclosure divides and signals information on the number of subsequent subgroups through a data unit header and transmits it to the decoder of the receiving device. For example, for the number of subsequent subgroups N, it may be divided and transmitted as N = N1 (number referenced in layer-group 1) + N2 (number referenced in layer-group 2). Furthermore, the decoder of the receiving device that does not decode layer-group #2 may release the context early when the context has been used N1 times. In addition, in the case of full decoding, the context may be released after being used N1 + N2 times. Here, subsequent subgroups may be subsequent data units.
[0578] FIGS. 41 to 45 illustrate the syntax of parameter information included in a bitstream. Specifically, FIGS. 41 to 45 describe a method for transmitting the number of times each data unit is referenced within each layer group and subgroup identification information (subgroup id). In the case where a plurality of child subgroups exist in a single parent subgroup, the present disclosure adds bounding box information of the referenced subgroup, etc., thereby transmitting the number of context references more accurately in the case of partial decoding based on ROI, so that the context memory (i.e., context buffer) can be released at the correct time in the decoder.
[0579] FIG. 41 shows an example of the syntax structure of a sequence parameter set (SPS) according to the embodiments.
[0580] In FIG. 41, if the value of layer_group_enabled_flag is 1, it specifies that the geometry bitstream of a slice is contained in multiple slices which is matched to a group of coding layers or its subgroup. If the value of layer_group_enabled_flag is 0, it specifies that the geometry bitstream is contained in a single slice.
[0581] Adding 1 to num_layer_groups_minus1 specifies the number of layer-groups. In this case, the layer-group represents a group of consecutive tree layers that are part of the geometry coding tree structure. num_layer_groups_minus1 is in the range from 0 to the number of coding tree layers.
[0582] layer_group_id specifies the indicator of a layer-group of a slice. The range of layer_group_id is between 0 and num_layer_groups_minus1.
[0583] num_layers_minus1 plus 1 specifies the number of coding layers contained in the i-th layer-group. The total number of layer-groups can be derived by adding all (num_layers_minus1[i] + 1), where i has a value between 0 and num_layer_groups_minus1.
[0584] If the value of subgroup_enabled_flag is 1, it indicates that the i-th layer-group is divided into two or more subgroups. In this case, the set of points in the subgroups of a layer-group is identical to the set of points in the layer-group. If the value of subgroup_enabled_flag for the i-th layer-group is 1, then the subgroup_enabled_flag for the j-th subgroup becomes 1, where j is greater than or equal to i. If the value of subgroup_enabled_flag is 0, it indicates that the current layer-group is not subdivided into multiple subgroups and is contained in a single slice.
[0585] Add 1 to subgroup_bbox_origin_bits_minus1 to represent the length of the syntax elements subgroup_bbox_origin in bits.
[0586] Add subgroup_bbox_size_bits_minus1 to represent the length of the syntax elements subgroup_bbox_size in bits.
[0587] Add 1 to num_subgroups_minus1 to represent the number of subgroups within the i-th layer group.
[0588] num_subsequent_data_units represents the number of subsequent dependent data units that reference the j-th data unit or the j-th dependent data unit of the i-th layer group.
[0589] subsequent_data_unit_id represents the index of the k-th subsequent data unit that references the j-th data unit or the j-th dependent data unit of the i-th layer group.
[0590] FIG. 42 shows an example of the syntax structure of a geometry data unit header according to embodiments.
[0591] FIG. 43 shows an example of the syntax structure of a dependent geometry data unit header according to embodiments.
[0592] FIG. 44 shows an example of the syntax structure of an attribute data unit header according to embodiments.
[0593] FIG. 45 shows an example of the syntax structure of a dependent attribute data unit header according to embodiments.
[0594] In FIGS. 42 through 45, dgsh_geometry_parameter_set_id is an identifier for identifying a geometry parameter set. In one embodiment, dgsh_geometry_parameter_set_id may be information identifying a parameter set for a dependent geometry data unit.
[0595] dgsh_slice_id is a slice identifier. In one embodiment, dgsh_slice_id may be a slice identifier associated with a dependent geometry data unit.
[0596] layer_group_id is a layer group identifier. In one embodiment, layer_group_id may be an identifier of the layer group associated with a dependent geometry data unit.
[0597] subgroup_id is a subgroup identifier. In one embodiment, subgroup_id may be an identifier of a subgroup associated with a dependent geometry data unit.
[0598] subgroup_bbox_origin[i] represents the origin position of the subgroup bounding box.
[0599] subgroup_bbox_size[i] represents the size of the subgroup bounding box.
[0600] ref_layer_group_id is a reference layer group identifier. In one embodiment, ref_layer_group_id may be an identifier of a layer group referenced by a dependent geometry data unit.
[0601] ref_subgroup_id is a referenced subgroup identifier. In one embodiment, ref_subgroup_id may be the identifier of a subgroup that a layer group references for a layer group identifier.
[0602] context_reference_indication_flag is a flag indicating whether a context is referenced. In one embodiment, depending on the value of context_reference_indication_flag, it may be determined whether to store a new context state generated during the decoding process of the corresponding data unit in the context buffer.
[0603] num_subsequent_data_units indicates the number of subsequent dependent data units that reference the current data unit or dependent data unit. A data unit may be a subgroup or a slice. A slice may be an FGS for a subgroup within a layer group according to the embodiments.
[0604] If the value of num_sdu_per_layer_group_present_flag is 1, it indicates that the number of subsequent data units in each layer-group is present. If the value of num_sdu_per_layer_group_present_flag is 0, it indicates that the number of subsequent data units in each layer-group is not present. In other words, num_sdu_per_layer_group_present_flag is a flag used to indicate whether to pass the referenced number and data_unit_id per layer-group. If the value of num_sdu_per_layer_group_present_flag is 0, only num_subsequent_data_units is passed.
[0605] If the value of sdu_present_flag is 1, it indicates that the successor data unit of the current data unit exists in the i-th layer group. If the value of sdu_present_flag is 0, it indicates that the successor data unit of the current data unit does not exist in the i-th layer group.
[0606] num_sdu_per_layer_group indicates the number of subsequent data units within the i-th layer group referenced by the current data unit or dependent data unit.
[0607] subsequent_data_unit_id represents the index of the subsequent data unit referenced by the current data unit or dependent data unit.
[0608] For syntax elements not described in FIGS. 42 to 45, refer to the descriptions in FIGS. 23 to 26 or FIGS. 35 to 38.
[0609] And, taking Fig. 39 as an example, the geometry data unit to which the geometry_data_unit_header of Fig. 42 is applied corresponds to the data unit included in layer-group #0, and the dependent data unit to which the dependent_geometry_data_unit_header of Fig. 43 is applied corresponds to the data unit included in layer-group #1 to layer-group #3.
[0610] FIG. 46 (a) to (c) shows another example of a context buffer management method according to the embodiments.
[0611] Fine granularity slicing (FGS) can mitigate coding loss caused by discontinuity between adjacent nodes or coding layers by using context inheritance between slices. However, as the number of slices or layer groups increases, the number of context states in memory (i.e., context buffer) may increase.
[0612] As shown in FIG. 46, to help the decoder manage the context buffer, signaling information can be transmitted so that the current context can be used later in subsequent slices. By doing so, the context state known to be used in subsequent slices can be saved, thereby reducing the overall size of the context buffer.
[0613] Additionally, stored context states may be released (referred to as release or deletion) using information regarding the number of subgroups referencing the current subgroup. However, when a partial layer or partial region of the bitstream is decoded, the number of subsequent subgroups outside the Region of Interest (ROI) cannot be counted, and thus the number of coded subsequent subgroups may not reach the signaled value (e.g., the number of signaled subsequent subgroups), and the said stored contexts may not be released (i.e., removed).
[0614] Since Figures 46 (a) to (c) are identical to Figures 34 (a) to (c), a detailed description will be omitted here and refer to Figures 34 (a) to (c).
[0615] FIGS. 47 (a) to (c) illustrates another example of a context buffer management method according to embodiments. In particular, FIGS. 47 (a) to (c) is intended to solve problems that may occur when a partial layer or partial region of a bitstream is decoded.
[0616] The method according to the embodiments can solve the aforementioned issue by managing a subsequent subgroup list that overlaps with the ROI for each layer group to consider partial decoding cases, as shown in FIG. 47 (a) to (c).
[0617] The method according to the embodiments can release the context buffer using the method proposed in FIG. 47 (a) to (c).
[0618] In FIG. 47 (a) to (c), the context buffer is managed per data unit (e.g., FGS). According to embodiments, the context buffer may be divided into three storage areas per data unit, for example, a first storage area (60010) for storing the context state, a second storage area (60030) for storing a subsequent subgroup list (list) that references the current subgroup, and a third storage area (60050) for storing a related coded subgroup list using the context state of the current subgroup.
[0619] That is, looking at FIG. 47 (a) through (c), there are two lists in the context buffer. The first list is a list of subsequent subgroups that reference the current subgroup, and the second list is a list of coded subgroups related to each context, that is, a list of coded subgroups related to each context, using the context state of the current subgroup. The first list may be provided for each data unit, and the second list may be updated in the decoder. The present disclosure may refer to the first list as List A and the second list as List B.
[0620] According to the embodiments, when an ROI is provided for partial regional decoding, the region of subgroups in the list is compared with the ROI, and the subgroups with the ROI overlapped region are retained in the list. In other words, they are kept in the list. Additionally, when the number of skipped layer groups is set for partial layer group decoding, a list of subsequent subgroups within the unskipped layer groups is retained in that list.
[0621] In FIG. 47(a), for example, the ROI is indicated by the line (60000) on the FGS having one skipped layer group. In FIG. 47(a) through (c), it is assumed that layer group 2 is skipped. Using this information, the FGS decoder of the receiving device decodes FGS 0 for layer group 0, FGS 1 and 2 for layer group 1, and does not decode anything for layer group 2. Based on this, a list of selected subsequent subgroups (i.e., List A) is generated as a subset of the given list of subsequent subgroups.
[0622] Referring to Fig. 47(b), when FGS 1 is decoded, the context state is initialized to the stored context state of FGS 0 (context state (0, 0)) and updated to a list of coded subsequent subgroups (i.e., list B).
[0623] Referring to FIG. 47(c), when FGS 2 is decoded, List B is updated to FGS 2(1, 1), and since List B and List A are identical, the context state (0, 0) can be released under the condition that the decoder does not decode a subgroup outside the ROI.
[0624] In FIG. 47 (a) to (c), the first list (List A) is a list of subsequent subgroups that reference the current subgroup, and the second list (List B) is a list of related coded subgroups using the context state of the current subgroup. While decoding the FGS of the layer group and subgroup related to the ROI for partial decoding, List B is updated according to the coded subgroup, and when the subsequent subgroup list (List A) and the updated List B become identical, the ROI-related partial decoding is completed, so there is no need to continue storing the context state stored in the context buffer. Therefore, the decoder has the effect of efficiently controlling the context buffer by releasing the context buffer.
[0625] The encoding / decoding method according to the embodiments may store a context state in a context buffer and a list of subsequent subgroups in a context buffer if the context reference indication flag (context_reference_indication_flag) is true. When encoding (or decoding) FGS 0, a context state (0, 0) is stored in a buffer, and subsequent subgroups (FGS(1, 0), FGS(1, 1) to FSG(1, N-1)) are stored in a buffer. When encoding (or decoding) FGS 1 (1, 0), if the reference layer group ID (ref_layer_group_id) is 0 and the reference subgroup ID (ref_subgroup_id) is 0, the context state (0, 0) stored in the context buffer is loaded. If the context reference indication flag (context_reference_inidcation_flag), ROI overlap area, and skipped layer group are all true, the context state (1, 0) is not stored in a buffer. When encoding (or decoding) FGS 2(1, 1), if the reference layer group ID is 0 and the reference subgroup ID is 0, the context state (0, 0) stored in the buffer is loaded. If the context reference inidation flag (context_reference_inidation_flag), ROI duplicate area, and skipped layer group are all true, the context state (1, 1) is not stored in the buffer.
[0626] The encoding method according to the embodiments can transmit by generating subgroup indexes and bounding box information for each layer group in the data unit header as signaling information (parameter information) based on the proposed context memory management mechanism, including them in the bitstream. The decoding method according to the embodiments can decode point cloud data based on context memory management-related parameter information.
[0627] FIG. 48 shows another example of the syntax structure of a sequence parameter set according to the embodiments.
[0628] If the subsequently_subgroups_info_present_flag is 1, it indicates that additional information for subsequent subgroups is delivered. If subsequently_subgroups_info_present_flag is 0, it indicates that additional information for subsequent subgroups is not delivered.
[0629] FIG. 49 shows another example of the syntax structure of a geometry data unit header included in a bitstream according to embodiments.
[0630] In FIG. 49, the number of subsequent subgroups (num_subsequent_subgroups): represents the number of subsequent dependent data units that reference the current data unit or dependent data unit.
[0631] Subsequently_subgroup_id: Represents the subgroup index of the subsequent dependent data unit of the i-th layer group that references the current data unit or dependent data unit.
[0632] subsequent_subgroup_bbox_origin: Represents the origin of the subsequent subgroup's subgroup bounding box, which references the context state of the current subgroup.
[0633] Subsequently_subgroup_bbox_size: Represents the size of the subgroup bounding box of the subsequent subgroup that references the context state of the current subgroup.
[0634] FIG. 50 shows another example of the syntax structure of a dependent geometry data unit header included in a bitstream according to the embodiments.
[0635] Geometry parameter set ID (dgsh_geometry_parameter_set_id): Represents the geometry parameter set ID referenced by the dependent geometry data unit header.
[0636] Slice ID (dgsh_slice_id): Represents the slice ID of the geometry data unit.
[0637] Layer group ID (layer_group_id): Represents the layer group ID of the geometry data unit.
[0638] Subgroup ID (subgroup_id): Represents the subgroup ID of the geometry data unit.
[0639] Subgroup bounding box origin (subgroup_bbox_origin[i]): Represents the origin of the bounding box belonging to the subgroup of the geometry data unit.
[0640] Subgroup bounding box size (subgroup_bbox_size[i]): Represents the size of the bounding box belonging to the subgroup of the geometry data unit.
[0641] Reference layer group ID (ref_layer_group_id): Represents the ID of the layer group referenced by the geometry data unit.
[0642] Reference subgroup ID (ref_subgroup_id): Represents the ID of the subgroup referenced by the geometry data unit.
[0643] Context reference indication flag (context_reference_indication_flag): A flag indicating whether a context reference exists for a geometry data unit.
[0644] Count of subsequent subgroups (num_subsequent_subgroups): Indicates the number of subsequent dependent data units that reference the current data unit or dependent data unit.
[0645] Subsequently_subgroup_id: Represents the subgroup index of the subsequent dependent data unit of the i-th layer group that references the current data unit or dependent data unit.
[0646] subsequent_subgroup_bbox_origin: Represents the origin of the subsequent subgroup's subgroup bounding box, which references the context state of the current subgroup.
[0647] Subsequently_subgroup_bbox_size: Represents the size of the subgroup bounding box of the subsequent subgroup that references the context state of the current subgroup.
[0648] Meanwhile, the present disclosure may perform an early release of context memory (e.g., context buffer) without additional signaling when performing partial decoding.
[0649] In other words, the present disclosure proposes an efficient context memory management method and signaling in partial coding / decoding situations for multiple slices, and in particular, proposes a method for managing a list internally within a decoder without additional signaling and early releasing a context buffer by determining whether decoding related to a region of interest (ROI) is complete.
[0650] At this time, the decoder may perform part or all of the operation of the point cloud video decoder (10006) of FIG. 1, the decoding (20003) of FIG. 2, the point cloud video decoder of FIG. 8, the point cloud video decoder of FIG. 9, the decoding device of FIG. 28, the decoding method of FIG. 29, the decoding method of FIG. 31, or the decoding method of FIG. 54.
[0651] FIGS. 51 (a) to (c) illustrates another example of a context buffer management method according to embodiments. That is, FIGS. 51 (a) to (c) is an example of partial decoding for a region of interest (ROI) (60000), where skip layer group is 1 (e.g., layer group 2), and decoding can be performed on the region of interest of the decoder.
[0652] That is, in a layer group structure divided into three layer groups as shown in (a) to (c) of FIG. 51, the decoder of the receiving device can perform partial decoding without using the last layer group (i.e., layer-group #2). In addition, partial decoding can be performed only on the ROI region corresponding to reference numeral 60000.
[0653] In FIG. 51 (a) to (c), the context buffer is managed per data unit (e.g., FGS). According to the embodiments, the context buffer may be divided into four storage areas per data unit, for example, a context state storage area (61010), a reference count information storage area (61020), a counter information storage area (61030), and a list information storage area (61040). At this time, the reference count information storage area (61020), the counter information storage area (61030), and the list information storage area (61040) may each be further divided into sub-storage areas corresponding to the number of layer groups within the layer group structure. Since FIG. 51 includes three layer groups, the reference count information storage area (61020), the counter information storage area (61030), and the list information storage area (61040) in the context buffer may each be divided into three sub-storage areas. In FIG. 51 (a) to (c), the top layer group, i.e., the root layer group, is referred to as layer-group #1, the next layer group is referred to as layer-group #2, and the next layer group is referred to as layer-group #3.
[0654] Referring to FIG. 51(a), the context state storage area (61010) of the context buffer stores the context state of FGS0 (0,0) (i.e., context state (0,0)) after FGS0 (0,0) is encoded or decoded. Then, the reference count information storage area (61020) stores the number of data units (i.e., subgroups) that reference the context state (0,0) in each layer group, i.e., three layer groups (layer-group #0 - layer-group #2). At this time, assuming that the context state (0,0) references only the N data units of layer-group #0 (i.e., FGS1 (1,0) ~ FGSN (1,N-1)) and not references it in other layer groups, the value {- | N | 0} is stored in the reference count information storage area (61020). Here, "-" means no reference. That is, for the context state (0,0), it indicates that there is no reference in layer-group #0, it is referenced N times in layer-group #1, and it is not referenced at all in layer-group #2. Also, in FIG. 51(a), since the context state (0,0) has not yet been referenced, the value {- | 0 | 0} is stored in the counter information storage area (61030). Additionally, the list information storage area (61040) stores a list of coded subgroups related to each context, that is, a list of coded subgroups related to the context state of the current subgroup. At this time, since layer-group #2 has not yet referenced layer-group #1, the value {- | - | -} is stored in the list information storage area (61040).
[0655] That is, in FIG. 51 (a), when FGS 0 / subgroup (0,0) is decoded (FGS 0(0,0)), the context state storage area (61010) of the context buffer stores the context state (0,0), and the reference count information storage area (61020) can store NumSubsequentSubgroups (corresponding to num_subsequent_data_units in the signaling information of FIG. 41 to 45, FIG. 49, FIG. 50) transmitted through the data unit header for each subsequent layer group. In this example, the value {- | N | 0} can be stored.
[0656] In other words, when the context_reference_indication_flag is enabled, the number of subsequent (i.e., following) subgroups (num_subsequent_subgroups or num_subsequent_data_units) is signaled, and as shown in FIG. 51 (a), the context state storage area (61010) stores the context state (0,0), and the corresponding number for each layer group is stored in the reference count information storage area (61020). At this time, the counter information storage area (61030) is set to {- | 0 | 0}, and the list information storage area (61040) is set to {- | - | -}.
[0657] These rules apply equally to other data units.
[0658] In the present disclosure, the loading process of the context state (i.e., context information) can be performed by referring to the ref_layer_group_id and ref_subgroup_id parameters. That is, the decoder can prepare for decoding by analyzing the data unit header for each data unit (i.e., subgroup or slice). At this time, the context state used for decoding can be retrieved from the context buffer through the reference layer group ID (ref_layer_group_id) and the reference subgroup ID (ref_subgroup_id). Based on the retrieved context state, the decoder initializes the context buffer for the current data unit and then performs decoding.
[0659] When processing (encoding or decoding) FGS 1 (1,0) in (b) of Fig. 51, the encoder / decoder can initialize the context state of FGS 1 (1,0) based on the context state (0,0) corresponding to ref_layer_group_id=0 and ref_subgroup_id=0. That is, the context state can be initialized based on the context state (0,0).
[0660] That is, in FIG. 51(b), context information (context states (0,0)) of FGS 0 (0,0) can be loaded to process (encode or decode) FGS 1 (1,0). Then, context states (0,0) are used to process FGS 1 (1,0), and the counter value increases by 1. That is, the value {- | 1} is stored in the counter information storage area (61030) corresponding to FGS 0 (0,0). Then, the list in the list information storage area (61040) corresponding to FGS 0 (0,0) is updated to {- | (1,0)}. In the present disclosure, as an embodiment, the list information storage area (61040) stores and updates the list as a pair of layer group index and subgroup index. Additionally, as an embodiment, the list in the list information storage area (61040) is generated by the decoder of the receiving device.
[0661] At this time, since layer-group #2 is assumed to be skipped for decoding, FGS 1 (1,0) is a data unit belonging to the last layer group, and therefore the context state (1,0) of FGS 1 (1,0) is not stored in the context buffer. That is, since FGS 1 (1,0) satisfies the context reference indication flag (context_reference_indication_flag) & under skipLayerGroup condition, the context state after encoding / decoding is not stored in the context buffer. In other words, if the current data unit is a data unit of the last layer group due to the skipped layer group and / or there is no data unit referencing the current data unit, the context state of the encoded or decoded current data unit is not stored in the context buffer.
[0662] In addition, after processing (encoding or decoding) FGS 1 (1,0), the ROI is checked to see if there are any remaining data units to process (encode or decode) in the corresponding layer group. Taking FIG. 51(a) as an example, it can be seen that FGS 2 (1,1) of layer-group #1 is also within the ROI range.
[0663] In this case, encoding or decoding for FGS 2 (1,1) is performed in the same way as FGS 1 (1,0) as in Fig. 51 (c). When processing (encoding or decoding) FGS 2 (1,1) in Fig. 51 (c), the encoder / decoder can initialize the context state of FGS 2 (1,1) based on the context state (0,0) corresponding to ref_layer_group_id=0 and ref_subgroup_id=0.
[0664] That is, in FIG. 51(c), context information (context states (0,0)) of FGS 0 (0,0) can be loaded to process (encode or decode) FGS 2 (1,1). Then, context state (0,0) is used to process FGS 2 (1,1), and the counter value is increased by 1. That is, the value {- | 2} is stored in the counter information storage area (61030) corresponding to FGS 0 (0,0). Then, the list in the list information storage area (61040) corresponding to FGS 0 (0,0) is updated to {- | (1,0), (1,1)}.
[0665] At this time, since it is assumed that layer-group #2 is skipped for decoding, FGS 2 (1,1) is a data unit belonging to the last layer group, and therefore the context state (1,1) of FGS 2 (1,1) is not stored in the context buffer. That is, since FGS 2 (1,1) satisfies the context reference indication flag (context_reference_indication_flag) & under skipLayerGroup condition, the context state after encoding / decoding is not stored in the context buffer.
[0666] As such, in FIG. 51 (a) to (c), depending on conditions such as when context_reference_indication_flag is 1, or for all cases, the context buffer may store the context state for the subgroup index (e.g., a pair of layer group index and subgroup index) that matches the FGS (i.e., data unit) being decoded. In this case, for partial decoding, num_subsequent_subgroups may not be stored because decoding can be completed before being referenced as many times as indicated by num_subsequent_subgroups. Instead, a list of coded subgroups related to each context may be created and stored to store the indices of subgroups that reference the context of the current subgroup (i.e., FGS or data unit).
[0667] And, when a referenced subsequent FGS that references the context of a specific FGS is decoded, the index of the referenced subsequent FGS may be stored in a list corresponding to the specific FGS. In the examples of the present disclosure, the case of storing as (1, 0) or (1, 1), which are pairs of layer-group index and subgroup index, is shown, and when an index of the FGS is defined, a value of 1 or 2 corresponding to fgs_id may be stored. When an index of the FGS is defined, fgs_id may be signaled to a geometry data unit and / or dependent geometry data unit.
[0668] In this case, the present disclosure provides an embodiment in which, as shown in FIG. 51(b) and FIG. 51(c), not only the ROI but also the Occupancy Map is checked together to determine whether the context buffer information corresponding to FGS 0, such as context state information, reference count information, counter information, and list information, is released (i.e., released or deleted). In the present disclosure, release means that the context state information, reference count information, counter information, and list information are deleted from the context buffer (or memory).
[0669] FIGS. 52(a) and FIGS. 52(b) are drawings showing examples of bounding boxes of subgroups and bounding boxes of ROIs according to embodiments.
[0670] That is, when a list corresponding to a specific FGS is updated for the context buffer array, the area defined in the ROI (roiBboxMin, roiBboxMax) and the area covered by the list stored in the list information storage area (61040) can be compared. This is a step of verifying whether the ROI is fully included in the areas defined by the bounding boxes (bbox) of the subgroups within the list. According to the example of FIG. 51 (b), the ROI may cover some of the ROI bounding boxes (ROI_bbox) as in FIG. 52 (a). However, as in the example of FIG. 51 (c), when the list is updated and a subgroup (Bbox[1][1]) having a bounding box (bbox) as in FIG. 52 (b) is additionally decoded, it can be seen that the ROI is fully covered by the subgroups within the list (e.g., (1,0), (1,1)). In other words, the bounding boxes (Bbox[1][0], Bbox[1][1]) of the subgroups within the list (e.g., (1,0), (1,1)) cover all of the bounding boxes (ROI_bbox) of the ROI.
[0671] For a context state referenced across multiple layer groups, it can be compared whether it is covered by the region of the ROI of the sub-list corresponding to each layer group, and if it is confirmed that it is covered, the corresponding context state (in the example of this disclosure, context state (0,0)) can be released. At this time, the release of the context state is possible only when the context is no longer used. That is, when additional partial decoding is performed on another region after progressive decoding or partial decoding, the release can be performed after decoding all FGS referenced within the slice through num_subsequent_subgroups, rather than by comparing the ROI and the region of the list.
[0672] However, due to the nature of point cloud data, there may be parts where no points exist. If a subgroup is not transmitted to the receiving device for areas where points exist, the area of the list may not overlap with the area of the ROI, as shown in Fig. 53 (a). That is, even if the list is updated and a subgroup (Bbox[1][1]) with a bounding box (bbox) is additionally decoded, a part of the ROI may not be covered by the subgroup within the list.
[0673] In this case, the present disclosure can determine whether the ROI is covered by the area of the list only for the area where the actual point exists. In the present disclosure, the occupancy map can determine whether each voxel is occupied based on the subgroup node position information of the subgroup (i.e., FGS) where the context state is stored. Examples of FIG. 51 (b) and (c) are examples of comparing the occupancy map of the subgroup (0, 0), the bounding box of the ROI (ROI_bbox), and the bounding box (bbox) of the subgroup belonging to the list when the referenced context state is (0, 0). In other words, when calculating the overlapping area (hatched region) where the ROI_bbox and the bounding boxes (bboxes) of the subgroups in the list overlap, there are parts of the ROI_bbox that are not covered; however, when the occupancy map is considered simultaneously, it can be confirmed that the ROI_bbox is fully covered for the occupied nodes. In this case, assuming there is no subsequent decoding using that context state, the corresponding context state (i.e., context state (0,0)) can be released from the context buffer.
[0674] FIGS. 53(a) and FIGS. 53(b) are drawings showing comparative examples of the bounding box of a subgroup, the bounding box of an ROI, and the occupied map according to the embodiments.
[0675] The following methods can be considered for releasing the stored context state to manage the context buffer through the aforementioned method.
[0676] For example, in the case of full decoding, it operates based on num_subsequent_data_units (sum of num_sdu_per_layer_group) (i.e., Case 1). As another example, in the case of partial depth, release is performed based on num_sdu_per_layer_group (i.e., Case 2). As yet another example, in the case of partial region, release can be performed by generating a list in the decoder that is signaled (e.g., list of subsequent subgroups) and / or without signaling (e.g., lost of coded subgroups related to each context) (i.e., Case 3).
[0677] In particular, for Case 3 above, the present disclosure can release the context state in the case of partial decoding based on the flowchart of FIG. 54.
[0678] FIGS. 54(a) and FIGS. 54(b) are flowcharts showing other examples of a context buffer management method according to embodiments. More specifically, FIG. 54(a) is an example of releasing a stored context state to manage a context buffer based on a signaled list (e.g., list of subsequent subgroups), and FIG. 54(b) is an example of releasing a stored context state to manage a context buffer by generating a list (e.g., lost of coded subgroups related to each context) in the decoder without signaling.
[0679] In the case of FIG. 54 (a), the process may include storing the signaled list in the storage space listOfSubregionsForROI, erasing the corresponding subgroup from listOfSubregionsForROI when the referenced subgroup is decoded, and releasing the context state when the subgroup index is lost from listOfSubregionsForROI. That is, in the process of storing the given list of FIG. 54 (a) in listOfSubregionsForROI, the given list may be a signaled list, for example, a subsequent subgroup list (list) that references the current subgroup. And, the process of finding a region that overlaps with the current subgroup [cur] within listOfSubregionsForROI (Find subgroup [cur] overlapped region in listOfsubregionsForROI [ref]) may be referred to as ROI check in the present disclosure. After performing the above ROI verification process, the overlapping, i.e., matching region (matchedRegion) in the list is removed (erease the matchedRegion from the list). Then, for all layer groups, if the occupied region within the ROI is covered by contextState[ref], contextState[ref] is released from the context buffer.
[0680] In the case of (b) of FIG. 54, the process may include initializing listOfSubregionsForROI as the current subgroup region (wherein it may consist of one region or may be included in the list as multiple sub-regions), erasing the bounding box (bbox) of the corresponding subgroup from listOfSubregionsForROI whenever a subsequent subgroup is received, and releasing the context state when there is no region in listOfSubregionsForROI or when the region remaining in listOfSubregionsForROI is a non-occupied region.
[0681] That is, when an ROI is established, the decoder generates a listOfSubregionsForROI for the context state of each referenced subgroup. Each listOfSubregionsForROI[ref] lists subregions based on the bounding box of the context referenced subgroup. In the list, the subregions are stored in a format of regional information with minimum and maximum positions. At this time, the initial value of listOfSubregionsForROI[ref] for a specific layer group (e.g., the k-th layer group) is set to the subgroup bounding box of the context reference divided by the unit bounding box of the current subgroup in the k-th layer group. In this case, considering that points are not uniformly distributed, only subregions occupied by one or more points can be stored in the list of subregions of the ROI for the context state of listOfSubregionsForROI[ref]. That is, in (b) of FIG. 54, listOfSubregionsForROI is initially initialized with the area of the subgroup currently being decoded by the decoder, and the list is updated whenever a subsequent subgroup (i.e., a child subgroup or grandchild subgroup referencing me) is decoded. In other words, the decoder constructs the list (listOfSubregionsForROI) and initializes the list with the area of the current subgroup. Since the entire area cannot be initialized, it is initialized with the area of the current subgroup. Additionally, whenever a subsequent subgroup comes in, the bounding box (bbox) of that subgroup is removed from the list.
[0682] In addition, a process of checking the occupancy of each region within the list is performed. That is, for each region within the list, the existence of a point is checked by referring to the occupancy map. Here, as an example, the list (listOfSubregionsForROI) is a list of coded subgroups related to each context. Then, a process of finding a region that overlaps with the current subgroup [cur] within listOfSubregionsForROI (Find subgroup [cur] overlapped region in listOfsubregionsForROI [ref]) is performed, and the present disclosure may refer to this process as ROI check.
[0683] During the above ROI verification process, the bounding box of the subgroup[cur] currently being decoded is compared with each region within listOfSubregionsForROI[ref]. If there is an overlapping region (i.e., matchedRegion) and the current subgroup's region (i.e., bounding box) is not larger than the overlapping region (matchedRegion), the overlapping region is divided and pushed back into listOfsubregionsForROI. In other words, if the current subgroup's region (i.e., bounding box) is larger than the overlapping region (matchedRegion), it means that the current subgroup's region covers the entire overlapping region; otherwise, it means that it covers only a portion of the overlapping region.
[0684] After performing the ROI verification process above or after dividing the overlapping region above, the overlapping region (matchedRegion) is erased from the list. That is, if there is an overlapping region (i.e., a matched region), that region is erased. If the matched region is of the same size, it is simply erased, and if it is not of the same size, i.e., if some parts are not matched, only the matched region is erased and the rest are left in the list.
[0685] Then, for all layer groups, if the occupied area within the ROI is covered by contextState[ref], contextState[ref] is released from the context buffer.
[0686] That is, when the context state of subgroup index ref is referenced by the subsequent subgroup index cur, the bounding box of subgroup cur is compared with the bounding box of the sub-regions in listOfSubregionsForROI[ref]. If there is a sub-region that matches the bounding box of subgroup cur, that matching sub-region is erased from the list. If there is a sub-region that overlaps the bounding box of subgroup cur, that sub-region is divided into the unit bounding box of subgroup cur. Then, the divided sub-regions are added to listOfSubregionsForROI[ref][k] only if they are not occupied by one or more points and overlap with an ROI. And, whenever the context state of a subgroup index ref is referenced by a subsequent subgroup, listOfSubregionsForROI[ref] is updated, and if listOfSubregionsForROI[ref] is empty for all layer groups, the context state of that index ref can be released.
[0687] As described above, there may be cases where a portion of the bounding box of an ROI is not covered by the bounding box of a subgroup in the list (i.e., the area of the decoded child subgroup). To address this, the present disclosure determines whether to release a context state by simultaneously considering the occupancy map. For example, if the bounding box of an ROI in an overlapping area is covered by an occupied node, the context state (i.e., context state (0,0)) may be released from the context buffer under the assumption that there is no subsequent decoding using that context state. That is, for partial decoding of a spatial area, the present disclosure creates a list of sub-areas for an ROI in the decoder and determines whether to release the context state by checking whether all occupied areas within the ROI are covered by the decoded subgroup.
[0688] By doing so, the decoder of the present disclosure enables the early release (i.e., release) of the context state without additional signaling when performing partial decoding. Therefore, even when partial decoding is performed based on an ROI (region of interest), unnecessary memory occupation of the context state can be reduced.
[0689] As described above, when an ROI for partial region decoding is provided, the decoder requires a process of selecting a data unit associated with the ROI. In this disclosure, slice, subgroup, data unit, and FGS may be used interchangeably as having the same meaning. In this disclosure, the data unit selected for decoding may be referred to as FGS, FGS data unit, etc.
[0690] The following explains how to select an efficient data unit in partial decoding situations.
[0691] According to embodiments, the decoder of the present disclosure can perform decoding by selecting a data unit that overlaps with the ROI.
[0692] In the present disclosure, a data unit can be selected through a selection method based on two conditions. One condition is a loosely overlapped condition, and the other condition is a strictly overlapped condition.
[0693] First, here is an explanation of the loosely overlapped condition. In other words, the loosely overlapped condition is a method of selecting a data unit (i.e., FGS) when the ROI overlaps with the region. That is to say, the loosely overlapped condition is a method of selecting a data unit (i.e., FGS) that includes an overlapping region by comparing the subgroup bounding box with the ROI if there is an overlapping region.
[0694] According to the loosely overlapping condition of the present disclosure, when selectively decoding FGS based on ROI (i.e., region of interest), if there is an overlapping region when comparing the subgroup bounding box with the region of the ROI, the FGS corresponding to that subgroup can be selected and decoded. That is, the regions of the subgroups are compared with the ROI, and the subgroup(s) having (or having) an overlapping region with the ROI are selected and decoded.
[0695] The following code shows the process of selecting data units to decode based on loosely overlapping conditions.
[0696] In the code below, for each axis (x, y, z), if all of the following conditions are satisfied (i.e., ROI_min[i] < BB_max[i] and ROI_max[i] ≥ BB_min[i], (i ∈{x, y, z})), the two regions are determined to be overlapping. That is, if the above conditions are true for all axes, the ROI and the subgroup bounding box are determined to be in an overlapping state. In other words, when the ROI is enabled (_roi_enabled_flag = true), the minimum / max coordinates of the ROI (_roi_min, _roi_max) and the minimum / max coordinates of the subgroup bounding box (bboxOrigin, bboxOrigin + bboxSize) are compared to determine whether an overlap exists in each axis direction.
[0697] bool isRequiredDataunit = (!_sps->subgroup_enabled_flag[_dep_gbh.layer_group_id] || _gHandler.checkRoi(_dep_gbh.subgroupBboxOrigin, _dep_gbh.subgroupBboxSize));
[0698] bool LayerGroupHandler::checkRoi(Vec3 <int>bboxorigin, Vec3 <int>bboxSize){
[0699] bool isRequiredSubGroup = true;
[0700] auto curBboxMin = bboxorigin;
[0701] auto curBboxMax = bboxorigin + bboxSize;
[0702] if (_roi_enabled_flag) {
[0703] for (int i = 0; i < 3; i++) {
[0704] if ((_roi_min[i] < curBboxMax[i] && _roi_max[i] >= curBboxMin[i])) {
[0705] continue;
[0706] } else {
[0707] isRequiredSubGroup = false;
[0708] break;
[0709] }
[0710] }
[0711] }
[0712] return isRequiredSubGroup;
[0713] }
[0714] The following is an explanation of the strictly overlapped condition. In this disclosure, the strictly overlapped condition is a method of selecting a data unit (i.e., FGS) containing a region only when there are points / nodes in the region that overlaps with the ROI. That is, under the strictly overlapped condition, if there are actually no nodes or points in the corresponding overlapping region, it can be determined that the data unit is not actually necessary from the perspective of the ROI, and thus the decoding of that data unit can be selected not to be performed. In other words, for each subgroup bounding box, the decoder can determine that a data unit (i.e., FGS) is related to the ROI only when the region overlaps with the ROI and points / nodes exist in the overlapping region, and then perform decoding of that data unit. In other words, subgroups in which no points exist within the overlapping region are considered unnecessary data units from the perspective of the ROI, and decoding is omitted.
[0715] According to the strictly overlapping condition of the present disclosure, when selectively decoding FGS based on an ROI (i.e., region of interest), if there is an overlapping area when comparing a subgroup bounding box with the area of the ROI, the FGS corresponding to that subgroup can be selected and decoded only if points and / or nodes exist in the overlapping area. In this way, the strictly overlapping condition further verifies whether actual points / nodes exist within the overlapping area.
[0716] The following code shows the process of selecting data units to be decoded based on strictly overlapping conditions.
[0717] In the code below, the FGS of the corresponding subgroup is determined to be decoded only when the ROI and the subgroup bounding box overlap (checkRoi) and an actual point exists within that overlapping area (checkRoiHasPoint). Specifically, the checkRoiHasPoint() function references the point cloud (subgroupPointCloud) of the upper layer group (refLayerGroupIdx = layer_group_id - 1) and checks whether each point coordinate (x, y, z) is included in both the ROI and the subgroup bounding box as follows.
[0718] if ((pos.x() < bboxMax.x() && pos.x() >= bboxMin.x() &&
[0719] pos.y() < bboxMax.y() && pos.y() >= bboxMin.y() &&
[0720] pos.z() < bboxMax.z() && pos.z() >= bboxMin.z()) &&
[0721] (pos.x() < _roi_max.x() && pos.x() >= _roi_min.x() &&
[0722] pos.y() < _roi_max.y() && pos.y() >= _roi_min.y() &&
[0723] pos.z() < _roi_max.z() && pos.z() >= _roi_min.z()))
[0724] If there is at least one point satisfying the above condition, it is determined that an actual point exists within the overlapped region, and the FGS containing that region is selected.
[0725] bool isRequiredLayer_gHandler.isRequiredLayer(_dep_gbh.layer_group_id);
[0726] bool isRequiredDataunit = (!_sps->subgroup_enabled_flag[_dep_gbh.layer_group_id] ||
[0727] (_gHandler.checkRoi(_dep_gbh.subgroupBboxOrigin, _dep_gbh.subgroupBboxSize)
[0728] && _gHandler.checkRoiHasPoint(_dep_gbh.layer_group_id, _dep_gbh.subgroup_id, curBboxMin, curBboxMax, subgroup PointCloud)));
[0729] if(!(isRequiredLayer && isRequiredDataunit)){
[0730] return 0;
[0731] }
[0732]
[0733] bool LayerGroupHandler::checkRoiHasPoint{
[0734] int layerGroupId,
[0735] int subgroupId,
[0736] Vec3 <int>& bboxMin,
[0737] Vec3 <int>& bboxMax,
[0738] std::vector<std::unique_ptr <pccpointset3>>& subgroupPointCloud)
[0739] {
[0740] bool isRequired = true;
[0741] if (_roi_enabled_flag) (
[0742] isRequired false;
[0743] pcc:: LayerGroupkey parentkey = { 0, 0);
[0744] if (layerGroupId > 0) {
[0745] int refLayerGroupIdx= layerGroupId 1;
[0746] parentkey checkBox (bboxMin, bboxMax, refLayerGroupIdx);
[0747] assert(parentKey.subgroupID >= 0);
[0748] }
[0749] if (layerGroupIdxToSavedArrayIdx.count(parentKey)) {
[0750] int parentArrayIdx = 0;
[0751] if (layerGroupId > 0)
[0752] parentArrayIdx= _layerGroupIdxToSavedArrayIdx[parentKey];
[0753] if (_available_geom[parentArrayIdx]) {
[0754] for (int i=_numDCMPointsSubgroup [parentArrayIdx]; i < subgroupPointCloud [parentArrayIdx]->getPointCount(); i++) {
[0755] auto& pointCloud = subgroupPointCloud [parentArrayIdx];
[0756] auto& pos pointCloud[i];
[0757] if ((pos.x() < bboxMax.x() && pos.x() >= bboxMin.x()
[0758] && pos.y() < bboxMax.y() && pos.y() >= bboxMin.y() / point is in the current bbox
[0759] && pos.z() < bboxMax.z() && pos.z() >= bboxMin.z())
[0760] && (pos.x() <_roi_max.x() && pos.x() >= _roi_min.x() / point is in the ROI
[0761] && pos.y() <_roi_max.y() && pos.y() >= _roi_min.y()
[0762] && pos.z() <_roi_max.z() && pos.z() >= _roi_min.z())) {
[0763] isRequired+true;
[0764] }
[0765] }
[0766] }
[0767] }
[0768] return isRequired;
[0769] }
[0770] Based on the above, the decoder of the present disclosure can selectively decode FGS for partial decoding situations as follows.
[0771] The following is a description of partial density decoding according to the embodiments.
[0772] partial density decoding
[0773] In the present disclosure, the decoder can generate a lower density slice (or FGS) point cloud. That is, the decoder can generate a lower density FGS point cloud by selectively decoding only some layer groups instead of decoding all layer groups of the entire point cloud.
[0774] According to the embodiments, a low-density FGS point cloud is defined by the following variables.
[0775] The variable SkippedLayerGroup is an application-specific variable representing the number of skipped layer-groups for partial decoding in the direction of the density.
[0776] The value of the variable SkippedLayerGroup can be in the range of 0 to num_layer_groups_minus1.
[0777] The variable MinNodeSizeLog2 represents the minimum occupancy tree node size defined by the SkippedLayerGroup.
[0778] The array SubgroupNodePos[layerGroupIdx][subgroupIdx][ptIdx][k] represents the subgroup output nodes corresponding to the layer group index (layerGroupIdx) and the subgroup index (subgroupIdx).
[0779] The array SubgroupNodeCnt[layerGroupIdx][subgroupIdx] represents the number of nodes in the subgroup output nodes of the layer-group index layerGroupIdx and the subgroup index subgroupIdx.
[0780] Selection of FGS
[0781] If the variable SkippedLayerGroup is greater than 0, only layer groups with indices greater than or equal to 0 and less than or equal to OutLayerGroup are selected for decoding.
[0782] The maximum value of the layer-group index of partial decoding (OutLayerGroup) is defined as the total number of layer-groups minus SkippedLayerGroup.
[0783] According to the embodiments, OutLayerGroup is calculated as num_layer_groups_minus1 - SkippedLayerGroup as in the code below. Then, if the layer group ID is 0, GDU or ADU is decoded, and if the layer group ID is less than or equal to OutLayerGroup, DGDU or DADU is decoded, and for other layer groups, DGDU or DADU is skipped.
[0784] OutLayerGroup := num_layer_groups_minus1 - SkippedLayerGroup
[0785] if (layer_group_id == 0)
[0786] decode GDU or ADU
[0787] else if (layer_group_id ≤ OutLayerGroup)
[0788] decode DGDU or DADU
[0789] else
[0790] skip DGDU or DADU
[0791] Consequently, the PartialDepth of the geometry occupancy tree in partial decoding is derived from the sum of the number of layers within each layer group with indices ranging from 0 to OutLayerGroup. That is, the cumulative value obtained by adding (num_layers_minus1[i] + 1) for each layer group becomes the PartialDepth as follows.
[0792] PartialDepth = 0
[0793] for (i=0; i ≤ OutLayerGroup; i++)
[0794] PartialDepth += num_layers_minus1[i] + 1
[0795] Geometry position compensation
[0796] The maximum depth (TotalDepth) of the geometry occupancy tree when decoding all layer groups is derived from the sum of the number of layers within each layer group with indices from 0 to num_layer_groups_minus1 as follows.
[0797] TotalDepth = 0
[0798] for (i=0; i< num_layer_groups_minus1; i++)
[0799] TotalDepth += num_layers_minus1[i] + 1
[0800] According to the embodiments, MinNodeSizeLog2 is derived from the difference between occtreeMaxDepthMinus1 and PartialDepth as follows.
[0801] MinNodeSizeLog2 = occtreeMaxDepthMinus1 + 1 - PartialDepth
[0802] If MinNodeSizeLog2 is greater than 1, points shall be centered within their corresponding block as follows:
[0803] for (ptIdx = 0; ptIdx < SubgroupNodeCnt[layerGroupIdx][subgroupIdx]; ptIdx++)
[0804] for (k = 0; k < 3; k++)
[0805] SubgroupNode[layerGroupIdx][subgroupIdx][ptIdx][k] |= (MinNodeSizeLog2 > 1) << (MinNodeSizeLog2 - 1)
[0806] In other words, the coordinates of each point are corrected relative to the block center according to the minimum node size (Log2 scale).
[0807] As such, the decoder controls the decoding range using the SkippedLayerGroup variable, which represents the number of layer groups skipped during partial decoding. SkippedLayerGroup is defined as an integer value between 0 and num_layer_groups_minus1. Furthermore, the minimum occupancy tree node size (MinNodeSizeLog2) determined by SkippedLayerGroup is calculated by reflecting the depth difference of the upper layer groups that are not decoded. That is, it is defined as MinNodeSizeLog2 = occtreeMaxDepthMinus1 + 1 - PartialDepth, which corresponds to the difference between the maximum depth (PartialDepth) of the decoded partial occupancy tree and the maximum depth (TotalDepth) of the entire occupancy tree.
[0808] According to embodiments, if SkippedLayerGroup is greater than 0, only layer groups with indices from 0 to OutLayerGroup are selected for decoding. In this case, OutLayerGroup is calculated as num_layer_groups_minus1 - SkippedLayerGroup. That is, the decoder processes each data unit in the following procedure: if the layer group ID is 0, the GDU or ADU is decoded. if the layer group ID is less than or equal to OutLayerGroup, the DGDU or DADU is decoded. Other layer groups are skipped.
[0809] Through such condition control, the decoder performs partial decoding within a specified density range.
[0810] In addition, the depth of the partial occupancy tree is determined by the number of layer groups being decoded and the number of layers in each group. If MinNodeSizeLog2 is greater than 1, each point is restored by aligning it to the center position of the corresponding block. That is, the position coordinates of the points are corrected relative to the block center through the following bit shift operation. Through this, geometry alignment correction is performed so that points are located at the block center even during partial decoding.
[0811] The following is a description of partial region decoding according to the embodiments.
[0812] Partial region decoding
[0813] According to embodiments, the decoder can generate a point cloud of a partial region of a slice. That is, the decoder can generate a point cloud of a partial region limited to the required spatial portion by selectively decoding only a portion of the region corresponding to the ROI, rather than the entire slice.
[0814] In the present disclosure, the partial region FGS point cloud is defined by the following variables.
[0815] The arrays RoiBBoxMin and RoiBBoxMax are application-specific arrays that specify the region of interest as the minimum and the maximum position of the bounding box. That is, the decoder uses RoiBBoxMin (the minimum coordinate of the ROI) and RoiBBoxMax (the maximum coordinate of the ROI) to specify the 3D spatial range of the ROI. These two arrays constitute the spatial bounding box of the ROI.
[0816] The array SubgroupNodePos[layerGroupIdx][subgroupIdx] represents the subgroup output nodes of the layer-group index layerGroupIdx and the subgroup index subgroupIdx.
[0817] The array SubgroupNodeCnt[layerGroupIdx][subgroupIdx] represents the number of nodes within the subgroup output nodes of the layer group index (layerGroupIdx) and subgroup index (subgroupIdx).
[0818] That is, corresponding to each layer group and subgroup, the coordinates and number of subgroup output nodes are represented as SubgroupNodePos[layerGroupIdx][subgroupIdx] and SubgroupNodeCnt[layerGroupIdx][subgroupIdx], respectively.
[0819] Selection of FGS
[0820] According to embodiments, the decoder selects for decoding the subgroup(s) whose subgroup bounding box overlaps with the bounding box of the ROI, where RoiBBoxMin and RoiBBoxMax exist.
[0821] The following indicates the selection criteria for the subgroup(s) to be decoded. That is, if the layer group index (layerGroupIdx) is 0, decode the GDU or ADU; if the ROI and the subgroup bounding box overlap and there is occupancy within the overlap area, decode the DGDU or DADU; otherwise, skip the corresponding DGDU or DADU.
[0822] if (layerGroupIdx == 0)
[0823] decode GDU or ADU
[0824] else if ((RoiBBoxMin[0] < SubgroupBBoxMax[layerGroupIdx][subgroupIdx][0] &&
[0825] RoiBBoxMin[1] < SubgroupBBoxMax[layerGroupIdx][subgroupIdx][1] &&
[0826] RoiBBoxMin[2] < SubgroupBBoxMax[layerGroupIdx][subgroupIdx][2]) &&
[0827] (RoiBBoxMax[0] > SubgroupBBoxMin[layerGroupIdx][subgroupIdx][0] &&
[0828] RoiBBoxMax[1] > SubgroupBBoxMin[layerGroupIdx][subgroupIdx][1] &&
[0829] RoiBBoxMax[2] > SubgroupBBoxMin[layerGroupIdx][subgroupIdx][2]) && occupied)
[0830] decode DGDU or DADU
[0831] else
[0832] skip DGDU or DADU
[0833] In the code above, 'occupied' represents the occupancy of the ROI overlapped region within the subgroup bounding box, which is estimated as follows:
[0834] occupied = false
[0835] for(i=0; i< SubgroupNodeCnt[layerGroupIdx][subgroupIdx]; i++) {
[0836] if (RoiBBoxMin[0] ≤ SubgroupNodePos[layerGroupIdx][subgroupIdx][i][0]
[0837] && RoiBBoxMin[1] ≤ SubgroupNodePos[layerGroupIdx][subgroupIdx][i][1]
[0838] && RoiBBoxMin[2] ≤ SubgroupNodePos[layerGroupIdx][subgroupIdx][i][2]
[0839] && RoiBBoxMax[0] > SubgroupNodePos[layerGroupIdx][subgroupIdx][i][0]
[0840] && RoiBBoxMax[1] > SubgroupNodePos[layerGroupIdx][subgroupIdx][i][1]
[0841] && RoiBBoxMax[2] > SubgroupNodePos[layerGroupIdx][subgroupIdx][i][2]) {
[0842] occupied = true
[0843] break
[0844] }
[0845] }
[0846] That is, if the coordinates of each node within a subgroup are inside the ROI bounding box, occupied = true is set, and if even one exists, the corresponding subgroup is selected as the target for decoding.
[0847] More specifically, in the code above, the overlap determination is performed under the following conditions.
[0848] RoiBBoxMin[k] < SubgroupBBoxMax[k] ∧ RoiBBoxMax[k] > SubgroupBBoxMin[k] (k ∈ {0,1,2})
[0849] That is, if the bounding boxes of the ROI and the subgroup overlap in all axis directions, it is determined to be an overlap.
[0850] Then, to determine whether actual points / nodes exist in the overlap area, the decoder examines the output node locations (i.e., coordinates) of each subgroup.
[0851] According to the embodiments, if there is at least one node satisfying all of the following conditions, the decoder determines that the overlap area of the corresponding subgroup is occupied = true.
[0852] RoiBBoxMin[k] ≤ SubgroupNodePos[layerGroupIdx][subgroupIdx][i][k] < RoiBBoxMax[k] (for all k ∈{0,1,2})
[0853] In other words, if a node is contained within a space greater than or equal to the minimum coordinates of the ROI and less than the maximum coordinates, that region is considered occupied.
[0854] According to the embodiments, the decoder decodes the DGDU or DADU of the corresponding subgroup when it is determined that occupied = true, and otherwise skips decoding.
[0855] To summarize, when decoding partial regions based on ROI, the decoder sets the bounding box of the ROI by referencing RoiBBoxMin and RoiBBoxMax, and determines whether there is an overlap between the ROI and the subgroup bounding box. If an overlap is determined, it checks whether actual points / nodes exist within the overlapping area and decodes subgroups where occupied = true (i.e., DGDU / DADU). Decoding is skipped for subgroups that are not overlapping or have no occupancy.
[0856] By doing this, data in subgroups that do not overlap with or are not occupied by the ROI area can be skipped, thereby reducing decoding computation and transmission volume, and only partial point clouds corresponding to the ROI can be quickly restored, maximizing random access and real-time rendering efficiency.
[0857] Meanwhile, the present disclosure may perform a memory release in relation to partial decoding as described above. In particular, FIGS. 51 to 54 disclose an early release of context memory (e.g., context buffer) without additional signaling when performing partial decoding.
[0858] In addition, the present disclosure can perform a point / node-based early release in relation to partial decoding.
[0859] According to the embodiments, if there are no points / nodes stored for a subgroup, a context memory (e.g., context buffer) release and a saved node release may be performed. Additionally, if there are no points / nodes in an area overlapping with an ROI, a context memory release and a saved node release may be performed.
[0860] That is, the present disclosure can perform point / node-based early release as memory management for partial decoding.
[0861] According to the embodiments, context buffer management (or stored node management) can be effectively performed based on a strictly overlapped condition as described above. If there are no nodes or points in the area overlapping with the ROI for the subgroup bounding box of the decoded FGS (for example, in the case of a direct coding node with no children, i.e., a direct coding node without child nodes), it can be expected that the child subgroups of that subgroup will not be selected as decoding targets. In other words, decoding of the child subgroups of that subgroup may be skipped. In this case, since the decoder can assume that the context state and node of that subgroup will no longer be used, they can be released from the context buffer immediately without being stored. That is, the context information of that subgroup stored in the context buffer can be released (deleted).
[0862] According to embodiments, if a point / node does not exist in an area overlapping with the ROI, the decoder may release the context buffer as follows.
[0863] if (_roi_enabled_flag
[0864] && !checkRoiHasPoint(groupIndex, subgroupIndex, _bboxMinVector[curArrayIdx], _bboxMaxVector[curArrayIdx], subgroupPointCloud)) {
[0865] releaseNodes(curArrayIdx);
[0866] releaseCtxForGeometry(curArrayIdx);
[0867] }
[0868] In the code above, checkRoiHasPoint() is a function that checks whether an actual point or node exists within the overlapping area of the ROI and the subgroup. That is, if there is no point within the area overlapping the ROI (checkRoiHasPoint() == false), the decoder releases the nodes of the corresponding subgroup and the geometry-related context from the context buffer.
[0869] Similarly, the decoder tracks the nodes used during the child subgroup decoding process as shown in FIG. 55a and FIG. 55b, and can release the memory (e.g., context buffer) for storing the nodes of the corresponding subgroup when there are no more nodes to be used for child subgroup decoding. That is, even during the process of decoding a child subgroup, the nodes used for the decoding are tracked, and if they are no longer used for the next step of child subgroup decoding, the context buffer for storing the nodes of the corresponding subgroup can be released early. For example, if the value of _numRamainingNodesForChildSubgroups[refArrayIdx4Parent] is 0, the nodes corresponding to refArrayIdx4Parent are released. Here, _numRamainingNodesForChildSubgroups is an array that tracks the number of remaining nodes among the child subgroups held by each parent subgroup that have not yet finished decoding or are being referenced, and refArrayIdx4Parent represents the reference index of the parent subgroup currently being processed. That is, the decoder tracks the number of nodes in the child subgroups managed by the parent subgroup (_numRamainingNodesForChildSubgroups), and when the value becomes 0 and is no longer used for child subgroup decoding, it releases the geometry node corresponding to the parent subgroup from the context buffer.
[0870] FIGS. 55a and 55b illustrate examples of code implementations for performing context buffer management according to embodiments. The code in FIGS. 55a and 55b defines the geometry decoder resource release procedure (releaseGeometryDecoderResource) performed within the LayerGroupHandler class. That is, it is a logic that dynamically manages node and context resources on a unit basis for each subgroup, parent, and reference layer group, taking into account ROI conditions, the existence of child subgroups, context reference status, etc.
[0871] According to the embodiments, if there are no remaining nodes in the child subgroup for a parent subgroup, the geometry nodes of that parent subgroup are immediately released (releaseNodes(refArrayIdx4Parent)). Additionally, releaseGeometryDecoderResource() partially or completely releases the geometry decoder resources (context, node, point set, etc.) corresponding to each subgroup (curArrayIdx) after decoding is complete, depending on the situation.
[0872] That is, when performing ROI-based partial decoding, the decoder performs resource management logic that comprehensively determines whether there are nodes in child subgroups, context reference status, ROI overlap, etc., for each subgroup and layer group, and releases geometry nodes and contexts in stages.
[0873] So far, we have explained how to manage the context buffer during geometry decoding.
[0874] According to the embodiments, in the case of attribute decoding, multiple attributes may exist for a single subgroup, and in this case, as shown in FIGS. 56a to 56c, the geometry node and attribute context may be released after waiting until all related attributes are decoded (_numAttrsForReleaseAttrSubgroups[curArrayIdx], allChildAttrsDecodedFlag).
[0875] FIGS. 56a to 56c illustrate examples of code implementations for managing the context buffer of attributes according to embodiments. According to embodiments, when attribute decoding is finished, the decoder evaluates for each subgroup the number of remaining nodes of the child subgroup, whether decoding of the child subgroup per attribute is complete, the number of subsequent subgroups of the reference context, and ROI-based early release conditions, and can release the geometry nodes and attribute contexts of the corresponding subgroup (and parent / reference).
[0876] FIGS. 57a and FIGS. 57b are drawings showing examples of code implementations for releasing a context state stored in a context buffer according to embodiments.
[0877] Referring to FIG. 57a and FIG. 57b, an example of an implementation for releasing a context state stored in a context buffer is as follows.
[0878] As part of the initialization process, if the ROI area list for the context reference is empty, the overlapping area between the bounding box of the context reference and the ROI is treated as a sub-region and added to the ROI area list (_listOfSubregionsForRoi or _listOfSubregionsForRoi_attr).
[0879] If the bounding box of a sub-region in the list matches that of a decoded subgroup (curArrayIdx), remove that sub-region from the ROI list.
[0880] If there is an overlapping area between the sub-area in the list and the bounding box of the decoded subgroup (curArrayIdx), the sub-area with the overlap is divided into sub-areas, the overlap is removed, and the remaining sub-areas are included in the list.
[0881] And, if all areas in the list are erased, it can be assumed that the context state for that subgroup is no longer used.
[0882] FIGS. 58a and FIGS. 58b are drawings showing examples of code implementation for releasing a parent node stored for decoding according to embodiments.
[0883] Referring to FIG. 58a and FIG. 58b, an example of an implementation for releasing a parent node stored for decoding is as follows.
[0884] As part of the initialization process, if the area list for the parent node is empty, the overlapping area between the parent subgroup's bounding box and the ROI is treated as a single sub-area and added to the ROI area list (_listOfSubregionsForRoi_parent or _listOfSubregionsForRoi_parent_attr).
[0885] If the bounding box (bboxMin, bboxMax) of a decoded subgroup matches a sub-region in the list, delete that sub-region from the list.
[0886] If there is an overlapping area between the sub-area in the list and the bounding box (bboxMin, bboxMax) of the decoded subgroup, the sub-area with the overlap is divided into detailed areas, the overlap is removed, and the remaining detailed areas are included in the list.
[0887] And, if all areas in the list are erased, it can be assumed that the corresponding parent node is no longer in use.
[0888] FIGS. 59a and 59b are drawings showing examples of code implementation for generating a list of sub-regions for an area overlapping with an ROI according to embodiments.
[0889] That is, when generating a list of sub-regions for areas that overlap with the ROI, as shown in FIG. 59a and FIG. 59b, additional consideration is given to whether there are points / nodes in each area, and only those cases where there is at least one point / node can be included in the list.
[0890] The following is a description of how to fix an error (or bug) regarding the number of subsequent subgroups (numSubsequentSubgroups) that may occur in the encoder.
[0891] Encoder bugfix for numSubsequentSubgroups
[0892] When point cloud data is input to an encoder according to the embodiments, the encoder configures a layer group structure and obtains relevant parameters. Then, based on the layer group structure (or already by external input), it establishes reference relationships between subgroups. At this time, for each subgroup (or FGS), it investigates whether it is used as a reference in subsequent subgroup(s) (or FGS(s)), and if it is used as a reference, it may signal the number of subsequent subgroups (numSubsequentSubgroups) referencing the said subgroup to a data unit header or a dependent data unit header. In this disclosure, the number of subsequent subgroups (numSubsequentSubgroups) may be referred to as the number of subsequent data units. The number of subsequent data units is signaled to a data unit header and / or a dependent data unit header as shown in FIGS. 41 through 45, FIGS. 49, and FIGS. 50 with the syntax element name num_subsequent_subgroups or num_subsequent_data_units. In this disclosure, num_subsequent_data_units, num_subsequent_subgroups, and numSubsequentSubgroups are used interchangeably with the same meaning. That is, the number of subsequent data units (num_subsequent_data_units) represents the number of subsequent dependent data units that reference the current data unit or dependent data unit. The data unit may be an FGS for a subgroup within a layer group according to the embodiments.
[0893] FIG. 60(a) is a flowchart showing an example of an encoding operation according to embodiments. That is, the FGS geometry of a specific subgroup is encoded, and the number of nodes within the subgroup is checked. If the number of nodes within the subgroup is greater than 0, that is, if there is one or more nodes in the output geometry FGS, the data unit header and the encoded FGS geometry are written to the bitstream. If the number of nodes within the subgroup is not greater than 0, that is, if there are no nodes in the output geometry FGS, the data unit header and the encoded FGS geometry are not written to the bitstream (i.e., skipped). Then, it is checked whether all subgroups have been encoded, and if it is confirmed that all subgroups have not been encoded, the process proceeds to the step of encoding the FGS geometry and the above process is repeated.
[0894] That is, if the encoder has no nodes in the output geometry FGS, it does not write the FGS to the bitstream.
[0895] In this case, since the value of numSubsequentSubgroups is determined before the FGS skip is decided, a discrepancy may occur between the value of numSubsequentSubgroups and the actual number of subsequent subgroups. For example, the value of numSubsequentSubgroups is 10, but due to the FGS skip, the actual number of subsequent subgroups may be 9.
[0896] To modify this, the present disclosure recalculates / modifies numSubsequentSubgroups when an empty FGS exists, and then writes the data unit header and data unit to the bitstream. Here, an empty FGS is when there are no nodes in the subgroup, that is, when there are no nodes in the output geometry FGS.
[0897] FIG. 60(b) is a flowchart showing another example of an encoding operation according to embodiments. That is, the FGS geometry of a specific subgroup is encoded, and it is checked whether all subgroups have been coded. If it is confirmed that all subgroups have not been coded, the process proceeds to the step of coding the FGS geometry and the above process is repeated. If it is confirmed that all subgroups have been coded and there is an empty FGS geometry, the number of subsequent subgroups (numSubsequentSubgroups) is recalculated. For example, if the number of subsequent subgroups is determined to be 10 but there is one empty FGS among the subsequent subgroups, the number of subsequent subgroups is modified to 9. That is, the value of the number of subsequent subgroups (numSubsequentSubgroups) is corrected so that the signaled number of subsequent subgroups (numSubsequentSubgroups) matches the actual number of subsequent subgroups.
[0898] A detailed explanation of the recalculation / modification of the number of subsequent subgroups will be provided later.
[0899] After the above processes are performed, the number of nodes in the subgroup is checked. If the number of nodes in the subgroup is greater than 0, that is, if there is one or more nodes in the output geometry FGS, the data unit header and the coded FGS geometry are written to the bitstream. If the number of nodes in the subgroup is not greater than 0, that is, if there are no nodes in the output geometry FGS, the data unit header and the coded FGS geometry are not written to the bitstream (i.e., skipped). Then, it is checked whether all FGSs have been coded, and if it is confirmed that all FGSs have not been coded, the process proceeds to the step of checking the number of nodes in the subgroup and the above process is repeated.
[0900] By doing this, the numSubsequentSubgroups value written to the bitstream can be exactly matched to the actual number of subsequent subgroups even when empty FGS are skipped.
[0901] According to the embodiments, when there is empty FGS geometry, numSubsequentSubgroups is recalculated as in the following code.
[0902] if (_gHandler._available_geom.size() != codedFGS_list.size()) {
[0903] for (int curArrayIdx = 0; curArrayIdx < _gHandler._available_geom.size(); curArrayIdx++) {
[0904] / when an FGS is not encoded, update numSubsequent Subgroups
[0905] if (!_gHandler._available_geom[curArrayIdx]) {
[0906] LayerGroupKey key = _gHandler.getLayerGroupIds (curArrayIdx);
[0907] int layerGroupId = key.layerGroupID;
[0908] int subgroupId = key.subgroupID;
[0909] int refArrayIdx = _gHandler.getReferenceIdx(curArrayIdx);
[0910] LayerGroupkey key_ref = _gHandler.getLayerGroupIds (refArrayIdx);
[0911] int layerGroupId_ref = key_ref.layerGroupID;
[0912] int subgroupId_ref = key_ref.subgroupID;
[0913] if (layerGroupId_ref == 0) {
[0914] gbh.numSubsequent Subgroups [layerGroupId]--;
[0915] }
[0916] else {
[0917] / find encoding order
[0918] int codingOrder_ref = 0;
[0919] for (int i = 0; i < codedFGS_list.size(); i++) {
[0920] if (codedFGS_list[i] == refArrayIdx) {
[0921] codingOrder_ref = i;
[0922] break;
[0923] }
[0924] }
[0925] if (codingOrder_ref) {
[0926] dep_gbh_array [codingOrder_ref].numSubsequent Subgroups [layerGroupId]--;
[0927] }
[0928] }
[0929] _gHandler._available_geom [curArrayIdx] = false;
[0930] }
[0931] }
[0932] }
[0933] The encoder according to the embodiments performs logic to recalculate / modify the number of subsequent subgroups (numSubsequentSubgroups) as in the code above when there is an empty subgroup (i.e., FGS) among the geometry FGSs. In the code above, _available_geom is an array representing the availability status of all geometry FGSs, and codedFGS_list is a list of FGSs actually written to the bitstream. At this time, the fact that the sizes of the two lists are different means that there is an empty FGS, and the encoder can recognize that it needs to correct the discrepancy in the number of subsequent subgroups (numSubsequentSubgroups). Specifically, the encoder determines whether an empty FGS exists by comparing the size of the availability flag (_available_geom) for all FGS candidates with the size of the actual encoded list (codedFGS_list). If an empty FGS is detected, the layer group ID and reference index of the corresponding FGS are queried to identify the reference relationship. If the reference layer group is a base layer group (i.e., layerGroupId_ref == 0), numSubsequentSubgroups[layerGroupId] of gbh is decremented, and if the reference is a dependent layer group, the encoding order is searched and numSubsequentSubgroups[layerGroupId] of the dependent header (dep_gbh_array[codingOrder_ref]) is decremented. Through this correction procedure, the present disclosure ensures that the number of subsequent subgroups signaled in the bitstream matches the actual number of subsequent subgroups.
[0934] In this way, the encoder recalculates numSubsequentSubgroups to prevent a discrepancy between the value of numSubsequentSubgroups and the actual number of subsequent subgroups that may occur when the output geometry FGS is skipped from the bitstream recording target if no nodes exist in the FGS. Subsequently, a data unit header containing information about the corrected number of subsequent subgroups is written to the bitstream.
[0935] FIGS. 61(a), FIGS. 61(b) and FIGS. 62(a), FIGS. 62(b) are drawings for illustrating an ROI bounding box adjusted according to embodiments.
[0936] FIGS. 61(a) and FIGS. 61(b) show the case where the node size (nodeSizeLog2) of the parent node is 3, and FIGS. 62(a) and FIGS. 62(b) show the case where the node size (nodeSizeLog2) of the parent node is 2.
[0937] In partial decoding of FGS-based G-PCC bitstreams, to select an FGS belonging to an ROI, it is possible to determine whether each FGS overlaps with the ROI. In this case, when examining the region where the bounding boxes (Bboxes) of the ROI and the FGS overlap (ROI overlapped region), the determination can be made based on whether a node of the parent subgroup exists in the ROI overlapped region.
[0938] To determine whether to select the current FGS for decoding, node information from an already decoded parent subgroup can be utilized. If the parent subgroup's node position is within the ROI bounding box, the current FGS (or subgroup) can be selected. If the parent subgroup's node position is not within the ROI bounding box, the current FGS can be skipped.
[0939] At this time, there may be a difference between the geometry resolution (or unit geometry node size) used to set the ROI and the actual resolution (or node size) of the parent node. For example, as shown in Fig. 61(a), even though a portion of the parent node overlaps with the ROI, it may be determined that there is no ROI overlap region because the parent node position is not within the ROI bounding box. Alternatively, as shown in Fig. 62(a), even though a child node is included in the ROI, it may be determined that there is no ROI overlap region because the parent node position is not within the ROI bounding box. In this case, a problem may arise where FGS containing the child node is skipped by determining that there are no nodes in the ROI overlapped region with the parent subgroup bounding box. In other words, a problem may occur where points within the ROI cannot be decoded.
[0940] As a method to solve this, the resolution of the ROI can be matched to the resolution of the parent node, as shown in FIGS. 61(b) and FIGS. 62(b). That is, the ROI bounding box can be adjusted to be aligned with the boundary of the parent node based on the node size of the parent node.
[0941] In this case, the size of the ROI bounding box (ROI Bbox) can vary depending on the node size of the layer group being considered.
[0942] Referring to FIG. 61(a) and FIG. 61(b), when the parent subgroup is layer group 0, the ROI bounding box (ROI Bbox) is shifted right and then left according to the node size (nodeSizeLog2), and its size can be changed to match the voxel size of the parent node. To explain further, the minimum position of the ROI bounding box (ROI Bbox) is shifted right by 3 units according to the node size (nodeSizeLog2) of the parent node and then left, so that it can be aligned with the boundary of the parent node. That is, the minimum position of the ROI bounding box (ROI Bbox) can be adjusted to the minimum position of the parent node that includes the minimum position of the ROI bounding box (ROI Bbox). And the maximum position of the ROI bounding box (ROI Bbox) plus 1 can be aligned with the boundary of the parent node by performing a right shift of 3 and then a left shift according to the parent node's node size (nodeSizeLog2). That is, the maximum position of the ROI bounding box (ROI Bbox) can be adjusted to the maximum position of the parent node that includes the maximum position of the ROI bounding box (ROI Bbox). Based on the adjusted ROI bounding box as shown in Fig. 61(b), it can be determined that the parent node position exists within the ROI. That is, it can be determined that the parent node is occupied.
[0943] Figures 62(a) and 62(b) show the case where the parent subgroup exists in layer group 1, and since the node size (nodeSIzeLog2) is 2, the size of the ROI bounding box (ROI Bbox) can be adapted more finely compared to the above case.
[0944] Referring to FIG. 62(a), even though the child node is included in the ROI, the parent node position is not included in the ROI, so the FGS containing the child node may not be selected or decoded. To solve this problem, referring to FIG. 62(b), the ROI bounding box can be adjusted based on the size of the parent node so that the ROI bounding box can be adjusted to include all areas of the parent subgroup that partially overlap with the ROI bounding box. In the case of FIG. 62(a) and FIG. 62(b), the parent node size is 2, so the ROI can be adjusted more finely than in the case of FIG. 61(a) and FIG. 61(b).
[0945] In order to select and decode an FGS included in an ROI during partial decoding, the decoder according to the embodiments may first determine whether the FGS overlaps with the ROI. The area where the FGS and the ROI overlap may be referred to as the ROI overlap area. The decoder may also determine whether a node of the parent subgroup exists in the ROI overlap area.
[0946] In the tree structure according to the embodiments, a node represents a rectangular space, and when occupied, the tree node can identify the existence of at least one point contained within the volume of the axis-aligned rectangular space. The node size corresponds to the length of each axis and can be represented as an integer in the form of a power of 2. In the embodiments, the existence of a node may mean that the node is occupied and contains one or more points, and the non-existence of a node may mean that the node is not occupied and does not contain points.
[0947] In the embodiments, since the unit node size for setting the ROI may not match the node size of the parent subgroup area, the boundary of the ROI area may not match the boundary of the parent subgroup's node size. If a part of the parent subgroup area overlaps with the ROI, it may be necessary to adjust the ROI area. This is because, since the position of the parent subgroup node is set to a single coordinate (e.g., bottom corner), it may be determined that no node exists in the ROI overlap area even if a part of the occupied parent subgroup volume overlaps with the ROI; in such cases, a problem may occur where the child FGS is skipped even if it is included in the ROI.
[0948] Referring to FIGS. 61(a) and FIGS. 61(b), the parent node (PN) is included in layer group 0, and the node size (nodeSIzeLog2) is 3, i.e., 2 3 It can have a length of 8. And the child nodes are included in layer group 1, and the node size (nodeSIzeLog2) is 2, i.e., 2 2 It can have a length of 4. After decoding the parent node according to the embodiments, the decoder can determine whether the position of the parent node is included in the ROI. In the case of FIG. 61(a), the decoder can skip the parent subgroup because, even though part of the occupied parent subgroup is included in the ROI, the node position is not included in the ROI. In the case of FIG. 61(b), the position minimum of the ROI is adjusted to match the node boundary of the parent subgroup, so the decoder can determine that the occupied parent subgroup is included in the ROI.
[0949] Referring to FIGS. 62(a) and FIGS. 62(b), the ROI area can be adjusted based on the node size of the parent subgroup. If the parent node size (nodeSIzeLog2) is 2, i.e., 2 2 If it has a length of 4, the ROI can be adjusted in units of 4 and aligned to the boundaries of the parent node.
[0950] This explains the Selection of FGS.
[0951] If ROI bounding box minimum value (RoiBBoxMin) and ROI bounding box maximum value (RoiBBoxMax) exist, the bounding box of the region of interest (ROI) and the subgroup bounding box overlap, and the subgroup that occupies the overlapping area is selected to be decoded.
[0952] The code below represents the FGS selection process.
[0953] PrtDepth = 0
[0954] for (i=0; i ≤ PrtLayerGroupIdx; i++)
[0955] PrtDepth += num_layers_minus1[i] + 1
[0956] PrtNodeSizeLog2 = occtreeMaxDepthMinus1 + 1 - PrtDepth
[0957] for(k = 0; k < 3; k++) {
[0958] AdjustedRoiMin[k] = (RoiBBoxMin[k] >> PrtNodeSizeLog2) << PrtNodeSizeLog2
[0959] AdjustedRoiMin [k] |= (PrtNodeSizeLog2 > 1) << PrtNodeSizeLog2 - 1
[0960] AdjustedRoiMax[k] = (RoiBBoxMax[k]+1 >> PrtNodeSizeLog2) << PrtNodeSizeLog2
[0961] AdjustedRoiMax [k] |= (PrtNodeSizeLog2 > 1) << PrtNodeSizeLog2 - 1
[0962] }
[0963]
[0964] occupied = false
[0965] for(i=0; i< SubgroupNodeCnt[PrtLayerGroupIdx][PrtSubgroupIdx]; i++) {
[0966] for(k = 0; k < 3; k++)
[0967] pos[k] = SubgroupNodePos[PrtLayerGroupIdx][PrtSubgroupIdx ][i][k] << PrtNodeSizeLog2
[0968] pos [k] |= (PrtNodeSizeLog2> 1) << PrtNodeSizeLog2- 1
[0969] if (AdjustedRoiMin[0] ≤ pos[0] && AdjustedRoiMax[0] > pos[0]
[0970] AdjustedRoiMin[1] ≤ pos[1] && AdjustedRoiMax[1] > pos[1]
[0971] AdjustedRoiMin[2] ≤ pos[2] && AdjustedRoiMax[2] > pos[2]) {
[0972] occupied = true
[0973] break
[0974] }
[0975] }
[0976]
[0977] if (layerGroupIdx == 0)
[0978] decode GDU or ADU
[0979] else if (RoiBBoxMin < SubgroupBBoxMax[layerGroupIdx][subgroupIdx] &&
[0980] RoiBBoxMax > SubgroupBBoxMin[layerGroupIdx][subgroupIdx] && occupied)
[0981] decode DGDU or DADU
[0982] else
[0983] skip DGDU or DADU
[0984] The decoder according to the embodiments can derive the parent node depth (PrtDepth) by adding the number of layers (num_layers_minus1+1) from 0 to the parent layer group index (PrtLayerGroupIdx).
[0985] The decoder according to the embodiments can derive the parent node size (PrtNodeSizeLog2) based on the value obtained by subtracting the parent node depth (PrtDepth) from the maximum depth of the occupancy tree (occtreeMaxDepthMinus1 + 1).
[0986] The decoder according to the embodiments obtains the adjusted minimum value of the bounding box of the region of interest (AdjustedRoiMin) by adjusting the minimum value of the bounding box of the region of interest (RoiBBoxMin) for each x, y, and z axis based on the value regarding the parent node size (PrtNodeSizeLog2).
[0987] The decoder according to the embodiments can adjust the minimum value of the adjusted region of interest bounding box (AdjustedRoiMin) to the center coordinates of the node by adding half of the parent node size when the value regarding the parent node size (PrtNodeSizeLog2) is greater than 1.
[0988] The decoder according to the embodiments obtains the maximum value of the region of interest bounding box (AdjustedRoiMax) by adjusting the maximum value of the region of interest bounding box (RoiBBoxMax) or the maximum value of the region of interest bounding box (RoiBBoxMax) plus 1 based on the value regarding the parent node size (PrtNodeSizeLog2).
[0989] The decoder according to the embodiments can adjust the maximum value of the adjusted region of interest bounding box (AdjustedRoiMax) to the center coordinates of the node by adding half of the parent node size when the value regarding the parent node size (PrtNodeSizeLog2) is greater than 1.
[0990] To explain further, the decoder according to the embodiments obtains the maximum value of the region of interest bounding box adjusted for the x, y, and z axes by adjusting the maximum value of the region of interest bounding box based on the value regarding the parent node size.
[0991] The decoder according to the embodiments derives the position values of points for each x, y, and z axis for each subgroup identified by the parent subgroup index (PrtSubgroupIdx) and the parent layer group index (PrtLayerGroupIdx).
[0992] Here, the point's location value is derived by adjusting the node location of the subgroup identified by the parent subgroup index (PrtSubgroupIdx) and parent layer group index (PrtLayerGroupIdx) based on the parent node size.
[0993] If the point's position value is included within the range of the minimum and maximum values of the adjusted region of interest bounding box, and the point exists within the range, the Occupancy is derived as true from fals.
[0994] According to the embodiments, the ROI minimum value (RoiBBoxMin) and ROI maximum value (RoiBBoxMax) can be adjusted based on the parent node size (PrtNodeSizeLog2) (AdjustedRoiMin, AdjustedRoiMax). The ROI minimum value (RoiBBoxMin) can be adjusted by right-shifting by the parent node size (PrtNodeSizeLog2) and left-shifting by the parent node size. The adjusted ROI minimum value (AdjustedRoiMin) can be rounded down to a multiple of the parent node size so that the ROI boundary can be aligned with the parent node boundary. And the ROI maximum value (RoiBBoxMax) + 1 can be adjusted by right-shifting by the parent node size (PrtNodeSizeLog2) and left-shifting by the parent node size. Unlike the ROI minimum value (RoiBBoxMin), the ROI maximum value (RoiBBoxMax) can be processed by +1 to include the maximum value. And the AdjustedRoiMax can be aligned with the boundaries of the parent node.
[0995] num_layers_minus1 + 1 represents the number of partial occupancy tree depths for each layer group. occtreeMaxDepthMinus1 represents th...
Claims
1. A step of decoding geometry data of point cloud data within a bitstream; and A step of decoding attribute data of the above point cloud data; comprising Decoding method.
2. In Paragraph 1, The above geometry data is divided and included in subgroups of a layer group structure, and At least one of the above subgroups is a parent subgroup, and The above parent subgroup includes at least one child subgroup, and The subgroup bounding box of each of the above subgroups provides the boundaries of the nodes included in that subgroup, and A decoding method in which at least one child subgroup exists within the subgroup boundary of the parent subgroup.
3. In paragraph 2, the step of decoding the geometry data is, A decoding method for deriving the minimum point position and maximum point position of the subgroup bounding box of each child subgroup based on the minimum node size of the partial occupancy tree of the parent subgroup.
4. In paragraph 3, the step of decoding the geometry data is, A decoding method for selecting at least one subgroup based on the overlap between the bounding box of each subgroup and the region of interest, which is set based on the minimum point position and maximum point position of the subgroup bounding box of each child subgroup derived above, and decoding geometry data associated with the selected at least one subgroup.
5. In Paragraph 2, Each of the above child subgroups is identified by a pair of layer group index and subgroup index, and The above bitstream further includes signaling information, and The above signaling information includes origin information and size information of the subgroup bounding box of the subgroup identified by the pair of the layer group index and the subgroup index, and A decoding method in which the origin information and size information of the above subgroup bounding box are expressed in units of the minimum node size of the partial occupancy tree of the above parent subgroup.
6. Memory; and At least one processor connected to the memory; comprising, The above at least one processor is: Decoding geometry data of point cloud data within a bitstream; and Decoding attribute data of the above point cloud data; configured to do so, Decoding device.
7. In Paragraph 6, The above geometry data is divided and included in subgroups of a layer group structure, and At least one of the above subgroups is a parent subgroup, and The above parent subgroup includes at least one child subgroup, and The subgroup bounding box of each of the above subgroups provides the boundaries of the nodes included in that subgroup, and The above at least one child subgroup is a decoding device existing within the subgroup boundary of the above parent subgroup.
8. In paragraph 7, the above at least one processor, A decoding device that derives the minimum point position and the maximum point position of the subgroup bounding box of each child subgroup based on the minimum node size of the partial occupancy tree of the parent subgroup.
9. In paragraph 8, the above at least one processor, A decoding device that selects at least one subgroup based on the overlap between the bounding box of each subgroup and the region of interest, which is set based on the minimum point position and the maximum point position of the subgroup bounding box of each child subgroup derived above, and decodes geometry data associated with the selected at least one subgroup.
10. In Paragraph 7, Each of the above child subgroups is identified by a pair of layer group index and subgroup index, and The above bitstream further includes signaling information, and The above signaling information includes origin information and size information of the subgroup bounding box of the subgroup identified by the pair of the layer group index and the subgroup index, and A decoding device in which the origin information and size information of the above subgroup bounding box are expressed in units of the minimum node size of the partial occupancy tree of the above parent subgroup.
11. A step of encoding the geometry data of the point cloud data; and A step of encoding attribute data of the above point cloud data; comprising Encoding method.
12. In Paragraph 11, The above geometry data is divided and included in subgroups of a layer group structure, and At least one of the above subgroups is a parent subgroup, and The above parent subgroup includes at least one child subgroup, and The subgroup bounding box of each of the above subgroups provides the boundaries of the nodes included in that subgroup, and The above-mentioned at least one child subgroup is an encoding method existing within the subgroup boundary of the above-mentioned parent subgroup.
13. In Paragraph 12, Each of the above child subgroups is identified by a pair of layer group index and subgroup index, and The bitstream including the above-mentioned encoded geometry data and the above-mentioned encoded attribute data further includes signaling information, and The above signaling information includes origin information and size information of the subgroup bounding box of the subgroup identified by the pair of the layer group index and the subgroup index, and An encoding method in which the origin information and size information of the above subgroup bounding box are expressed in units of the minimum node size of the partial occupancy tree of the above parent subgroup.
14. Memory; and At least one processor connected to the memory; comprising, wherein the at least one processor: Encoding the geometry data of the point cloud data; and Configured to encode the attribute data of the above point cloud data; Encoding device.
15. A computer-readable storage medium for storing a bitstream generated by the method according to paragraph 11.
16. Step for acquiring a bitstream for point cloud data, The bitstream is generated based on the step of encoding geometry data of the point cloud data; and the step of encoding attribute data of the point cloud data; and A method comprising the step of transmitting data including the bitstream above.