Point cloud data transmission device, point cloud data transmission method, point cloud data reception device, and point cloud data reception method

By dividing point cloud data into multiple 3D spatial regions and using the V-PCC encoding method, the efficiency and complexity issues of point cloud data transmission and reception are solved, enabling low-latency, high-efficiency point cloud content access and rendering, and supporting services such as autonomous driving.

CN115668938BActive Publication Date: 2026-01-13LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202180035883.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-18
Filing Date
2021-01-19
Publication Date
2026-01-13
Estimated Expiration
2041-01-19

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently send and receive point cloud data, and suffer from latency and encoding/decoding complexity issues, thus failing to effectively provide optimized point cloud content.

Method used

By dividing point cloud data into multiple 3D spatial regions and sending signals in the signaling information to notify the spatial region information, including identification information, anchor point location information, tile identification information, priority and dependency information, a video-based point cloud compression (V-PCC) encoding and decoding method is adopted to achieve efficient encoding and decoding of point cloud data.

Benefits of technology

It achieves low-latency, high-efficiency point cloud data transmission and reception, supports services such as autonomous driving, provides optimized point cloud content access and rendering, and reduces the complexity of encoding and decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115668938B_ABST
    Figure CN115668938B_ABST
Patent Text Reader

Abstract

The point cloud data transmission method according to the embodiments can include the steps of encoding point cloud data, encapsulating a bitstream including the encoded point cloud data and signaling data into a file, and transmitting the file, wherein the bitstream includes geometry data, attribute data, occupancy map data, and atlas data, and is stored in a single track or multiple tracks of the file, and the signaling data can include spatial region information about the point cloud data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The implementation provides a method for providing point cloud content to offer users various services such as virtual reality (VR), augmented reality (AR), mixed reality (MR), and autonomous driving services. Background Technology

[0002] A point cloud is a collection of points in three-dimensional (3D) space. Because of the large number of points in 3D space, it is difficult to generate point cloud data.

[0003] High throughput is required to send and receive point cloud data. Summary of the Invention

[0004] Technical issues

[0005] The purpose of this disclosure is to provide a point cloud data transmitting device, a point cloud data transmitting method, a point cloud data receiving device, and a point cloud data receiving method for efficiently transmitting and receiving point clouds.

[0006] Another objective of this disclosure is to provide a point cloud data transmission device, a point cloud data transmission method, a point cloud data receiving device, and a point cloud data receiving method for solving latency and encoding / decoding complexity.

[0007] Another object of this disclosure is to provide a point cloud data transmitting device, a point cloud data transmitting method, a point cloud data receiving device, and a point cloud data receiving method for providing optimized point cloud content to a user by signaling viewport-related signaling of point cloud data.

[0008] Another object of this disclosure is to provide a point cloud data transmission apparatus, a point cloud data transmission method, a point cloud data receiving apparatus, and a point cloud data receiving method for providing optimized point cloud content to a user by allowing the transmission of viewport information, recommended viewport information, and initial viewing orientation (i.e., viewport) for data processing and rendering in a V-PCC bitstream.

[0009] Additional advantages, objects, and features of this disclosure will be set forth in part in the description which follows, and will in part become apparent to those skilled in the art upon examination of the following or may be learned from practice of this disclosure. The objects and other advantages of this disclosure may be realized and obtained from the structures specifically pointed out in the draft specification and its claims, as well as in the accompanying drawings.

[0010] Technical solution

[0011] To achieve these objectives and other advantages and in accordance with the purposes of this disclosure, as specifically implemented and broadly described herein, a method for transmitting point cloud data may include the steps of: encoding the point cloud data; encapsulating a bitstream including the encoded point cloud data into a file; and transmitting the file.

[0012] According to the implementation method, the point cloud data may include at least geometric data, attribute data or occupancy map data, the bit stream may be stored in multiple tracks of the file, the file may also include signaling data, and the signaling data may include spatial region information of the point cloud data.

[0013] According to the implementation method, point cloud data can be divided into one or more 3D spatial regions, and the spatial region information can include at least identification information for identifying each 3D spatial region or the location information of anchor points of each 3D spatial region.

[0014] According to the implementation method, spatial area information may be signaled in at least some or all of the sample entries of the orbits in the signaling information or in the sample of the metadata orbits associated with the orbits.

[0015] According to the implementation, the spatial region information that signals in the sample entry may also include tile identification information for identifying one or more tiles associated with each 3D spatial region.

[0016] According to the implementation method, the sample may also include priority information and dependency information related to the 3D spatial region.

[0017] According to an implementation, a point cloud data transmission device may include: an encoder for encoding point cloud data; a packager for packaged bit streams including the encoded point cloud data into a file; and a transmitter for transmitting the file.

[0018] According to the implementation method, the point cloud data may include at least geometric data, attribute data or occupancy map data, the bit stream may be stored in multiple tracks of the file, the file may also include signaling data, and the signaling data may include spatial region information of the point cloud data.

[0019] According to the implementation method, point cloud data can be divided into one or more 3D spatial regions, and the spatial region information can include at least identification information for identifying each 3D spatial region or the location information of anchor points of each 3D spatial region.

[0020] According to the implementation method, spatial area information may be signaled in at least some or all of the sample entries of the orbits in the signaling information or in the sample of the metadata orbits associated with the orbits.

[0021] According to the implementation, the spatial region information that signals in the sample entry may also include tile identification information for identifying one or more tiles associated with each 3D spatial region.

[0022] According to the implementation method, the sample may also include priority information and dependency information related to the 3D spatial region.

[0023] According to an implementation, a point cloud data receiving device may include: a receiver for receiving a file; a decapsulator for decapsulating the file into a bitstream including point cloud data, the bitstream being stored in multiple tracks of the file and the file also including signaling data; a decoder for decoding the point cloud data based on the signaling data; and a renderer for rendering the decoded point cloud data based on the signaling data.

[0024] According to the implementation method, point cloud data may include at least geometric data, attribute data or occupancy map data, and signaling data may include spatial region information of point cloud data.

[0025] According to the implementation method, point cloud data can be divided into one or more 3D spatial regions, and the spatial region information can include at least identification information for identifying each 3D spatial region or the location information of anchor points of each 3D spatial region.

[0026] According to the implementation method, spatial area information may be signaled in at least some or all of the sample entries of the orbits in the signaling information or in the sample of the metadata orbits associated with the orbits.

[0027] According to the implementation, the spatial region information that signals in the sample entry may also include tile identification information for identifying one or more tiles associated with each 3D spatial region.

[0028] According to the implementation method, the sample may also include priority information and dependency information related to the 3D spatial region.

[0029] Beneficial effects

[0030] The point cloud data transmission method, point cloud data transmission device, point cloud data reception method, and point cloud data reception device according to the embodiments can provide point cloud services of good quality.

[0031] The point cloud data transmission method, point cloud data transmission device, point cloud data reception method, and point cloud data reception device according to the embodiments can implement various video encoding and decoding methods.

[0032] The point cloud data transmission method, point cloud data transmission device, point cloud data reception method, and point cloud data reception device according to the embodiments can provide general point cloud content such as autonomous driving services.

[0033] By utilizing the point cloud data transmission method, point cloud data transmission apparatus, point cloud data reception method, and point cloud data reception apparatus according to the embodiments, a V-PCC bitstream can be configured, and files can be sent, received, and stored. Therefore, optimal point cloud content services can be provided.

[0034] By utilizing the point cloud data transmission method, point cloud data transmission apparatus, point cloud data reception method, and point cloud data reception apparatus according to the embodiments, metadata for data processing and rendering in the V-PCC bitstream can be transmitted and received in the V-PCC bitstream. Therefore, optimal point cloud content services can be provided.

[0035] By utilizing the point cloud data transmission method, point cloud data transmission device, point cloud data reception method, and point cloud data reception device according to the embodiments, players and the like can achieve spatial or partial access to point cloud objects / content based on the user's viewport. Therefore, point cloud bitstreams can be accessed and processed efficiently based on the user's viewport.

[0036] Using the point cloud data transmission method, point cloud data transmission apparatus, point cloud data reception method, and point cloud data reception apparatus according to embodiments, a bounding box for partial and / or spatial access to point cloud content and its signaling information can be provided. Therefore, it is possible to access point cloud content in various ways on the receiving side, assuming either a player environment or a user environment.

[0037] The point cloud data transmission method and transmission apparatus according to the embodiments can provide 3D region information about the point cloud content and 2D region-related metadata in the associated video or Atlas frame to support spatial / partial access to the point cloud content according to the user viewport.

[0038] The point cloud data transmission method and transmission apparatus according to the embodiments can process signaling regarding 3D region information of point clouds in point cloud bitstreams and related 2D region metadata in video or Atlas frames associated with them.

[0039] By utilizing the point cloud data receiving method and receiving apparatus according to the embodiments, point cloud content can be efficiently accessed based on the storage and signaling of 3D region information of the point cloud in the point cloud bitstream and 2D region-related metadata in the associated video or Atlas frame.

[0040] Using the point cloud data receiving method and receiving apparatus according to the embodiments, point cloud content can be provided based on 3D region information of point clouds associated with image items in a file and 2D region information in video or Atlas frames associated with them, taking into account the user environment.

[0041] The point cloud data transmission method and apparatus according to the embodiments, as well as the point cloud data receiving method and apparatus, process point cloud data by dividing space into multiple 3D spatial regions, thereby performing encoding and transmission operations in real time on the transmitting side and decoding and rendering operations on the receiving side, and processing these operations with low latency.

[0042] The point cloud data transmission method according to the embodiments encodes spatially segmented 3D blocks (e.g., 3D spatial regions) independently or dependently, thereby enabling parallel encoding and random access in the 3D space occupied by the point cloud data.

[0043] The point cloud data transmission method and apparatus according to the embodiments, as well as the point cloud data reception method and apparatus, perform point cloud data encoding and decoding independently or dependently on spatially divided 3D blocks (e.g., 3D spatial regions), thereby preventing errors from accumulating during the encoding and decoding process. Attached Figure Description

[0044] The accompanying drawings are included to provide a further understanding of this disclosure and are incorporated in and constitute a part of this application. The drawings illustrate embodiments of the disclosure and, together with the description, serve to illustrate the principles of the disclosure. In the drawings:

[0045] Figure 1 An exemplary structure of a sending / receiving system for providing point cloud content according to an embodiment is shown.

[0046] Figure 2 The capture of point cloud data according to an embodiment is shown.

[0047] Figure 3 Exemplary point cloud, geometry, and texture images are shown according to an embodiment.

[0048] Figure 4 An exemplary V-PCC encoding process according to an implementation is shown.

[0049] Figure 5 An example of the tangent plane and normal vector of a surface according to an embodiment is shown.

[0050] Figure 6 An exemplary bounding box of a point cloud according to an implementation method is shown.

[0051] Figure 7 An example of determining the position of each patch on the occupancy map according to an embodiment is shown.

[0052] Figure 8 An exemplary relationship between the normal axis, tangential axis, and double tangential axis is shown according to an embodiment.

[0053] Figure 9Exemplary configurations of the minimum and maximum modes of the projection mode according to the implementation are shown.

[0054] Figure 10 An exemplary EDD code according to an implementation method is shown.

[0055] Figure 11 An example of recoloring based on the color values ​​of neighboring points according to an implementation method is shown.

[0056] Figure 12 An example of a push-pull background fill according to an implementation method is shown.

[0057] Figure 13 An exemplary possible traversal order of a 4x4 block according to an implementation is shown.

[0058] Figure 14 An exemplary optimal traversal order according to the implementation method is shown.

[0059] Figure 15 An exemplary 2D video / image encoder according to an embodiment is shown.

[0060] Figure 16 An exemplary V-PCC decoding process according to an implementation is shown.

[0061] Figure 17 An exemplary 2D video / image decoder according to an embodiment is shown.

[0062] Figure 18 This is a flowchart illustrating the operation of a transmitting device according to an embodiment of the present disclosure.

[0063] Figure 19 This is a flowchart illustrating the operation of the receiving device according to an embodiment.

[0064] Figure 20 An exemplary architecture for V-PCC-based storage and streaming of point cloud data, according to an embodiment, is shown.

[0065] Figure 21 This is an exemplary block diagram of an apparatus for storing and transmitting point cloud data according to an embodiment.

[0066] Figure 22 This is an exemplary block diagram of a point cloud data receiving device according to an embodiment.

[0067] Figure 23 An exemplary structure is shown that can be operated in conjunction with a point cloud data transmission / reception method / apparatus according to an embodiment.

[0068] Figure 24An exemplary association is shown between a portion of a 3D region of point cloud data and one or more 2D regions in a video frame, according to an embodiment.

[0069] Figure 25 An exemplary structure of a V-PCC bitstream according to an embodiment is shown;

[0070] Figure 26 This illustrates exemplary data carried by a sample stream V-PCC cell in a V-PCC bitstream according to an embodiment;

[0071] Figure 27 An exemplary syntax structure of a sample stream V-PCC header included in a V-PCC bitstream according to an embodiment is shown;

[0072] Figure 28 An exemplary syntax structure of a sample stream V-PCC unit according to an implementation is shown;

[0073] Figure 29 An exemplary syntax structure of a V-PCC unit according to an embodiment is shown;

[0074] Figure 30 An exemplary syntax structure of the V-PCC unit header according to an implementation is shown;

[0075] Figure 31 An exemplary type of V-PCC unit assigned to the vuh_unit_type field according to an implementation is shown;

[0076] Figure 32 An exemplary syntax structure for a V-PCC cell payload according to an embodiment is shown;

[0077] Figure 33 An exemplary syntax structure for the V-PCC parameter set according to an implementation is shown;

[0078] Figure 34 An exemplary structure of an Atlas substream according to an implementation is shown;

[0079] Figure 35 An exemplary syntax structure of a sample stream NAL header included in an Atlas substream is shown according to an implementation.

[0080] Figure 36 An exemplary syntax structure of a sample stream NAL unit according to an implementation is shown;

[0081] Figure 37 An exemplary syntax structure for the Atlas sequence parameter set according to an implementation is shown;

[0082] Figure 38An exemplary syntax structure for the Atlas frame parameter set according to an implementation is shown;

[0083] Figure 39 An exemplary syntax structure for Atlas frame chunk information according to an implementation is shown;

[0084] Figure 40 An exemplary syntax structure for Supplemental Enhancement Information (SEI) according to an implementation method is shown;

[0085] Figure 41 An exemplary syntax structure for 3D bounding box information SEI according to an implementation is shown;

[0086] Figure 42 An exemplary syntax structure for 3D region mapping information (SEI) according to an implementation method is shown;

[0087] Figure 43 An exemplary syntax structure for volumetric tile information (SEI) according to an implementation method is shown;

[0088] Figure 44 An exemplary syntax structure for volume tile information tagging information according to an embodiment is shown;

[0089] Figure 45 An exemplary syntax structure for volume tile information object information according to an implementation method is shown;

[0090] Figure 46 An exemplary structure of a V-PCC sample entry according to an implementation method is shown;

[0091] Figure 47 An exemplary structure of a moov box according to an embodiment and an exemplary structure of a sample entry are shown;

[0092] Figure 48 Exemplary track alternatives and track groupings according to an implementation method are shown;

[0093] Figure 49 An exemplary structure for encapsulating non-timing V-PCC data according to an embodiment is shown;

[0094] Figure 50 The overall structure of the sample stream V-PCC unit according to the embodiment is shown;

[0095] Figure 51 This is a flowchart of a file-level signaling method according to an implementation method;

[0096] Figure 52 This is a flowchart of a signaling information acquisition method in a receiving device according to an embodiment;

[0097] Figure 53This is a flowchart of a point cloud data transmission method according to an implementation method; and

[0098] Figure 54 This is a flowchart of a point cloud data receiving method according to an implementation method. Detailed Implementation

[0099] Preferred embodiments of the present disclosure will now be described in detail, examples of which are illustrated in the accompanying drawings. The detailed description given below with reference to the drawings is intended to illustrate exemplary embodiments of the present disclosure, and not to show only embodiments that can be implemented according to the present disclosure. The following detailed description includes specific details in order to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure can be practiced without these specific details.

[0100] Although most of the terms used in this disclosure are selected from commonly used terms in the art, some terms have been arbitrarily chosen by the applicant and their meanings are explained in detail in the following description as needed. Therefore, this disclosure should be understood based on the intended meaning of the terms rather than their simple names or meanings.

[0101] Figure 1 An exemplary structure of a sending / receiving system for providing point cloud content according to an embodiment is shown.

[0102] This disclosure provides a method for providing point cloud content to offer users various services such as virtual reality (VR), augmented reality (AR), mixed reality (MR), and autonomous driving. The point cloud data representation according to the embodiments represents objects as points and may be referred to as point cloud, point cloud data, point cloud video data, point cloud image data, etc.

[0103] The point cloud data transmission device 10000 according to an embodiment may include a point cloud video acquisition unit 10001, a point cloud video encoder 10002, a file / fragment encapsulation module (file / fragment encapsulator) 10003, and / or a transmitter (or communication module) 10004. The transmission device according to an embodiment can acquire and process point cloud video (or point cloud content) and transmit it. According to an embodiment, the transmission device may include a fixed station, a base transceiver system (BTS), a network, an artificial intelligence (AI) device and / or system, a robot, and an AR / VR / XR device and / or server. According to an embodiment, the transmission device 10000 may include devices configured to communicate with base stations and / or other wireless devices using radio access technologies (e.g., 5G New RAT (NR), Long Term Evolution (LTE)), robots, vehicles, AR / VR / XR devices, portable devices, home appliances, Internet of Things (IoT) devices, and AI devices / servers.

[0104] According to the embodiment, the point cloud video acquisition unit 10001 acquires point cloud video through processes of capturing, synthesizing, or generating point cloud video.

[0105] The point cloud video encoder 10002 according to the embodiment encodes the point cloud video data acquired from the point cloud video acquisition unit 10001. According to the embodiment, the point cloud video encoder 10002 may be referred to as a point cloud encoder, point cloud data encoder, encoder, etc. The point cloud compression encoding (encoding) according to the embodiment is not limited to the above embodiment. The point cloud video encoder can output a bitstream including the encoded point cloud video data. The bitstream may include not only the encoded point cloud video data, but also signaling information related to the encoding of the point cloud video data.

[0106] The point cloud video encoder 10002 according to the embodiments can support geometry-based point cloud compression (G-PCC) encoding schemes and / or video-based point cloud compression (V-PCC) encoding schemes. Furthermore, the point cloud video encoder 10002 can encode point clouds (referred to as point cloud data or points) and / or signaling data associated with point clouds.

[0107] The term V-PCC used in this paper to refer to video-based point cloud compression has the same meaning as visual volumetric video coding (V3C), and they can be used in conjunction with each other.

[0108] The file / fragment encapsulation module 10003 according to the embodiment encapsulates point cloud data in the form of files and / or fragments. The point cloud data transmission method / apparatus according to the embodiment can transmit point cloud data in the form of files and / or fragments.

[0109] According to the embodiment, the transmitter (or communication module) 10004 transmits encoded point cloud video data in the form of a bitstream. According to the embodiment, files or segments can be transmitted to a receiving device via a network or stored in a digital storage medium (e.g., USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.). The transmitter according to the embodiment is capable of wired / wireless communication with the receiving device (or receiver) via networks such as 4G, 5G, and 6G. Furthermore, the transmitter can perform necessary data processing operations according to the network system (e.g., a 4G, 5G, or 6G communication network system). The transmitting device can transmit encapsulated data on demand.

[0110] The point cloud data receiving device 10005 according to an embodiment may include a receiver 10006, a file / fragment decapsulator (or file / fragment decapsulator module) 10007, a point cloud video decoder 10008, and / or a renderer 10009. According to an embodiment, the receiving device may include devices, robots, vehicles, AR / VR / XR devices, portable devices, home appliances, Internet of Things (IoT) devices, and AI devices / servers configured to communicate with base stations and / or other wireless devices using radio access technologies (e.g., 5G New RAT (NR), Long Term Evolution (LTE)).

[0111] According to an embodiment, receiver 10006 receives a bitstream containing point cloud video data. According to an embodiment, receiver 10006 can send feedback information to point cloud data transmitting device 10000.

[0112] The file / fragment decapsulation module 10007 decapsulates files and / or fragments containing point cloud data.

[0113] The point cloud video decoder 10008 decodes the received point cloud video data.

[0114] Renderer 10009 renders the decoded point cloud video data. According to one embodiment, renderer 10009 can send feedback information obtained at the receiving side to point cloud video decoder 10008. The point cloud video data, according to one embodiment, can carry the feedback information to receiver 10006. According to one embodiment, the feedback information received by the point cloud transmitting device can be provided to point cloud video encoder 10002.

[0115] The arrows indicated by dashed lines in the diagram represent the transmission paths of the feedback information acquired by the receiving device 10005. The feedback information reflects the interactivity of the user consuming the point cloud content and includes user information (e.g., head orientation information, viewport information, etc.). Specifically, when the point cloud content is for a service requiring user interaction (e.g., autonomous driving service, etc.), the feedback information can be provided to the content sender (e.g., the sending device 10000) and / or the service provider. According to embodiments, the feedback information can be used in both the receiving device 10005 and the sending device 10000, and may not be provided at all.

[0116] According to the embodiment, head orientation information is information about the user's head position, orientation, angle, movement, etc. The receiving device 10005 according to the embodiment can calculate viewport information based on the head orientation information. Viewport information can be information about the area of ​​the point cloud video that the user is viewing. The viewpoint (or orientation) is the point from which the user views the point cloud video and can refer to the center point of the viewport area. That is, the viewport is the area centered on the viewpoint, and the size and shape of the area can be determined by the field of view (FOV). In other words, the viewport is determined based on the position and viewpoint (or orientation) of the visual camera or the user, and the point cloud data is rendered in the viewport based on the viewport information. Therefore, in addition to head orientation information, the receiving device 10005 can also extract viewport information based on the vertical or horizontal FOV supported by the device. Furthermore, the receiving device 10005 performs gaze analysis to examine how the user consumes the point cloud, the area the user gazes at in the point cloud video, the gaze duration, etc. According to the embodiment, the receiving device 10005 can send feedback information, including the gaze analysis results, to the transmitting device 10000. The feedback information according to the embodiments can be obtained during the rendering and / or display process. The feedback information according to the embodiments can be obtained by one or more sensors included in the receiving device 10005. Furthermore, according to the embodiments, the feedback information can be obtained by the renderer 10009 or by a separate external component (or device, assembly, etc.). Figure 1 The dashed lines in the diagram represent the process of sending feedback information obtained by the renderer 10009. The point cloud content providing system can process (encode / decode) point cloud data based on the feedback information. Therefore, the point cloud video decoder 10008 can perform decoding operations based on the feedback information. The receiving device 10005 can send feedback information to the transmitting device. The transmitting device (or the point cloud video encoder 10002) can perform encoding operations based on the feedback information. Therefore, the point cloud content providing system can efficiently process necessary data (e.g., point cloud data corresponding to the user's head position) based on the feedback information instead of processing (encoding / decoding) all point cloud data, and provide point cloud content to the user.

[0117] According to the implementation, the transmitting device 10000 may be referred to as an encoder, transmitting device, transmitter, etc., and the receiving device 10005 may be referred to as a decoder, receiving device, receiver, etc.

[0118] According to the implementation method Figure 1 Point cloud data processed in a point cloud content provision system (through a series of processes including acquisition, encoding, transmission, decoding, and rendering) can be referred to as point cloud content data or point cloud video data. Depending on the implementation, point cloud content data can be used as a concept encompassing metadata or signaling information related to point cloud data.

[0119] Figure 1The components of the point cloud content provided by the system can be implemented by hardware, software, processors, and / or combinations thereof.

[0120] Implementations may provide a method for providing point cloud content to offer users various services such as virtual reality (VR), augmented reality (AR), mixed reality (MR), and autonomous driving.

[0121] To provide point cloud content services, point cloud video can first be acquired. The acquired point cloud video can be sent to the receiving side through a series of processes, and the receiving side can process the received data back into the original point cloud video and render the processed point cloud video. Thus, the point cloud video can be provided to the user. This implementation provides a method for efficiently performing this series of processes.

[0122] All processing used to provide point cloud content services (point cloud data sending methods and / or point cloud data receiving methods) may include acquisition processing, encoding processing, transmission processing, decoding processing, rendering processing, and / or feedback processing.

[0123] According to an implementation, the processing of providing point cloud content (or point cloud data) can be referred to as point cloud compression processing. According to an implementation, point cloud compression processing can represent video-based point cloud compression (V-PCC) processing.

[0124] The individual components of the point cloud data transmitting device and the point cloud data receiving device according to the embodiments may be hardware, software, processor and / or combinations thereof.

[0125] A point cloud compression system may include a transmitting device and a receiving device. According to different embodiments, the transmitting device may be referred to as an encoder, transmitting device, transmitter, point cloud data transmitting device, etc. According to different embodiments, the receiving device may be referred to as a decoder, receiving device, receiver, point cloud data receiving device, etc. The transmitting device can encode the point cloud video to output a bitstream and transmit it to the receiving device as a file or stream (streaming segment) via a digital storage medium or network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD.

[0126] The transmitting device may include a point cloud video acquisition unit, a point cloud video encoder, a file / segment encapsulator, and a transmitting unit (transmitter), such as Figure 1 As shown. The receiving device may include a receiver, a file / fragment decapsulator, a point cloud video decoder, and a renderer, such as... Figure 1As shown. The encoder may be referred to as a point cloud video / image / frame encoder, and the decoder may be referred to as a point cloud video / image / frame decoder. The renderer may include a display. The renderer and / or display may be configured as separate devices or external components. The transmitting and receiving devices may also include separate internal or external modules / units / components for feedback processing. According to the implementation, each element in the transmitting and receiving devices may be configured by hardware, software, and / or a processor.

[0127] According to the implementation method, the operation of the receiving device can be the reverse processing of the operation of the transmitting device.

[0128] Point cloud video acquirers can perform point cloud video acquisition processing by capturing, orchestrating, or generating point cloud video. During acquisition processing, 3D position (x, y, z) / attribute (color, reflectivity, transparency, etc.) data for multiple points can be generated, such as polygon file format (PLY) (or Stanford triangle format) files. For videos with multiple frames, one or more files can be acquired. Point cloud-related metadata (e.g., capture-related metadata) can be generated during capture processing.

[0129] The point cloud data transmission apparatus according to the embodiments may include an encoder configured to encode point cloud data and a transmitter configured to transmit point cloud data or a bit stream including point cloud data.

[0130] The point cloud data receiving apparatus according to the embodiments may include a receiver configured to receive a bit stream including point cloud data, a decoder configured to decode the point cloud data, and a renderer configured to render the point cloud data.

[0131] The method / apparatus according to the embodiments represents a point cloud data transmitting device and / or a point cloud data receiving device.

[0132] Figure 2 The capture of point cloud data according to an embodiment is shown.

[0133] The point cloud data (point cloud video data) according to the embodiments can be acquired by a camera or the like. The capture techniques according to the embodiments may include, for example, inward-facing and / or outward-facing.

[0134] In the inward orientation according to the implementation, one or more cameras facing the point cloud data of the object can capture images of the object from outside the object.

[0135] In the outward-facing configuration according to the embodiment, one or more cameras can capture the object of the point cloud data. For example, according to the embodiment, four cameras may be present.

[0136] According to the embodiments, point cloud data or point cloud content can be video or still images of objects / environments represented in various types of 3D space. According to the embodiments, point cloud content can include video / audio / images of objects.

[0137] As a device for capturing point cloud content, a combination of a camera device capable of acquiring depth (a combination of an infrared pattern projector and an infrared camera) and an RGB camera capable of extracting color information corresponding to the depth information can be configured. Alternatively, depth information can be extracted using a LiDAR radar system that measures the position coordinates of a reflector by emitting laser pulses and measuring the return time. Geometry composed of points in 3D space can be extracted from the depth information, and attributes representing the color / reflectivity of each point can be extracted from the RGB information. Point cloud content can include information about position (x, y, z) and the color (YCbCr or RGB) or reflectivity (r) of the points. For point cloud content, outward-facing techniques for capturing the external environment and inward-facing techniques for capturing the central object can be used. In VR / AR environments, when an object (e.g., a core object such as a character, player, thing, or actor) is configured within point cloud content that the user can view from any direction (360 degrees), the configuration of the capture camera can be based on inward-facing techniques. When the current surrounding environment is configured within the point cloud content in vehicle modes such as autonomous driving, the configuration of the capture camera can be based on outward-facing techniques. Since point cloud content can be captured by multiple cameras, camera calibration may be required to configure the camera's global coordinate system before capturing the content.

[0138] Point cloud content can be video or still images of objects / environments existing in various types of 3D space.

[0139] Furthermore, in point cloud content acquisition methods, any point cloud video can be orchestrated based on the captured point cloud video. Alternatively, when providing point cloud video of a computer-generated virtual space, capture using an actual camera may not be performed. In this case, the capture processing can be simply replaced by processing that generates relevant data.

[0140] Post-processing of captured point cloud video may be necessary to improve content quality. During video capture processing, the maximum / minimum depth can be adjusted within the range provided by the camera device. Even after adjustment, unwanted areas of point data may still exist. Therefore, post-processing can be performed to remove unwanted areas (e.g., background) or to identify connected spaces and fill in spatial holes. Additionally, point clouds extracted from cameras in a shared spatial coordinate system can be integrated into a single piece of content by transforming individual points to a global coordinate system based on the position coordinates of each camera obtained through calibration processing. This can generate a single point cloud content with a wide range, or it can acquire point cloud content with high-density points.

[0141] The point cloud video encoder 10002 can encode an input point cloud video into one or more video streams. A point cloud video can include multiple frames, each frame corresponding to a still image / picture. In this specification, point cloud video can include point cloud images / frames / pictures / video / audio. Additionally, the term "point cloud video" is used interchangeably with point cloud images / frames / pictures. The point cloud video encoder 10002 can perform video-based point cloud compression (V-PCC) processing. For compression and encoding efficiency, the point cloud video encoder can perform a series of processes such as prediction, transform, quantization, and entropy coding. The encoded data (encoded video / image information) can be output as a bitstream. Based on V-PCC processing, the point cloud video encoder can encode the point cloud video by dividing it into geometric video, attribute video, occupancy map video, and auxiliary information (or auxiliary data) (described later). Geometric video can include geometric images, attribute video can include attribute images, and occupancy map video can include occupancy map images. Auxiliary information can include auxiliary patch information. Attribute video / images can include texture video / images.

[0142] The file / fragment encapsulator (file / fragment encapsulation module) 10003 can encapsulate encoded point cloud video data and / or metadata related to the point cloud video in, for example, the form of a file. Here, the metadata related to the point cloud video can be received from a metadata processor. The metadata processor can be included in the point cloud video encoder 10002, or can be configured as a separate component / module. The file / fragment encapsulator 10003 can encapsulate data in a file format such as ISOBMFF or process data in the form of DASH fragments, etc. According to an embodiment, the file / fragment encapsulator 10003 can include point cloud video-related metadata in a file format. The point cloud video metadata can be included in various levels of boxes in, for example, the ISOBMFF file format, or as data in a separate track within a file. According to an embodiment, the file / fragment encapsulator 10003 can encapsulate point cloud video-related metadata into a file. The transmission processor can perform transmission processing on the point cloud video data encapsulated according to the file format. The transmission processor can be included in the transmitter 10004, or can be configured as a separate component / module. The transmission processor can process the point cloud video data according to a transmission protocol. The transmission process may include processing transmitted via a broadcast network and processing transmitted via broadband. According to one implementation, the transmitting processor may receive point cloud video-related metadata from a metadata processor along with the point cloud video data, and perform processing on the point cloud video data for transmission.

[0143] Transmitter 10004 can transmit encoded video / image information or data, output in bitstream form, to receiver 10006 of receiving device in the form of a file or stream via digital storage medium or network. Digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitter may include elements for generating media files in a predetermined file format and may include elements for transmission via broadcast / communication network. Receiver can extract the bitstream and send the extracted bitstream to decoding device.

[0144] Receiver 10006 can receive point cloud video data transmitted by the point cloud video transmitting device according to this disclosure. Depending on the transmission channel, the receiver can receive the point cloud video data via a broadcast network or via broadband. Alternatively, the point cloud video data can be received via a digital storage medium.

[0145] The receiving processor can process the received point cloud video data according to the transmission protocol. The receiving processor can be included in the receiver 10006 or configured as a separate component / module. The receiving processor can reverse the above-described processing of the transmitting processor, such that the processing corresponds to the transmission processing performed on the transmitting side. The receiving processor can transmit the acquired point cloud video data to the file / fragment decapsulator 10007 and transmit the acquired point cloud video-related metadata to the metadata processor (not shown). The point cloud video-related metadata acquired by the receiving processor can be in the form of a signaling table.

[0146] The file / fragment decapsulator (file / fragment decapsulation module) 10007 can decapsulate point cloud video data received from the receiver processor in file form. The file / fragment decapsulator 10007 can decapsulate files according to ISOBMFF or similar formats and can obtain point cloud video bitstreams or point cloud video-related metadata (metadata bitstreams). The obtained point cloud video bitstream can be transmitted to the point cloud video decoder 10008, and the obtained point cloud video-related metadata (metadata bitstreams) can be transmitted to a metadata processor (not shown). The point cloud video bitstream may include metadata (metadata bitstreams). The metadata processor may be included in the point cloud video decoder 10008 or can be configured as a separate component / module. The point cloud video-related metadata obtained by the file / fragment decapsulator 10007 may take the form of boxes or tracks in a file format. When needed, the file / fragment decapsulator 10007 can receive the metadata required for decapsulation from the metadata processor. The metadata related to point cloud video can be sent to point cloud video decoder 10008 and used in point cloud video decoding processing, or it can be sent to renderer 10009 and used in point cloud video rendering processing.

[0147] The point cloud video decoder 10008 can receive bitstreams and decode video / images by performing operations corresponding to those of the point cloud video encoder. In this case, the point cloud video decoder 10008 can decode the point cloud video by dividing it into geometric video, attribute video, occupancy map video, and auxiliary information, as described below. Geometric video may include geometric images, attribute video may include attribute images, occupancy map video may include occupancy map images, and auxiliary information may include auxiliary patch information. Attribute video / images may include texture video / images.

[0148] 3D geometry can be reconstructed based on decoded geometric images, occupancy maps, and auxiliary patch information, and then subjected to smoothing. Color point cloud images / pictures can be reconstructed by assigning color values ​​to the smoothed 3D geometry based on texture images. Renderer 10009 can render the reconstructed geometry and color point cloud images / pictures. The rendered video / images can be displayed on a monitor (not shown). Users can view all or part of the rendering results via VR / AR displays or typical displays.

[0149] Feedback processing may include a decoder that transmits various types of feedback information, which can be obtained during rendering / display processing, to the sending or receiving side. Interactivity can be provided through feedback processing when consuming point cloud video. According to one embodiment, head orientation information, viewport information indicating the area the user is currently viewing, etc., may be transmitted to the sending side during feedback processing. According to another embodiment, the user can interact with objects realized in a VR / AR / MR / autonomous driving environment. In this case, information related to the interaction may be transmitted to the sending side or service provider during feedback processing. According to yet another embodiment, feedback processing may be skipped.

[0150] Head orientation information represents the position, angle, and movement of the user's head. Based on this information, information about the region of the point cloud video currently being viewed by the user (i.e., viewport information) can be calculated.

[0151] Viewport information can be information about the region of a point cloud video currently being viewed by the user. Viewport information can be used to perform gaze analysis to examine how the user consumes the point cloud video, the region of the point cloud video the user is gazing at, and how long the user is gazing at that region. Gaze analysis can be performed on the receiving side, and the analysis results can be transmitted to the transmitting side via a feedback channel. Devices such as VR / AR / MR displays can extract the viewport region based on the user's head position / orientation, the vertical or horizontal FOV supported by the device, etc.

[0152] According to the implementation method, the aforementioned feedback information can be transmitted not only to the sending side but also consumed at the receiving side. That is, decoding and rendering processing at the receiving side can be performed based on the aforementioned feedback information. For example, point cloud video of the area currently being viewed by the user can be decoded and rendered preferentially only based on head orientation information and / or viewport information.

[0153] Here, the viewport or viewport region can represent the area of ​​the point cloud video currently being viewed by the user. The viewpoint is the point in the point cloud video that the user is viewing, and can represent the center point of the viewport region. That is, the viewport is the area surrounding the viewpoint, and the size and shape of the area can be determined by the field of view (FOV).

[0154] This disclosure relates to point cloud video compression as described above. For example, the methods / implementations disclosed in this disclosure can be applied to the Moving Picture Experts Group (MPEG) point cloud compression or point cloud coding (PCC) standard or next-generation video / image coding standards.

[0155] As used in this article, a picture / frame can typically represent a unit representing an image within a specific time interval.

[0156] A pixel, or image unit, can be the smallest unit that makes up a picture (or image). Additionally, the term "sample" can be used as the counterpart to a pixel. A sample can typically represent a pixel or pixel value. It can represent only the pixel / pixel value of the luminance component, only the pixel / pixel value of the chrominance component, or only the pixel / pixel value of the depth component.

[0157] A unit can represent a basic unit of image processing. A unit may include a specific region of an image and at least one of the information associated with that region. In some cases, a unit may be used interchangeably with terms such as block, region, or module. In general, an M×N block may include samples (or sample arrays) or a set (or array) of transform coefficients arranged in M ​​columns and N rows.

[0158] Figure 3 Examples of point clouds, geometric images, and texture images according to embodiments are shown.

[0159] The point cloud, according to the implementation method, can be input into what will be described below. Figure 4 The V-PCC encoding process generates geometric and texture images. Depending on the implementation, the point cloud can have the same meaning as the point cloud data.

[0160] exist Figure 3 The left image shows a point cloud that can represent point cloud objects in 3D space as bounding boxes. Figure 3The middle image shows the geometric image, and the right image shows the texture image (non-filled). That is, the 3D bounding box can be specified as a volume defined as a cube with six rectangular faces placed at right angles. In this specification, the geometric image is also referred to as a geometric patch frame / picture or geometry frame / picture. Similarly, the texture image is also referred to as an attribute patch frame / picture or attribute frame / picture.

[0161] Video-based point cloud compression (V-PCC), according to the implementation method, is a method for compressing 3D point cloud data based on 2D video codecs such as High Efficiency Video Coding (HEVC) or Multi-Functional Video Coding (VVC). The data and information that can be generated during V-PCC compression processing are as follows:

[0162] Occupancy map: This is a binary map that uses values ​​of 0 or 1 to indicate whether data exists at corresponding locations in the 2D plane when the points that make up a point cloud are divided into patches and mapped to a 2D plane. The occupancy map can represent a 2D array corresponding to an atlas, and the values ​​of the occupancy map can indicate whether each sample location in the atlas corresponds to a 3D point.

[0163] An Atlas is composed of patches and refers to an object that includes information about the 2D patches for each point cloud frame. For example, an Atlas may include the 2D arrangement and size of the patches, the position of the corresponding 3D regions within the 3D points, the projection plane, and level of detail parameters. In other words, an Atlas can be divided into patch bundles of the same size.

[0164] Furthermore, an Atlas is a collection of 2D bounding boxes and associated information, which is placed in a rectangular frame and corresponds to a 3D bounding box (i.e., volume) in 3D space where volume data is rendered.

[0165] An Atlas bitstream is a sequence of bits that forms the representation of one or more Atlas frames that constitute an Atlas.

[0166] An Atlas frame is a 2D rectangular array of Atlas samples, onto which patches are projected.

[0167] An Atlas sample is the position of a rectangular frame onto which a patch associated with Atlas is projected.

[0168] An Atlas sequence is a collection of Atlas frames.

[0169] According to the implementation method, an Atlas frame can be segmented into tiles. A tile is a unit for segmenting a 2D frame. That is, a tile is a unit used to segment signaling information in point cloud data called Atlas.

[0170] A patch is a set of points that make up a point cloud. Points belonging to the same patch are adjacent to each other in 3D space and are mapped in the same direction in the six bounding box planes during the mapping to a 2D image. A patch is the unit for dividing tiles. In other words, a patch is the signaling information about the construction of point cloud data.

[0171] A tile represents an independent, décodeable rectangular region of an Atlas frame.

[0172] The receiving device according to the embodiment can recover attribute video data, geometric video data, and occupancy video data based on Atlas (patch or patch), which are actual video data with the same presentation time.

[0173] Geometric Image: This is a depth map-like image that presents the positional information (geometric) of the individual points that make up the point cloud patch by patch. A geometric image can consist of pixel values ​​from a single channel. The geometric representation is a set of coordinates associated with a point cloud frame.

[0174] Texture image: This is an image that represents color information about the individual points that make up a point cloud, patch by patch. A texture image may consist of pixel values ​​from multiple channels (e.g., R, G, and B channels). Texture is included in attributes. Depending on the implementation, texture and / or attributes may be interpreted as the same object and / or have an inclusion relationship.

[0175] Auxiliary patch information: This indicates the metadata required to reconstruct the point cloud using the individual patches. Auxiliary patch information may include information about the location, size, etc. of the patches in 2D / 3D space.

[0176] Point cloud data, depending on the implementation method, such as V-PCC components, may include atlases, occupancy maps, geometry, and attributes.

[0177] An atlas represents a set of 2D bounding boxes. That is, an atlas can be a set of patches, for example, patches projected onto a rectangular frame corresponding to a 3D bounding box in 3D space, which can represent a subset of a point cloud. In this case, a patch can represent a rectangular region within an atlas corresponding to a rectangular region in a planar projection. Furthermore, patch data can represent data from 2D to 3D that requires patch transformation. Additionally, a set of patch data is also referred to as an atlas.

[0178] Attributes can represent scalars or vectors associated with each point in a point cloud. For example, attributes can include color, reflectivity, surface normal, timestamp, and material ID.

[0179] The point cloud data according to the implementation represents PCC data based on a video-based point cloud compression (V-PCC) scheme. The point cloud data may include multiple components. For example, it may include occupancy maps, patches, geometry, and / or textures.

[0180] Figure 4 An example of a point cloud video encoder according to an embodiment is shown.

[0181] Figure 4 The V-PCC encoding process used to generate and compress occupancy maps, geometric images, texture images, and auxiliary patch information is shown. Figure 4 V-PCC encoding processing can be performed by Figure 1 The point cloud video encoder 10002 is used for processing. Figure 4 Each component can be executed by software, hardware, processor, and / or a combination thereof.

[0182] Patch generator 14000 receives point cloud frames (which may be in the form of a bitstream containing point cloud data). Patch generator 14000 generates patches based on the point cloud data. In addition, patch information including information about patch generation is generated.

[0183] Patch packing or patch packer 14001 packs one or more patches. In addition, patch packer 14001 generates a occupancy map containing information about patch packing.

[0184] The geometric image generator 14002 generates a geometric image based on point cloud data, patch information (or auxiliary information), and / or occupancy map information. A geometric image refers to data containing geometry related to the point cloud data (i.e., the 3D coordinates of points) and refers to the geometric framework.

[0185] Texture image generation or texture image generator 14003 generates texture images based on point cloud data, patches, packed patches, patch information (or auxiliary information), and / or smoothed geometry. A texture image refers to an attribute frame. That is, texture images can also be generated based on smoothed geometry generated through smoothing processing based on patch information.

[0186] Smoothing or smoother 14004 can mitigate or eliminate errors contained in image data. For example, the reconstructed geometric image is smoothed based on patch information. That is, portions that may cause errors between data can be smoothly filtered out to generate smooth geometry.

[0187] The auxiliary patch information compressor 14005 can compress auxiliary patch information associated with the patch information generated during patch generation. Furthermore, the auxiliary patch information compressed in the auxiliary patch information compressor 14005 can be sent to the multiplexer 14013. The auxiliary patch information can be used in the geometry image generator 14002.

[0188] Image fillers 14006 and 14007 can fill geometric images and texture images respectively. Fill data can be applied to both geometric and texture images.

[0189] Group expansion or group expander 14008 can add data to a texture image in a manner similar to image padding. Auxiliary patch information can be inserted into the texture image.

[0190] Video compressors 14009, 14010, and 14011 can compress filled geometry images, filled texture images, and / or occupancy maps, respectively. In other words, video compressors 14009, 14010, and 14011 can compress input geometry frames, attribute frames, and / or occupancy map frames to output video bitstreams of the geometry image, texture image, and occupancy map, respectively. Video compression can encode geometric information, texture information, and occupancy information.

[0191] Entropy compression, or entropy compressor 14012, can compress occupancy graphs based on entropy schemes.

[0192] According to the implementation method, entropy compression and / or video compression can be performed on the occupied frames depending on whether the point cloud data is lossless and / or lossy.

[0193] Multiplexer 14013 multiplexes the compressed geometric video bitstream, the compressed texture image video bitstream, the compressed occupancy map video bitstream, and the compressed auxiliary patch information bitstream from each compressor into a single bitstream.

[0194] The aforementioned blocks can be omitted or replaced by blocks with similar or identical functions. Furthermore, Figure 4 Each block shown can be used as at least one of a processor, software, and hardware.

[0195] According to the implementation method Figure 4 The detailed operation descriptions for each process are as follows.

[0196] Patch generation (14000)

[0197] Patch generation refers to the process of dividing a point cloud into patches (mapping units) to map the point cloud onto a 2D image. Patch generation can be divided into three steps: normal value calculation, segmentation, and patch segmentation.

[0198] Reference Figure 5 Describe the normal value calculation process in detail.

[0199] Figure 5 An example of the tangent plane and normal vector of a surface according to an embodiment is shown.

[0200] exist Figure 4The patch generator 14000 for V-PCC encoding processing is used as follows: Figure 5 The surface.

[0201] Normal calculation related to patch generation

[0202] Each point in a point cloud has its own orientation, represented by a 3D vector called the normal vector. Using the neighbors of each point obtained through methods such as KD-trees, we can obtain... Figure 5 The diagram shows the tangent planes and normal vectors of each point on the surface that constitutes the point cloud. The search range applied to the neighbor search process can be defined by the user.

[0203] A tangent plane is a plane that passes through a point on a surface and completely includes the tangent to a curve on the surface.

[0204] Figure 6 An exemplary bounding box of a point cloud according to an implementation method is shown.

[0205] According to the implementation method, the bounding box refers to the box used to divide point cloud data into units based on hexahedrons in 3D space.

[0206] The method / apparatus according to the implementation (e.g., patch generator 14000) may use bounding boxes in the process of generating patches from point cloud data.

[0207] Bounding boxes can be used in processing where a target object of point cloud data is projected onto the planes of the flat faces of a hexahedron in 3D space. Bounding boxes can be defined by... Figure 1 The point cloud video acquisition unit 10001 and the point cloud video encoder 10002 generate and process the data. Furthermore, based on the bounding box, execution is possible. Figure 4 The V-PCC encoding process generates patches 14000, patches packing 14001, geometric images 14002, and texture images 14003.

[0208] Segments related to patch generation

[0209] The segmentation is divided into two processes: initial segmentation and refined segmentation.

[0210] The point cloud video encoder 10002, according to the embodiment, projects points onto a face of a bounding box. Specifically, as... Figure 6 As shown, each point that makes up the point cloud is projected onto one of the six faces of the bounding box surrounding the point cloud. The initial segmentation is the process of determining one of the flat faces of the bounding box to which each point will be projected.

[0211] It is the normal value corresponding to each of the six flat planes, as defined below:

[0212] (1.0,0.0,0.0), (0.0,1.0,0.0), (0.0,0.0,1.0), (-1.0,0.0,0.0), (0.0,-1.0,0.0), (0.0,0.0,-1.0).

[0213] As shown in the following formula, the normal vectors of each point are obtained in the normal value calculation process. and The plane with the largest dot product value is determined as the projection plane of the corresponding point. That is, the plane whose direction is most similar to the normal vector of the point is determined as the projection plane of the point.

[0214]

[0215] The determined plane can be identified by a cluster index (one of 0 to 5).

[0216] Refining the segmentation involves considering the enhancement of the projection planes of neighboring points, which are determined in the initial segmentation process for each point constituting the point cloud. In this process, scoring normals and scoring smoothing can be considered together. The scoring normal represents the similarity between the normal vectors of each point considered when determining the projection planes in the initial segmentation process and the normals of each flat surface of the bounding box. Scoring smoothing indicates the similarity between the projection plane of the current point and the projection planes of its neighboring points.

[0217] Score smoothing can be considered by assigning weights to the scoring method. In this case, the weight values ​​can be defined by the user. Refining the segments can be performed repeatedly, and the number of repetitions can also be defined by the user.

[0218] Patch segmentation related to patch generation

[0219] Patch segmentation is a process that divides the entire point cloud into patches (sets of neighboring points) based on the projection plane information of each point constituting the point cloud obtained in the initial / refining segmentation process. Patch segmentation may include the following steps:

[0220] ① Use KD-trees or similar methods to calculate the neighboring points of each point in the point cloud. The maximum number of neighbors can be defined by the user;

[0221] ② When neighboring points are projected onto the same plane as the current point (when they have the same cluster index),

[0222] Extract the current point and its neighboring points as a patch;

[0223] ③ Calculate the geometric values ​​of the extracted patch.

[0224] ④ Repeat steps ② to ③ until there are no more points left to be extracted.

[0225] The occupancy map, geometric image, and texture image of each patch, as well as the size of each patch, are determined through patch segmentation processing.

[0226] Figure 7 An example is shown of determining the positions of each patch on the occupancy map according to an implementation method.

[0227] The point cloud video encoder 10002 according to the implementation method can perform patch packaging and generate occupancy map.

[0228] Patch packing and occupancy map generation (14001)

[0229] This is the process of determining the positions of individual patches in a 2D image to map segmented patches onto the 2D image. As a 2D image, an occupancy map is a binary image that uses values ​​of 0 or 1 to indicate the presence or absence of data at corresponding locations. An occupancy map consists of blocks, and its resolution is determined by the block size. For example, when the block is 1x1, pixel-level resolution is obtained. The size of the occupancy block can be determined by the user.

[0230] The process for determining the position of each patch on the occupancy map can be configured as follows:

[0231] ① Set all positions on the occupied map to 0;

[0232] ② Place the patch at point (u,v) in the occupied plane, where the horizontal coordinate is within the range of (0,occupancySizeU-patch.sizeU0) and the vertical coordinate is within the range of (0,occupancySizeV-patch.sizeV0).

[0233] ③ Set the point (x, y) in the patch plane whose horizontal coordinate is in the range of (0, patch.sizeU0) and whose vertical coordinate is in the range of (0, patch.sizeV0) as the current point;

[0234] ④ Change the position of point (x, y) in raster order. If the value of coordinate (x, y) on the patch occupancy map is 1 (data exists at the point in the patch) and the value of coordinate (u+x, v+y) on the global occupancy map is 1 (the occupancy map is filled with the previous patch), then repeat operations ③ and ④. Otherwise, proceed to operation ⑥.

[0235] ⑤ Change the position of (u,v) according to the grating sequence and repeat operations ③ to ⑤;

[0236] ⑥ Determine (u,v) as the location of the patch and copy the occupancy map data of the patch to the corresponding part of the global occupancy map; and

[0237] ⑦ Repeat steps ② to ⑥ for the next patch.

[0238] occupancySizeU: Indicates the width of the occupancy map. Its unit is the size of the occupancy block.

[0239] occupancySizeV: Indicates the height of the occupancy map. Its unit is the size of the occupied block.

[0240] patch.sizeU0: Indicates the width of the patch. Its unit is the size of the pack block it occupies.

[0241] patch.sizeV0: Indicates the height of the patch. Its unit is the size of the pack block it occupies.

[0242] For example, such as Figure 7 As shown, in the box corresponding to the block occupying the package size, there exists a box corresponding to the patch with the patch size, and the point (x,y) can be located in that box.

[0243] Figure 8 An exemplary relationship between the normal axis, tangential axis, and double tangential axis is shown according to an embodiment.

[0244] The point cloud video encoder 10002 according to the embodiment can generate a geometric image. A geometric image refers to image data that includes geometric information about the point cloud. The geometric image generation process can employ the three axes (normal, tangential, and bitangential) of the patch in path 8.

[0245] Geometric Image Generation (14002)

[0246] In this process, the depth values ​​of the geometric images constituting each patch are determined, and the entire geometric image is generated based on the patch positions determined in the patch packing process described above. The process for determining the depth values ​​of the geometric images constituting each patch can be configured as follows.

[0247] ① Calculate parameters related to the position and size of each patch. These parameters may include the following information. According to the implementation, the position of the patch is included in the patch information.

[0248] The normal index of the normal axis is obtained in the previous patch generation process. The tangential axis is the axis perpendicular to the normal axis that coincides with the horizontal axis u of the patch image, and the bitangential axis is the axis perpendicular to the normal axis that coincides with the vertical axis v of the patch image. The three axes can be... Figure 8 As shown.

[0249] Figure 9 Exemplary configurations of the minimum and maximum modes of the projection mode according to the implementation are shown.

[0250] The point cloud video encoder 10002 according to the embodiment can perform patch-based projection to generate a geometric image, and the projection modes according to the embodiment include a minimum mode and a maximum mode.

[0251] The 3D spatial coordinates of the patch can be calculated based on the bounding box of the minimum size surrounding the patch. For example, the 3D spatial coordinates can include the minimum tangential value of the patch (on the 3D displacement tangential axis of the patch), the minimum bitangential value of the patch (on the 3D displacement bitangential axis of the patch), and the minimum normal value of the patch (on the 3D displacement normal axis of the patch).

[0252] The 2D size of the patch indicates the horizontal and vertical sizes of the patch when it is packed into a 2D image. The horizontal size (pattern 2D size u) can be obtained as the difference between the maximum and minimum tangent values ​​of the bounding box, and the vertical size (pattern 2D size v) can be obtained as the difference between the maximum and minimum bitangent values ​​of the bounding box.

[0253] ② Determine the projection mode of the patch. The projection mode can be either the minimum mode or the maximum mode. Geometric information about the patch is represented using depth values. When the points constituting the patch are projected onto the normal of the patch, two layers of images can be generated: an image constructed using the maximum depth value and an image constructed using the minimum depth value.

[0254] In minimum mode, when generating two layers of images d0 and d1, a minimum depth can be configured for d0, and a maximum depth within the surface thickness of the minimum depth can be configured for d1, such as... Figure 9 As shown.

[0255] For example, when the point cloud is located in 2D (such as...) Figure 9 As shown, there are multiple patches comprising multiple points. The figure indicates that points marked with the same style of shading can belong to the same patch. The figure illustrates the processing of patches with projected blank points.

[0256] When projecting blank points to the left or right, the depth can be incremented by 1 relative to the left, for 0, 1, 2, ..., 6, 7, 8, 9, and the number used to calculate the depth of the point can be marked on the right.

[0257] The same projection mode can be applied to all point clouds, or different projection modes can be applied to individual frames or patches based on user definitions. When applying different projection modes to individual frames or patches, the projection mode that enhances compression efficiency or minimizes missing points can be adaptively selected.

[0258] ③ Calculate the depth value of each point.

[0259] In minimum mode, image d0 is constructed using depth0, which is obtained by subtracting the minimum normal value of the patch (on the patch's 3D shift normal axis) calculated in operation 1) from the minimum normal value of the patch at each point (on the patch's 3D shift normal axis). If another depth value exists at the same location within the range between depth0 and the surface thickness, that value is set to depth1. Otherwise, the value of depth0 is assigned to depth1. Image d1 is constructed using the value of depth1.

[0260] For example, when calculating the depth of points in image d0, the minimum value can be calculated (4 2 4 4 0 6 0 0 9 9 0 80). When calculating the depth of points in image d1, the larger value among two or more points can be calculated. When only one point exists, its value can be calculated (4 4 4 4 6 6 6 8 9 9 8 8 9). In the processing of points in the encoding and reconstruction patch, some points may be missing (e.g., eight points are missing in the image).

[0261] In maximum mode, image d0 is constructed using depth0, which is obtained by subtracting the minimum normal value of the patch (on the 3D shifted normal axis of the patch) calculated in operation 1) from the minimum normal value of the patch (on the 3D shifted normal axis of the patch) for each point using the maximum normal value. If another depth value exists at the same location within the range between depth0 and the surface thickness, that value is set to depth1. Otherwise, the value of depth0 is assigned to depth1. Image d1 is constructed using the value of depth1.

[0262] For example, when calculating the depth of points in image d0, a minimum value can be calculated (4 4 4 4 6 6 6 8 9 9 8 89). When calculating the depth of points in image d1, the smaller value among two or more points can be calculated. When only one point exists, its value can be calculated (4 2 4 4 5 6 0 6 9 9 0 8 0). In the processing of points in the encoding and reconstruction patch, some points may be missing (e.g., six points are missing in the image).

[0263] The entire geometric image can be generated by placing the geometric images of each patch generated through the above processing onto the entire geometric image based on the patch position information determined in the patch packing process.

[0264] The generated layer d1 of the entire geometric image can be encoded using various methods. The first method (absolute d1 encoding) encodes the depth values ​​of the previously generated image d1. The second method (differential encoding) encodes the difference between the depth values ​​of the previously generated image d1 and the depth values ​​of image d0.

[0265] In the encoding method described above that uses depth values ​​of two layers d0 and d1, if there is another point between the two depths, the geometric information about that point is lost during the encoding process. Therefore, Enhanced Incremental Depth (EDD) codes can be used for lossless encoding.

[0266] The following will refer to Figure 10 Describe the EDD code in detail.

[0267] Figure 10 An exemplary EDD code according to an implementation method is shown.

[0268] In some / all of the processing of point cloud video encoder 10002 and / or V-PCC encoding (e.g., video compression 14009), geometric information about points can be encoded based on EOD codes.

[0269] like Figure 10 As shown, EDD codes are used for binary encoding of the positions of all points within a surface thickness range including d1. For example, in Figure 10 In the diagram, since points exist at the first and fourth positions on D0 and the second and third positions are empty, the points included in the second column on the left can be represented by the EDD code 0b1001 (=9). When the EDD code is encoded and transmitted together with D0, the receiving terminal can recover the geometric information about all points without loss.

[0270] For example, the value is 1 when there is a point above the reference point, and 0 when there is no point. Therefore, the code can be represented using 4 bits.

[0271] Smoothing (14004)

[0272] Smoothing is an operation used to eliminate discontinuities that may appear at patch boundaries due to image quality degradation that occurs during compression processing. Smoothing can be performed by the point cloud video encoder 10002 or the smoother 14004.

[0273] ① Reconstructing a point cloud from a geometric image. This operation can be the reverse of the geometric image generation described above. For example, it can be the reverse processing of the encoded data;

[0274] ② Use KD trees and other methods to calculate the neighboring points of each point in the reconstructed point cloud;

[0275] ③ Determine whether each point is located on the patch boundary. For example, when there are neighboring points with a different projection plane (cluster index) than the current point, it can be determined that the point is located on the patch boundary;

[0276] ④ If a point exists on the patch boundary, move that point to the centroid of a neighboring point (located at the average x, y, z coordinates of the neighboring point). That is, change the geometry. Otherwise, maintain the previous geometry.

[0277] Figure 11 An example of recoloring based on the color values ​​of neighboring points according to an implementation method is shown.

[0278] According to the embodiments, the point cloud video encoder 10002 or texture image generator 14003 can generate texture images based on recoloring.

[0279] Texture image generation (14003)

[0280] Similar to the geometric image generation process described above, the texture image generation process involves generating texture images for each patch and generating the entire texture image by arranging the texture images in defined positions. However, in the operation of generating texture images for each patch, instead of using depth values ​​for geometry generation, an image with color values ​​(e.g., R, G, and B values) of the points that constitute the point cloud corresponding to the position is generated.

[0281] When estimating the color values ​​of the individual points that make up the point cloud, the geometry previously obtained through smoothing can be used. In a smoothed point cloud, the positions of some points may have shifted relative to the original point cloud, so a recoloring process may be needed to find colors suitable for the changed positions. Recoloring can be performed using the color values ​​of neighboring points. For example, as... Figure 11 As shown, the color values ​​of the nearest neighbor and neighboring points can be considered to calculate the new color value.

[0282] For example, refer to Figure 11 In recoloring, the appropriate color value for the changed position can be calculated based on the average of the attribute information about the point's nearest original point and / or the average of the attribute information about the point's nearest original point.

[0283] Similar to a geometric image generated from two layers d0 and d1, a texture image can also be generated from two layers t0 and t1.

[0284] Auxiliary patch information compression (14005)

[0285] The point cloud video encoder 10002 or auxiliary patch information compressor 14005 according to the embodiment can compress auxiliary patch information (auxiliary information about the point cloud).

[0286] The auxiliary patch information compressor 14005 compresses the auxiliary patch information generated during the patch generation, patch packaging, and geometry generation processes described above. The auxiliary patch information may include the following parameters:

[0287] An index (cluster index) used to identify the projection plane (normal plane);

[0288] The 3D spatial position of the patch, namely, the minimum tangential value of the patch (on the 3D displacement tangential axis of the patch), the minimum double tangential value of the patch (on the 3D displacement double tangential axis of the patch), and the minimum normal value of the patch (on the 3D displacement normal axis of the patch).

[0289] The 2D spatial position and size of the patch, namely, horizontal size (pattern 2D size u), vertical size (pattern 2D size v), minimum horizontal value (pattern 2D displacement u), and minimum vertical value (pattern 2D displacement u); and

[0290] Regarding the mapping information for each block and patch, there are candidate indices (when patches are set sequentially based on information about their 2D spatial location and size, multiple patches can be mapped to a block in an overlapping manner. In this case, the mapped patches constitute a candidate list, and the candidate index indicates the sequential position of the patch whose data exists within the block) and local patch indices (indicating the index of a patch existing in a frame). Table 1 shows the pseudocode representing the process of matching between blocks and patches based on the candidate list and local patch indices.

[0291] The maximum number of candidates can be defined by the user.

[0292] [Table 1]

[0293]

[0294] Figure 12 This illustrates a push-pull background fill according to an embodiment.

[0295] Image padding and group expansion (14006, 14007, 14008)

[0296] The image filler according to the implementation method can fill the space outside the patch area with meaningless supplementary data based on the push-pull background fill technology.

[0297] Image padding 14006 and 14007 uses meaningless data to fill the space outside the patch area to improve compression efficiency. For image padding, pixel values ​​from columns or rows near the boundaries of the patch can be copied to fill the blank space. Alternatively, such as Figure 12 As shown, a push-pull background filling method can be used. According to this method, the blank space is filled using pixel values ​​from the low-resolution image during the process of gradually reducing the resolution of the unfilled image and then increasing the resolution again.

[0298] Group expansion 14008 is a process that fills the blank spaces of a geometric image and a texture image configured with two layers, d0 / d1 and t0 / t1, respectively. In this process, the blank spaces of the two layers are calculated by averaging the values ​​at the same location using image filling.

[0299] Figure 13 An exemplary possible traversal order of a 4x4 block according to an implementation is shown.

[0300] Occupancy graph compression (14012, 14011)

[0301] The occupancy map compressor according to the implementation can compress previously generated occupancy maps. Specifically, two methods can be used: video compression for lossy compression and entropy compression for lossless compression. Video compression is described below.

[0302] Entropy compression can be performed using the following methods.

[0303] ① If the block constituting the occupancy map is fully occupied, then encode 1 and repeat the same operation for the next block of the occupancy map. Otherwise, encode 0 and perform operations 2) through 5).

[0304] ② Determine the optimal traversal order for performing run-length encoding on the occupied pixels of the block. Figure 13 This shows four possible traversal orders for a 4x4 block.

[0305] Figure 14 An exemplary optimal traversal order according to the implementation method is shown.

[0306] As described above, the entropy compressor according to the implementation can be based on Figure 14 The traversal order scheme shown encodes the blocks.

[0307] For example, the optimal traversal order with the minimum number of runs is selected from the possible traversal orders, and its index is encoded. The diagram illustrates the selection process. Figure 13 The case of the third traversal order. In the case shown, the number of runs can be minimized to 2, therefore the third traversal order can be chosen as the optimal traversal order.

[0308] ③ Encode the number of runs. Figure 14 In the example, there are two runs, so 2 is encoded.

[0309] ④ Encode the occupancy of the first run. Figure 14 In the example, 0 is encoded because the first run corresponds to an unoccupied pixel.

[0310] ⑤ Encode the length of each run (the total number of runs). Figure 14In the example, the lengths of the first run and the second run, 6 and 10, are encoded in sequence.

[0311] Video compression (14009, 14010, 14011)

[0312] According to the embodiments, the video compressors 14009, 14010, and 14011 use 2D video codecs such as HEVC or VVC to encode sequences of geometric images, texture images, occupancy map images, etc., generated in the above operations.

[0313] Figure 15 An exemplary 2D video / image encoder according to an embodiment is shown. According to the embodiment, the 2D video / image encoder may be referred to as an encoding device.

[0314] This indicates the application of embodiments of the video compressors 14009, 14010, and 14011 described above. Figure 15 This is a schematic block diagram of a 2D video / image encoder 15000 configured to encode video / image signals. The 2D video / image encoder 15000 may be included in the point cloud video encoder 10002 described above, or may be configured as an internal / external component. Figure 15 Each component can correspond to software, hardware, processor, and / or a combination thereof.

[0315] Here, the input image can include one of the aforementioned geometric images, texture images (attribute images), and occupancy map images. When Figure 15 When the 2D video / image encoder is applied to the video compressor 14009, the image input to the 2D video / image encoder 15000 is a filled geometric image, and the bitstream output from the 2D video / image encoder 15000 is a compressed geometric image bitstream. Figure 15 When the 2D video / image encoder is applied to the video compressor 14010, the image input to the 2D video / image encoder 15000 is a filled texture image, and the bitstream output from the 2D video / image encoder 15000 is a compressed texture image bitstream. Figure 15 When the 2D video / image encoder is applied to the video compressor 14011, the image input to the 2D video / image encoder 15000 is a occupancy map image, and the bitstream output from the 2D video / image encoder 15000 is a bitstream of the compressed occupancy map image.

[0316] Inter-frame predictor 15090 and intra-frame predictor 15100 can be collectively referred to as predictors. That is, the predictor may include inter-frame predictor 15090 and intra-frame predictor 15100. Transformer 15030, quantizer 15040, inverse quantizer 15050 and inverse transformer 15060 can be collectively referred to as residual processors. The residual processor may also include subtractor 15020. According to the embodiment, Figure 15 The image segmenter 15010, subtractor 15020, transformer 15030, quantizer 15040, inverse quantizer 15050, inverse transformer 15060, adder 15200, filter 15070, inter-frame predictor 15090, intra-frame predictor 15100, and entropy encoder 15110 can be configured by a single hardware component (e.g., an encoder or processor). Additionally, memory 15080 may include a decoded picture buffer (DPB) and can be configured by a digital storage medium.

[0317] Image segmenter 15010 can segment an image (or picture or frame) input to encoder 15000 into one or more processing units. For example, a processing unit may be referred to as a coding unit (CU). In this case, the CU may be recursively segmented from coding tree units (CTUs) or maximum coding units (LCUs) according to a quadtree-binary tree (QTBT) structure. For example, a CU may be segmented into multiple CUs of lower depth based on a quadtree structure and / or a binary tree structure. In this case, for example, a quadtree structure may be applied first, followed by a binary tree structure. Alternatively, a binary tree structure may be applied first. The coding process according to this disclosure can be performed based on a final CU that is no longer segmented. In this case, an LCU may be used as the final CU based on coding efficiency according to the characteristics of the image. If necessary, the CU may be recursively segmented into lower-depth CUs, and the CU of optimal size may be used as the final CU. Here, the coding process may include prediction, transformation, and reconstruction (described later). As another example, the processing unit may also include a prediction unit (PU) or a transform unit (TU). In this case, PU and TU can be separated or divided from the final CU mentioned above. PU can be a unit for sample prediction, and TU can be a unit for deriving transform coefficients and / or deriving residual signals from transform coefficients.

[0318] The term "unit" is used interchangeably with terms such as block, region, or module. In general, an M×N block can represent a set of samples or transform coefficients arranged in M ​​columns and N rows. Samples typically represent pixels or pixel values ​​and can indicate only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component. "Sample" can be used as a term corresponding to a pixel or image within a frame (or image).

[0319] The subtractor 15020 of encoder 15000 generates a residual signal (residual block or residual sample array) by subtracting the prediction signal (prediction block or prediction sample array) output from inter-frame predictor 15090 or intra-frame predictor 15100 from the input image signal (original block or original sample array), and the generated residual signal is sent to converter 15030. In this case, as shown, the unit in encoder 15000 that subtracts the prediction signal (prediction block or prediction sample array) from the input image signal (original block or original sample array) can be referred to as subtractor 15020. The predictor can perform prediction on the processing target block (hereinafter referred to as the current block) and generate a prediction block including the prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. As described later in the description of the various prediction modes, the predictor can generate various types of information about the prediction (e.g., prediction mode information) and transmit the generated information to entropy encoder 15110. Information about the prediction can be encoded by the entropy encoder 15110 and output as a bit stream.

[0320] The intra-frame predictor 15100 of the predictor can refer to samples in the current frame to predict the current block. Depending on the prediction mode, the samples can be near or far from the current block. Under intra-frame prediction, the prediction modes can include multiple non-directional modes and multiple directional modes. Non-directional modes can include, for example, DC mode and planar mode. Depending on the granularity of the prediction direction, the directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is only an example, and more or fewer directional prediction modes can be used depending on the settings. The intra-frame predictor 15100 can determine the prediction mode to be applied to the current block based on the prediction modes applied to neighboring blocks.

[0321] The inter-frame predictor 15090 of the predictor can deduce the prediction block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference frame. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted based on each block, sub-block, or sample based on the correlation between motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference frame indices. Motion information may also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current frame and temporally neighboring blocks existing in the reference frame. The reference frame including the reference block may be the same as or different from the reference frame including the temporally neighboring block. The temporally neighboring block may be referred to as a juxtaposed reference block or juxtaposed CU (colCU), and the reference frame including the temporally neighboring block may be referred to as a juxtaposed frame (colPic). For example, the inter-frame predictor 15090 can configure a motion information candidate list based on neighboring blocks and generate information indicating candidates to be used to deduce the motion vector and / or reference frame index of the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip and merge modes, the inter-frame predictor 15090 can use motion information about neighboring blocks as motion information about the current block. In skip mode, unlike merge mode, residual signals may not be sent. In motion vector prediction (MVP) mode, motion vectors from neighboring blocks can be used as motion vector predictors, and motion vector differences can be signaled to indicate the motion vector of the current block.

[0322] The prediction signal generated by the inter-frame predictor 15090 or the intra-frame predictor 15100 can be used to generate the reconstructed signal or the residual signal.

[0323] Transformer 15030 can generate transform coefficients by applying transform techniques to the residual signal. For example, the transform technique may include at least one of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen–Loève Transform (KLT), Graphical Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, GBT refers to a transform obtained from a graph depicting the relationships between pixels. CNT refers to a transform obtained based on a prediction signal generated from all previously reconstructed pixels. Furthermore, the transform operation can be applied to square pixel blocks of the same size, or to blocks of variable size other than square.

[0324] The quantizer 15040 quantizes the transform coefficients and sends them to the entropy encoder 15110. The entropy encoder 15110 encodes the quantized signal (information about the quantized transform coefficients) and outputs a bitstream of the encoded signal. The information about the quantized transform coefficients can be referred to as residual information. The quantizer 15040 rearranges the block-form quantized transform coefficients in the form of a one-dimensional vector based on the coefficient scan order, and generates information about the quantized transform coefficients based on the one-dimensional vector form of the quantized transform coefficients.

[0325] The entropy encoder 15110 can employ various coding techniques such as, for example, exponential Golomb, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoder 15110 can encode information required for video / image reconstruction (e.g., values ​​of syntax elements) together with or separately from quantized transform coefficients. The encoded information (e.g., encoded video / image information) can be transmitted or stored as a bitstream based on Network Abstraction Layer (NAL) units.

[0326] The bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcast network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. A transmitter (not shown) that transmits the signal output from the entropy encoder 15110 and / or a storage unit (not shown) that stores the signal may be configured as internal / external components of the encoder 15000. Alternatively, the transmitter may be included within the entropy encoder 15110.

[0327] The quantized transform coefficients output from quantizer 15040 can be used to generate a prediction signal. For example, inverse quantization and inverse transform can be applied to the quantized transform coefficients via inverse quantizer 15050 and inverse transformer 15060 to reconstruct the residual signal (residual block or residual sample). Adder 15200 adds the reconstructed residual signal to the prediction signal output from inter-frame predictor 15090 or intra-frame predictor 15100. This generates a reconstructed signal (reconstructed frame, reconstructed block, reconstructed sample array). When there is no residual signal for the processing target block, as in the case of applying skip mode, the prediction block can be used as a reconstructed block. Adder 15200 can be referred to as a reconstructor or reconstructed block generator. As described below, the generated reconstructed signal can be used for intra-frame prediction of the next processing target block in the current frame, or it can be filtered for inter-frame prediction of the next frame.

[0328] Filter 15070 can improve subjective / objective image quality by applying filtering to the reconstructed signal output from adder 15200. For example, filter 15070 can generate a modified reconstructed image by applying various filtering techniques to the reconstructed image, and the modified reconstructed image can be stored in memory 15080 (specifically, the DPB of memory 15080). Various filtering techniques may include, for example, deblocking filtering, sample adaptive offsetting, adaptive loop filtering, and bilateral filtering. As described below in the description of filtering techniques, filter 15070 can generate various types of information about the filtering and transmit the generated information to entropy encoder 15110. The information about the filtering can be encoded by entropy encoder 15110 and output as a bitstream.

[0329] The modified reconstructed frame stored in memory 15080 can be used as a reference frame by inter-frame predictor 15090. Therefore, when applying inter-frame prediction, the encoder can avoid prediction mismatch between encoder 15000 and decoder and improve coding efficiency.

[0330] The DPB of memory 15080 can store modified reconstructed frames for use as reference frames by inter-frame predictor 15090. Memory 15080 can store motion information about blocks in the current frame that have been deduced (or encoded) and / or motion information about already reconstructed blocks in the frame. The stored motion information can be transmitted to inter-frame predictor 15090 for use as motion information about spatially adjacent blocks or temporally adjacent blocks. Memory 15080 can store reconstructed samples of reconstructed blocks in the current frame and transmit the reconstructed samples to intra-frame predictor 15100.

[0331] At least one of the above prediction, transformation, and quantization processes can be skipped. For example, for a block that has applied Pulse Code Modulation (PCM), the prediction, transformation, and quantization processes can be skipped, and the values ​​of the original samples can be encoded and output as a bitstream.

[0332] Figure 16 An exemplary V-PCC decoding process according to an implementation is shown.

[0333] V-PCC decoding processing or V-PCC decoder can follow Figure 4 The reverse processing of V-PCC encoding (or encoder). Figure 16 Each component can correspond to software, hardware, processor, and / or a combination thereof.

[0334] The demultiplexer 16000 demultiplexes the compressed bitstream to output compressed texture images, compressed geometric images, compressed occupancy maps, and compressed auxiliary patch information, respectively.

[0335] Video decompression or video decompressor 16001, 16002 decompresses each of compressed texture images and compressed geometric images.

[0336] Occupancy map decompression or occupancy map decompression compressor 16003 decompresses compressed occupancy map images.

[0337] The auxiliary patch information decompression or auxiliary patch information decompression device 16004 decompresses the compressed auxiliary patch information.

[0338] The geometric reconstruction, or geometric reconstructor 16005, recovers (reconstructs) geometric information based on the decompressed geometric image, the decompressed occupancy map, and / or the decompressed auxiliary patch information. For example, geometry altered during encoding processing can be reconstructed.

[0339] The smoother or smoother16006 can smooth the reconstructed geometry. For example, a smoothing filter can be applied.

[0340] Texture reconstruction or texture reconstructor 16007 reconstructs textures from decompressed texture images and / or smoothed geometry.

[0341] Color smoothing, or the color smoother 16008, smooths color values ​​from the reconstructed texture. For example, a smoothing filter can be applied.

[0342] As a result, reconstructed point cloud data can be generated.

[0343] Figure 16 This illustrates the V-PCC decoding process for reconstructing a point cloud by decompressing (decoding) the compressed occupancy map, geometric image, texture image, and auxiliary patch information.

[0344] Figure 16 Each unit shown can be used as at least one of a processor, software, and hardware. According to the implementation method... Figure 16 The detailed operation descriptions for each unit are as follows.

[0345] Video decompression (16001, 16002)

[0346] Video decompression is the inverse process of video compression described above. It involves using a 2D video codec such as HEVC or VVC to decode the bitstream of the geometric image, the bitstream of the compressed texture image, and / or the bitstream of the compressed occupancy map image generated in the above process.

[0347] Figure 17 An exemplary 2D video / image decoder, also referred to as a decoding device, is shown according to an embodiment.

[0348] 2D video / image decoders can follow Figure 15 The reverse processing of the operation of the 2D video / image encoder.

[0349] Figure 17 The 2D video / image decoder is Figure 16 Implementation methods of video decompressors 16001 and 16002. Figure 17 This is a schematic block diagram of a 2D video / image decoder 17000 that decodes video / image signals. The 2D video / image decoder 17000 may be included in the point cloud video decoder 10008 described above, or may be configured as an internal / external component. Figure 17 Each component can correspond to software, hardware, processor, and / or a combination thereof.

[0350] Here, the input bitstream can be one of the following: a bitstream of a geometric image, a bitstream of a texture image (attribute image), or a bitstream of a occupancy map image. When Figure 17 When the 2D video / image decoder is applied to the video decompressor 16001, the bitstream input to the 2D video / image decoder is a bitstream of compressed texture images, and the reconstructed image output from the 2D video / image decoder is a decompressed texture image. Figure 17 When the 2D video / image decoder is applied to the video decompressor 16002, the bitstream input to the 2D video / image decoder is the bitstream of the compressed geometric image, and the reconstructed image output from the 2D video / image decoder is the decompressed geometric image. Figure 17 The 2D video / image decoder can receive a bitstream of compressed occupancy map images and decompress it in alignment. The reconstructed image (or output image or decoded image) can represent the reconstructed image of the aforementioned geometric image, texture image (attribute image), and occupancy map image.

[0351] Reference Figure 17 The inter-frame predictor 17070 and the intra-frame predictor 17080 can be collectively referred to as predictors. That is, the predictor may include the inter-frame predictor 17070 and the intra-frame predictor 17080. The inverse quantizer 17020 and the inverse transformer 17030 can be collectively referred to as residual processors. That is, according to the implementation, the residual processor may include the inverse quantizer 17020 and the inverse transformer 17030. Figure 17 The entropy decoder 17010, inverse quantizer 17020, inverse transformer 17030, adder 17040, filter 17050, inter-frame predictor 17070, and intra-frame predictor 17080 can be configured by a single hardware component (e.g., a decoder or processor). Additionally, the memory 17060 can include a decoded picture buffer (DPB) or can be configured by a digital storage medium.

[0352] When the input contains a bitstream containing video / image information, the decoder 17000 can... Figure 15The encoder processes video / image information, corresponding to the image reconstruction process. For example, the decoder 17000 can use the processing unit applied in the encoder to perform decoding. Therefore, the decoding processing unit can be, for example, a CU. The CU can be segmented from a CTU or LCU along a quadtree structure and / or a binary tree structure. The reconstructed video signal decoded and output by the decoder 17000 can then be played by a player.

[0353] Decoder 17000 can receive signals output from encoder in the form of a bitstream, and the received signals can be decoded by entropy decoder 17010. For example, entropy decoder 17010 can parse the bitstream to derive information (e.g., video / image information) required for image reconstruction (or picture reconstruction). For example, entropy decoder 17010 can decode the information in the bitstream based on coding techniques such as exponential Golomb coding, CAVLC, or CABAC, outputting quantized values ​​of the transform coefficients of the syntax elements and residuals required for image reconstruction. More specifically, in CABAC entropy decoding, bins corresponding to individual syntax elements in the bitstream can be received, and a context model can be determined based on information about the target syntax elements and decoding information about neighboring target blocks or information about symbols / bins decoded in previous steps. The probability of bin occurrence can then be predicted based on the determined context model, and arithmetic decoding of the bins can be performed to generate symbols corresponding to the values ​​of the individual syntax elements. According to CABAC entropy decoding, after determining the context model, the context model can be updated based on information about the symbols / cells decoded for the next symbol / cell. Information about prediction from the information decoded by the entropy decoder 17010 can be provided to the predictors (inter-frame predictor 17070 and intra-frame predictor 17080), and the residual values ​​(i.e., quantized transform coefficients and related parameter information) from the entropy decoding performed by the entropy decoder 17010 can be input to the inverse quantizer 17020. Additionally, information about filtering from the information decoded by the entropy decoder 17010 can be provided to the filter 17050. A receiver (not shown) configured to receive the signal output from the encoder can also be configured as an internal / external element of the decoder 17000. Alternatively, the receiver can be a component of the entropy decoder 17010.

[0354] The inverse quantizer 17020 outputs transform coefficients by inverse quantizing the quantized transform coefficients. The inverse quantizer 17020 can rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement can be performed based on the coefficient scan order implemented by the encoder. The inverse quantizer 17020 can perform inverse quantization on the quantized transform coefficients and obtain the transform coefficients using quantization parameters (e.g., quantization step size information).

[0355] The inverse converter 17030 obtains the residual signal (residual block and residual sample array) by transforming the transform coefficients.

[0356] The predictor can perform prediction on the current block and generate a prediction block that includes prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on information about the prediction output from the entropy decoder 17010, and can determine a specific intra-frame / inter-frame prediction mode.

[0357] The intra-predictor 17080 of the predictor can refer to samples in the current frame to predict the current block. Depending on the prediction mode, the samples can be near or far from the current block. In intra-prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The intra-predictor 17080 can use the prediction modes applied to neighboring blocks to determine the prediction mode applied to the current block.

[0358] The inter-frame predictor 17070 of the predictor can deduce the prediction block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference frame. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted based on each block, sub-block, or sample based on the correlation between motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference frame indices. Motion information may also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current frame and temporally neighboring blocks existing in the reference frame. For example, the inter-frame predictor 17070 can configure a motion information candidate list based on neighboring blocks and deduce the motion vector and / or reference frame index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes. Information about the prediction may include information indicating the inter-frame prediction mode of the current block.

[0359] Adder 17040 adds the residual signal obtained from inverse transformer 17030 to the prediction signal (prediction block or prediction sample array) output from inter-frame predictor 17070 or intra-frame predictor 17080 to generate a reconstruction signal (reconstructed frame, reconstruction block, or reconstruction sample array). When there is no residual signal for processing the target block, as in the case of applying skip mode, the prediction block can be used as a reconstruction block.

[0360] The adder 17040 can be referred to as a reconstructor or reconstruction block generator. The generated reconstructed signal can be used for intra-frame prediction of the next processing target block in the current frame, or it can be used for inter-frame prediction of the next frame by filtering, as described below.

[0361] Filter 17050 can improve subjective / objective image quality by applying filtering to the reconstructed signal output from adder 17040. For example, filter 17050 can generate a modified reconstructed image by applying various filtering techniques to the reconstructed image, and the modified reconstructed image can be sent to memory 17060 (specifically, the DPB of memory 17060). For example, various filtering methods may include deblocking filtering, sample adaptive shifting, adaptive loop filtering, and bilateral filtering.

[0362] The reconstructed frame stored in the DPB of memory 17060 can be used as a reference frame in the inter-frame predictor 17070. Memory 17060 can store motion information about blocks in the current frame whose motion information has been derived (or decoded) and / or about blocks in already reconstructed frames. The stored motion information can be transmitted to the inter-frame predictor 17070 as motion information about spatially adjacent blocks or about temporally adjacent blocks. Memory 17060 can store reconstructed samples of reconstructed blocks in the current frame and transmit the reconstructed samples to the intra-frame predictor 17080.

[0363] In this disclosure, regarding Figure 15 The implementations described for the encoder 15000's filter 15070, inter-frame predictor 15090, and intra-frame predictor 15100 can be applied to the decoder 17000's filter 17050, inter-frame predictor 17070, and intra-frame predictor 17080 in the same or corresponding manner.

[0364] At least one of the above prediction, inverse transform, and inverse quantization processes can be skipped. For example, for a block that has applied pulse code modulation (PCM), the prediction, inverse transform, and inverse quantization processes can be skipped, and the values ​​of the decoded samples can be used as samples for reconstructing the image.

[0365] Occupancy diagram compression (16003)

[0366] This is the reverse process of the occupancy graph compression described above. Occupancy graph decompression is the process of reconstructing the occupancy graph by decompressing the occupancy graph bitstream.

[0367] Decompress auxiliary patch information (16004)

[0368] The auxiliary patch information can be reconstructed by performing the inverse process of the above auxiliary patch information compression and decoding the compressed auxiliary patch information bitstream.

[0369] Geometric Reconstruction (16005)

[0370] This is the inverse process of the aforementioned geometric image generation. Initially, patches are extracted from the geometric image using the reconstructed occupancy map, 2D position / size information about the patches included in the auxiliary patch information, and information about the mapping between blocks and patches. Then, a point cloud is reconstructed in 3D space based on the extracted patch's geometric image and the 3D position information about the patches included in the auxiliary patch information. When the geometric value corresponding to a point (u,v) within the patch is g(u,v), and the patch's position coordinates on the normal, tangential, and bitangential axes in 3D space are (δ0,s0,r0), the normal, tangential, and bitangential coordinates δ(u,v), s(u,v), and r(u,v) mapped to the position of point (u,v) in 3D space can be expressed as follows.

[0371] δ(u,v)=δ0+g(u,v)

[0372] s(u,v)=s0+u

[0373] r(u,v)=r0+v

[0374] Smoothing (16006)

[0375] Similar to smoothing in the encoding process described above, smoothing is a process used to eliminate discontinuities that may appear at patch boundaries due to image quality degradation that occurs during compression.

[0376] Texture reconstruction (16007)

[0377] Texture reconstruction is a process of reconstructing a color point cloud by assigning color values ​​to each point that makes up the smooth point cloud. This can be performed by assigning the color values ​​corresponding to the texture image pixels at the same locations in the 2D geometric image to the points in the point cloud that correspond to the same locations in the 3D space, based on the mapping information between the reconstructed geometric image and the point cloud from the geometric reconstruction process described above.

[0378] Color smoothing (16008)

[0379] Color smoothing is similar to the geometric smoothing process described above. Color smoothing is used to eliminate discontinuities that may appear at patch boundaries due to image quality degradation that occurs during compression. Color smoothing can be performed through the following operations:

[0380] ① Use KD-trees or similar methods to calculate the neighboring points of each point in the reconstructed point cloud. The neighboring point information calculated in the geometric smoothing process described above can be used.

[0381] ② Determine whether each point lies on the patch boundary. These operations can be performed based on the boundary information calculated in the geometric smoothing process described above.

[0382] ③ Examine the distribution of color values ​​of neighboring points of points existing on the boundary and determine whether smoothing should be performed. For example, when the entropy of the brightness value is less than or equal to a threshold local entry (where many similar brightness values ​​exist), it can be determined that the corresponding part is not an edge part, and smoothing can be performed. As a smoothing method, the color value of a point can be replaced with the average of the color values ​​of its neighboring points.

[0383] Figure 18 This is a flowchart illustrating the operation of a transmission apparatus for compressing and transmitting V-PCC-based point cloud data according to an embodiment of the present disclosure.

[0384] The transmitting device according to the embodiment can correspond to Figure 1 The transmitting device Figure 4 Encoding processing, and Figure 15 A 2D video / image encoder, or some / all of which perform its operations. Each component of the transmitting device may correspond to software, hardware, a processor, and / or a combination thereof.

[0385] The operation of compressing and transmitting point cloud data using V-PCC at the sending terminal can be performed as shown in the figure.

[0386] The point cloud data transmitting device according to the embodiments may be referred to as a transmitting device or a transmitting system.

[0387] Regarding the patch generator 18000, it generates patches for 2D image mapping of the point cloud based on the input point cloud data. As a result of patch generation, patch information and / or auxiliary patch information are generated. The generated patch information and / or auxiliary patch information can be used in geometric image generation, texture image generation, smoothing, and geometric reconstruction for smoothing.

[0388] Patch packer 18001 performs patch packing processing, mapping patches generated by patch generator 18000 to a 2D image. For example, one or more patches can be packed. As a result of patch packing, an occupancy map can be generated. The occupancy map can be used in geometry image generation, geometry image filling, texture image filling, and / or for smooth geometry reconstruction.

[0389] The geometric image generator 18002 generates a geometric image based on point cloud data, patch information (or auxiliary patch information), and / or occupancy map. The generated geometric image is preprocessed by the encoding preprocessor 18003 and then encoded into a bitstream by the video encoder 18006.

[0390] The encoding preprocessor 18003 may include an image filling process. In other words, some spaces in the generated geometric image and the generated texture image may be filled with meaningless data. The encoding preprocessor 18003 may also include a group expansion process for the generated texture image or a texture image that has already undergone image filling.

[0391] The geometry reconstructor 18010 reconstructs a 3D geometric image based on the geometric bitstream encoded by the video encoder 18006, auxiliary patch information, and / or occupancy map.

[0392] Smoother 18009 smooths the 3D geometric image reconstructed and output by geometry reconstructor 18010 based on auxiliary patch information, and outputs the smoothed 3D geometric image to texture image generator 18004.

[0393] Texture image generator 18004 can generate texture images based on smooth 3D geometry, point cloud data, patches (or packed patches), patch information (or auxiliary patch information), and / or occupancy maps. The generated texture images can be preprocessed by encoding preprocessor 18003 and then encoded into a video bitstream by video encoder 18006.

[0394] The metadata encoder 18005 can encode auxiliary patch information into a metadata bitstream.

[0395] The video encoder 18006 can encode geometric and texture images output from the encoding preprocessor 18003 into corresponding video bitstreams, and can also encode occupancy maps into a video bitstream. According to an embodiment, the video encoder 18006 applies... Figure 15 The 2D video / image encoder encodes each input image.

[0396] Multiplexer 18007 multiplexes the video bitstream of geometry, the video bitstream of texture image, the video bitstream of occupancy map output from video encoder 18006, and the bitstream of metadata (including auxiliary patch information) output from metadata encoder 18005 into a single bitstream.

[0397] Transmitter 18008 sends the bit stream output from multiplexer 18007 to the receiving side. Alternatively, a file / fragment encapsulator can be provided between multiplexer 18007 and transmitter 18008, and the bit stream output from multiplexer 18007 can be encapsulated into files and / or fragments and output to transmitter 18008.

[0398] Figure 18The patch generator 18000, patch packer 18001, geometry image generator 18002, texture image generator 18004, metadata encoder 18005, and smoother 18009 can correspond to patch generation 14000, patch packing 14001, geometry image generation 14002, texture image generation 14003, auxiliary patch information compression 14005, and smoothing 14004, respectively. Figure 18 The encoding preprocessor 18003 may include Figure 4 Image fillers 14006 and 14007 and group expander 14008, and Figure 18 The video encoder 18006 may include Figure 4 Video compressors 14009, 14010, and 14011 and / or entropy compressor 14012. For those without reference... Figure 18 The described part, refer to Figures 4 to 15 The description is as follows. The above block can be omitted or replaced by a block with similar or identical functionality. Furthermore, Figure 18 Each block shown can be used as at least one of a processor, software, or hardware. Alternatively, the generated geometry, texture image, video bitstream of the occupancy map, and metadata bitstream of auxiliary patch information can be formed into one or more track data in a file, or encapsulated into fragments and sent to the receiving side via a transmitter.

[0399] The process of operating the receiving device

[0400] Figure 19 This is a flowchart illustrating the operation of a receiving apparatus for receiving and recovering V-PCC-based point cloud data according to an embodiment.

[0401] The receiving device according to the embodiment can correspond to Figure 1 The receiving device Figure 16 Decoding processing and Figure 17 A 2D video / image encoder, or some / all of which perform its operations. Each component of the receiving device may correspond to software, hardware, a processor, and / or a combination thereof.

[0402] The operation of receiving and reconstructing point cloud data using V-PCC at the receiving terminal can be performed as shown in the figure. The operation of the V-PCC receiving terminal can follow... Figure 18 The reverse processing of the operation of the V-PCC transmitting terminal.

[0403] The point cloud data receiving device according to the implementation method may be referred to as a receiving device, receiving system, etc.

[0404] The receiver receives the bitstream of the point cloud (i.e., the compressed bitstream), and the demultiplexer 19000 demultiplexes the bitstreams of the texture image, geometry image, and occupancy map image, as well as the bitstream of metadata (i.e., auxiliary patch information), from the received point cloud bitstream. The demultiplexed bitstreams of the texture image, geometry image, and occupancy map image are output to the video decoder 19001, and the bitstream of metadata is output to the metadata decoder 19002.

[0405] According to the implementation method, when Figure 18 When the sending device is equipped with a file / fragment encapsulator, the file / fragment decapsulator is set to... Figure 19 The receiving device is located between the receiver and the demultiplexer 19000. In this case, the transmitting device encapsulates and transmits the point cloud bitstream in the form of files and / or fragments, and the receiving device receives and decapsulates the files and / or fragments containing the point cloud bitstream.

[0406] The video decoder 19001 decodes the bitstreams of the geometric image, the texture image, and the occupancy map image into geometric images, texture images, and occupancy map images, respectively. According to an embodiment, the video decoder 19001 applies... Figure 17 A 2D video / image decoder is used to perform decoding operations on each input bitstream. The metadata decoder 19002 decodes the metadata bitstream into auxiliary patch information and outputs this information to the geometry reconstructor 19003.

[0407] The geometry reconstructor 19003 reconstructs 3D geometry based on the geometric images, occupancy maps, and / or auxiliary patch information output from the video decoder 19001 and the metadata decoder 19002.

[0408] Smoother 19004 smooths the 3D geometry reconstructed by geometry reconstructor 19003.

[0409] Texture reconstructor 19005 reconstructs the texture using the texture image and / or smoothed 3D geometry output from video decoder 19001. That is, texture reconstructor 19005 reconstructs the color point cloud image / screen by assigning color values ​​to smoothed 3D geometry using the texture image. Subsequently, to improve objective / subjective visual quality, additional color smoothing processing can be performed on the color point cloud image / screen by color smoother 19006. After rendering processing in point cloud renderer 19007, the modified point cloud image / screen derived through the above operations is displayed to the user. In some cases, color smoothing processing can be omitted.

[0410] The aforementioned blocks can be omitted or replaced by blocks with similar or identical functions. Furthermore, Figure 19Each block shown can be used as at least one of a processor, software, and hardware.

[0411] Figure 20 An exemplary architecture for V-PCC-based storage and streaming of point cloud data, according to an embodiment, is shown.

[0412] Figure 20 The system may include part / all of it Figure 1 Transmitting and receiving devices, Figure 4 Encoding processing, Figure 15 2D video / image encoder, Figure 16 Decoding processing, Figure 18 The transmitting device and / or Figure 19 Some or all of the receiving device. Each component in the diagram may correspond to software, hardware, a processor, and / or a combination thereof.

[0413] Figure 20 This diagram illustrates the overall architecture for storing or streaming point cloud data compressed using Video-Based Point Cloud Compression (V-PCC). The processing for storing and streaming point cloud data may include acquisition processing, encoding processing, transmission processing, decoding processing, rendering processing, and / or feedback processing.

[0414] The implementation proposes a method for efficiently providing point cloud media / content / data.

[0415] To efficiently deliver point cloud media / content / data, the point cloud acquirer 20000 can acquire point cloud video. For example, one or more cameras can acquire point cloud data by capturing, orchestrating, or generating point clouds. This acquisition process allows the acquisition of point cloud video that includes the 3D positions of each point (represented by x, y, and z position values, etc.) (hereinafter referred to as geometry) and the attributes of each point (color, reflectivity, transparency, etc.). For example, a Polygon file format (PLY) (or Stanford triangle format) file containing the point cloud video can be generated. For point cloud data with multiple frames, one or more files can be acquired. During this process, point cloud-related metadata (e.g., metadata related to the capture, etc.) can be generated.

[0416] Captured point cloud videos may require post-processing to improve content quality. During video capture processing, the maximum / minimum depth can be adjusted within the range provided by the camera device. Even after adjustment, unwanted areas of point data may still exist. Therefore, post-processing can be performed to remove unwanted areas (e.g., background) or to identify connected spaces and fill in spatial holes. Additionally, point clouds extracted from cameras in a shared spatial coordinate system can be integrated into a single content by transforming individual points to a global coordinate system based on the position coordinates of each camera obtained through calibration processing. This results in a point cloud video with high-density points.

[0417] The point cloud preprocessor 20001 can generate one or more frames of a point cloud video. Generally, a frame can be a unit representing an image at specific time intervals. Furthermore, when dividing the points constituting the point cloud video into one or more patches and mapping them to a 2D plane, the point cloud preprocessor 20001 can generate occupancy map frames with values ​​of 0 or 1, which are binary maps indicating the presence or absence of data at corresponding positions in the 2D plane. Here, a patch is a set of points constituting the point cloud video, where points belonging to the same patch are adjacent to each other in 3D space and are mapped to the same face among the flat faces of a 6-face bounding box when mapped to a 2D image. Additionally, the point cloud preprocessor 20001 can generate geometric frames in the form of depth maps representing information about the position (geometry) of each point constituting the point cloud video per patch. The point cloud preprocessor 20001 can also generate texture frames representing color information about each point constituting the point cloud video per patch. In this process, metadata required to reconstruct the point cloud from the individual patches can be generated. Metadata can contain information about patches (auxiliary information or auxiliary patch information), such as the position and size of each patch in 2D / 3D space. These images / frames can be generated sequentially in time to construct a video stream or metadata stream.

[0418] The point cloud video encoder 20002 can encode one or more video streams associated with point cloud video. A video may include multiple frames, and a frame may correspond to a still image / picture. In this disclosure, point cloud video may include point cloud images / frames / pictures, and the term "point cloud video" is used interchangeably with point cloud video / frames / pictures. The point cloud video encoder 20002 can perform video-based point cloud compression (V-PCC) processing. For compression and encoding efficiency, the point cloud video encoder 20002 can perform a series of processes such as prediction, transform, quantization, and entropy coding. The encoded data (encoded video / image information) can be output as a bitstream. Based on V-PCC processing, as described below, the point cloud video encoder 20002 can encode point cloud video by dividing it into geometric video, attribute video, occupancy map video, and metadata (e.g., information about patches). Geometric video may include geometric images, attribute video may include attribute images, and occupancy map video may include occupancy map images. Patch data as auxiliary information may include patch-related information. Attribute videos / images can include textured videos / images.

[0419] The point cloud image encoder 20003 can encode one or more images associated with a point cloud video. The point cloud image encoder 20003 can perform video-based point cloud compression (V-PCC) processing. For compression and encoding efficiency, the point cloud image encoder 20003 can perform a series of processes such as prediction, transform, quantization, and entropy coding. The encoded image can be output as a bitstream. Based on V-PCC processing, as described below, the point cloud image encoder 20003 can encode the point cloud image by dividing it into a geometric image, an attribute image, an occupancy map image, and metadata (e.g., information about patches).

[0420] According to the implementation, the point cloud video encoder 20002, point cloud image encoder 20003, point cloud video decoder 20006, and point cloud image decoder 20008 can be executed by one encoder / decoder as described above, and can be executed along separate paths, as shown in the figure.

[0421] In the file / fragment encapsulator 20004, encoded point cloud data and / or point cloud-related metadata can be encapsulated into files or fragments for streaming. Here, the point cloud-related metadata can be received from a metadata processor (not shown), etc. The metadata processor can be included in the point cloud video / image encoder 20002 / 20003, or can be configured as a separate component / module. The file / fragment encapsulator 20004 can encapsulate the corresponding video / image / metadata in a file format such as ISOBMFF or in the form of DASH fragments, etc. According to an embodiment, the file / fragment encapsulator 20004 can include point cloud metadata in a file format. Point cloud-related metadata can be included in boxes of various levels, such as ISOBMFF file format, or as data in a separate track within a file. According to an embodiment, the file / fragment encapsulator 20004 can encapsulate point cloud-related metadata into a file.

[0422] The file / fragment encapsulator 20004 according to the implementation can store a bitstream or individual bitstreams into one or more tracks in a file, and can also encapsulate signaling information used for this operation. Furthermore, atlas streams (or patch streams) included in the bitstream can be stored as tracks in the file, and associated signaling information can be stored. Additionally, SEI messages present in the bitstream can be stored as tracks in the file, and associated signaling information can be stored.

[0423] A transmission processor (not shown) can perform transmission processing of encapsulated point cloud data according to a file format. The transmission processor can be included in a transmitter (not shown) or configured as a separate component / module. The transmission processor can process the point cloud data according to a transmission protocol. Transmission processing can include processing via broadcast networks and processing via broadband transmission. According to one embodiment, the transmission processor can receive point cloud-related metadata and point cloud data from a metadata processor and perform transmission processing of point cloud video data.

[0424] The transmitter can send a point cloud bitstream or a file / fragment including the bitstream to a receiver (not shown) via a digital storage medium or network. For transmission, processing according to any transmission protocol can be performed. The data processed for transmission can be transmitted via a broadcast network and / or broadband. Data can be transmitted to the receiving side on demand. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitter can include elements for generating media files in a predetermined file format and may include elements for transmission via a broadcast / communication network. The receiver can extract the bitstream and send the extracted bitstream to a decoder.

[0425] The receiver can receive point cloud data transmitted by the point cloud data transmitting device according to this disclosure. Depending on the transmission channel, the receiver can receive the point cloud data via a broadcast network or via broadband. Alternatively, the point cloud data can be received via a digital storage medium. The receiver may include processing for decoding the received data and rendering the data according to the user's viewport.

[0426] A receiving processor (not shown) can perform processing on the received point cloud video data according to the transmission protocol. The receiving processor can be included in the receiver or configured as a separate component / module. The receiving processor can conversely perform the processing of the transmitting processor described above, corresponding to the transmission processing performed on the transmitting side. The receiving processor can transmit the acquired point cloud video to the file / fragment decapsulator 20005 and transmit the acquired point cloud-related metadata to the metadata parser.

[0427] File / fragment decapsulator 20005 can decapsulate point cloud data received from a receiving processor in file form. File / fragment decapsulator 20005 can decapsulate files according to ISOBMFF or similar formats and can acquire point cloud bitstreams or point cloud-related metadata (or separate metadata bitstreams). The acquired point cloud bitstreams can be transmitted to point cloud video decoder 20006 and point cloud image decoder 20008, and the acquired point cloud video-related metadata (metadata bitstreams) can be transmitted to a metadata processor (not shown). The point cloud bitstreams may include metadata (metadata bitstreams). The metadata processor may be included in the point cloud video decoder 20006 or may be configured as a separate component / module. The point cloud video-related metadata acquired by file / fragment decapsulator 20005 may take the form of boxes or tracks in a file format. If necessary, file / fragment decapsulator 20005 can receive metadata required for decapsulation from the metadata processor. Point cloud-related metadata can be transmitted to point cloud video decoder 20006 and / or point cloud image decoder 20008 and used in point cloud decoding processing, or it can be transmitted to renderer 20009 and used in point cloud rendering processing.

[0428] The point cloud video decoder 20006 can receive a bitstream and decode the video / image by performing operations corresponding to those of the point cloud video encoder 20002. In this case, as described below, the point cloud video decoder 20006 can decode the point cloud video by dividing it into geometric video, attribute video, occupancy map video, and auxiliary patch information. The geometric video may include a geometric image, the attribute video may include an attribute image, and the occupancy map video may include an occupancy map image. The auxiliary information may include auxiliary patch information. The attribute video / image may include texture video / image.

[0429] The point cloud image decoder 20008 can receive a bitstream and perform the inverse processing corresponding to the operations of the point cloud image encoder 20003. In this case, the point cloud image decoder 20008 can divide the point cloud image into a geometric image, an attribute image, an occupancy map image, and metadata (which is, for example, auxiliary patch information) for decoding.

[0430] 3D geometry can be reconstructed based on decoded geometric video / images, occupancy maps, and auxiliary patch information, and then subjected to smoothing. Color point cloud images / pictures can be reconstructed by assigning color values ​​to the smoothed 3D geometry based on textured video / images. Renderer 20009 can render the reconstructed geometry and color point cloud images / pictures. The rendered video / images can be displayed on a monitor. All or part of the rendering results can be displayed to the user on a VR / AR monitor or a typical monitor.

[0431] The sensor / tracker (sensing / tracking) 20007 acquires orientation information and / or user viewport information from the user or receiving side and transmits the orientation information and / or user viewport information to the receiver and / or transmitter. Orientation information may represent information about the position, angle, movement, etc., of the user's head, or information about the position, angle, movement, etc., of the device through which the user is viewing video / images. Based on this information, information about the area currently being viewed by the user in 3D space (i.e., viewport information) can be calculated.

[0432] Viewport information can be information about the area in 3D space that the user is currently viewing through the device or HMD. Devices such as displays can extract the viewport area based on orientation information, the vertical or horizontal field of view supported by the device, etc. Orientation or viewport information can be extracted or calculated on the receiving side. The orientation or viewport information analyzed on the receiving side can be transmitted to the transmitting side on the feedback channel.

[0433] Based on the orientation information and / or viewport information indicating the area currently viewed by the user acquired by the sensor / tracker 20007, the receiver can effectively extract or decode media data from the file only for a specific area (i.e., the area indicated by the orientation information and / or viewport information). Additionally, based on the orientation information and / or viewport information acquired by the sensor / tracker 20007, the transmitter can effectively encode, or generate and transmit the file only for media data in that specific area (i.e., the area indicated by the orientation information and / or viewport information).

[0434] Renderer 20009 can render decoded point cloud data in 3D space. Rendered videos / images can be displayed on a monitor. Users can view all or part of the rendering results through VR / AR displays or typical monitors.

[0435] Feedback processing may include a decoder that transmits various feedback information, which can be obtained during rendering / display processing, to the sending or receiving side. Feedback processing enables interactivity when consuming point cloud data. According to one implementation, head orientation information, viewport information indicating the area the user is currently viewing, etc., may be transmitted to the sending side during feedback processing. According to another implementation, the user can interact with content implemented in a VR / AR / MR / autonomous driving environment. In this case, information related to the interaction may be transmitted to the sending side or service provider during feedback processing. According to yet another implementation, feedback processing may be skipped.

[0436] According to the implementation method, the aforementioned feedback information can be sent not only to the sending side but also consumed at the receiving side. That is, the decapsulation, decoding, and rendering processes at the receiving side can be performed based on the aforementioned feedback information. For example, point cloud data about the area currently being viewed by the user can be preferentially decapsulated, decoded, and rendered based on orientation information and / or viewport information.

[0437] Figure 21 This is an exemplary block diagram of an apparatus for storing and transmitting point cloud data according to an embodiment.

[0438] Figure 21 A point cloud system according to an embodiment is shown. Figure 21 The system may include part / all of it Figure 1 Transmitting and receiving devices, Figure 4 Encoding processing, Figure 15 2D video / image encoder, Figure 16 Decoding processing, Figure 18 The transmitting device and / or Figure 19 Some or all of the receiving devices. Furthermore, it can include or correspond to... Figure 20 Part / all of the system.

[0439] The point cloud data transmission device according to the embodiment can be configured as shown in the figure. The various components of the transmission device can be modules / units / components / hardware / software / processors.

[0440] The geometry, attributes, occupancy map, auxiliary data (auxiliary information), and mesh data of a point cloud can each be configured as a separate stream or stored in different tracks within a file. Furthermore, they can be included in separate fragments.

[0441] Point cloud acquirer 21000 acquires point clouds. For example, one or more cameras can acquire point cloud data by capturing, arranging, or generating point clouds. Through this acquisition process, point cloud data including the 3D position of each point (which can be represented by x, y, and z position values, etc.) (hereinafter referred to as geometry) and the attributes of each point (color, reflectivity, transparency, etc.) can be acquired. For example, a Polygon file format (PLY) (or Stanford triangle format) file including point cloud data can be generated. For point cloud data with multiple frames, one or more files can be acquired. In this process, point cloud-related metadata (e.g., metadata related to the capture, etc.) can be generated. Patch generator 21001 generates patches from point cloud data. Patch generator 21001 generates one or more frames from point cloud data or point cloud video. A frame can typically represent a unit representing an image at specific time intervals. When the points constituting a point cloud video are divided into one or more patches (a set of points constituting the point cloud video, where points belonging to the same patch are adjacent to each other in 3D space and mapped in the same direction between the flat faces of a 6-sided bounding box when mapped to a 2D image) and mapped to a 2D plane, a binary occupancy map frame can be generated, indicating the presence of data at the corresponding location in the 2D plane with 0 or 1. Additionally, a geometric frame in the form of a depth map representing the position (geometry) of each point constituting the point cloud video can be generated patch by patch. A texture frame representing the color information of each point constituting the point cloud video can also be generated patch by patch. During this process, metadata required to reconstruct the point cloud from the individual patches can be generated. The metadata may include information about the patches, such as the position and size of each patch in 2D / 3D space. These frames can be generated sequentially in time to construct a video stream or a metadata stream.

[0442] Additionally, patches can be used for 2D image mapping. For example, point cloud data can be projected onto the faces of a cube. After patch generation, geometric images, one or more attribute images, occupancy maps, auxiliary data, and / or mesh data can be generated based on the generated patches.

[0443] The generation of geometric images, attribute images, occupancy maps, auxiliary data, and / or mesh data is performed by the point cloud preprocessor 20001 or a controller (not shown). The point cloud preprocessor 20001 may include a patch generator 21001, a geometric image generator 21002, an attribute image generator 21003, an occupancy map generator 21004, an auxiliary data generator 21005, and a mesh data generator 21006.

[0444] The geometry image generator 21002 generates a geometry image based on the results of patch generation. The geometry represents the position of points in 3D space. An occupancy map is used to generate the geometry image, which includes information related to the packing of the 2D image of the patch, auxiliary data (including patch data), and / or patch-based mesh data. The geometry image is related to information such as the depth (e.g., near, far) of the patch generated after patch generation.

[0445] The attribute image generator 21003 generates an attribute image. For example, an attribute may represent a texture. The texture may be a color value matched to individual points. According to an embodiment, an image including multiple attributes of the texture (e.g., color and reflectance) (N attributes) can be generated. The multiple attributes may include material information and reflectance. According to an embodiment, the attributes may additionally include information indicating color, which may vary depending on the viewing angle and light, even for the same texture.

[0446] Occupancy map generator 21004 generates an occupancy map from the patch. The occupancy map includes information indicating whether data exists in pixels (e.g., corresponding to a geometric or attribute image).

[0447] The auxiliary data generator 21005 generates auxiliary data (or auxiliary information) that includes information about the patch. That is, the auxiliary data represents metadata about the patch of a point cloud object. For example, it may represent information such as the patch's normal vector. Specifically, the auxiliary data may include information needed to reconstruct the point cloud from the patch (e.g., information about the patch's position, size, etc. in 2D / 3D space, as well as projection (normal) plane identification information, patch mapping information, etc.).

[0448] The mesh data generator 21006 generates mesh data from patches. A mesh represents the connection between neighboring points. For example, it can represent data in the shape of a triangle. Mesh data refers to the connectivity between points.

[0449] The point cloud preprocessor 20001 or controller generates metadata related to patch generation, geometric image generation, attribute image generation, occupancy map generation, auxiliary data generation, and mesh data generation.

[0450] The point cloud transmitting device performs video encoding and / or image encoding in response to the results generated by the point cloud preprocessor 20001. The point cloud transmitting device can generate point cloud image data and point cloud video data. According to embodiments, the point cloud data may contain only video data, only image data, and / or both video data and image data.

[0451] The video encoder 21007 performs geometric video compression, attribute video compression, occupancy graph video compression, auxiliary data compression, and / or mesh data compression. The video encoder 21007 generates a video stream containing encoded video data.

[0452] Specifically, in geometric video compression, point cloud geometric video data is encoded. In attribute video compression, point cloud attribute video data is encoded. In auxiliary data compression, auxiliary data associated with the point cloud video data is encoded. In mesh data compression, the mesh data of the point cloud video data is encoded. The various operations of the point cloud video encoder can be executed in parallel.

[0453] The image encoder 21008 performs geometric image compression, attribute image compression, occupancy map image compression, auxiliary data compression, and / or mesh data compression. The image encoder generates an image containing encoded image data.

[0454] Specifically, in geometric image compression, point cloud geometric image data is encoded. In attribute image compression, point cloud attribute image data is encoded. In auxiliary data compression, auxiliary data associated with the point cloud image data is encoded. In mesh data compression, mesh data associated with the point cloud image data is encoded. The various operations of the point cloud image encoder can be executed in parallel.

[0455] The video encoder 21007 and / or the image encoder 21008 can receive metadata from the point cloud preprocessor 20001. The video encoder 21007 and / or the image encoder 21008 can perform individual encoding processes based on the metadata.

[0456] File / fragment wrapper 21009 encapsulates video streams and / or images in the form of files and / or fragments. File / fragment wrapper 21009 performs video track encapsulation, metadata track encapsulation, and / or image encapsulation.

[0457] In video track encapsulation, one or more video streams can be encapsulated into one or more tracks.

[0458] In metadata track encapsulation, metadata related to the video stream and / or images can be encapsulated in one or more tracks. Metadata includes data related to the content of the point cloud data. For example, it may include initial viewing orientation metadata. Depending on the implementation, metadata may be encapsulated into a metadata track, or it may be encapsulated together in a video track or an image track.

[0459] In image encapsulation, one or more images can be encapsulated into one or more tracks or projects.

[0460] For example, according to an implementation, when four video streams and two images are input to the encapsulator, the four video streams and two images can be encapsulated in one file.

[0461] The file / fragment wrapper 21009 can receive metadata from the point cloud preprocessor 20001. The file / fragment wrapper 21009 can perform wrapping based on the metadata.

[0462] Files and / or fragments generated by the file / fragment encapsulator 21009 are sent by a point cloud sending device or transmitter. For example, fragments may be transmitted according to a DASH-based protocol.

[0463] The transmitter can send point cloud bitstreams or files / fragments including bitstreams to a receiver of a receiving device via digital storage media or a network. For transmission, processing according to any transmission protocol can be performed. Processed data can be transmitted via broadcast networks and / or broadband. Data can be transmitted to the receiving side on demand. Digital storage media can include a variety of storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD.

[0464] The file / fragment encapsulator 21009 according to the implementation can divide and store a bitstream or individual bitstreams into one or more tracks in a file, and can encapsulate signaling information used therein. Furthermore, patch (or atlas) streams included in the bitstream can be stored as tracks in the file, and associated signaling information can be stored therein. Additionally, SEI messages present in the bitstream can be stored as tracks in the file, and associated signaling information can be stored therein.

[0465] The transmitter may include elements for generating media files in a predetermined file format and may include elements for transmission via a broadcast / communication network. The transmitter receives orientation information and / or viewport information from a receiver. The transmitter may transmit the acquired orientation information and / or viewport information (or user-selected information) to a point cloud preprocessor 20001, a video encoder 21007, an image encoder 21008, a file / fragment encapsulator 21009, and / or a point cloud encoder. Based on the orientation information and / or viewport information, the point cloud encoder may encode all point cloud data or point cloud data indicated by the orientation information and / or viewport information. Based on the orientation information and / or viewport information, the file / fragment encapsulator may encapsulate all point cloud data or point cloud data indicated by the orientation information and / or viewport information. Based on the orientation information and / or viewport information, the transmitter may transmit all point cloud data or point cloud data indicated by the orientation information and / or viewport information.

[0466] For example, the point cloud preprocessor 21001 can perform the above operations on all point cloud data or on point cloud data indicated by orientation information and / or viewport information. The video encoder 21007 and / or the image encoder 21008 can perform the above operations on all point cloud data or on point cloud data indicated by orientation information and / or viewport information. The file / fragment encapsulator 21009 can perform the above operations on all point cloud data or on point cloud data indicated by orientation information and / or viewport information. The transmitter can perform the above operations on all point cloud data or on point cloud data indicated by orientation information and / or viewport information.

[0467] Figure 22 This is an exemplary block diagram of a point cloud data receiving device according to an embodiment.

[0468] Figure 22 A point cloud system according to an embodiment is shown. Figure 22 The system may include part / all of it Figure 1 Transmitting and receiving devices, Figure 4 Encoding processing, Figure 15 2D video / image encoder, Figure 16 Decoding processing, Figure 18 The transmitting device and / or Figure 19 Some or all of the receiving devices. Furthermore, it can include or correspond to... Figure 20 and Figure 21 Part / all of the system.

[0469] The various components of the receiving device can be modules / units / components / hardware / software / processors. The transmitting client can receive point cloud data, point cloud bitstreams, or files / fragments, including bitstreams transmitted by the point cloud data transmitting device according to the embodiment. Depending on the channel used for transmission, the receiver can receive point cloud data via a broadcast network or via broadband. Alternatively, point cloud data can be received via a digital storage medium. The receiver can include processing for decoding the received data and rendering the received data according to a user viewport. The transmitting client (receiving processor) 22006 can perform processing on the received point cloud data according to a transmission protocol. The receiving processor can be included in the receiver or configured as a separate component / module. The receiving processor can conversely perform the processing of the transmitting processor described above to correspond to the transmission processing performed on the transmitting side. The receiving processor can transmit the acquired point cloud data to the file / fragment decapsulator 22000 and the acquired point cloud-related metadata to the metadata processor (not shown).

[0470] Sensor / tracker 22005 acquires orientation information and / or viewport information. Sensor / tracker 22005 can transmit the acquired orientation information and / or viewport information to transmission client 22006, file / fragment decapsulator 22000, point cloud decoders 22001 and 22002, and point cloud processor 22003.

[0471] The transmission client 22006 can receive all point cloud data or point cloud data indicated by the orientation information and / or viewport information based on orientation information and / or viewport information. The file / fragment decapsulator 22000 can decapsulate all point cloud data or point cloud data indicated by the orientation information and / or viewport information based on orientation information and / or viewport information. The point cloud decoder (video decoder 22001 and / or image decoder 22002) can decode all point cloud data or point cloud data indicated by the orientation information and / or viewport information based on orientation information and / or viewport information. The point cloud processor 22003 can process all point cloud data or point cloud data indicated by the orientation information and / or viewport information based on orientation information and / or viewport information.

[0472] File / fragment decapsulator 22000 performs video track decapsulation, metadata track decapsulation, and / or image decapsulation. File / fragment decapsulator 22000 can decapsulate point cloud data received from a receiving processor in file format. File / fragment decapsulator 22000 can decapsulate files or fragments according to ISOBMFF, etc., to obtain point cloud bitstreams or point cloud-related metadata (or separate metadata bitstreams). The obtained point cloud bitstreams can be transmitted to point cloud decoders 22001 and 22002, and the obtained point cloud-related metadata (or metadata bitstreams) can be transmitted to a metadata processor (not shown). The point cloud bitstreams may include metadata (metadata bitstreams). The metadata processor may be included in the point cloud video decoder or may be configured as a separate component / module. The point cloud-related metadata obtained by file / fragment decapsulator 22000 may take the form of boxes or tracks in a file format. If necessary, file / fragment decapsulator 22000 may receive metadata required for decapsulation from the metadata processor. Point cloud-related metadata can be transmitted to point cloud decoders 22001 and 22002 and used in point cloud decoding processing, or it can be transmitted to renderer 22004 and used in point cloud rendering processing. File / fragment decapsulator 22000 can generate metadata related to point cloud data.

[0473] In video track decapsulation performed by the file / segment decapsulator 22000, video tracks contained in files and / or segments are decapsulated. Video streams including geometric video, attribute video, occupancy maps, auxiliary data, and / or grid data are decapsulated.

[0474] In the metadata track decapsulation performed by the file / fragment decapsulator 22000, the bitstream including metadata related to point cloud data and / or auxiliary data is decapsulated.

[0475] In image decapsulation performed by file / fragment decapsulator 22000, images including geometric images, attribute images, occupancy maps, auxiliary data, and / or mesh data are decapsulated.

[0476] The file / fragment decapsulator 22000 according to the embodiments can store a bitstream or individual bitstreams into one or more tracks in a file, and can also decapsulate the signaling information used therein. Furthermore, streams included in the bitstream or atlas (patch) streams can be decapsulated based on tracks in the file, and the associated signaling information can be parsed. Additionally, SEI messages present in the bitstream can be decapsulated based on tracks in the file, and the associated signaling information can also be obtained.

[0477] The video decoder 22001 performs geometric video decompression, attribute video decompression, occupancy map decompression, auxiliary data decompression, and / or mesh data decompression. The video decoder 22001 decodes the geometric video, attribute video, auxiliary data, and / or mesh data in a process corresponding to the process performed by the video encoder of the point cloud transmitting apparatus according to the embodiment.

[0478] The image decoder 22002 performs geometric image decompression, attribute image decompression, occupancy map decompression, auxiliary data decompression, and / or mesh data decompression. The image decoder 22002 decodes the geometric image, attribute image, auxiliary data, and / or mesh data in a process corresponding to the process performed by the image encoder of the point cloud transmitting apparatus according to the embodiment.

[0479] According to the embodiments, video decoder 22001 and video decoder 22002 can be processed by a video / image decoder as described above, and can be executed along separate paths, as shown in the figure.

[0480] The video decoder 22001 and / or the image decoder 22002 can generate metadata related to video data and / or image data.

[0481] In the point cloud processor 22003, geometric reconstruction and / or attribute reconstruction are performed.

[0482] In geometric reconstruction, geometric video and / or geometric images are reconstructed from decoded video data and / or decoded image data based on occupancy maps, auxiliary data, and / or grid data.

[0483] In attribute reconstruction, attribute videos and / or attribute images are reconstructed from decoded attribute videos and / or decoded attribute images based on occupancy maps, auxiliary data, and / or mesh data. According to an implementation, for example, the attribute can be a texture. According to an implementation, the attribute can represent multiple attribute information. When multiple attributes exist, the point cloud processor 22003 according to an implementation performs multiple attribute reconstructions.

[0484] The point cloud processor 22003 can receive metadata from the video decoder 22001, the image decoder 22002, and / or the file / fragment decapsulator 22000, and process the point cloud based on the metadata.

[0485] Point cloud renderer 22004 renders the reconstructed point cloud. Point cloud renderer 22004 can receive metadata from video decoder 22001, image decoder 22002 and / or file / fragment decapsulator 22000, and render the point cloud based on the metadata.

[0486] The monitor displays the rendered results on the actual display device.

[0487] According to the method / apparatus of the embodiment, such as Figures 20 to 22 As shown, the transmitting side can encode point cloud data into a bitstream, encapsulate the bitstream into files and / or fragments, and send it. The receiving side can decapsulate the files and / or fragments into a bitstream containing point clouds, and can decode the bitstream back into point cloud data. For example, the point cloud data apparatus according to the embodiment can encapsulate point cloud data based on files. The file may include a V-PCC track containing point cloud parameters, a geometry track containing geometry, an attribute track containing attributes, and an occupancy track containing an occupancy map.

[0488] Furthermore, the point cloud data receiving device according to the embodiment decapsulates point cloud data based on a file. The file may include a V-PCC track containing point cloud parameters, a geometry track containing geometry, an attribute track containing attributes, and an occupancy track containing an occupancy map.

[0489] The above packaging operation can be performed by Figure 20 File / Fragment Wrapper 20004 Figure 21 The file / fragment wrapper 21009 performs the above decapsulation operation. Figure 20 File / fragment decapsulator 20005 or Figure 22 The file / fragment decapsulator 22000 is executed.

[0490] Figure 23 An exemplary structure is shown that can be combined with the point cloud data transmission / reception method / apparatus according to the embodiments.

[0491] In the structure according to the embodiment, at least one of the following is connected to the cloud network 23000: server 23600, robot 23100, self-driving vehicle 23200, XR device 23300, smartphone 23400, home appliance 23500, and / or head-mounted display (HMD) 23700. Here, robot 23100, self-driving vehicle 23200, XR device 23300, smartphone 23400, or home appliance 23500 may be referred to as a device. Additionally, XR device 23300 may correspond to a point cloud data (PCC) device according to the embodiment, or may be operatively connected to a PCC device.

[0492] Cloud Network 23000 can refer to a network that forms part of or exists within a cloud computing infrastructure. Here, Cloud Network 23000 can be configured using a 3G network, a 4G or Long Term Evolution (LTE) network, or a 5G network.

[0493] Server 23600 can be connected via cloud network 23000 to at least one of robot 23100, self-driving vehicle 23200, XR device 23300, smartphone 23400, home appliance 23500 and / or HMD 23700, and can assist at least a portion of the processing of the connected devices 23100 to 23700.

[0494] HMD 23700 represents one of the implementation types of an XR device and / or PCC device according to an embodiment. An HMD-type device according to an embodiment includes a communication unit, a control unit, a memory, an I / O unit, a sensor unit, and a power supply unit.

[0495] Hereinafter, various embodiments of the apparatus 23100 to 23500 that apply the above-described technology will be described. Figure 23 The devices 23100 to 23500 shown can be operatively connected to / coupled to the point cloud data transmission and reception device according to the above embodiments.

[0496] <PCC+XR>

[0497] The XR / PCC device 23300 can employ PCC technology and / or XR (AR+VR) technology, and can be implemented as an HMD, a head-up display (HUD) installed in a vehicle, a television, a mobile phone, a smartphone, a computer, a wearable device, a home appliance, a digital signage, a vehicle, a stationary robot, or a mobile robot.

[0498] The XR / PCC device 23300 can analyze 3D point cloud data or image data obtained through various sensors or from external devices and generate position data and attribute data regarding the 3D points. Thus, the XR / PCC device 23300 can obtain information about the surrounding space or a real object and render and output an XR object. For example, the XR / PCC device 23300 can match an XR object including auxiliary information regarding the recognized object with the recognized object and output the matched XR object.

[0499] <PCC + Autonomous Driving + XR>

[0500] The autonomous driving vehicle 23200 can be implemented as a mobile robot, a vehicle, an unmanned aerial vehicle, etc. by applying PCC technology and XR technology.

[0501] The autonomous driving vehicle 23200 applying the XR / PCC technology can represent an autonomous vehicle provided with means for providing an XR image or an autonomous vehicle as a control / interaction target in the XR image. Specifically, the autonomous driving vehicle 23200 as a control / interaction target in the XR image can be distinguished from the XR device 23300 and operably connected thereto.

[0502] The autonomous driving vehicle 23200 having means for providing an XR / PCC image can obtain sensor information from sensors including a camera and output the generated XR / PCC image based on the obtained sensor information. For example, the autonomous driving vehicle can have a HUD and output the XR / PCC image thereto to provide an XR / PCC object corresponding to a real object or an object existing on the screen to a passenger.

[0503] In this case, when the XR / PCC object is output to the HUD, at least a part of the XR / PCC object can be output to overlap with the real object pointed at by the passenger's eyes. On the other hand, when the XR / PCC object is output on a display provided inside the autonomous driving vehicle, at least a part of the XR / PCC object can be output to overlap with the object on the screen. For example, the autonomous driving vehicle can output an XR / PCC object corresponding to objects such as a road, another vehicle, a traffic light, a traffic sign, a two-wheeler, a pedestrian, and a building.

[0504] According to an embodiment, virtual reality (VR) technology, augmented reality (AR) technology, mixed reality (MR) technology, and / or point cloud compression (PCC) technology are applicable to various devices.

[0505] In other words, VR technology is a display technology that only provides real-world objects, backgrounds, etc., as CG images. On the other hand, AR technology refers to the technology of displaying CG images virtually created on top of real-world object images. MR technology is similar to AR technology in that the virtual objects to be displayed are mixed and combined with the real world. However, MR technology differs from AR technology. AR technology clearly distinguishes between real objects and virtual objects created as CG images and uses virtual objects as supplementary objects to real objects, while MR technology treats virtual objects as objects with the same characteristics as real objects. More specifically, an example of MR technology application is holographic services.

[0506] Recently, VR, AR, and MR technologies have often been referred to as extended reality (XR) technologies rather than clearly distinguished from each other. Therefore, the embodiments of this disclosure are applicable to all VR, AR, MR, and XR technologies. For these technologies, encoding / decoding based on PCC, V-PCC, and G-PCC technologies can be applied.

[0507] The PCC method / apparatus according to the embodiments can be applied to self-driving vehicles 23200 that provide self-driving services.

[0508] The self-driving vehicle 23200, which provides self-driving services, connects to the PCC device for wired / wireless communication.

[0509] When the point cloud data compression transmitting and receiving device (PCC device) according to the embodiment is connected to the autonomous vehicle 23200 for wired / wireless communication, the device can receive and process content data related to AR / VR / PCC services that can be provided with the autonomous driving service and transmit the processed content data to the autonomous vehicle 23200. When the point cloud data transmitting and receiving device is installed in the vehicle, the device can receive and process content data related to the AR / VR / PCC service based on user input signals input through a user interface device and provide the processed content data to the user. The autonomous vehicle 23200 or the user interface device according to the embodiment can receive user input signals. The user input signals according to the embodiment may include signals indicating autonomous driving services.

[0510] As mentioned above, Figure 1 , Figure 4 , Figure 18 , Figure 20 or Figure 21The V-PCC-based point cloud video encoder projects 3D point cloud data (or content) into 2D space to generate patches. Patches are generated in 2D space by dividing the data into geometric images representing positional information (called geometric frames or geometric patch frames) and texture images representing color information (called attribute frames or attribute patch frames). For each frame, the geometric and texture images are video compressed, outputting a video bitstream of the geometric image (called the geometric bitstream) and a video bitstream of the texture image (called the attribute bitstream). Furthermore, auxiliary patch information (also called patch information, metadata, or Atlas data), including projection plane information and patch size information for each patch (these are required for decoding 2D patches on the receiving side), is also video compressed, and a bitstream of auxiliary patch information is output. Additionally, an occupancy map, indicating the presence / absence of each pixel as 0 or 1, is entropy-compressed or video-compressed depending on whether it is in lossless or lossy mode, and a video bitstream of the occupancy map (or occupancy map bitstream) is output. The structure of a V-PCC bitstream is formed by multiplexing the compressed geometric bitstream, the compressed attribute bitstream, the compressed auxiliary patch information bitstream (also known as the Atlas bitstream), and the compressed occupancy map bitstream.

[0511] According to the implementation method, the V-PCC bitstream can be sent to the receiving side as is, or it can be... Figure 1 , Figure 18 , Figure 20 or Figure 21 The file / fragment encapsulator encapsulates files / fragments and transmits them to a receiving device or stores them in a digital storage medium (e.g., USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.). According to embodiments of this disclosure, the file is in ISOBMFF file format.

[0512] Depending on the implementation, the V-PCC bitstream can be sent via multiple tracks in a file, or via a single track. Details will be described later.

[0513] In this document, point cloud data (i.e., V-PCC data) represents the volumetric encoding of a point cloud consisting of a series of point cloud frames. In a point cloud sequence (i.e., a sequence of point cloud frames), each point cloud frame comprises a set of points. Each point can have a 3D location (i.e., geometric information) as well as multiple attributes, such as color, reflectivity, and surface normal. That is, each point cloud frame refers to a set of 3D points specified by the Cartesian coordinates (x, y, z) (i.e., location) of the 3D points and zero or more attributes for a specific time instance.

[0514] The video-based point cloud compression (V-PCC) described in this document is the same as the visual volumetric video encoding (V3C). V-PCC according to the implementation can be used interchangeably with V3C.

[0515] According to the implementation method, point cloud content (also known as V-PCC content or V3C content) refers to volumetric media encoded using V-PCC.

[0516] According to the implementation method, a volumetric scene can refer to 3D data and can consist of one or more objects. That is, a volumetric scene is a region or unit composed of one or more objects that constitute the volumetric medium. Furthermore, when the V-PCC bitstream is encapsulated into a file format for transmission, the region obtained by dividing the bounding box of the entire volumetric medium according to a spatial reference is called a 3D spatial region. According to the implementation method, the 3D spatial region can be referred to as a 3D region or a spatial region.

[0517] According to the implementation, an object can refer to point cloud data, volumetric media, or V3C content. An object can be divided into several objects based on a spatial reference, and in this specification, each divided object is referred to as a sub-object or simply an object. According to the implementation, a 3D bounding box can be information representing the position of an object in 3D space, and a 2D bounding box can represent a rectangular region surrounding a patch corresponding to an object in a 2D frame. That is, the data generated after the process of projecting an object onto a 2D plane (which is a process of encoding an object) is a patch, and the box surrounding the patch can be called a 2D bounding box. In other words, since an object can be composed of several patches during the encoding process, an object is associated with patches. An object in 3D space is represented by 3D bounding box information surrounding the object. Since an Atlas frame includes patch information corresponding to each object and 3D bounding box information in 3D space, the object is associated with a 3D bounding box, a tile, or a 3D spatial region.

[0518] According to the implementation, when a V-PCC bitstream is encapsulated in a file format, the 3D bounding box of the point cloud data can be segmented into one or more 3D regions, and each segmented 3D region can include one or more objects. When patches are packed into Atlas frames, the patches collect and pack (mapped to) one or more Atlas tile regions in the Atlas frame on an object-by-object basis. That is, a 3D region (i.e., file level) can be associated with one or more objects (i.e., bitstream level), and an object can be associated with one or more 3D regions. Since each object is associated with one or more Atlas tiles, a 3D region can be associated with one or more Atlas tiles. For example, object #1 can correspond to Atlas tile #1 and Atlas tile #2, object #2 can correspond to Atlas tile #3, and object #3 can correspond to Atlas tile #4 and Atlas tile #5.

[0519] According to the implementation, the 3D regions can overlap each other. As an example, 3D (spatial) region #1 may include object #1, and 3D region #2 may include object #2 and object #3. However, in another example, 3D region #1 may include object #1 and object #2, and 3D region #2 may include object #2 and object #3. In other words, object #2 can be associated with both 3D region #1 and 3D region #2. Furthermore, since the same object (e.g., object #2) can be included in different 3D regions (e.g., 3D region #1 and 3D region #2), a patch corresponding to object #2 can be assigned to (included in) different 3D regions (3D region #1 and 3D region #2).

[0520] For partial access to point cloud data, it is necessary to access portions of the point cloud data based on 3D (spatial) regions or objects. To this end, this specification signals the association between 3D regions and tiles, or between objects and tiles. The signaling for these associations will be described in detail below.

[0521] According to the implementation method, Atlas data is signaling information including the Atlas Sequence Parameter Set (ASPS), the Atlas Frame Parameter Set (AFPS), Atlas Patch Group Information (or Atlas Patch Information), and SEI messages, and may be referred to as metadata about Atlas.

[0522] According to the implementation, Atlas represents a set of 2D bounding boxes and can be a patch projected onto a rectangular frame.

[0523] According to the implementation, an Atlas frame is a 2D rectangular array of Atlas samples on which patches are projected. An Atlas sample is the position of a rectangular frame on which a patch associated with an Atlas is projected.

[0524] According to the implementation, an Atlas frame can be divided into one or more rectangular tiles. That is, a tile is a unit used to segment a 2D frame. In other words, a tile is a unit used to segment signaling information (called an Atlas) from point cloud data. According to the implementation, tiles do not overlap within an Atlas frame, and an Atlas frame can include regions unrelated to tiles. Furthermore, the height and width of each tile included in an Atlas can be different for each tile.

[0525] According to the implementation method, a 3D region can correspond to an Atlas frame, so a 3D region can be associated with multiple 2D regions.

[0526] According to the implementation, a patch is a set of points that constitute a point cloud. Points belonging to the same patch are adjacent to each other in 3D space and are mapped in the same direction in the six bounding box planes during the mapping to a 2D image. The patch is signaling information about the construction of the point cloud data.

[0527] The receiving device according to the embodiment can recover attribute video data, geometric video data, and occupancy video data based on Atlas (patch or patch), which are actual video data with the same presentation time.

[0528] Furthermore, when the user zooms in or changes the viewport, a portion of the point cloud object / data, excluding the overall point cloud object / data, can be rendered or displayed on the user's viewport. In this case, for the PCC decoder / player, it is efficient to decode or process the video or atlas data associated with the portion of the point cloud data rendered or displayed on the user's viewport, rather than decoding or processing the video or atlas data associated with the point cloud data of the unrendered or undisplayed portions / areas.

[0529] Therefore, partial access to point cloud data needs to be supported.

[0530] In this scenario, partial point cloud data corresponding to a specific 3D spatial region of the overall point cloud data can be associated with one or more 2D regions. According to the implementation, a 2D region refers to one or more video frames or Atlas frames that include data associated with the point cloud data within the corresponding 3D region.

[0531] Figure 24 An exemplary association is shown between a portion of a 3D region of point cloud data according to an embodiment and one or more 2D regions in a video frame.

[0532] like Figure 24 As shown, data associated with a portion of the 3D region of the point cloud data can be associated with video data of one or more 2D regions in a video frame. That is, data associated with a portion of the 3D region 24020 of the point cloud object 24010 (or bounding box) can be associated with video data of one or more 2D regions 24040, 24050, and 24060 in video frame 24030.

[0533] In the case of dynamic point cloud data (i.e., data on the number of points or the position of points in the point cloud that changes over time), the point cloud data displayed in the same 3D area can change over time.

[0534] Therefore, in order to enable spatial or partial access to point cloud data rendered / displayed on the user's viewport within the PCC decoder / player of the receiving device, the transmitting device can send information about one or more 2D regions in a video frame that are associated with a 3D region of the point cloud that can change over time, either via a V-PCC bitstream or as signaling or metadata in a file. In this specification, the information sent via a V-PCC bitstream or as signaling or metadata in a file is referred to as signaling information.

[0535] The method / apparatus according to the embodiment can encode point cloud data that changes over time (i.e., dynamic point cloud data), transmit and receive dynamic point cloud data, and decode dynamic point cloud data. In this case, point cloud data displayed in the same 3D area can change over time.

[0536] When a user zooms in or changes the viewport, the receiving device and renderer, according to the implementation method, can render or display a portion of the point cloud object / data on the user's viewport.

[0537] For efficient processes, the PCC decoder / player, according to the implementation method, can decode or process video or Atlas data associated with partial point cloud data rendered or displayed on the user's viewport.

[0538] For efficient processes, the PCC decoder / player, according to the implementation method, may not decode or process video or Atlas data associated with point cloud data of unrendered or undisplayed portions / regions.

[0539] In other words, signaling information is needed so that the receiving device can extract a portion of the point cloud data corresponding to a specific 3D spatial region from the total point cloud data in the file, and decode and / or render the extracted portion of the point cloud data. That is, signaling information is needed to support partial access by the receiving device.

[0540] In addition, when the receiving device renders point cloud data based on the user viewport, it needs signaling information to extract partial point cloud data from the overall point cloud data in the file and to decode and / or render the partial point cloud data.

[0541] According to the implementation method, the signaling information required to extract only a portion of the point cloud data from the total point cloud data in the file and decode that portion (i.e., the signaling information for partial access) may include 3D bounding box information, 3D spatial region information, 2D region information, and 3D region mapping information. The signaling information may be stored in samples within a track, sample entries within a track, sample groups within a track, track groups, or individual metadata tracks. Specifically, some signaling information may be stored in sample entries in the form of boxes or entire boxes. The storage and signaling of the 3D bounding box information, 3D spatial region information, 2D region information, and 3D region mapping information included in the signaling information will be described in detail later.

[0542] According to the implementation method, the signaling information used to support partial access can be generated by the metadata generation unit of the transmitting device (e.g., Figure 18 The metadata encoding unit 18005 generates the signaling information, which is then signaled by the file / fragment encapsulation module or unit in samples within a track, sample entries within a track, sample groups within a track, track groups, or individual metadata tracks. Alternatively, the signaling information may be generated by the file / fragment encapsulation module and then signaled in samples within a track, sample entries within a track, sample groups within a track, track groups, or individual metadata tracks. In this specification, the signaling information may include metadata about the point cloud data (e.g., setting values). Depending on the application, the signaling information may also be defined on the system side (e.g., file format, Dynamic Adaptive Streaming over HTTP (DASH), or MPEG Media Transport (MMT)) or on the wired interface side (e.g., High Definition Multimedia Interface (HDMI), DisplayPort, Video Electronics Standards Association (VESA), or CTA).

[0543] The method / apparatus according to the implementation can signal 3D region information or 2D region-related information in a video Atlas frame associated with 3D region information to the V-PCC bitstream (e.g., see...). Figure 25 This is used to perform sending and receiving.

[0544] The method / apparatus according to the embodiments can signal 3D region information or 2D region-related information in a video or Atlas frame associated with 3D region information to a file to perform sending and receiving.

[0545] The method / apparatus according to the embodiments can signal 3D region information associated with an image item or 2D region-related information in a video or Atlas frame associated with 3D region information to a file to perform sending and receiving.

[0546] The method / apparatus according to the embodiments can group tracks including data associated with 3D regions and generate track-related signaling information to perform transmission and reception.

[0547] The method / apparatus according to the embodiments can group tracks including data associated with 2D regions and generate track-related signaling information to perform transmission and reception.

[0548] Figure 25 Examples of V-PCC bitstream structures according to other embodiments of this disclosure are shown. In the embodiments, Figure 25 The V-PCC bitstream is composed of Figure 1 , Figure 4 , Figure 18 , Figure 20 or Figure 21 The generation and output of point cloud video encoder based on V-PCC.

[0549] According to the implementation method, a V-PCC bitstream containing a coded point cloud sequence (CPCS) can be composed of sample stream V-PCC units. The sample stream V-PCC unit carries V-PCC parameter set (also known as VPS) data, an Atlas bitstream, a 2D video-coded occupancy map bitstream, a 2D video-coded geometry bitstream, and zero or more 2D video-coded attribute bitstreams.

[0550] exist Figure 25 In this context, the V-PCC bitstream may include a sample stream V-PCC header 40010 and one or more sample stream V-PCC units 40020. For simplicity, one or more sample stream V-PCC units 40020 may be referred to as the sample stream V-PCC payload. That is, the sample stream V-PCC payload may be referred to as a collection of sample stream V-PCC units.

[0551] The Sample Stream V-PCC header 40010 can specify the precision (in bytes) of the ssvu_vpcc_unit_size element in all Sample Stream V-PCC units.

[0552] Each sample stream V-PCC unit 40021 may include V-PCC unit size information 40030 and V-PCC unit 40040. The V-PCC unit size information 40030 indicates the size of the corresponding V-PCC unit 40040. For simplicity, the V-PCC unit size information 40030 may be referred to as the sample stream V-PCC unit header, and the V-PCC unit 40040 may be referred to as the sample stream V-PCC unit payload.

[0553] Each V-PCC unit 40040 may include a V-PCC unit header 40041 and a V-PCC unit payload 40042.

[0554] In this disclosure, the data contained in the V-PCC unit payload 40042 is distinguished by the V-PCC unit header 40041. For this purpose, the V-PCC unit header 40041 contains type information indicating the type of the V-PCC unit. Based on the type information in the V-PCC unit header 40041, each V-PCC unit payload 40042 may contain geometric video data (i.e., a geometric bitstream encoded in 2D video), attribute video data (i.e., an attribute bitstream encoded in 2D video), occupancy video data (i.e., an occupancy graph bitstream encoded in 2D video), Atlas data, or a V-PCC parameter set (VPS).

[0555] The VPS implemented according to the method is also called a Sequence Parameter Set (SPS). These two terms are used interchangeably.

[0556] Figure 26 An example of data carried by a sample stream V-PCC cell in a V-PCC bitstream is shown according to an embodiment.

[0557] exist Figure 26 In the example, the V-PCC bitstream includes sample stream V-PCC units carrying V-PCC parameter sets (VPS), sample stream V-PCC units carrying Atlas data (AD), sample stream V-PCC units carrying occupied video data (OVD), sample stream V-PCC units carrying geometric video data (GVD), and sample stream V-PCC units carrying attribute video data (AVD).

[0558] That is, each sample stream V-PCC unit contains one type of V-PCC unit among VPS, AD, OVD, GVD, and AVD.

[0559] Fields that are used as terms in the syntax described later in this disclosure may have the same meaning as parameters or syntax elements.

[0560] Figure 27 An example of the syntax structure of the sample stream V-PCC header contained in the V-PCC bitstream according to an implementation is shown.

[0561] According to the implementation, sample_stream_v-pcc_header() may include the ssvh_unit_size_precision_bytes_minus1 field and the ssvh_reserved_zero_5bits field.

[0562] Increasing the value of the ssvh_unit_size_precision_bytes_minus1 field by 1 specifies the precision (in bytes) of the ssvu_vpcc_unit_size element in all sample stream V-PCC units. The value of this field can range from 0 to 7.

[0563] The ssvh_reserved_zero_5bits field is a reserved field for future use.

[0564] Figure 28 An example of the syntax structure of a sample stream V-PCC unit (sample_stream_vpcc_unit()) according to an implementation is shown.

[0565] The contents of each sample stream V-PCC cell are associated with the same access cell as the V-PCC cells contained in the sample stream V-PCC cell.

[0566] According to the implementation, sample_stream_vpcc_unit() may include the ssvu_vpcc_unit_size field and vpcc_unit(ssvu_vpcc_unit_size).

[0567] The ssvu_vpcc_unit_size field corresponds to Figure 25 The V-PCC unit size information is 40030, and the size of the subsequent vpcc_unit is specified (in bytes). The number of bits used to represent the ssvu_vpcc_unit_size field is equal to (ssvh_unit_size_precision_bytes_minus1+1)*8.

[0568] vpcc_unit(ssvu_vpcc_unit_size) has the length corresponding to the value of the ssvu_vpcc_unit_size field and carries one of VPS, AD, OVD, GVD and AVD.

[0569] Figure 29 An example of the syntax structure of a V-PCC unit according to an embodiment is shown. A V-PCC unit consists of a V-PCC unit header (vpcc_unit_header()) and a V-PCC unit payload (vpcc_unit_payload()). A V-PCC unit according to an embodiment may contain more data. In this case, it may also include a trailing_zero_8bits field. The trailing_zero_8bits field according to an embodiment is a byte corresponding to 0x00.

[0570] Figure 30 An example of the syntax structure of the V-PCC unit header according to an implementation is shown. In the implementation, Figure 30 The `vpcc_unit_header()` function includes a `vuh_unit_type` field. The `vuh_unit_type` field indicates the type of the corresponding V-PCC unit. Depending on the implementation, the `vuh_unit_type` field is also referred to as the `vpcc_unit_type` field.

[0571] Figure 31 An example of a V-PCC unit type assigned to the vuh_unit_type field according to an implementation is shown.

[0572] Reference Figure 31 According to the implementation, a vuh_unit_type field set to 0 indicates that the data included in the V-PCC unit payload of the V-PCC unit is a V-PCC parameter set (VPCC_VPS). A vuh_unit_type field set to 1 indicates that the data is Atlas data (VPCC_AD). A vuh_unit_type field set to 2 indicates that the data is occupancy video data (VPCC_OVD). A vuh_unit_type field set to 3 indicates that the data is geometric video data (VPCC_GVD). A vuh_unit_type field set to 4 indicates that the data is attribute video data (VPCC_AVD).

[0573] Those skilled in the art can easily change the meaning, order, deletion, addition, etc. of the values ​​assigned to the vuh_unit_type field, therefore this disclosure is not limited to the above-described embodiments.

[0574] When the vuh_unit_type field indicates VPCC_AVD, VPCC_GVD, VPCC_OVD, or VPCC_AD, the V-PCC unit header according to the implementation may also include the vuh_vpcc_parameter_set_id field and the vuh_atlas_id field.

[0575] The vuh_vpcc_parameter_set_id field specifies the value of vps_vpcc_parameter_set_id for the active V-PCCVPS.

[0576] The vuh_atlas_id field specifies the index (or identifier) ​​of the atlas in the current V-PCC cell.

[0577] When the vuh_unit_type field indicates VPCC_AVD, the V-PCC unit header according to the implementation may also include the vuh_attribute_index field, the vuh_attribute_dimension_index field, the vuh_map_index field, and the vuh_raw_video_flag field.

[0578] The vuh_attribute_index field indicates the index of the attribute data loaded in the attribute video data unit.

[0579] The `vuh_attribute_dimension_index` field indicates the index of the attribute dimension group loaded in the attribute video data unit.

[0580] When present, the vuh_map_index field can indicate the graph index of the current geometry or attribute flow.

[0581] The `vuh_raw_video_flag` field indicates whether RAW encoded points are included. For example, a `vuh_raw_video_flag` field set to 1 indicates that the associated attribute video data unit contains only RAW encoded points. As another example, a `vuh_raw_video_flag` field set to 0 indicates that the associated attribute video data unit may contain RAW encoded points. When the `vuh_raw_video_flag` field is not present, its value can be inferred to be equal to 0. According to implementation, RAW encoded points are also referred to as Pulse Code Modulation (PCM) encoded points.

[0582] When the vuh_unit_type field indicates VPCC_GVD, the V-PCC unit header according to the implementation may also include the vuh_map_index field, the vuh_raw_video_flag field, and the vuh_reserved_zero_12bits field.

[0583] When present, the vuh_map_index field indicates the index of the current geometry flow.

[0584] The `vuh_raw_video_flag` field indicates whether RAW encoded points are included. For example, a `vuh_raw_video_flag` field set to 1 indicates that the associated geometric video data unit contains only RAW encoded points. As another example, a `vuh_raw_video_flag` field set to 0 indicates that the associated geometric video data unit may contain RAW encoded points. When the `vuh_raw_video_flag` field is not present, its value can be inferred to be equal to 0. According to the implementation, RAW encoded points are also referred to as PCM encoded points.

[0585] The vuh_reserved_zero_12bits field is a reserved field for future use.

[0586] If the vuh_unit_type field indicates VPCC_OVD or VPCC_AD, the V-PCC unit header, depending on the implementation, may also include a vuh_reserved_zero_17bits field. Otherwise, the V-PCC unit header may also include a vuh_reserved_zero_27bits field.

[0587] The vuh_reserved_zero_17bits and vuh_reserved_zero_27bits fields are reserved for future use.

[0588] Figure 32 An exemplary syntax structure for a V-PCC unit payload (vpcc_unit_payload()) according to an implementation is shown.

[0589] In other words, a V-PCC bitstream is a collection of V-PCC components (e.g., atlases, occupancy maps, geometry, and attributes). An atlas component (or atlas frame) can be divided into one or more pieces (or groups of pieces) and can be encapsulated in a NAL unit. In one implementation, a collection of one or more pieces in an atlas frame can be referred to as a piece group. In another implementation, a collection of one or more pieces in an atlas frame can also be referred to as a piece. For example, an atlas frame can be divided into one or more rectangular partitions, and one or more rectangular partitions can be referred to as pieces.

[0590] According to the implementation, the encoded V-PCC video components are referred to as video bitstreams and the Atlas components as Atlas bitstreams. The video bitstream can be subdivided into smaller units, such as video sub-bitstreams, and the Atlas bitstream can also be subdivided into smaller units, such as Atlas sub-bitstreams.

[0591] Figure 32 The payload of a V-PCC unit can contain one of the following, depending on the value of the vuh_unit_type field in the V-PCC unit header: the V-PCC parameter set (vpcc_parameter_set()), the atlas sub-bitstream (atlas_sub_bitstream()), or the video sub-bitstream (video_sub_bitstream()).

[0592] For example, when the `vuh_unit_type` field indicates `VPCC_VPS`, the V-PCC unit payload includes `vpcc_parameter_set()`, which contains overall encoding information about the bitstream. When the `vuh_unit_type` field indicates `VPCC_AD`, the V-PCC unit payload includes `atlas_sub_bitstream()` carrying Atlas data. Furthermore, according to an implementation, when the `vuh_unit_type` field indicates `VPCC_OVD`, the V-PCC unit payload includes an occupied video sub-bitstream (`video_sub_bitstream()`) carrying occupied video data. When the `vuh_unit_type` field indicates `VPCC_GVD`, the V-PCC unit payload includes a geometric video sub-bitstream (`video_sub_bitstream()`) carrying geometric video data. When the `vuh_unit_type` field indicates `VPCC_AVD`, the V-PCC unit payload includes an attribute video sub-bitstream (`video_sub_bitstream()`) carrying attribute video data.

[0593] According to the implementation, the Atlas sub-bitstream can be referred to as the Atlas sub-stream, and the occupied video sub-bitstream can be referred to as the occupied video sub-stream. The geometric video sub-bitstream can be referred to as the geometric video sub-stream, and the attribute video sub-bitstream can be referred to as the attribute video sub-stream. The V-PCC unit payload according to the implementation conforms to the format of the High-Efficiency Video Coding (HEVC) Network Abstraction Layer (NAL) unit.

[0594] Figure 33 An exemplary syntax structure for the V-PCC parameter set (VPS) according to an implementation is shown.

[0595] Depending on the implementation, a VPS may include the profile_tier_level(), vps_vpcc_parameter_set_id, and sps_bounding_box_present_flag fields.

[0596] `profile_tier_level()` specifies limitations on the bitstream. For example, `profile_tier_level()` specifies limitations on the capabilities required to decode the bitstream. Profiles, tiers, and levels can also be used to indicate interoperability points between different decoder implementations.

[0597] The vps_vpcc_parameter_set_id field can provide an identifier for V-PCC VPS for reference by other syntax elements.

[0598] The `sps_bounding_box_present_flag` field represents a flag that indicates whether information about the overall (whole) bounding box of the point cloud object / content exists in the bitstream (the overall bounding box can be a bounding box that includes all bounding boxes that change over time). For example, an `sps_bounding_box_present_flag` field equal to 1 can indicate the overall bounding box offset and size information of the point cloud content carried in the bitstream.

[0599] According to the implementation method, when the sps_bounding_box_present_flag field is equal to 1, the VPS may also include the sps_bounding_box_offset_x field, the sps_bounding_box_offset_y field, the sps_bounding_box_offset_z field, the sps_bounding_box_size_width field, the sps_bounding_box_size_height field, the sps_bounding_box_size_depth field, the sps_bounding_box_changed_flag field, and the sps_bounding_box_info_flag field.

[0600] The `sps_bounding_box_offset_x` field indicates the x-offset of the overall bounding box, serving as information about the size of the point cloud content carried in the bitstream in Cartesian coordinates. When it is absent, the value of `sps_bounding_box_offset_x` can be inferred to be 0.

[0601] The `sps_bounding_box_offset_y` field indicates the y-offset of the overall bounding box, serving as information about the size of the point cloud content carried in the bitstream in Cartesian coordinates. When it does not exist, the value of `sps_bounding_box_offset_y` can be inferred to be 0.

[0602] The `sps_bounding_box_offset_z` field indicates the z-offset of the overall bounding box offset, serving as information about the size of the point cloud content carried in the bitstream in Cartesian coordinates. When it does not exist, the value of `sps_bounding_box_offset_z` can be inferred to be 0.

[0603] The `sps_bounding_box_size_width` field indicates the width of the overall bounding box offset, serving as information about the size of the point cloud content carried in the bitstream in Cartesian coordinates. When it does not exist, the value of `sps_bounding_box_size_width` can be inferred to be 1.

[0604] The `sps_bounding_box_size_height` field indicates the height of the overall bounding box offset, serving as information about the size of the point cloud content carried in the bitstream in Cartesian coordinates. When it does not exist, the value of `sps_bounding_box_size_height` can be inferred to be 1.

[0605] The `sps_bounding_box_size_depth` field indicates the depth of the overall bounding box offset, serving as information about the size of the point cloud content carried in the bitstream in Cartesian coordinates. When it does not exist, the value of `sps_bounding_box_size_depth` can be inferred to be 1.

[0606] The `sps_bounding_box_changed_flag` field can indicate whether the bounding boxes of point cloud data included in the bitstream have changed over time. For example, a `sps_bounding_box_changed_flag` field equal to 1 indicates that the bounding boxes of the point cloud data have changed over time.

[0607] The `sps_bounding_box_info_flag` field can indicate whether a bounding box information (SEI) including point cloud data exists in the bitstream. For example, a `sps_bounding_box_info_flag` field equal to 1 can indicate that an SEI including bounding box information (3D bounding box SEI) is included in the bitstream. In this case, this can instruct the PCC player corresponding to the method / apparatus according to the implementation to acquire and use the information included in the corresponding SEI.

[0608] The VPS, depending on the implementation, may also include a vps_atlas_count_minus1 field. Increasing the vps_atlas_count_minus1 field by 1 indicates the total number of Atlases supported in the current bitstream.

[0609] According to the implementation, the VPS may include a first iteration statement that repeats the value of the vps_atlas_count_minus1 field as many times. The first iteration statement may include the vps_frame_width[j] field, the vps_frame_height[j] field, the vps_map_count_minus1[j] field, and the vps_raw_patch_enabled_flag[j] field. In the implementation, index j is initialized to 0 and incremented by 1 each time the first iteration statement is executed, and the first iteration statement is repeated until the value of j becomes the value of the vps_atlas_count_minus1 field.

[0610] The `vps_frame_width[j]` field indicates the width of the V-PCC frame in terms of integer luminance samples for the Atlas with index `j`. This frame width is the nominal width associated with all V-PCC components of the Atlas with index `j`.

[0611] The `vps_frame_height[j]` field indicates the height of a V-PCC frame in terms of integer luminance samples for an Atlas with index `j`. This frame height is the nominal height associated with all V-PCC components of the Atlas with index `j`.

[0612] Increasing the value of the vps_map_count_minus1[j] field by 1 indicates the number of graphs used to encode the geometry and attribute data of the Atlas at index j.

[0613] According to the implementation method, when the value of the vps_map_count_minus1[j] field is greater than 0, the first iteration statement may also include the vps_multiple_map_streams_present_flag[j] field and vps_map_absolute_coding_enabled_flag[j][0] = 1.

[0614] The `vps_multiple_map_streams_present_flag[j]` field being equal to 0 indicates that all geometry or attribute maps of the Atlas with index j are placed in a single geometry or attribute video stream. The `vps_multiple_map_streams_present_flag[j]` field being equal to 1 indicates that all geometry or attribute maps of the Atlas with index j are placed in a separate video stream.

[0615] According to the implementation, the first iteration statement includes a second iteration statement that repeats the same number of times as the value of the vps_map_count_minus1[j] field. In the implementation, the index i is initialized to 0 and incremented by 1 each time the second iteration statement is executed, and the second iteration statement is repeated until the value of i becomes the value of the vps_map_count_minus1 field.

[0616] The second iteration statement may also include the vps_map_absolute_coding_enabled_flag[j][i] field and / or the vps_map_predictor_index_diff[j][i] field, depending on the value of the vps_multiple_map_streams_present_flag[j].

[0617] According to the implementation method, if the value of the vps_multiple_map_streams_present_flag[j] field is 1, the second iteration statement may also include the vps_map_absolute_coding_enabled_flag[j][i] field; otherwise, the vps_map_absolute_coding_enabled_flag[j][i] field may be equal to 1.

[0618] The `vps_map_absolute_coding_enabled_flag[j][i]` field being equal to 1 indicates that the geometry at index i of the Atlas at index j is encoded without any form of graph prediction. `vps_map_absolute_coding_enabled_flag[j][i]` being equal to 0 indicates that the geometry at index i of the Atlas at index j was first predicted from an earlier encoded graph before encoding.

[0619] The field `vps_map_absolute_coding_enabled_flag[j][0]` equals 1, indicating that the geometry at index 0 is encoded without graph prediction.

[0620] If the value of the vps_map_absolute_coding_enabled_flag[j][i] field is 0 and i is greater than 0, the second iteration statement can also include the vps_map_predictor_index_diff[j][i] field; otherwise, the value of vps_map_predictor_index_diff[j] field [i] can be changed to 0.

[0621] When vps_map_absolute_coding_enabled_flag[j][i] equals 0, the value of the vps_map_predictor_index_diff[j][i] field can be used to calculate the prediction of the geometry at index i for the atlas at index j.

[0622] According to the implementation method, the second iteration statement may also include the vps_raw_patch_enabled_flag[j] field.

[0623] The vps_raw_patch_enabled_flag[j] field being equal to 1 indicates that a patch with the raw code point of the Atlas at index j can exist in the bitstream.

[0624] According to the implementation, if the value of the vps_raw_patch_enabled_flag[j] field is true, the second iteration statement may also include the vps_raw_separate_video_present_flag[j] field, occupancy_information(j), geometry_information(j), and attribute_information(j).

[0625] The field vps_raw_separate_video_present_flag[j] equal to 1 indicates that the geometry and attribute information of the raw encoding of the Atlas with index j can be stored in a separate video stream.

[0626] occupancy_information(j) includes a set of parameters related to the occupancy video of the Atlas with index j.

[0627] geometry_information(j) includes the set of parameters associated with the geometry video of the Atlas at index j.

[0628] The attribute_information(j) includes the set of parameters associated with the attribute video of the Atlas at index j.

[0629] Figure 34 An exemplary structure of an Atlas substream according to an embodiment is shown. In the embodiment, Figure 34 The Atlas substream conforms to the format of HEVC NAL cells.

[0630] According to the implementation, the Atlas substream 41000 may include a sample stream NAL header 41010 and one or more sample stream NAL units 41020.

[0631] One or more sample stream NAL units 41020 according to the implementation may include: a sample stream NAL unit 41030 including ASPS, a sample stream NAL unit 41040 including AFPS, one or more sample stream NAL units 41050 including information about one or more Atlas tiles (or tile groups), and / or one or more sample stream NAL units 41060 including one or more SEI messages.

[0632] Figure 35 An example of the syntax structure of the sample stream NAL header (sample_stream_nal_header()) contained in an Atlas substream according to an implementation is shown.

[0633] According to the implementation, sample_stream_nal_header() may include the ssnh_unit_size_precision_bytes_minus1 field and the ssnh_reserved_zero_5bits field.

[0634] Increasing the value of the ssnh_unit_size_precision_bytes_minus1 field by 1 specifies the precision (in bytes) of the ssnu_nal_unit_size elements in all sample stream NAL units. The value of this field can range from 0 to 7.

[0635] The ssnh_reserved_zero_5bits field is a reserved field for future use.

[0636] Figure 36 An example of the syntax structure of the sample stream NAL unit (sample_stream_nal_unit()) according to an implementation is shown.

[0637] The sample_stream_nal_unit() method, depending on the implementation, may include the ssnu_nal_unit_size field and nal_unit(ssnu_nal_unit_size).

[0638] The `ssnu_nal_unit_size` field specifies the size of the subsequent `NAL_unit` (in bytes). The number of bits used to represent the `ssnu_nal_unit_size` field is equal to (`ssnh_unit_size_precision_bytes_minus1+1`)*8.

[0639] The `nal_unit(ssnu_nal_unit_size)` has a length corresponding to the value of the `ssnu_nal_unit_size` field and carries one of the Atlas Sequence Parameter Set (ASPS), Atlas Frame Parameter Set (AFPS), Atlas Patch Group Information, and SEI message. That is, each sample stream NAL unit can contain ASPS, AFPS, Atlas Patch Group Information, or SEI message. According to the implementation, ASPS, AFPS, Atlas Patch Group Information, and SEI message are referred to as Atlas data (or Atlas metadata).

[0640] According to the implementation method, the SEI message can assist in processing related to decoding, reconstruction, display, or other purposes.

[0641] Each SEI message according to the implementation method consists of an SEI message header and an SEI message payload (sei_payload). The SEI message header may contain payload type information (payloadType) and payload size information (payloadSize).

[0642] The payload type indicates the type of payload in the SEI message. For example, the payloadType can be used to identify whether an SEI message is a prefix SEI message or a suffix SEI message.

[0643] The payload size indicates the payload size of the SEI message.

[0644] Figure 37 An exemplary syntax structure for `atlas_sequence_parameter_set()` according to an implementation is shown. When the type of the NAL unit is an Atlas sequence parameter, Figure 37 ASPS can be included in a sample stream NAL cell conforming to the HEVC NAL cell format.

[0645] ASPS may include syntax elements applied to zero or one or more fully encoded Atlas sequences (CAS) determined by the content of syntax elements in ASPS, which are referenced in the syntax elements of each piece group (or piece) header.

[0646] According to the implementation, ASPS may include the following fields: asps_atlas_sequence_parameter_set_id, asps_frame_width, asps_frame_height, asps_log2_patch_packing_block_size, asps_log2_max_atlas_frame_order_cnt_lsb_minus4, asps_max_dec_atlas_frame_buffering_minus1, asps_long_term_ref_atlas_frames_flag, asps_num_ref_atlas_frame_lists_in_asps, asps_use_eight_orientations_flag, and asps_45degree_projection_patch_present_flag. , asps_normal_axis_limits_quantization_enabled_flag field, asps_normal_axis_max_delta_value_enabled_flag field, asps_remove_duplicate_point_enabled_flag field, asps_pixel_deinterleaving_flag field, asps_patch_prec edence_order_flag field, asps_patch_size_quantizer_present_flag field, asps_enhanced_occupancy_map_for_depth_flag field, asps_point_local_reconstruction_enabled_flag field, and asps_vui_parameters_present_flagfield.

[0647] The asps_atlas_sequence_parameter_set_id field can provide an identifier for the Atlas sequence parameter set for reference by other syntax elements.

[0648] The asps_frame_width field indicates the Atlas frame width in terms of integer luminance samples for the current Atlas.

[0649] The asps_frame_height field indicates the Atlas frame height in terms of integer luminance samples for the current Atlas.

[0650] The asps_log2_patch_packing_block_size field specifies the value of the variable PatchPackingBlockSize, which is used for the horizontal and vertical placement of patches within the Atlas.

[0651] The asps_log2_max_atlas_frame_order_cnt_lsb_minus4 field specifies the value of the variable MaxAtlasFrmOrderCntLsb, which is used in the decoding process of Atlas frame order counting.

[0652] The asps_max_dec_atlas_frame_buffering_minus1 field plus 1 specifies the maximum required size of the Atlas frame buffer for CAS decoding, in units of Atlas frame storage buffers.

[0653] A value of 0 for the `asps_long_term_ref_atlas_frames_flag` field indicates that no long-term reference atlas frames are used for inter-frame prediction of any encoded atlas frames in CAS. A value of 1 for the `asps_long_term_ref_atlas_frames_flag` field indicates that long-term reference atlas frames can be used for inter-frame prediction of one or more encoded atlas frames in CAS.

[0654] The asps_num_ref_atlas_frame_lists_in_asps field specifies the number of ref_list_struct(rlsIdx) syntax structures included in the Atlas sequence parameter set.

[0655] The ref_list_struct(i) can be included in the Atlas sequence parameter set based on the value of the asps_num_ref_atlas_frame_lists_in_asps field.

[0656] The `asps_use_eight_orientations_flag` field being equal to 0 specifies that the patch orientation index `pdu_orientation_index[i][j]` of the patch with index `j` in the frame with index `i` is in the range of 0 to 1 (inclusive). The `asps_use_eight_orientations_flag` field being equal to 1 specifies that the patch orientation index `pdu_orientation_index[i][j]` of the patch with index `j` in the frame with index `i` is in the range of 0 to 7 (inclusive).

[0657] The `asps_45degree_projection_patch_present_flag` field being equal to 0 indicates that no signaling of tile projection information is sent for the current Atlas tile (or Atlas tile group). The `asps_45degree_projection_present_flag` field being equal to 1 indicates that no signaling of tile projection information is sent for the current Atlas tile (or Atlas tile group).

[0658] The `asps_normal_axis_limits_quantization_enabled_flag` field being equal to 1 indicates that quantization parameters should be signaled and used to quantize normal-axis related elements of patch data units, merged patch data units, or inter-patch data units. If the `asps_normal_axis_limits_quantization_enabled_flag` field is equal to 0, quantization is not applied to any normal-axis related elements of patch data units, merged patch data units, or inter-patch data units. When `asps_normal_axis_limits_quantization_enabled_flag` is 1, the `atgh_pos_min_z_quantizer` field can be included in the Atlas tile group (or tile) header.

[0659] The `asps_normal_axis_max_delta_value_enabled_flag` field being equal to 1 specifies that the maximum nominal shift value of the normal axis that can exist in the geometry of the patch at index i in a frame with index j will be indicated in the bitstream of each patch data unit, merged patch data unit, or inter-patch data unit. If the `asps_normal_axis_max_delta_value_enabled_flag` field is equal to 0, the maximum nominal shift value of the normal axis that can exist in the geometry of the patch at index i in a frame with index j should not be indicated in the bitstream of each patch data unit, merged patch data unit, or inter-patch data unit. When the `asps_normal_axis_max_delta_value_enabled_flag` field is equal to 1, the `atgh_pos_delta_max_z_quantizer` field can be included in the Atlas patch group (or patch) header.

[0660] A value of 1 for the `asps_remove_duplicate_point_enabled_flag` field indicates that duplicate points are not reconstructed for the current Atlas, where a duplicate point is a point with the same 2D and 3D geometric coordinates as another point from a lower index graph. A value of 0 for the `asps_remove_duplicate_point_enabled_flag` field indicates that all points are reconstructed.

[0661] The asps_max_dec_atlas_frame_buffering_minus1 field plus 1 specifies the maximum required size of the Atlas frame buffer for CAS decoding, in units of Atlas frame storage buffers.

[0662] The `asps_pixel_deinterleaving_flag` field being equal to 1 indicates that the decoded geometry and attribute video for the current Atlas contains spatially interleaved pixels from two graphs. The `asps_pixel_deinterleaving_flag` field being equal to 0 indicates that the decoded geometry and attribute video corresponding to the current Atlas contains pixels from only a single graph.

[0663] A value of 1 in the `asps_patch_precedence_order_flag` field indicates that the patch priority of the current Atlas is the same as the decoding order. A value of 0 in the `asps_patch_precedence_order_flag` field indicates that the patch priority of the current Atlas is the opposite of the decoding order.

[0664] A value of 1 for the `asps_patch_size_quantizer_present_flag` field indicates that the patch size quantization parameter exists in the Atlas patch (or patch group) header. A value of 0 for the `asps_patch_size_quantizer_present_flag` field indicates that the patch size quantization parameter does not exist. When the `asps_patch_size_quantizer_present_flag` field is equal to 1, the `atgh_patch_size_x_info_quantizer` and `atgh_patch_size_y_info_quantizer` fields can be included in the Atlas patch group (or patch) header.

[0665] A value of 1 for the `asps_enhanced_occupancy_map_for_depth_flag` field indicates that the current Atlas decoded occupancy map video contains information related to whether intermediate depth positions between two depth maps are occupied. A value of 0 for the `asps_eom_patch_enabled_flag` field indicates that the decoded occupancy map video does not contain information related to whether intermediate depth positions between two depth maps are occupied.

[0666] A value of 1 for the `asps_point_local_reconstruction_enabled_flag` field indicates that point local reconstruction mode information can exist in the current Atlas bitstream. A value of 0 for the `asps_point_local_reconstruction_enabled_flag` field indicates that no information related to point local reconstruction mode exists in the current Atlas bitstream.

[0667] The asps_map_count_minus1 field can be included in ASPS when the asps_enhanced_occupancy_map_for_depth_flag field or the asps_point_local_reconstruction_enabled_flag field is equal to 1.

[0668] The asps_map_count_minus1 field incremented by 1 indicates the number of maps that can be used to encode the geometry and attribute data of the current Atlas.

[0669] The asps_enhanced_occupancy_map_for_depth_flag field can be included in ASPS when the asps_enhanced_occupancy_map_fix_bit_count_minus1 field is set to 0.

[0670] The asps_enhanced_occupancy_map_fix_bit_count_minus1 field incremented by 1 indicates the bit size of the EOM codeword.

[0671] When the asps_point_local_reconstruction_enabled_flag field is equal to 1, ASPS can include ASPS point local reconstruction information (asps_point_local_reconstruction_information(asps_map_count_minus1)).

[0672] The asps_surface_thickness_minus1 field can be included in ASPS when the asps_pixel_deinterleaving_flag field (or the asps_pixel_interleaving_flag field) or the asps_point_local_reconstruction_enabled_flag field is equal to 1.

[0673] The asps_surface_thickness_minus1 field plus 1 specifies the maximum absolute difference between the explicitly encoded depth value and the interpolated depth value.

[0674] The `asps_vui_parameters_present_flag` field being equal to 1 indicates that the `vui_parameters()` syntax structure exists in ASPS. The `asps_vui_parameters_present_flag` field being equal to 0 indicates that the `vui_parameters()` syntax structure does not exist in ASPS. In other words, if the value of the `asps_vui_parameters_present_flag` field is 1, then the `vui_parameters()` syntax is included in ASPS.

[0675] Figure 38 An exemplary syntax structure for the Atlas Frame Parameter Set (AFPS) according to an implementation is shown. When the type of the NAL unit is an Atlas Frame Parameter, Figure 38 AFPS can be included in the sample stream NAL cell conforming to the HEVC NAL cell format.

[0676] AFPS includes a syntax structure that includes syntax elements applied to zero or one or more fully encoded Atlas frames.

[0677] According to the implementation, AFPS may include the following fields: afps_atlas_frame_parameter_set_id, afps_atlas_sequence_parameter_set_id, atlas_frame_tile_information(), afps_num_ref_idx_default_active_minus1, afps_additional_lt_afoc_lsb_len, afps_2d_pos_x_bit_count_minus1, afps_2d_pos_y_bit_count_minus1, afps_3d_pos_x_bit_count_minus1, afps_3d_pos_y_bit_count_minus1, afps_lod_bit_count, afps_override_eom_for_depth_flag, and afps_raw_3d_pos_bit_count_explicit_mode_flag.

[0678] The afps_atlas_frame_parameter_set_id field specifies an identifier used to identify AFPS for reference by other syntax elements.

[0679] The afps_atlas_sequence_parameter_set_id field specifies the value of asps_atlas_sequence_parameter_set_id for the active ASPS.

[0680] Reference Figure 39 Provide a detailed description of atlas_frame_tile_information().

[0681] Increasing the afps_num_ref_idx_default_active_minus1 field by 1 specifies the inferred value of the variable NumRefIdxActive for a tile (or tile group) where the atgh_num_ref_idx_active_override_flag field equals 0.

[0682] The afps_additional_lt_afoc_lsb_len field specifies the value of the variable MaxLtAtlasFrmOrderCntLsb used in the decoding process to reference the Atlas frame list.

[0683] The afps_2d_pos_x_bit_count_minus1 field plus 1 specifies the number of bits in the fixed-length representation of pdu_2d_pos_x[j] of the patch with index j in the Atlas patch (or patch group) that references afps_atlas_frame_parameter_set_id.

[0684] The afps_2d_pos_y_bit_count_minus1 field plus 1 specifies the number of bits in the fixed-length representation of pdu_2d_pos_y[j] of the patch with index j in the Atlas patch (or patch group) that references afps_atlas_frame_parameter_set_id.

[0685] The afps_3d_pos_x_bit_count_minus1 field plus 1 specifies the number of bits in the fixed-length representation of pdu_3d_pos_x[j] of the patch with index j in the Atlas patch (or patch group) that references afps_atlas_frame_parameter_set_id.

[0686] The afps_3d_pos_y_bit_count_minus1 field plus 1 specifies the number of bits in the fixed-length representation of pdu_3d_pos_y[j] of the patch with index j in the Atlas patch (or patch group) that references afps_atlas_frame_parameter_set_id.

[0687] The afps_lod_bit_count field specifies the number of bits in the fixed-length representation of pdu_lod[j] of the patch with index j in the Atlas patch (or patch group) that references the afps_atlas_frame_parameter_set_id field.

[0688] A value of 1 for the `afps_override_eom_for_depth_flag` field indicates that the values ​​of the `afps_eom_number_of_patch_bit_count_minus1` and `afps_eom_max_bit_count_minus1` fields are explicitly present in the bitstream. A value of 0 for the `afps_override_eom_for_depth_flag` field indicates that the values ​​of the `afps_eom_number_of_patch_bit_count_minus1` and `afps_eom_max_bit_count_minus1` fields are implicitly inferred.

[0689] The afps_eom_number_of_patch_bit_count_minus1 field plus 1 specifies the number of bits used to represent the number of geometric patches associated with the current EOM attribute patch.

[0690] The afps_eom_max_bit_count_minus1 field plus 1 specifies the number of bits used to represent the number of EOM points for each geometric patch associated with the current EOM attribute patch.

[0691] The afps_raw_3d_pos_bit_count_explicit_mode_flag field being equal to 1 indicates that the bit counts of the rpdu_3d_pos_x, rpdu_3d_pos_y, and rpdu_3d_pos_z fields are explicitly encoded in the piece group header that references the afps_atlas_frame_parameter_set_id field.

[0692] Figure 39 An exemplary syntax structure for Atlas frame tile information (atlas_frame_tile_information) according to an implementation is shown.

[0693] Figure 39 Show Figure 38 The implementation of the syntax for Atlas Frame Tiling Information (AFTI) included in AFPS.

[0694] According to the implementation, AFTI may include the fields afti_single_tile_in_atlas_frame_flag, afti_num_tiles_in_atlas_frame_minus1, afti_num_tile_groups_in_atlas_frame_minus1, or afti_signalled_tile_group_id_flag.

[0695] The `afti_single_tile_in_atlas_frame_flag` field equal to 1 indicates that there is only one tile in each Atlas frame referencing AFPS. The `afti_single_tile_in_atlas_frame_flag` field equal to 0 indicates that there are multiple tiles in each Atlas frame referencing AFPS.

[0696] If the value of the afti_single_tile_in_atlas_frame_flag field is false (e.g., 0), then the afti_uniform_tile_spacing_flag field can be included in AFTI.

[0697] The `afti_uniform_tile_spacing_flag` field being equal to 1 specifies that the tile column and row boundaries are evenly distributed across the Atlas frame, and is signaled using syntax elements such as `afti_tile_cols_width_minus1` and `afti_tile_rows_height_minus1`, respectively. The `afti_uniform_tile_spacing_flag` field being equal to 0 specifies that the tile column and row boundaries are either evenly distributed or not distributed across the entire Atlas frame, and is signaled using syntax elements such as `afti_num_tile_columns_minus1`, `afti_num_tile_rows_minus1`, and a list of syntax element pairs `afti_tile_column_width_minus1` and `afti_tile_row_height_minus1`.

[0698] If the value of the afti_uniform_tile_spacing_flag field is true (e.g., 1), then the afti_tile_cols_width_minus1 and afti_tile_rows_height_minus1 fields can be included in AFTI.

[0699] The value of the afti_tile_cols_width_minus1 field plus 1 specifies the width of the tile column, excluding the rightmost tile column, in an Atlas frame of 64 samples.

[0700] The value of the afti_tile_rows_height_minus1 field plus 1 specifies the height of the tile rows, excluding the bottom tile row, in an Atlas frame of 64 samples.

[0701] If the value of the afti_uniform_tile_spacing_flag field is false (e.g., 0), then the afti_num_tile_columns_minus1 and afti_num_tile_rows_minus1 fields can be included in AFTI.

[0702] The value of the afti_num_tile_columns_minus1 field plus 1 specifies the number of tile columns used to divide the Atlas frame.

[0703] The value of the afti_num_tile_rows_minus1 field plus 1 specifies the number of tile rows used to divide the Atlas frame.

[0704] AFTI may include an afti_tile_column_width_minus1[i] field as many as the afti_num_tile_columns_minus1 field, and may include an afti_tile_row_height_minus1[i] field as many as the afti_num_tile_rows_minus1 field.

[0705] The value of the afti_tile_column_width_minus1[i] field plus 1 specifies the width of the i-th tile column, which is based on 64 samples.

[0706] Increasing the value of the afti_tile_row_height_minus1[i] field by 1 specifies the height of the i-th tile row, which is based on 64 samples.

[0707] The value of the afti_num_tiles_in_atlas_frame_minus1 field plus 1 specifies the number of tiles in each Atlas frame referencing AFPS.

[0708] AFTI can include as many afti_tile_idx[i] fields as the afti_num_tiles_in_atlas_frame_minus1 field.

[0709] The afti_tile_idx[i] field specifies the tile index of the i-th tile in each Atlas frame referencing AFPS.

[0710] The `afti_single_tile_per_tile_group_flag` field equal to 1 indicates that each tile group (or tile) referencing AFPS includes one tile. The `afti_single_tile_per_tile_group_flag` field equal to 0 indicates that a tile group (or tile) referencing AFPS may include more than one tile.

[0711] If the value of the afti_single_tile_per_tile_group_flag field is false (e.g., 0), then the afti_num_tile_groups_in_atlas_frame_minus1 field can be included in AFTI.

[0712] The value of the `afti_num_tile_groups_in_atlas_frame_minus1` field plus 1 specifies the number of tile groups (or tiles) in each Atlas frame referencing AFPS. The value of the `afti_num_tile_groups_in_atlas_frame_minus1` field can be in the range of 0 to `NumTilesInAtlasFrame-1` (inclusive). If the `afti_num_tile_groups_in_atlas_frame_minus1` field does not exist and `afti_single_tile_per_tile_group_flag` is equal to 1, then it can be inferred that the value of `afti_num_tile_groups_in_atlas_frame_minus1` is equal to `NumTilesInAtlasFrame-1`.

[0713] AFTI can include the same number of fields as the afti_top_left_tile_idx[i] and afti_bottom_right_tile_idx_delta[i] fields as the afti_num_tile_groups_in_atlas_frame_minus1 field.

[0714] The `afti_top_left_tile_idx[i]` field can specify the tile index of the top-left tile in the i-th tile group (or tile). For any `i` not equal to `j`, the value of the `afti_top_left_tile_idx[i]` field is not equal to the value of the `afti_top_left_tile_idx[j]` field. If the `afti_top_left_tile_idx[j]` field does not exist, the value of `afti_top_left_tile_idx[i]` can be inferred to be equal to `i`. The length of the `afti_top_left_tile_idx[i]` field can be Ceil(Log2(NumTilesInAtlasFrame)) bits.

[0715] The `afti_bottom_right_tile_idx_delta[i]` field specifies the difference between the tile index of the bottom-right tile in the i-th tile group (or tile) and the `afti_top_left_tile_idx[i]` field. If the `afti_single_tile_per_tile_group_flag` field is equal to 1, then the value of the `afti_bottom_right_tile_idx_delta[i]` field can be inferred to be equal to 0. The length of the `afti_bottom_right_tile_idx_delta[i]` field can be Ceil(Log2(NumTilesInAtlasFrame-afti_top_left_tile_idx[i]))` bits.

[0716] The `afti_signalled_tile_group_id_flag` field being equal to 1 indicates that the tile group ID or the tile ID of each tile group is signaled.

[0717] If the value of the afti_signalled_tile_group_id_flag field is 1, then the afti_signalled_tile_group_id_length_minus1 field can be included in AFTI.

[0718] The value of the `afti_signalled_tile_group_id_length_minus1` field plus 1 specifies the number of bits used to represent the syntax element `afti_tile_group_id[i]`. If the `afti_tile_group_id[i]` field exists, the syntax element `atgh_address` may exist in the tile (or tile group) header.

[0719] AFTI can include as many afti_tile_group_id[i] fields as the afti_signalled_tile_group_id_length_minus1 field.

[0720] The afti_tile_group_id[i] field specifies the tile group (or tile) ID of the i-th tile group. The length of the afti_tile_group_id[i] field is afti_signalled_tile_group_id_length_minus1+1 bits.

[0721] Figure 40 An exemplary syntax structure for Supplemental Enhancement Information (SEI) according to an implementation is shown. Figure 40 The SEI can be included in the sample stream NAL cell that conforms to the HEVC NAL cell.

[0722] The receiving method / apparatus and system according to the embodiments can be configured to decode, recover and display point cloud data based on SEI messages.

[0723] According to the implementation method, the SEI message may include a prefix SEI message or a suffix SEI message. The payload of each SEI message signals the information corresponding to the payload type information (payloadType) and payload size information (payloadSize) through the SEI message payload (sei_payload(payloadType,payloadSize)).

[0724] For example, when payloadType indicates 13, the payload may include 3D region mapping (3d_region_mapping(payloadSize)) information.

[0725] According to an embodiment, when the psd_unit_type field indicates the prefix (PSD_PREFIX_SEI), the SEI may include buffering_period(payloadSize), pic_timing(payloadSize), filler_payload(payloadSi ze), user_data_registered_itu_t_t35(payloadSize), user_data_unregistered(payloadSize), recovery_point(payloadSize), no_display(payloadS ize), time_code(payloadSize), regional_nesting(payloadSize), sei_manifest(payloadSize), sei_prefix_indication(payloadSize), geometry_tra nsformation_params(payloadSize), 3d_bounding_box_info(payloadSize) and 3d_region_mapping(payloadSize), reserved_sei_message(payloadSize).

[0726] According to the implementation, when the psd_unit_type field indicates a suffix (PSD_SUFFIX_SEI), the SEI can include filler_payload(payloadSize), user_data_registered_itu_t_t35(payloadSize), user_data_unregistered(payloadSize), decoded_pcc_hash(payloadSize), and reserved_sei_message(payloadSize).

[0727] Figure 41 An exemplary syntax structure for the 3D bounding box information (3d_bounding_box_info(payloadSize)) SEI according to an implementation is shown.

[0728] According to the implementation, if the psd_unit_type field indicates a prefix (PSD_PREFIX_SEI) and the payload type (payloadType) is 12, then the payload of the SEI message (or SEI) may include 3D bounding box information (3d_bounding_box_info(payloadSize)). A payload type (payloadType) of 12 is an example, and this disclosure is not limited thereto, as the value of payloadType can be easily changed by those skilled in the art.

[0729] According to the implementation method, the 3D bounding box information may include the 3dbi_cancel_flag field. If the value of the 3dbi_cancel_flag field is false, the 3D bounding box information may also include the object_id field, the 3d_bounding_box_x field, the 3d_bounding_box_y field, the 3d_bounding_box_z field, the 3d_bounding_box_delta_x field, the 3d_bounding_box_delta_y field, and the 3d_bounding_box_delta_z field.

[0730] The 3dbi_cancel_flag field being equal to 1 indicates that the 3D bounding box information SEI message cancels the persistence of any previous 3D bounding box information SEI messages in the output order.

[0731] The object_id field specifies the identifier of the point cloud object / content carried in the bitstream.

[0732] The 3d_bounding_box_x field indicates the X-coordinate value of the origin of the object's 3D bounding box.

[0733] The 3d_bounding_box_y field indicates the Y-coordinate value of the origin of the object's 3D bounding box.

[0734] The 3d_bounding_box_z field indicates the Z-coordinate value of the origin of the object's 3D bounding box.

[0735] The 3d_bounding_box_delta_x field indicates the size of the bounding box on the X-axis of the object.

[0736] The 3d_bounding_box_delta_y field indicates the size of the bounding box on the Y-axis of the object.

[0737] The 3d_bounding_box_delta_z field indicates the size of the bounding box on the Z-axis of the object.

[0738] Figure 42 An exemplary syntax structure for 3D region mapping information (3d_region_mapping(payloadSize)) SEI according to an implementation is shown.

[0739] According to the implementation, if the psd_unit_type field indicates a prefix (PSD_PREFIX_SEI) and the payload type (payloadType) is 13, then the payload of the SEI message (or SEI) may include 3D region mapping information (3d_region_mapping(payloadSize)). A payload type (payloadType) of 13 is an example, and this disclosure is not limited thereto, as the value of payloadType can be easily changed by those skilled in the art.

[0740] According to the implementation method, the 3D region mapping information may include the 3dmi_cancel_flag field. If the value of the 3dmi_cancel_flag field is false, the 3D bounding box information may also include the num_3d_regions field.

[0741] The 3dmi_cancel_flag field being equal to 1 indicates that the 3D Region Mapping Information SEI message cancels the persistence of any previous 3D Region Mapping Information SEI messages in the output order.

[0742] The num_3d_regions field can indicate the number of 3D regions that are signaled in the corresponding SEI.

[0743] According to one implementation, the 3D region mapping information may include an iterative statement that repeats as many times as the value of the num_3d_regions field. In this implementation, the index i is initialized to 0 and incremented by 1 each time an iterative statement is executed, and the iterative statement is repeated until the value of i becomes the value of the num_3d_regions field. This iterative statement may include the 3d_region_idx[i] field, the 3d_region_anchor_x[i] field, the 3d_region_anchor_y[i] field, the 3d_region_anchor_z[i] field, the 3d_region_type[i] field, and the num_2d_regions[i] field.

[0744] The 3d_region_idx[i] field can indicate the identifier of the i-th 3D region.

[0745] The fields 3d_region_anchor_x[i], 3d_region_anchor_y[i], and 3d_region_anchor_z[i] can respectively indicate the x, y, and z coordinate values ​​of the anchor point of the i-th 3D region. For example, when the 3D region is a cube, the anchor point can be the origin of the cube. That is, the fields 3d_region_anchor_x[i], 3d_region_anchor_y[i], and 3d_region_anchor_z[i] can respectively indicate the x, y, and z coordinate values ​​of the origin position of the cube in the i-th 3D region.

[0746] The 3d_region_type[i] field can indicate the type of the i-th 3D region and has type values ​​ranging from 0x01 to cube.

[0747] The 3d_region_type[i] field being equal to 1 indicates that the type of the 3D region is a cube. If the value of the 3d_region_type[i] field is 1, then the 3d_region_delta_x[i], 3d_region_delta_y[i], and 3d_region_delta_z[i] fields can be included in the 3D region mapping information.

[0748] The 3d_region_delta_x[i], 3d_region_delta_y[i], and 3d_region_delta_y[i] fields can respectively indicate the differences on the x, y, and z axes of the i-th 3D region.

[0749] The `num_2d_regions[i]` field can indicate the number of 2D regions in a frame containing video or Atlas data associated with the i-th 3D region. Depending on the implementation, the 2D regions may correspond to Atlas frames.

[0750] According to one implementation, the 3D region mapping information may include an iterative statement that repeats as many times as the value of the num_2d_regions[i] field. In one implementation, the index j is initialized to 0 and incremented by 1 each time an iterative statement is executed, and the iterative statement is repeated until the value of j becomes the value of the num_2d_regions[i] field. This iterative statement may include the 2d_region_idx[j] field, the 2d_region_top[j] field, the 2d_region_left[j] field, the 2d_region_width[j] field, the 2d_region_height[j] field, and the num_tiles[j] field.

[0751] The 2d_region_idx[j] field can indicate the identifier of the j-th 2D region of the i-th 3D region.

[0752] The 2d_region_top[j] and 2d_region_left[j] fields can respectively indicate the vertical and horizontal coordinate values ​​of the top-left position of the j-th 2D region in the i-th 3D region within the frame.

[0753] The 2d_region_width[j] and 2d_region_height[j] fields can respectively include the horizontal width and vertical height of the j-th 2D region in the frame of the i-th 3D region.

[0754] The num_tiles[j] field can indicate the number of Atlas tiles or video tiles associated with the j-th 2D region of the i-th 3D region.

[0755] According to one implementation, the 3D region mapping information may include an iterative statement that repeats as many times as the value of the num_tiles[j] field. In one implementation, the index k is initialized to 0 and incremented by 1 each time an iterative statement is executed, and the iterative statement is repeated until the value of k becomes the value of the num_tiles[j] field. The iterative statement may include the tile_idx[k] field and the num_tile_groups[k] field.

[0756] The tile_idx[k] field can indicate the identifier of the k-th Atlas tile or video tile associated with the j-th 2D region.

[0757] The `num_tile_groups[k]` field indicates the number of the k-th Atlas tile group or video tile group associated with the j-th 2D region. This value can correspond to the number of Atlas tiles or video tiles.

[0758] According to one implementation, the 3D region mapping information may include an iterative statement that is repeated as many times as the value of the num_tile_groups[k] field. In this implementation, the index k is initialized to 0 and incremented by 1 each time an iterative statement is executed, and the iterative statement is repeated until the value of k becomes the value of the num_tile_groups[k] field. This iterative statement may include the tile_group_idx[m] field.

[0759] The `tile_group_idx[m]` field can indicate the identifier of the `m`th Atlas tile group or video tile group associated with the `j`th 2D region. This value can correspond to the tile index.

[0760] According to the signaling method of the implementation method, the receiving method / apparatus can monitor the mapping relationship between the 3D region and one or more Atlas tiles (2D regions) and obtain the corresponding data.

[0761] Figure 43 An exemplary syntax structure for a volumetric tiling information (payloadSize) SEI message according to an embodiment is shown. According to the embodiment, the volumetric tiling information may include identification information, 3D bounding box information, and / or 2D bounding box information for each spatial region. The volumetric tiling information may also include priority information, dependency information, and hiding information. Therefore, when rendering point cloud data, the receiving device can display each spatial region (or spatial object) on a display device in a form suitable for the corresponding information.

[0762] According to the implementation, if the psd_unit_type field indicates a prefix (PSD_PREFIX_SEI) and the payload type (payloadType) is 14, then the payload of the SEI message (or SEI) may include volumetric tiling information (volumetric_tiling_info(payloadSize)). A payload type (payloadType) of 14 is an example, and this disclosure is not limited thereto, as the value of payloadType can be easily changed by those skilled in the art.

[0763] According to the implementation, the volumetric tile information (SEI) message can instruct the V-PCC decoder (e.g., a point cloud video decoder) to avoid decoding different characteristics of the point cloud, including associations with objects in 2D Atlas and 3D space, associations with regions, labeling and relationships of regions, and correspondences of regions.

[0764] The persistence of this SEI message can be for the remainder of the bitstream, or until a new volume concatenation SEI message is encountered. Only the corresponding parameters specified in the SEI message can be updated. Previously defined parameters from earlier SEI messages can be persistent if the parameters are not modified and if the value of vti_cancel_flag is not equal to 1.

[0765] According to the implementation method, the volumetric tile information may include the vti_cancel_flag field. If the value of the vti_cancel_flag field is false, the volumetric tile information may include the vti_object_label_present_flag field, the vti_3d_bounding_box_present_flag field, the vti_object_priority_present_flag field, the vti_object_hidden_present_flag field, the vti_object_collision_shape_present_flag field, the vti_object_dependency_present_flag field, and the volumetric_tiling_info_objects() field.

[0766] The vti_cancel_flag field being equal to 1 indicates that the volume tile information SEI message cancels the persistence of any previous volume tile information SEI messages in the output order.

[0767] A value of 1 for the `vti_object_label_present_flag` field indicates that the object label information exists in the current volume tile information SEI message. A value of 0 for the `vti_object_label_present_flag` field indicates that the object label information does not exist.

[0768] A value of 1 for the `vti_3d_bounding_box_present_flag` field indicates that 3D bounding box information exists in the current volume tile information SEI message. A value of 0 for the `vti_3d_bounding_box_present_flag` field indicates that 3D bounding box information does not exist.

[0769] A value of 1 for the `vti_object_priority_present_flag` field indicates that object priority information exists in the current volume tile information SEI message. A value of 0 for the `vti_object_priority_present_flag` field indicates that object priority information does not exist.

[0770] A value of 1 for the `vti_object_hidden_present_flag` field indicates that hidden object information exists in the current volume tile information SEI message. A value of 0 for the `vti_object_hidden_present_flag` field indicates that hidden object information does not exist.

[0771] A value of 1 for the `vti_object_collision_shape_present_flag` field indicates that the object's collision shape information exists in the current volume tile information SEI message. A value of 0 for the `vti_object_collision_shape_present_flag` field indicates that the object's collision shape information does not exist.

[0772] A value of 1 for the `vti_object_dependency_present_flag` field indicates that object dependency information exists in the current volume tile information SEI message. A value of 1 for the `vti_object_dependency_present_flag` field indicates that object dependency information does not exist.

[0773] If the value of the vti_object_label_present_flag field is 1, then the volume tiling information SEI message includes volume tiling information labels (volumetric_tiling_info_labels()).

[0774] Reference Figure 44 Describe the volume tiling information labels (volumetric_tiling_info_labels()) in detail.

[0775] If the value of the vti_3d_bounding_box_present_flag field is 1, the volume tile information SEI message may include the vti_bounding_box_scale_log2 field, the vti_3d_bounding_box_scale_log2 field, and the vti_3d_bounding_box_precision_minus8 field.

[0776] The `vti_bounding_box_scale_log2` field indicates the scale to be applied to the 2D bounding box parameters that can be specified for the object.

[0777] The vti_3d_bounding_box_scale_log2 field indicates the scale to be applied to the 3D bounding box parameters that can be specified for the object.

[0778] The value of the vti_3d_bounding_box_precision_minus8 field plus 8 indicates the precision of the 3D bounding box parameters that can be specified for an object.

[0779] When each of the vti_object_label_present_flag, vti_3d_bounding_box_present_flag, vti_object_priority_present_flag, vti_object_hidden_present_flag, vti_object_collision_shape_present_flag, and vti_object_dependency_present_flag fields has a value of 1, volumetric_tiling_info_objects() specifies additional information.

[0780] Reference Figure 45 Detailed description volumetric_tiling_info_objects (vti_object_label_present_flag, vti_3d_bounding_box_present_flag, vti_object_priority_present_flag, vti_object_hidden_present_flag, vti_object_collision_shape_present_flag and vti_object_dependency_present_flag)

[0781] Figure 44 This illustrates an exemplary syntax structure for volumetric tiling information labels (volumetric_tiling_info_labels()) according to an implementation. In this implementation, if the value of the vti_object_label_present_flag field is 1, such as... Figure 43 As shown, the volumetric tiling information labels (volumetric_tiling_info_labels()) can be included in the volumetric tiling information SEI message. In another embodiment, a separate payload type can be assigned to the volumetric tiling information labels (volumetric_tiling_info_labels()) information, thereby including the volumetric_tiling_info_labels() information in the sample stream NAL unit as a separate SEI message.

[0782] According to the implementation method, the volumetric tiling information label (volumetric_tiling_info_labels()) information may include the vti_object_label_language_present_flag field and the vti_num_object_label_updates field.

[0783] A value of 1 for the `vti_object_label_language_present_flag` field indicates that the object label language information exists in the current volume tile information SEI message. A value of 0 for the `vti_object_label_language_present_flag` field indicates that the object label language information does not exist.

[0784] If the value of the vti_object_label_language_present_flag field is 1, then the volumetric_tiling_info_labels() information can include the vti_bit_equal_to_zero field and the vti_object_label_language field.

[0785] The value of the vti_bit_equal_to_zero field is equal to 0.

[0786] The `vti_object_label_language` field contains the language label, followed by a null terminator byte equal to 0x00. The length of the `vti_object_label_language` field can be less than or equal to 255 bytes (excluding the null terminator byte).

[0787] The `vti_num_object_label_updates` field indicates the number of object labels to be updated by the current SEI.

[0788] The vti_label_idx[i] field, vti_label_cancel_flag field, vti_bit_equal_to_zero field, and / or vti_label[i] field, which have the same value as the vti_num_object_label_updates field, can be included in the volumetric_tiling_info_labels() information.

[0789] The vti_label_idx[i] field indicates the label index of the i-th label to be updated.

[0790] A value of 1 for the `vti_label_cancel_flag` field indicates that the label with the same index as the `vti_label_idx[i]` field is canceled and set to an empty string. A value of 0 for `vti_label_cancel_flag` indicates that the label with the same index as the `vti_label_idx[i]` field is updated with the information following that element.

[0791] The value of the vti_bit_equal_to_zero field is equal to 0.

[0792] The vti_label[i] field indicates the label of the i-th label. The length of the vti_label[i] field can be less than or equal to 255 bytes (excluding the null terminator byte).

[0793] Figure 45 An exemplary syntax structure for volumetric tiling information objects (volumetric_tiling_info_objects()) according to an embodiment is shown. In the embodiment, the volumetric tiling information objects (volumetric_tiling_info_objects()) information includes information based on... Figure 43 Additional information about the values ​​of the vti_object_label_present_flag, vti_3d_bounding_box_present_flag, vti_object_priority_present_flag, vti_object_hidden_present_flag, vti_present_flag_collision_shape_present_flag, and vti_object_dependency_present_flag fields.

[0794] According to the implementation, the volumetric_tiling_info_objects() information includes the vti_num_object_updates field.

[0795] The vti_num_object_updates field indicates the number of objects to be updated by the current SEI.

[0796] According to the implementation, the volumetric_tiling_info_objects() information includes the same number of fields as the vti_num_object_updates field: vti_object_idx[i], vti_num_object_tile_groups[i], and vti_object_cancel_flag[i].

[0797] The vti_object_idx[i] field indicates the object index of the i-th object to be updated.

[0798] The vti_num_object_tile_groups[i] field indicates the number of Atlas tile groups (or tiles) associated with the object identified by the vti_object_idx[i] field (i.e., the object with index i).

[0799] According to the implementation, the volumetric_tiling_info_objects() information includes the same number of vti_num_object_tile_group_id[k] fields as the vti_num_object_tile_groups[i] fields.

[0800] The vti_num_object_tile_group_id[k] field indicates the identifier of the k-th Atlas tile group (or tile) of the object with index i.

[0801] In the implementation, the value of the vti_num_object_tile_group_id field is equal to the identifier of the associated Atlas tile group that was signaled in the Atlas frame tile information (atlas_frame_tile_information()) of AFPS (i.e., Figure 39 (The value of the atti_tile_group_id field).

[0802] The `vti_object_cancel_flag[i]` field being equal to 1 indicates that the object at index `i` is canceled, and the parameter `ObjectTracked[i]` is set to 0. The 2D and 3D bounding box parameters of the object can be set to 0. The `vti_object_cancel_flag[i]` field being equal to 0 indicates that the object at index `vti_object_idx[i]` is updated with the information following this field, and the parameter `ObjectTracked[i]` is set to 1.

[0803] If the value of the vti_object_cancel_flag[i] field is false (i.e., 0), then the number of objects updated based on the values ​​of the vti_object_label_present_flag, vti_3d_bounding_box_present_flag, vti_object_priority_present_flag, vti_object_hidden_present_flag, vti_object_collision_shape_present_flag, and vti_object_dependency_present_flag fields is added to the object-related information.

[0804] In other words, a value of 1 for the vti_bounding_box_update_flag[vti_object_idx[i]] field indicates the existence of 2D bounding box information for an object with index i. A value of 0 for the vti_bounding_box_update_flag[vti_object_idx[i]] field indicates the absence of 2D bounding box information.

[0805] If the value of the vti_bounding_box_update_flag[vti_object_idx[i]] field is 1, then the vti_bounding_box_top[vti_object_idx[i]], vti_bounding_box_left[vti_object_idx[i]], vti_bounding_box_width[vti_object_idx[i]], and vti_bounding_box_height[vti_object_idx[i]] fields can be included in the volumetric_tiling_info_objects() information.

[0806] The `vti_bounding_box_top[vti_object_idx[i]]` field indicates the vertical coordinate value of the top-left position of the bounding box of the object with index `i` in the current Atlas frame.

[0807] The `vti_bounding_box_left[vti_object_idx[i]]` field indicates the horizontal coordinate value of the top-left position of the bounding box of the object with index `i` in the current Atlas frame.

[0808] The `vti_bounding_box_width[vti_object_idx[i]]` field indicates the width of the bounding box of the object at index `i`.

[0809] The `vti_bounding_box_height[vti_object_idx[i]]` field indicates the height of the bounding box of the object with index `i`.

[0810] According to the implementation method, if vti3dBoundingBoxPresentFlag (i.e., Figure 43 If the value of the `vti_3d_bounding_box_present_flag` field is 1, then the `vti_3d_bounding_box_update_flag[vti_object_idx[i]]` field can be included in the `volumetric_tiling_info_objects()` information. Furthermore, if the `vti_3d_bounding_box_update_flag[vti_object_idx[i]]` field is 1, then the `vti_3d_bounding_box_x[vti_object_idx[i]]` field, `vti_3d_bounding_b`, and `vti_object_idx[i]]` fields are also included. The fields ox_y[vti_object_idx[i]], vti_3d_bounding_box_z[vti_object_idx[i]], vti_3d_bounding_box_delta_x[vti_object_idx[i]], vti_3d_bounding_box_delta_y[vti_object_idx[i]], and vti_3d_bounding_box_delta_z[vti_object_idx[i]] can be included in the volumetric_tiling_info_objects() information.

[0811] The field `vti_3d_bounding_box_update_flag[vti_object_idx[i]]` equals 1, indicating that there is 3D bounding box information for the object at index `i`. The field `vti_3d_bounding_box_update_flag[vti_object_idx[i]]` equals 0, indicating that there is no 3D bounding box information.

[0812] The vti_3d_bounding_box_x[vti_object_idx[i]] field indicates the X-coordinate value of the origin of the 3D bounding box of the object with index i.

[0813] The `vti_3d_bounding_box_y[vti_object_idx[i]]` field indicates the Y-coordinate of the origin of the 3D bounding box of the object with index `i`.

[0814] The vti_3d_bounding_box_z[vti_object_idx[i]] field indicates the Z-coordinate value of the origin of the 3D bounding box of the object with index i.

[0815] The vti_3d_bounding_box_delta_x[vti_object_idx[i]] field indicates the size of the bounding box of the object with index i on the X-axis.

[0816] The field vti_3d_bounding_box_delta_y[vti_object_idx[i]] represents the size of the bounding box on the Y-axis for the object with index i.

[0817] The field vti_3d_bounding_box_delta_z[vti_object_idx[i]] represents the size of the bounding box of the object with index i on the Z-axis.

[0818] According to the implementation method, if vtiObjectPriorityPresentFlag (i.e., Figure 43 If the value of the vti_object_priority_present_flag field is 1, then the vti_object_priority_update_flag[vti_object_idx[i]] field can be included in the volumetric_tiling_info_objects() information. And if the vti_object_priority_update_flag[vti_object_idx[i]] field is 1, then the vti_object_priority_value[vti_object_idx[i]] field can be included in the volumetric_tiling_info_objects() information.

[0819] The field `vti_object_priority_update_flag[vti_object_idx[i]]` equal to 1 indicates that there is object priority update information for the object with index `i`. The field `vti_object_priority_update_flag[vti_object_idx[i]]` equal to 0 indicates that there is no object priority update information.

[0820] The `vti_object_priority_value[vti_object_idx[i]]` field indicates the priority of the object at index `i`. A lower priority value indicates a higher priority.

[0821] According to the implementation method, if vtiObjectHiddenPresentFlag (i.e., Figure 43 If the value of the vti_object_hidden_present_flag field is 1, then the vti_object_hidden_flag[vti_object_idx[i]] field can be included in the volumetric_tiling_info_objects() information.

[0822] The field `vti_object_hidden_flag[vti_object_idx[i]]` equals 1, indicating that the object at index `i` is hidden. The field `vti_object_hidden_flag[vti_object_idx[i]]` equals 0, indicating that an object with index `i` exists.

[0823] According to the implementation method, if vtiObjectLabelPresentFlag (i.e., Figure 43 If the value of the vti_object_label_present_flag field is 1, then the vti_object_label_update_flag[vti_object_idx[i]] field can be included in the volumetric_tiling_info_objects() information, and if the value of the vti_object_label_update_flag[vti_object_idx[i]] field is 1, then the vti_object_label_idx[vti_object_idx[i]] field can be included in the volumetric_tiling_info_objects() information.

[0824] The field `vti_object_label_update_flag[vti_object_idx[i]]` equal to 1 indicates that there is an object label update for the object with index `i`. The field `vti_object_label_update_flag[vti_object_idx[i]]` equal to 0 indicates that there is no object label update.

[0825] The vti_object_label_idx[vti_object_idx[i]] field indicates the label index of the object with index i.

[0826] According to the implementation method, if vtiObjectCollisionShapePresentFlag(i.e., Figure 43 If the value of the vti_object_collision_shape_present_flag field is 1, then the vti_object_collision_shape_update_flag[vti_object_idx[i]] field can be included in the volumetric_tiling_info_objects() information. And if the vti_object_collision_shape_update_flag[vti_object_[i]] field is 1, then the vti_object_collision_shape_id[vti_object_idx[i]] field can be included in the volumetric_tiling_info_objects() information.

[0827] The field `vti_object_collision_shape_update_flag[vti_object_idx[i]]` equals 1, indicating that there is object collision shape update information for the object with index `i`. `vti_object_collision_shape_update_flag[i]` equals 0, indicating that there is no object collision shape update information.

[0828] The vti_object_collision_shape_id[vti_object_idx[i]] field indicates the collision shape ID of the object with index i.

[0829] According to the implementation method, if vtiObjectDependencyPresentFlag (i.e., Figure 43If the value of vti_object_dependency_present_flag is 1, then the vti_object_dependency_update_flag[vti_object_idx[i]] field can be included in the volumetric_tiling_info_objects() information. And if the value of vti_object_dependency_update_flag[vti_object_idx[i]] field is 1, then the vti_object_num_dependencies[vti_object_idx[i]] field can be included in the volumetric_tiling_info_objects() information.

[0830] The field `vti_object_dependency_update_flag[vti_object_idx[i]]` equal to 1 indicates that there is object dependency update information for the object with index `i`. The field `vti_object_dependency_update_flag[vti_object_idx[i]]` equal to 0 indicates that there is no object dependency update information.

[0831] The `vti_object_num_dependencies[vti_object_idx[i]]` field indicates the number of dependencies of the object at index `i`.

[0832] The same number of vti_object_dependency_idx[vti_object_idx[i]][j] fields as the vti_object_num_dependencies[vti_object_idx[i]] field can be included in the volumetric_tiling_info_objects() information.

[0833] The `vti_object_dependency_idx[vti_object_idx[i]][j]` field indicates the index of the `j`th object that is dependent on the object with index `i`.

[0834] In addition, it has the following characteristics Figure 25 The V-PCC bitstream shown can be sent to the receiving side, or it can be transmitted by... Figure 1 , Figure 18 , Figure 20 or Figure 21 The file / fragment encapsulation module or unit is encapsulated in ISOBMFF file format and then sent to the receiving side.

[0835] In the latter case, V-PCC streams can be sent via multiple tracks or a single track of a file. In this case, it can be done via... Figure 20 or Figure 22 The file / fragment decapsulation module of the receiving device decapsulates the file into a V-PCC bitstream.

[0836] For example, a V-PCC bitstream carrying a V-PCC parameter set, geometry bitstream, occupancy map bitstream, attribute bitstream, and / or Atlas data bitstream can be transmitted via... Figure 20 or Figure 21 The file / fragment encapsulation module encapsulates the data. In this implementation, the V-PCC bitstream is stored in one or more tracks within an ISOBMFF-based file.

[0837] According to the implementation method, an ISOBMFF-based file may be referred to as a container, container file, media file, V-PCC file, etc. More specifically, the file may consist of boxes and / or information, which may be referred to as ftyp, meta, moov, or mdat.

[0838] The ftyp box (file type box) can provide information related to file compatibility or file type. The receiving side can refer to the ftyp box to identify files.

[0839] Meta boxes can include vpcg{0,1,2,3} boxes (V-PCC group boxes).

[0840] The mdat box, also known as a media data box, may include the actual media data. According to an implementation, the video-coded geometry bitstream, video-coded attribute bitstream, video-coded occupancy graph bitstream, and / or Atlas data bitstream are included in a sample of the mdat box within the file. According to an implementation, the sample may be referred to as a V-PCC sample.

[0841] A moov box, also known as a movie box, can contain metadata about the file's media data (e.g., geometric bitstreams, attribute bitstreams, occupancy graph bitstreams, etc.). For example, it can contain information needed to decode and play the media data, as well as information about samples of the file. A moov box can serve as a container for all metadata. The moov box can be the highest-level box among metadata-related boxes. Depending on the implementation, only one moov box may exist in a file.

[0842] The box according to the embodiment may include a track box that provides information related to the track of a file. The track box may include a media box that provides media information about the track and a track reference container (tref) for referencing the track and samples of the file corresponding to the track.

[0843] The mdia box may include a media information container (minf) box that provides information about the media data and a handler box that indicates the type of stream.

[0844] The minf box can include a sample table (stbl) box that provides metadata related to samples of the mdat box.

[0845] The stbl box may include a sample description (stsd) box, which provides information about the encoding type used and the initialization information required for that encoding type.

[0846] According to an implementation, the STSD box may include sample entries for storing tracks of the V-PCC bitstream.

[0847] In this disclosure, the track in the document that carries some or all of the V-PCC bitstream may be referred to as a V-PCC track or a volume track.

[0848] In order to store the V-PCC bitstream according to the implementation in a single track or multiple tracks in a file, this disclosure defines a volumetric visual track, a volumetric visual media header, a volumetric sample entry, a volumetric sample, and samples and simple entries for the V-PCC track.

[0849] The term V-PCC used in this paper is the same as that used for visual volumetric video coding (V3C). These two terms can be used to complement each other.

[0850] According to the implementation method, video-based point cloud compression (V-PCC) represents the volumetric encoding of point cloud visual information.

[0851] In other words, the minf box within the trak box of the moov box can also include a volumetric visual media head box. The volumetric visual media head box contains information about the volumetric visual track containing the volumetric visual scene.

[0852] Each volumetric vision scene can be represented by a unique volumetric vision track. An ISOBMFF file can contain multiple scenes, and therefore multiple volumetric vision tracks can exist within an ISOBMFF file.

[0853] According to the implementation, a volumetric visual track can be identified by the volumetric visual media handler type 'volv' included in the handler box of the MediaBox and / or by the volumetric visual media header (vvhd) in the minf box of the MediaBox. The minf box is referred to as a media information container or media information box. The minf box is included in an mdia box, the mdia box is included in a trak box, and the trak box is included in the moov box of the file. A single volumetric visual track or multiple volumetric visual tracks may exist in the file.

[0854] According to one implementation, the volumetric visual track can use the volumetric visual media headbox (VolumetricVisualMediaHeaderBox) within the MediaInformationBox. The MediaInformationBox is referred to as the minf box, and the VolumetricVisualMediaHeaderBox is referred to as the vvhd box. According to one implementation, the vvhd box can be defined as follows.

[0855] Box type: 'vvhd'

[0856] Container: MediaInformationBox

[0857] Mandatory: Yes

[0858] Quantity: Exactly one

[0859] The syntax for a volumetric visual media headbox (i.e., a box of type vvhd) is as follows.

[0860] aligned(8) class VolumetricVisualMediaHeaderBox

[0861] Extend FullBox('vvhd', version = 0, 1){

[0862] }

[0863] version can be an integer indicating the version of the box.

[0864] According to the implementation method, the volumetric visual track can use volumetric visual sample entries to transmit signaling information as follows.

[0865] Class VolumetricVisualSampleEntry(codingname)

[0866] Extend SampleEntry(codingname){

[0867] unsigned int(8)

[32] compressor_name;

[0868] }

[0869] `compressor_name` is the name used for informational purposes. It is formatted as a fixed 32-byte field, where the first byte is set to the number of bytes to be displayed, followed by the number of bytes of displayable data encoded in UTF-8, and then padding to complete a total of 32 bytes (including size bytes). This field can be set to 0.

[0870] The format of the volumetric visual sample according to the implementation method can be defined by the encoding system.

[0871] According to the implementation method, the V-PCC unit header box can exist in the V-PCC track included in the sample entry and in the V-PCC component tracks included in all video encodings in the scheme information. The V-PCC unit header box can contain V-PCC unit headers for data carried by the respective tracks.

[0872] aligned(8) class VPCCUnitHeaderBox extends FullBox('vunt', version = 0, 0) {

[0873] vpcc_unit_header()unit_header;

[0874] }

[0875] In other words, VPCCUnitHeaderBox can include vpcc_unit_header(). Figure 30 An example of the syntax structure of (vpcc_unit_header()) is shown.

[0876] According to the implementation, the sample entries inherited by the Volumetric Visual Sample Entry (i.e., the higher class of Volumetric Visual Sample Entry) include the VPCC Decoder Configuration Box.

[0877] According to the implementation, VPCCConfigurationBox may include a VPCC decoder configuration record (VPCCDecoderConfigurationRecord) as shown below.

[0878] aligned(8) class VPCCDDecoderConfigurationRecord{

[0879] unsigned int(8)configurationVersion=1;

[0880] unsigned int(3)sampleStreamSizeMinusOne;

[0881] unsigned int(5)numOfVPCCParameterSets;

[0882] for(i=0; i <numOfVPCCParameterSets;i++){

[0883] sample_stream_vpcc_unit VPCCParameterSet;

[0884] }

[0885] unsigned int(8)numOfAtlasSetupUnits;

[0886] for(i=0; i <numOfAtlasSetupUnits;i++){

[0887] sample_stream_vpcc_unit atlas_setupUnit;

[0888] }

[0889] }

[0890] The `configurationVersion` field included in `VPCCDDecoderConfigurationRecord` indicates the version. Incompatible changes to the record are indicated by changes to the version number.

[0891] Increasing 1 to sampleStreamSizeMinusOne indicates the precision of the ssvu_vpcc_unit_size element in all sample stream V-PCC units in this configuration record or in V-PCC samples in the stream to which this configuration record applies.

[0892] numOfVPCCParameterSets specifies the number of V-PCC parameter sets (VPS) signaled in VPCCDecoderConfigurationRecord.

[0893] VPCCParameterSet is a sample_stream_vpcc_unit() instance of a V-PCC unit of type VPCC_VPS. A V-PCC unit can include vpcc_parameter_set() (see...). Figure 33 In other words, the VPCCParameterSet array can include vpcc_parameter_set(). Figure 28 An example of the syntax structure of a sample stream V-PCC unit (sample_stream_vpcc_unit()) is shown.

[0894] numOfAtlasSetupUnits indicates the number of Atlas stream setup arrays signaled in VPCCDecoderConfigurationRecord.

[0895] Atlas_setupUnit is a sample_stream_vpcc_unit() instance containing an Atlas sequence parameter set, an Atlas frame parameter set, an Atlas tile (or tile group), or an SEI Atlas NAL unit (see See also) Figures 34 to 45 ). Figure 28 An example of the syntax structure of a sample stream V-PCC unit (sample_stream_vpcc_unit()) is shown.

[0896] Specifically, the atlas_setupUnit array can include atlas parameter sets that are constant for the stream pointed to by the sample entry of the VPCCDecoderConfigurationRecord and the atlas stream SEI message (see [link]). Figures 40 to 45 According to the implementation method, atlas_setupUnit can be simply referred to as the setup unit.

[0897] According to other implementations, VPCCDDecoderConfigurationRecord can be represented as follows.

[0898] aligned(8) class VPCCDDecoderConfigurationRecord{

[0899] unsigned int(8)configurationVersion=1;

[0900] unsigned int(3)sampleStreamSizeMinusOne;

[0901] bit(2)reserved = 1;

[0902] unsigned int(3)lengthSizeMinusOne;

[0903] unsigned int(5)numOfVPCCParameterSets;

[0904] for(i=0; i <numOfVPCCParameterSets;i++){

[0905] sample_stream_vpcc_unit VPCCParameterSet;

[0906] }

[0907] unsigned int(8)numOfSetupUnitArrays;

[0908] for(j=0;j <numOfSetupUnitArrays;j++){

[0909] bit(1)array_completeness;

[0910] bit(1)reserved = 0;

[0911] unsigned int(6)NAL_unit_type;

[0912] unsigned int(8)numNALUnits;

[0913] for(i=0; i <numNALUnits;i++){

[0914] sample_stream_nal_unit setupUnit;

[0915] }

[0916] }

[0917] }

[0918] `configurationVersion` is the version field. Incompatible changes to a record are indicated by changes to the version number.

[0919] The value of lengthSizeMinusOne plus 1 indicates the precision (in bytes) of the ssnu_nal_unit_size element in all sample stream NAL units of the V-PCC samples in the stream to which VPCCDecoderConfigurationRecord or VPCCDecoderConfigurationRecord is applied. Figure 36 An example of the syntax structure for a sample stream NAL unit (sample_stream_nal_unit()) including the ssnu_nal_unit_size field is shown.

[0920] numOfVPCCParameterSets specifies the number of V-PCC parameter sets (VPS) signaled in VPCCDecoderConfigurationRecord.

[0921] VPCCParameterSet is a sample_stream_vpcc_unit() instance of a V-PCC unit of type VPCC_VPS. A V-PCC unit can include vpcc_parameter_set(). That is, a VPCCParameterSet array can include vpcc_parameter_set(). Figure 28 An example of the syntax structure of sample_stream_vpcc_unit() is shown.

[0922] numOfSetupUnitArrays indicates the number of arrays of Atlas NAL units of the indicated type.

[0923] Repeated iteration statements with the same number of values ​​as numOfSetupUnitArrays can include array_completeness.

[0924] An array_completeness of 1 indicates that all Atlas NAL units of the given type are in the array below, and none are in the stream. An array_completeness of 0 indicates that additional Atlas NAL units of the specified type may be in the stream. Default values ​​and allowed values ​​are limited by the sample entry names.

[0925] NAL_unit_type indicates the type of Atlas NAL unit in the following arrays. NAL_unit_type is restricted to one of the values ​​indicating NAL_ASPS, NAL_PREFIX_SEI, or NAL_SUFFIX_SEI Atlas NAL units.

[0926] `numNALUnits` indicates the number of Atlas NAL units of the indicator type contained in the `VPCCDecoderConfigurationRecord` for the stream to which the `VPCCDecoderConfigurationRecord` is applied. The SEI array should contain only SEI messages of a "declarative" nature, i.e., SEI messages that provide information about the entire stream. An example of such an SEI could be a user data SEI.

[0927] `setupUnit` is an instance of `sample_stream_nal_unit()`, containing an Atlas sequence parameter set, an Atlas frame parameter set, or a declarative SEI Atlas NAL unit.

[0928] sample group

[0929] According to the implementation method, Figure 20 or Figure 21 The file / fragment encapsulation unit can generate sample groups by grouping one or more samples. According to the implementation, Figure 20 or Figure 21 The file / fragment encapsulation unit or metadata processing unit can signal the signaling information associated with the sample group to the sample, sample group, or sample entry. That is, sample group information associated with the sample group can be added to the sample, sample group, or sample entry. The sample group information and the corresponding sample group description will be described below. According to the implementation, the sample group information may include V-PCC Atlas parameter set sample group information, V-PCC SEI sample group information, V-PCC bounding box sample group information, and V-PCC 3D region mapping sample group information.

[0930] V-PCC Atlas parameter set sample group

[0931] According to the implementation method, one or more samples that can be applied to the same V-PCC Atlas parameter set can be grouped, and the sample group can be called the V-PCC Atlas parameter sample group.

[0932] According to the implementation method, the syntax of the V-PCC Atlas Param Sample Group information (VPCCAtlasParamSampleGroupDescriptionEntry) associated with the V-PCC Atlas Param Sample Group can be defined as follows.

[0933] aligned(8) class VPCCAtlasParamSampleGroupDescriptionEntry() extends SampleGroupDescriptionEntry('vaps'){

[0934] unsigned int(8)numOfAtlasParameterSets;

[0935] for(i=0; i <numOfAtlasParameterSets;i++){

[0936] sample_stream_vpcc_unit atlasParameterSet;

[0937] }

[0938] }

[0939] According to the implementation, the grouping_type of “vaps” used for sample grouping indicates the assignment of samples in the V-PCC orbit to the set of Atlas parameters carried in the V-PCC Atlas parameter sample group.

[0940] According to the implementation method, a V-PCC track can contain at most one SampleToGroupBox with grouping_type equal to 'vaps'.

[0941] According to the implementation, if a SampleToGroupBox with a grouping_type equal to 'vaps' exists, then an accompanying SampleGroupDescriptionBox with the same grouping type exists and contains the ID of the group of samples.

[0942] The VPCCAtlasParamSampleGroupDescriptionEntry with group type "vaps" can include numOfAtlasParameterSets.

[0943] numOfAtlasParameterSets indicates the number of atlas parameter sets signaled in the sample group description.

[0944] The atlasParameterSet corresponding to the value of numOfAtlasParameterSets can be included in VPCCAtlasParamSampleGroupDescriptionEntry.

[0945] atlasParameterSet is an instance of a sample stream VPCC unit (sample_stream_vpcc_unit()) that includes the ASPS or AFPS associated with the set of samples.

[0946] According to another implementation, the syntax of VPCCAtlasParamSampleGroupDescriptionEntry associated with the V-PCC Atlas parameter sample group can be defined as follows.

[0947] aligned(8) class VPCCAtlasParamSampleGroupDescriptionEntry() extends SampleGroupDescriptionEntry('vaps'){

[0948] unsigned int(3)lengthSizeMinusOne;

[0949] unsigned int(5)numOfAtlasParameterSets;

[0950] for(i=0; i <numOfAtlasParameterSets;i++){

[0951] sample_stream_nal_unit atlasParameterSetNALUnit;

[0952] }

[0953] }

[0954] The increment of lengthSizeMinusOne indicates the precision (in bytes) of the ssnu_nal_unit_size element in all sample stream NAL units that are signaled in the corresponding sample group description.

[0955] numOfAtlasParameterSets indicates the number of Atlas parameter sets that are signaled in the sample group description.

[0956] The atlasParameterSetNALUnit corresponding to the value of numOfAtlasParameterSets can be included in VPCCAtlasParamSampleGroupDescriptionEntry.

[0957] atlasParameterSetNALUnit is a sample_stream_nal_unit() instance that includes the ASPS or AFPS associated with the set of samples.

[0958] V-PCC SEI sample group

[0959] According to the implementation method, one or more samples that can be applied to the same V-PCC SEI can be grouped, and the sample group can be called the V-PCC SEI sample group.

[0960] According to the implementation method, the syntax of the V-PCC SEI sample group information (VPCCSEISampleGroupDescriptionEntry) associated with the V-PCC SEI sample group can be defined as follows.

[0961] aligned(8) class VPCCSEISampleGroupDescriptionEntry() extends SampleGroupDescriptionEntry('vsei'){

[0962] unsigned int(8) numOfSEIs;

[0963] for(i=0; i <numOfSEISets;i++){

[0964] sample_stream_vpcc_unit_sei;

[0965] }

[0966] }

[0967] According to the implementation, the grouping_type of “vsei” used for sample grouping indicates the assignment of samples in the V-PCC track to the SEI carried in the V-PCC SEI sample group.

[0968] According to the implementation, a V-PCC track can contain at most one SampleToGroupBox with grouping_type equal to 'vsei'.

[0969] According to the implementation, if a SampleToGroupBox with a grouping_type equal to 'vsei' exists, then an accompanying SampleGroupDescriptionBox with the same grouping type exists and contains the ID of the group of samples.

[0970] numOfSEI indicates the number of V-PCC SEIs that are signaled in the corresponding sample group description.

[0971] The "sei" corresponding to the value of numOfSEI can be included in VPCCSEISampleGroupDescriptionEntry.

[0972] 'sei' is a sample_stream_vpcc_unit() instance that includes the SEI associated with this set of samples.

[0973] According to another implementation, the syntax of VPCCSEISampleGroupDescriptionEntry associated with a V-PCC SEI sample group can be defined as follows.

[0974] aligned(8) class VPCCSEISampleGroupDescriptionEntry() extends SampleGroupDescriptionEntry('vsei'){

[0975] unsigned int(3)lengthSizeMinusOne;

[0976] unsigned int(5)numOfSEIs;

[0977] for(i=0; i <numOfSEIs;i++){

[0978] sample_stream_nal_unit seiNALUnit;

[0979] }

[0980] }

[0981] The lengthSizeMinusOne plus 1 indicates the precision (in bytes) of the ssnu_nal_unit_size element in all sample stream NAL units signaled in this sample group description.

[0982] numOfSEI indicates the number of V-PCC SEIs that are signaled in the corresponding sample group description.

[0983] The atlasParameterSetNALUnit corresponding to the value of numOfSEI can be included in VPCCSEISampleGroupDescriptionEntry.

[0984] seiNALUnit is a sample_stream_nal_unit() instance of the SEI associated with this set of samples.

[0985] V-PCC bounding box sample group

[0986] According to the implementation method, one or more samples that can be applied to the same V-PCC bounding box can be grouped, and the sample group can be called the V-PCC bounding box sample group.

[0987] According to the implementation method, the syntax of the V-PCC bounding box sample group information (VPCC3DBoundingBoxSampleGroupDescriptionEntry) associated with the V-PCC bounding box sample group can be defined as follows.

[0988] aligned(8) class VPCC3DBoundingBoxSampleGroupDescriptionEntry() extends SampleGroupDescriptionEntry('vpbb'){

[0989] 3DBoundingBoxInfoStruct();

[0990] }

[0991] According to the implementation, the grouping_type of “vpbb” used for sample grouping indicates the assignment of samples in the V-PCC track to 3D bounding box information carried in the V-PCC bounding box sample group.

[0992] According to the implementation, a V-PCC track can contain at most one SampleToGroupBox with grouping_type equal to 'vpbb'.

[0993] According to the implementation method, if a SampleToGroupBox with grouping_type equal to 'vpbb' exists, then an accompanying SampleGroupDescriptionBox with the same grouping type exists and contains the ID of the sample in that group.

[0994] The details included in 3DBoundingBoxInfoStruct() in the above syntax will be described below.

[0995] V-PCC 3D Region Mapping Sample Group

[0996] According to the implementation method, one or more samples that can be applied to the same V-PCC 3D region mapping can be grouped, and the sample group can be called the V-PCC 3D region mapping sample group.

[0997] According to the implementation method, the syntax of the V-PCC 3D Region Mapping Sample Group information (VPCC3DRegionMappingSampleGroupDescriptionEntry) associated with the V-PCC 3D Region Mapping Sample Group can be defined as follows.

[0998] aligned(8) class VPCC3DRegionMappingSampleGroupDescriptionEntry() extends SampleGroupDescriptionEntry('vpsr'){

[0999] VPCC3DRegionMappingBox3d_region_mapping;

[1000] }

[1001] According to the implementation, the grouping_type of “vpsr” used for sample grouping indicates the assignment of samples in the V-PCC track to the 3D region mapping information carried in the sample group.

[1002] According to the implementation, a V-PCC track can contain at most one SampleToGroupBox with grouping_type equal to 'vpsr'.

[1003] According to the implementation, if a SampleToGroupBox with grouping_type equal to 'vpsr' exists, then an accompanying SampleGroupDescriptionBox with the same grouping type exists and contains the ID of the sample in that group.

[1004] The details included in VPCC3DRegionMappingBox in the above syntax will be described below.

[1005] Orbital group

[1006] According to the implementation method, Figure 20 or Figure 21 The file / fragment encapsulation unit can generate track groups by grouping one or more tracks. According to the implementation, Figure 20 or Figure 21The file / fragment encapsulation unit or metadata processing unit can signal the signaling information associated with the track group to the sample, track group, or sample entry. That is, track group information associated with the track group can be added to the sample, track group, or sample entry. The track group information and the corresponding track group description will be described below. According to the implementation, the track group information may include 3D region track group information and 2D region track group information.

[1007] 3D regional orbit group

[1008] According to the implementation method, one or more tracks that can be applied to the same 3D spatial region can be grouped, and the track group can be referred to as a 3D region track group.

[1009] According to the implementation method, the syntax of the 3D region track group information (SpatialRegionGroupBox) associated with the 3D region track group can be defined as follows.

[1010] The aligned(8) class SpatialRegionGroupBox extends TrackGroupTypeBox('3drg'){

[1011] 3DRegionInfoStruct()

[1012] }

[1013] According to the implementation, a TrackGroupTypeBox with track_group_type equal to "3drg" indicates that the track belongs to a set of V-PCC component tracks corresponding to a 3D spatial region.

[1014] According to the implementation method, tracks belonging to the same 3D spatial region have the same track_group_id value for track_group_type of "3drg", and the track_group_id of a track in one 3D spatial region is different from the track_group_id of a track in another 3D spatial region.

[1015] According to the implementation method, tracks with the same track_group_id value within a TrackGroupTypeBox with track_group_type equal to "3drg" belong to the same 3D spatial region. Therefore, the track_group_id in the TrackGroupTypeBox with track_group_type equal to "3drg" is used as an identifier for the 3D spatial region.

[1016] SpatialRegionGroupBox can include 3DSpatialRegionStruct() instead of 3DRegionInfoStruct() above.

[1017] 3DRegionInfoStruct() and 3DSpatialRegionStruct() contain 3D region information for the orbits applied to a 3D region orbit group. The details included in 3DRegionInfoStruct() and 3DSpatialRegionStruct() are described below.

[1018] 2D regional orbital group

[1019] According to the implementation method, one or more tracks that can be applied to the same 2D region can be grouped, and the track group can be called a 2D region track group.

[1020] According to the implementation method, the syntax of the 2D region group information (RegionGroupBox) associated with the 2D region group can be defined as follows.

[1021] The aligned(8) class RegionGroupBox extends TrackGroupTypeBox('2drg'){

[1022] 2DRegionInfoStruct()

[1023] }

[1024] According to the implementation, a TrackGroupTypeBox with track_group_type equal to "2drg" can indicate that the track belongs to a set of V-PCC component tracks corresponding to a 2D region.

[1025] According to the implementation method, tracks belonging to the same 2D region have the same track_group_id value for track_group_type of "2drg", and the track_group_id of a track in one 2D region is different from the track_group_id of a track in another 2D region.

[1026] According to the implementation, tracks with the same track_group_id value within a TrackGroupTypeBox where track_group_type is equal to "2drg" belong to the same 2D region. Therefore, the track_group_id in a TrackGroupTypeBox where track_group_type is equal to "2drg" is used as an identifier for the 2D region.

[1027] 2DRegionInfoStruct() includes 2D region information for the orbits applied to a 2D region orbital group. The details included in 2DRegionInfoStruct() are described below.

[1028] As mentioned above, V-PCC bitstreams can be stored in a single track or multiple tracks and then transmitted.

[1029] Next, we will describe a multitrack container for V-PCC bitstreams associated with multiple tracks.

[1030] According to the implementation, in the general layout of a multi-track container (also known as a multi-track ISOBMFF V-PCC container), V-PCC cells in the V-PCC elementary stream can be mapped to individual tracks within the container file according to their type. There are two types of tracks in the multi-track ISOBMFF V-PCC container according to the implementation: V-PCC tracks and V-PCC component tracks.

[1031] According to the implementation, the V-PCC track is a track that carries volumetric visual information in the V-PCC bitstream, which includes an Atlas sub-bitstream and a sequence parameter set (or V-PCC parameter set).

[1032] The V-PCC component track according to the implementation is a restricted video scheme track that carries 2D video encoded data for the occupancy map, geometry, and attribute sub-bitstreams of the V-PCC bitstream. Furthermore, the V-PCC component track can satisfy the following conditions:

[1033] a) A new box was inserted into the sample entry, which records the role of the video stream contained in that track in the V-PCC system;

[1034] b) Introduce orbital references from V-PCC orbits to V-PCC component orbits to establish the membership of V-PCC component orbits in a specific point cloud represented by V-PCC orbits;

[1035] c) Set the track head flag to 0 to indicate that the track does not directly contribute to the overall layout of the film, but does contribute to the V-PCC system.

[1036] Tracks belonging to the same V-PCC sequence can be time-aligned. V-PCC component tracks across different video codes and samples of V-PCC tracks contributing to the same point cloud frame have the same rendering time. The decoding time of the V-PCC Atlas sequence parameter set and Atlas frame parameter set used for such samples is equal to or earlier than the synthesis time of the point cloud frame. Furthermore, all tracks belonging to the same V-PCC sequence have the same implicit or explicit edit list.

[1037] Note: Synchronization between basic streams in component tracks is handled by the ISOBMFF track timing structure (stts, ctts, and cslg) or equivalent mechanisms in movie clips.

[1038] Based on this layout, the V-PCC ISOBMFF container may include the following:

[1039] The -V-PCC track contains samples of V-PCC parameter sets (in the sample entries) and payloads carrying V-PCC units (unit type VPCC_VPS) and Atlas V-PCC units (unit type VPCC_AD). This track also includes track references for other tracks carrying payloads of V-PCC units that carry video compression (i.e., unit types VPCC_OVD, VPCC_GVD, and VPCC_AVD).

[1040] - A constrained video scheme track in which samples contain access units for the video-coded basic stream of the occupies graph data (i.e., payloads of V-PCC units of type VPCC_OVD).

[1041] - One or more restricted video scheme tracks, wherein the samples contain access units for the video-coded basic stream of geometric data (i.e., payloads of V-PCC units of type VPCC_GVD).

[1042] - Zero or more restricted video scheme tracks, where the samples contain access units for the video-coded basic stream of attribute data (i.e., payloads of V-PCC units of type VPCC_AVD).

[1043] Next, the V-PCC orbital will be described.

[1044] The syntax structure configuration of the V-PCC track sample entries according to the implementation method is as follows.

[1045] Sample entry types: "vpc1", "vpcg"

[1046] Container: SampleDescriptionBox

[1047] Mandatory: The sample entries “vpc1” or “vpcg” are mandatory.

[1048] Quantity: There may be one or more sample entries.

[1049] The sample entry type is "vpc1" or "vpcg".

[1050] Under the “vpc1” sample entry, all Atlas sequence parameter sets, Atlas frame parameter sets, or V-PCC SEIs are in the setupUnit array (i.e., the sample entry).

[1051] Under the “vpcg” sample entry, the Atlas sequence parameter set, the Atlas frame parameter set, and the V-PCC SEI can exist in the array (i.e., the sample entry) or in the stream (i.e., the sample).

[1052] An optional BitRateBox may be present in the VPCC volume sample entry to signal the bit rate information of the V-PCC track.

[1053] As described below, V-PCC tracks use V-PCC sample entries (VPCCSampleEntry) that inherit from VolumetricVisualSampleEntry. VPCCSampleEntry includes a V-PCC configuration box (VPCCConfigurationBox), a V-PCC unit header box (VPCCUnitHeaderBox), and / or VPCCBoundingInformationBox(). VPCCConfigurationBox contains a V-PCC decoder configuration record (VPCCDecoderConfigurationRecord).

[1054] Volume sequence:

[1055] The class VPCCConfigurationBox extends Box('vpcC'){

[1056] VPCCDecoderConfigurationRecord()VPCCConfig;

[1057] }

[1058] The aligned(8) class VPCCSampleEntry() extends VolumetricVisualSampleEntry('vpc1'){

[1059] VPCCConfigurationBoxconfig;

[1060] VPCCUnitHeaderBox unit_header;

[1061] VPCCBoundingInformationBox();

[1062] }

[1063] Figure 46 An exemplary structure of a V-PCC sample entry according to an implementation method is shown. Figure 46 In V-PCC sample entries, a VPS may be included, and optionally, an ASPS, AFPS, or SEI. That is, depending on the sample entry type (i.e., vpc1 or vpcg), an ASPS, AFPS, or SEI may be included in the sample entry or sample.

[1064] According to the implementation method, the V-PCC sample entry may also include a sample stream V-PCC header, a sample stream NAL header, and a V-PCC unit header box.

[1065] Figure 47 An exemplary structure of a moov box and an exemplary structure of a sample entry are shown according to an implementation. Specifically, the structure of a sample entry when the sample entry type is vpc1 is shown.

[1066] exist Figure 47 In the moov box, the stbl box may include a sample description (stsd) box, and the stsd box may include sample entries for storing tracks of the V-PC...

Claims

1. A point cloud data transmitting method, the point cloud data transmitting method comprising the steps of: encoding point cloud data; encapsulating a bitstream including the encoded point cloud data into a file; and transmitting the file, wherein the point cloud data includes at least geometry data, attribute data, or occupancy map data, wherein the bitstream is composed of first precision information and first units, wherein each of the first units is composed of first size information and second units, wherein the first precision information includes information for specifying precision of the first size information in the first units, wherein the first size information includes information for specifying size of the second units, wherein the second units are composed of headers and payloads, wherein the headers include type information for indicating type of data in the payloads, wherein the payloads include one of the geometry data, the attribute data, the occupancy map data, and atlas data, wherein the atlas data is composed of second precision information and third units, wherein each of the third units is composed of second size information and fourth units, wherein the second precision information includes information for specifying precision of the second size information in the third units, wherein the second size information includes information for specifying size of the fourth units, wherein the fourth units include an atlas frame parameter set including first tile identification information for identifying each of one or more tiles in an atlas frame, wherein the bitstream is stored in a plurality of tracks of the file, wherein the file further includes signaling data, wherein the signaling data includes spatial region information of the point cloud data, wherein the spatial region information is at least static spatial region information that does not vary over time or dynamic spatial region information that varies over time, wherein the point cloud data is divided into one or more 3-dimensional (3D) spatial regions, and wherein the static spatial region information includes region number information for identifying a number of the one or more 3D spatial regions, region identification information for identifying each 3D spatial region, tile number information for identifying a number of one or more tiles associated with each 3D spatial region, and second tile identification information for identifying each of the one or more tiles associated with each 3D spatial region. 2.The point cloud data transmitting method of claim 1, a value of the first tile identification information is equal to a value of the second tile identification information. wherein, 3.The point cloud data transmitting method of claim 1, the dynamic spatial region information includes region number information for identifying a number of the one or more 3D spatial regions, region identification information for identifying each 3D spatial region, and priority information and dependency information related to a 3D spatial region. wherein 4.A point cloud data transmitting apparatus, the point cloud data transmitting apparatus comprising: an encoder that encodes point cloud data; ​ an encapsulator that encapsulates a bitstream including encoded point cloud data into a file; and a transmitter that transmits the file, wherein the point cloud data includes at least geometry data, attribute data, or occupancy map data, wherein the bitstream is composed of first precision information and first units, wherein each of the first units is composed of first size information and second units, wherein the first precision information includes information for specifying precision of the first size information in the first units, wherein the first size information includes information for specifying size of the second units, wherein the second units are composed of a header and a payload, wherein the header includes type information for indicating a type of data in the payload, wherein the payload includes one of the geometry data, the attribute data, the occupancy map data, and atlas data, wherein the atlas data is composed of second precision information and third units, wherein each of the third units is composed of second size information and fourth units, wherein the second precision information includes information for specifying precision of the second size information in the third units, wherein the second size information includes information for specifying size of the fourth units, wherein the fourth units include an atlas frame parameter set including first tile identification information for identifying each of one or more tiles in an atlas frame, wherein the bitstream is stored in a plurality of tracks of the file, wherein the file further includes signaling data, wherein the signaling data includes spatial region information of the point cloud data, wherein the spatial region information is at least static spatial region information that does not change over time or dynamic spatial region information that changes over time, wherein the point cloud data is divided into one or more 3-dimensional (3D) spatial regions, and wherein the static spatial region information includes region number information for identifying a number of the one or more 3D spatial regions, region identification information for identifying each 3D spatial region, tile number information for identifying a number of one or more tiles associated with each 3D spatial region, and second tile identification information for identifying each of the one or more tiles associated with each 3D spatial region.

5. The point cloud data transmitting apparatus according to claim 4, wherein, a value of the first tile identification information is equal to a value of the second tile identification information.

6. The point cloud data transmitting apparatus according to claim 4, wherein, the dynamic spatial region information includes region number information for identifying a number of the one or more 3D spatial regions, region identification information for identifying each 3D spatial region, and priority information and dependency information related to a 3D spatial region.

7. A point cloud data receiving method including the steps of: receiving a file, decapsulating the file into a bitstream comprising point cloud data, wherein the bitstream is stored in a plurality of tracks of the file, and wherein the file further comprises signaling data; and decoding the point cloud data based on the signaling data; wherein the point cloud data comprises at least geometry data, attribute data, or occupancy map data, wherein the bitstream is composed of first precision information and first units, wherein each of the first units is composed of first size information and second units, wherein the first precision information comprises information for specifying precision of the first size information in the first units, wherein the first size information comprises information for specifying size of the second units, wherein the second units are composed of a header and a payload, wherein the header comprises type information for indicating type of data in the payload, wherein the payload comprises one of the geometry data, the attribute data, the occupancy map data, and atlas data, wherein the atlas data is composed of second precision information and third units, wherein each of the third units is composed of second size information and fourth units, wherein the second precision information comprises information for specifying precision of the second size information in the third units, wherein the second size information comprises information for specifying size of the fourth units, wherein the fourth units comprise an atlas frame parameter set comprising first tile identification information for identifying each of one or more tiles in an atlas frame, wherein the signaling data comprises spatial region information of the point cloud data, wherein the spatial region information is at least static spatial region information that does not vary over time or dynamic spatial region information that varies over time, wherein the point cloud data is partitioned into one or more 3-dimensional (3D) spatial regions, and wherein the static spatial region information comprises region number information for identifying a number of the one or more 3D spatial regions, region identification information for identifying each 3D spatial region, tile number information for identifying a number of one or more tiles associated with each 3D spatial region, and second tile identification information for identifying each of the one or more tiles associated with each 3D spatial region.

8. The point cloud data reception method of claim 7, wherein, a value of the first tile identification information is equal to a value of the second tile identification information.

9. The point cloud data reception method of claim 7, wherein, the dynamic spatial region information comprises region number information for identifying a number of the one or more 3D spatial regions, region identification information for identifying each 3D spatial region, and priority information and dependency information related to a 3D spatial region.

10. A point cloud data reception apparatus comprising: a receiver that receives a file, a decoder that decodes the point cloud data based on the signaling data, a decapsulator that decapsulates the file into a bitstream comprising point cloud data, wherein the bitstream is stored in a plurality of tracks of the file, and wherein the file further comprises signaling data; and a decoder that decodes the point cloud data based on the signaling data; wherein the point cloud data comprises at least geometry data, attribute data, or occupancy map data, wherein the bitstream is composed of first precision information and first units, wherein each of the first units is composed of first size information and second units, wherein the first precision information comprises information for specifying precision of the first size information in the first units, wherein the first size information comprises information for specifying size of the second units, wherein the second units are composed of a header and a payload, wherein the header comprises type information for indicating type of data in the payload, wherein the payload comprises one of the geometry data, the attribute data, the occupancy map data, and atlas data, wherein the atlas data is composed of second precision information and third units, wherein each of the third units is composed of second size information and fourth units, wherein the second precision information comprises information for specifying precision of the second size information in the third units, wherein the second size information comprises information for specifying size of the fourth units, wherein the fourth units comprise atlas frame parameter sets comprising first tile identification information for identifying each of one or more tiles in an atlas frame, wherein the signaling data comprises spatial region information of the point cloud data, wherein the spatial region information is at least static spatial region information that does not vary over time or dynamic spatial region information that varies over time, wherein the point cloud data is partitioned into one or more 3-dimensional, 3D, spatial regions, and wherein the static spatial region information comprises region number information for identifying a number of the one or more 3D spatial regions, region identification information for identifying each 3D spatial region, tile number information for identifying a number of one or more tiles associated with each 3D spatial region, and second tile identification information for identifying each of the one or more tiles associated with each 3D spatial region.

11. The point cloud data reception apparatus according to claim 10, wherein, a value of the first tile identification information is equal to a value of the second tile identification information.

12. The point cloud data reception apparatus according to claim 10, wherein, the dynamic spatial region information comprises region number information for identifying a number of the one or more 3D spatial regions, region identification information for identifying each 3D spatial region, and priority information and dependency information related to a 3D spatial region.