Point cloud data transmitting device and method, point cloud data receiving device and method

CN116684666BActive Publication Date: 2026-08-14LG ELECTRONICS INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-23
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

因为3D空间中的点的数量大,所以难以生成点云数据

Benefits of technology

[0013]The point cloud data transmission method, point cloud data transmission device, point cloud data reception method, and point cloud data reception device according to the embodiments can provide high-quality point cloud services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116684666B_ABST
    Figure CN116684666B_ABST
Patent Text Reader

Abstract

A point cloud data transmitting apparatus and method, and a point cloud data receiving apparatus and method are disclosed. The point cloud data transmitting method according to an embodiment includes the following steps: encoding point cloud data; encapsulating point cloud data; and transmitting the point cloud data. The point cloud data receiving apparatus according to an embodiment includes: a receiver for receiving point cloud data; a decapsulator for decapsulating the point cloud data; and a decoder for decoding the point cloud data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the original invention patent application No. 202080092244.0 (International Application No.: PCT / KR2020 / 012847, Application Date: September 23, 2020, Invention Title: Point Cloud Data Transmitting Device, Point Cloud Data Transmitting Method, Point Cloud Data Receiving Device and Point Cloud Data Receiving Method). Technical Field

[0002] The implementation provides a method for providing point cloud content to offer users various services such as virtual reality (VR), augmented reality (AR), mixed reality (MR), and autonomous driving services. Background Technology

[0003] A point cloud is a collection of points in three-dimensional (3D) space. Because of the large number of points in 3D space, it is difficult to generate point cloud data.

[0004] High throughput is required to send and receive point cloud data. Summary of the Invention

[0005] Technical issues

[0006] The purpose of this disclosure is to provide a point cloud data transmitting device, a point cloud data transmitting method, a point cloud data receiving device, and a point cloud data receiving method for effectively transmitting and receiving point clouds.

[0007] Another objective of this disclosure is to provide a point cloud data transmission device, a point cloud data transmission method, a point cloud data receiving device, and a point cloud data receiving method for addressing latency and encoding / decoding complexity.

[0008] The implementation methods are not limited to the above objectives, and the scope of the implementation methods can be extended to other objectives that can be inferred by those skilled in the art based on the entire contents of this disclosure.

[0009] Technical solution

[0010] To achieve these objectives and other advantages, and in one aspect of this disclosure, a method for transmitting point cloud data may include the steps of: encoding the point cloud data; encapsulating the point cloud data; and transmitting the point cloud data.

[0011] In another aspect of this disclosure, an apparatus for receiving point cloud data may include: a receiver configured to receive point cloud data; a decapsulator configured to decapsulate the point cloud data; and a decoder configured to decode the point cloud data.

[0012] Beneficial effects

[0013] The point cloud data transmission method, point cloud data transmission device, point cloud data reception method, and point cloud data reception device according to the embodiments can provide high-quality point cloud services.

[0014] The point cloud data transmission method, point cloud data transmission device, point cloud data reception method, and point cloud data reception device according to the embodiments can implement various video encoding and decoding methods.

[0015] The point cloud data transmission method, point cloud data transmission device, point cloud data reception method, and point cloud data reception device according to the embodiments can provide general point cloud content such as autonomous driving services. Attached Figure Description

[0016] The accompanying drawings are included to provide a further understanding of this disclosure and are incorporated in and constitute a part of this application. The drawings illustrate embodiments of the disclosure and, together with the description, serve to illustrate the principles of the disclosure. In the drawings:

[0017] Figure 1 An exemplary structure of a sending / receiving system for providing point cloud content according to an embodiment is shown.

[0018] Figure 2 The capture of point cloud data according to an embodiment is shown.

[0019] Figure 3 Exemplary point cloud, geometry, and texture images are shown according to an embodiment.

[0020] Figure 4 An exemplary V-PCC encoding process according to an implementation is shown.

[0021] Figure 5 An example of the tangent plane and normal vector of a surface according to an embodiment is shown.

[0022] Figure 6 An exemplary bounding box of a point cloud according to an implementation method is shown.

[0023] Figure 7 An example of determining the position of each patch on the occupancy map according to an embodiment is shown.

[0024] Figure 8 An exemplary relationship between the normal axis, tangential axis, and double tangential axis is shown according to an embodiment.

[0025] Figure 9 Exemplary configurations of the minimum and maximum modes of the projection mode according to the implementation are shown.

[0026] Figure 10 An exemplary EDD code according to an implementation method is shown.

[0027] Figure 11 An example of recoloring based on the color values ​​of neighboring points according to an implementation method is shown.

[0028] Figure 12 An example of a push-pull background fill according to an implementation method is shown.

[0029] Figure 13 An exemplary possible traversal order of a 4x4 block according to an implementation is shown.

[0030] Figure 14 An exemplary optimal traversal order is shown according to the implementation method.

[0031] Figure 15 An exemplary 2D video / image encoder according to an embodiment is shown.

[0032] Figure 16 An exemplary V-PCC decoding process according to an implementation is shown.

[0033] Figure 17 An exemplary 2D video / image decoder according to an embodiment is shown.

[0034] Figure 18 This is a flowchart illustrating the operation of a transmitting device according to an embodiment of the present disclosure.

[0035] Figure 19 This is a flowchart illustrating the operation of the receiving device according to an embodiment.

[0036] Figure 20 An exemplary architecture for V-PCC-based storage and streaming of point cloud data, according to an embodiment, is shown.

[0037] Figure 21 This is an exemplary block diagram of an apparatus for storing and transmitting point cloud data according to an embodiment.

[0038] Figure 22 This is an exemplary block diagram of a point cloud data receiving device according to an embodiment.

[0039] Figure 23 An exemplary structure is shown that can be operated in conjunction with a point cloud data transmission / reception method / apparatus according to an embodiment.

[0040] Figure 24 The structure of the encapsulated V-PCC data container according to an embodiment is shown.

[0041] Figure 25 The encapsulated V-PCC data container structure according to an embodiment is shown.

[0042] Figure 26 The structure of a bitstream containing point cloud data according to an embodiment is shown.

[0043] Figure 27 The configuration of the sample stream V-PCC unit according to an embodiment is shown.

[0044] Figure 28 The V-PCC unit and V-PCC unit header according to an embodiment are shown.

[0045] Figure 29 The payload of a V-PCC unit according to an embodiment is shown.

[0046] Figure 30 The V-PCC parameter set according to the implementation method is shown.

[0047] Figure 31 The diagram shows the tiles according to the implementation method.

[0048] Figure 32 The structure of the atlas bitstream according to an embodiment is shown.

[0049] Figure 33 The NAL unit according to an embodiment is shown.

[0050] Figure 34 The type of NAL unit according to the implementation is shown.

[0051] Figure 35 The atlas sequence parameter set according to the implementation method is shown.

[0052] Figure 36 The atlas frame parameter set according to the implementation method is shown.

[0053] Figure 37 The atlas_frame_tile_information is shown according to the implementation method.

[0054] Figure 38 The atlas adaptation parameter set (atlas_adaptation_parameter_set_rbsp()) according to the implementation method is shown.

[0055] Figure 39 The atlas_camera_parameters are shown according to the implementation method.

[0056] Figure 40 The atlas_tile_group_layer and atlas_tile_group_header according to the implementation are shown.

[0057] Figure 41 The reference list structure (ref_list_struct) according to the implementation method is shown.

[0058] Figure 42 This shows the atlas tile group data (atlas_tile_group_data_unit) according to the implementation method.

[0059] Figure 43 The patch information data (patch_information_data) according to the implementation method is shown.

[0060] Figure 44 The patch_data_unit according to the implementation is shown.

[0061] Figure 45 The rotation and offset relative to the patch orientation are shown according to the embodiment.

[0062] Figure 46 The scene object information (scene_object_information) is shown according to the implementation method.

[0063] Figure 47 The object label information is shown according to the implementation method.

[0064] Figure 48 Information about the patch is shown according to an embodiment.

[0065] Figure 49 Information about the volume rectangle according to the embodiment is shown.

[0066] Figure 50 The configuration of the sample stream vpcc unit according to an embodiment is shown.

[0067] Figure 51 This illustrates the configuration of atlas tile groups (or tiles) according to an embodiment.

[0068] Figure 52 The structure of the V-PCC spatial region box according to an embodiment is shown.

[0069] Figure 53 The DynamicSpatialRegionSample is shown according to an implementation method.

[0070] Figure 54 The diagram illustrates a structure for encapsulating non-timing V-PCC data according to an embodiment.

[0071] Figure 55 This is a flowchart of a point cloud data transmission method according to an implementation method.

[0072] Figure 56 This is a flowchart of a point cloud data receiving method according to an implementation method.

[0073] Figure 57The file format structure according to the implementation method is shown.

[0074] Figure 58 The document-level signaling according to the implementation method is shown.

[0075] Figure 59 This illustrates the relationship between 3D regions of a point cloud and regions in a video frame according to an embodiment.

[0076] Figure 60 The parameter set according to the implementation method is shown.

[0077] Figure 61 The Atlas Sequence Parameter Set (ASPS) according to the implementation method is shown.

[0078] Figure 62 The Atlas Frame Parameter Set (AFPS) according to the implementation method is shown.

[0079] Figure 63 The atlas frame tile information (atlas_frame_tile_information) is shown according to the implementation method.

[0080] Figure 64 Supplemental Enhancement Information (SEI) according to the implementation method is shown.

[0081] Figure 65 The 3D bounding box SEI according to the embodiment is shown.

[0082] Figure 66 The 3D region mapping information SEI message according to the implementation method is shown.

[0083] Figure 67 The volume tiling information according to the embodiment is shown.

[0084] Figure 68 The volume tiling information object is shown according to the implementation method.

[0085] Figure 69 The volume tiling information label is shown according to the implementation method.

[0086] Figure 70 A sample entry for V-PCC according to an implementation method is shown.

[0087] Figure 71 The track replacement and grouping are shown according to the implementation method.

[0088] Figure 72 The structure of V-PCC 3D region mapping information according to an embodiment is shown.

[0089] Figure 73 The structure of a bitstream according to an embodiment is shown.

[0090] Figure 74 A transmission method according to an embodiment is shown.

[0091] Figure 75 A receiving method according to an embodiment is shown. Detailed Implementation

[0092] Preferred embodiments of the present disclosure will now be described in detail, examples of which are illustrated in the accompanying drawings. The detailed description given below with reference to the drawings is intended to illustrate exemplary embodiments of the present disclosure, and not to show only embodiments that can be implemented according to the present disclosure. The following detailed description includes specific details in order to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure can be practiced without these specific details.

[0093] Although most of the terms used in this disclosure are selected from general terms widely used in the art, some terms are arbitrarily chosen by the applicant, and their meanings are explained in detail as needed in the following description. Therefore, this disclosure should be understood based on the intended meaning of the terms rather than their simple names or meanings.

[0094] Figure 1 An exemplary structure of a sending / receiving system for providing point cloud content according to an embodiment is shown.

[0095] This disclosure provides a method for providing point cloud content to offer users various services such as virtual reality (VR), augmented reality (AR), mixed reality (MR), and autonomous driving. According to the embodiments, the point cloud content represents data representing objects as points and may be referred to as point cloud, point cloud data, point cloud video data, point cloud image data, etc.

[0096] The point cloud data transmission device 10000 according to an embodiment may include a point cloud video acquirer 10001, a point cloud video encoder 10002, a file / fragment encapsulation module 10003, and / or a transmitter (or communication module) 10004. The transmission device according to an embodiment can secure and process point cloud video (or point cloud content) and transmit it. According to an embodiment, the transmission device may include a fixed station, a base transceiver system (BTS), a network, an artificial intelligence (AI) device and / or system, a robot, and an AR / VR / XR device and / or server. According to an embodiment, the transmission device 10000 may include a device robot, vehicle, AR / VR / XR device, portable device, home appliance, Internet of Things (IoT) device, and AI device / server configured to communicate with a base station and / or other wireless devices using radio access technologies (e.g., 5G New RAT (NR), Long Term Evolution (LTE)).

[0097] The point cloud video acquirer 10001 according to the embodiment acquires point cloud video through processes of capturing, synthesizing or generating point cloud video.

[0098] The point cloud video encoder 10002 according to the embodiment encodes point cloud video data. According to the embodiment, the point cloud video encoder 10002 may be referred to as a point cloud encoder, point cloud data encoder, encoder, etc. The point cloud compression encoding (encoding) according to the embodiment is not limited to the above embodiment. The point cloud video encoder can output a bitstream containing encoded point cloud video data. The bitstream may include not only the encoded point cloud video data, but also signaling information related to the encoding of the point cloud video data.

[0099] The encoder according to the embodiment can support both geometry-based point cloud compression (G-PCC) encoding schemes and / or video-based point cloud compression (V-PCC) encoding schemes. Furthermore, the encoder can encode point clouds (referring to point cloud data or points) and / or signaling data associated with point clouds. Specific encoding operations according to the embodiment will be described below.

[0100] As used herein, the term V-PCC may stand for Video-Based Point Cloud Compression (V-PCC). The term V-PCC may be synonymous with Visual Volumetric Video Coding (V3C). These terms may be used complementaryly.

[0101] The file / fragment encapsulation module 10003 according to the embodiment encapsulates point cloud data in the form of files and / or fragments. The point cloud data transmission method / apparatus according to the embodiment can transmit point cloud data in the form of files and / or fragments.

[0102] According to the embodiment, the transmitter (or communication module) 10004 transmits encoded point cloud video data in the form of a bitstream. According to the embodiment, files or segments can be transmitted to a receiving device via a network or stored in a digital storage medium (e.g., USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.). The transmitter according to the embodiment is capable of wired / wireless communication with the receiving device (or receiver) via a 4G, 5G, 6G, or other network. Furthermore, the transmitter can perform necessary data processing operations according to the network system (e.g., a 4G, 5G, or 6G communication network system). The transmitting device can transmit encapsulated data on demand.

[0103] The point cloud data receiving device 10005 according to an embodiment may include a receiver 10006, a file / fragment decapsulation module 10007, a point cloud video decoder 10008, and / or a renderer 10009. According to an embodiment, the receiving device may include a device robot, vehicle, AR / VR / XR device, portable device, home appliance, Internet of Things (IoT) device, and AI device / server, configured to communicate with base stations and / or other wireless devices using radio access technologies (e.g., 5G New RAT (NR), Long Term Evolution (LTE)).

[0104] According to an embodiment, receiver 10006 receives a bitstream containing point cloud video data. According to an embodiment, receiver 10006 can send feedback information to point cloud data transmitting device 10000.

[0105] The file / fragment decapsulation module 10007 decapsulates files and / or fragments containing point cloud data. According to the embodiment, the decapsulation module can perform the reverse processing of the capsulation process according to the embodiment.

[0106] The point cloud video decoder 10008 decodes the received point cloud video data. The decoder according to the embodiment can perform inverse processing of the encoding according to the embodiment.

[0107] Renderer 10009 renders decoded point cloud video data. According to one embodiment, renderer 10009 can send feedback information obtained at the receiving side to point cloud video decoder 10008. The point cloud video data, according to one embodiment, can carry feedback information to the receiver. According to one embodiment, the feedback information received by the point cloud transmitting device can be provided to the point cloud video encoder.

[0108] The arrows indicated by the dashed lines in the figure represent the transmission paths of the feedback information acquired by the receiving device 10005. The feedback information reflects the interactivity of the user consuming the point cloud content and includes user information (e.g., head orientation information, viewport information, etc.). Specifically, when the point cloud content is the content of a service requiring user interaction (e.g., autonomous driving service, etc.), the feedback information can be provided to the content sending side (e.g., sending device 10000) and / or the service provider. According to embodiments, the feedback information can be used in both the receiving device 10005 and the sending device 10000, and may not be provided at all.

[0109] According to the embodiment, head orientation information is information about the user's head position, orientation, angle, movement, etc. The receiving device 10005 according to the embodiment can calculate viewport information based on the head orientation information. The viewport information can be information about the point cloud video region that the user is viewing. The viewpoint is the point in the point cloud video that the user is viewing, and can refer to the center point of the viewport region. That is, the viewport is the region centered on the viewpoint, and the size and shape of the region can be determined by the field of view (FOV). Therefore, in addition to head orientation information, the receiving device 10005 can extract viewport information based on the vertical or horizontal FOV supported by the device. Furthermore, the receiving device 10005 performs gaze analysis to examine how the user consumes the point cloud, the region the user gazes at in the point cloud video, the gaze duration, etc. According to the embodiment, the receiving device 10005 can send feedback information including the gaze analysis results to the transmitting device 10000. The feedback information according to the embodiment can be obtained during rendering and / or display processing. The feedback information according to the embodiment can be ensured by one or more sensors included in the receiving device 10005. In addition, according to the implementation, feedback information can be ensured by the renderer 10009 or by a separate external component (or device, assembly, etc.). Figure 1 The dashed lines in the diagram represent the processing of feedback information ensured by the sender renderer 10009. The point cloud content providing system can process (encode / decode) point cloud data based on the feedback information. Therefore, the point cloud video decoder 10008 can perform decoding operations based on the feedback information. The receiving device 10005 can send feedback information to the sending device. The sending device (or the point cloud video data encoder 10002) can perform encoding operations based on the feedback information. Therefore, the point cloud content providing system can effectively process necessary data (e.g., point cloud data corresponding to the user's head position) based on the feedback information, rather than processing (encoding / decoding) all point cloud data, and provide point cloud content to the user.

[0110] According to the implementation, the transmitting device 10000 may be referred to as an encoder, transmitting device, transmitter, etc., and the receiving device 10005 may be referred to as a decoder, receiving device, receiver, etc.

[0111] According to the implementation method Figure 1 The point cloud data processed in the point cloud content provision system (through a series of processes including acquisition, encoding, transmission, decoding, and rendering) can be referred to as point cloud content data or point cloud video data. Depending on the implementation, point cloud content data can be used as a concept encompassing metadata or signaling information related to point cloud data.

[0112] Figure 1 The components of the point cloud content provided by the system can be implemented by hardware, software, processors, and / or combinations thereof.

[0113] Implementations may provide a method for providing point cloud content to offer users various services such as virtual reality (VR), augmented reality (AR), mixed reality (MR), and autonomous driving.

[0114] To provide point cloud content services, point cloud video can first be acquired. The acquired point cloud video can be sent through a series of processes, and the receiving side can process the received data back into the original point cloud video and render the processed point cloud video. Thus, the point cloud video can be provided to the user. This implementation provides a method for efficiently performing this series of processes.

[0115] All processing used to provide point cloud content services (point cloud data sending methods and / or point cloud data receiving methods) may include acquisition processing, encoding processing, transmission processing, decoding processing, rendering processing, and / or feedback processing.

[0116] According to an implementation, the processing of providing point cloud content (or point cloud data) can be referred to as point cloud compression processing. According to an implementation, point cloud compression processing can represent geometry-based point cloud compression processing.

[0117] The individual components of the point cloud data transmitting device and the point cloud data receiving device according to the embodiments may be hardware, software, processor and / or combinations thereof.

[0118] To provide point cloud content services, point cloud video can be acquired. The acquired point cloud video is sent after a series of processing steps, and the receiving side can process the received data back into the original point cloud video and render the processed point cloud video. Thus, the point cloud video can be provided to the user. This implementation provides a method for efficiently performing this series of processing steps.

[0119] All processing used to provide point cloud content services may include acquisition processing, encoding processing, transmission processing, decoding processing, rendering processing, and / or feedback processing.

[0120] A point cloud compression system may include a transmitting device and a receiving device. The transmitting device can encode the point cloud video to output a bitstream and transmit it to the receiving device as a file or stream (streaming segment) via digital storage media or a network. Digital storage media may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD.

[0121] The transmitting device may include a point cloud video acquirer, a point cloud video encoder, a file / segment encapsulator, and a transmitter. The receiving device may include a receiver, a file / segment decapsulator, a point cloud video decoder, and a renderer. The encoder may be referred to as a point cloud video / image / image / frame encoder, and the decoder may be referred to as a point cloud video / image / image / frame decoder. The transmitter may be included in the point cloud video encoder. The receiver may be included in the point cloud video decoder. The renderer may include a display. The renderer and / or the display may be configured as separate devices or external components. The transmitting and receiving devices may also include separate internal or external modules / units / components for feedback processing.

[0122] According to the implementation method, the operation of the receiving device can be the reverse processing of the operation of the transmitting device.

[0123] Point cloud video acquirers can perform point cloud video acquisition processing by capturing, orchestrating, or generating point cloud video. During acquisition processing, 3D position (x, y, z) / attribute (color, reflectivity, transparency, etc.) data for multiple points can be generated, such as polygon file format (PLY) (or Stanford triangle format) files. For videos with multiple frames, one or more files can be acquired. Point cloud-related metadata (e.g., capture-related metadata) can be generated during capture processing.

[0124] The point cloud data transmission apparatus according to the embodiments may include an encoder configured to encode point cloud data and a transmitter configured to transmit point cloud data. The data may be transmitted in the form of a bitstream containing point clouds.

[0125] The point cloud data receiving apparatus according to the embodiments may include a receiver configured to receive point cloud data, a decoder configured to decode point cloud data, and a renderer configured to render point cloud data.

[0126] The method / apparatus described in the embodiments represents a point cloud data transmitting device and / or a point cloud data receiving device.

[0127] Figure 2 The capture of point cloud data according to an embodiment is shown.

[0128] Point cloud data according to the embodiments can be acquired by cameras or the like. The capture techniques according to the embodiments may include, for example, inward-facing and / or outward-facing.

[0129] In the inward orientation according to the implementation, one or more cameras facing the point cloud data of the object can capture images of the object from outside the object.

[0130] In the outward-facing configuration according to the embodiment, one or more cameras can capture the object of the point cloud data. For example, according to the embodiment, four cameras may be present.

[0131] According to the embodiments, point cloud data or point cloud content can be video or still images of objects / environments represented in various types of 3D space. According to the embodiments, point cloud content can include video / audio / images of objects.

[0132] To capture point cloud content, a combination of a camera device capable of acquiring depth (a combination of an infrared pattern projector and an infrared camera) and an RGB camera capable of extracting color information corresponding to the depth information can be configured. Alternatively, depth information can be extracted using a LiDAR radar system that measures the position coordinates of a reflector by emitting laser pulses and measuring their return time. The geometry composed of points in 3D space can be extracted from the depth information, and attributes representing the color / reflectivity of each point can be extracted from the RGB information. Point cloud content can include information about position (x, y, z) and the color (YCbCr or RGB) or reflectivity (r) of the points. For point cloud content, outward-facing techniques for capturing the external environment and inward-facing techniques for capturing the central object can be used. In VR / AR environments, when an object (e.g., a core object such as a character, player, thing, or actor) is configured within point cloud content that the user can view from any direction (360 degrees), the configuration of the capture camera can be based on inward-facing techniques. When the current surrounding environment is configured within the point cloud content in vehicle modes such as autonomous driving, the configuration of the capture camera can be based on outward-facing techniques. Since point cloud content can be captured by multiple cameras, camera calibration may be required to configure the camera's global coordinate system before capturing the content.

[0133] Point cloud content can be video or still images of objects / environments existing in various types of 3D space.

[0134] Furthermore, in point cloud content acquisition methods, any point cloud video can be orchestrated based on the captured point cloud video. Alternatively, when providing point cloud video of a computer-generated virtual space, capture using an actual camera may not be performed. In this case, the capture processing can be simply replaced by processing that generates relevant data.

[0135] Post-processing of captured point cloud video may be necessary to improve content quality. During video capture processing, the maximum / minimum depth can be adjusted within the range provided by the camera device. Even after adjustment, unwanted areas of point data may still exist. Therefore, post-processing can be performed to remove unwanted areas (e.g., background) or to identify connected spaces and fill in spatial holes. Additionally, point clouds extracted from cameras in a shared spatial coordinate system can be integrated into a single piece of content by transforming individual points to a global coordinate system based on the position coordinates of each camera obtained through calibration processing. This can generate a single point cloud content with a wide range, or it can acquire point cloud content with high-density points.

[0136] A point cloud video encoder can encode an input point cloud video into one or more video streams. A video may include multiple frames, each frame corresponding to a still image / picture. In this specification, point cloud video may include point cloud images / frames / pictures / video / audio. Additionally, the term "cloud video" can perform video-based point cloud compression (V-PCC) processing. For compression and encoding efficiency, the point cloud video encoder can perform a series of processes such as prediction, transform, quantization, and entropy coding. The encoded data (encoded video / image information) can be output as a bitstream. Based on V-PCC processing, the point cloud video encoder can encode the point cloud video by dividing it into geometric video, attribute video, occupancy map video, and auxiliary information (described later). Geometric video may include geometric images, attribute video may include attribute images, and occupancy map video may include occupancy map images. Auxiliary information may include auxiliary patch information. Attribute video / images may include texture video / images.

[0137] The encapsulation processor (file / fragment encapsulation module) 10003 can encapsulate encoded point cloud video data and / or metadata related to the point cloud video in, for example, file format. Here, the metadata related to the point cloud video can be received from a metadata processor. The metadata processor can be included in the point cloud video encoder or configured as a separate component / module. The encapsulation processor can encapsulate data in a file format such as ISOBMFF or process data in the form of DASH fragments, etc. According to an embodiment, the encapsulation processor can include point cloud video-related metadata in a file format. The point cloud video metadata can be included in various levels of frames in, for example, ISOBMFF file format, or as data in a separate track within a file. According to an embodiment, the encapsulation processor can encapsulate point cloud video-related metadata into a file. The transmission processor can perform transmission processing on the point cloud video data encapsulated according to the file format. The transmission processor can be included in a transmitter or configured as a separate component / module. The transmission processor can process the point cloud video data according to a transmission protocol. The transmission processing can include processing for transmission via a broadcast network and processing for transmission via broadband. According to the implementation method, the transmitting processor can receive point cloud video-related metadata from the metadata processor along with the point cloud video data, and perform point cloud video data processing for transmission.

[0138] Transmitter 10004 can transmit encoded video / image information or data, output in bitstream form, to receiver of receiving device in the form of a file or stream via digital storage medium or network. Digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitter may include elements for generating media files in a predetermined file format and may include elements for transmission via broadcast / communication network. Receiver can extract the bitstream and send the extracted bitstream to decoding device.

[0139] Receiver 10006 can receive point cloud video data transmitted by the point cloud video transmitting device according to this disclosure. Depending on the transmission channel, the receiver can receive the point cloud video data via a broadcast network or via broadband. Alternatively, the point cloud video data can be received via a digital storage medium.

[0140] The receiving processor can process the received point cloud video data according to the transmission protocol. The receiving processor can be included in the receiver or configured as a separate component / module. The receiving processor can reverse the above-described processing of the transmitting processor, such that the processing corresponds to the transmission processing performed on the transmitting side. The receiving processor can transmit the acquired point cloud video data to the decapsulation processor and transmit the acquired point cloud video-related metadata to the metadata parser. The point cloud video-related metadata acquired by the receiving processor can be in the form of a signaling table.

[0141] The decapsulation processor (file / fragment decapsulation module) 10007 decapsulates point cloud video data received from the receiving processor in file form. The decapsulation processor can decapsulate files according to ISOBMFF, etc., and can acquire point cloud video bitstreams or point cloud video-related metadata (metadata bitstreams). The acquired point cloud video bitstream can be transmitted to the point cloud video decoder, and the acquired point cloud video-related metadata (metadata bitstreams) can be transmitted to the metadata processor. The point cloud video bitstream may include metadata (metadata bitstreams). The metadata processor may be included in the point cloud video decoder or can be configured as a separate component / module. The point cloud video-related metadata acquired by the decapsulation processor may take the form of boxes or tracks in a file format. When needed, the decapsulation processor can receive the metadata required for decapsulation from the metadata processor. The point cloud video-related metadata may be transmitted to the point cloud video decoder and used in the point cloud video decoding process, or it may be transmitted to the renderer and used in the point cloud video rendering process.

[0142] A point cloud video decoder can receive a bitstream and decode the video / image by performing operations corresponding to those of a point cloud video encoder. In this case, the point cloud video decoder can decode the point cloud video by dividing it into geometric video, attribute video, occupancy map video, and auxiliary information, as described below. Geometric video may include geometric images, and attribute video may include attribute images. Occupancy map video may include occupancy map images. Auxiliary information may include auxiliary patch information. Attribute video / images may include texture video / images.

[0143] 3D geometry can be reconstructed based on decoded geometric images, occupancy maps, and auxiliary patch information, and then subjected to smoothing. Color point cloud images / pictures can be reconstructed by assigning color values ​​to the smoothed 3D geometry based on texture images. The renderer can render the reconstructed geometry and color point cloud images / pictures. The rendered video / images can be displayed on a monitor. Users can view all or part of the rendering results through VR / AR displays or typical monitors.

[0144] Feedback processing may include transmitting various types of feedback information, which can be obtained during rendering / display processing, to a decoder on the sending or receiving side. Interactivity can be provided through feedback processing when consuming point cloud video. According to one embodiment, head orientation information, viewport information indicating the area the user is currently viewing, etc., may be transmitted to the sending side during feedback processing. According to another embodiment, the user may interact with objects realized in a VR / AR / MR / autonomous driving environment. In this case, information related to the interaction may be transmitted to the sending side or service provider during feedback processing. According to yet another embodiment, feedback processing may be skipped.

[0145] Head orientation information represents the position, angle, and movement of the user's head. Based on this information, information about the region of the point cloud video currently being viewed by the user (i.e., viewport information) can be calculated.

[0146] Viewport information can be information about the region of a point cloud video currently being viewed by the user. Viewport information can be used to perform gaze analysis to examine how the user consumes the point cloud video, the region of the point cloud video the user is gazing at, and how long the user is gazing at that region. Gaze analysis can be performed on the receiving side, and the analysis results can be transmitted to the transmitting side via a feedback channel. Devices such as VR / AR / MR displays can extract the viewport region based on the user's head position / orientation, the vertical or horizontal FOV supported by the device, etc.

[0147] According to the implementation method, the aforementioned feedback information can be transmitted not only to the sending side but also consumed at the receiving side. That is, decoding and rendering processing at the receiving side can be performed based on the aforementioned feedback information. For example, point cloud video of the area currently being viewed by the user can be decoded and rendered preferentially only based on head orientation information and / or viewport information.

[0148] Here, the viewport or viewport region can represent the area of ​​the point cloud video currently being viewed by the user. The viewpoint is the point in the point cloud video that the user is viewing, and can represent the center point of the viewport region. That is, the viewport is the area surrounding the viewpoint, and the size and shape of the area can be determined by the field of view (FOV).

[0149] This disclosure relates to point cloud video compression as described above. For example, the methods / implementations disclosed in this disclosure can be applied to the Moving Picture Experts Group (MPEG) point cloud compression or point cloud coding (PCC) standard or next-generation video / image coding standards.

[0150] As used in this article, a picture / frame can typically represent a unit representing an image within a specific time interval.

[0151] A pixel, or image unit, can be the smallest unit that makes up a picture (or image). Additionally, the term "sample" can be used as the counterpart to a pixel. A sample can typically represent a pixel or pixel value, or it can represent only the pixel / pixel value of the luminance component, only the pixel / pixel value of the chrominance component, or only the pixel / pixel value of the depth component.

[0152] A unit can represent a basic unit of image processing. A unit may include a specific region of an image and at least one of the information associated with that region. In some cases, a unit may be used interchangeably with terms such as block or region. In general, an M×N block may include samples (or sample arrays) or a set (or array) of transform coefficients arranged in M ​​columns and N rows.

[0153] Figure 3Examples of point clouds, geometric images, and texture images according to embodiments are shown.

[0154] The point cloud, according to the implementation method, can be input into what will be described later. Figure 4 The V-PCC encoding process generates geometric and texture images. Depending on the implementation, the point cloud may have the same meaning as the point cloud data.

[0155] As shown in the figure, the left side shows a point cloud, where objects are located in 3D space and can be represented by bounding boxes, etc. The middle part shows the geometry, and the right side shows a texture image (non-filled image).

[0156] Video-based point cloud compression (V-PCC) according to the implementation provides a method for compressing 3D point cloud data based on 2D video codecs such as HEVC or VVC. The data and information generated in the V-PCC compression process are as follows:

[0157] Occupancy Map: This is a binary map that uses values ​​of 0 or 1 to indicate the presence of data at corresponding locations in a 2D plane when the points constituting a point cloud are divided into patches and mapped onto a 2D plane. An occupancy map can represent a 2D array corresponding to an atlas, and its values ​​indicate whether each sample location in the atlas corresponds to a 3D point. An atlas refers to an object that includes information about the 2D patches of individual point cloud frames. For example, an atlas may include the 2D arrangement and size of the patches, the location of the corresponding 3D regions within 3D points, the projection plane, and the level of detail parameters.

[0158] An atlas is a collection of 2D bounding boxes corresponding to 3D bounding boxes in the 3D space of the rendered volume data, located within a rectangular frame, along with related information.

[0159] A map atlas bitstream is a bitstream of one or more map atlas frames and associated data that make up a map atlas.

[0160] An atlas frame is a 2D rectangular array of atlas samples onto which a patch is projected.

[0161] A map sample is the position of a rectangular frame onto which a patch associated with the map is projected.

[0162] Atlas frames can be divided into tiles. A tile is a unit in which a 2D frame is divided. That is, a tile is a unit used to divide the signaling information of point cloud data called an atlas.

[0163] Patch: A set of points that make up a point cloud, indicating that points belonging to the same patch are adjacent to each other in 3D space and are mapped in the same direction between 6 bounding box planes in the process of mapping to a 2D image.

[0164] A patch is a unit into which a mosaic is divided. A patch contains signaling information about the configuration of point cloud data.

[0165] According to the implementation method, the receiving device can recover attribute video data, geometric video data, and occupied video data (actual video data with the same presentation time) based on atlases (tiles, patches).

[0166] Geometric Image: This is a depth map-like image that presents the positional information (geometric) of the individual points that make up the point cloud patch by patch. A geometric image can consist of pixel values ​​from a single channel. The geometric representation is the set of coordinates associated with a point cloud frame.

[0167] Texture image: This is an image that represents color information about the individual points that make up a point cloud, patch by patch. A texture image may consist of pixel values ​​from multiple channels (e.g., R, G, and B channels). Texture is included in attributes. Depending on the implementation, texture and / or attributes may be interpreted as the same object and / or have an inclusion relationship.

[0168] Auxiliary patch information: This indicates the metadata required to reconstruct the point cloud using the individual patches. Auxiliary patch information may include information about the location, size, etc. of the patches in 2D / 3D space.

[0169] Point cloud data (e.g., V-PCC components) according to the implementation may include atlases, accuracy maps, geometry, and attributes.

[0170] An atlas represents a collection of 2D bounding boxes. It can be a patch, for example, a patch projected onto a rectangular frame. An atlas can correspond to 3D bounding boxes in 3D space and can represent a subset of a point cloud.

[0171] Attributes can represent scalars or vectors associated with individual points in a point cloud. For example, attributes can include color, reflectivity, surface normal, timestamp, and material ID.

[0172] The point cloud data according to the implementation represents PCC data based on a video-based point cloud compression (V-PCC) scheme. The point cloud data may include multiple components. For example, it may include occupancy maps, patches, geometry, and / or textures.

[0173] Figure 4 The V-PCC encoding process according to the implementation method is shown.

[0174] This diagram illustrates the V-PCC encoding process used to generate and compress occupancy maps, geometric images, texture images, and auxiliary patch information. Figure 4 V-PCC encoding processing can be performed by Figure 1 The point cloud video encoder 10002 is used for processing. Figure 4 The various components can be executed by software, hardware, processors, and / or combinations thereof.

[0175] The patch generator 40000 receives point cloud frames (which can be in the form of a bitstream containing point cloud data). The patch generator 40000 generates patches from the point cloud data. Additionally, patch information including information about the patch generation is generated.

[0176] Patch packing, or patch packer 40001, is used for patch packing of point cloud data. For example, one or more patches can be packed. Additionally, the patch packer generates an occupancy map containing information about the patch packing.

[0177] The geometric image generator 40002 generates geometric images based on point cloud data, patches, and / or packed patches. A geometric image is data that contains geometry related to point cloud data.

[0178] Texture image generation or texture image generator 40003 generates texture images based on point cloud data, patches, and / or packed patches. Additionally, texture images can be further generated based on smoothed geometry generated through smoothing processing based on patch information.

[0179] Smoothing or smoother 40004 can mitigate or eliminate errors contained in image data. For example, in a patch-based reconstructed geometry image, parts of the data that could cause errors can be smoothly filtered out to generate smooth geometry.

[0180] The auxiliary patch information is compressed, or the auxiliary patch information compressor 40005 compresses the auxiliary patch information related to the patch information generated during patch generation. Additionally, the compressed auxiliary patch information can be sent to a multiplexer. The auxiliary patch information can be generated in the geometry image generation 40002.

[0181] Image fillers or image fillers 40006 and 40007 can fill geometric images and texture images respectively. Fill data can be applied to both geometric and texture images.

[0182] Group expansion or group expander 40008 can add data to a texture image in a manner similar to image filling. The added data can be inserted into the texture image.

[0183] Video compression or video compressors 40009, 40010, and 40011 can compress filled geometric images, filled texture images, and / or occupancy maps, respectively. Compression can encode geometric information, texture information, occupancy information, etc.

[0184] Entropy compression or entropy compressor 40012 can compress (e.g., encode) occupancy graphs based on entropy schemes.

[0185] According to the implementation method, entropy compression and / or video compression can be performed based on whether the point cloud data is lossless and / or lossy.

[0186] Multiplexer 40013 multiplexes compressed geometric images, compressed texture images, and compressed occupancy maps into a bitstream.

[0187] The following description Figure 4 Specific operations in each process.

[0188] 40,000 patches generated

[0189] Patch generation refers to the process of dividing a point cloud into patches (mapping units) to map the point cloud onto a 2D image. Patch generation can be divided into three steps: normal value calculation, segmentation, and patch segmentation.

[0190] Reference Figure 5 Describe the normal value calculation process in detail.

[0191] Figure 5 An example of the tangent plane and normal vector of a surface according to an embodiment is shown.

[0192] Figure 5 The surface is used as follows Figure 4 The patch generation process of V-PCC encoding is in 40000.

[0193] Normal calculations related to patch generation:

[0194] Each point in a point cloud has its own orientation, represented by a 3D vector called a normal vector. Using the neighbors of each point obtained through methods such as KD-trees, the tangent planes and normal vectors of each point on the surface constituting the point cloud, as shown in the figure, can be obtained. The search range used for neighbor search can be defined by the user.

[0195] A tangent plane is a plane that passes through a point on a surface and completely includes the tangent to a curve on the surface.

[0196] Figure 6 An exemplary bounding box of a point cloud according to an implementation method is shown.

[0197] The method / apparatus according to the implementation (e.g., patch generation) may employ bounding boxes when generating patches from point cloud data.

[0198] Bounding boxes can be used in processing where a target object of point cloud data is projected onto the planes of the flat faces of a hexahedron in 3D space. Bounding boxes can be defined by... Figure 1 The point cloud video acquirer 10001 and point cloud video encoder 10002 generate and process the data. Furthermore, based on bounding boxes, the following can be performed: Figure 2 The V-PCC encoding process generates patch 40000, patch packing 40001, geometric image generation 40002, and texture image generation 40003.

[0199] Segments related to patch generation

[0200] The segmentation is divided into two processes: initial segmentation and refined segmentation.

[0201] The point cloud video encoder 10002 according to the embodiment projects points onto one face of a bounding box. Specifically, each point constituting the point cloud is projected onto one of the six faces of the bounding box surrounding the point cloud. Initial segmentation is a process of determining one of the flat faces of the bounding box to which each point is to be projected.

[0202] It is the normal value corresponding to each of the six flat planes, as defined below:

[0203] (1.0,0.0,0.0), (0.0,1.0,0.0), (0.0,0.0,1.0), (-1.0,0.0,0.0), (0.0,-1.0,0.0), (0.0,0.0,-1.0).

[0204] As shown in the following formula, the normal vectors of each point are obtained in the normal value calculation process. and The plane with the largest dot product value is determined as the projection plane of the corresponding point. That is, the plane whose direction is most similar to the normal vector of the point is determined as the projection plane of the point.

[0205]

[0206] The determined plane can be identified by a cluster index (one of 0 to 5).

[0207] Refining the segmentation involves considering the projection plane enhancement of each point constituting the point cloud, determined during the initial segmentation process. In this process, scoring normals and scoring smoothing can be considered together. The scoring normal represents the similarity between the normal vectors of each point considered when determining the projection planes in the initial segmentation process and the normals of each flat surface of the bounding box. Scoring smoothing indicates the similarity between the projection plane of the current point and the projection planes of its neighboring points.

[0208] Score smoothing can be considered by assigning weights to the scoring method. In this case, the weight values ​​can be defined by the user. Refining the segments can be performed repeatedly, and the number of repetitions can also be defined by the user.

[0209] Patch segmentation related to patch generation

[0210] Patch segmentation is a process that divides the entire point cloud into patches (sets of neighboring points) based on the projection plane information of each point constituting the point cloud obtained in the initial / refining segmentation process. Patch segmentation may include the following steps:

[0211] 1) Use KD-trees or similar methods to calculate the neighboring points of each point in the point cloud. The maximum number of neighbors can be defined by the user.

[0212] 2) When neighboring points are projected onto the same plane as the current point (when they have the same cluster index), extract the current point and neighboring points as a patch.

[0213] 3) Calculate the geometric values ​​of the extracted patch. Details are described in Section 1.3.

[0214] 4) Repeat steps 2) to 4) until there are no more unextracted points.

[0215] The occupancy map, geometric image, and texture image of each patch, as well as the size of each patch, are determined through patch segmentation processing.

[0216] Figure 7 An example is shown of determining the positions of each patch on the occupancy map according to an implementation method.

[0217] The point cloud video encoder 10002 according to the implementation method can perform patch packaging and generate a precision map.

[0218] Patch packing and occupancy map generation (40001)

[0219] This is the process of determining the positions of individual patches in a 2D image to map segmented patches onto the 2D image. As a 2D image, an occupancy map is a binary image that uses values ​​of 0 or 1 to indicate the presence or absence of data at corresponding locations. An occupancy map consists of blocks, and its resolution is determined by the block size. For example, when the block is 1x1, pixel-level resolution is obtained. The size of the occupancy block can be determined by the user.

[0220] The process for determining the position of each patch on the occupancy map can be configured as follows:

[0221] 1) Set all positions on the occupied map to 0;

[0222] 2) Place the patch at point (u,v) in the occupied plane with a horizontal coordinate in the range of (0,occupancySizeU-patch.sizeU0) and a vertical coordinate in the range of (0,occupancySizeV-patch.sizeV0);

[0223] 3) Set the point (x, y) in the patch plane whose horizontal coordinate is in the range (0, patch.sizeU0) and whose vertical coordinate is in the range (0, patch.sizeV0) as the current point;

[0224] 4) Change the position of point (x, y) in raster order, and if the value of coordinate (x, y) on the patch occupancy map is 1 (data exists at the point in the patch) and the value of coordinate (u+x, v+y) on the global occupancy map is 1 (the occupancy map is filled with the previous patch), then repeat operations 3) and 4). Otherwise, proceed to operation 6).

[0225] 5) Change the position of (u,v) according to the grating sequence and repeat operations 3) to 5);

[0226] 6) Determine (u,v) as the location of the patch and copy the occupancy map data for the patch to the corresponding portion of the global occupancy map; and

[0227] 7) Repeat steps 2) to 7) for the next patch.

[0228] occupancySizeU: Indicates the width of the occupancy map. Its unit is the size of the occupancy block.

[0229] occupancySizeV: Indicates the height of the occupancy map. Its unit is the size of the occupied block.

[0230] patch.sizeU0: Indicates the width of the occupied graph. Its unit is the size of the occupied packing block.

[0231] patch.sizeV0: Indicates the height of the occupied map. Its unit is the size of the occupied block.

[0232] For example, such as Figure 7 As shown, there exists a box corresponding to a patch with a patch size within the box corresponding to the occupied package size, and the point (x,y) can be located within this box.

[0233] Figure 8 An exemplary relationship between the normal axis, tangential axis, and double tangential axis is shown according to an embodiment.

[0234] The point cloud video encoder 10002 according to the embodiment can generate geometric images. A geometric image refers to image data that includes geometric information about the point cloud. The geometric image generation process can employ... Figure 8 The patch has three axes (normal, tangential, and bitangential).

[0235] Geometric Image Generation (40002)

[0236] In this process, the depth values ​​of the geometric images constituting each patch are determined, and the entire geometric image is generated based on the patch positions determined in the patch packing process described above. The process for determining the depth values ​​of the geometric images constituting each patch can be configured as follows.

[0237] 1) Calculate the parameters related to the position and size of each patch. The parameters may include the following information.

[0238] The normal index of the indicator normal axis is obtained in the previous patch generation process. The tangential axis is the axis perpendicular to the normal axis that coincides with the horizontal axis u of the patch image, and the double tangential axis is the axis perpendicular to the normal axis that coincides with the vertical axis v of the patch image. The three axes are shown in the figure.

[0239] Figure 9 Exemplary configurations of the minimum and maximum modes of the projection mode according to the implementation are shown.

[0240] The point cloud video encoder 10002 according to the embodiment can perform patch-based projection to generate a geometric image, and the projection modes according to the embodiment include a minimum mode and a maximum mode.

[0241] The 3D spatial coordinates of the patch can be calculated based on the bounding box surrounding the minimum size of the patch. For example, the 3D spatial coordinates may include the minimum tangential value of the patch (on the 3D displacement tangential axis of the patch), the minimum bitangential value of the patch (on the 3D displacement bitangential axis of the patch), and the minimum normal value of the patch (on the 3D displacement normal axis of the patch).

[0242] The 2D dimensions of the patch indicate the horizontal and vertical dimensions of the patch when it is packed into a 2D image. The horizontal dimension (pattern 2D dimension u) can be obtained as the difference between the maximum and minimum tangent values ​​of the bounding box, and the vertical dimension (pattern 2D dimension v) can be obtained as the difference between the maximum and minimum bitangent values ​​of the bounding box.

[0243] 2) Determine the projection mode of the patch. The projection mode can be either the minimum mode or the maximum mode. Geometric information about the patch is represented using depth values. When the points constituting the patch are projected onto the normal of the patch, two layers of images can be generated: an image constructed using the maximum depth value and an image constructed using the minimum depth value.

[0244] In minimum mode, when generating two layers of images d0 and d1, a minimum depth can be configured for d0, and a maximum depth within the surface thickness of the minimum depth can be configured for d1, as shown in the figure.

[0245] For example, when a point cloud is located in 2D as shown in the figure, multiple patches containing multiple points can exist. As shown in the figure, points marked with the same shading style can belong to the same patch. This figure illustrates the processing of patches with projected blank points.

[0246] When projecting blank points to the left / right, the depth can be increased by 1 relative to the left side as 0, 1, 2, ..., 6, 7, 8, 9, and the number used to calculate the depth of the point can be marked on the right side.

[0247] The same projection mode can be applied to all point clouds, or different projection modes can be applied to individual frames or patches based on user definitions. When applying different projection modes to individual frames or patches, the projection mode that enhances compression efficiency or minimizes missing points can be adaptively selected.

[0248] 3) Calculate the depth value of each point.

[0249] In minimum mode, image d0 is constructed using depth0, which is obtained by subtracting the minimum normal value of the patch (on the patch's 3D shift normal axis) calculated in operation 1) from the minimum normal value of the patch at each point (on the patch's 3D shift normal axis). If another depth value exists at the same location within the range between depth0 and the surface thickness, that value is set to depth1. Otherwise, the value of depth0 is assigned to depth1. Image d1 is constructed using the value of depth1.

[0250] For example, the minimum value (4 2 4 4 0 6 0 0 9 9 0 8 0) can be calculated when determining the depth of a point in image d0. When determining the depth of a point in image d1, the larger value among two or more points can be calculated. When only one point exists, its value (4 4 4 4 6 6 68 9 9 8 8 9) can be calculated. During the processing of points in the encoding and reconstruction patch, some points may be lost (e.g., eight points are lost in the figure).

[0251] In maximum mode, image d0 is constructed using depth0, which is obtained by subtracting the minimum normal value of the patch (on the 3D shifted normal axis of the patch) calculated in operation 1) from the minimum normal value of the patch (on the 3D shifted normal axis of the patch) for each point using the maximum normal value. If another depth value exists at the same location within the range between depth0 and the surface thickness, that value is set to depth1. Otherwise, the value of depth0 is assigned to depth1. Image d1 is constructed using the value of depth1.

[0252] For example, the maximum value (4 4 4 4 6 6 6 8 9 9 8 8 9) can be calculated when determining the depth of a point in image d0. Conversely, when determining the depth of a point in image d1, the lower value among two or more points can be calculated. When only one point exists, its value (4 24 4 5 6 0 6 9 9 0 8 0) can be calculated. During the processing of points in the encoded and reconstructed patch, some points may be lost (e.g., six points are lost in the image).

[0253] The entire geometric image can be generated by placing the geometric images of each patch generated through the above processing onto the entire geometric image based on the patch position information determined in the patch packing process.

[0254] Layer d1 of the generated entire geometric image can be encoded using various methods. The first method (absolute d1 method) encodes the depth values ​​of the previously generated image d1. The second method (differential method) encodes the difference between the depth values ​​of the previously generated image d1 and the depth values ​​of image d0.

[0255] In the encoding method described above that uses depth values ​​of two layers d0 and d1, if there is another point between the two depths, the geometric information about that point is lost during the encoding process. Therefore, Enhanced Incremental Depth (EDD) codes can be used for lossless encoding.

[0256] The following will refer to Figure 10 Describe the EDD code in detail.

[0257] Figure 10 An exemplary EDD code according to an implementation method is shown.

[0258] In some / all of the processing of point cloud video encoder 10002 and / or V-PCC encoding (e.g., video compression 40009), geometric information about points can be encoded based on EOD codes.

[0259] As shown in the figure, the EDD code is used for binary encoding of the positions of all points within the surface thickness range of d1. For example, in the figure, since points exist at the first and fourth positions on D0 and the second and third positions are empty, the points included in the second column on the left can be represented by the EDD code 0b1001 (=9). When the EDD code is encoded and transmitted together with D0, the receiving terminal can recover the geometric information about all points without loss.

[0260] For example, the value is 1 when there is a point above the reference point, and 0 when there is no point. Therefore, the code can be represented using 4 bits.

[0261] Smoothing (40004)

[0262] Smoothing is an operation used to eliminate discontinuities that may appear at patch boundaries due to image quality degradation that occurs during compression processing. Smoothing can be performed by a point cloud encoder or a smoother.

[0263] 1) Reconstructing point clouds from geometric images. This operation can be the reverse of the geometric image generation described above; for example, the inverse processing of reconstructable codes.

[0264] 2) Use KD trees and other methods to calculate the neighboring points of each point in the reconstructed point cloud.

[0265] 3) Determine whether each point is located on the patch boundary. For example, when there are neighboring points with a different projection plane (cluster index) than the current point, it can be determined that the point is located on the patch boundary.

[0266] 4) If a point exists on the patch boundary, move that point to the centroid of a neighboring point (located at the average x, y, z coordinates of the neighboring point). That is, change the geometry. Otherwise, maintain the previous geometry.

[0267] Figure 11 An example of recoloring based on the color values ​​of neighboring points according to an implementation method is shown.

[0268] The point cloud encoder or texture image generator 40003 according to the implementation can generate texture images based on recoloring.

[0269] Texture image generation (40003)

[0270] Similar to the geometric image generation process described above, the texture image generation process involves generating texture images for each patch and generating the entire texture image by arranging the texture images in defined positions. However, in the operation of generating texture images for each patch, instead of using depth values ​​for geometry generation, an image with color values ​​(e.g., R, G, and B values) of the points that constitute the point cloud corresponding to the position is generated.

[0271] When estimating the color values ​​of the individual points that make up a point cloud, the geometry previously obtained through smoothing can be used. In a smoothed point cloud, the positions of some points may have shifted relative to the original point cloud, so a recoloring process may be needed to find colors suitable for the changed positions. Recoloring can be performed using the color values ​​of neighboring points. For example, as shown in the figure, the color values ​​of the nearest neighbor and neighboring points can be considered to calculate new color values.

[0272] For example, referring to the accompanying figure, during recoloring, the appropriate color value for the changed position can be calculated based on the average of attribute information about the nearest original point and / or the average of attribute information about the nearest original location.

[0273] Similar to a geometric image generated from two layers d0 and d1, a texture image can also be generated from two layers t0 and t1.

[0274] Auxiliary patch information compression (40005)

[0275] The point cloud encoder or auxiliary patch information compressor according to the implementation method can compress auxiliary patch information (auxiliary information about the point cloud).

[0276] The auxiliary patch information compressor compresses the auxiliary patch information generated during the patch generation, patch packaging, and geometry generation processes described above. The auxiliary patch information may include the following parameters:

[0277] An index (cluster index) used to identify the projection plane (normal plane);

[0278] The 3D spatial position of the patch, namely, the minimum tangential value of the patch (on the 3D displacement tangential axis of the patch), the minimum double tangential value of the patch (on the 3D displacement double tangential axis of the patch), and the minimum normal value of the patch (on the 3D displacement normal axis of the patch).

[0279] The 2D spatial position and dimensions of the patch, namely, the horizontal dimension (pattern 2D dimension u), the vertical dimension (pattern 2D dimension v), the minimum horizontal value (pattern 2D displacement u), and the minimum vertical value (pattern 2D displacement u); and

[0280] Regarding the mapping information for each block and patch, there are candidate indices (when patches are set sequentially based on information about their 2D spatial location and size, multiple patches can be mapped to a block in an overlapping manner. In this case, the mapped patches constitute a candidate list, and the candidate index indicates the sequential position of the patch whose data exists within the block) and local patch indices (indicating the index of a patch existing in a frame). Table X shows pseudocode representing the process of matching between blocks and patches based on the candidate list and local patch indices.

[0281] The maximum number of candidates can be defined by the user.

[0282] Table 1-1 Pseudocode for mapping blocks to patches

[0283]

[0284] Figure 12 This illustrates a push-pull background fill according to an embodiment.

[0285] Image fill and group expansion (40006, 40007, 40008)

[0286] According to the implementation method, the image filler can fill the space other than the patch area using meaningless supplementary data based on push-pull background fill technology.

[0287] Image padding is a process that fills the space outside the patch area with meaningless data to improve compression efficiency. For image padding, pixel values ​​from columns or rows near the boundaries of the patch can be copied to fill the blank space. Alternatively, as shown in the figure, a push-pull background padding method can be used. According to this method, pixel values ​​from the low-resolution image are used to fill the blank space in a process of gradually reducing the resolution of the unpadded image and then increasing the resolution again.

[0288] Group dilation is a process that fills the blank spaces in a geometric image and a texture image configured with two layers, d0 / d1 and t0 / t1, respectively. In this process, the blank spaces of the two layers are calculated by averaging the values ​​at the same location using image filling.

[0289] Figure 13 An exemplary possible traversal order of a 4x4 block according to an implementation is shown.

[0290] Occupancy map compression (40012, 40011)

[0291] The occupancy map compressor according to the implementation can compress previously generated occupancy maps. Specifically, two methods can be used: video compression for lossy compression and entropy compression for lossless compression. Video compression is described below.

[0292] Entropy compression can be performed using the following methods.

[0293] 1) If the block constituting the occupancy map is fully occupied, encode 1 and repeat the same operation for the next block of the occupancy map. Otherwise, encode 0 and perform operations 2) through 5).

[0294] 2) Determine the optimal traversal order for performing run-length encoding on the occupied pixels of the block. The figure shows four possible traversal orders for a 4x4 block.

[0295] Figure 14 An exemplary optimal traversal order is shown according to the implementation method.

[0296] As described above, the entropy compressor according to the implementation can encode blocks based on the traversal order scheme as described above.

[0297] For example, the optimal traversal order with the minimum number of runs is selected from the possible traversal orders, and its index is encoded. The diagram illustrates the selection process. Figure 13 The case of the third traversal order. In the case shown, the number of runs can be minimized to 2, therefore the third traversal order can be chosen as the optimal traversal order.

[0298] 3) Encode the number of runs. Figure 14 In the example, there are two runs, so 2 is encoded.

[0299] 4) Encode the occupancy of the first run. Figure 14 In the example, 0 is encoded because the first run corresponds to an unoccupied pixel.

[0300] 5) Encode the length of each run (the total number of runs). Figure 14 In the example, the lengths of the first run and the second run, 6 and 10, are encoded in sequence.

[0301] Video compression (40009, 40010, 40011)

[0302] The video compressor according to the embodiment uses a 2D video codec such as HEVC or VVC to encode sequences of geometric images, texture images, occupancy map images, etc. generated in the above operations.

[0303] Figure 15 An exemplary 2D video / image encoder according to an embodiment is shown.

[0304] This figure illustrates an implementation of the video compression or video compressors 40009, 40010, and 40011 described above, and is a schematic block diagram of a 2D video / image encoder 15000 configured to encode video / image signals. The 2D video / image encoder 15000 may be included in the aforementioned point cloud video encoder, or may be configured as an internal / external component. ​ The individual components may correspond to software, hardware, processor, and / or combinations thereof.

[0305] Here, the input images may include the geometric images, texture images (attribute images), and occupancy map images mentioned above. The output bitstream of the point cloud video encoder (i.e., the point cloud video / image bitstream) may include the output bitstreams of each input image (i.e., the geometric image, texture image (attribute image), occupancy map image, etc.).

[0306] Inter-frame predictor 15090 and intra-frame predictor 15100 can be collectively referred to as predictors. That is, the predictor may include inter-frame predictor 15090 and intra-frame predictor 15100. Transformer 15030, quantizer 15040, inverse quantizer 15050, and inverse transformer 15060 may be included in the residual processor. The residual processor may also include subtractor 15020. According to an embodiment, the image segmenter 15010, subtractor 15020, transformer 15030, quantizer 15040, inverse quantizer 15050, inverse transformer 15060, adder 155, filter 15070, inter-frame predictor 15090, intra-frame predictor 15100, and entropy encoder 15110 may be configured by a single hardware component (e.g., encoder or processor). In addition, memory 15080 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium.

[0307] Image segmenter 15010 can segment an image (or picture or frame) input to encoder 15000 into one or more processing units. For example, a processing unit may be referred to as a coding unit (CU). In this case, the CU may be recursively segmented from coding tree units (CTUs) or maximum coding units (LCUs) according to a quadtree-binary tree (QTBT) structure. For example, a CU may be segmented into multiple CUs of lower depth based on a quadtree structure and / or a binary tree structure. In this case, for example, a quadtree structure may be applied first, followed by a binary tree structure. Alternatively, a binary tree structure may be applied first. The coding process according to this disclosure can be performed based on a final CU that is no longer segmented. In this case, an LCU may be used as the final CU based on coding efficiency according to the characteristics of the image. If necessary, the CU may be recursively segmented into lower-depth CUs, and the CU of optimal size may be used as the final CU. Here, the coding process may include prediction, transform, and reconstruction (described later). As another example, the processing unit may also include a prediction unit (PU) or a transform unit (TU). In this case, PU and TU can be separated or divided from the final CU mentioned above. PU can be a unit for sample prediction, and TU can be a unit for deriving transform coefficients and / or deriving residual signals from transform coefficients.

[0308] The term "unit" is used interchangeably with terms such as block or region. In general, an M×N block can represent a set of samples or transform coefficients arranged in M ​​columns and N rows. A sample typically represents a pixel or pixel value, and may indicate only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component. "Sample" can be used as a term corresponding to a pixel or image within a frame (or image).

[0309] The encoder 15000 generates a residual signal (residual block or residual sample array) by subtracting the prediction signal (prediction block or prediction sample array) output from the inter-frame predictor 15090 or intra-frame predictor 15100 from the input image signal (original block or original sample array), and the generated residual signal is sent to the converter 15030. In this case, as shown, the unit in the encoder 15000 that subtracts the prediction signal (prediction block or prediction sample array) from the input image signal (original block or original sample array) may be referred to as the subtractor 15020. The predictor can perform prediction on the processing target block (hereinafter referred to as the current block) and generate a prediction block that includes the prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. As described later in the description of the various prediction modes, the predictor can generate various types of information about the prediction (e.g., prediction mode information) and transmit the generated information to the entropy encoder 15110. The information about the prediction can be encoded by the entropy encoder 15110 and output as a bitstream.

[0310] The intra-frame predictor 15100 can predict the current block by referencing samples in the current frame. Depending on the prediction mode, the samples can be near or far from the current block. Under intra-frame prediction, the prediction modes can include multiple non-directional modes and multiple directional modes. Non-directional modes can include, for example, DC mode and planar mode. Depending on the granularity of the prediction direction, the directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is only an example, and more or fewer directional prediction modes can be used depending on the settings. The intra-frame predictor 15100 can determine the prediction mode to be applied to the current block based on the prediction modes applied to neighboring blocks.

[0311] The inter-frame predictor 15090 can derive the prediction block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference frame. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted based on each block, sub-block, or sample based on the correlation between motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference frame indices. Motion information may also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current frame and temporally neighboring blocks existing in the reference frame. The reference frame including the reference block may be the same as or different from the reference frame including the temporally neighboring block. The temporally neighboring block may be referred to as a juxtaposed reference block or juxtaposed CU (colCU), and the reference frame including the temporally neighboring block may be referred to as a juxtaposed frame (colPic). For example, the inter-frame predictor 15090 can configure a motion information candidate list based on neighboring blocks and generate information indicating candidates to be used for deriving the motion vector and / or reference frame index of the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip and merge modes, the inter-frame predictor 15090 can use motion information about neighboring blocks as motion information about the current block. In skip mode, unlike merge mode, residual signals may not be sent. In motion vector prediction (MVP) mode, motion vectors from neighboring blocks can be used as motion vector predictors, and the motion vector difference can be signaled to indicate the motion vector of the current block.

[0312] The prediction signal generated by the inter-frame predictor 15090 or the intra-frame predictor 15100 can be used to generate the reconstructed signal or the residual signal.

[0313] Transformer 15030 can generate transform coefficients by applying transform techniques to the residual signal. For example, the transform techniques may include at least one of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen–Loève Transform (KLT), Graphical Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, GBT refers to a transform obtained from a graph depicting the relationships between pixels. CNT refers to a transform obtained based on a prediction signal generated from all previously reconstructed pixels. Furthermore, the transform operation can be applied to square pixel blocks of the same size, or to blocks of variable size other than square.

[0314] The quantizer 15040 quantizes the transform coefficients and sends them to the entropy encoder 15110. The entropy encoder 15110 encodes the quantized signal (information about the quantized transform coefficients) and outputs a bitstream of the encoded signal. The information about the quantized transform coefficients can be referred to as residual information. The quantizer 15040 rearranges the quantized transform coefficients in block form according to the coefficient scan order in the form of a one-dimensional vector, and generates information about the quantized transform coefficients based on the one-dimensional vector form of the quantized transform coefficients. The entropy encoder 15110 can employ various coding techniques such as, for example, exponential Golomb, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoder 15110 can encode information required for video / image reconstruction (e.g., values ​​of syntactic elements) together with or separately from the quantized transform coefficients. The encoded information (e.g., encoded video / image information) can be sent or stored as a bitstream based on Network Abstraction Layer (NAL) units. The bitstream can be sent via a network or stored in a digital storage medium. Here, the network may include a broadcast network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. A transmitter (not shown) that sends a signal output from the entropy encoder 15110 and / or a storage unit (not shown) that stores the signal may be configured as internal / external elements of the encoder 15000. Alternatively, the transmitter may be included within the entropy encoder 15110.

[0315] The quantized transform coefficients output from quantizer 15040 can be used to generate a prediction signal. For example, inverse quantization and inverse transform can be applied to the quantized transform coefficients via inverse quantizer 15050 and inverse transformer 15060 to reconstruct the residual signal (residual block or residual sample). Adder 155 adds the reconstructed residual signal to the prediction signal output from inter-frame predictor 15090 or intra-frame predictor 15100. This generates a reconstructed signal (reconstructed frame, reconstructed block, reconstructed sample array). When there is no residual signal for the processing target block, as in the case of applying skip mode, the prediction block can be used as a reconstructed block. Adder 155 can be referred to as a reconstructor or reconstructed block generator. As described below, the generated reconstructed signal can be used for intra-frame prediction of the next processing target block in the current frame, or it can be filtered for inter-frame prediction of the next frame.

[0316] Filter 15070 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 15070 can generate a modified reconstructed image by applying various filtering techniques to the reconstructed image, and the modified reconstructed image can be stored in memory 15080 (specifically, the DPB of memory 15080). Various filtering techniques may include, for example, deblocking filtering, sample adaptive offsetting, adaptive loop filtering, and bilateral filtering. As described below in the description of filtering techniques, filter 15070 can generate various types of information about the filtering and transmit the generated information to entropy encoder 15110. The information about the filtering can be encoded by entropy encoder 15110 and output as a bitstream.

[0317] The modified reconstructed frame sent to memory 15080 can be used as a reference frame by inter-frame predictor 15090. Therefore, when applying inter-frame prediction, the encoder can avoid prediction mismatch between encoder 15000 and decoder and improve coding efficiency.

[0318] The DPB of memory 15080 can store modified reconstructed frames for use as reference frames by inter-frame predictor 15090. Memory 15080 can store motion information about blocks in the current frame that have been deduced (or encoded) and / or motion information about already reconstructed blocks in the frame. The stored motion information can be transmitted to inter-frame predictor 15090 for use as motion information about spatially adjacent blocks or temporally adjacent blocks. Memory 15080 can store reconstructed samples of reconstructed blocks in the current frame and transmit the reconstructed samples to intra-frame predictor 15100.

[0319] At least one of the above prediction, transformation, and quantization processes can be skipped. For example, for a block using Pulse Code Mode (PCM), the prediction, transformation, and quantization processes can be skipped, and the values ​​of the original samples can be encoded and output as a bitstream.

[0320] ​ An exemplary V-PCC decoding process according to an implementation is shown.

[0321] V-PCC decoding processing or V-PCC decoder can follow ​ The inverse processing of V-PCC encoding (or encoder). ​ The individual components can correspond to software, hardware, processors, and / or combinations thereof.

[0322] The demultiplexer 16000 demultiplexes the compressed bitstream to output a compressed texture image, a compressed geometric image, a compressed occupancy map, and compressed auxiliary patch information.

[0323] Video decompression or video decompressor 16001, 16002 decompresses (or decodes) each of the compressed texture image and compressed geometric image.

[0324] Occupancy diagram compression or occupancy diagram compressor 16003 will compress the occupancy diagram.

[0325] The auxiliary patch information is decompressed, or the auxiliary patch information decompressor 16004 decompresses the auxiliary patch information.

[0326] The geometric reconstruction, or geometric reconstructor 16005, recovers (reconstructs) geometric information based on the decompressed geometric image, the decompressed occupancy map, and / or the decompressed auxiliary patch information. For example, geometry altered during encoding processing can be reconstructed.

[0327] The smoother or smoother16006 can smooth the reconstructed geometry. For example, a smoothing filter can be applied.

[0328] Texture reconstruction or texture reconstructor 16007 reconstructs textures from decompressed texture images and / or smoothed geometry.

[0329] Color smoothing, or the color smoother 16008, smooths color values ​​from the reconstructed texture. For example, a smoothing filter can be applied.

[0330] As a result, reconstructed point cloud data can be generated.

[0331] The figure illustrates the V-PCC decoding process for reconstructing a point cloud by decoding the compressed occupancy map, geometric image, texture image, and auxiliary patch information. The various processes according to the implementation are as follows.

[0332] Video decompression (1600, 16002)

[0333] Video decompression is the inverse process of video compression described above. In video decompression, a 2D video codec such as HEVC or VVC is used to decode the compressed bitstream containing the geometric images, texture images, and occupancy map images generated in the above process.

[0334] ​ An exemplary 2D video / image decoder according to an embodiment is shown.

[0335] 2D video / image decoders can follow ​ Inverse processing of a 2D video / image encoder.

[0336] ​ The 2D video / image decoder is ​ A video decompressor or video decompressor. ​ This is a schematic block diagram of a 2D video / image decoder 17000 that performs video / image signal decoding. The 2D video / image decoder 17000 may be included in... ​ In point cloud video decoders, it can be configured as an internal / external component. ​ The individual components can correspond to software, hardware, processors, and / or combinations thereof.

[0337] Here, the input bitstream may include the bitstreams of the aforementioned geometric image, texture image (attribute image), and occupancy map image. The reconstructed image (or output image or decoded image) may represent the reconstructed image of the aforementioned geometric image, texture image (attribute image), and occupancy map image.

[0338] Referring to the figure, the inter-frame predictor 17070 and the intra-frame predictor 17080 can be collectively referred to as predictors. That is, the predictor may include the inter-frame predictor 17070 and the intra-frame predictor 17080. The inverse quantizer 17020 and the inverse transformer 17030 can be collectively referred to as residual processors. That is, according to the embodiment, the residual processor may include the inverse quantizer 17020 and the inverse transformer 17030. The entropy decoder 17010, inverse quantizer 17020, inverse transformer 17030, adder 17040, filter 17050, inter-frame predictor 17070, and intra-frame predictor 17080 described above may be configured by a single hardware component (e.g., a decoder or a processor). In addition, the memory 170 may include a decoded frame buffer (DPB) or may be configured by a digital storage medium.

[0339] When the input contains a bitstream of video / image information, the decoder 17000 can reconstruct the image in a process corresponding to the encoder's processing of video / image information in Figure 0.2-1. For example, the decoder 17000 can use a processing unit applied in the encoder to perform decoding. Therefore, the decoding processing unit can be, for example, a CU. The CU can be segmented from a CTU or LCU along a quadtree structure and / or a binary tree structure. The reconstructed video signal decoded and output by the decoder 17000 can then be played back by a player.

[0340] Decoder 17000 can receive signals output from encoder in the form of a bitstream, and the received signals can be decoded by entropy decoder 17010. For example, entropy decoder 17010 can parse the bitstream to derive information (e.g., video / image information) required for image reconstruction (or picture reconstruction). For example, entropy decoder 17010 can decode the information in the bitstream based on coding techniques such as exponential Golomb coding, CAVLC, or CABAC, outputting quantized values ​​of the syntactic elements and transform coefficients of the residuals required for image reconstruction. More specifically, in CABAC entropy decoding, bins corresponding to individual syntactic elements in the bitstream can be received, and a context model can be determined based on information about the target syntactic elements and decoding information about neighboring target blocks or information about symbols / bins decoded in previous steps. The probability of bin occurrence can then be predicted based on the determined context model, and arithmetic decoding of the bins can be performed to generate symbols corresponding to the values ​​of the individual syntactic elements. According to CABAC entropy decoding, after determining the context model, the context model can be updated based on information about the symbols / cells decoded for the next symbol / cell. Information about prediction from the information decoded by the entropy decoder 17010 can be provided to the predictors (inter-frame predictor 17070 and intra-frame predictor 17080), and the residual values ​​(i.e., quantized transform coefficients and related parameter information) from the entropy decoding performed by the entropy decoder 17010 can be input to the inverse quantizer 17020. Additionally, information about filtering from the information decoded by the entropy decoder 17010 can be provided to the filter 17050. A receiver (not shown) configured to receive the signal output from the encoder can also be configured as an internal / external element of the decoder 17000. Alternatively, the receiver can be a component of the entropy decoder 17010.

[0341] The inverse quantizer 17020 outputs transform coefficients by inverse quantizing the quantized transform coefficients. The inverse quantizer 17020 can rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement can be performed based on the coefficient scan order implemented by the encoder. The inverse quantizer 17020 can perform inverse quantization on the quantized transform coefficients and obtain the transform coefficients using quantization parameters (e.g., quantization step size information).

[0342] The inverse converter 17030 obtains the residual signal (residual block and residual sample array) by performing an inverse transformation on the transform coefficients.

[0343] The predictor can perform prediction on the current block and generate a prediction block that includes prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on information about the prediction output from the entropy decoder 17010, and can determine a specific intra-frame / inter-frame prediction mode.

[0344] Intra-predictor 265 can refer to samples in the current frame to predict the current block. Depending on the prediction mode, the samples may be near or far from the current block. In intra-prediction, the prediction mode may include multiple non-directional modes and multiple directional modes. Intra-predictor 17080 can use the prediction modes applied to neighboring blocks to determine the prediction mode applied to the current block.

[0345] The inter-frame predictor 17070 can deduce the prediction block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference frame. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted based on the correlation between motion information of neighboring blocks and the current block, on a per-block, sub-block, or sample basis. Motion information may include motion vectors and reference frame indices. Motion information may also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current frame and temporally neighboring blocks existing in the reference frame. For example, the inter-frame predictor 17070 can configure a motion information candidate list based on neighboring blocks and deduce the motion vector and / or reference frame index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes. Information about the prediction may include information indicating the inter-frame prediction mode of the current block.

[0346] Adder 17040 can add the acquired residual signal to the prediction signal (prediction block or prediction sample array) output from inter-frame predictor 17070 or intra-frame predictor 17080 to generate a reconstruction signal (reconstructed frame, reconstruction block, or reconstruction sample array). When there is no residual signal for processing the target block, as in the case of applying skip mode, the prediction block can be used as a reconstruction block.

[0347] The adder 17040 can be referred to as a reconstructor or reconstruction block generator. The generated reconstructed signal can be used for intra-frame prediction of the next processing target block in the current frame, or it can be used for inter-frame prediction of the next frame by filtering, as described below.

[0348] Filter 17050 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 17050 can generate a modified reconstructed image by applying various filtering techniques to the reconstructed image, and can send the modified reconstructed image to memory 250 (specifically, the DPB of memory 17060). For example, various filtering methods may include deblocking filtering, sample adaptive shifting, adaptive loop filtering, and bilateral filtering.

[0349] The reconstructed frame stored in the DPB of memory 17060 can be used as a reference frame in the inter-frame predictor 17070. Memory 17060 can store motion information about blocks in the current frame whose motion information has been derived (or decoded) and / or about blocks in already reconstructed frames. The stored motion information can be transmitted to the inter-frame predictor 17070 as motion information about spatially adjacent blocks or about temporally adjacent blocks. Memory 17060 can store reconstructed samples of reconstructed blocks in the current frame and transmit the reconstructed samples to the intra-frame predictor 17080.

[0350] In this disclosure, the embodiments described with respect to the filter 160, inter-frame predictor 180, and intra-frame predictor 185 of the encoder 100 can be applied to the filter 17050, inter-frame predictor 17070, and intra-frame predictor 17080 of the decoder 17000 in the same or corresponding manner.

[0351] At least one of the above prediction, transformation, and quantization processes can be skipped. For example, for a block with Pulse Code Mode (PCM) applied, the prediction, transformation, and quantization processes can be skipped, and the values ​​of the decoded samples can be used as samples for reconstructing the image.

[0352] Occupancy diagram compression (16003)

[0353] This is the reverse process of the occupancy graph compression described above. Occupancy graph decompression is the process of reconstructing the occupancy graph by decompressing the occupancy graph bitstream.

[0354] Decompress auxiliary patch information (16004)

[0355] The auxiliary patch information can be reconstructed by performing the inverse processing of the above auxiliary patch information compression and decoding the compressed auxiliary patch information bit stream.

[0356] Geometric Reconstruction (16005)

[0357] This is the inverse process of generating the aforementioned geometric image. Initially, patches are extracted from the geometric image using the reconstructed occupancy map, 2D position / size information about the patches included in the auxiliary patch information, and information about the mapping between blocks and patches. Then, a point cloud is reconstructed in 3D space based on the extracted patch's geometric image and the 3D position information about the patches included in the auxiliary patch information. When the geometric value corresponding to a point (u,v) within the patch is g(u,v), and the patch's position coordinates on the normal, tangential, and bitangential axes in 3D space are (d0,s0,r0), the normal, tangential, and bitangential coordinates d(u,v), s(u,v), and r(u,v) mapped to the position of point (u,v) in 3D space can be represented as follows:

[0358] δ(u,v)=δ0+g(u,v)

[0359] s(u,v)=s0+u

[0360] r(u,v)=r0+v.

[0361] Smoothing (16006)

[0362] Similar to smoothing in the encoding process described above, smoothing is a process used to eliminate discontinuities that may appear at patch boundaries due to image quality degradation that occurs during compression.

[0363] Texture reconstruction (16007)

[0364] Texture reconstruction is a process of reconstructing a color point cloud by assigning color values ​​to each point that makes up the smooth point cloud. This can be performed by assigning the color values ​​corresponding to the texture image pixels at the same locations in the geometric image in 2D space to the points in the point cloud corresponding to the same locations in 3D space, based on the mapping information between the geometric image and the point cloud in the geometric reconstruction process described above.

[0365] Color smoothing (16008)

[0366] Color smoothing is similar to the geometric smoothing process described above. Color smoothing is used to eliminate discontinuities that may appear at patch boundaries due to image quality degradation that occurs during compression. Color smoothing can be performed through the following operations:

[0367] 1) Calculate the neighboring points of each point in the reconstructed point cloud using methods such as KD-trees. The neighboring point information calculated in the geometric smoothing process described in Section 2.5 can be used.

[0368] 2) Determine whether each point lies on the patch boundary. These operations can be performed based on the boundary information calculated in the geometric smoothing process described above.

[0369] 3) Examine the distribution of color values ​​of neighboring points of a point existing on the boundary and determine whether smoothing should be performed. For example, when the entropy of a brightness value is less than or equal to a threshold local entry (where many similar brightness values ​​exist), it can be determined that the corresponding part is not an edge part, and smoothing can be performed. As a smoothing method, the color value of a point can be replaced with the average of the color values ​​of its neighboring points.

[0370] ​ This is a flowchart illustrating the operation of a transmitting device according to an embodiment of the present disclosure.

[0371] The transmitting device according to the embodiment may correspond to ​ The transmitting device ​ Encoding processing and ​ A 2D video / image encoder, or performs some / all of its operations. The various components of the transmitting device may correspond to software, hardware, a processor, and / or a combination thereof.

[0372] The operation of compressing and transmitting point cloud data using V-PCC at the sending terminal can be performed as shown in the figure.

[0373] The point cloud data transmitting device according to the implementation method may be referred to as a transmitting device.

[0374] Regarding patch generator 18000, it generates patches for 2D image mapping of point clouds. Auxiliary patch information is generated as a result of patch generation. The generated information can be used in geometric image generation, texture image generation, and processing for smooth geometric reconstruction.

[0375] Regarding patch packer 18001, it performs patch packing processing that maps the generated patches to a 2D image. As a result of patch packing, an occupancy map is generated. The occupancy map can be used in geometry image generation, texture image generation, and processing for smooth geometry reconstruction.

[0376] The geometric image generator 18002 generates geometric images based on auxiliary patch information and occupancy maps. The generated geometric images are encoded into a bitstream using video coding.

[0377] The encoding preprocessor 18003 may include image padding processing. The regenerated geometric image, obtained by decoding the generated geometric image or the encoded geometric bitstream, can be used for 3D geometric reconstruction and then subjected to smoothing processing.

[0378] The texture image generator 18004 can generate texture images based on (smoothed) 3D geometry, point clouds, auxiliary patch information, and occupancy maps. The generated texture images can be encoded into a video bitstream.

[0379] The metadata encoder 18005 can encode auxiliary patch information into a metadata bitstream.

[0380] The 18006 video encoder can encode a occupancy map into a video bitstream.

[0381] The multiplexer 18007 can multiplex the generated geometric image, texture image, and occupancy map video bitstream and the metadata bitstream of auxiliary patch information into a single bitstream.

[0382] The transmitter 18008 can send a bitstream to a receiving terminal. Alternatively, the generated video bitstream of the geometric image, texture image, and occupancy map, as well as the metadata bitstream of the auxiliary patch information, can be processed into a file of one or more track data or encapsulated into fragments, and can be sent to the receiving terminal via the transmitter.

[0383] ​ This is a flowchart illustrating the operation of the receiving device according to an embodiment.

[0384] The receiving device according to the embodiment may correspond to ​ The receiving device ​ Decoding processing and ​ A 2D video / image encoder, or performs some / all of its operations. The various components of the receiving device may correspond to software, hardware, a processor, and / or a combination thereof.

[0385] The operation of receiving and reconstructing point cloud data using V-PCC at the receiving terminal can be performed as shown in the figure. The operation of the V-PCC receiving terminal can follow... ​ The reverse processing of the operation of the V-PCC transmitting terminal.

[0386] The point cloud data receiving device according to the implementation method may be referred to as a receiving device.

[0387] The received point cloud bitstream, after file / fragment decapsulation, is demultiplexed by demultiplexer 19000 into a compressed geometric image, texture image, occupancy map video bitstream, and metadata bitstream of auxiliary patch information. Video decoder 19001 and metadata decoder 19002 decode the demultiplexed video bitstream and metadata bitstream. The 3D geometry is reconstructed by geometry reconstructor 19003 based on the decoded geometric image, occupancy map, and auxiliary patch information, and then undergoes smoothing processing performed by smoother 19004. The color point cloud image / picture can be reconstructed by texture reconstructor 19005 by assigning color values ​​to the smoothed 3D geometry based on the texture image. Subsequently, additional color smoothing processing can be performed to improve subjective / objective visual quality, and the modified point cloud image / picture derived through color smoothing processing is displayed to the user through rendering processing (e.g., by a point cloud renderer). In some cases, color smoothing processing can be skipped.

[0388] ​An exemplary architecture for V-PCC-based storage and streaming of point cloud data, according to an embodiment, is shown.

[0389] ​ The system may include part or all of it. ​ Transmitting and receiving devices, ​ Encoding processing, ​ 2D video / image encoder, ​ Decoding processing, ​ The transmitting device and / or ​ Some or all of the receiving device. The various components in the figure may correspond to software, hardware, processor, and / or combinations thereof.

[0390] ​ This diagram illustrates the structure of the system further connected to the transmitting and receiving devices according to the embodiment. The transmitting and receiving devices of the system according to the embodiment may be referred to as the transmitting / receiving devices according to the embodiment.

[0391] exist ​ In the device shown according to the embodiment, with ​ The corresponding transmitting device can generate a container suitable for a data format used to transmit bit streams containing encoded point cloud data.

[0392] The V-PCC system according to the implementation can create containers that include point cloud data, and can also add additional data required for effective transmission / reception to the containers.

[0393] The receiving device according to the embodiment can be based on ​ The system shown is used to receive and parse containers. ​ The corresponding receiving devices can decode the parsed bitstream and recover the point cloud data.

[0394] The diagram illustrates the overall architecture for storing or streaming point cloud data compressed using Video-Based Point Cloud Compression (V-PCC). The processing for storing and streaming point cloud data may include acquisition processing, encoding processing, transmission processing, decoding processing, rendering processing, and / or feedback processing.

[0395] The implementation proposes a method for efficiently providing point cloud media / content / data.

[0396] To efficiently deliver point cloud media / content / data, the point cloud acquirer 20000 can acquire point cloud video. For example, one or more cameras can acquire point cloud data by capturing, orchestrating, or generating point clouds. This acquisition process allows the acquisition of point cloud video that includes the 3D positions of each point (represented by x, y, and z position values, etc.) (hereinafter referred to as geometry) and the attributes of each point (color, reflectivity, transparency, etc.). For example, a Polygon file format (PLY) (or Stanford triangle format) file containing the point cloud video can be generated. For point cloud data with multiple frames, one or more files can be acquired. During this process, point cloud-related metadata (e.g., metadata related to the capture, etc.) can be generated.

[0397] Captured point cloud videos may require post-processing to improve content quality. During video capture processing, the maximum / minimum depth can be adjusted within the range provided by the camera device. Even after adjustment, unwanted areas of point data may still exist. Therefore, post-processing can be performed to remove unwanted areas (e.g., background) or to identify connected spaces and fill in spatial holes. Additionally, point clouds extracted from cameras in a shared spatial coordinate system can be integrated into a single content by transforming individual points to a global coordinate system based on the position coordinates of each camera obtained through calibration processing. This results in a point cloud video with high-density points.

[0398] The point cloud preprocessor 20001 can generate one or more frames / pictures of a point cloud video. Here, a frame / picture can typically represent a unit representing an image at specific time intervals. When the points constituting the point cloud video are divided into one or more patches (a set of points constituting the point cloud video, where points belonging to the same patch are adjacent to each other in 3D space and are mapped in the same direction between the flat faces of a 6-sided bounding box when mapped to a 2D image) and mapped to a 2D plane, a binary occupancy map frame / picture can be generated, indicating the presence of data at the corresponding location in the 2D plane with a value of 0 or 1. Additionally, a geometric frame / picture in the form of a depth map representing information about the position (geometry) of each point constituting the point cloud video per patch can be generated. A texture frame / picture representing color information about each point constituting the point cloud video per patch can also be generated. In this process, metadata required to reconstruct the point cloud from the individual patches can be generated. The metadata may include information about the patches, such as the position and size of each patch in 2D / 3D space. These images / frames can be generated sequentially in time to construct a video stream or metadata stream.

[0399] The point cloud video encoder 20002 can encode one or more video streams associated with point cloud video. A video may include multiple frames, and a frame may correspond to a still image / picture. In this disclosure, point cloud video may include point cloud images / frames / pictures, and the term "point cloud video" is used interchangeably with point cloud video / frames / pictures. The point cloud video encoder can perform video-based point cloud compression (V-PCC) processing. For compression and encoding efficiency, the point cloud video encoder can perform a series of processes such as prediction, transform, quantization, and entropy coding. The encoded data (encoded video / image information) may be output as a bitstream. Based on V-PCC processing, as described below, the point cloud video encoder can encode point cloud video by dividing it into geometric video, attribute video, occupancy map video, and metadata (e.g., information about patches). Geometric video may include geometric images, attribute video may include attribute images, and occupancy map video may include occupancy map images. Patch data as auxiliary information may include patch-related information. Attribute video / images may include texture video / images.

[0400] The point cloud image encoder 20003 can encode one or more images associated with a point cloud video. The point cloud image encoder can perform video-based point cloud compression (V-PCC) processing. For compression and encoding efficiency, the point cloud image encoder can perform a series of processes such as prediction, transform, quantization, and entropy coding. The encoded image can be output as a bitstream. Based on V-PCC processing, as described below, the point cloud image encoder can encode the point cloud image by dividing it into a geometric image, an attribute image, an occupancy map image, and metadata (e.g., information about patches).

[0401] The point cloud video encoder and / or point cloud image encoder according to the embodiments can generate PCC bitstreams (G-PCC and / or V-PCC bitstreams) according to the embodiments.

[0402] According to the implementation, the video encoder 2002, the image encoder 20003, the video decoder 20006, and the image decoder can be performed by a single encoder / decoder as described above, and can be performed along separate paths as shown in the figure.

[0403] In file / fragment encapsulation 20004, encoded point cloud data and / or point cloud-related metadata can be encapsulated into files or fragments for streaming. Here, the point cloud-related metadata can be received from a metadata processor, etc. The metadata processor can be included in the point cloud video / image encoder or can be configured as a separate component / module. The encapsulation processor can encapsulate the corresponding video / image / metadata in a file format such as ISOBMFF or in the form of DASH fragments, etc. According to an embodiment, the encapsulation processor can include point cloud metadata in a file format. Point cloud-related metadata can be included in various levels of frames in, for example, ISOBMFF file format, or as data in a separate track within a file. According to an embodiment, the encapsulation processor can encapsulate point cloud-related metadata into a file.

[0404] The encapsulation or encapsulator according to the implementation can divide the G-PCC / V-PCC bitstream into one or more tracks and store them in a file, and can also encapsulate signaling information used for this operation. Additionally, atlas streams included in the G-PCC / V-PCC bitstream can be stored as tracks in the file, and related signaling information can be stored therein. Furthermore, SEI messages present in the G-PCC / V-PCC bitstream can be stored in tracks in the file, and related signaling information can be stored therein.

[0405] The transmission processor can perform transmission processing of encapsulated point cloud data according to the file format. The transmission processor can be included in the transmitter or configured as a separate component / module. The transmission processor can process the point cloud data according to the transmission protocol. Transmission processing can include processing via broadcast network and processing via broadband transmission. According to one implementation, the transmission processor can receive point cloud-related metadata and point cloud data from a metadata processor and perform transmission processing of point cloud video data.

[0406] The transmitter can send a point cloud bitstream or a file / fragment including the bitstream to a receiver of the receiving device via a digital storage medium or network. For transmission, processing according to any transmission protocol can be performed. Data processed for transmission can be transmitted via a broadcast network and / or broadband. Data can be transmitted to the receiving side on demand. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitter may include elements for generating media files in a predetermined file format and may include elements for transmission via a broadcast / communication network. The receiver can extract the bitstream and send the extracted bitstream to a decoder.

[0407] The receiver can receive point cloud data transmitted by the point cloud data transmitting device according to this disclosure. Depending on the transmission channel, the receiver can receive the point cloud data via a broadcast network or via broadband. Alternatively, the point cloud data can be received via a digital storage medium. The receiver may include processing for decoding the received data and rendering the data according to the user's viewport.

[0408] The receiving processor can perform processing on the received point cloud video data according to the transmission protocol. The receiving processor can be included in the receiver or configured as a separate component / module. The receiving processor can conversely perform the processing of the transmitting processor described above, corresponding to the transmission processing performed on the transmitting side. The receiving processor can transmit the acquired point cloud video to the decapsulation processor and transmit the acquired point cloud-related metadata to the metadata parser.

[0409] The decapsulation processor (file / fragment decapsulation) 20005 decapsulates point cloud data received from the receiving processor in file form. The decapsulation processor can decapsulate files according to ISOBMFF, etc., and can acquire point cloud bitstreams or point cloud-related metadata (or separate metadata bitstreams). The acquired point cloud bitstream can be transmitted to the point cloud decoder, and the acquired point cloud video-related metadata (metadata bitstream) can be transmitted to the metadata processor. The point cloud bitstream may include metadata (metadata bitstream). The metadata processor may be included in the point cloud decoder or can be configured as a separate component / module. The point cloud video-related metadata acquired by the decapsulation processor may take the form of frames or tracks in a file format. If necessary, the decapsulation processor can receive metadata required for decapsulation from the metadata processor. The point cloud-related metadata may be transmitted to the point cloud decoder and used in point cloud decoding processing, or it may be transmitted to the renderer and used in point cloud rendering processing.

[0410] The point cloud video decoder 20006 can receive bitstreams and decode video / images by performing operations corresponding to those of the point cloud video encoder. In this case, as described below, the point cloud video decoder can decode the point cloud video by dividing it into geometric video, attribute video, occupancy map video, and auxiliary patch information. The geometric video may include geometric images, the attribute video may include attribute images, and the occupancy map video may include occupancy map images. The auxiliary information may include auxiliary patch information. The attribute video / image may include texture video / images.

[0411] 3D geometry can be reconstructed based on decoded geometric images, occupancy maps, and auxiliary patch information, and then subjected to smoothing. Color point cloud images / pictures can be reconstructed by assigning color values ​​to the smoothed 3D geometry based on texture images. The renderer can render the reconstructed geometry and color point cloud images / pictures. The rendered video / images can be displayed on a monitor. All or part of the rendering results can be displayed to the user via VR / AR displays or typical displays.

[0412] The sensor / tracker (sensing / tracking) 20007 acquires orientation information and / or user viewport information from the user or receiving side and transmits the orientation information and / or user viewport information to the receiver and / or transmitter. Orientation information may represent information about the position, angle, movement, etc., of the user's head, or information about the position, angle, movement, etc., of the device through which the user is viewing video / images. Based on this information, information about the area currently being viewed by the user in 3D space (i.e., viewport information) can be calculated.

[0413] Viewport information can be information about the area in 3D space that the user is currently viewing through the device or HMD. Devices such as displays can extract the viewport area based on orientation information, the vertical or horizontal field of view supported by the device, etc. Orientation or viewport information can be extracted or calculated on the receiving side. The orientation or viewport information analyzed on the receiving side can be transmitted to the transmitting side on the feedback channel.

[0414] Based on orientation information acquired by sensors / trackers and / or viewport information indicating the area currently being viewed by the user, the receiver can effectively extract or decode media data from a file only for a specific area (i.e., the area indicated by the orientation information and / or viewport information). Additionally, based on the orientation information and / or viewport information acquired by sensors / trackers, the transmitter can effectively encode, or generate and transmit its file only for media data in that specific area (i.e., the area indicated by the orientation information and / or viewport information).

[0415] The renderer can render decoded point cloud data in 3D space. The rendered video / images can be displayed on a monitor. Users can view all or part of the rendering results through VR / AR displays or typical monitors.

[0416] Feedback processing may include transmitting various feedback information, which can be obtained during rendering / display processing, to a decoder on the sending or receiving side. Feedback processing enables interactivity when consuming point cloud data. According to one implementation, head orientation information, viewport information indicating the area the user is currently viewing, etc., may be transmitted to the sending side during feedback processing. According to another implementation, the user can interact with content implemented in a VR / AR / MR / autonomous driving environment. In this case, information related to the interaction may be transmitted to the sending side or service provider during feedback processing. According to yet another implementation, feedback processing may be skipped.

[0417] According to the implementation method, the aforementioned feedback information can be sent not only to the sending side but also consumed at the receiving side. That is, the decapsulation, decoding, and rendering processes at the receiving side can be performed based on the aforementioned feedback information. For example, point cloud data about the area currently being viewed by the user can be preferentially decapsulated, decoded, and rendered based on orientation information and / or viewport information.

[0418] The method for transmitting point cloud data according to the embodiments may include: encoding the point cloud data; encapsulating the point cloud data; and transmitting the point cloud data.

[0419] The method for receiving point cloud data according to the embodiments may include: receiving point cloud data; decapsulating the point cloud data; and decoding the point cloud data.

[0420] ​ This is an exemplary block diagram of an apparatus for storing and transmitting point cloud data according to an embodiment.

[0421] ​ A point cloud system according to an embodiment is shown. Part / all of the system may include... ​ Transmitting and receiving devices, ​ Encoding processing, ​ 2D video / image encoder, ​ Decoding processing, ​ The transmitting device and / or ​ Some or all of the receiving devices. Additionally, it may include or correspond to... ​ Part / all of the system.

[0422] The point cloud data transmission device according to the embodiment can be configured as shown in the figure. The various components of the transmission device can be modules / units / components / hardware / software / processors.

[0423] The geometry, attributes, auxiliary data, and mesh data of a point cloud can each be configured as a separate stream or stored in different tracks within a file. Furthermore, they can be included in separate fragments.

[0424] The point cloud acquirer (point cloud acquisition) 21000 acquires point clouds. For example, one or more cameras can acquire point cloud data by capturing, arranging, or generating point clouds. This acquisition process acquires point cloud data including the 3D positions of each point (represented by x, y, and z position values, etc.) (hereinafter referred to as geometry) and the attributes of each point (color, reflectivity, transparency, etc.). For example, a Polygon file format (PLY) (or Stanford triangle format) file containing the point cloud data can be generated. For point cloud data with multiple frames, one or more files can be acquired. During this process, point cloud-related metadata (e.g., metadata related to the capture, etc.) can be generated.

[0425] The patch generator (or patch generator) 21002 generates patches from point cloud data. The patch generator generates one or more frames from point cloud data or point cloud video. A frame typically represents a unit representing an image at specific time intervals. When the points constituting a point cloud video are divided into one or more patches (a set of points constituting the point cloud video, where points belonging to the same patch are adjacent to each other in 3D space and mapped in the same direction between the flat faces of a 6-sided bounding box when mapped to a 2D image) and mapped to a 2D plane, a binary occupancy map frame is generated, indicating the presence of data at the corresponding location in the 2D plane with 0 or 1. Additionally, a geometric frame in the form of a depth map representing the position (geometry) of each point constituting the point cloud video per patch can be generated. A texture frame representing the color information of each point constituting the point cloud video per patch can also be generated. In this process, metadata required to reconstruct the point cloud from the individual patches can be generated. Metadata can include information about the patches, such as the location and size of each patch in 2D / 3D space. These images / frames can be generated sequentially in time to construct a video stream or a metadata stream.

[0426] Additionally, patches can be used for 2D image mapping. For example, point cloud data can be projected onto the faces of a cube. After patch generation, geometric images, one or more attribute images, occupancy maps, auxiliary data, and / or mesh data can be generated based on the generated patches.

[0427] Geometric image generation, attribute image generation, occupancy map generation, auxiliary data generation, and / or mesh data generation are performed by a preprocessor or controller.

[0428] In geometry image generation 21002, a geometry image is generated based on the results of patch generation. The geometry represents points in 3D space. An occupancy map is used to generate the geometry image, which includes information related to the packing of 2D images of the patches, auxiliary data (pattern data), and / or patch-based mesh data. The geometry image is related to information such as the depth (e.g., near, far) of the patches generated after patch generation.

[0429] In attribute image generation 21003, an attribute image is generated. For example, an attribute may represent a texture. The texture may be a color value matched to individual points. According to an embodiment, an image including multiple attributes of the texture (e.g., color and reflectivity) (N attributes) may be generated. The multiple attributes may include material information and reflectivity. According to an embodiment, the attributes may additionally include information indicating color, which may vary depending on the viewing angle and light, even for the same texture.

[0430] In occupancy map generation 21004, an occupancy map is generated from the patch. The occupancy map includes information indicating whether data exists in pixels (e.g., corresponding to a geometric or attribute image).

[0431] In the auxiliary data generation 21005, auxiliary data including information about the patch is generated. That is, the auxiliary data represents metadata about the patch of the point cloud object. For example, it may represent information such as the patch's normal vector. Specifically, the auxiliary data may include information needed to reconstruct the point cloud from the patch (e.g., information about the patch's position, size, etc. in 2D / 3D space, as well as projection (normal) plane identification information, patch mapping information, etc.).

[0432] In grid data generation 21006, grid data is generated from patches. A grid represents the connection between neighboring points. For example, it can represent triangular-shaped data. For instance, grid data refers to the connectivity between points.

[0433] The point cloud preprocessor or controller generates metadata related to patch generation, geometric image generation, attribute image generation, occupancy map generation, auxiliary data generation, and mesh data generation.

[0434] The point cloud transmitting device performs video encoding and / or image encoding in response to the results generated by the preprocessor. The point cloud transmitting device can generate point cloud image data and point cloud video data. According to embodiments, the point cloud data may contain only video data, only image data, and / or both video data and image data.

[0435] The video encoder 21007 performs geometric video compression, attribute video compression, occupancy graph compression, auxiliary data compression, and / or mesh data compression. The video encoder generates a video stream containing encoded video data.

[0436] Specifically, in geometric video compression, point cloud geometric video data is encoded. In attribute video compression, point cloud attribute video data is encoded. In auxiliary data compression, auxiliary data associated with the point cloud video data is encoded. In mesh data compression, the mesh data of the point cloud video data is encoded. The various operations of the point cloud video encoder can be executed in parallel.

[0437] The image encoder 21008 performs geometric image compression, attribute image compression, occupancy map compression, auxiliary data compression, and / or mesh data compression. The image encoder generates an image containing encoded image data.

[0438] Specifically, in geometric image compression, point cloud geometric image data is encoded. In attribute image compression, point cloud attribute image data is encoded. In auxiliary data compression, auxiliary data associated with the point cloud image data is encoded. In mesh data compression, mesh data associated with the point cloud image data is encoded. The various operations of the point cloud image encoder can be executed in parallel.

[0439] Video encoders and / or image encoders may receive metadata from the preprocessor. Based on this metadata, video encoders and / or image encoders may perform individual encoding processes.

[0440] The File / Fragment Encapsulator (File / Fragment Encapsulator) 21009 encapsulates video streams and / or images as files and / or fragments. The File / Fragment Encapsulator performs video track encapsulation, metadata track encapsulation, and / or image encapsulation.

[0441] In video track encapsulation, one or more video streams can be encapsulated into one or more tracks.

[0442] In metadata track encapsulation, metadata related to the video stream and / or images can be encapsulated in one or more tracks. Metadata includes data related to the content of the point cloud data. For example, it may include initial viewing orientation metadata. Depending on the implementation, metadata may be encapsulated into a metadata track, or it may be encapsulated together in a video track or an image track.

[0443] In image encapsulation, one or more images can be encapsulated into one or more tracks or projects.

[0444] For example, according to an implementation, when four video streams and two images are input to the encapsulator, the four video streams and two images can be encapsulated in one file.

[0445] The point cloud video encoder and / or point cloud image encoder according to the embodiments can generate G-PCC / V-PCC bitstreams according to the embodiments.

[0446] The file / fragment wrapper can receive metadata from the preprocessor. The file / fragment wrapper can perform wrapping based on the metadata.

[0447] Files and / or fragments generated through file / fragment encapsulation are sent by a point cloud sending device or transmitter. For example, fragments may be transmitted according to a DASH-based protocol.

[0448] The encapsulation or encapsulator according to the implementation can divide a V-PCC bitstream into one or more tracks and store it in a file, and can also encapsulate signaling information used for this operation. Additionally, atlas streams included on the V-PCC bitstream can be stored as tracks in the file, and related signaling information can be stored therein. Furthermore, SEI messages present in the V-PCC bitstream can be stored in tracks in the file, and related signaling information can be stored therein.

[0449] The transmitter can send a point cloud bitstream or a file / fragment including the bitstream to a receiver of the receiving device via a digital storage medium or network. For transmission, processing according to any transmission protocol can be performed. Processed data can be transmitted via a broadcast network and / or broadband. Data can be transmitted to the receiving side on demand. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitter can include elements for generating media files in a predetermined file format and can include elements for transmission via a broadcast / communication network. The transmitter receives orientation information and / or viewport information from the receiver. The transmitter can transmit the acquired orientation information and / or viewport information (or user-selected information) to a preprocessor, video encoder, image encoder, file / fragment encapsulator, and / or point cloud encoder. Based on the orientation information and / or viewport information, the point cloud encoder can encode all point cloud data or the point cloud data indicated by the orientation information and / or viewport information. Based on orientation information and / or viewport information, the file / fragment encapsulator can encapsulate all point cloud data or point cloud data indicated by the orientation information and / or viewport information. Based on orientation information and / or viewport information, the transmitter can transmit all point cloud data or point cloud data indicated by the orientation information and / or viewport information.

[0450] For example, a preprocessor may perform the above operations on all point cloud data or on point cloud data indicated by orientation information and / or viewport information. A video encoder and / or image encoder may perform the above operations on all point cloud data or on point cloud data indicated by orientation information and / or viewport information. A file / fragment encapsulator may perform the above operations on all point cloud data or on point cloud data indicated by orientation information and / or viewport information. A transmitter may perform the above operations on all point cloud data or on point cloud data indicated by orientation information and / or viewport information.

[0451] ​ This is an exemplary block diagram of a point cloud data receiving device according to an embodiment.

[0452] ​ A point cloud system according to an embodiment is shown. Part / all of the system may include... ​ Transmitting and receiving devices, ​ Encoding processing, ​2D video / image encoder, ​ Decoding processing, ​ The transmitting device and / or ​ Some or all of the receiving devices. Additionally, it may include or correspond to... ​ and ​ Part / all of the system.

[0453] The various components of the receiving device can be modules / units / components / hardware / software / processors. The transmitting client can receive point cloud data, point cloud bitstreams, or files / fragments, including bitstreams transmitted by the point cloud data transmitting device according to the embodiment. Depending on the channel used for transmission, the receiver can receive point cloud data via a broadcast network or via broadband. Alternatively, point cloud video data can be received via a digital storage medium. The receiver may include processing for decoding the received data and rendering the received data according to a user viewport. The receiving processor can perform processing on the received point cloud data according to a transmission protocol. The receiving processor may be included in the receiver or configured as a separate component / module. The receiving processor may conversely perform the processing of the transmitting processor described above to correspond to the transmission processing performed on the transmitting side. The receiving processor may transmit the acquired point cloud data to a decapsulation processor and transmit the acquired point cloud-related metadata to a metadata parser.

[0454] Sensors / trackers (sensing / tracking) acquire orientation and / or viewport information. The acquired orientation and / or viewport information can then be transmitted to a delivery client, a file / fragment decapsulator, and a point cloud decoder.

[0455] The delivery client can receive all point cloud data or point cloud data indicated by orientation information and / or viewport information based on orientation information and / or viewport information. The file / fragment decapsulator can decapsulate all point cloud data or point cloud data indicated by orientation information and / or viewport information based on orientation information and / or viewport information. The point cloud decoder (video decoder and / or image decoder) can decode all point cloud data or point cloud data indicated by orientation information and / or viewport information based on orientation information and / or viewport information. The point cloud processor can process all point cloud data or point cloud data indicated by orientation information and / or viewport information based on orientation information and / or viewport information.

[0456] The file / fragment decapsulator (file / fragment decapsulator) 22000 performs video track decapsulation, metadata track decapsulation, and / or image decapsulation. The decapsulation processor (file / fragment decapsulator) can decapsulate point cloud data received from the receiving processor as a file. The decapsulation processor (file / fragment decapsulator) can decapsulate files or fragments according to ISOBMFF, etc., to obtain point cloud bitstreams or point cloud-related metadata (or separate metadata bitstreams). The obtained point cloud bitstream can be transmitted to the point cloud decoder, and the obtained point cloud-related metadata (or metadata bitstream) can be transmitted to the metadata processor. The point cloud bitstream may include metadata (metadata bitstream). The metadata processor may be included in the point cloud video decoder or can be configured as a separate component / module. The point cloud-related metadata obtained by the decapsulation processor may take the form of boxes or tracks in a file format. If necessary, the decapsulation processor can receive the metadata required for decapsulation from the metadata processor. The point cloud-related metadata may be transmitted to the point cloud decoder and used in the point cloud decoding process, or it may be transmitted to the renderer and used in the point cloud rendering process. The file / fragment decapsulator can generate metadata related to point cloud data.

[0457] In video track decapsulation, video tracks contained in files and / or segments are decapsulated. Video streams including geometric video, attribute video, occupancy maps, auxiliary data, and / or mesh data are decapsulated.

[0458] During metadata track decapsulation, the bitstream containing metadata related to point cloud data and / or auxiliary data is decapsulated.

[0459] In image decapsulation, images including geometric images, attribute images, occupancy maps, auxiliary data, and / or grid data are decapsulated.

[0460] According to the implementation method, the decapsulation or decapsulator can divide and parse (decapsulate) the G-PCC / V-PCC bitstream based on one or more tracks in the file, and can also decapsulate its signaling information. Additionally, atlas streams included in the G-PCC / V-PCC bitstream can be decapsulated based on tracks in the file, and related signaling information can be parsed. Furthermore, SEI messages present in the G-PCC / V-PCC bitstream can be decapsulated based on tracks in the file, and related signaling information can also be obtained.

[0461] The video decoder or video decoder 22001 performs geometric video decompression, attribute video decompression, occupancy map decompression, auxiliary data decompression, and / or mesh data decompression. The video decoder decodes the geometric video, attribute video, auxiliary data, and / or mesh data in a process corresponding to the process performed by the video encoder of the point cloud transmitting apparatus according to the embodiment.

[0462] Image decoder 22002 performs geometric image decompression, attribute image decompression, occupancy map decompression, auxiliary data decompression, and / or mesh data decompression. The image decoder decodes the geometric image, attribute image, auxiliary data, and / or mesh data in a process corresponding to the process performed by the image encoder of the point cloud transmitting apparatus according to the embodiment.

[0463] According to the implementation method, video decoding and image decoding can be processed by a single video / image decoder as described above, and can be performed along separate paths as shown in the figure.

[0464] Video decoding and / or image decoding can generate metadata related to video data and / or image data.

[0465] The point cloud video encoder and / or point cloud image encoder according to the embodiments can decode G-PCC / V-PCC bitstreams according to the embodiments.

[0466] In point cloud processing 22003, geometric reconstruction and / or attribute reconstruction are performed.

[0467] In geometric reconstruction, geometric video and / or geometric images are reconstructed from decoded video data and / or decoded image data based on occupancy maps, auxiliary data, and / or grid data.

[0468] In attribute reconstruction, attribute videos and / or attribute images are reconstructed from decoded attribute videos and / or decoded attribute images based on occupancy maps, auxiliary data, and / or mesh data. According to an implementation, the attribute may be a texture. According to an implementation, the attribute may represent multiple attribute pieces of information. When multiple attributes exist, the point cloud processor according to the implementation performs multiple attribute reconstructions.

[0469] The point cloud processor can receive metadata from video decoders, image decoders, and / or file / fragment decapsulators, and process point clouds based on the metadata.

[0470] Point cloud rendering or point cloud renderer rendering reconstructed point clouds. The point cloud renderer can receive metadata from video decoders, image decoders and / or file / fragment decapsulators, and render point clouds based on the metadata.

[0471] The display will actually show the rendered results on the screen.

[0472] like ​ As shown, in the method / apparatus according to the embodiment, such as ​ As shown, after encoding / decoding the point cloud data, the bit stream containing the point cloud data can be encapsulated and / or decapsulated in the form of files and / or fragments.

[0473] For example, the point cloud data apparatus according to the embodiments may encapsulate point cloud data based on files. The files may include V-PCC tracks containing point cloud parameters, geometry tracks containing geometry, attribute tracks containing attributes, and occupancy tracks containing occupancy maps.

[0474] Furthermore, the point cloud data receiving device according to the embodiment decapsulates point cloud data based on a file. The file may include a V-PCC track containing point cloud parameters, a geometry track containing geometry, an attribute track containing attributes, and an occupancy track containing an occupancy map.

[0475] The above operations can be performed by ​ File / fragment wrapper 20004, 20005, ​ File / fragment wrapper 21009 and ​ The file / fragment decapsulator 22000 is executed.

[0476] ​ An exemplary structure is shown that can be combined with the point cloud data transmission / reception method / apparatus according to the embodiments.

[0477] In the structure according to the embodiment, at least one of the following is connected to the cloud network 2300: server 2360, robot 2310, self-driving vehicle 2320, XR device 2330, smartphone 2340, home appliance 2350, and / or head-mounted display (HMD) 2370. Here, robot 2310, self-driving vehicle 2320, XR device 2330, smartphone 2340, or home appliance 2350 may be referred to as a device. Additionally, XR device 1730 may correspond to a point cloud data (PCC) device according to the embodiment, or may be operatively connected to a PCC device.

[0478] Cloud network 2300 can refer to a network that forms part of or exists within a cloud computing infrastructure. Here, cloud network 2300 can be configured using a 3G network, a 4G or Long Term Evolution (LTE) network, or a 5G network.

[0479] Server 2360 can be connected via cloud network 2300 to at least one of robot 2310, self-driving vehicle 2320, XR device 2330, smartphone 2340, home appliance 2350 and / or HMD 2370, and can assist at least a portion of the processing of the connected devices 2310 to 2370.

[0480] HMD 2370 represents one of the implementation types of an XR device and / or PCC device according to an embodiment. An HMD-type device according to an embodiment includes a communication unit, a control unit, a memory, an I / O unit, a sensor unit, and a power supply unit.

[0481] Hereinafter, various embodiments of apparatuses 2310 to 2350 to which the above technologies are applied will be described. ​ The apparatuses 2310 to 2350 shown can be operatively connected / linked to the point cloud data sending and receiving apparatus according to the above embodiments.

[0482]

[0483] The XR / PCC apparatus 2330 can adopt PCC technology and / or XR (AR + VR) technology, and can be implemented as an HMD, a head-up display (HUD) provided in a vehicle, a television, a mobile phone, a smartphone, a computer, a wearable device, a household appliance, a digital signage, a vehicle, a stationary robot, or a mobile robot.

[0484] The XR / PCC apparatus 2330 can analyze 3D point cloud data or image data obtained through various sensors or from external apparatuses and generate position data and attribute data regarding 3D points. Thereby, the XR / PCC apparatus 2330 can obtain information regarding the surrounding space or real objects, and render and output XR objects. For example, the XR / PCC apparatus 2330 can cause an XR object including auxiliary information regarding the identified object to match the identified object and output the matched XR object.

[0485] <PCC + XR + Mobile Phone>

[0486] The XR / PCC apparatus 2330 can be implemented as a mobile phone 2340 by applying PCC technology.

[0487] The mobile phone 2340 can decode and display point cloud content based on PCC technology.

[0488]

[0489] The self-driving vehicle 2320 can be implemented as a mobile robot, a vehicle, an unmanned aerial vehicle, etc. by applying PCC technology and XR technology.

[0490] The self-driving vehicle 2320 to which XR / PCC technology is applied can represent an autonomous vehicle provided with means for providing an XR image, or an autonomous vehicle as a control / interaction target in an XR image. Specifically, the self-driving vehicle 2320 as a control / interaction target in an XR image can be distinguished from the XR apparatus 2330 and can be operatively connected thereto.

[0491] The autonomous vehicle 2320, equipped with means for providing XR / PCC images, can acquire sensor information from sensors including cameras and output generated XR / PCC images based on the acquired sensor information. For example, the autonomous vehicle may have a HUD and output XR / PCC images to it to provide passengers with XR / PCC objects corresponding to real objects or objects present in the image.

[0492] In this scenario, when an XR / PCC object is output to a HUD, at least a portion of the XR / PCC object can be output to overlap with the real object being pointed at by the passenger's eyes. Conversely, when an XR / PCC object is output to a display provided inside the autonomous vehicle, at least a portion of the XR / PCC object can be output to overlap with objects on the screen. For example, the autonomous vehicle can output XR / PCC objects corresponding to objects such as roads, other vehicles, traffic lights, traffic signs, two-wheeled vehicles, pedestrians, and buildings.

[0493] Virtual reality (VR), augmented reality (AR), mixed reality (MR), and / or point cloud compression (PCC) technologies according to the implementation methods are applicable to various devices.

[0494] In other words, VR technology is a display technology that only provides real-world objects, backgrounds, etc., as CG images. On the other hand, AR technology refers to the technology of displaying CG images virtually created on top of real-world object images. MR technology is similar to AR technology in that the virtual objects to be displayed are mixed and combined with the real world. However, MR technology differs from AR technology. AR technology clearly distinguishes between real objects and virtual objects created as CG images and uses virtual objects as supplementary objects to real objects, while MR technology treats virtual objects as objects with the same characteristics as real objects. More specifically, an example of MR technology application is holographic services.

[0495] Recently, VR, AR, and MR technologies have often been referred to as extended reality (XR) technologies rather than clearly distinguished from each other. Therefore, the embodiments of this disclosure are applicable to all VR, AR, MR, and XR technologies. For these technologies, encoding / decoding based on PCC, V-PCC, and G-PCC technologies can be applied.

[0496] The PCC method / apparatus according to the embodiments can be applied to vehicles that provide autonomous driving services.

[0497] Vehicles providing autonomous driving services connect to the PCC device for wired / wireless communication.

[0498] When the point cloud data transmitting and receiving device (PCC device) according to the embodiment is connected to a vehicle for wired / wireless communication, the device can receive and process content data related to AR / VR / PCC services that can be provided with autonomous driving services and transmit the processed content data to the vehicle. When the point cloud data transmitting and receiving device is installed in the vehicle, the device can receive and process content data related to AR / VR / PCC services based on user input signals input through a user interface device and provide the processed content data to the user. The vehicle or user interface device according to the embodiment can receive user input signals. User input signals according to the embodiment may include signals indicating autonomous driving services.

[0499] The transmitting device according to the embodiment is configured to transmit point cloud data, and the receiving device according to the embodiment is configured to receive point cloud data.

[0500] The method / apparatus according to the embodiments represents a method / apparatus for transmitting and receiving point cloud data according to the embodiments, a point cloud encoder and decoder included in the transmitting / receiving device, a device configured to generate and parse data to transmit and receive point cloud data, a processor and / or a corresponding method.

[0501] A point cloud data transmission apparatus according to an embodiment may include a point cloud data encoder and a transmitter configured to transmit point cloud data. The point cloud data transmission apparatus may also include a point cloud data encapsulator capable of configuring the point cloud data in a format for efficient transmission. The encoder configured to compress the point cloud data and the encapsulator configured to perform encapsulation for transmission may be collectively referred to as a point cloud data system. In this specification, the above components may be simply referred to as the method / apparatus according to an embodiment.

[0502] A point cloud data receiving apparatus according to an embodiment may include a point cloud data decoder and a receiver configured to receive point cloud data. The point cloud data receiving apparatus may also include a decapsulator configured to parse point cloud data from a data structure in a format suitable for effective reception of the point cloud data. The decoder configured to recover the point cloud data and the decapsulator configured to perform decapsulation for reception / parsing may be collectively referred to as a point cloud data system. In this specification, the above components may be simply referred to as the method / apparatus according to an embodiment.

[0503] The video-based point cloud compression (V-PCC) described in this specification is the same as the visual volumetric video encoding (V3C). The terms V-PCC and V3C are used interchangeably according to the implementation and may have the same meaning.

[0504] The method / apparatus according to the implementation can generate a file format for dynamic point cloud objects and provide its signaling method (file encapsulation of dynamic point cloud objects).

[0505] ​ The structure of the encapsulated V-PCC data container according to an embodiment is shown.

[0506] ​ The encapsulated V-PCC data container structure according to an embodiment is shown.

[0507] ​ The point cloud video encoder 10002 of the transmitting device 10000 ​ and ​ encoder, ​ The transmitting device ​ Video / image encoders 20002 and 20003, ​ Processors and encoders 21000 to 21008 and ​ The XR device 2330 generates a bitstream containing point cloud data according to the embodiment.

[0508] ​ File / fragment wrapper 10003 ​ File / Fragment Wrapper 20004 ​ File / fragment wrapper 21009 and ​ The XR device formats the bitstream as ​ and ​ The file structure.

[0509] Similarly, ​ The receiving device 10005's file / fragment decapsulation module 10007, ​ File / fragment decapsulators 20005, 21004, and 22000 and ​ The XR device 2330 receives and decapsulates files and parses the bitstream. The bitstream is generated by... ​ Point cloud video decoder 10008 ​ and ​ decoder ​ The receiving device ​ Video / image decoders 20006, 22001 and 22002 and ​ The XR device 2330 decodes to recover point cloud data.

[0510] ​ and ​ The structure of a point cloud data container according to the ISOBMFF file format is shown.

[0511] ​ and​ The structure of a container for transporting point clouds based on multiple orbits is shown.

[0512] The method / apparatus according to the implementation can send / receive container files including point cloud data and additional data related to the point cloud data based on multiple tracks.

[0513] Track 1 24000 is an attribute track and can contain, for example... ​ , ​ , ​ , ​ The attribute data 24040 is shown in the figure.

[0514] Track 2 24010 is an occupied track and may contain, for example, ​ , ​ , ​ , ​ The geometric data 24050 is shown in the figure.

[0515] Orbit 3 24020 is a geometric orbital and may contain, for example, ​ , ​ , ​ , ​ The data occupancy shown in the figure is 24060.

[0516] Track 4 24030 is a v-pcc (v3c) track and may contain an atlas bitstream 27070 that includes data related to point cloud data.

[0517] Each track consists of sample entries and samples. A sample is a unit corresponding to a frame. To decode the Nth frame, a sample or sample entry corresponding to the Nth frame is required. A sample entry may contain information describing the sample.

[0518] ​ yes ​ Detailed structural diagram.

[0519] v3c track 25000 corresponds to track 4 24030. Data contained in v3c track 25000 may have a data container format called a box. v3c track 25000 contains reference information about V3C component tracks 25010 to 25030.

[0520] The receiving method / apparatus according to the embodiment can receive, such as ​ The container (which may be called a file) containing point cloud data is shown, and the V3C orbital is parsed. The occupancy data, geometric data, and attribute data can be decoded and recovered based on the reference information contained in the V3C orbital.

[0521] Occupied track 25010 corresponds to track 2 24010 and contains occupancy data. Geometric track 25020 corresponds to track 3 24020 and contains geometric data. Attribute track 25030 corresponds to track 1 24000 and contains attribute data.

[0522] The following will describe in detail ​ and ​ The syntax of the data structures included in the file.

[0523] ​ The structure of a bitstream containing point cloud data according to an embodiment is shown.

[0524] ​ The structure of a bitstream containing point cloud data to be encoded or decoded, according to an embodiment, is shown in reference to [reference]. ​ and ​ Described.

[0525] A bitstream of dynamic point cloud objects is generated according to the method / apparatus of the implementation. In this regard, a file format for the bitstream is proposed, and a signaling scheme for it is provided.

[0526] The method / apparatus according to the embodiments includes a transmitter, a receiver and / or a processor for providing a point cloud content service that efficiently stores a V-PCC (=V3C) bitstream in a file track and provides its signaling.

[0527] The method / apparatus according to the embodiments provides a data format for storing V-PCC bitstreams containing point cloud data. Therefore, the receiving method / apparatus according to the embodiments provides a data storage and signaling method for receiving point cloud data and efficiently accessing the point cloud data. Thus, based on the storage technology of the file containing point cloud data for efficient access, the transmitter and / or receiver can provide point cloud content services.

[0528] The method / apparatus according to the embodiments efficiently stores a point cloud bitstream (V-PCC bitstream) in a file track. It generates signaling information regarding the efficient storage technique and stores it in the file. To support efficient access to the V-PCC bitstream stored in the file, in addition to (or by modifying / combining) the file storage technique according to the embodiments, a technique for dividing the V-PCC bitstream into one or more tracks and storing it in the file can be provided.

[0529] The terms used in this document are defined as follows:

[0530] VPS: V-PCC parameter set; AD: Atlas data; OVD: Occupied video data; GVD: Geometric video data; AVD: Attribute video data; ACL: Atlas coding layer; AAPS: Atlas adaptation parameter set; ASPS: Atlas sequence parameter set, which may be a syntactic structure containing syntactic elements according to the implementation method, these syntactic elements are applied to zero or more complete coded atlas sequences (CAS) determined by the contents of syntactic elements found in the ASPS referenced by the syntactic elements found in the respective piece group headers.

[0531] AFPS: Atlas Frame Parameter Set, which may include a syntactic structure containing syntactic elements applied to zero or more complete coded atlas frames determined by the contents of the syntactic elements found in the tile group header.

[0532] SEI: Supplemental Enhancement Information.

[0533] Atlas: A collection of 2D bounding boxes, for example, patches projected onto rectangular frames corresponding to 3D bounding boxes in 3D space. Atlases can represent subsets of point clouds.

[0534] Atlas Sub-Bitstream: A sub-bitstream extracted from a V-PCC bitstream that contains a portion of the Atlas NAL bitstream.

[0535] V-PCC content: Point cloud based on V-PCC (V3C) encoding.

[0536] V-PCC track: The volumetric visual track of the atlas bitstream carrying the V-PCC bitstream.

[0537] V-PCC Component Track: A video track that carries 2D video encoded data from any of the V-PCC bitstream's occupancy graph, geometry, or attribute component video bitstreams.

[0538] This section describes an implementation scheme for supporting partial access to dynamic point cloud objects. The implementation includes atlas tile group information associated with some data of V-PCC objects included in various spatial regions at the file system level. Furthermore, the implementation includes an extended signaling scheme for tag and / or patch information included in each atlas tile group.

[0539] ​ The structure of the point cloud bitstream included in the data sent and received by the method / apparatus according to the embodiment is shown.

[0540] The method for compressing and decompressing point cloud data according to the implementation method refers to the volumetric encoding and decoding of point cloud visual information.

[0541] The point cloud bitstream containing the encoded point cloud sequence (CPCS) (which may be referred to as the V-PCC bitstream or V3C bitstream) 26000 may include sample stream V-PCC units 26010. The sample stream V-PCC unit 26010 may carry V-PCC parameter set (VPS) data 26020, atlas bitstream 26030, 2D video encoded occupancy graph bitstream 26040, 2D video encoded geometry bitstream 26050, and zero or one or more 2D video encoded attribute bitstreams 26060.

[0542] The point cloud bitstream 26000 may include the sample stream VPCC header 26070.

[0543] ssvh_unit_size_precision_bytes_minus1: The value obtained by adding 1 to this value specifies the precision (in bytes) of the ssvu_vpcc_unit_size elements in all sample stream V-PCC units. ssvh_unit_size_precision_bytes_minus1 can be in the range of 0 to 7.

[0544] The syntax 26080 of the sample stream V-PCC unit 26010 is configured as follows. Each sample stream V-PCC unit may include one of the V-PCC unit types of VPS, AD, OVD, GVD, and AVD. The content of each sample stream V-PCC unit may be associated with the same access unit as the V-PCC units included in the sample stream V-PCC unit.

[0545] ssvu_vpcc_unit_size: Specifies the size (in bytes) of the subsequent vpcc_unit. The number of bits used to represent ssvu_vpcc_unit_size is equal to (ssvh_unit_size_precision_bytes_minus1+1)*8.

[0546] Received according to the method / apparatus of the embodiment ​ A bitstream containing encoded point cloud data, and generated by wrapper 20004 or 21009 as follows. ​ and ​ The file shown.

[0547] According to the method / apparatus of the embodiment, receive such ​ and ​ The file shown is used to decode the point cloud data using a decapsulator such as 22000.

[0548] VPS26020 and / or AD 26030 are packaged in track 4 (V3C track) 24030.

[0549] OVD 26040 is encapsulated in track 2 (occupied track) 24010.

[0550] The GVD 26050 is packaged in track 3 (geometric track) 24020.

[0551] AVD 26060 is encapsulated in track 1 (attribute track) 24000.

[0552] ​ The configuration of the sample stream V-PCC unit according to an embodiment is shown.

[0553] ​ The bitstream 27000 corresponds to ​ The bitstream is 26000.

[0554] The sample stream V-PCC unit included in the bit stream 27000 associated with the point cloud data according to the embodiment may include V-PCC unit size 27010 and V-PCC unit 27020.

[0555] The abbreviations are defined as follows: VPS (V-PCC Parameter Set); AD (Atlas Data); OVD (Occupied Video Data); GVD (Geometric Video Data); AVD (Attribute Video Data).

[0556] Each V-PCC unit 27020 may include a V-PCC unit header 27030 and a V-PCC unit payload 27040. The V-PCC unit header 27030 may describe the V-PCC unit type. The V-PCC unit header for attribute video data may describe the attribute type, its index, multiple instances of the same attribute type supported, etc.

[0557] The cell payloads 27050, 27060, and 27070 for occupancy, geometry, and attribute video data may correspond to video data cells. For example, occupancy video data, geometry video data, and attribute video data 27050, 27060, and 27070 may be HEVC NAL cells. This video data may be decoded by a video decoder according to an embodiment.

[0558] ​ The V-PCC unit and V-PCC unit header according to an embodiment are shown.

[0559] ​ The above reference is shown. ​ The syntax of V-PCC unit 27020 and V-PCC unit header 27030 is described.

[0560] According to the implementation method, the V-PCC bitstream may contain a series of V-PCC sequences.

[0561] The value of `vuh_unit_type` equals the vpcc unit type of VPCC_VPS, which can be expected to be the first V-PCC unit type in the V-PCC sequence. All other V-PCC unit types follow this unit type without any additional restrictions on their encoding order. The V-PCC unit payload carrying V-PCC units that are occupies video, attribute video, or geometric video consists of one or more NAL units.

[0562] A VPCC unit may include a head and a payload.

[0563] The VPCC cell header can include the following information based on the VUH cell type.

[0564] The vuh_unit_type indicates the type of V-PCC unit 27020.

[0565] 0 ​ ​ ​ 1 ​ ​ ​ 2 ​ ​ ​ 3 ​ ​ ​ 4 ​ ​ ​ 5…31 ​ ​ -

[0566] When `vuh_unit_type` indicates attribute video data (VPCC_AVD), geometric video data (VPCC_GVD), occupancy video data (VPCC_OVD), or atlas data (VPCC_AD), `vuh_vpcc_parameter_setID` and `vuh_atlas_id` are carried in the unit header. The parameter set ID and atlas ID associated with the V-PCC unit can be transmitted.

[0567] When the cell type is atlas video data, the cell header can carry the attribute index (vuh_attribute_index), attribute partition index (vuh_attribute_partition_index), graph index (vuh_map_index), and auxiliary video flag (vuh_auxiliary_video_flag).

[0568] When the cell type is geometric video data, it can carry vuh_map_index and vuh_auxiliary_video_flag.

[0569] When the cell type is occupied video data or atlas data, the cell header may include additional reserved bits.

[0570] The `vuh_vpcc_parameter_set_id` specifies the value of the `vps_vpcc_parameter_set_id` of the active V-PCC VPS. The `vpcc_parameter_set_id` in the header of the current V-PCC unit reveals the ID of the VPS parameter set and allows declaration of the relationship between the V-PCC unit and the V-PCC parameter set.

[0571] The `vuh_atlas_id` specifies the index of the atlas corresponding to the current V-PCC cell. The index of the atlas can be determined from the `vuh_atlas_id` in the header of the current V-PCC cell, and the atlas corresponding to the V-PCC cell can be declared.

[0572] vuh_attribute_index indicates the index of the attribute data carried in the attribute video data unit.

[0573] vuh_attribute_partition_index indicates the index of the attribute dimension group carried in the attribute video data unit.

[0574] When present, vuh_map_index indicates the graph index of the current geometry or attribute flow.

[0575] A value of 1 for `vuh_auxiliary_video_flag` indicates that the associated geometric or attribute video data unit is only RAW and / or EOM coded point video. A value of 0 for `vuh_auxiliary_video_flag` indicates that the associated geometric or attribute video data unit may contain RAW and / or EOM coded points.

[0576] ​ The payload of a V-PCC unit according to an embodiment is shown.

[0577] ​ The syntax of the V-PCC unit payload 27040 is shown.

[0578] When vuh_unit_type is the V-PCC parameter set (VPCC_VPS), the V-PCC unit payload includes vpcc_parameter_set().

[0579] When vuh_unit_type is V-PCC atlas data (VPCC_AD), the V-PCC unit payload contains atlas_sub_bitstream().

[0580] When vuh_unit_type is V-PCC cumulative video data (VPCC_OVD), geometric video data (VPCC_GVD), or attribute video data (VPCC_AVD), the V-PCC unit payload includes video_sub_bitstream().

[0581] ​ The V-PCC parameter set according to the implementation method is shown.

[0582] ​This illustrates that when the payload 27040 of the bitstream unit 27020 according to the embodiment includes, as shown, ​ The parameter set shown is the syntax of the parameter set.

[0583] `profile_tier_level()` specifies limitations on the bitstream, and therefore limitations on the capabilities required to decode the bitstream. Profiles, tiers, and levels can also be used to indicate interoperability points between individual decoder implementations.

[0584] vps_vpcc_parameter_set_id provides an identifier for the V-PCC VPS for reference by other syntax elements.

[0585] Incrementing `vps_atlas_count_minus1` by 1 indicates the total number of atlases supported in the current bitstream.

[0586] Depending on the number of atlases, the following parameters may be further included in the parameter set.

[0587] `vps_frame_width[j]` indicates the V-PCC frame width based on integer luminance samples from the atlas at index `j`. This frame width is the nominal width associated with all V-PCC components in the atlas at index `j`.

[0588] `vps_frame_height[j]` indicates the V-PCC frame height based on integer luminance samples from the atlas at index `j`. This frame height is the nominal height associated with all V-PCC components in the atlas at index `j`.

[0589] The increment of `vps_map_count_minus1[j]` indicates the number of maps used to encode the geometric and attribute data of the atlas at index `j`.

[0590] When vps_map_count_minus1[j] is greater than 0, the following parameters may be further included in the parameter set.

[0591] Based on the value of vps_map_count_minus1[j], the following parameters may be further included in the parameter set.

[0592] A value of 0 for `vps_multiple_map_streams_present_flag[j]` indicates that all geometry or attribute maps of the atlas at index `j` are placed in a single geometry or attribute video stream. A value of 1 for `vps_multiple_map_streams_present_flag[j]` indicates that all geometry or attribute maps of the atlas at index `j` are placed in a separate video stream.

[0593] If vps_multiple_map_streams_present_flag[j] indicates 1, then vps_map_absolute_coding_enabled_flag[j][i] may be further included in the parameter set. Otherwise, vps_map_absolute_coding_enabled_flag[j][i] may have a value of 1.

[0594] `vps_map_absolute_coding_enabled_flag[j][i]` equal to 1 indicates that the geometry of index i in the atlas of index j is encoded without any form of graph prediction. `vps_map_absolute_coding_enabled_flag[j][i]` equal to 0 indicates that the geometry of index i in the atlas of index j is first predicted from another previously encoded graph before encoding.

[0595] The value of vps_map_absolute_coding_enabled_flag[j][0] equal to 1 indicates that the geometry at index 0 is encoded without graph prediction.

[0596] If vps_map_absolute_coding_enabled_flag[j][i] is 0 and i is greater than 0, then vps_map_predictor_index_diff[j][i] can be further included in the parameter set. Otherwise, vps_map_predictor_index_diff[j][i] can be 0.

[0597] When vps_map_absolute_coding_enabled_flag[j][i] equals 0, vps_map_predictor_index_diff[j][i] is used to compute the predictor of the geometry of the atlas at index i for index j.

[0598] A value of 1 for `vps_auxiliary_video_present_flag[j]` indicates that the auxiliary information of the atlas at index `j` (i.e., RAW or EOM patch data) can be stored in a separate video stream (called the auxiliary video stream). A value of 0 for `vps_auxiliary_video_present_flag[j]` indicates that the auxiliary information of the atlas at index `j` is not stored in a separate video stream.

[0599] occupancy_information() includes information related to video occupancy.

[0600] geometry_information() includes geometry-related video information.

[0601] attribute_information() includes video-related information.

[0602] A value of 1 for `vps_extension_present_flag` indicates that the syntax element `vps_extension_length` exists in the `vpcc_parameter_set` syntax structure. A value of 0 for `vps_extension_present_flag` indicates that the syntax element `vps_extension_length` does not exist.

[0603] Increasing `vps_extension_length_minus1` by 1 specifies the number of `vps_extension_data_byte` elements that follow this syntax element.

[0604] Based on vps_extension_length_minus1, the extension data can be further included in the parameter set.

[0605] vps_extension_data_byte can have any value.

[0606] ​ The diagram shows the tiles according to the implementation method.

[0607] ​ Showing includes by ​ Point cloud video encoder 10002 ​ encoder, ​ encoder, ​ The transmitting device ​ and ​ The atlas frame is composed of tiles encoded by systems such as [system name missing]. The atlas illustrates [something missing] composed of [something missing] ​ Point cloud video decoder 10008 ​ and ​ decoder ​ The receiving device ​ The system decodes the atlas frames of the mosaic pieces.

[0608] An atlas frame can be divided into one or more tile rows and one or more tile columns. A tile is a rectangular area within the atlas frame. A tile group comprises multiple tiles from the atlas frame. Tiles and tile groups can be indistinguishable from each other, and a tile group can correspond to a single tile. Only rectangular tile groups are supported. In this mode, a tile group (or tile) can collectively comprise multiple tiles from the atlas frame within a rectangular area. ​ This illustrates the division of atlas frames into tiles or groups of tiles according to an embodiment.​ The image frame shows an atlas frame comprising 24 tiles (6 tile columns and 4 tile rows) and 9 rectangular tile groups. According to an implementation, a tile group can be used as a term corresponding to a tile, without distinguishing between tile groups and tiles.

[0609] That is, according to the implementation method, tile group 31000 may correspond to tile 31010 and may be referred to as tile 31010. In addition, tile 31010 may correspond to tile partition and may be referred to as tile partition. The name of the signaling information may also be changed according to the complementary correspondence.

[0610] ​ The structure of the atlas bitstream according to an embodiment is shown.

[0611] ​ Show ​ Example of a sub-bit stream 32000 carrying a payload 27040 of a bit stream 27000, with unit 27020.

[0612] The payload of a V-PCC cell carrying a map sub-bit stream may include one or more sample streams NAL cells 32010.

[0613] According to the implementation, the atlas sub-bitstream 32000 may include a sample stream NAL header 32020 and one or more sample stream NAL units 32010.

[0614] The sample stream NAL header 32020 may include ssnh_unit_size_precision_bytes_minus1. ssnh_unit_size_precision_bytes_minus1, incremented by 1, specifies the precision (in bytes) of the ssnu_nal_unit_size elements in all sample stream NAL units. ssnh_unit_size_precision_bytes_minus1 can range from 0 to 7.

[0615] The sample stream NAL unit 32010 may include ssnu_nal_unit_size.

[0616] `ssnu_nal_unit_size` specifies the size (in bytes) of the subsequent `NAL_unit`. The number of bits used to represent `ssnu_nal_unit_size` can be equal to `(ssnh_unit_size_precision_bytes_minus1+1)*8`.

[0617] ​ The NAL unit according to an embodiment is shown.

[0618] ​ Show ​ Syntax of NAL units.

[0619] NAL units may include nal_unit_header() and (NumBytesInRbsp).

[0620] NumBytesInRbsp indicates the byte corresponding to the payload of the NAL unit, and its initial value is set to 0.

[0621] The nal_unit_header() function can include nal_forbidden_zero_bit, nal_unit_type, nal_layer_id, and nal_temporal_id_plus1.

[0622] nal_forbidden_zero_bit is a field used for error detection in NAL cells and should be 0.

[0623] nal_unit_type indicates, for example ​ The types of RBSP data structures included in the NAL unit shown.

[0624] nal_layer_id specifies the identifier of the layer to which the ACL NAL unit belongs or the identifier of the layer to which a non-ACL NAL unit is applied.

[0625] nal_temporal_id_plus1 minus 1 specifies the time identifier of the NAL unit.

[0626] ​ The type of NAL unit according to the implementation is shown.

[0627] ​ Show ​ The type of nal_unit_type included in the NAL unit header of the sample stream NAL unit 32010.

[0628] NAL_TRAIL: Encoded tile groups of non-TSA, non-STSA ending atlas frames can be included in NAL units. The RBSP syntax structure of a NAL unit is atlas_tile_group_layer_rbsp() or atlas_tile_layer_rbsp(). The type class of a NAL unit is ACL. Depending on the implementation, a tile group may correspond to a tile.

[0629] NAL TSA: Encoded tile groups of TSA atlas frames can be included in NAL units. The RBSP syntax structure of a NAL unit is atlas_tile_group_layer_rbsp() or atlas_tile_layer_rbsp(). The type class of a NAL unit is ACL.

[0630] NAL_STSA: Encoded tile groups of STSA atlas frames can be included in NAL units. The RBSP syntax structure of a NAL unit is atlas_tile_group_layer_rbsp() or atlas_tile_layer_rbsp(). The type class of a NAL unit is ACL.

[0631] NAL_RADL: Encoded tile groups of RADL atlas frames can be included in NAL units. The RBSP syntax structure of a NAL unit is atlas_tile_group_layer_rbsp() or atlas_tile_layer_rbsp(). The type class of a NAL unit is ACL.

[0632] NAL_RASL: Encoded tile groups of RASL atlas frames can be included in NAL units. The RBSP syntax structure of a NAL unit is atlas_tile_group_layer_rbsp() or aatlas_tile_layer_rbsp(). The type class of a NAL unit is ACL.

[0633] NAL_SKIP: Encoded tile groups that skip atlas frames can be included in NAL units. The RBSP syntax structure of a NAL unit is either atlas_tile_group_layer_rbsp() or atlas_tile_layer_rbsp(). The type class of a NAL unit is ACL.

[0634] NAL_RSV_ACL_6 to NAL_RSV_ACL_9: Reserved for non-IRAP ACLs. NAL cell types can be included in NAL cells. The type class of NAL cells is ACL.

[0635] NAL_BLA_W_LP, NAL_BLA_W_RADL, NAL_BLA_N_LP: Encoded tile groups of BLA atlas frames can be included in NAL units. The RBSP syntax structure of a NAL unit is atlas_tile_group_layer_rbsp() or atlas_tile_layer_rbsp(). The type class of a NAL unit is ACL.

[0636] NAL_GBLA_W_LP, NAL_GBLA_W_RADL, NAL_GBLA_N_LP: Encoded tile groups of GBLA atlas frames can be included in NAL units. The RBSP syntax structure of a NAL unit is atlas_tile_group_layer_rbsp() or atlas_tile_layer_rbsp(). The type class of a NAL unit is ACL.

[0637] NAL_IDR_W_RADL and NAL_IDR_N_LP: Encoded tile groups of IDR atlas frames can be included in NAL units. The RBSP syntax structure of a NAL unit is atlas_tile_group_layer_rbsp() or atlas_tile_layer_rbsp(). The type class of a NAL unit is ACL.

[0638] NAL_GIDR_W_RADL and NAL_GIDR_N_LP: Encoded tile groups of GIDR atlas frames can be included in NAL units. The RBSP syntax structure of a NAL unit is atlas_tile_group_layer_rbsp() or atlas_tile_layer_rbsp(). The type class of a NAL unit is ACL.

[0639] NAL_CRA: Encoded tile groups of CRA atlas frames can be included in NAL units. The RBSP syntax structure of a NAL unit is atlas_tile_group_layer_rbsp() or atlas_tile_layer_rbsp(). The type class of a NAL unit is ACL.

[0640] NAL_GCRA: Encoded tile groups of GCRA atlas frames can be included in NAL units. The RBSP syntax structure of a NAL unit is atlas_tile_group_layer_rbsp() or atlas_tile_layer_rbsp(). The type class of a NAL unit is ACL.

[0641] NAL_IRAP_ACL_22 and NAL_IRAP_ACL_23: Reserved IRAP ACL NAL cell types can be included in NAL cells. The type class of NAL cells is ACL.

[0642] NAL_RSV_ACL_24 to NAL_RSV_ACL_31: Reserved for non-IRAP ACLs. NAL cell types can be included in NAL cells. The type class of NAL cells is ACL.

[0643] NAL_ASPS: Atlas sequence parameter sets can be included in NAL units. The RBSP syntax structure of a NAL unit is atlas_sequence_parameter_set_rbsp(). The type class of a NAL unit is non-ACL.

[0644] NAL_AFPS: Atlas frame parameter sets can be included in NAL units. The RBSP syntax structure of a NAL unit is atlas_frame_parameter_set_rbsp(). The type class of a NAL unit is non-ACL.

[0645] NAL_AUD: The access unit delimiter can be included in a NAL unit. The RBSP syntax structure of a NAL unit is access_unit_delimiter_rbsp(). The type class of a NAL unit is non-ACL.

[0646] NAL_VPCC_AUD: The V-PCC access unit delimiter can be included in a NAL unit. The RBSP syntax structure of a NAL unit is access_unit_delimiter_rbsp(). The type class of a NAL unit is non-ACL.

[0647] NAL_EOS: The NAL unit type can be the end of a sequence. The RBSP syntax structure of the NAL unit is end_of_seq_rbsp(). The type class of the NAL unit is non-ACL.

[0648] NAL_EOB: The NAL unit type can be the end of the bitstream. The RBSP syntax structure for a NAL unit is end_of_atlas_sub_bitstream_rbsp(). The type class of the NAL unit is non-ACL.

[0649] NAL_FD Filler: The NAL cell type can be filler_data_rbsp(). The type class of the NAL cell is non-ACL.

[0650] NAL_PREFIX_NSEI, NAL_SUFFIX_NSEI: NAL unit types can be optional supplementary enhancement information. The RBSP syntax structure of a NAL unit is sei_rbsp(). The type class of a NAL unit is non-ACL.

[0651] NAL_PREFIX_ESEI, NAL_SUFFIX_ESEI: NAL unit types can be necessary supplementary enhancement information. The RBSP syntax structure of a NAL unit is sei_rbsp(). The type class of a NAL unit is non-ACL.

[0652] NAL_AAPS: The NAL cell type can be an atlas adaptation parameter set. The RBSP syntax structure of the NAL cell is atlas_adaptation_parameter_set_rbsp(). The type class of the NAL cell is non-ACL.

[0653] NAL_RSV_NACL_44 to NAL_RSV_NACL_47: NAL cell types can be reserved non-ACL NAL cell types. The type class of NAL cells is non-ACL.

[0654] NAL_UNSPEC_48 to NAL_UNSPEC_63: NAL cell types can be unspecified non-ACL NAL cell types. The type class of NAL cells is non-ACL.

[0655] ​ The atlas sequence parameter set according to the implementation method is shown.

[0656] ​ This shows the syntax of the RBSP data structures included in the NAL unit when the NAL unit type is a graph sequence parameter.

[0657] Each sample stream NAL unit may contain a set of atlas parameters (e.g., ASPS, AAPS, or AFPS), one or more atlas tile groups, and one of the SEIs.

[0658] ASPS may contain zero or more complete coded atlas sequences (CAS) whose contents are determined by the syntactic elements found in the ASPS referenced by the syntactic elements found in the header of each piece group.

[0659] asps_atlas_sequence_parameter_set_id can provide an identifier for the set of parameters of the atlas sequence for reference by other syntax elements.

[0660] asps_frame_width indicates the atlas frame width based on an integer number of samples, where the samples correspond to the luminance samples of the video components.

[0661] asps_frame_height indicates the atlas frame height based on an integer number of samples, where the samples correspond to the luminance samples of the video components.

[0662] asps_log2_patch_packing_block_size specifies the value of the variable PatchPackingBlockSize, which is used for the horizontal and vertical placement of patches within the atlas.

[0663] asps_log2_max_atlas_frame_order_cnt_lsb_minus4 specifies the value of the variable MaxAtlasFrmOrderCntLsb used for decoding processing of atlas frame order counting.

[0664] `asps_max_dec_atlas_frame_buffering_minus1` incremented by 1 specifies the maximum required size of the atlas frame buffer used for CAS decoding, in units of atlas frame storage buffers.

[0665] A value of 0 for `asps_long_term_ref_atlas_frames_flag` indicates that no long-term reference atlas frames are used for inter-frame prediction of any coded atlas frames in CAS. A value of 1 for `asps_long_term_ref_atlas_frames_flag` indicates that long-term reference atlas frames are available for inter-frame prediction of one or more coded atlas frames in CAS.

[0666] asps_num_ref_atlas_frame_lists_in_asps specifies the number of ref_list_struct(rlsIdx) syntax structures included in the atlas sequence parameter set.

[0667] When asps_num_ref_atlas_frame_lists_in_asps is greater than 0, atgh_ref_atlas_frame_list_sps_flag can be included in the atlas tile group (tile) header.

[0668] When asps_num_ref_atlas_frame_lists_in_asps is greater than 1, atgh_ref_atlas_frame_list_idx can be included in the atlas tile group (tile) header.

[0669] For each value of NumLtrAtlasFrmEntries, atgh_additional_afoc_lsb_present_flag[j] may be included in the atlas tile group (tile) header.

[0670] The atlas sequence parameter set may include as many ref_list_struct(i) as there are ASPS reference atlas frame lists (asps_num_ref_atlas_frame_lists_in_asps).

[0671] `asps_use_eight_orientations_flag` equal to 0 specifies that the patch orientation index `pdu_orientation_index[i][j]` of the patch at index `j` in frame `i` is in the range of 0 to 1 (inclusive). `asps_use_eight_orientations_flag` equal to 1 specifies that the patch orientation index `pdu_orientation_index[i][j]` of the patch at index `j` in frame `i` is in the range of 0 to 7 (inclusive).

[0672] `asps_extended_projection_enabled_flag` equal to 0 indicates that no signal is used to notify patch projection information for the current atlas tile group. `asps_extended_projection_enabled_flag` equal to 1 indicates that signal is used to notify patch projection information for the current atlas tile group.

[0673] When atgh_type is not SKIP_TILE_GRP, the following elements may be included in the atlas tile group (or tile) header.

[0674] An asps_normal_axis_limits_quantization_enabled_flag value of 1 specifies that the application signal notifies the quantization parameters and is used to quantize the normal axis-dependent elements of patch data units, merged patch data units, or inter-patch data units. If asps_normal_axis_limits_quantization_enabled_flag value of 0, quantization is not applied to any normal axis-dependent elements of patch data units, merged patch data units, or inter-patch data units.

[0675] When asps_normal_axis_limits_quantization_enabled_flag is 1, atgh_pos_min_z_quantizer can be included in the atlas tile group (or tile) header.

[0676] If `asps_normal_axis_max_delta_value_enabled_flag` is equal to 1, the maximum nominal shift value of the normal axis that can exist in the geometry of the patch at index i in the frame of index j will be indicated in the bitstream of the individual patch data unit, merged patch data unit, or inter-patch data unit. If `asps_normal_axis_max_delta_value_enabled_flag` is equal to 0, the maximum nominal shift value of the normal axis that can exist in the geometry of the patch at index i in the frame of index j should not be indicated in the bitstream of the individual patch data unit, merged patch data unit, or inter-patch data unit.

[0677] When asps_normal_axis_max_delta_value_enabled_flag is 1, atgh_pos_delta_max_z_quantizer can be included in the atlas tile group (or tile) header.

[0678] `asps_remove_duplicate_point_enabled_flag` equal to 1 indicates that duplicate points are not reconstructed for the current atlas, where a duplicate point is a point with the same 2D and 3D geometric coordinates as another point from a lower-indexed atlas. `asps_remove_duplicate_point_enabled_flag` equal to 0 indicates that all points are reconstructed.

[0679] `asps_max_dec_atlas_frame_buffering_minus1` incremented by 1 specifies the maximum required size of the atlas frame buffer used for CAS decoding, in units of atlas frame storage buffers.

[0680] `asps_pixel_deinterleaving_flag` equal to 1 indicates that the decoded geometry and attribute video of the current atlas contains spatially interleaved pixels from two atlases. `asps_pixel_deinterleaving_flag` equal to 0 indicates that the decoded geometry and attribute video corresponding to the current atlas contains pixels from only a single atlas.

[0681] `asps_patch_precedence_order_flag` equal to 1 indicates that the patch priority order of the current atlas is the same as the decoding order. `asps_patch_precedence_order_flag` equal to 0 indicates that the patch priority order of the current atlas is the reverse of the decoding order.

[0682] `asps_patch_size_quantizer_present_flag` equal to 1 indicates that the patch size quantizer parameter exists in the atlas patch group header. `asps_patch_size_quantizer_present_flag` equal to 0 indicates that the patch size quantizer parameter does not exist.

[0683] When asps_patch_size_quantizer_present_flag equals 1, atgh_patch_size_x_info_quantizer and atgh_patch_size_y_info_quantizer can be included in the header of the atlas patch group (or patch).

[0684] `asps_eom_patch_enabled_flag` equal to 1 indicates that the decoded occupancy map video of the current atlas contains information about whether intermediate depth positions between two depth maps are occupied. `asps_eom_patch_enabled_flag` equal to 0 indicates that the decoded occupancy map video does not contain information about whether intermediate depth positions between two depth maps are occupied.

[0685] `asps_raw_patch_enabled_flag` equal to 1 indicates that the decoded geometry and attribute video of the current atlas contains information related to RAW encoded points. `asps_raw_patch_enabled_flag` equal to 0 indicates that the decoded geometry and attribute video does not contain information related to RAW encoded points.

[0686] When asps_eom_patch_enabled_flag or asps_raw_patch_enabled_flag is equal to 1, asps_auxiliary_video_enabled_flag can be included in the atlas sequence parameter set syntax.

[0687] `asps_auxiliary_video_enabled_flag` equal to 1 indicates that information associated with RAW and EOM patch types can be placed in the auxiliary video sub-bitstream. `asps_auxiliary_video_enabled_flag` equal to 0 indicates that information associated with RAW and EOM patch types can be placed only in the main video sub-bitstream.

[0688] A value of 1 for `asps_point_local_reconstruction_enabled_flag` indicates that point local reconstruction mode information is available in the bitstream of the current atlas. A value of 0 for `asps_point_local_reconstruction_enabled_flag` indicates that no information related to point local reconstruction mode is available in the bitstream of the current atlas.

[0689] When asps_point_local_reconstruction_enabled_flag equals 1, asps_point_local_reconstruction_information can be carried in the atlas sequence parameter set.

[0690] The increment of asps_map_count_minus1 indicates the number of maps available for encoding the geometry and attribute data of the current atlas.

[0691] `asps_pixel_deinterleaving_enabled_flag` equal to 1 indicates that the decoded geometry and attribute video of the current atlas contains spatially interleaved pixels. `asps_pixel_deinterleaving_flag` equal to 0 indicates that the decoded geometry and attribute video corresponding to the current atlas contains only pixels from a single atlas.

[0692] `asps_pixel_deinterleaving_map_flag[i]` equal to 1 indicates that the decoded geometry and attribute video corresponding to the map at index i in the current set contains pixels that are spatially interleaved with the two maps. `asps_pixel_deinterleaving_map_flag[i]` equal to 0 indicates that the decoded geometry and attribute video corresponding to the map at index i in the current set contains pixels that correspond to a single map.

[0693] When asps_pixel_deinterleaving_map_flag[i] equals 1, asps_pixel_deinterleaving_map_flag[i] can be carried in the atlas sequence parameter set according to the value of asps_map_count_minus1.

[0694] When asps_eom_patch_enabled_flag and asps_map_count_minus1 are equal to 1, asps_eom_fix_bit_count_minus1 can be carried in the atlas sequence parameter set.

[0695] The increment of asps_eom_fix_bit_count_minus1 indicates the size (in bits) of the EOM codeword.

[0696] When asps_pixel_deinterleaving_enabled_flag or asps_point_local_reconstruction_enabled_flag is equal to 1, asps_surface_thickness_minus1 can be carried in the atlas sequence parameter set.

[0697] When asps_pixel_deinterleaving_enabled_flag or asps_point_local_reconstruction_enabled_flag equals 1, asps_surface_thickness_minus1 plus 1 specifies the maximum absolute difference between the explicit encoded depth value and the interpolated depth value.

[0698] `asps_vui_parameters_present_flag` equal to 1 indicates that the `vui_parameters()` syntax structure exists. `asps_vui_parameters_present_flag` equal to 0 indicates that the `vui_parameters()` syntax structure does not exist.

[0699] An asps_extension_flag value of 0 indicates that the asps_extension_data_flag syntax element does not exist in the ASPS RBSP syntax structure.

[0700] The `asps_extension_data_flag` flag indicates that the ASPS RBSP syntax structure includes data for extension.

[0701] The rbsp_trailing_bits function is used to pad the remaining bits with 0s for byte alignment after adding a 1 (stop bit) to indicate the end of the RBSP data.

[0702] ​ The atlas frame parameter set according to the implementation method is shown.

[0703] ​ As shown ​ The syntax shown is for the atlas frame parameter set contained in the NAL unit when the NAL unit type is NAL_AFPS.

[0704] The Atlas Frame Parameter Set (AFPS) contains the syntactic structure including syntactic elements applied to all zero or more complete encoded atlas frames.

[0705] The `afps_atlas_frame_parameter_set_id` identifier identifies the atlas frame parameter set for reference by other syntax elements. Identifiers that other syntax elements can reference are provided through the AFPS atlas frame parameter set.

[0706] afps_atlas_sequence_parameter_set_id specifies the value of asps_atlas_sequence_parameter_set_id for the activity atlas sequence parameter set.

[0707] atlas_frame_tile_information() will refer to ​ describe.

[0708] A value of 1 for `afps_output_flag_present_flag` indicates that the `atgh_frame_output_flag` or `ath_frame_output_flag` syntax element exists in the associated concatenation header. A value of 0 for `afps_output_flag_present_flag` indicates that the `atgh_frame_output_flag` or `ath_frame_output_flag` syntax element does not exist in the associated concatenation header.

[0709] Increasing afps_num_ref_idx_default_active_minus1 by 1 specifies the inferred value of the variable NumRefIdxActive for the tile group where atgh_num_ref_idx_active_override_flag is equal to 0.

[0710] afps_additional_lt_afoc_lsb_len specifies the value of the variable MaxLtAtlasFrmOrderCntLsb used for decoding processing of reference atlas frames.

[0711] The increment of 1 in afps_3d_pos_x_bit_count_minus1 specifies the number of bits in the fixed-length representation of pdu_3d_pos_x[j] of the patch at index j in the atlas tile group of reference afps_atlas_frame_parameter_set_id.

[0712] The increment of 1 in afps_3d_pos_y_bit_count_minus1 specifies the number of bits in the fixed-length representation of pdu_3d_pos_y[j] of the patch at index j in the atlas tile group of reference afps_atlas_frame_parameter_set_id.

[0713] A value of 1 for `afps_lod_mode_enabled_flag` indicates that LOD parameters can exist in the patch. A value of 0 for `afps_lod_mode_enabled_flag` indicates that LOD parameters do not exist in the patch.

[0714] `afps_override_eom_for_depth_flag` equal to 1 indicates that the values ​​of `afps_eom_number_of_patch_bit_count_minus1` and `afps_eom_max_bit_count_minus1` explicitly exist in the bitstream. `afps_override_eom_for_depth_flag` equal to 0 indicates that the values ​​of `afps_eom_number_of_patch_bit_count_minus1` and `afps_eom_max_bit_count_minus1` are implicitly inferred.

[0715] afps_eom_number_of_patch_bit_count_minus1 plus 1 specifies the number of bits used to represent the number of geometric patches associated with the EOM attribute patches in the atlas frame associated with this atlas frame parameter set.

[0716] afps_eom_max_bit_count_minus1 plus 1 specifies the number of bits used to represent the number of EOM points per geometric patch associated with the EOM attribute patch in the atlas frame associated with this atlas frame parameter set.

[0717] A value of 1 for `afs_raw_3d_pos_bit_count_explicit_mode_flag` indicates that the number of bits in the fixed-length representation of `rpdu_3d_pos_x`, `rpdu_3d_pos_y`, and `rpdu_3d_pos_z` is explicitly encoded by `atgh_raw_3d_pos_axis_bit_count_minus1` in the atlas tile group header referencing `afps_atlas_frame_parameter_set_id`. A value of 0 for `afs_raw_3d_pos_bit_count_explicit_mode_flag` indicates that the value of `atgh_raw_3d_pos_axis_bit_count_minus1` is implicitly deduced.

[0718] When afps_raw_3d_pos_bit_count_explicit_mode_flag equals 1, atgh_raw_3d_pos_axis_bit_count_minus1 can be included in the atlas tile group (or tile) header.

[0719] A value of 0 for afps_extension_flag indicates that the afps_extension_data_flag syntax element does not exist in the AFPS RBSP syntax structure.

[0720] afps_extension_data_flag can contain extension-related data.

[0721] ​ The atlas_frame_tile_information is shown according to the implementation method.

[0722] ​ Showing includes ​ The syntax of atlas_frame_tile_information.

[0723] `afti_single_tile_in_atlas_frame_flag` equal to 1 indicates that only one tile exists in each atlas frame of the reference AFPS. `afti_single_tile_in_atlas_frame_flag` equal to 0 indicates that more than one tile exists in each atlas frame of the reference AFPS.

[0724] `afti_uniform_tile_spacing_flag` equal to 1 specifies that tile column and row boundaries are evenly distributed across the atlas frame and are signaled using the syntax elements `afti_tile_cols_width_minus1` and `afti_tile_rows_height_minus1`, respectively. `afti_uniform_tile_spacing_flag` equal to 0 specifies that tile column and row boundaries may or may not be evenly distributed across the atlas frame and are signaled using the syntax elements `afti_num_tile_columns_minus1` and `afti_num_tile_rows_minus1`, as well as a list of syntax element pairs `afti_tile_column_width_minus1` and `afti_tile_row_height_minus1`.

[0725] When afti_uniform_tile_spacing_flag equals 1, afti_tile_cols_width_minus1 plus 1 specifies the width of tile columns other than the rightmost tile column of the atlas frame in units of 64 samples.

[0726] When afti_uniform_tile_spacing_flag equals 1, afti_tile_rows_height_minus1 plus 1 specifies the height of the tile rows, excluding the bottom tile rows of the atlas frame, in units of 64 samples.

[0727] When afti_uniform_tile_spacing_flag equals 0, afti_num_tile_columns_minus1 plus 1 specifies the number of tile columns used to divide the atlas frame.

[0728] When pti_uniform_tile_spacing_flag equals 0, afti_num_tile_rows_minus1 plus 1 specifies the number of tile rows to divide the atlas frame.

[0729] afti_tile_column_width_minus1[i] increments by 1 to specify the width of the i-th tile column in units of 64 samples.

[0730] afti_tile_row_height_minus1[i] increments by 1 to specify the height of the i-th tile row in units of 64 samples.

[0731] `afti_single_tile_per_tile_group_flag` equal to 1 indicates that each tile group referencing this AFPS consists of one tile. `afti_single_tile_per_tile_group_flag` equal to 0 indicates that a tile group referencing this AFPS may contain more than one tile.

[0732] When afti_single_tile_per_tile_group_flag equals 0, afti_num_tile_groups_in_atlas_frame_minus1 is carried in the atlas frame tile information. Based on afti_num_tile_groups_in_atlas_frame_minus1, afti_top_left_tile_idx[i] and afti_bottom_right_tile_idx_delta[i] can be carried in the atlas frame tile information.

[0733] afti_num_tile_groups_in_atlas_frame_minus1 plus 1 specifies the number of tile groups in each atlas frame of the reference AFPS.

[0734] afti_top_left_tile_idx[i] specifies the tile index of the top-left tile in the i-th tile group.

[0735] afti_bottom_right_tile_idx_delta[i] specifies the difference between the tile index of the tile located at the bottom right corner of the i-th tile group and afti_top_left_tile_idx[i].

[0736] A value of 1 for afti_signalled_tile_group_id_flag indicates that the tile group ID is signaled to each tile group.

[0737] When afti_signalled_tile_group_id_flag is 1, afti_signalled_tile_group_id_length_minus1 and afti_tile_group_id[i] can be carried in the atlas frame tile information.

[0738] afti_signalled_tile_group_id_length_minus1 plus 1 specifies the number of bits used to represent the syntax element afti_tile_group_id[i] (if it exists) and the syntax element atgh_address in the tile group header.

[0739] afti_tile_group_id[i] specifies the tile group ID of the i-th tile group. The length of the afti_tile_group_id[i] syntax element is afti_signalled_tile_group_id_length_minus1+1 bits.

[0740] ​ The atlas adaptation parameter set (atlas_adaptation_parameter_set_rbsp()) according to the implementation method is shown.

[0741] ​ This shows the syntax of the atlas adaptation parameter set carried by the NAL unit when the NAL unit type is NAL_AAPS.

[0742] An AAPS RBSP includes parameters that can be referenced by the NAL units of one or more coded atlas frames (or tiles). During the decoding process, at most one AAPS RBSP is considered active at any given time, and the activation of any particular AAPS RBSP causes the previously active AAPS RBSP to be deactivated.

[0743] aaps_atlas_adaptation_parameter_set_id identifies the atlas adaptation parameter set for reference by other syntax elements.

[0744] aaps_atlas_sequence_parameter_set_id specifies the value of asps_atlas_sequence_parameter_set_id for the activity atlas sequence parameter set.

[0745] A value of 1 for `aps_camera_parameters_present_flag` indicates that camera parameters exist in the current atlas adaptation parameter set. A value of 0 for `aps_camera_parameters_present_flag` indicates that camera parameters do not exist in the current adaptation parameter set.

[0746] A value of 0 for aaps_extension_flag indicates that the aaps_extension_data_flag syntax element does not exist in the AAPS RBSP syntax structure.

[0747] aaps_extension_data_flag can contain extension-related data.

[0748] ​The atlas_camera_parameters are shown according to the implementation method.

[0749] ​ Show ​ The detailed syntax of atlas_camera_parameters.

[0750] acp_camera_model indicates the camera model of the point cloud frame associated with the current set of adaptation parameters.

[0751] For example, an acp_camera_model of 0 indicates that the camera model is not specified.

[0752] A value of 1 for acp_camera_model indicates that the camera model is an orthographic camera model.

[0753] When acp_camera_model is 2-255, a camera model can be reserved.

[0754] When acp_camera_model equals 1, the following elements related to scaling, offset, and rotation can be included in the atlas camera parameters.

[0755] An acp_scale_enabled_flag value of 1 indicates that scaling parameters exist for the current camera model. An acp_scale_enabled_flag value of 0 indicates that scaling parameters do not exist for the current camera model.

[0756] When acp_scale_enabled_flag equals 1, the values ​​of d in the atlas camera parameters may include acp_scale_on_axis[d].

[0757] An acp_offset_enabled_flag value of 1 indicates that an offset parameter exists for the current camera model. An acp_offset_enabled_flag value of 0 indicates that an offset parameter does not exist for the current camera model.

[0758] When acp_offset_enabled_flag equals 1, the values ​​of d can include the acp_offset_on_axis[d] element in the atlas camera parameters.

[0759] A value of 1 for `acp_rotation_enabled_flag` indicates that rotation parameters exist for the current camera model. A value of 0 for `acp_rotation_enabled_flag` indicates that rotation parameters do not exist for the current camera model.

[0760] The `acp_scale_on_axis[d]` directive specifies the scaling value of the current camera model along the d-axis: `Scale[d]`. The value of `d` can be in the range of 0 to 2 (inclusive), where the values ​​0, 1, and 2 correspond to the X, Y, and Z axes, respectively.

[0761] `acp_offset_on_axis[d]` indicates the offset of the current camera model along the d-axis by the value of `Offset[d]`, where `d` can be in the range of 0 to 2 (inclusive). Values ​​of `d` equal to 0, 1, and 2 correspond to the X, Y, and Z axes, respectively.

[0762] When acp_rotation_enabled_flag equals 1, the following rotation values ​​can be included in the atlas camera parameters.

[0763] acp_rotation_qx uses a quaternion to represent the x-component qX of the rotation of the current camera model.

[0764] acp_rotation_qy uses a quaternion to represent the y-component qY of the rotation of the current camera model.

[0765] acp_rotation_qz uses a quaternion to represent the z-component qZ of the rotation of the current camera model.

[0766] ​ The atlas_tile_group_layer and atlas_tile_group_header according to the implementation are shown.

[0767] ​ Showing according to such ​ The NAL cell type shown contains the syntax of atlas_tile_group_layer carried in the NAL cell, as well as the syntax atlas_tile_group_header contained in atlas_tile_group_layer.

[0768] According to the implementation, a tile group may correspond to a tile. In this disclosure, the term "tile group" may be referred to as the term "tile". Similarly, the term "atgh" may be interpreted as the term "ath".

[0769] The atlas_tile_group_layer or atlas_tile_layer can contain either the atlas_tile_group_header or the atlas_tile_header.

[0770] When atgh_type is not SKIP_TILE_GRP, atlas tile group (or tile) data can be contained in atlas_tile_group_layer or atlas_tile_layer.

[0771] atgh_atlas_frame_parameter_set_id specifies the value of afps_atlas_frame_parameter_set_id for the active atlas frame parameter set of the current atlas tile group.

[0772] atgh_atlas_adaptation_parameter_set_id specifies the value of aaps_atlas_adaptation_parameter_set_id for the active atlas adaptation parameter set of the current atlas tile group.

[0773] `atgh_address` specifies the tile group address of a tile group. If it does not exist, the value of `atgh_address` is inferred to be 0. The tile group address is the tile group ID of the tile group. The length of `atgh_address` is `afti_signalled_tile_group_id_length_minus1+1` bits. If `afti_signalled_tile_group_id_flag` is equal to 0, then the value of `atgh_address` is in the range of 0 to `afti_num_tile_groups_in_atlas_frame_minus1` (inclusive). Otherwise, the value of `atgh_address` is in the range of 0 to 2(afti_signalled_tile_group_id_length_minus1+1)-1 (inclusive).

[0774] atgh_type specifies the encoding type of the current atlas tile group (tile).

[0775] When the value of atgh_type is 0, the type of atlas tile group or atlas tile is P_TILE_GRP (inter-frame atlas tile group (or tile)).

[0776] When the value of atgh_type is 1, the type of atlas tile group or atlas tile is I_TILE_GRP (intra-frame atlas tile group (or tile)).

[0777] When the value of atgh_type is 2, the type of atlas tile group or atlas tile is SKIP_TILE_GRP (skip atlas tile group (or tile)).

[0778] When the value of atgh_type is 3, the type of atlas tile group or atlas tile can have a reserved value.

[0779] atgh_atlas_output_flag affects the output and removal of the decoded atlas.

[0780] atgh_atlas_frm_order_cnt_lsb specifies the atlas frame order count modulo MaxAtlasFrmOrderCntLsb for the current atlas tile group.

[0781] When afps_output_flag_present_flag equals 1, the atlas tile group (tile) header can contain atgh_atlas_output_flag and atgh_atlas_frm_order_cnt_lsb.

[0782] `atgh_ref_atlas_frame_list_sps_flag` equal to 1 specifies the reference atlas frame list for the current atlas tile group, which is deduced based on one of the `ref_list_struct(rlsIdx)` syntax structures in the active ASPS. `atgh_ref_atlas_frame_list_sps_flag` equal to 0 specifies the reference atlas frame list for the current atlas tile list, which is deduced based on the `ref_list_struct(rlsIdx)` syntax structure directly included in the tile group header of the current atlas tile group.

[0783] When atgh_ref_atlas_frame_list_sps_flag equals 0, the atlas tile group (tile) header may include ref_list_struct(asps_num_ref_atlas_frame_lists_in_asps).

[0784] atgh_ref_atlas_frame_list_idx specifies the index of the ref_list_struct(rlsIdx) syntax structure used to derive the reference atlas frame list for the current atlas tile group within the list of ref_list_struct(rlsIdx) syntax structures included in the active ASPS.

[0785] The value of atgh_additional_afoc_lsb_present_flag[j] being equal to 1 indicates that atgh_additional_afoc_lsb_val[j] exists for the current atlas tile group or atlas tile. The value of atgh_additional_afoc_lsb_present_flag[j] being equal to 0 indicates that atgh_additional_afoc_lsb_val[j] does not exist.

[0786] When atgh_additional_afoc_lsb_present_flag[j] equals 1, atgh_additional_afoc_lsb_val[j] can be included in the atlas tile group (tile) header.

[0787] atgh_additional_afoc_lsb_val[j] specifies the value of FullAtlasFrmOrderCntLsbLt[RlsIdx][j] for the current atlas tile group or tile.

[0788] `atgh_pos_min_z_quantizer` specifies the quantizer to be applied to the value of `pdu_3d_pos_min_z[p]` of patch `p`. If `atgh_pos_min_z_quantizer` does not exist, its value can be inferred to be equal to 0.

[0789] The atgh_pos_delta_max_z_quantizer specifies the quantizer to be applied to the pdu_3d_pos_delta_max_z[p] value of the patch at index p. If atgh_pos_delta_max_z_quantizer does not exist, its value can be inferred to be equal to 0.

[0790] `atgh_patch_size_x_info_quantizer` specifies the value of the quantizer `PatchSizeXQuantizer` for the patches `pdu_2d_size_x_minus1[p]`, `mpdu_2d_delta_size_x[p]`, `ipdu_2d_delta_size_x[p]`, `rpdu_2d_size_x_minus1[p]`, and `epdu_2d_size_x_minus1[p]` to be applied to the patch of variable index `p`. If `atgh_patch_size_x_info_quantizer` does not exist, its value can be inferred to be equal to `asps_log2_patch_packing_block_size`.

[0791] `atgh_patch_size_y_info_quantizer` specifies the value of the quantizer `PatchSizeYQuantizer` for the variables `pdu_2d_size_y_minus1[p]`, `mpdu_2d_delta_size_y[p]`, `ipdu_2d_delta_size_y[p]`, `rpdu_2d_size_y_minus1[p]`, and `epdu_2d_size_y_minus1[p]` to be applied to the patch at index `p`. If `atgh_patch_size_y_info_quantizer` does not exist, its value can be inferred to be equal to `asps_log2_patch_packing_block_size`.

[0792] The increment of 1 in `atgh_raw_3d_pos_axis_bit_count_minus1` specifies the number of bits in the fixed-length representation of `rpdu_3d_pos_x`, `rpdu_3d_pos_y`, and `rpdu_3d_pos_z`.

[0793] When atgh_type is P_TILE_GRP and num_ref_entries[RlsIdx] is greater than 1, atgh_num_ref_idx_active_override_flag can be included in the atlas tile group (or tile) header. Additionally, when atgh_num_ref_idx_active_override_flag is equal to 1, atgh_num_ref_idx_active_minus1 can be included in the atlas tile group (tile) header.

[0794] `atgh_num_ref_idx_active_override_flag` equal to 1 indicates that the syntax element `atgh_num_ref_idx_active_minus1` exists for the current atlas tile group. `atgh_num_ref_idx_active_override_flag` equal to 0 indicates that the syntax element `atgh_num_ref_idx_active_minus1` does not exist. If `atgh_num_ref_idx_active_override_flag` does not exist, its value can be inferred to be 0.

[0795] `atgh_num_ref_idx_active_minus1` specifies the maximum reference index of the list of reference atlas frames that can be used to decode the current atlas tile group. When the value of `NumRefIdxActive` is equal to 0, no reference index of the list of reference atlas frames is available for decoding the current atlas tile group.

[0796] byte_alignment is used to align bytes by padding the remaining bits with 0s after adding a 1 (stop bit) to indicate the end of the data.

[0797] ​ The reference list structure (ref_list_struct) according to the implementation method is shown.

[0798] ​ Show ​ Atlas parameter set, ​ The syntax of a reference list structure that may be included in the header of a set of atlases or atlases.

[0799] num_ref_entries[rlsIdx] specifies the number of entries in the ref_list_struct(rlsIdx) syntax structure.

[0800] The following elements, which are the same number as the value of num_ref_entries[rlsIdx], can be included in the reference list structure.

[0801] When asps_long_term_ref_atlas_frames_flag equals 1, the reference atlas frame flag (st_ref_atlas_frame_flag[rlsIdx][i]) can be included in the reference list structure.

[0802] A value of 1 for `st_ref_atlas_frame_flag[rlsIdx][i]` indicates that the i-th entry in the `ref_list_struct(rlsIdx)` syntax structure is a short-term reference atlas frame entry. A value of 0 for `st_ref_atlas_frame_flag[rlsIdx][i]` indicates that the i-th entry in the `ref_list_struct(rlsIdx)` syntax structure is a long-term reference atlas frame entry. When `st_ref_atlas_frame_flag[rlsIdx][i]` does not exist, its value can be inferred to be 1.

[0803] When st_ref_atlas_frame_flag[rlsIdx][i] equals 1, abs_delta_afoc_st[rlsIdx][i] can be included in the reference list structure.

[0804] When the i-th entry is the first short-term reference atlas frame entry in the ref_list_struct(rlsIdx) syntax structure, abs_delta_afoc_st[rlsIdx][i] specifies the absolute difference between the atlas frame order count of the current atlas tile group and the atlas frame referenced by the i-th entry. Alternatively, when the i-th entry is a short-term reference atlas frame entry but not the first short-term reference atlas frame entry in the ref_list_struct(rlsIdx) syntax structure, it specifies the absolute difference between the atlas frame order count of the i-th entry and the atlas frame referenced by the previous short-term reference atlas frame entry in the ref_list_struct(rlsIdx) syntax structure.

[0805] When abs_delta_afoc_st[rlsIdx][i] has a value greater than 0, the entry sign flag (strpf_entry_sign_flag[rlsIdx][i]) may be included in the reference list structure.

[0806] The value of `strpf_entry_sign_flag[rlsIdx][i]` being equal to 1 indicates that the i-th entry in the syntax structure `ref_list_struct(rlsIdx)` has a value greater than or equal to 0. The value of `strpf_entry_sign_flag[rlsIdx][i]` being equal to 0 indicates that the i-th entry in the syntax structure `ref_list_struct(rlsIdx)` has a value less than 0. When `strpf_entry_sign_flag[rlsIdx][i]` does not exist, the value of `strpf_entry_sign_flag[rlsIdx][i]` can be inferred to be equal to 1.

[0807] `afoc_lsb_lt[rlsIdx][i]` specifies the atlas frame order count modulo `MaxAtlasFrmOrderCntLsb` of the atlas frame referenced by the i-th entry in the `ref_list_struct(rlsIdx)` syntax structure. The length of the `afoc_lsb_lt[rlsIdx][i]` syntax element is `asps_log2_max_atlas_frame_order_cnt_lsb_minus4+4` bits.

[0808] ​ This shows the atlas tile group data (atlas_tile_group_data_unit) according to the implementation method.

[0809] ​ Show ​The syntax of atlas tile group data included in the atlas tile group layer (or atlas tile layer). Atlas tile group data can correspond to atlas tile data, and a tile group can be called a tile.

[0810] As p increases from 0 to 1, the atlas-related elements based on index p can be included in the atlas tile group (or tile) data.

[0811] `atgdu_patch_mode[p]` indicates the patch mode of the patch at index `p` in the current atlas patch group. A patch group with `atgh_type = SKIP_TILE_GRP` means that the entire patch group information is copied directly from the patch group that has the same `atgh_address` as the current patch group corresponding to the first reference atlas frame.

[0812] When atgdu_patch_mode[p] is not I_END and atgdu_patch_mode[p] is not P_END, each index p may include patch_information_data and atgdu_patch_mode[p] in the atlas patch group data (or atlas patch data).

[0813] The patch mode type of the I_TILE_GRP type atlas tile group can be represented as follows.

[0814] atgdu_patch_mode equal to 0 indicates a non-predictive patch mode with the I_INTRA identifier.

[0815] atgdu_patch_mode equal to 1 indicates the RAW point patch mode with the I_RAW identifier.

[0816] atgdu_patch_mode equal to 2 indicates the EOM point patch mode with the I_EOM identifier.

[0817] The value of atgdu_patch_mode is equal to 3 to 13, indicating a reserved mode with the I_RESERVED identifier.

[0818] atgdu_patch_mode equal to 14 indicates the patch termination mode with the I_END identifier.

[0819] The patch mode type of the P_TILE_GRP type atlas tile group (or tile) can be represented as follows.

[0820] atgdu_patch_mode equal to 0 indicates the patch skip mode with the P_SKIP identifier.

[0821] atgdu_patch_mode equal to 1 indicates the patch merging mode with the P_MERGE identifier.

[0822] atgdu_patch_mode equal to 2 indicates the inter-frame prediction patch mode with the P_INTER identifier.

[0823] atgdu_patch_mode equal to 3 indicates a non-predictive patch mode with the identifier P_INTRA.

[0824] atgdu_patch_mode equal to 4 indicates the RAW point patch mode with the P_RAW identifier.

[0825] atgdu_patch_mode equal to 5 indicates the EOM point patch mode with the P_EOM identifier.

[0826] The value of atgdu_patch_mode is between 6 and 13, indicating a reserved mode with the P_RESERVED identifier.

[0827] atgdu_patch_mode equal to 14 indicates the patch termination mode with the P_END identifier.

[0828] The patch mode type of the SKIP_TILE_GRP type atlas tile group (or tile) can be represented as follows.

[0829] atgdu_patch_mode equal to 0 indicates the patch skip mode with the P_SKIP identifier.

[0830] AtgduTotalNumberOfPatches is the number of patches and can be set as the final value of p.

[0831] ​ The patch information data (patch_information_data) according to the implementation method is shown.

[0832] ​ Show ​ The syntax of patch_information_data included in the atlas patch group (or patch) data unit.

[0833] If atgh_type is Skip Atlas Patch Group (or Skip Atlas Patch) (SKIP_TILE_GR), then skip_patch_data_unit (patchIdx) is included as patch information data.

[0834] If `atgh_type` is inter-atlas patch group (or inter-atlas patch) (`P_TILE_GR`) and `patchMode` is patch skip mode (`P_SKIP`), then `skip_patch_data_unit` (`patchIdx`) is included as patch information data. If `patchMode` is patch merge mode (`P_MERGE`), then `merge_patch_data_unit` (`patchIdx`) is included as patch information data. If `patchMode` is non-predictive patch mode (`P_INTRA`), then `patch_data_unit` (`patchIdx`) (hereinafter referred to as patch information data) is included in this syntax structure. If `patchMode` is inter-frame predictive patch mode (`P_INTER`), then `inter_patch_data_unit` (`patchIdx`) is included in this syntax structure. If `patchMode` is RAW point patch mode (`P_RAW`), then `raw_patch_data_unit` (`patchIdx`) is included in this syntax structure. If patchMode is EOM dot patch mode (P_EOM), then eom_patch_data_unit(patchIdx) is included in this syntax structure.

[0835] If `atgh_type` is an in-atlas patch group (`I_TILE_GR`) and `patchMode` is a non-predictive patch mode (`I_INTRA`), then `patch_data_unit(patchIdx)` is included in this syntax structure. If `patchMode` is a RAW point patch mode (`I_RAW`), then `raw_patch_data_unit(patchIdx)` is included in this syntax structure. If `patchMode` is an EOM point patch mode (`I_EOM`), then `eom_patch_data_unit(patchIdx)` is included in this syntax structure.

[0836] ​ The patch_data_unit according to the implementation is shown.

[0837] ​ Show ​ The syntax of patch_data_unit included in it.

[0838] `pdu_2d_pos_x[p]` specifies the x-coordinate (or left offset) of the top-left corner of the patch bounding box of patch `p` in the current atlas tile group (or tile). Atlas tile groups can have tile group indices, and atlas tiles can have tile indices. These indices can be represented as multiples of `PatchPackingBlockSize`.

[0839] `pdu_2d_pos_y[p]` specifies the y-coordinate (or top offset) of the top-left corner of the patch bounding box of patch `p` in the current atlas tile group. Atlas tile groups can have tile group indices, and atlas tiles can have tile indices. These indices can be represented as multiples of `PatchPackingBlockSize`.

[0840] pdu_2d_size_x_minus1[p] + 1 specifies the quantization width value of the patch at index p in the current atlas tileGroupIdx or the current atlas tile tiledx.

[0841] pdu_2d_size_y_minus1[p] + 1 specifies the quantization height value of the patch at index p in the current atlas tileGroupIdx or the current atlas tile tiledx.

[0842] pdu_3d_pos_x[p] specifies the shift of the reconstructed patch point in the patch of the current atlas tile group along the tangential axis to be applied to the index p.

[0843] pdu_3d_pos_y[p] specifies the shift of the reconstructed patch point in the patch of index p to be applied to the current atlas patch group along the double tangent axis.

[0844] pdu_3d_pos_min_z[p] specifies the shift of the reconstructed patch point in the patch of index p to be applied to the current atlas patch group along the normal axis.

[0845] When asps_normal_axis_max_delta_value_enabled_flag equals 1, pdu_3d_pos_delta_max_z[patchIdx] can be included in the patch data unit.

[0846] If present, pdu_3d_pos_delta_max_z[p] specifies the nominal maximum value of the shift expected to exist in the patch of index p of the current atlas tile group along the normal axis after reconstructing the bit depth patch geometry sample after conversion to its nominal representation.

[0847] pdu_projection_id[p] specifies the projection mode and index value of the projection plane normal of the patch of the current atlas tile group at index p.

[0848] pdu_orientation_index[p] indicates the patch orientation index of the patch of the current atlas tile group.

[0849] When afps_lod_mode_enabled_flag equals 1, pdu_lod_enabled_flag[patchIndex] can be included in the patch data unit. When pdu_lod_enabled_flag[patchIndex] is greater than 0, pdu_lod_scale_x_minus1[patchIndex] and pdu_lod_scale_y[patchIndex] can be included in the patch data unit.

[0850] If pdu_lod_enabled_flag[p] equals 1, it indicates that an LOD parameter exists for the current patch p. If pdu_lod_enabled_flag[p] equals 0, then no LOD parameter exists for the current patch.

[0851] pdu_lod_scale_x_minus1[p] specifies the LOD scaling factor for the local x-coordinates of points in the patch of the current atlas tile group at index p, before adding them to the patch coordinates Patch3dPosX[p].

[0852] pdu_lod_scale_y[p] specifies the LOD scaling factor for the local y-coordinates of points in the patch of the current atlas tile group at index p, before adding them to the patch coordinates Patch3dPosY[p].

[0853] When asps_point_local_reconstruction_enabled_flag equals 1, point_local_reconstruction_data(patchIdx) can be included in the patch data unit.

[0854] The point_local_reconstruction_data(patchIdx) function can contain information that allows the decoder to recover points lost due to compression loss, etc.

[0855] ​ The rotation and offset relative to the patch orientation are shown according to the embodiment.

[0856] ​ Show​ The rotation matrix and offset of the orientation index.

[0857] The method / apparatus according to the embodiments can perform orientation operations on point cloud data, and uses identifiers, rotations, and offsets for such operations, such as... ​ As shown.

[0858] ​ The scene object information (scene_object_information) is shown according to the implementation method.

[0859] ​ Show ​ The syntax of the SEI message of the NAL unit of the sample stream contained in the bit stream 32000 shown.

[0860] According to the implementation method, the object is a point cloud object. Furthermore, an object is a concept that even includes (partial) objects that constitute or identify an object.

[0861] The method / apparatus according to the embodiments may provide partial access based on scene object information according to the embodiments, based on a 3D spatial region including one or more scene objects.

[0862] According to the implementation, the SEI message includes information about processing related to decoding, reconstruction, display, or other purposes.

[0863] According to the implementation method, SEI messages include two types: necessary and unnecessary.

[0864] Decoding does not require unnecessary SEI messages. The decoder does not need to process this information to conform to the output order.

[0865] According to the implementation method, the volume annotation information (volume annotation SEI message) including scene object information, object label information, patch information and volume rectangle information can be a non-essential SEI message.

[0866] According to the implementation method, the above information can be carried in the necessary SEI message.

[0867] Necessary SEI messages are an integral part of the V-PCC bitstream and should not be removed from it. Necessary SEI messages can be categorized into two types:

[0868] Type A Required SEI Messages: These SEIs may contain information needed to check bitstream consistency and output timing decoder consistency. The V-PCC decoder according to the implementation does not discard any relevant Type A Required SEI messages. The V-PCC decoder according to the implementation may consider this information for bitstream consistency and output timing decoder consistency.

[0869] Type B Required SEI Messages: V-PCC decoders that conform to a specific reconstruction profile may not discard any relevant Type B Required SEI messages and may consider them for 3D point cloud reconstruction and consistency purposes.

[0870] ​ The volume annotation SEI message is displayed.

[0871] According to the implementation method, the V-PCC bitstream definition is as follows: ​ The volume annotation SEI message shown is related to partial access.

[0872] soi_cancel_flag equal to 1 indicates that the Scene Object Information SEI message cancels the persistence of any previous Scene Object Information SEI messages in the output order.

[0873] soi_num_object_updates indicates the number of objects to be updated by the current SEI.

[0874] When soi_num_object_updates is greater than 0, the following elements can be included in the scene object information.

[0875] A soi_simple_objects_flag value of 1 indicates that no signal will be used to notify the user of additional information about updated or newly introduced objects. A soi_simple_objects_flag value of 0 indicates that a signal will be used to notify the user of additional information about updated or newly introduced objects.

[0876] When soi_simple_objects_flag equals 0, the following elements can be included in the scene object information.

[0877] If soi_simple_objects_flag is not equal to 0, then the flags can be set as soi_object_label_present_flag=0, soi_priority_present_flag=0, soi_object_hidden_present_flag=0, soi_object_dependency_present_flag=0, soi_visibility_cones_present_flag=0, soi_3d_bounding_box_present_flag=0, soi_collision_shape_present_flag=0, soi_point_style_present_flag=0, soi_material_id_present_flag=0, and soi_extension_present_flag=0.

[0878] A soi_object_label_present_flag value of 1 indicates that object label information exists in the current scene's object information SEI message. A soi_object_label_present_flag value of 0 indicates that object label information does not exist.

[0879] `soi_priority_present_flag` equal to 1 indicates that priority information exists in the current scene object information SEI message. `soi_priority_present_flag` equal to 0 indicates that priority information does not exist.

[0880] A soi_object_hidden_present_flag value of 1 indicates that hidden object information exists in the current scene object information SEI message. A soi_object_hidden_present_flag value of 0 indicates that hidden object information does not exist.

[0881] `soi_object_dependency_present_flag` equal to 1 indicates that object dependency information exists in the current scene's object information SEI message. `soi_object_dependency_present_flag` equal to 0 indicates that object dependency information does not exist.

[0882] A soi_visibility_cones_present_flag value of 1 indicates that visible cone information exists in the current scene object information (SEI) message. A soi_visibility_cones_present_flag value of 0 indicates that visible cone information does not exist.

[0883] A `soi_3d_bounding_box_present_flag` value of 1 indicates that 3D bounding box information exists in the current scene object information (SEI) message. A `soi_3d_bounding_box_present_flag` value of 0 indicates that 3D bounding box information does not exist.

[0884] A soi_collision_shape_present_flag value of 1 indicates that a collision exists in the current scene object information (SEI) message. A soi_collision_shape_present_flag value of 0 indicates that a collision does not exist.

[0885] `soi_point_style_present_flag` equal to 1 indicates that point style information exists in the current scene object information (SEI) message. `soi_point_style_present_flag` equal to 0 indicates that point style information does not exist.

[0886] `soi_material_id_present_flag` equal to 1 indicates that material ID information exists in the current scene object information SEI message. `soi_material_id_present_flag` equal to 0 indicates that material ID information does not exist.

[0887] `soi_extension_present_flag` equal to 1 indicates that additional extension information should exist in the current scene object information (SEI) message. `soi_extension_present_flag` equal to 0 indicates that additional extension information does not exist. The bitstream consistency requirement of this version of the specification is that `soi_extension_present_flag` should be equal to 0.

[0888] When soi_3d_bounding_box_present_flag equals 1, the following elements can be included in the scene object information.

[0889] soi_3d_bounding_box_scale_log2 indicates the scaling to be applied to the 3D bounding box parameters that can be specified for the object.

[0890] soi_3d_bounding_box_precision_minus8 plus 8 indicates the precision of the 3D bounding box parameters that can be specified for an object.

[0891] soi_log2_max_object_idx_updated specifies the number of bits used in the current scene object information SEI message to signal the value of the object index.

[0892] When soi_object_dependency_present_flag equals 1, the following elements can be included in the scene object information.

[0893] soi_log2_max_object_dependency_idx specifies the number of bits in the current scene object information (SEI) message used to signal the value of the dependent object index.

[0894] The following elements, which are the same number as the soi_num_object_updates value, can be included in the scene object information.

[0895] `soi_object_idx[i]` indicates the object index of the i-th object to be updated. The number of bits used to represent `soi_object_idx[i]` is equal to `soi_log2_max_object_idx_updated`. When `soi_object_idx[i]` does not exist in the bitstream, its value can be inferred to be equal to 0.

[0896] `soi_object_cancel_flag` equal to 1 indicates that the object with index `i` can be cancelled, and the variable `ObjectTracked[i]` should be set to 0. Furthermore, all associated parameters (including object label, 3D bounding box parameters, priority information, hidden flag, dependency information, visible cone, conflict shape, point style, and material ID) can be reset to their default values. `soi_object_cancel_flag` equal to 0 also indicates that the object with index `soi_object_idx[i]` should be updated using the information following that element, and the variable `ObjectTracked[i]` can be set to 1.

[0897] When soi_object_cancel_flag[k] is not equal to 1 and soi_object_label_present_flag is equal to 1, the following elements may be included in the scene object information.

[0898] A value of 1 for soi_object_label_update_flag[i] indicates that object label update information exists for object index i. A value of 0 for soi_object_label_update_flag[i] indicates that object label update information does not exist.

[0899] When soi_object_label_update_flag[k] equals 1, the following elements can be included in the scene object information.

[0900] soi_object_label_idx[i] indicates the label index of the object with index i.

[0901] When soi_priority_present_flag equals 1, the following elements can be included in the scene object information.

[0902] `soi_priority_update_flag[i]` equal to 1 indicates that priority update information exists for the object at index `i`. `soi_priority_update_flag[i]` equal to 0 indicates that priority information does not exist for the object.

[0903] When soi_priority_update_flag[k] equals 1, the following elements can be included in the scene object information.

[0904] `soi_priority_value[i]` indicates the priority of the object at index `i`. The lower the priority value, the higher the priority.

[0905] When soi_object_hidden_present_flag is 1, the following elements can be included in the scene object information.

[0906] A value of 1 for soi_object_hidden_flag[i] indicates that the object at index i should be hidden. A value of 0 for soi_object_hidden_flag[i] indicates that the object at index i should be present.

[0907] When soi_object_dependency_present_flag equals 1, the following elements can be included in the scene object information.

[0908] A value of 1 for soi_object_dependency_update_flag[i] indicates that object dependency update information exists for object index i. A value of 0 for soi_object_dependency_update_flag[i] indicates that object dependency update information does not exist.

[0909] When soi_object_dependency_update_flag[k] equals 1, the following elements can be included in the scene object information.

[0910] soi_object_num_dependencies[i] indicates the number of dependencies of the object at index i.

[0911] Based on the value of soi_object_num_dependencies[k], the following elements can be included in the scene object information.

[0912] soi_object_dependency_idx[i][j] indicates the index of the j-th object that has a dependency on the object at index i.

[0913] When soi_visibility_cones_present_flag equals 1, the following elements can be included in the scene object information.

[0914] `soi_visibility_cones_update_flag[i]` equal to 1 indicates that visual cone update information exists for the object at index `i`. `soi_visibility_cones_update_flag[i]` equal to 0 indicates that visual cone update information does not exist.

[0915] When soi_visibility_cones_update_flag[k] equals 1, the following elements can be included in the scene object information.

[0916] `soi_direction_x[i]` indicates the normalized x-component value of the direction vector of the visible cone of the object at index `i`. When it does not exist, the value of `soi_direction_x[i]` can be assumed to be equal to 1.0.

[0917] `soi_direction_y[i]` indicates the normalized y-component value of the direction vector of the visible cone of the object at index `i`. When it does not exist, the value of `soi_direction_y[i]` can be assumed to be equal to 1.0.

[0918] `soi_direction_z[i]` indicates the normalized z-component value of the direction vector of the visible cone of the object at index `i`. When it does not exist, the value of `soi_direction_z[i]` can be assumed to be equal to 1.0.

[0919] soi_angle[i] indicates the angle (in degrees) of the visible cone along the direction vector. When it does not exist, the value of soi_angle[i] can be assumed to be equal to 180°.

[0920] When soi_3d_bounding_box_present_flag equals 1, the following elements can be included in the scene object information.

[0921] A value of 1 for soi_3d_bounding_box_update_flag[i] indicates that 3D bounding box information exists for the object at index i. A value of 0 for soi_3d_bounding_box_update_flag[i] indicates that 3D bounding box information does not exist.

[0922] `soi_3d_bounding_box_x[i]` indicates the x-coordinate of the origin of the 3D bounding box of the object at index `i`. The default value of `soi_3d_bounding_box_x[i]` can be 0.

[0923] `soi_3d_bounding_box_y[i]` indicates the y-coordinate of the origin of the 3D bounding box of the object at index `i`. The default value of `soi_3d_bounding_box_y[i]` can be 0.

[0924] `soi_3d_bounding_box_z[i]` indicates the z-coordinate of the origin of the 3D bounding box of the object at index `i`. The default value of `soi_3d_bounding_box_z[i]` can be 0.

[0925] `soi_3d_bounding_box_delta_x[i]` indicates the size of the bounding box of the object at index `i` on the x-axis. The default value of `soi_3d_bounding_box_delta_x[i]` can be equal to 0.

[0926] `soi_3d_bounding_box_delta_y[i]` indicates the size of the bounding box of the object at index `i` on the y-axis. The default value of `soi_3d_bounding_box_delta_y[i]` can be equal to 0.

[0927] `soi_3d_bounding_box_delta_z[i]` indicates the size of the bounding box of the object at index `i` on the z-axis. The default value of `soi_3d_bounding_box_delta_z[i]` can be equal to 0.

[0928] When soi_collision_shape_present_flag equals 1, the following elements can be included in the scene object information.

[0929] A value of 1 for soi_collision_shape_update_flag[i] indicates that there is a conflicting shape update for the object at index i. A value of 0 for soi_collision_shape_update_flag[i] indicates that there is no conflicting shape update.

[0930] When soi_collision_shape_update_flag[k]] equals 1, the following elements can be included in the scene object information.

[0931] soi_collision_shape_id[i] indicates the collision shape id of the object at index i.

[0932] When soi_point_style_present_flag equals 1, the following elements can be included in the scene object information.

[0933] A value of 1 for soi_point_style_update_flag[i] indicates that point style update information exists for the object at index i. A value of 0 for soi_point_style_update_flag[i] indicates that point style update information does not exist.

[0934] When soi_point_style_update_flag[k]] equals 1, the following elements can be included in the scene object information.

[0935] soi_point_shape_id[i] indicates the point shape id of the object at index i. The default value of soi_point_shape_id[i] can be equal to 0.

[0936] soi_point_size[i] indicates the point size of the object at index i. The default value of soi_point_size[i] can be equal to 1.

[0937] When soi_material_id_present_flag equals 1, the following elements can be included in the scene object information.

[0938] `soi_material_id_update_flag[i]` equal to 1 indicates that there is material ID update information for the object at index `i`. `soi_point_style_update_flag[i]` equal to 0 indicates that point style update information does not exist.

[0939] When soi_material_id_update_flag[k] equals 1, the following elements can be included in the scene object information.

[0940] soi_material_id[i] indicates the material ID of the object at index i. The default value of soi_material_id[i] can be equal to 0.

[0941] ​ The object label information is shown according to the implementation method.

[0942] Figure 47 Show Figure 32 The syntax of the object tag information of the SEI message of the NAL unit in the sample stream contained in the bit stream 32000 shown.

[0943] The value of oli_cancel_flag equal to 1 indicates that the object label information SEI message cancels the persistence of any previous object label information SEI messages in the output order.

[0944] When oli_cancel_flag is not equal to 1, the following elements can be included in the object label information.

[0945] A value of 1 for `oli_label_language_present_flag` indicates that label language information exists in the object label information SEI message. A value of 0 for `oli_label_language_present_flag` indicates that label language information does not exist.

[0946] When oli_label_language_present_flag equals 1, the following elements can be included in the object label information.

[0947] oli_bit_equal_to_zero can be equal to 0.

[0948] The `oli_label_language` string contains the language label specified in IETF RFC 5646, followed by a null terminator byte equal to 0x00. The length of the `oli_label_language` syntax element can be less than or equal to 255 bytes (excluding the null terminator byte).

[0949] `oli_num_label_updates` indicates the number of labels that the current SEI needs to update.

[0950] The following elements, which are as many as the value of oli_num_label_updates, can be included in the object label information.

[0951] oli_label_idx[i] indicates the label index of the i-th label to be updated.

[0952] An oli_label_cancel_flag value of 1 indicates that the label at index oli_label_idx[i] should be canceled and set to an empty string. An oli_label_cancel_flag value of 0 indicates that the label at index oli_label_idx[i] should be updated using the information following that element.

[0953] When oli_label_cancel_flag is not equal to 1, the following elements can be included in the object label information.

[0954] oli_bit_equal_to_zero equals 0.

[0955] oli_label[i] indicates the label of the i-th label. The length of the vti_label[i] syntax element should be less than or equal to 255 bytes (excluding null terminator bytes).

[0956] Figure 48 Information about the patch is shown according to an embodiment.

[0957] Figure 48 Show Figure 32 Syntax of the patch information of the SEI message of the NAL unit in the sample stream contained in the bit stream 32000 shown.

[0958] A value of 1 for pi_cancel_flag indicates that the SEI message cancels the persistence of any previous SEI messages in the output order and all entries in the SEI message table should be removed.

[0959] pi_num_tile_group_updates indicates the number of tile groups to be updated in the tile information table via the current SEI message.

[0960] When pi_num_tile_group_updates is greater than 0, the following elements can be included in the patch information.

[0961] pi_log2_max_object_idx_tracked specifies the number of bits in the current patch information SEI message used to signal the value of the tracked object index.

[0962] pi_log2_max_patch_idx_updated specifies the number of bits in the current patch information SEI message used to signal the update of the patch index.

[0963] The following elements, which are as many as the value of pi_num_tile_group_updates, can be included in the patch information.

[0964] pi_tile_group_address[i] specifies the tile group address of the i-th updated tile group in the current SEI message.

[0965] A value of 1 for `pi_tile_group_cancel_flag[i]` indicates that the tile group at index `i` should be reset and all patches previously assigned to that tile group will be removed. A value of 0 for `pi_tile_group_cancel_flag[i]` indicates that all patches previously assigned to the tile group at index `i` will be retained.

[0966] pi_num_patch_updates[i] indicates the number of patches to be updated by the current SEI within the patch information table at index i of the patch group.

[0967] The following elements, which are the same number as the value of pi_num_patch_updates, can be included in the patch information.

[0968] `pi_patch_idx[i][j]` indicates the patch index of the j-th patch in the patch information table, where index i is to be updated. The number of bits used to represent `pi_patch_idx[i]` is equal to `pi_log2_max_patch_idx_updated`. When `pi_patch_idx[i]` does not exist in the bitstream, its value can be inferred to be 0.

[0969] The value pi_patch_cancel_flag[i][j] equal to 1 indicates that the patch at index j in the patch information group of index i should be removed from the patch information table.

[0970] When pi_patch_cancel_flag[j][p] is not equal to 1, the following elements may be included in the patch information.

[0971] pi_patch_number_of_objects_minus1[i][j] indicates the number of objects to be associated with the patch at index j in the patch group of index i.

[0972] m can be set to m = pi_patch_number_of_objects_minus1[j][p] + 1, and the following elements, as many as the value of m, can be included in the patch information.

[0973] `pi_patch_object_idx[i][j][k]` indicates the index of the k-th object associated with the j-th patch in the patch group at index i. The number of bits used to represent `pi_patch_object_idx[i]` can be equal to `pi_log2_max_object_idx_tracked`. When `pi_patch_object_idx[i]` does not exist in the bitstream, its value can be inferred to be 0.

[0974] Figure 49 Information about the volume rectangle according to the embodiment is shown.

[0975] Figure 49 Show Figure 32 Syntax of the volume rectangle information of the SEI message of the NAL unit in the sample stream contained in the bit stream 32000 shown.

[0976] A value of 1 for `vri_cancel_flag` indicates that the Volume Rectangle Information SEI message cancels the persistence of any previous Volume Rectangle Information SEI messages in the output order, and all entries in the Volume Rectangle Information table should be removed.

[0977] vri_num_rectangles_updates indicates the number of volume rectangles to be updated via the current SEI.

[0978] When vri_num_rectangles_updates is greater than 0, the following elements can be included in the volume rectangle information.

[0979] vri_log2_max_object_idx_tracked specifies the number of bits in the current volume rectangle information SEI message used to signal the value of the tracked object index.

[0980] vri_log2_max_rectangle_idx_updated specifies the number of bits in the current volume rectangle information SEI message used to signal the update of the volume rectangle index.

[0981] The following elements, as many as the value of vri_num_rectangles_updates, can be included in the volume rectangle information.

[0982] `vri_rectangle_idx[i]` indicates the index of the i-th volume rectangle to be updated in the volume rectangle information table. The number of bits used to represent `vri_rectangle_idx[i]` can be equal to `vri_log2_max_rectangle_idx_updated`. When `vri_rectangle_idx[i]` does not exist in the bitstream, its value can be inferred to be equal to 0.

[0983] A value of 1 for vri_rectangle_cancel_flag[i] indicates that the volume rectangle at index i can be removed from the volume rectangle information table.

[0984] When vri_rectangle_cancel_flag[p] is not equal to 1, the following elements can be included in the volume rectangle information.

[0985] A value of 1 for `vri_bounding_box_update_flag[i]` indicates that the 2D bounding box information of the volume rectangle at index `i` should be updated. A value of 0 for `vti_bounding_box_update_flag[i]` indicates that the 2D bounding box information of the volume rectangle at index `i` should not be updated.

[0986] When vri_bounding_box_update_flag[p] equals 1, the following elements can be included in the volume rectangle information.

[0987] vri_bounding_box_top[i] indicates the vertical coordinate of the top-left position of the bounding box of the i-th volume rectangle within the current atlas frame. The default value of vri_bounding_box_top[i] can be equal to 0.

[0988] vri_bounding_box_left[i] indicates the horizontal coordinate of the top-left position of the bounding box of the i-th volume rectangle within the current atlas frame. The default value of vri_bounding_box_left[i] can be equal to 0.

[0989] vri_bounding_box_width[i] indicates the width of the bounding box of the i-th volume rectangle. The default value of vri_bounding_box_width[i] can be equal to 0.

[0990] vri_bounding_box_height[i] indicates the height of the bounding box of the i-th volume rectangle. The default value of vri_bounding_box_height[i] can be equal to 0.

[0991] vri_rectangle_number_of_objects_minus1[i] indicates the number of objects to be associated with the i-th volume rectangle.

[0992] The value of m can be set to vri_rectangle_number_of_objects_minus1[p]+1. For the value of m, the following elements can be included in the volume rectangle information.

[0993] `vri_rectangle_object_idx[i][j]` indicates the index of the `j`-th object associated with the `i`-th volume rectangle. The number of bits used to represent `vri_rectangle_object_idx[i]` can be equal to `vri_log2_max_object_idx_tracked`. When `vri_rectangle_object_idx[i]` is not present in the bitstream, its value can be inferred to be 0.

[0994] Figure 50 The configuration of the sample stream vpcc unit according to an embodiment is shown.

[0995] Figure 50 Show Figure 26 26000 bitstreams Figure 27 27000 bitstreams Figure 32 The specific hierarchical relationship between bitstreams such as 32000, etc.

[0996] Figure 50 The hierarchical structure of SEI messages in the atlas sub-bitstream is shown.

[0997] The sending method / apparatus according to the implementation method can generate, such as Figure 50 The atlas shown is a sub-bit stream.

[0998] Establish the relationships between the NAL unit data that constitute the sub-bit stream of the atlas.

[0999] In the SEI messages added to the atlas sub-bitstream, the information corresponding to `volumetric_tiling_info_objects()` is `scene_object_information()`. There are also `patch_information()` indicating the relationship to objects belonging to each patch, and `volumetric_rectangle_information()` which can be assigned to one or more objects.

[1000] Bitstream 50000 corresponds to Figure 26 The bitstream is 26000.

[1001] The NAL unit 50010 contained in the payload of the v-pcc unit in the atlas data of bitstream 50000 corresponds to... Figure 32 The bitstream is 32000.

[1002] Figure 50 The atlas frame parameter set 50020 corresponds to Figure 36 The set of image frame parameters.

[1003] Figure 50 The atlas tile group (or tile) layer 50030 corresponds to Figure 40 The atlas of tile groups (or tile layers).

[1004] Figure 50 SEI message 50040 corresponds to Figures 46 to 49 SEI message.

[1005] Figure 50 The atlas frame tile information 50050 corresponds to Figure 37 The atlas frame mosaic information.

[1006] Figure 50 The atlas tile group (or tile) header 50060 corresponds to Figure 40 The image set of puzzle pieces (or puzzle pieces) header.

[1007] Figure 50 Scene object information 50070 corresponds to Figure 46 Scene object information.

[1008] Figure 50 Object tag information 50080 corresponds to Figure 47 Object tag information.

[1009] Figure 50 The patch information 50090 corresponds to Figure 48 Information on replacement lenses.

[1010] Figure 50 The volume rectangle information 50100 corresponds to Figure 49 The volume rectangle information.

[1011] The atlas frame tile information 50050 can be identified by the atlas tile group (or tile) ID and can be included in the atlas frame parameter set 50020.

[1012] The atlas tile group (or tile) layer 50030 may include an atlas tile group (or tile) header 50060. The atlas tile group (or tile) header may be identified by the atlas tile group (or tile) address.

[1013] Scene object information 50070 can be identified by object index and object tag index.

[1014] Object tag information 50080 can be identified by the tag index.

[1015] Patch information 50090 can be identified by the tile group (or tile) address and patch object index.

[1016] The volume rectangle information 50100 can be identified by the rectangle object index.

[1017] The transmission method / apparatus according to the embodiment can encode point cloud data and generate, for example, Figure 50 The reference / hierarchical relationship information shown is used to generate the bitstream.

[1018] The receiving method / apparatus according to the embodiment can receive, such as Figure 50 The bitstream shown is used to recover the point cloud data contained within it. Additionally, data can be extracted from the bitstream. Figure 50 The atlas data is used to effectively encode and recover point cloud data.

[1019] Figure 51 This illustrates the configuration of atlas tile groups (or tiles) according to an embodiment.

[1020] Figure 51 The diagram illustrates the relationships between video frames, atlases, patches, and objects of point cloud data presented in a bitstream and signaled according to the method / apparatus of the embodiment.

[1021] Atlas frames can be generated and decoded by metadata processors 18005 and 19002 of the transmitting and receiving devices according to the embodiment. Thereafter, atlas bitstreams representing atlas frames can be formed according to the format according to the embodiment and transmitted / received by encapsulators / decapsulators 20004, 20005, 21009, and 22000 of the transmitting / receiving devices according to the embodiment.

[1022] An object is a target represented as point cloud data.

[1023] According to the implementation method, the position and / or size of the object can be changed dynamically. In this case, the changed atlas tile group or the configuration of the atlas tiles can be as follows: Figure 49 As shown in the figure.

[1024] Patches P1 to P3 can be configured with multiple scene objects O1 to O2 that constitute one or more objects. Frame 1 (video frame 1) can be composed of three atlas tile groups.

[1025] According to the implementation method, the atlas tile group may be referred to as an atlas tile. Atlas tile groups 1 to 3 may correspond to atlas tiles 1 to 3.

[1026] Atlas tile group 1 may include 3 patches. Patch 1 (P1) may include three objects O1 to O3. Patch 2 (P2) may contain one object (O2). Patch 3 (P3) may contain one object O1.

[1027] The method / apparatus according to the implementation can be based on patch information ( Figure 48 The field (pi_patch_object_idx[j]) corresponding to the object ID in patch_information() represents the mapping relationship between the patch and the object.

[1028] Atlas tile group (or tile) 2 may include two patches P1 and P2. Patch 1 (P1) may contain one object O1, and patch 2 (P2) may contain two objects O2.

[1029] The method / apparatus according to the implementation can be based on patch information ( Figure 48 The mapping relationship is represented by the field (pi_patch_object_idx[j]) corresponding to the object ID in patch_information.

[1030] Atlas tile group (or tile) 3 may include three patches P1, P2 and P3. Patch 1 (P1) may contain an object O2.

[1031] The method / apparatus according to the implementation can be based on patch information ( Figure 48 The mapping relationship is indicated by the field (pi_patch_object_idx[j]) corresponding to the object ID in patch_information.

[1032] Frame 2 can consist of three atlas tile groups (or tiles).

[1033] Atlas tile group (or tile) 1 (49000) may include two patches P1 and P2. Patch 1 may contain two objects O1 and O2. Patch 2 may contain one object O1.

[1034] The method / apparatus according to the implementation can be based on patch information ( Figure 46 The mapping relationship is indicated by the field (pi_patch_object_idx[j]) corresponding to the object ID in patch_information.

[1035] Atlas tile group (or tile) 2 (49000) may include two patches P1 and P2. Patch 1 (P1) may contain one object (O2). Patch 2 (P1) may contain two objects.

[1036] The method / apparatus according to the implementation can be based on patch information ( Figure 48 The mapping relationship is indicated by the field (pi_patch_object_idx[j]) corresponding to the object ID in patch_information.

[1037] Atlas tile group (or tile) 3 may include two patches P1 and P2. Patch 1 (P1) may contain an object O1.

[1038] The method / apparatus according to the implementation can be based on patch information ( Figure 48 The mapping relationship is indicated by the field (pi_patch_object_idx[j]) corresponding to the object ID in patch_information.

[1039] The data structures generated and transmitted / received by the V-PCC (V3C) system included in or connected to the point cloud data transmission / reception method / apparatus according to the embodiment will be described.

[1040] The following is for reference Figure 24 and Figure 25 The described method / apparatus can generate a file according to the corresponding implementation, and generate and send / receive the following data in the file.

[1041] The transmitting method / apparatus according to the embodiments can generate and transmit the following data structure based on a bit stream containing encoded point cloud data, and the receiving method / apparatus according to the embodiments can receive and parse the following data structure and recover the point cloud data contained in the bit stream.

[1042] Video-based point cloud compression represents the volumetric encoding of visual information from point clouds. The V-PCC bitstream containing the encoded point cloud sequence (CPCS) consists of V-PCC units carrying V-PCC parameter set (VPS) data, an encoded atlas bitstream, a 2D video encoded occupancy graph bitstream, a 2D video encoded geometry bitstream, and zero or more 2D video encoded attribute bitstreams.

[1043] Volumetric visual track

[1044] A volumetric visual track can be identified by the volumetric visual media handler type "volv" in the MediaBox's HandlerBox and the volumetric visual media header. Multiple volumetric visual tracks can exist in a file.

[1045] Volumetric Visual Media Head

[1046] Box type: "vvhd"

[1047] Container: MediaInformationBox

[1048] Mandatory: Yes

[1049] Quantity: Exactly one

[1050] For volumetric visual media headers, the box type is "vvhd" and the container is a MediaInformationBox. This is mandatory information and can exist as a header.

[1051] The volumetric visual track can be created using the VolumetricVisualMediaHeaderBox in the MediaInformationBox.

[1052] The structure of VolumetricVisualMediaHeaderBox is configured as follows.

[1053] aligned(8)class VolumetricVisualMediaHeaderBox

[1054] extends FullBox('vvhd',version=0,1){

[1055] }

[1056] "version" is an integer that specifies the version of the box.

[1057] Volumetric visual sample entries

[1058] Volumetric visual tracks can be created using VolumetricVisualSampleEntry.

[1059] The structure of VolumetricVisualSampleEntry can be configured as follows.

[1060] class VolumetricVisualSampleEntry(codingname)extends SampleEntry(codingname){

[1061] unsigned int(8)

[32] compressor_name;

[1062] }

[1063] `compressor_name` is the name used to provide information. It is formatted as a fixed 32-byte field, where the first byte is set to the number of bytes to be displayed, followed by the number of bytes of displayable data encoded in UTF-8, and then padded to complete a total of 32 bytes (including the size bytes). This field can be set to 0.

[1064] Volumetric visual samples

[1065] The format of volumetric visual samples is defined by the encoding system.

[1066] Next, we will describe the common data structures generated by the V-PCC system.

[1067] V-PCC unit head frame

[1068] This box exists in both the V-PCC track (in the sample entry) and all video-encoded V-PCC component tracks (in the scheme information). This box can contain the V-PCC unit header of the data carried by each track.

[1069] The structure of the V-PCC unit head frame can be configured as follows.

[1070] aligned(8)class VPCCUnitHeaderBox extends FullBox('vunt',version=0,0){

[1071] vpcc_unit_header()unit_header;

[1072] }

[1073] This box can contain something like vpcc_unit_header().

[1074] V-PCC Decoder Configuration Box

[1075] The V-PCC decoder configuration box may include VPCCDecoderConfigurationRecord.

[1076] class VPCCConfigurationBox extends Box('vpcC'){

[1077] VPCCDecoderConfigurationRecord()VPCCConfig;

[1078] }

[1079] This record contains a version field. The specification defines this record as version 1. Changes to the version number can indicate incompatible changes to the record.

[1080] VPCCParameterSet can include vpcc_parameter_set().

[1081] The SetupUnit array can be constant for the stream referenced by the sample entry. Decoder configuration records and atlas substream SEI messages exist.

[1082]

[1083] `configurationVersion` is a version field. Changes in the version number can indicate incompatible changes to a record.

[1084] lengthSizeMinusOne plus 1 indicates the length (in bytes) of the NALUnitLength field in the V-PCC sample of the stream to which this configuration record is applied.

[1085] For example, the size of a byte can be indicated by the value 0. The value of this field can be equal to ssnh_unit_size_precision_bytes_minus1 in the sample_stream_nal_header of the atlas substream.

[1086] numOfVPCCParameterSets specifies the number of V-PCC parameter set units that are signaled in the decoder configuration record.

[1087] VPCCParameterSetLength indicates the size (in bytes) of the vpccParameterSet field.

[1088] vpccParameterSet is a V-PCC unit of type VPCC_VPS that carries vpcc_parameter_set().

[1089] numOfSetupUnitArrays indicates the number of arrays of NAL units of the specified atlas type.

[1090] `array_completeness` equal to 1 indicates that all atlas NAL cells of the given type are in the following array, and none are in the stream. `array_completeness` equal to 0 indicates that additional atlas NAL cells of the indicated type are available in the stream. Default values ​​and allowed values ​​are constrained by the sample entry name.

[1091] NAL_unit_type indicates the type of atlas NAL unit in the following arrays. It can be one of the values ​​indicating an atlas NAL_ASPS, NAL_PREFIX_SEI, or NAL_SUFFIX_SEI NAL unit.

[1092] numNALUnits indicates the number of atlas NAL units of the indicated type included in the configuration record of the stream to which this configuration record is applied. SEI arrays may contain only SEI messages.

[1093] SetupUnitLength indicates the size (in bytes) of the setupUnit field. The length field may include the size of both the NAL unit header and the NAL unit payload, but not the length field itself.

[1094] A setupUnit can contain NAL units of type NAL_ASPS, NAL_AFPS, NAL_PREFIX_ESEI, NAL_PREFIX_NSEI, NAL_SUFFIX_ESEI, or NAL_SUFFIX_NSEI. When present in a setupUnit, NAL_PREFIX_ESEI, NAL_PREFIX_NSEI, NAL_SUFFIX_ESEI, or NAL_SUFFIX_NSEI contain SEI messages, that is, those that provide information about the stream as a whole. An example of such an SEI could be a user data SEI.

[1095] Figure 52 The structure of the V-PCC spatial region box according to an embodiment is shown.

[1096] Files generated and transmitted / received by the point cloud data transmitting / receiving device according to the embodiment and the system included in or connected to the transmitting / receiving device (see reference) Figure 24 and Figure 25 The boxes included in ) include the VPCC spatial region boxes.

[1097] This frame may contain information such as 3D bounding box information about the VPCC spatial region and label information related to the spatial region, atlas type group ID (or atlas tile ID), and patch IDs that may be included in the atlas tile group (or tile).

[1098] Additionally, this box may contain V-PCC component orbital group information related to the space region.

[1099] The V-PCC spatial region box can be included in the sample entries of the V-PCC orbital.

[1100] num_regions indicates the number of 3D spatial regions in the point cloud.

[1101] num_region_tile_groups indicates the number of atlas tile groups (or tiles) associated with some data of V-PCC objects included in a spatial region.

[1102] num_patch_updates indicates the number of patches in each atlas patch group (or patch) that belong to atlas patch groups (or patches) associated with some data included in the V-PCC object in the corresponding spatial region.

[1103] patch_id indicates the patch ID of a patch within a patch of an atlas patch group (or patch) that belongs to a patch associated with some data contained in a V-PCC object in the corresponding spatial region.

[1104] num_track_groups indicates the number of track groups associated with a 3D spatial region.

[1105] track_group_id indicates the track group of the V-PCC component that carries the relevant 3D spatial region.

[1106] label_id indicates the label ID associated with an atlas tile group (or tile) that is related to some data of a V-PCC object included in the corresponding spatial region.

[1107] label_language can indicate the language information of the label associated with a set of atlas tiles (or tiles) that are related to some data of V-PCC objects included in the corresponding spatial region.

[1108] label_name indicates the label name information associated with an atlas tile group (or tile) that is related to some data of a V-PCC object included in the corresponding spatial region.

[1109] The method / apparatus for transmitting or receiving point cloud data according to the embodiments and the system included in the transmitting / receiving apparatus can generate sample sets.

[1110] V-PCC Atlas Parameter Set Sample Group

[1111] The "vaps" grouping_type used for sample grouping indicates the V-PCC orbital to which samples are assigned to the atlas parameter set carried in that sample group. When a SampleToGroupBox with grouping_type equal to "vaps" exists, an accompanying SampleGroupDescriptionBox exists, containing the ID of the group to which the sample belongs.

[1112] A V-PCC track can contain at most one SampleToGroupBox with grouping_type equal to "vaps".

[1113]

[1114] numOfAtlasParameterSets specifies the number of atlas parameter sets that are signaled in the sample group description.

[1115] The atlasParameterSet is a sample_stream_vpcc_unit() instance that contains the atlas sequence parameter set and atlas frame parameter set associated with the set of samples.

[1116] The description entries of the V-PCC atlas parameter sample group according to the implementation method can be represented as follows.

[1117]

[1118] lengthSizeMinusOne plus 1 indicates the precision (in bytes) of the ssnu_nal_unit_size element in all sample stream NAL units notified by signals in the sample group description.

[1119] The atlasParameterSetNALUnit is a sample_stream_nal_unit() instance that contains the atlas sequence parameter set and atlas frame parameter set associated with the set of samples.

[1120] The point cloud data transmission / reception method / apparatus according to the embodiments and the system included in the point cloud data transmission / reception apparatus can generate dynamic spatial region sample groups.

[1121] The “dysr” grouping_type used for sample grouping indicates that samples in V-PCC orbits are assigned to the spatial region bounding boxes carried in the sample group.

[1122] When a SampleToGroupBox with grouping_type equal to "dysr" exists, there is an accompanying SampleGroupDescriptionBox with the same grouping type, containing the ID of the group to which the sample belongs.

[1123] A V-PCC track can contain at most one SampleToGroupBox with grouping_type equal to "dysr".

[1124]

[1125] The method / apparatus for transmitting / receiving point cloud data and the system included in the point cloud data transmitting / receiving apparatus, according to the embodiments, may provide track grouping as follows.

[1126] Space region orbit grouping

[1127] A TrackGroupTypeBox with track_group_type equal to "3drg" indicates that the track belongs to a group of V-PCC component tracks corresponding to a 3D spatial region. Tracks belonging to the same spatial region have the same track_group_id value for track_group_type "3drg", and the track_group_id of a track from one spatial region is different from the track_group_id of a track from any other spatial region.

[1128] aligned(8)class SpatialRegionGroupBox

[1129] extends TrackGroupTypeBox('3drg'){

[1130] }

[1131] The point cloud data transmission / reception method / apparatus according to the embodiments and the system included in the transmission / reception apparatus can provide a multi-track container for V-PCC bitstreams as follows.

[1132] General layout of multi-track ISOBMFF V-PCC container

[1133] V-PCC cells in a V-PCC bitstream are mapped to individual tracks within the container file based on their type. In a multi-track ISOBMFF V-PCC container, there are two types of tracks: V-PCC tracks and V-PCC component tracks.

[1134] A V-PCC component track is a video scheme track that carries 2D video-coded data of the occupancy map, geometry, and attribute sub-bitstreams of the V-PCC bitstream. Additionally, the following conditions must be met for a V-PCC component track to function:

[1135] a) In a sample entry, a new box can be inserted in the V-PCC system, which records the effect of the video stream contained in that track;

[1136] b) Orbit references can be introduced from V-PCC orbits to V-PCC component orbits to establish membership relationships of V-PCC component orbits in a specific point cloud represented by V-PCC orbits;

[1137] c) The track head flag can be set to 0 to indicate that the track does not directly contribute to the overall laying of the film, but rather to the V-PCC system.

[1138] Tracks belonging to the same V-PCC sequence can be time-aligned. Samples across different video coding V-PCC component tracks and V-PCC tracks that contribute to the same point cloud frame have the same rendering time. The decoding time of the V-PCC atlas sequence parameter set and atlas frame parameter set used for these samples is equal to or earlier than the orchestration time of the point cloud frame. Furthermore, all tracks belonging to the same V-PCC sequence have the same implicit or explicit editlist.

[1139] Synchronization between the basic streams in the component tracks can be handled by the ISOBMFF track timing structure (stts, ctts, and cslg) or an equivalent mechanism in the movie fragments.

[1140] Synchronization samples in V-PCC tracks and V-PCC component tracks may or may not be time-aligned. Without time alignment, random access may involve pre-rolling various tracks from different synchronization start times to allow for start at the desired time. With time alignment (e.g., as required by a V-PCC profile such as the basic toolset profile defined in [VPCC]), synchronization samples of V-PCC tracks can be considered random access points to V-PCC content, and random access can be accomplished by referring only to the synchronization sample information of the V-PCC track.

[1141] Based on this layout, the V-PCC ISOBMFF container may include the following:

[1142] • The V-PCC track contains the V-PCC parameter set and the atlas sub-bitstream parameter set (in the sample entries) as well as samples carrying the atlas sub-bitstream NAL units. This track also includes track references for other tracks that carry payloads of video compression V-PCC units (i.e., unit types VPCC_OVD, VPCC_GVD, and VPCC_AVD).

[1143] • Restricted video scheme track, where the samples contain access units of the video-coded basic stream of occupied graph data (i.e., payloads of V-PCC units of type VPCC_OVD).

[1144] • One or more restricted video scheme tracks, where the samples contain access units of the video-coded basic stream (i.e., payloads of V-PCC units of type VPCC_GVD).

[1145] • Zero or more restricted video scheme tracks, where the samples contain access units of the video-coded basic stream (i.e., payloads of V-PCC units of type VPCC_AVD) containing attribute data.

[1146] The point cloud data transmission / reception method / apparatus and the system included in the transmission / reception apparatus according to the embodiments may provide V-PCC tracks as follows.

[1147] V-PCC Track Sample Entries

[1148] Sample entry types: "vpc1", "vpcg"

[1149] Container: SampleDescriptionBox

[1150] Mandatory: Sample entries for "vpc1" or "vpcg" are mandatory.

[1151] Quantity: There may be one or more sample entries.

[1152] For V-PCC track sample entries, the sample entry type can be "vpc1" or "vpcg", and the container can be a SampleDescriptionBox. "vpc1" or "vpcg" sample entries can be mandatory, and one or more sample entries can exist.

[1153] The V-PCC track uses a VPCCSampleEntry of VolumetricVisualSampleEntry with an extended sample entry type of "vpc1" or "vpcg". The VPCC track sample entry contains a VPCCConfigurationBox.

[1154] Under the "vpc1" sample entry, all atlas sequence parameter sets, atlas frame parameter sets, or V-PCC SEIs are located in the setupUnit array. Under the "vpcg" sample entry, atlas sequence parameter sets, atlas frame parameter sets, and V-PCC SEIs can exist in this array or in the stream.

[1155] An optional BitRateBox can be present in the VPCC volume sample entry to signal the bit rate information of the V-PCC track.

[1156]

[1157] V-PCC Track Sample Format

[1158] Each sample in a V-PCC track corresponds to a single coded atlas access unit. Samples in various component tracks corresponding to the same frame may have the same orchestration time as V-PCC track samples. Each V-PCC sample may contain only one V-PCC unit payload of type VPCC_AD, which may include one or more atlas NAL units.

[1159]

[1160] nalUnit contains a single atlas NAL unit in the NAL unit sample stream format.

[1161] V-PCC track synchronization sample

[1162] Synchronization samples in the V-PCC track are samples containing Intra-Frame Random Access Point (IRAP) coded atlas access units. If needed, atlas sub-bitstream parameter sets such as ASPS, AAPS, AFPS, and SEI messages can be repeated at the synchronization samples to allow random access.

[1163] The point cloud data transmission / reception method / apparatus and the system included in the transmission / reception apparatus according to the embodiments may provide a video encoding V-PCC component track as follows.

[1164] Since it is meaningless to display decoded frames from attribute, geometry, or occupancy map tracks without reconstructing the point cloud on the player side, a restricted video scheme type is defined for these video encoding tracks.

[1165] Limited video solutions

[1166] The V-PCC component video track is represented as restricted video in the file and is identified by the “pccv” in the scheme_type field of the SchemeTypeBox of the RestrictedSchemeInfoBox of its restricted video sample entry.

[1167] There are no restrictions on the video codecs used to encode the attribute, geometry, and occupancy map V-PCC components. Furthermore, these components can be encoded using different video codecs.

[1168] Solution Information

[1169] The SchemeInformationBox can exist and contain a VPCCUnitHeaderBox.

[1170] The point cloud data transmission / reception method / apparatus and the system included in the transmission / reception apparatus according to the embodiments can provide a method for referencing V-PCC component tracks as follows.

[1171] To link a V-PCC track to a component video track, add three TrackReferenceTypeBoxes to the TrackReferenceBox inside the TrackBox of the V-PCC track, one for each component.

[1172] The TrackReferenceTypeBox contains an array of track_IDs for the video tracks that specify the V-PCC track references.

[1173] The `reference_type` property of `TrackReferenceTypeBox` identifies the component type, such as occupancy map, geometry, attribute, or occupancy map. These track reference types are as follows:

[1174] • "pcco": The reference track contains the Video Encoding Occupancy Map (V-PCC) component.

[1175] • “pccg”: The reference track contains the video coding geometry V-PCC component.

[1176] • “pcca”: The reference track contains the V-PCC component of the video encoding properties.

[1177] The type of V-PCC component carried by the reference-restricted video track and notified by a signal in the track's RestrictedSchemeInfoBox can match the reference type of the track reference from the V-PCC track.

[1178] The point cloud data transmission / reception method / apparatus according to the embodiments and the system included in the transmission / reception apparatus can provide a single track container of V-PCC bit streams as follows.

[1179] Single-track encapsulation of V-PCC data requires V-PCC encoding of the basic bitstream, which is represented by a single-track declaration.

[1180] In the case of a simple ISOBMFF encapsulation of a V-PCC encoded bitstream, single-track encapsulation of PCC data can be used. This bitstream can be directly stored as a single track without further processing. The V-PCC cell header data structure can be maintained within the bitstream. The single-track container of V-PCC data can be provided to the media workflow for further processing (e.g., multitrack file generation, transcoding, and DASH segmentation).

[1181] The point cloud data transmission / reception method / apparatus and the system included in the transmission / reception apparatus according to the embodiments may provide V-PCC bitstream tracks as follows.

[1182] Sample entry types: "vpe1", "vpeg"

[1183] Container: SampleDescriptionBox

[1184] Forced: The number of "vpe1" or "vpeg" sample entries is forced: one or more sample entries may exist.

[1185] The sample entry type is "vpe1" or "vpeg", and the container is SampleDescriptionBox. "vpe1" or "vpeg" sample entries are mandatory, and one or more sample entries may exist.

[1186] V-PCC bitstream tracks use VolumetricVisualSampleEntry with sample entry type "vpe1" or "vpeg".

[1187] The VPCC bitstream sample entry contains a VPCCConfigurationBox.

[1188] Under the "vpe1" sample entry, all atlas sequence parameter sets, atlas frame parameter sets, and SEIs are in the setupUnit array. Under the "vpeg" sample entry, atlas sequence parameter sets, atlas frame parameter sets, and SEIs may exist in this array or in the stream.

[1189] aligned(8)class VPCCBitStreamSampleEntry()extendsVolumetricVisualSampleEntry('vpe1'){

[1190] VPCCConfigurationBoxconfig;

[1191] }

[1192] V-PCC bitstream sample format

[1193] A V-PCC bitstream sample contains one or more V-PCC units (i.e., a V-PCC access unit) belonging to the same presentation time. Samples can be decoded either from themselves or from other samples in the V-PCC bitstream track.

[1194] V-PCC bitstream synchronization samples

[1195] V-PCC bitstream synchronization samples satisfy all of the following conditions:

[1196] • Can be decoded independently.

[1197] • The decoding of samples following the synchronization sample in the decoding order does not depend on any samples preceding the synchronization sample.

[1198] • All samples following the synchronized sample in the decoding order can be successfully decoded.

[1199] V-PCC bit flow sample

[1200] V-PCC bitstream subsamples are V-PCC units contained in V-PCC bitstream samples.

[1201] The V-PCC bitstream track should contain a SubSampleInformationBox in its SampleTableBox or in the TrackFragmentBox of its individual MovieFragmentBox, which lists the V-PCC bitstream subsamples.

[1202] The 32-bit cell header representing the V-PCC cell of a subsample can be copied to the 32-bit codec_specific_parameters field of the subsample entry in the SubSampleInformationBox. The V-PCC cell type of each subsample can be identified by parsing the codec_specific_parameters field of the subsample entry in the SubSampleInformationBox.

[1203] Because of the signaling scheme according to the embodiment, the receiving method / apparatus according to the embodiment can identify how the tiles are configured as 3D regions. The object according to the embodiment can represent an object or part of an object that is a target of the point cloud data. A tile is a unit that divides an atlas frame (2D). The receiving method / apparatus according to the embodiment can identify objects that match the tiles. As a result, the receiving method / apparatus according to the embodiment can effectively perform partial access to the point cloud data.

[1204] Figure 53 The DynamicSpatialRegionSample is shown according to an implementation method.

[1205] The point cloud data transmission / reception method / apparatus according to the embodiments and the system included in the transmission / reception apparatus can generate and transmit / receive information about dynamic spatial regions in a file. Figure 24 and Figure 25 This allows for the use of signals to notify dynamic spatial regions of point cloud data. References can be created and sent / received.

[1206] The point cloud data transmission / reception method / apparatus according to the embodiments and the system included in the transmission / reception apparatus can add dynamic spatial regions to timing metadata tracks or sample entries.

[1207] According to the implementation method, the timing metadata track can be transmitted as a separate track referenced by V3C track 25000.

[1208] If a V-PCC track has an associated timing metadata track of sample entry type "dysr", then the spatial region defined by the point cloud stream carried by the V-PCC track is considered a dynamic region.

[1209] The associated timing metadata track may contain a “cdsc” track reference for the V-PCC track that carries the atlas stream.

[1210] The point cloud data transmission / reception method / apparatus according to the embodiments and the system included in the transmission / reception apparatus can, for example, Figure 51 The VPCCSpatialRegionsBox data structure is added to the sample entry.

[1211] The DynamicSpatialRegionSampleEntry, which extends MetaDataSampleEntry, includes VPCCSpatialRegionsBox.

[1212] The specific syntax of VPCCSpatialRegionsBox is as follows: Figure 50 The configuration described in [the document / document].

[1213] The DynamicSpatialRegionSampleEntry includes DynamicSpatialRegionSample such as Figure 51 The elements shown. For a detailed definition of the elements, please refer to the reference. Figure 50 The description. Besides Figure 50 In addition, Figure 51 The elements can represent information related to dynamic spatial regions.

[1214] An atlas frame according to an embodiment may include multiple tiles. A method / apparatus according to an embodiment can generate, for example... Figure 51 The signaling information shown provides partial access to point cloud data at the tile level. Therefore, the receiving method / apparatus according to the embodiment can identify how tiles are configured for each region. Additionally, it can identify the objects to which tiles are matched.

[1215] Figure 54 The diagram illustrates a structure for encapsulating non-timing V-PCC data according to an embodiment.

[1216] The point cloud data transmission / reception method / apparatus according to the embodiments and the system included in the transmission / reception apparatus may be as follows: Figure 54 The diagram shows the encapsulation and transmission / reception of non-timing V-PCC data.

[1217] Non-timed V-PCC data is stored as image items in files. Two new item types (V-PCC item and V-PCC cell item) are defined to encapsulate non-timed V-PCC data.

[1218] Define a new handler type 4CC code "vpcc" and store it in the HandlerBox of the MetaBox to indicate the existence of V-PCC items, V-PCC cell items and other V-PCC encoded content representation information.

[1219] V-PCC Item 52000: A V-PCC item represents an independently decodeable V-PCC access unit. The item type "vpci" is defined to identify a V-PCC item. A V-PCC item stores the V-PCC unit payload of the atlas sub-bitstream. If a PrimaryItemBox exists, the item_id in that box is set to indicate the V-PCC item.

[1220] V-PCC Cell Item 52010: A V-PCC cell item is an item that represents V-PCC cell data. A V-PCC cell item stores the V-PCC cell payload, including occupancy, geometry, and attribute video data units. A V-PCC cell item stores only one V-PCC access cell-related data.

[1221] The item type of a V-PCC unit item is set according to the codec used to encode the corresponding video data unit. A V-PCC unit item is associated with the corresponding V-PCC unit header item properties and the codec-specific configuration item properties. V-PCC unit items are marked as hidden items because displaying them independently is meaningless.

[1222] To indicate the relationship between V-PCC projects and V-PCC unit projects, the following three project reference types are used. The "From" V-PCC Project to the Related V-PCC Unit Project defines the project reference.

[1223] • "pcco": Refer to the V-PCC unit project which contains occupied video data units.

[1224] “pccg”: Refer to the V-PCC unit project which contains geometric video data units.

[1225] “pcca”: Refer to the V-PCC unit project which contains attribute video data units.

[1226] V-PCC Configuration Project Nature 52020

[1227] Frame type: "vpcp"

[1228] Nature type: Descriptive project nature

[1229] Container: ItemPropertyContainerBox

[1230] Forced (per project): For V-PCC projects of type "vpci", it is

[1231] Quantity (per project): For V-PCC projects of type "vpci", one or more

[1232] For V-PCC configuration item properties, the box type is "vpcp", and the property type is descriptive item property. The container is ItemPropertyContainerBox. For V-PCC items of type "vpci", per item is mandatory. For V-PCC items of type "vpci", each item can have one or more properties.

[1233] The V-PCC parameter set is stored as the descriptive project property and associated with the V-PCC project.

[1234]

[1235] vpcc_unit_payload_size specifies the size (in bytes) of vpcc_unit_payload().

[1236] aligned(8)class VPCCConfigurationProperty extends ItemProperty('vpcc'){

[1237] vpcc_unit_payload_struct()[];

[1238] }

[1239] vpcc_unit_paylod() includes V-PCC units of type VPCC_VPS.

[1240] V-PCC Unit Header Item Type 52030

[1241] Box type: "vunt"

[1242] Nature type: Descriptive project nature

[1243] Container: ItemPropertyContainerBox

[1244] Forced (per project): For V-PCC projects of type "vpci" and for V-PCC unit projects, it is

[1245] Quantity (per item): one

[1246] For V-PCC cell header item properties, the box type is "vunt", the property type is descriptive item property, and the container is ItemPropertyContainerBox. For V-PCC items of type "vpci" and for V-PCC cell items, per item is mandatory. Each item can have one property.

[1247] The V-PCC cell header is stored as a descriptive project and is associated with the V-PCC project and the V-PCC cell project.

[1248] aligned(8)class VPCCUnitHeaderProperty()extends ItemFullProperty('vunt',version=0,0){

[1249] vpcc_unit_header();

[1250] }

[1251] Figure 55 This is a flowchart of a point cloud data transmission method according to an implementation method.

[1252] The point cloud data transmission apparatus according to the embodiments may include a file / fragment encapsulator (hereinafter referred to as an encapsulator) and / or a transmitter, such as Figure 55 As shown. The point cloud data encoder, file / fragment encapsulator, and transmitter according to the embodiments can be collectively referred to as the point cloud data transmission apparatus and / or point cloud data system according to the embodiments. In this document, they can be simply referred to as the method / apparatus according to the embodiments.

[1253] The transmitting / receiving apparatus according to the implementation may include an ISOBMFF module. The ISOBMFF module is a module configured to provide APIs for creating / modifying / deleting frames that constitute ISOBMFF format files.

[1254] The transmitting / receiving apparatus according to the implementation may include a VPCCBitstream module. The VPCCBitstream module generates / parses data in VPCC bitstream format.

[1255] The wrapper according to the embodiment can generate the frame structure required to configure and encode the V-PCC bitstream in ISOBMFF file format. The transmitter can send the generated data. A detailed flowchart of the operation of the wrapper according to the embodiment is configured as follows. Each operation is performed by the wrapper, method / apparatus, etc. according to the embodiment.

[1256] 0. Input the V-PCC encoded bit stream from the point cloud data transmission device (or encoder).

[1257] Can create ISOM files. Can add VpccTrack(AD) to a file. Informs the system's ISOBMFF level to create a new track. Can receive tracks from the ISOBMFF level. Can add vpcc_sample_entry(track) to a file.

[1258] According to the implementation method, the gf function represents the box required to create an ISOBMFF format file, that is, the API for adding / modifying / deleting tracks, samples, etc.

[1259] 1. V-PCC tracks can be created based on the ISOBMFF file structure.

[1260] 1-1. In the case of dynamic space regions, timed metadata tracks can be created.

[1261] You can add TimedMetaTrack (AD) to a file. You can notify the ISOBMFF level of the existence of a new track. You can receive tracks from the ISOBMFF level.

[1262] 2. Sample entries can be created in the V-PCC track created in step 1 above.

[1263] 3. VPCC parameter set information (VPS info) can be obtained from the input bit stream.

[1264] Sample stream V-PCC cells (VPS) can be requested. Requests can be made within the V-PCC bitstream based on the V-PCC buffer (position). Sample stream V-PCC cells can be obtained.

[1265] 4. VPS information (info) can be added to the sample entry.

[1266] V-PCC decoder configuration (track, sample stream V-PCC unit) can be requested from the ISOBMFF level.

[1267] 5. Atlas scene object information can be obtained from the input bitstream.

[1268] You can request new object information (AD). You can request it in the V-PCC bitstream based on the V-PCC buffer (position). You can obtain new object information.

[1269] 5-1. Atlas object label information can be obtained from the input bitstream.

[1270] Object tag information (AD) can be requested. It can be requested within the V-PCC bitstream based on the V-PCC buffer (position). Object tag information can be obtained.

[1271] 5-2. Atlas patch information can be obtained from the input bitstream.

[1272] You can request patch information (AD). You can request it in the V-PCC bitstream based on the V-PCC buffer (position). You can obtain patch information.

[1273] 6. Based on the atlas volume tiling information, a VPCCSpatialRegionsBox structure suitable for the point cloud system file format can be created and added to sample entries. Alternatively, in the case of dynamic spatial regions, the VPCCSpatialRegionsBox structure can be added to sample entries in the timing metadata track.

[1274] You can request V-PCC decoder configuration (track, V-PCC spatial region box).

[1275] 7. The VPCC cell header information of the atlas can be obtained from the input bit stream.

[1276] You can request a V-PCC cell header frame (AD). You can request it within the V-PCC bitstream based on the V-PCC buffer (position). You can obtain the cell header.

[1277] 8. VPCC cell header information can be added to sample entries.

[1278] You can request the V-PCC unit header (track, unit header).

[1279] 9. NAL cell data from the atlas can be added to sample entries or samples based on nalType.

[1280] Under the loop-size operation (LOOP), the following operations can be performed: V-PCC cells (ADs) can be requested in the bitstream, and NAL cells in the sample stream can be obtained.

[1281] Under the loop of NAL counting, the following operations can be performed on the atlas data.

[1282] When the NAL type is NAL_ASPS, the sample stream NAL unit ASPS can be o...

Claims

1. A method for transmitting point cloud data, the method comprising the following steps: The occupancy data, geometric data, and attribute data of the point cloud data are encoded based on a video-based point cloud compression scheme. Encapsulate the coded occupancy data, coded geometric data, and coded attribute data in a file; as well as Send the file, The file includes a first component track, a second component track, and a third component track. The first component track includes the occupancy data, the second component track includes the geometric data, and the third component track includes the attribute data. The file also includes a track containing spatial region information for partial access to the point cloud data, and The spatial region information includes identifiers for the spatial regions of the point cloud data, information indicating the number of atlas tiles associated with the spatial region, and tile identifiers for identifying the atlas tiles associated with the spatial region.

2. The method according to claim 1, in, The spatial region information includes static information related to the spatial region or dynamic information related to the spatial region over time.

3. A device for receiving point cloud data, the device comprising: A receiver configured to receive a file including the point cloud data; A decapsulator configured to decapsulate the file; as well as A decoder configured to decode the point cloud data based on a video-based point cloud decompression scheme. The file includes a first component track, a second component track, and a third component track. The first component track includes occupancy data, the second component track includes geometric data, and the third component track includes attribute data. The file also includes a track containing spatial region information for partial access to the point cloud data, and The spatial region information includes identifiers for the spatial regions of the point cloud data, information indicating the number of atlas tiles associated with the spatial region, and tile identifiers for identifying the atlas tiles associated with the spatial region.

4. The device according to claim 3, in, The spatial region information includes static information related to the spatial region or dynamic information related to the spatial region over time.

5. A device for transmitting point cloud data, the device comprising: The encoder is configured to encode the occupancy data, geometric data, and attribute data of the point cloud data based on a video-based point cloud compression scheme; A wrapper configured to encapsulate encoded occupancy data, encoded geometry data, and encoded attribute data in a file; as well as A sender configured to send the file. The file includes a first component track, a second component track, and a third component track. The first component track includes the occupancy data, the second component track includes the geometric data, and the third component track includes the attribute data. The file also includes a track containing spatial region information for partial access to the point cloud data, and The spatial region information includes identifiers for the spatial regions of the point cloud data, information indicating the number of atlas tiles associated with the spatial region, and tile identifiers for identifying the atlas tiles associated with the spatial region.

6. A method for receiving point cloud data, the method comprising the following steps: Receive a file containing the point cloud data; Decapsulate the file; as well as The point cloud data is decoded based on a video-based point cloud decompression scheme. The file includes a first component track, a second component track, and a third component track. The first component track includes occupancy data, the second component track includes geometric data, and the third component track includes attribute data. The file also includes a track containing spatial region information for partial access to the point cloud data, and The spatial region information includes identifiers for the spatial regions of the point cloud data, information indicating the number of atlas tiles associated with the spatial region, and tile identifiers for identifying the atlas tiles associated with the spatial region.

Citation Information

Patent Citations

  • Method, apparatus and stream for immersive video format

    CN110383342A

  • XR device and method for controlling the same

    US20190392647A1