Point cloud data transmitting device, point cloud data transmitting method, point cloud data receiving device, and point cloud data receiving method

Efficient transmission and reception of point cloud data through encoding, encapsulation, and decoding processes address latency and complexity issues, enabling high-quality services in VR, AR, MR, and autonomous driving.

JP7749058B2Active Publication Date: 2025-10-03LG ELECTRONICS INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024067219
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-04-12
Filing Date
2024-04-18
Publication Date
2025-10-03
Estimated Expiration
2040-12-15

AI Technical Summary

Technical Problem

The large number of points in 3D space and the high processing requirements pose challenges in efficiently transmitting and receiving point cloud data, leading to latency and encoding/decoding complexity.

Method used

A method involving encoding, encapsulating, and transmitting point cloud data, followed by decapsulating and decoding it, utilizing devices and methods such as point cloud video acquisition, encoding, file/segment encapsulation, and rendering to facilitate efficient transmission and reception.

Benefits of technology

Enables high-quality point cloud services, supporting various video codec methods and applications like VR, AR, MR, and autonomous driving, while optimizing processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007749058000007
    Figure 0007749058000007
  • Figure 0007749058000008
    Figure 0007749058000008
  • Figure 0007749058000009
    Figure 0007749058000009
Patent Text Reader

Abstract

To provide a point cloud data transmission and reception device and a point cloud data transmission and reception method for efficiently transmitting and receiving a point cloud.SOLUTION: A point cloud data transmission method according to an embodiment includes encoding point cloud data and transmitting the point cloud data. A point cloud data reception method according to an embodiment includes receiving point cloud data, decoding the point cloud data, and rendering the point cloud data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The embodiment provides a method for providing point cloud content to provide users with various services such as VR (Virtual Reality), AR (Augmented Reality), MR (Mixed Reality) and autonomous driving services. [Background technology]

[0002] A point cloud is a collection of points in 3D space. There is a problem that the number of points in 3D space is large, making it difficult to generate point cloud data.

[0003] The problem with sending and receiving point cloud data is that it requires a large amount of processing. Summary of the Invention [Problem to be solved by the invention]

[0004] The technical objective of the embodiments is to provide a point cloud data transmitting device, a transmitting method, a point cloud data receiving device, and a receiving method for efficiently transmitting and receiving point clouds in order to solve the above problems.

[0005] A technical problem related to the embodiments is to provide a point cloud data transmitting device, transmitting method, and point cloud data receiving device and receiving method that solve the latency and encoding / decoding complexity.

[0006] However, the scope of the invention is not limited to the above technical problems, but can be extended to other technical problems that can be derived by a person skilled in the art based on the entire contents of this specification. [Means for solving the problem]

[0007] To achieve the above object and other advantages, a point cloud data transmission method according to an embodiment includes the steps of encoding point cloud data, encapsulating the point cloud data, and transmitting the point cloud data.

[0008] A method for receiving point cloud data according to an embodiment includes receiving point cloud data, decapsulating the point cloud data, and decoding the point cloud data. [Effects of the Invention]

[0009] The point cloud data transmitting method, transmitting device, point cloud data receiving method and receiving device according to the embodiments can provide high-quality point cloud services.

[0010] The point cloud data transmitting method, transmitting device, point cloud data receiving method and receiving device according to the embodiments can achieve various video codec methods.

[0011] The point cloud data transmitting method, transmitting device, point cloud data receiving method and receiving device according to the embodiments can provide general-purpose point cloud content such as an autonomous driving service. [Brief explanation of the drawings]

[0012] The drawings are attached to provide a further understanding of the embodiments and together with the description of the embodiments illustrate the embodiments.

[0013]

Figure 1

[0014]

Figure 2

[0015]

Figure 3

[0016]

Figure 4

[0017]

Figure 5

[0018]

Figure 6

[0019]

Figure 7

[0020]

Figure 8

[0021]

Figure 9

[0022]

Figure 10

[0023]

Figure 11

[0024]

Figure 12

[0025]

Figure 13

[0026]

Figure 14

[0027]

Figure 15

[0028]

Figure 16

[0029]

Figure 17

[0030]

Figure 18

[0031]

Figure 19

[0032]

Figure 20

[0033]

Figure 21

[0034]

Figure 22

[0035]

Figure 23

[0036]

Figure 24

[0037]

Figure 25

[0038]

Figure 26

[0039]

Figure 27

[0040]

Figure 28

[0041]

Figure 29

[0042]

Figure 30

[0043]

Figure 31

[0044]

Figure 32

[0045]

Figure 33

[0046]

Figure 34

[0047]

Figure 35

[0048]

Figure 36

[0049]

Figure 37

[0050]

Figure 38

[0051]

Figure 39

[0052]

Figure 40

[0053]

Figure 41

[0054]

Figure 42

[0055]

Figure 43

[0056]

Figure 44

[0057]

Figure 45

[0058]

Figure 46

[0059]

Figure 47

[0060]

Figure 48

[0061]

Figure 49

[0062]

Figure 50

[0063] The syntax of the information contained in the atlas bitstream of Figure 50 will be explained below.

[0064]

Figure 51

[0065]

Figure 52

[0066]

Figure 53

[0067]

Figure 54

[0068]

Figure 55

[0069]

Figure 56

[0070]

Figure 57

[0071]

Figure 58

[0072]

Figure 59

[0073]

Figure 60

[0074]

Figure 61

[0075]

Figure 62

[0076]

Figure 63

[0077]

Figure 64

[0078]

Figure 65

[0079] Preferred embodiments will now be described in detail with reference to the accompanying drawings. The following detailed description, with reference to the accompanying drawings, is intended to illustrate preferred embodiments rather than merely illustrating embodiments that may be implemented by the embodiments. The following detailed description includes details to provide a thorough understanding of the embodiments. However, it will be apparent to those skilled in the art that embodiments may be practiced without such details.

[0080] Although most of the terms used in the examples are common and widely used in the relevant fields, some of them have been arbitrarily selected by the applicant, and their meanings will be explained in detail below as necessary. Therefore, the examples should be understood based on the intended meaning of the terms, rather than the simple names or meanings of the terms.

[0081] FIG. 1 illustrates an example of a transmitting / receiving system architecture for providing point cloud content according to an embodiment.

[0082] This document provides a method for providing point cloud content to provide users with various services such as VR (Virtual Reality), AR (Augmented Reality), MR (Mixed Reality), and autonomous driving services. The point cloud content according to the embodiment refers to data representing objects using points, and is also called a point cloud, point cloud data, point cloud video data, point cloud image data, etc.

[0083] A point cloud data transmission device 10000 according to an embodiment includes a point cloud video acquisition unit 10001, a point cloud video encoder 10002, a file / segment encapsulation unit 10003, and / or a transmitter (or communication module) 10004. The transmission device according to an embodiment can acquire, process, and transmit point cloud video (or point cloud content). According to an embodiment, the transmission device includes a fixed station, a base transceiver system (BTS), a network, an AI (artificial intelligence) device and / or system, a robot, an AR / VR / XR device and / or server, etc. In addition, depending on the embodiment, the transmitting device 10000 may include devices that communicate with base stations and / or other wireless devices using wireless connection technologies (e.g., 5G NR (New RAT), LTE (Long Term Evolution)), robots, vehicles, AR / VR / XR devices, mobile devices, home appliances, IoT (Internet of Things) devices, AI devices / servers, etc.

[0084] According to the embodiment, a point cloud video acquisition unit 10001 acquires a point cloud video through a process of capturing, synthesizing or generating a point cloud video.

[0085] A point cloud video encoder 10002 according to an embodiment encodes point cloud video data. Depending on the embodiment, the point cloud video encoder 10002 may be referred to as a point cloud encoder, a point cloud data encoder, an encoder, etc. Furthermore, point cloud compression coding (encoding) according to an embodiment is not limited to the above-described embodiments. The point cloud video encoder outputs a bitstream including encoded point cloud video data. The bitstream includes not only the encoded point cloud video data but also signaling information related to the encoding of the point cloud video data.

[0086] The encoder according to the embodiment supports both a Geometry-based Point Cloud Compression (G-PCC) encoding method and / or a Video-based Point Cloud Compression (V-PCC) encoding method. The encoder can also encode point clouds (referring to point cloud data or all points) and / or signaling data related to point clouds. Specific operations of encoding according to the embodiment will be described later.

[0087] Meanwhile, the term V-PCC used in this specification means Video-based point Cloud Compression (V-PCC), and the term V-PCC is the same as Visual Volumetric Video-based Coding (V3C), and can be referred to as complementary terms.

[0088] A File / Segment Encapsulation module 10003 according to the embodiment encapsulates point cloud data in the form of a file and / or a segment. A point cloud data transmission method / apparatus according to the embodiment can transmit point cloud data in the form of a file and / or a segment.

[0089] According to an embodiment, a transmitter (or communication module) 10004 transmits encoded point cloud video data in the form of a bitstream. Depending on the embodiment, the file or segment may be transmitted to a receiving device via a network or stored on a digital storage medium (e.g., USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.). According to an embodiment, the transmitter can communicate with a receiving device (or receiver) via wired / wireless communication via a network such as 4G, 5G, or 6G. The transmitter can also perform necessary data processing operations depending on the network system (e.g., a communication network system such as 4G, 5G, or 6G). The transmitter can also transmit encapsulated data on an on-demand basis.

[0090] A point cloud data receiving device 10005 according to an embodiment includes a receiver 10006, a file / segment decapsulator 10007, a point cloud video decoder 10008, and / or a renderer 10009. According to an embodiment, the receiving device includes a device, robot, vehicle, AR / VR / XR device, mobile device, home appliance, IoT (Internet of Things) device, AI device / server, etc. that communicates with a base station and / or other wireless device using a wireless connection technology (e.g., 5G NR (New RAT), LTE (Long Term Evolution)).

[0091] A receiver 10006 according to an embodiment receives a bitstream including point cloud video data, and transmits feedback information to a point cloud data transmitting device 10000 according to an embodiment.

[0092] A File / Segment Decapsulation module 10007 decapsulates files and / or segments containing point cloud data. The decapsulation module according to the embodiment performs the reverse process of the encapsulation process according to the embodiment.

[0093] A point cloud video decoder 10007 decodes the received point cloud video data. The decoder according to the embodiment performs the reverse process of the encoding according to the embodiment.

[0094] A Renderer 10007 renders the decoded point cloud video data. According to an embodiment, the Renderer 10007 transmits feedback information acquired at the receiving end to a point cloud video decoder 10006. According to an embodiment, the point cloud video data transmits feedback information to a receiver. According to an embodiment, the feedback information received by the point cloud transmitting device is provided to a point cloud video encoder.

[0095] In the drawing, dotted arrows indicate the transmission path of feedback information acquired by the receiving device 10005. The feedback information is information for reflecting interaction with a user consuming point cloud content and includes user information (e.g., head orientation information, viewport information, etc.). In particular, if the point cloud content is for a service requiring interaction with a user (e.g., an autonomous driving service), the feedback information is transmitted to a content transmitting side (e.g., the transmitting device 10000) and / or a service provider. Depending on the embodiment, the feedback information may be used not only by the transmitting device 10000 but also by the receiving device 10005, or may not be provided.

[0096] According to an embodiment, head orientation information is information regarding the position, direction, angle, movement, etc. of the user's head. According to an embodiment, the receiving device 10005 calculates viewport information based on the head orientation information. The viewport information is information regarding the area of ​​the point cloud video that the user is viewing. The viewpoint is the point at which the user is viewing the point cloud video and refers to the center of the viewport area. That is, the viewport is an area centered on the viewpoint, and the size, shape, etc. of the area are determined by the FOV (Field of View). Therefore, the receiving device 10004 can extract viewport information based on the vertical or horizontal FOV supported by the device in addition to the head orientation information. Furthermore, the receiving device 10005 performs gaze analysis, etc. to confirm the user's point cloud consumption method, the point cloud video area the user is gazing at, the gaze time, etc. According to an embodiment, the receiving device 10005 transmits feedback information including the results of the gaze analysis to the transmitting device 10000. According to an embodiment, the feedback information is obtained during the rendering and / or display process. According to an embodiment, the feedback information is obtained by one or more sensors included in the receiving device 10005. According to another embodiment, the feedback information is obtained by the renderer 10009 or another external element (or device, component, etc.). The dotted lines shown in FIG. 1 indicate the transmission process of the feedback information obtained by the renderer 10009. The point cloud content providing system processes (encodes / decodes) the point cloud data based on the feedback information. Therefore, the point cloud video data decoder 10008 can perform a decoding operation based on the feedback information. The receiving device 10005 also transmits the feedback information to the transmitting device. The transmitting device (or point cloud video data encoder 10002) performs an encoding operation based on the feedback information.Therefore, the point cloud content providing system does not process (encode / decode) all point cloud data, but efficiently processes necessary data (e.g., point cloud data corresponding to the user's head position) based on feedback information, and can provide point cloud content to the user.

[0097] In the embodiment, the sending device 10000 is referred to as an encoder, a sending device, a transmitter, etc., and the receiving device 10004 is referred to as a decoder, a receiving device, a receiver, etc.

[0098] 1 according to an embodiment (processed in a series of processes of acquisition / encoding / transmission / decoding / rendering), the point cloud data is also called point cloud content data or point cloud video data. According to an embodiment, the point cloud content data can be used as a concept including metadata or signaling information related to the point cloud data.

[0099] The elements of the point cloud content providing system shown in FIG. 1 may be embodied in hardware, software, a processor, and / or a combination thereof.

[0100] The embodiment provides point cloud content to provide users with various services such as VR (Virtual Reality), AR (Augmented Reality), MR (Mixed Reality), and autonomous driving services.

[0101] To provide a point cloud content service, a point cloud video is first obtained. The obtained point cloud video is transmitted through a series of processes, and the received data is then processed and rendered back into the original point cloud video at the receiving end. This allows the point cloud video to be provided to users. The embodiment provides a solution required to effectively perform this series of processes.

[0102] The overall process for providing point cloud content services (point cloud data transmitting method and / or point cloud data receiving method) includes an acquisition process, an encoding process, a transmission process, a decoding process, a rendering process and / or a feedback process.

[0103] According to an embodiment, the process of providing point cloud content (or point cloud data) is also referred to as a point cloud compression process. According to an embodiment, the point cloud compression process refers to a geometry-based point cloud compression process.

[0104] Each element of the point cloud data transmitting device and the point cloud data receiving device according to the embodiment means hardware, software, a processor, and / or a combination thereof.

[0105] To provide a point cloud content service, a point cloud video is first obtained. The obtained point cloud video is then transmitted through a series of processes, and the received data is then reprocessed and rendered back into the original point cloud video at the receiving end. This allows the point cloud video to be provided to users. The present invention provides a solution for effectively performing this series of processes.

[0106] All processes for providing the point cloud content service include an acquisition process, an encoding process, a transmission process, a decoding process, a rendering process and / or a feedback process.

[0107] The point cloud compression system includes a transmitting device and a receiving device. The transmitting device encodes the point cloud video and outputs a bitstream, which is then transmitted to the receiving device via a digital storage medium or a network in the form of a file or streaming (streaming segments). The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.

[0108] The transmitting device generally includes a point cloud video acquisition unit, a point cloud video encoder, a file / segment encapsulation unit, and a transmitter. The receiving device generally includes a receiver, a file / segment decapsulation unit, a point cloud video decoder, and a renderer. The encoder is also called a point cloud video / image / picture / frame encoder, and the decoder is also called a point cloud video / image / picture / frame decoder. The transmitter may be included in the point cloud video encoder. The receiver may be included in the point cloud video decoder. The renderer may include a display unit, and the renderer and / or the display unit may be configured as separate devices or external components. The transmitting device and the receiving device may further include another internal or external module / unit / component for a feedback process.

[0109] The operation of the receiving device according to the embodiment follows the reverse process of the operation of the transmitting device.

[0110] The point cloud video acquisition unit performs a process to acquire a point cloud video through a point cloud video capture, synthesis, or generation process. The acquisition process generates 3D position (x, y, z) / attribute (color, reflectance, transparency, etc.) data for a large number of points, such as a PLY (Polygon File format or the Stanford Triangle format) file. For videos with multiple frames, one or more files can be obtained. Metadata related to the point cloud (e.g., metadata about the capture) is generated during the capture process.

[0111] The point cloud data transmitting device according to the embodiment may include an encoder for encoding the point cloud data and a transmitter for transmitting the point cloud data, which may be transmitted in the form of a bitstream containing the point cloud.

[0112] A point cloud data receiving device according to an embodiment includes a receiving unit that receives point cloud data, a decoder that decodes the point cloud data, and a renderer that renders the point cloud data.

[0113] The method / apparatus according to the embodiment represents a point cloud data transmitting device and / or a point cloud data receiving device.

[0114] FIG. 2 illustrates an example of capturing point cloud data according to an embodiment.

[0115] In the embodiment, the point cloud data is obtained by a camera, etc. In the embodiment, the capturing method may be, for example, an inward-facing method and / or an outward-facing method.

[0116] The inward-facing method according to the embodiment is a capture method in which one or more cameras capture an object of point cloud data from the outside to the inside of the object.

[0117] In the outward facing method according to the embodiment, one or more cameras capture an object of point cloud data from the inside to the outside of the object. For example, in the embodiment, there are four cameras.

[0118] According to an embodiment, point cloud data or point cloud content may be video or still images of objects / environments represented in various forms in 3D space. According to an embodiment, point cloud content may include video / audio / images of objects.

[0119] Point cloud content capture is composed of a combination of camera equipment (a combination of an infrared pattern projector and an infrared camera) that can obtain depth and an RGB camera that can extract color information corresponding to the depth information. Depth information can also be extracted using LiDAR, a radar system that measures the position coordinates of reflectors by emitting a laser pulse and measuring the time it takes for the pulse to reflect back. Geometry, consisting of points in 3D space, can be extracted from the depth information, and attributes representing the color / reflectance of each point can be extracted from the RGB information. Point cloud content consists of position (x, y, z), color (YCbCr or RGB), or reflectance (r) information for each point. Point cloud content can be created in two ways: outward-facing, which captures the external environment, and inward-facing, which captures a central object. When creating point cloud content that allows users to freely view an object (e.g., a character, player, object, actor, or other core object) in a VR / AR environment in 360°, the capture camera is configured in an inward-facing manner. In addition, when the current surrounding environment of the vehicle is configured as point cloud content, such as in autonomous driving, the capture camera configuration uses an outward-facing method.Since point cloud content is captured by multiple cameras, a camera calibration process may be required before capturing content to set a global coordinate system between the cameras.

[0120] Point cloud content is video or still footage of objects / environments shown in various forms in 3D space.

[0121] Alternatively, the point cloud content can be obtained by synthesizing any point cloud video based on the captured point cloud video. Alternatively, if you want to provide a point cloud video for a computer-generated virtual space, actual camera capture may not be performed. In this case, the capture process can simply be replaced by a process that generates the relevant data.

[0122] The captured point cloud video requires post-processing to improve the quality of the content. Although the maximum and minimum depth values ​​can be adjusted within the range provided by the camera equipment during the video capture process, point data from undesired areas may still be included. Therefore, post-processing can be performed to remove undesired areas (e.g., background) or to fill spatial holes by recognizing connected spaces. Point clouds extracted from cameras sharing a spatial coordinate system can be integrated into a single content by converting each point to a global coordinate system based on the position coordinates of each camera acquired through a calibration process. This allows for the generation of a single point cloud content covering a wide area, or for the acquisition of point cloud content with a high point density.

[0123] A point cloud video encoder encodes input point cloud video into one or more video streams. A point cloud video may contain multiple frames, with each frame corresponding to a still image / picture. In this document, point cloud video includes point cloud video / frame / picture / video / audio / image, and point cloud video can be mixed with point cloud video / frame / picture. A point cloud video encoder performs a video-based point cloud compression (V-PCC) procedure. For compression and coding efficiency, the point cloud video encoder performs a series of procedures, including prediction, transformation, quantization, and entropy coding. The encoded data (encoded video / video information) is output in the form of a bitstream. Based on the V-PCC procedure, a point cloud video encoder encodes the point cloud video into geometry video, attribute video, occupancy map video, and auxiliary information, as described below. The geometry video contains geometry images, the attribute video contains attribute images, the occupancy map video contains occupancy map images, the auxiliary information contains auxiliary patch information, and the attribute video / image contains texture videos / images.

[0124] The encapsulation processing unit (file / segment encapsulation module) 10003 encapsulates the encoded point cloud video data and / or point cloud video-related metadata in a format such as a file. Here, the point cloud video-related metadata is transmitted from a metadata processing unit or the like. The metadata processing unit may be included in the point cloud video encoder or may be configured as a separate component / module. The encapsulation unit may encapsulate the data in a file format such as ISOBMFF, or may process the data in other formats such as DASH segments. According to an embodiment, the encapsulation processing unit may include the point cloud video-related metadata in a file format. For example, the point cloud video-related metadata may be included in boxes at various levels in the ISOBMFF file format, or may be included in separate tracks within the file. According to an embodiment, the encapsulation processing unit encapsulates the point cloud video-related metadata itself in a file. The transmission processing unit may perform processing for transmission on the point cloud video data encapsulated in the file format. The transmission processing unit may be included in the transmission unit or may be configured as a separate component / module. The transmission processing unit processes the point cloud video data according to any transmission protocol. The processing for transmission includes processing for transmission via a broadcast network and processing for transmission via broadband. The transmission processing unit according to the embodiment performs processing for transmission of not only the point cloud video data but also the point cloud video related metadata transmitted from the metadata processing unit.

[0125] The transmitter 10004 transmits the encoded video / image information or data output in the form of a bitstream to a receiver of a receiving device via a digital storage medium or network in the form of a file or streaming. Digital storage media include USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter includes elements for generating a media file in a predetermined file format and elements for transmission via a broadcast / communication network. The receiver extracts the bitstream and transmits it to a decoding device.

[0126] The receiver 10003 receives the point cloud video data transmitted by the point cloud video transmitting device of the present invention. Depending on the channel transmitted, the receiver may receive the point cloud video data via a broadcast network, via broadband, or via a digital storage medium.

[0127] The receiving processing unit processes the received point cloud video data in accordance with the transmission protocol. The receiving processing unit may be included in the receiving unit or may be configured as a separate component / module. Corresponding to the processing for transmission performed on the transmitting side, the receiving processing unit performs the reverse process of the transmitting processing unit described above. The receiving processing unit transmits the acquired point cloud video data to the decapsulation processing unit, and transmits the acquired point cloud video-related metadata to the metadata processing unit. The point cloud video-related metadata acquired by the receiving processing unit may be in the form of a signaling table.

[0128] The decapsulation processing unit (file / segment decapsulation module) 10007 decapsulates the point cloud video data in a file format transmitted from the receiving processing unit. The decapsulation processing unit decapsulates a file such as ISOBMFF to obtain a point cloud video bitstream or point cloud video-related metadata (metadata bitstream). The obtained point cloud video bitstream is transmitted to a point cloud video decoder, and the obtained point cloud video-related metadata (metadata bitstream) is transmitted to a metadata processing unit. The point cloud video bitstream may include metadata (metadata bitstream). The metadata processing unit may be included in the point cloud video decoder or may be configured as a separate component / module. The point cloud video-related metadata obtained by the decapsulation processing unit may be in the form of a box or track in a file format. If necessary, the decapsulation processing unit receives metadata required for decapsulation from the metadata processing unit. The point cloud video-related metadata is transmitted to the point cloud video decoder for use in the point cloud video decoding procedure or to the renderer for use in the point cloud video rendering procedure.

[0129] The point cloud video decoder receives a bitstream and performs operations corresponding to those of the point cloud video encoder to decode the video / image. In this case, the point cloud video decoder separates the point cloud video into geometry video, attribute video, occupancy map video, and auxiliary information, as described below, and decodes it. The geometry video includes geometry images, the attribute video includes attribute images, and the occupancy map video includes occupancy map images. The auxiliary information includes auxiliary patch information. The attribute video / image includes texture video / image.

[0130] The 3D geometry is reconstructed using the decoded geometry image, occupancy map, and additional patch information, and may then undergo a smoothing process. A color point cloud image / picture is reconstructed by assigning color values ​​to the smoothed 3D geometry using a texture image. A renderer renders the reconstructed geometry and color point cloud image / picture. The rendered video / image is displayed by a display unit. The user views all or part of the rendered result on a VR / AR display or a general display.

[0131] The feedback process may include a process of transmitting various feedback information obtainable in the rendering / display process to the transmitting side or to a decoder on the receiving side. The feedback process provides interactivity in the consumption of the point cloud video. According to an embodiment, head orientation information, viewport information indicating the area the user is currently viewing, etc. are transmitted in the feedback process. According to an embodiment, a user may interact with something embodied in a VR / AR / MR / autonomous driving environment, and in this case, information related to the interaction may be transmitted to the transmitting side or service provider in the feedback process. Depending on the embodiment, the feedback process may not be performed.

[0132] The head orientation information is information about the position, angle, movement, etc. of the user's head. Based on this information, information about the area the user is currently looking at in the point cloud video, i.e., viewport information, is calculated.

[0133] Viewport information is information about the area the user is currently looking at in the point cloud video. This allows for gaze analysis to determine how the user consumes the point cloud video, and how long they gaze at which area of ​​the point cloud video. Gaze analysis is performed on the receiving side and transmitted to the transmitting side via a feedback channel. Devices such as VR / AR / MR displays extract the viewport area based on the user's head position / direction, the vertical or horizontal FOV supported by the device, etc.

[0134] According to an embodiment, the above-mentioned feedback information may be not only transmitted to the transmitting side but also consumed by the receiving side. That is, the above-mentioned feedback information may be used to perform decoding and rendering processes on the receiving side. For example, head orientation information and / or viewport information may be used to prioritize decoding and rendering of only the point cloud video for the area currently viewed by the user.

[0135] Here, the viewport or viewport area refers to the area the user is looking at in the point cloud video. The viewpoint is the point the user is looking at in the point cloud video, which means the center of the viewport area. That is, the viewport is an area centered on the viewpoint, and the size and shape of the area are determined by the FOV (Field Of View).

[0136] As mentioned above, this specification relates to point cloud video compression. For example, the methods / embodiments disclosed in this specification may be applied to the Moving Picture Experts Group (MPEG) point cloud compression or point cloud coding (PCC) standard or next generation video / image coding standards.

[0137] In this specification, a picture / frame generally refers to a unit that shows one image in a specific time period.

[0138] A pixel or pel is the smallest unit that makes up a picture (or image). The term "sample" is also used as a counterpart to a pixel. A sample generally refers to a pixel or a pixel value, and may refer to only a pixel / pixel value of a luminance (luma) component, only a pixel / pixel value of a chroma component, or only a pixel / pixel value of a depth component.

[0139] A unit refers to a basic unit of image processing. A unit includes at least one of a specific region of a picture and information about that region. The term unit may be mixed with terms such as block, area, or module. In general, an MxN block includes a set (or array) of samples or transform coefficients consisting of M columns and N rows.

[0140] FIG. 3 shows an example of a point cloud, geometry, and texture image according to an embodiment.

[0141] The point cloud according to the embodiment is input to the V-PCC encoding process in Fig. 4, which will be described later, to generate a geometry image and a texture image. According to the embodiment, the point cloud is used interchangeably with point cloud data.

[0142] In Figure 3, the left side shows a point cloud, where point cloud objects are located in 3D space and are represented by bounding boxes, etc. The middle side of Figure 3 shows a geometry image, and the right side shows a texture image (non-padded).

[0143] Video-based Point Cloud Compression (V-PCC) is a method for compressing 3D point cloud data based on 2D video codecs such as HEVC (Efficiency Video Coding) and VVC (Versatile Video Coding). The V-PCC compression process generates the following data and information:

[0144] Occupancy map: When the points that make up a point cloud are divided into patches and mapped onto a 2D plane, this represents a binary map that indicates with a value of 0 or 1 whether data exists at that location on the 2D plane. The occupancy map represents a 2D array that corresponds to an atlas, and the values ​​of the occupancy map indicate whether each sample location in the atlas corresponds to a 3D point. An atlas is a collection of 2D bounding boxes and associated information located in a rectangular frame that corresponds to a 3D bounding box in the 3D space where volumetric data is rendered.

[0145] An atlas bitstream is a bitstream for one or more atlas frames and associated data that make up an atlas.

[0146] An atlas frame is a 2D rectangular array of atlas samples onto which multiple patches are projected.

[0147] An atlas sample is a position in a rectangular frame onto which a patch associated with an atlas is projected.

[0148] An atlas frame is divided into tiles, which are the units that divide a 2D frame. In other words, tiles are the units that divide the signaling information of the point cloud data called an atlas.

[0149] Patch: A set of points that make up a point cloud. Points that belong to the same patch are adjacent to each other in 3D space and map to the same direction in the six-sided bounding box plane during the mapping process to a 2D image.

[0150] A patch is a unit that divides a tile. A patch is signaling information about the configuration of point cloud data.

[0151] A receiving device according to the embodiment can restore the attribute video data, geometry video data, and occupation video data, which are actual video data having the same presentation time, based on the atlas (tiles, patches).

[0152] Geometry Image: A depth map-style image that represents the geometry of each point in a point cloud, in units of patches. A geometry image consists of a single channel of pixel values. The geometry represents a set of coordinates associated with the point cloud frame.

[0153] Texture image: An image that represents the color information of each point in a point cloud on a patch-by-patch basis. A texture image consists of pixel values ​​for multiple channels (e.g., three channels R, G, B). Textures are included in features. According to the embodiment, textures and / or features are interpreted as the same object and / or as containing features.

[0154] Auxiliary patch info: Indicates the metadata required to reconstruct a point cloud from individual patches. The auxiliary patch info includes information about the patch's position in 2D / 3D space, size, etc.

[0155] Point cloud data, eg, V-PCC components, according to an embodiment, include atlases, occupancy maps, geometry, features, etc.

[0156] An atlas represents a collection of 2D bounding boxes, i.e. patches projected onto a rectangular frame, and a corresponding 3D bounding box in 3D space, representing a subset of a point cloud.

[0157] An attribute is a scalar or vector associated with each point in the point cloud, such as color, reflectance, surface normal, time stamps, or material ID.

[0158] The point cloud data according to the embodiment represents PCC data in a video-based point cloud compression (V-PCC) format. The point cloud data includes multiple components, such as an occupancy map, patches, geometry, and / or textures.

[0159] FIG. 4 illustrates an example of a V-PCC encoding process according to an embodiment.

[0160] Figure 4 shows a V-PCC encoding process for generating and compressing an occupancy map, a geometry image, a texture image, and auxiliary patch information. The V-PCC encoding process in Figure 4 is processed by the point cloud video encoder 10002 in Figure 1. Each component in Figure 4 is implemented by software, hardware, a processor, and / or a combination thereof.

[0161] The patch generation unit 40000 or patch generator receives a point cloud frame (which may be in the form of a bitstream containing point cloud data). The patch generation unit 40000 generates a patch from the point cloud data. The patch generation unit 40000 also generates patch information containing information regarding the generation of the patch.

[0162] The patch packing 40001 or patch packer packs patches for point cloud data, e.g., packs one or more patches, and generates an occupancy map containing information about the patch packing.

[0163] The geometry image generation unit 40002 or geometry image generator generates a geometry image based on the point cloud data, patches, and / or packed patches. A geometry image refers to data including geometry related to the point cloud data.

[0164] The texture image generation unit 40003 or texture image generator generates a texture image based on point cloud data, patches, and / or packed patches, and also generates a texture image based on smoothed geometry generated by smoothing a reconstructed geometry image based on patch information.

[0165] The smoothing 40004 or smoothing unit reduces or removes errors contained in image data. For example, it generates smoothed geometry by gently filtering out parts of a reconstructed geometry image that may induce errors between data based on patch information.

[0166] The auxiliary patch info compression 40005 or auxiliary patch info compression unit compresses auxiliary patch information related to the patch information generated in the patch generation process. The compressed auxiliary patch information is transmitted to the multiplexer, and the geometry image generation unit 40002 also uses the auxiliary patch information.

[0167] Image padding 40006, 40007 or image padding units pad the geometry image and texture image, respectively, i.e., pad data is padded into the geometry image and texture image.

[0168] Group dilation 40008 or group dilation unit adds data to the texture image, similar to an image pad. Additional patch information is inserted into the texture image.

[0169] Video compression 40009, 40010, 40011 or video compression units compress the padded geometry image, padded texture image and / or occupancy map, respectively. The compression encodes geometry information, texture information, occupancy information, etc.

[0170] Entropy compression 40012 or entropy compressor compresses (eg, encodes) the occupancy map based on an entropy scheme.

[0171] According to an embodiment, if the point cloud data is lossless and / or lossy, entropy compression and / or video compression is performed on the occupancy map frames.

[0172] A multiplexer 40013 multiplexes the compressed geometry images, compressed texture images, and compressed occupancy maps into a bitstream.

[0173] The detailed operation of each process in FIG. 4 according to the embodiment will be described below.

[0174] Patch generation: 40000

[0175] The patch generation process is the process of dividing a point cloud into patches, which are the units of mapping, in order to map the point cloud to a 2D image. The patch generation process can be divided into three steps: normal value calculation, segmentation, and patch division as follows.

[0176] The process of calculating the normalized value will be specifically described with reference to FIG.

[0177] FIG. 5 shows an example of a tangent plane and a normal vector of a surface according to an embodiment.

[0178] The surface of FIG. 5 is used in the patch generation process 40000 of the V-PCC encoding process of FIG. 4 as follows.

[0179] Normal calculation related to patch generation

[0180] Each point (e.g., a point) that makes up a point cloud has a unique direction, which is represented by a 3D vector called a normal. Using the neighbors of each point, which are found using a KD tree or similar, the tangent plane and normal vector of each point that makes up the surface of the point cloud, as shown in Figure 5, are found. The search range in the process of finding neighboring points is defined by the user.

[0181] Tangen plane: A plane that passes through a point on a surface and completely contains the tangent to a curve on the surface.

[0182] FIG. 6 illustrates an example of a bounding box for a point cloud according to an embodiment.

[0183] The method / apparatus according to the embodiment, for example, the patch generator, uses bounding boxes in the process of generating patches from point cloud data.

[0184] The bounding box in this embodiment is a unit box that divides point cloud data based on a hexahedron in 3D space.

[0185] Bounding boxes are used in the process of projecting point cloud objects of point cloud data onto the planes of each hexahedron based on a hexahedron in 3D space. The bounding boxes are generated and processed by the point cloud video acquisition unit 10000 and point cloud video encoder 10002 in Figure 1. Furthermore, patch generation 40000, patch packing 40001, geometry image generation 40002, and texture image generation 40003 in the V-PCC encoding process in Figure 2 are performed based on the bounding boxes.

[0186] Segmentation in relation to patch generation

[0187] Segmentation consists of two processes: initial segmentation and refinement segmentation.

[0188] The point cloud video encoder 10002 according to the embodiment projects points onto one side of a bounding box. Specifically, each point constituting a point cloud is projected onto one side of the six faces of a bounding box that surrounds the point cloud, as shown in Figure 6. Initial segmentation is a process of determining one of the planes of the bounding box onto which each point is projected.

[0189] The normal values ​​corresponding to each of the six planes are JPEG0007749058000001.jpg814 is defined as follows:

[0190] (1.0,0.0,0.0), (0.0,1.0,0.0), (0.0,0.0,1.0), (-1.0,0.0,0.0), (0.0,-1.0,0.0), (0.0,0.0,-1.0).

[0191] The normal value of each point obtained from the normal value calculation process described above is as follows: The plane with the largest dot product of JPEG0007749058000002.jpg1031 is determined to be the projection plane of that plane. That is, the plane with a normal oriented most similarly to the normal of the point is determined to be the projection plane of that point.

[0192] JPEG0007749058000003.jpg1341

[0193] The determined plane is identified as one of the index-type values ​​(cluster index) from 0 to 5.

[0194] Refine segmentation is a process of improving the projection plane of each point of the point cloud determined in the above-mentioned initial segmentation process by considering the projection planes of adjacent points. In this process, score normal, which indicates the similarity between the normal of each point considered to determine the projection plane in the above-mentioned initial segmentation process and the normal value of each plane of the bounding box, and score smooth, which indicates the degree of agreement between the projection plane of the current point and the projection plane of the adjacent points, are simultaneously considered.

[0195] Score smooth can be considered by assigning weights to score normal, where the weights are defined by the user. The refinement division is performed iteratively, and the number of iterations is also user-defined.

[0196] Segment patches in relation to patch generation

[0197] Patch division is a process of dividing the entire point cloud into patches, which are sets of adjacent points, based on the projection plane information of each point that makes up the point cloud obtained in the initial / improved division process described above. Patch division consists of the following steps:

[0198] (1) Calculate the neighboring points of each point in the point cloud using a KD tree or similar. The maximum number of neighboring points is defined by the user.

[0199] (2) If the neighboring points project onto the same plane as the current point (have the same cluster index value), extract the current point and its neighboring points into one patch.

[0200] (3) Calculate the geometry value of the extracted patch. The detailed process will be described later.

[0201] (4) Repeat steps (2) and (3) until there are no more points left unextracted.

[0202] Through the process of patch division, the size of each patch and the occupancy map, geometry image, texture image, etc. of each patch are determined.

[0203] FIG. 7 illustrates an example of positioning individual patches of an occupancy map according to an embodiment.

[0204] The point cloud encoder 10002 according to the embodiment can generate patch packing and occupancy maps.

[0205] Patch packing & Occupancy map generation 40001

[0206] This process determines the positions of individual patches within a 2D image in order to map previously divided patches into a single 2D image. An occupancy map is a type of 2D image and is a binary map that indicates the presence or absence of data at that position with a value of 0 or 1. The occupancy map consists of blocks, and its resolution is determined according to the size of the blocks. For example, if the block size is 1*1, it has a resolution in pixels. The size of the blocks (occupancy packing block size) is determined by the user.

[0207] The process of determining the location of an individual patch in the occupancy map is as follows.

[0208] (1) Set all values ​​in the entire occupancy map to 0.

[0209] (2) Position the patch at a point (u, v) in the occupancy map plane whose horizontal coordinate is in the range [0, occupancySizeU-patch.sizeU0) and whose vertical coordinate is in the range [0, occupancySizeV-patch.sizeV0).

[0210] (3) Set the point (x, y) on the patch plane whose horizontal coordinate is in the range [0, patch.sizeU0) and whose vertical coordinate is in the range [0, patch.sizeV0) as the current point.

[0211] (4) For point (x, y), if the (x, y) coordinate value in the patch occupancy map is 1 (data exists at that point in the patch) and the (u+x, v+y) coordinate value in the global occupancy map is 1 (the occupancy map was filled by the previous patch), change the (x, y) position in raster order and repeat steps (3) to (4). Otherwise, perform step (6).

[0212] (5) Change the (u, v) position in the raster order and repeat the steps (3) to (5).

[0213] (6) (u, v) is determined as the location of the patch, and the occupancy map data of the patch is assigned (copied) to the corresponding part of the overall occupancy map.

[0214] (7) Repeat steps (2) to (7) for the next patch.

[0215] OccupancySizeU: Indicates the width of the occupancy map, and its unit is occupancy packing block size.

[0216] Occupancy Size V (occupancySizeV): Indicates the height of the occupancy map, and its unit is the occupancy packing block size.

[0217] Patch size U0 (patch.sizeU0): Indicates the width of the occupancy map, and its unit is the occupancy packing block size.

[0218] Patch size V0 (patch.sizeV0): Indicates the height of the occupancy map, in units of the occupancy packing block size.

[0219] For example, as shown in FIG. 7, there may be a box corresponding to a patch having an in-box patch size corresponding to the occupied packing size block, and a point (x, y) may be located within the box.

[0220] FIG. 8 shows an example of the relationship between the normal, tangent, and bitangent axes according to an embodiment.

[0221] The point cloud video encoder 10002 according to the embodiment can generate a geometry image. The geometry image means image data containing the geometry information of the point cloud. The generation process of the geometry image uses the three axes (normal, tangent, and bitangent) of the patch in Figure 8.

[0222] Geometry image generation 40002

[0223] In this process, the depth values ​​constituting the geometry image of each individual patch are determined, and the entire geometry image is generated based on the patch positions determined in the above-mentioned patch packing process. The process of determining the depth values ​​constituting the geometry image of each individual patch is configured as follows.

[0224] (1) Calculate parameters related to the position and size of individual patches. The parameters include the following information:

[0225] An index indicating the normal axis: The normal is calculated in the patch generation process described above, the tangent axis is the axis perpendicular to the normal that coincides with the horizontal (u) axis of the patch image, and the bitangent axis is the axis perpendicular to the normal that coincides with the vertical (v) axis of the patch image. The three axes are shown in the figure below.

[0226] FIG. 9 illustrates an example of minimum and maximum mode configurations of the projection mode according to an embodiment.

[0227] The point cloud video encoder 10002 according to the embodiment performs patch-based projection to generate a geometry image, and the projection modes according to the embodiment include a minimum mode and a maximum mode.

[0228] 3D space coordinates of a patch: Calculated by the minimum size of the bounding box that surrounds the patch. For example, the 3D space coordinates of a patch include the minimum value of the patch tangent direction (patch 3D shift tangent axis), the minimum value of the patch bitangent direction (patch 3D shift bitangent axis), the minimum value of the patch normal direction (patch 3D shift normal axis), etc.

[0229] Patch 2D size: indicates the horizontal and vertical size of the patch when it is packed into a 2D image. The horizontal size (patch 2D size u) is the difference between the maximum and minimum values ​​of the tangent direction of the bounding box, and the vertical size (patch 2D size v) is the difference between the maximum and minimum values ​​of the bitangent direction of the bounding box.

[0230] (2) Determine the projection mode of the patch. The projection mode can be either the minimum mode or the maximum mode. The geometry information of the patch is represented by depth values, and when each point of the patch is projected in the normal direction of the patch, two layer images are generated: one consisting of the maximum depth value and the other consisting of the minimum depth value.

[0231] When generating two layer images d0 and d1 in minimum mode, the minimum depth is set to d0 as shown in Figure 9, and the maximum depth within the surface thickness from the minimum depth is set to d1.

[0232] For example, if the point cloud is located in 2D as shown in the figure, there may be multiple patches containing multiple points. As shown in the figure, points shown with the same shading belong to the same patch. The process of projecting patches of points shown with blank spaces is shown.

[0233] When projecting a point indicated by a blank space to the left / right side, the depth is calculated by increasing the depth by one from the left side, such as 0, 1, 2, ..6, 7, 8, 9, and then writing the numbers for calculating the depth of the point on the right side.

[0234] The projection mode can be defined by the user, and the same method can be applied to all point clouds, or different methods can be applied to each frame or patch. When different projection modes are applied to each frame or patch, a projection mode that can improve compression efficiency or minimize missed points is adaptively selected.

[0235] (3) Calculate the depth value of each point.

[0236] In minimum mode, the d0 image is constructed with depth0, which is the minimum value of the normal axis at each point minus the minimum value of the normal direction of the patch (patch 3D shift normal axis) calculated in process (1). If there is another depth value within the range of depth0 and the surface thickness at the same position, this value is set to depth1. If not, the value of depth0 is also assigned to depth1. The d1 image is constructed with the value of Depth1.

[0237] For example, in determining the depth of the point d0, the minimum value is calculated (4 2 4 4 0 6 0 0 9 9 0 8 0). In determining the depth of the point d1, the largest value of two or more points is calculated, or if there is only one point, that value is calculated (4 4 4 4 6 6 6 8 9 9 8 8 9). In addition, in the process of encoding and reconstructing the points of the patch, some points are lost (for example, 8 points are lost in the figure).

[0238] In Max mode, the d0 image is constructed with depth0, which is the maximum value of the normal axis at each point minus the minimum value of the normal direction of the patch (patch 3D shift normal axis) calculated in process (1). If there is another depth value within the range of depth0 and the surface thickness at the same position, this value is set to depth1. If not, the value of depth0 is also assigned to depth1. The d1 image is constructed with the value of Depth1.

[0239] For example, when determining the depth of the point d0, the maximum value is calculated (4 4 4 4 6 6 6 8 9 9 8 8 9). When determining the depth of the point d1, the smaller value of two or more points is calculated, or if there is only one point, that value is calculated (4 2 4 4 5 6 0 6 9 9 0 8 0). Also, in the process of encoding and reconstructing the points of the patch, some points are lost (for example, six points are lost in the figure).

[0240] The overall geometry image can be generated by placing the geometry images of the individual patches generated from the above process into the overall geometry image using the position information of the individual patches generated through the patch packing process described above.

[0241] The d1 layer of the generated overall geometry image can be encoded using various methods. The first method is to encode the depth values ​​of the previously generated d1 image as they are (absolute d1 encoding method). The second method is to encode the difference between the depth values ​​of the previously generated d1 image and the depth values ​​of the d0 image (differential encoding method).

[0242] In this encoding method using the depth values ​​of two layers d0 and d1, if there is another point between the two depths, the geometry information of that point is lost in the encoding process, so for lossless coding, an Enhanced-Delta-Depth (EDD) code may be used.

[0243] The EDD code will be specifically described with reference to FIG.

[0244] FIG. 10 shows an example of an EDD code according to an embodiment.

[0245] The point cloud video encoder 10002 and / or some / all of the V-PCC encoding process (e.g., video compression 40009) can encode the geometry information of the points based on the EOD code.

[0246] The EDD code is a method of binary encoding the positions of all points within the surface thickness range, including d1, as shown in the figure. As an example, for the points in the second column from the left in the figure, there are points in the first and fourth positions above D0, and the second and third positions are empty, so they are represented by an EDD code of 0b1001 (=9). If the EDD code is encoded and transmitted along with D0, the receiving end can recover the geometry information of all points without any loss.

[0247] For example, if a point exists on the reference point, it is 1, and if no point exists, it is 0, and the code is expressed based on four bits.

[0248] Smoothing40004

[0249] Smoothing is the process of removing discontinuities that may occur at the point boundary due to image quality degradation resulting from the compression process, and is performed in the point cloud video encoder or the smoothing section.

[0250] (1) Reconstruct the point cloud from the geometry image. This process can be considered the reverse of the geometry image construction described above. For example, reconstruction is the reverse process of encoding.

[0251] (2) Calculate the neighboring points of each point that makes up the regenerated point cloud using a KD tree or similar.

[0252] (3) For each point, determine whether the point is located on the boundary surface of the patch. For example, if there is an adjacent point with a different projection plane (cluster index) from the current point, the point can be determined to be located on the boundary surface of the patch.

[0253] (4) If a patch boundary exists, move the point to the center of gravity of the neighboring points (located at the average x, y, z coordinates of the neighboring points), i.e., change the geometry value. If no boundary exists, keep the previous geometry value.

[0254] FIG. 11 shows an example of recoloring using color values ​​of neighboring points according to an embodiment.

[0255] The point cloud video encoder or texture image generator 40003 according to the embodiment can generate a texture image based on decolorization.

[0256] Texture image generation 40003

[0257] The texture image generation process is similar to the geometry image generation process described above, in that it generates texture images for individual patches and places them at predetermined positions to generate an overall texture image. However, in the process of generating texture images for individual patches, an image is generated that has the color values ​​(e.g., R, G, B) of the points that make up the point cloud corresponding to that position, instead of the depth values ​​used for geometry generation.

[0258] The process of calculating the color value of each point that makes up the point cloud uses the geometry that has undergone the smoothing process described above. Because the smoothed point cloud may contain points that have shifted in position compared to the original point cloud, a color restoration process is required to find a color that matches the changed position. Color restoration is performed using the color values ​​of neighboring points. For example, as shown in the figure, a new color value can be calculated by taking into account the color values ​​of the nearest neighbor and the color values ​​of the neighboring points.

[0259] For example, referring to the figure, the recoloring calculates an appropriate color value for the modified position based on the average of the characteristic information of the closest original points to the point and / or the average of the characteristic information of the closest original positions to the point.

[0260] Texture images are also generated in two layers, t0 / t1, just like geometry images, which are generated in two layers, d0 / d1.

[0261] Auxiliary patch info compression 40005

[0262] The point cloud video encoder or additional patch information compressor according to the embodiment can compress additional patch information (additional information about the point cloud).

[0263] The additional patch information compression unit compresses the additional patch information generated in the above-mentioned patch generation, patch packing, geometry generation processes, etc. The additional patch information includes the following parameters:

[0264] Cluster index that identifies the projection plane (normal)

[0265] 3D spatial position of the patch: minimum value of the patch tangent direction (patch 3D shift tangent axis), minimum value of the patch bitangent direction (patch 3D shift bitangent axis), minimum value of the patch normal direction (patch 3D shift normal axis)

[0266] 2D spatial position and size of the patch: horizontal size (patch 2D size u), vertical size (patch 2D size v), horizontal minimum value (patch 2D shift u), vertical minimum value (patch 2D shift u)

[0267] Mapping information for each block and patch: candidate index (when placing patches in order based on the 2D spatial position and size information of the patches mentioned above, multiple patches may be mapped to one block. In this case, the mapped patches form a candidate list, and the index indicates which patch's data is present in the block), local patch index (an index indicating one of the total patches present in the frame). Table 1 shows pseudocode showing the process of matching blocks and patches using the candidate list and local patch index.

[0268] The maximum number of candidate lists is user defined.

[0269] Table 1-1 Pseudo code for block and patch mapping

[0270] for(i=0;i <BlockCount;i++){

[0271] if(candidatePatches[i].size()==1){

[0272] blockToPatch[i]=candidatePatches[i][0]

[0273] } else {

[0274] candidate_index

[0275] if(candidate_index==max_candidate_count) {

[0276] blockToPatch[i]=local_patch_index

[0277] } else {

[0278] blockToPatch[i]=candidatePatches[i][candidate_index]

[0279] }

[0280] }

[0281] }

[0282] FIG. 12 illustrates an example of push-pull background filling according to an embodiment.

[0283] Image padding and group dilation 40006, 40007, 40008

[0284] The image padder according to the embodiment can fill spaces other than the patch area with meaningless additional data based on a push-pull background filling method.

[0285] Image padding is a process of filling spaces outside the patch area with meaningless data to improve compression efficiency. For image padding, pixel values ​​of the corresponding columns or rows on the boundary side of the patch are copied to fill the empty spaces. Alternatively, a push-pull background filling method may be used, in which the resolution of the unpadded image is gradually reduced and then increased again, filling the empty spaces with pixel values ​​from a lower-resolution image, as shown in the figure.

[0286] Group dilation is a method for filling empty spaces in a geometry or texture image consisting of two layers, d0 / d1 and t0 / t1. It is a process of filling the empty space values ​​of the two layers calculated by the image padding described above with the average value of the values ​​for the same position in the two layers.

[0287] FIG. 13 shows an example of a possible traversal order for a block of size 4*4 according to an embodiment.

[0288] Occupancy map compression 40012, 40011

[0289] In the embodiment, the occupancy map compression compresses the occupancy map. Specifically, there are two methods: video compression for lossy compression and entropy compression for lossless compression. Video compression will be described later.

[0290] The entropy compression process is as follows:

[0291] (1) For each block that makes up the occupancy map, if all blocks are filled, code 1, and repeat the same process for the next block. Otherwise, code 0, and repeat steps (2) to (5).

[0292] (2) Determine the best traversal order for run-length coding of the filled pixels of the block. The figure shows four possible traversal orders for a 4x4 size block as an example.

[0293] FIG. 14 shows an example of the best traversal order according to the embodiment.

[0294] As described above, the entropy compressor according to the embodiment codes (encodes) blocks based on the traversal order scheme as shown in the figure.

[0295] For example, among the possible traversal orders, the best traversal order with the smallest number of runs is selected and its index is encoded. For example, when selecting the third traversal order in FIG. 13, the number of runs can be minimized to 2, and this is selected as the best traversal order.

[0296] At this time, the number of runs is coded. In the example of Fig. 14, since there are two runs, 2 is coded.

[0297] (4) Encode the occupancy of the first run. In the example of Figure 14, the first run corresponds to an unfilled pixel, so 0 is encoded.

[0298] (5) The length of each run (the number of runs) is coded. In the example of Figure 14, the lengths of the first and second runs, 6 and 10, are coded sequentially.

[0299] Video compression 40009, 40010, 40011

[0300] The video compressor according to the embodiment encodes sequences of geometry images, texture images, occupancy map images, etc. generated by the above-described processes using a 2D video codec such as HEVC or VVC.

[0301] FIG. 15 illustrates an example of a 2D video / image encoder according to an embodiment.

[0302] The figure shows a schematic block diagram of a 2D video / image encoder 15000 in which the video / image signals are encoded, in an embodiment to which the above-mentioned video compression units 40009, 40010, and 40011 are applied. The 2D video / image encoder 15000 may be included in the above-mentioned point cloud video encoder 10002, or may consist of internal / external components. Each component in Figure 15 corresponds to software, hardware, a processor, and / or a combination thereof.

[0303] Here, the input images include the above-mentioned geometry images, texture images (feature images), occupancy map images, etc. The output bitstream of the point cloud video encoder (i.e., point cloud video / image bitstream) includes an output bitstream for each input image (geometry image, texture image (feature image), occupancy map image, etc.).

[0304] The inter prediction unit 15090 and the intra prediction unit 15100 are collectively referred to as a prediction unit. That is, the prediction unit includes the inter prediction unit 15090 and the intra prediction unit 15100. The transform unit 15030, the quantization unit 15040, the inverse quantization unit 15050, and the inverse transform unit 15060 are collectively referred to as a residual processing unit. The residual processing unit further includes a subtraction unit 15020. Depending on the embodiment, the above-mentioned image division unit 15010, the subtraction unit 15020, the transform unit 15030, the quantization unit 15040, the inverse quantization unit 15050, the inverse transform unit 15060, the addition unit 155, the filtering unit 15070, the inter prediction unit 15090, the intra prediction unit 15100, and the entropy coding unit 15110 may be configured as a single hardware component (e.g., an encoder or a processor). The memory 15080 also includes a DPB (decoded picture buffer) and is configured as a digital storage medium.

[0305] The image division unit 15010 divides an input image (or picture, frame) input to the encoding device 15000 into one or more processing units. For example, the processing units are also called coding units (CUs). In this case, the coding units are recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) according to a QTBT (quad-tree binary-tree) structure. For example, one coding unit is divided into multiple coding units of a deeper depth based on a quad-tree structure and / or a binary-tree structure. In this case, for example, a quad-tree may be applied first, followed by a binary-tree, or a binary-tree may be applied first. The coding procedure according to the present invention may be performed based on the final coding unit that is not further divided. In this case, the largest coding unit may be used as the final coding unit based on coding efficiency according to video characteristics, or the coding unit may be recursively divided into coding units of lower depth as needed, and a coding unit of an optimal size may be used as the final coding unit. Here, the coding procedure includes procedures such as prediction, transformation, and restoration, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, each of the prediction unit and the transform unit is divided or partitioned from the final coding unit. The prediction unit is a unit of sample prediction, and the transform unit is a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.

[0306] The term "unit" is sometimes used interchangeably with terms such as "block," "area," or "module." In general, an MxN block refers to a set of samples or transform coefficients consisting of M columns and N rows. A sample generally refers to a pixel or pixel value, and may refer to only a pixel / pixel value of the luma component, or only a pixel / pixel value of the chroma component. A sample is used as a term corresponding to one picture (or image) pixel or pel.

[0307] The encoding device 15000 subtracts a prediction signal (predicted block, prediction sample array) output from the inter prediction unit 15090 or the intra prediction unit 15100 from an input video signal (original block, original sample array) to generate a residual signal (residual signal, residual block, residual sample array), and the generated residual signal is transmitted to the transform unit 15030. In this case, as shown in the figure, a unit in the encoding device 15000 that subtracts a prediction signal (predicted block, prediction sample array) from an input video signal (original block, original sample array) is called a subtraction unit 15020. The prediction unit predicts a block to be processed (hereinafter referred to as a current block) and generates a predicted block including prediction samples for the current block. The prediction unit determines whether to apply intra prediction or inter prediction on a current block or CU basis. The prediction unit generates various information related to prediction, such as prediction mode information, as will be described below for each prediction mode, and transmits the information to the entropy encoding unit 15110. The prediction information is coded by the entropy coding unit 15110 and output in the form of a bitstream.

[0308] The intra prediction unit 15100 predicts the current block by referring to samples in the current picture. The referenced samples are located either neighboring or distant from the current block depending on the prediction mode. Prediction modes in intra prediction include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes include, for example, DC mode and planar mode. The directional modes include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the accuracy of the prediction direction. However, this is merely an example, and more or less directional prediction modes may be used depending on the settings. The prediction mode applied to the neighboring blocks of the intra prediction unit 15100 may be used to determine the prediction mode applied to the current block.

[0309] The inter prediction unit 15090 derives a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector in a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information is predicted in units of blocks, sub-blocks, or samples based on the correlation between the motion information of neighboring blocks and the current block. The motion information includes a motion vector and a reference picture index. The motion information also includes information on the inter prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring blocks may be the same or different. The temporal neighboring block is called a collocated reference block, collocated CU (colCU), etc., and the reference picture including the temporal neighboring block is also called a collocated picture (colPic). For example, the inter predictor 15090 constructs a motion information candidate list based on neighboring blocks and generates information indicating which candidate is used to derive a motion vector and / or a reference picture index for the current block. Inter prediction is performed based on various prediction modes, such as skip mode and merge mode, in which the inter predictor uses motion information of neighboring blocks as motion information for the current block. Unlike merge mode, skip mode may not transmit a residual signal. In motion vector prediction (MVP) mode, the motion vector of the neighboring block is used as a motion vector predictor, and the motion vector of the current block is indicated by signaling a motion vector difference.

[0310] The prediction signal generated by the inter prediction unit 15090 or the intra prediction unit 15100 is used to generate a reconstructed signal or to generate a residual signal.

[0311] The transform unit 15030 applies a transform method to the residual signal to generate transform coefficients. For example, the transform method may include at least one of the following: Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loeve Transform (KLT), Graph-Based Transform (GBT), or Conditionally Non-Linear Transform (CNT). Here, GBT refers to a transform obtained from a graph representing inter-pixel relationship information. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. The transform process may be applied to square pixel blocks of the same size or non-square variable-size blocks.

[0312] The quantizer 15040 quantizes the transform coefficients and transmits them to the entropy encoder 15110, which then encodes the quantized signal (information about the quantized transform coefficients) and outputs it as a bitstream. The information about the quantized transform coefficients is called residual information. The quantizer 15040 may rearrange the block-type quantized transform coefficients into a one-dimensional vector based on the coefficient scan order and generate information about the quantized transform coefficients based on the one-dimensional vector-type quantized transform coefficients. The entropy encoder 15110 performs various encoding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). In addition to the quantized transform coefficients, the entropy encoder 15110 may also encode information required for video / image reconstruction (e.g., values ​​of syntax elements) together with or separately from the quantized transform coefficients. The encoded information (e.g., encoded video / image information) is transmitted or stored in the form of a bitstream in units of network abstraction layer (NAL) units. The bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network includes a broadcast network and / or a communication network, and the digital storage medium includes various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) that transmits and / or a memory (not shown) that stores the signal output from the entropy encoder 15110 may be configured as an internal / external element of the encoding device 15000, or the transmitter may be included in the entropy encoder 15110.

[0313] The quantized transform coefficients output from the quantization unit 15040 are used to generate a prediction signal. For example, the inverse quantization unit 15040 and the inverse transform unit 15060 apply inverse quantization and inverse transform to the quantized transform coefficients to reconstruct a residual signal (residual block or residual samples). The adder 15200 generates a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter prediction unit 15090 or the intra prediction unit 15100. When there is no residual for the current block, such as when skip mode is applied, the predicted block is used as the reconstructed block. The adder 155 is referred to as a reconstruction unit or a reconstructed block generator. The generated reconstructed signal may be used for intra prediction of the next current block in the current picture, or may be used for inter prediction of the next picture after filtering, as described below.

[0314] The filtering unit 15070 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 15070 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and store the modified reconstructed picture in the memory 15080, specifically in the DPB of the memory 15080. Examples of various filtering methods include deblocking filtering, sample adaptive offset, an adaptive loop filter, and a bilateral filter. The filtering unit 15070 generates various information related to filtering, such as for each filtering method described below, and transmits the information to the entropy coding unit 15110. The entropy coding unit 15110 encodes the filtering information and outputs it in the form of a bitstream.

[0315] The modified reconstructed picture sent to the memory 15080 is used as a reference picture in the inter prediction unit 15090. This allows the encoding device to avoid prediction mismatches in the encoding device 15000 and the decoding device when inter prediction is applied, and also improves encoding efficiency.

[0316] The DPB of the memory 15080 stores the modified reconstructed picture for use as a reference picture in the inter predictor 15090. The memory 15080 stores motion information of a block from which motion information in the current picture is derived (or coded) and / or motion information of a block in an already reconstructed picture. The stored motion information is transmitted to the inter predictor 15090 to be used as motion information of a spatially adjacent block or a temporally adjacent block. The memory 15080 stores reconstructed samples of blocks reconstructed in the current picture and transmits them to the intra predictor 15100.

[0317] At least one of the above-described prediction, transformation, and quantization steps may be omitted. For example, for a block to which PCM (pulse code modulation) is applied, the prediction, transformation, and quantization steps may be omitted, and the original sample values ​​may be directly coded and output to a bitstream.

[0318] FIG. 16 illustrates an example of a V-PCC decoding process according to an embodiment.

[0319] The V-PCC decoding process or V-PCC decoder is the inverse process of the V-PCC encoding process (or encoder) in Figure 4. Each component in Figure 16 corresponds to software, hardware, a processor, and / or a combination thereof.

[0320] A demultiplexer 16000 demultiplexes the compressed bitstream and outputs a compressed texture image, a compressed geometry image, a compressed occupancy map image, and compressed additional patch information, respectively.

[0321] The video decompression 16001, 16002 or video decompression unit decompresses the compressed texture images and the compressed geometry images, respectively.

[0322] Occupancy map decompression 16003 or occupancy map decompressor decompresses the compressed occupancy map image.

[0323] The auxiliary patch information decompression 16004 or auxiliary patch information restoration unit restores the compressed auxiliary patch information.

[0324] The geometry reconstruction 16005 or geometry reconstruction unit reconstructs geometry information based on the reconstructed geometry image, the reconstructed occupancy map, and / or the reconstructed additional patch information, e.g., reconstructing geometry that has changed during the encoding process.

[0325] Smoothing 16006 or the smoothing unit applies smoothing to the reconstructed geometry, for example, smoothing filtering.

[0326] Texture reconstruction 16007 or texture reconstruction unit reconstructs texture from the restored texture image and / or smoothed geometry.

[0327] Color smoothing 16008 or color smoothing unit smooths the color values ​​from the reconstructed texture, e.g., applies smoothing filtering.

[0328] As a result, reconstructed point cloud data is generated.

[0329] The figure shows the V-PCC decoding process for decoding the compressed occupancy map, geometry image, texture image, and additional patch information to reconstruct a point cloud. The operation of each process according to the embodiment is as follows.

[0330] Video decompression 16001, 16002

[0331] This is the inverse process of the video compression described above, where the compressed bitstreams of geometry, texture, and occupancy map images generated in the above process are decoded using a 2D video codec such as HEVC or VVC.

[0332] FIG. 17 illustrates an example of a 2D video / image decoder according to an embodiment.

[0333] The 2D video / image decoder performs the reverse process of the 2D video / image encoder in FIG.

[0334] The 2D video / image decoder of Figure 17 is an embodiment of the video restoration or video restoration unit of Figure 16, and shows a schematic block diagram of a 2D video / image decoder 17000 in which video / image signals are decoded. The 2D video / image decoder 17000 may be included in the point cloud video decoder of Figure 1, or may be composed of internal or external components. Each component of Figure 17 corresponds to software, hardware, a processor, and / or a combination thereof.

[0335] Here, the input bitstream includes bitstreams for the above-mentioned geometry image, texture image (feature image), occupancy map image, etc. The restored image (or output image, decoded image) indicates the restored image for the above-mentioned geometry image, texture image (feature image), and occupancy map image.

[0336] Referring to the figure, the inter prediction unit 17070 and the intra prediction unit 17080 are collectively referred to as the prediction unit. That is, the prediction unit includes the inter prediction unit 180 and the intra prediction unit 185. The inverse quantization unit 17020 and the inverse transform unit 17030 are collectively referred to as the residual processing unit. That is, the residual processing unit includes the inverse quantization unit 17020 and the inverse transform unit 17030. Depending on the embodiment, the entropy decoding unit 17010, the inverse quantization unit 17020, the inverse transform unit 17030, the addition unit 17040, the filtering unit 17050, the inter prediction unit 17070, and the intra prediction unit 17080 may be configured as a single hardware component (e.g., a decoder or a processor). In addition, the memory 17060 may include a decoded picture buffer (DPB) or may be configured as a digital storage medium.

[0337] When a bitstream containing video / image information is input, the decoding device 17000 reconstructs the image corresponding to the process by which the video / image information was processed in the encoding device of Fig. 0.2-1. For example, the decoding device 17000 performs decoding using the processing unit applied in the encoding device. Therefore, the processing unit for decoding is, for example, a coding unit, and the coding unit is divided from the coding tree unit or the maximum coding unit according to a quad-tree structure and / or a binary-tree structure. Furthermore, the reconstructed video signal decoded and output by the decoding device 17000 is reproduced by a playback device.

[0338] The decoding device 17000 receives a signal output from an encoding device in the form of a bitstream, and the received signal is decoded by the entropy decoding unit 17010. For example, the entropy decoding unit 17010 parses the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). For example, the entropy decoding unit 17010 decodes information in the bitstream based on a coding method such as Exponential Golomb coding, CAVLC, or CABAC, and outputs values ​​of syntax elements necessary for image restoration and quantized values ​​of transform coefficients related to the residual. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using information on the syntax element to be decoded and decoded information on neighboring and current blocks or information on symbols / bins decoded in a previous step, predicts the occurrence probability of the bin according to the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values ​​of each syntax element. In this case, the CABAC entropy decoding method determines a context model and then updates the context model using information about the decoded symbol / bin for the context model of the next symbol / bin. Prediction-related information from the information decoded by the entropy decoding unit 17010 is provided to prediction units (inter prediction unit 17070 and intra prediction unit 17080), and residual values ​​entropy-decoded by the entropy decoding unit 17010, i.e., quantized transform coefficients and related parameter information, are input to the inverse quantization unit 17020. In addition, filtering-related information from the information decoded by the entropy decoding unit 17010 is provided to the filtering unit 17050. Meanwhile, a receiving unit (not shown) that receives a signal output from the encoding device may be further configured as an internal / external element of the decoding device 17000, and the receiving unit may be a component of the entropy decoding unit 17010.

[0339] The inverse quantization unit 17020 quantizes the quantized transform coefficients and outputs the transform coefficients. The inverse quantization unit 17020 rearranges the quantized transform coefficients into a two-dimensional block. In this case, the rearrangement is performed based on the coefficient scanning order performed by the encoding device. The inverse quantization unit 17020 performs inverse quantization on the quantized transform coefficients using a quantization parameter (e.g., quantization step size information) to obtain transform coefficients.

[0340] The inverse transform unit 17030 inversely transforms the transform coefficients to obtain a residual signal (residual block, residual sample array).

[0341] The prediction unit performs prediction on the current block and generates a predicted block including prediction samples for the current block. The prediction unit determines whether intra prediction or inter prediction is applied to the current block based on prediction information output from the entropy decoding unit 17010, and determines a specific intra / inter prediction mode.

[0342] The intra prediction unit 265 predicts the current block by referring to samples in the current picture. The referenced samples may be located neighboring or distant from the current block depending on the prediction mode. Prediction modes in intra prediction include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit 265 determines the prediction mode to be applied to the current block using the prediction modes applied to neighboring blocks.

[0343] The inter predictor 17070 derives a predicted block for the current block based on a reference block (reference sample array) identified by a motion vector in a reference picture. In this case, to reduce the amount of motion information transmitted in inter prediction mode, the motion information is predicted in units of blocks, sub-blocks, or samples based on the correlation between the motion information of neighboring blocks and the current block. The motion information includes a motion vector and a reference picture index. The motion information also includes information on an inter prediction method (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. For example, the inter predictor 17070 constructs a motion information candidate list based on the neighboring blocks and derives a motion vector and / or a reference picture index for the current block based on received candidate selection information. Inter prediction is performed based on various prediction modes, and the prediction information includes information indicating the inter prediction mode for the current block.

[0344] The adder 17040 generates a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the inter prediction unit 17070 or the intra prediction unit 265. When there is no residual for the current block, such as when skip mode is applied, the predicted block is used as the reconstructed block.

[0345] The adder 17040 is referred to as a reconstruction unit or a reconstruction block generator. The generated reconstruction signal may be used for intra prediction of the next block to be processed in the current picture, and may also be used for inter prediction of the next picture after filtering, as described below.

[0346] The filtering unit 17050 applies filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 17050 applies various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and transmits the modified reconstructed picture to the memory 17060, specifically to the DPB of the memory 17060. Examples of various filtering methods include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0347] The (modified) reconstructed picture stored in the DPB of the memory 17060 is used as a reference picture in the inter predictor 17070. The memory 17060 stores motion information of a block from which motion information in the current picture is derived (or decoded) and / or motion information of a block in an already reconstructed picture. The stored motion information is transmitted to the inter predictor 17070 to be used as motion information of a spatially neighboring block or a temporally neighboring block. The memory 17060 stores reconstructed samples of reconstructed blocks in the current picture and transmits them to the intra predictor 17080.

[0348] In this specification, the embodiments described for the filtering unit 160, inter prediction unit 180, and intra prediction unit 185 of the encoding device 100 can be applied in the same or corresponding manner to the filtering unit 17050, inter prediction unit 17070, and intra prediction unit 17080 of the decoding device 17000.

[0349] Meanwhile, at least one of the above-mentioned steps of prediction, inverse transform, and inverse quantization may be omitted. For example, for a block to which PCM (pulse code modulation) is applied, the steps of prediction, inverse transform, and inverse quantization may be omitted, and the values ​​of the decoded samples may be used directly as samples of the restored image.

[0350] Occupancy map decompression 16003

[0351] This is the reverse process of the occupancy map compression described above, and is the process of decoding the compressed occupancy map bitstream to restore the occupancy map.

[0352] Auxiliary patch info decompression 16004

[0353] This is the reverse process of the additional patch information compression described above, and is a process of decoding the compressed additional patch information bit stream to restore the additional patch information.

[0354] Geometry reconstruction 16005

[0355] This is the reverse process of generating the geometry image described above. First, a patch is extracted from the geometry image using the restored occupancy map and the 2D position / size information of the patch contained in the additional patch information, as well as the block-to-patch mapping information. Then, the point cloud is restored in 3D space using the extracted geometry image of the patch and the 3D position information of the patch contained in the additional patch information. The geometry value corresponding to an arbitrary point (u, v) within a patch is called g(u, v). If the coordinate values ​​of the normal, tangent, and bitangent axes of the patch's position in 3D space are (δ0, s0, r0), the coordinate values ​​of the normal, tangent, and bitangent axes of the position in 3D space mapped to the point (u, v), δ(u, v), s(u, v), and r(u, v), are expressed as follows:

[0356] d(u,v)=d0+g(u,v)

[0357] s(u,v)=s0+u

[0358] r(u,v)=r0+v

[0359] Smoothing16006

[0360] This is similar to the smoothing in the encoding process described above, and is a process for removing discontinuities that may arise from patch boundaries due to image quality degradation that occurs in the compression process.

[0361] Texture reconstruction 16007

[0362] This is the process of restoring a color point cloud by assigning a color value to each point that makes up the smoothed point cloud. Using the mapping information of the geometry image and point cloud reconstructed in the diorama reconstruction process described above, this is done by assigning the color value corresponding to the texture image pixel at the same position as the geometry image in 2D space to the point cloud point that corresponds to the same position in 3D space.

[0363] Color smoothing 16008

[0364] Similar to the geometry smoothing process described above, color smoothing is a process for removing discontinuities in color values ​​that may arise from patch boundaries due to image degradation resulting from the compression process.

[0365] (1) Calculate the neighboring points of each point constituting the reconstructed color point cloud using a KD tree, etc. The neighboring point information calculated in the above-mentioned geometry smoothing process may be used as is.

[0366] (2) For each point, determine whether the point is located on the patch boundary surface. The boundary surface information calculated in the above-mentioned geometry smoothing process may be used as is.

[0367] (3) The distribution of color values ​​of adjacent points of a point on the boundary surface is examined to determine whether to perform smoothing. For example, if the entropy of the brightness value is below the boundary value (threshold local entry) (if there are many similar brightness values), it is determined that it is not an edge and smoothing is performed. One method of smoothing is to replace the color value of that point with the average value of the adjacent points.

[0368] FIG. 18 shows an example of the flow of operations of the transmitting device according to the embodiment.

[0369] The transmitting device according to the embodiment may correspond to or perform some / all of the operations of the transmitting device in Fig. 1, the encoding process in Fig. 4, and the 2D video / image encoder in Fig. 15. Each component of the transmitting device corresponds to software, hardware, a processor, and / or a combination thereof.

[0370] The operation of the transmitting end for compressing and transmitting point cloud data using V-PCC is as shown in the figure.

[0371] The point cloud data transmission device according to the embodiment is referred to as a transmission device or the like.

[0372] The patch generator 18000 first generates patches for 2D image mapping of a point cloud. Additional patch information is generated as a result of the patch generation, and the information is used for geometry image generation, texture image generation, smoothing, or geometry restoration process for smoothing.

[0373] The patch packing unit 18001 performs the patch packing process, which maps the patches generated by the patch generation unit into a 2D image. An occupancy map is generated as a result of the patch packing, and the occupancy map is used in the geometry image generation, texture image generation, and geometry restoration process for smoothing.

[0374] The geometry image generator 18002 generates a geometry image using the additional patch information and the occupancy map, and the generated geometry image is encoded into one bitstream by video encoding.

[0375] The encoding pre-processing unit 18003 includes image padding. The generated geometry image or the geometry image regenerated by decoding the encoded geometry bitstream is used for 3D geometry decoding, and then a smoothing process is performed.

[0376] The texture image generator 18004 generates a texture image using the smoothed 3D geometry, point cloud data, additional patch information, and occupancy map. The generated texture image is encoded into a video bitstream.

[0377] The metadata encoding unit 18005 encodes the additional patch information into one metadata bitstream.

[0378] The video encoder 18006 encodes the occupancy map into a video bitstream.

[0379] The multiplexing unit 18007 multiplexes the generated geometry, texture image, and occupancy map video bitstreams and the additional patch information metadata bitstream into one bitstream.

[0380] The transmitter 18008 transmits the bitstream to the receiving end, or the generated video bitstream of geometry, texture image, and occupancy map and the additional patch information metadata bitstream are encapsulated in a segment to generate a file with one or more track data, and then transmitted from the transmitter to the receiving end.

[0381] FIG. 19 shows an example of the flow of operations of the receiving device according to the embodiment.

[0382] The receiving device according to the embodiment corresponds to or performs some / all of the operations of the receiving device in Fig. 1, the decoding process in Fig. 16, and the 2D video / image encoder in Fig. 17. Each component of the receiving device corresponds to software, hardware, a processor, and / or a combination thereof.

[0383] The operation process of the receiving end for receiving and restoring point cloud data using V-PCC is as shown in the figure. The operation of the V-PCC receiving end is the reverse process of the operation of the V-PCC transmitting end in FIG.

[0384] The point cloud data receiving device according to the embodiment is referred to as a receiving device or the like.

[0385] The received point cloud bitstream is demultiplexed by the demultiplexer 19000 into a video bitstream of compressed geometry images, texture images, and occupancy maps after file / segment de-encapsulation, and an additional patch information metadata bitstream. The video decoder 19001 and metadata decoder 19002 decode the demultiplexed video bitstream and metadata bitstream. The geometry restorer 19003 restores 3D geometry using the decoded geometry image, occupancy map, and additional patch information, and then a smoothing process is performed by the smoother 19004. The texture restorer 19005 restores a color point cloud image / picture by assigning color values ​​to the smoothed 3D geometry using a texture image. Then, a color smoothing process is further performed to improve objective / subjective visual quality, and the resulting modified point cloud image / picture is presented to the user after the rendering process. Note that the color smoothing process may be omitted in some cases.

[0386] FIG. 20 illustrates an example architecture for V-PCC-based point cloud data storage and streaming according to an embodiment.

[0387] Part or all of the system in Figure 20 includes part or all of the transmitting / receiving device in Figure 1, the encoding process in Figure 4, the 2D video / image encoder in Figure 15, the decoding process in Figure 16, the transmitting device in Figure 18, and / or the receiving device in Figure 19. Each component in the figure corresponds to software, hardware, a processor, or a combination thereof.

[0388] 20 to 22 show a structure in which a system is further connected to a transceiver according to an embodiment. The transceiver according to the embodiment and the system are collectively referred to as a transceiver according to an embodiment.

[0389] The apparatus according to the embodiment shown in FIGS. 20 to 22 generates a container conforming to a data format for transmitting a bitstream including encoded point cloud data, which is a transmitting apparatus corresponding to FIG. 18, etc.

[0390] The V-PCC system according to the embodiment generates a container containing point cloud data and further adds additional data to the container as required for efficient transmission and reception.

[0391] A receiving device according to an embodiment receives and parses a container based on the systems shown in Figures 20 to 22. A receiving device corresponding to Figure 19 etc. decodes and restores point cloud data from the parsed bitstream.

[0392] The figure shows an overall architecture for storing or streaming point cloud data compressed based on video-based point cloud compression (V-PCC). The process of storing and streaming point cloud data includes an acquisition process, an encoding process, a transmission process, a decoding process, a rendering process, and / or a feedback process.

[0393] The embodiment proposes a method for efficiently providing point cloud media / content / data.

[0394] To efficiently provide point cloud media / content / data, the point cloud acquisition unit 20000 first acquires a point cloud video. For example, point cloud data is acquired using one or more cameras through a point cloud capture, synthesis, or generation process. This acquisition process can acquire a point cloud video including the 3D position of each point (represented by x, y, z position values, etc., hereinafter referred to as geometry) and the characteristics of each point (color, reflectance, transparency, etc.). The acquired point cloud video can also be generated as a file containing this information, such as a PLY (Polygon File format or the Stanford Triangle format) file. In the case of point cloud data having multiple frames, one or more files can be acquired. Point cloud-related metadata (e.g., metadata related to the capture, etc.) can also be generated during this process.

[0395] The captured point cloud video may require post-processing to improve the quality of the content. During the video capture process, the maximum and minimum depth values ​​may be adjusted within the range provided by the camera equipment. However, even after adjustment, point data from undesired areas may still be included. Therefore, post-processing may be performed to remove undesired areas (e.g., background) or fill spatial holes by recognizing connected spaces. In addition, point clouds extracted from cameras sharing a spatial coordinate system may be integrated into a single content by converting each point to a global coordinate system based on the position coordinates of each camera acquired through calibration. This allows for the acquisition of a point cloud video with a high point density.

[0396] The point cloud pre-processing unit 20001 generates a point cloud video into one or more pictures / frames. Here, a picture / frame generally refers to a unit representing one image at a specific time period. Furthermore, when dividing the points constituting the point cloud video into one or more patches (a set of points constituting a point cloud, where points belonging to the same patch are adjacent to each other in 3D space and are mapped in the same direction on the plane of a six-sided bounding box in the mapping process to a 2D image) and mapping them to a 2D plane, an occupancy map picture / frame can be generated, which is a binary map indicating with a value of 0 or 1 whether data exists at that position on the 2D plane. Furthermore, a geometry picture / frame can be generated, which is a depth map-type picture / frame representing the geometry of each point constituting the point cloud video in units of patches. A texture picture / frame can be generated, which is a picture / frame representing the color information of each point constituting the point cloud video in units of patches. In this process, metadata necessary to reconstruct a point cloud from the individual patches can be generated, including information about each patch (called additional information or additional patch information), such as its position in 2D / 3D space, size, etc. Such pictures / frames are generated successively in time order, and can form a video stream or a metadata stream.

[0397] The point cloud video encoder 20002 can encode one or more video streams related to a point cloud video. A video includes multiple frames, and a frame corresponds to a still image / picture. In this specification, the term point cloud video includes point cloud images / frames / pictures, and may be used interchangeably with point cloud images / frames / pictures. The point cloud video encoder performs a video-based point cloud compression (V-PCC) procedure. The point cloud video encoder can perform a series of procedures, such as prediction, transformation, quantization, and entropy coding, for compression and coding efficiency. The encoded data (encoded video / image information) is output in bitstream format. Based on the V-PCC procedure, the point cloud video encoder can encode the point cloud video into a geometry video, an attribute video, an occupancy map video, and metadata, such as information about patches, as described below. The geometry video may include geometry images, the attribute video may include attribute images, and the occupancy map video may include occupancy map images. The additional information, patch data, may include information about the patch. The feature video / image may include texture video / image.

[0398] The point cloud image encoder 20003 can encode the point cloud into one or more images related to a point cloud video. The point cloud image encoder 20003 performs a video-based point cloud compression (V-PCC) procedure. The point cloud image encoder can perform a series of procedures such as prediction, transformation, quantization, and entropy coding for compression and coding efficiency. The encoded image is output in bitstream format. When based on the V-PCC procedure, the point cloud image encoder 20003 can encode the point cloud image into a geometry image, an attribute image, an occupancy map image, and metadata, such as information about patches, as described below.

[0399] The point cloud video encoder and / or point cloud image encoder according to the embodiment generates a PCC bitstream (G-PCC and / or V-PCC bitstream) according to the embodiment.

[0400] Depending on the embodiment, the video encoder 20002, image encoder 20003, video decoder 20006, and image decoder may be implemented by a single encoder / decoder as described above, or may be implemented by separate paths as shown in the drawing.

[0401] An encapsulation unit (file / segment encapsulation unit) 20004 encapsulates the encoded point cloud data and / or metadata related to the point cloud in the form of a file or a segment for streaming. Here, the metadata related to the point cloud may be transmitted from a metadata processing unit or the like. The metadata processing unit may be included in the point cloud video / image encoder or may be configured as a separate component / module. The encapsulation unit encapsulates the video / image / metadata into a file format such as ISOBMFF or processes it into a form such as a DASH segment. According to an embodiment, the encapsulation unit can include the metadata related to the point cloud in the file format. For example, the point cloud metadata is included in boxes at various levels in the ISOBMFF file format or in data in a separate track within the file. According to an embodiment, the encapsulation unit can encapsulate the point cloud-related metadata itself in a file.

[0402] The encapsulation and encapsulation unit according to the embodiment divides and stores the G-PCC / V-PCC bitstream into one or more tracks in a file, and encapsulates the corresponding signaling information. Also, the atlas stream included in the G-PCC / V-PCC bitstream may be stored in a track in the file, and the associated signaling information may be stored. Furthermore, the SEI message present in the G-PCC / V-PCC bitstream may be stored in a track in the file, and the associated signaling information may be stored.

[0403] The transmission processing unit may process the encapsulated point cloud data according to a file format for transmission. The transmission processing unit may be included in the transmission unit or may be configured as a separate component / module. The transmission processing unit may process the point cloud data according to any transmission protocol. The processing for transmission may include processing for transmission via a broadcast network and processing for transmission via broadband. According to an embodiment, the transmission processing unit may process not only the point cloud data but also point cloud-related metadata transmitted from the metadata processing unit for transmission.

[0404] The transmitter transmits the point cloud bitstream or a file / segment containing the bitstream to the receiver of the receiving device via a digital storage medium or a network. Any transmission protocol may be used for transmission. The processed data for transmission is transmitted via a broadcast network and / or broadband. This data is transmitted to the receiver on an on-demand basis. Digital storage media include various media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitter includes elements for generating a media file in a predetermined file format and may also include elements for transmission via a broadcast / communication network. The receiver extracts the bitstream and transmits it to the decoding device.

[0405] The receiving unit receives the point cloud data transmitted by the point cloud data transmitting device according to this specification. Depending on the transmitted channel, the receiving unit may receive the point cloud data via a broadcast network, may receive the point cloud data via broadband, or may receive the point cloud video data via a digital storage medium. The receiving unit may decode the received data and render it according to the user's viewport, etc.

[0406] The receiving processor processes the received point cloud video data according to the transmission protocol. The receiving processor may be included in the receiving unit or may be configured as a separate component / module. In response to the processing for transmission performed on the transmitting side, the receiving processor performs the reverse process of the transmitting processor. The receiving processor transmits the acquired point cloud video to the decapsulating unit and transmits metadata related to the acquired point cloud to the metadata processor.

[0407] The decapsulating unit (file / segment decapsulation unit) 20005 decapsulates the point cloud data in a file format transmitted from the receiving processing unit. The decapsulating unit can decapsulate a file in ISOBMFF format or the like to obtain a point cloud bitstream or point cloud-related metadata (or another metadata bitstream). The obtained point cloud bitstream is transmitted to a point cloud video decoder and a point cloud image decoder, and the obtained point cloud-related metadata (or metadata bitstream) is transmitted to a metadata processing unit. The point cloud bitstream may include metadata (metadata bitstream). The metadata processing unit may be included in the point cloud video decoder or may be configured as a separate component / module. The point cloud-related metadata obtained by the decapsulating unit may be in the form of a box or track in a file format. If necessary, the decapsulating processing unit may receive metadata required for decapsulation from the metadata processing unit. The point cloud-related metadata may be transmitted to a point cloud video decoder and / or a point cloud image decoder for use in point cloud decoding, or to a renderer for use in point cloud rendering.

[0408] The point cloud video decoder 20006 receives the bitstream and decodes the video / image by performing the reverse process corresponding to the operation of the point cloud video encoder. In this case, the point cloud video decoder 20006 can decode the point cloud video by separating it into a geometry video, an attribute video, an occupancy map video, and auxiliary patch information, as described below. The geometry video may include a geometry image, the attribute video may include an attribute image, and the occupancy map video may include an occupancy map image. The auxiliary information may include auxiliary patch information. The attribute video / image may include a texture video / image.

[0409] The 3D geometry is restored using the decoded geometry video / image, occupancy map, and additional patch information, followed by a smoothing process. A color point cloud video / picture is restored by assigning color values ​​to the smoothed 3D geometry using the texture video / image. The renderer renders the restored geometry and color point cloud video / picture. The rendered video / image is displayed on a display unit. The user can view all or part of the rendered result on a VR / AR display or a general display.

[0410] The sensing / tracking unit (Sensing / Tracking) 20007 acquires orientation information and / or user viewport information from the user or the receiving side and transmits it to the receiving unit and / or transmitting unit. The orientation information indicates information about the position, angle, movement, etc. of the user's head, or information about the position, angle, movement, etc. of the device the user is looking at. Based on this information, information about the area the user is currently looking at in 3D space, i.e., viewport information, is calculated.

[0411] The viewport information may be information about the area that a user is currently viewing in 3D space through a device or HMD. A device such as a display may extract the viewport area based on orientation information, a vertical or horizontal FOV supported by the device, etc. The orientation or viewport information is extracted or calculated at the receiving side. The orientation or viewport information analyzed at the receiving side may be transmitted to the transmitting side via a feedback channel.

[0412] The receiving unit efficiently extracts or decodes only media data of a specific region, i.e., a region indicated by the orientation information and / or viewport information, from a file using the orientation information and / or viewport information indicating a region where a user is currently viewing, acquired by the sensing / tracking unit. The transmitting unit can efficiently encode only media data of a specific region, i.e., a region indicated by the orientation information and / or viewport information, or generate and transmit a file using the orientation information and / or viewport information acquired by the sensing / tracking unit.

[0413] The renderer renders the decoded point cloud data in 3D space. The rendered video / image is displayed on a display unit. The user can view all or part of the rendered result on a VR / AR display or a general display.

[0414] The feedback process may include transmitting various feedback information obtained from the rendering / display process to the transmitting side or to a decoder on the receiving side. The feedback process may provide interactivity in the consumption of point cloud data. According to an embodiment, head orientation information, viewport information indicating the area the user is currently viewing, etc. may be transmitted during the feedback process. According to an embodiment, a user may interact with something embodied in a VR / AR / MR / autonomous driving environment. In this case, information regarding the interaction may be transmitted to the transmitting side and the service provider during the feedback process. According to an embodiment, the feedback process may be omitted.

[0415] According to an embodiment, the above-mentioned feedback information is not only transmitted to the transmitting side but can also be consumed by the receiving side. That is, the above-mentioned feedback information may be used to perform decapsulation, decoding, rendering, etc. on the receiving side. For example, orientation information and / or viewport information may be used to prioritize decapsulate, decode, and render point cloud data for the area currently viewed by the user.

[0416] FIG. 21 illustrates an example of the configuration of a point cloud data storage and transmission device according to an embodiment.

[0417] Figure 21 shows a point cloud system according to an embodiment, part or all of which may include part or all of the transceiver device of Figure 1, the encoding process of Figure 4, the 2D video / image encoder of Figure 15, the decoding process of Figure 16, the transmitter device of Figure 18 and / or the receiver device of Figure 19, etc. Also, part or all of which may be included in or correspond to part or all of the system of Figure 20.

[0418] The point cloud data transmitting device according to the embodiment is configured as shown in the drawing. Each component of the transmitting device may be a module, unit, component, hardware, software, processor, etc.

[0419] The geometry, characteristics, additional data (or additional information), mesh data, etc. of the point cloud may each be organized into separate streams or stored in different tracks in the file, and may even be contained in separate segments.

[0420] The point cloud acquisition unit 21000 acquires a point cloud. For example, point cloud data is acquired through a point cloud capture, synthesis, or generation process using one or more cameras. This acquisition process obtains point cloud data including the 3D position of each point (represented by x, y, z position values, etc., hereinafter referred to as geometry) and the characteristics of each point (color, reflectance, transparency, etc.), and this data can be generated as, for example, a PLY (Polygon File format or the Stanford Triangle format) file. In the case of point cloud data having multiple frames, one or more files can be acquired. During this process, point cloud-related metadata (e.g., metadata related to the capture, etc.) can be generated.

[0421] The patch generation unit 21001 generates patches from point cloud data. The patch generation unit 21001 generates point cloud data or a point cloud video in one or more pictures / frames. Generally, a picture / frame refers to a unit representing one image at a specific time period. When dividing the points constituting the point cloud video into one or more patches (a set of points constituting a point cloud, where points belonging to the same patch are adjacent to each other in 3D space and are mapped in the same direction on a six-sided bounding box plane in the mapping process to a 2D image) and mapping them to a 2D plane, an occupancy map picture / frame can be generated, which is a binary map indicating with a value of 0 or 1 whether data exists at that position on the 2D plane. The patch generation unit 21001 also generates a geometry picture / frame, which is a depth map-type picture / frame representing the geometry of each point constituting the point cloud video in units of patches. The patch generation unit 21001 can also generate a texture picture / frame, which is a picture / frame representing the color information of each point constituting the point cloud video in units of patches. In this process, the metadata necessary to reconstruct a point cloud from the individual patches can be generated, which may include information about the patches, such as their position in 2D / 3D space, size, etc. Such pictures / frames are generated successively in time order and can constitute a video stream or metadata stream.

[0422] The patches may also be used for 2D image mapping, for example, by projecting point cloud data onto each face of a cube. After generating the patches, a geometry image, one or more feature images, an occupancy map, additional data, and / or mesh data can be generated based on the generated patches.

[0423] The pre-processing unit or controller performs geometry image generation, attribute image generation, occupancy map generation, auxiliary data generation, and / or mesh data generation.

[0424] The Geometry Image Generation unit 21002 generates a geometry image based on the results of patch generation. Geometry represents points in 3D space. The geometry image is generated based on the patch using an occupancy map containing information related to 2D image packing of the patch, additional data (patch data), and / or mesh data. The geometry image is related to information such as the depth (e.g., closeness, distance) of the patch generated after patch generation.

[0425] The attribute image generation unit 21003 generates an attribute image. For example, the attribute may indicate texture. The texture may be a color value corresponding to each point. According to an embodiment, an image of multiple (N) attributes (color, reflectance, etc.) including texture is generated. The multiple attributes include material (information about material), reflectance, etc. Also, according to an embodiment, the attribute may further include information that the color changes depending on the visual sense or light even for the same texture.

[0426] The occupancy map generation unit 21004 generates an occupancy map from the patch. The occupancy map contains information indicating the presence or absence of data in a pixel, such as its geometry or feature image.

[0427] The auxiliary data generation unit 21005 generates auxiliary data including information about patches. That is, the auxiliary data indicates metadata about patches of point cloud objects. For example, it may indicate information such as normal vectors for the patches. Specifically, according to an embodiment, the auxiliary data includes information necessary for reconstructing a point cloud from the patches (e.g., information about the position and size of the patches in 2D / 3D space, projection plane (normal) identification information, patch mapping information, etc.).

[0428] A mesh data generation unit 21006 generates mesh data from the patches. The mesh represents connectivity information between adjacent points. For example, it may represent triangle data. For example, the mesh data according to the embodiment represents connectivity information between each point.

[0429] The point cloud preprocessor or controller generates metadata related to patch generation, geometry image generation, feature image generation, occupancy map generation, additional data generation, and mesh data generation.

[0430] The point cloud transmitting device performs video encoding and / or image encoding according to the result generated by the pre-processing unit. The point cloud transmitting device generates not only point cloud video data but also point cloud image data. According to an embodiment, the point cloud data may include only video data, only image data, and / or both video data and image data.

[0431] The video encoder 21007 performs geometry video compression, feature video compression, occupancy map video compression, additional data compression, and / or mesh data compression. The video encoder generates a video stream containing each of the encoded video data.

[0432] Specifically, geometry video compression encodes point cloud geometry video data, attribute video compression encodes attribute video data of point clouds, additional data compression encodes additional data related to point cloud video data, and mesh data compression encodes mesh data of point cloud video data. Each operation of the point cloud video encoding unit is performed in parallel.

[0433] The image encoder 21008 performs geometry image compression, feature image compression, occupancy map image compression, additional data compression, and / or mesh data compression. The image encoder generates an image containing each of the encoded image data.

[0434] Specifically, geometry image compression encodes point cloud geometry image data, feature image compression encodes feature image data of the point cloud, additional data compression encodes additional data associated with the point cloud image data, and mesh data compression encodes mesh data associated with the point cloud image data. Each operation of the point cloud image encoding unit is performed in parallel.

[0435] The video encoder and / or the image encoder receive the metadata from the pre-processor, and perform their respective encoding processes based on the metadata.

[0436] The File / Segment Encapsulation unit 21009 encapsulates video streams and / or images into files and / or segments. The File / Segment Encapsulation unit performs video track encapsulation, metadata track encapsulation, and / or image encapsulation.

[0437] A video track encapsulation can encapsulate one or more video streams into one or more tracks.

[0438] Metadata track encapsulation encapsulates metadata related to the video stream and / or images in one or more tracks. The metadata includes data related to the content of the point cloud data, such as initial viewing orientation metadata. Depending on the embodiment, the metadata may be encapsulated in a metadata track, or may be encapsulated together in a video track or an image track.

[0439] Image encapsulation encapsulates one or more images into one or more tracks or items.

[0440] For example, according to an embodiment, if four video streams and two images are input to the encapsulator, the four video streams and two images are encapsulated into one file.

[0441] The point cloud video encoder and / or point cloud image encoder according to the embodiment generates a G-PCC / V-PCC bitstream according to the embodiment.

[0442] The file / segment encapsulation unit receives the metadata from the preprocessing unit and performs encapsulation based on the metadata.

[0443] The files and / or segments generated by the file / segment encapsulation are transmitted by a point cloud transmission device or transmitter. For example, the segments can be delivered based on a DASH-based protocol.

[0444] The encapsulation unit according to the embodiment can divide and store a V-PCC bitstream into one or more tracks in a file, and encapsulate the signaling information for the V-PCC bitstream. Also, the atlas stream included in the V-PCC bitstream can be stored in a track in the file, and associated signaling information can be stored. Furthermore, the SEI message present in the V-PCC bitstream can be stored in a track in the file, and associated signaling information can be stored.

[0445] The delivery unit delivers the point cloud bitstream or a file / segment containing the bitstream to the receiving unit of the receiving device via a digital storage medium or a network. For transmission, the data is processed using any transmission protocol. After processing for transmission, the data is transmitted via a broadcast network and / or broadband. This data may also be transmitted to the receiving side on an on-demand basis. Digital storage media include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The delivery unit includes elements for generating a media file in a predetermined file format and elements for transmission via a broadcast / communication network. The delivery unit receives orientation information and / or viewport information from the receiving unit. The delivery unit transmits the acquired orientation information and / or viewport information (or information selected by the user) to the preprocessing unit, video encoding unit, image encoding unit, file / segment encapsulation unit, and / or point cloud encoding unit. Based on the orientation information and / or viewport information, the point cloud encoding unit encodes all point cloud data or encodes point cloud data indicated by the orientation information and / or viewport information. Based on the orientation information and / or viewport information, the file / segment encapsulation unit can encapsulate all point cloud data or encapsulate point cloud data indicated by the orientation information and / or viewport information. Based on the orientation information and / or viewport information, the transfer unit transfers all point cloud data or transfers point cloud data indicated by the orientation information and / or viewport information.

[0446] For example, the preprocessing unit may perform the above-described operations on all point cloud data, or may perform the operations on point cloud data indicated by the orientation information and / or viewport information. The video encoding unit and / or image encoding unit may perform the above-described operations on all point cloud data, or may perform the above-described operations on point cloud data indicated by the orientation information and / or viewport information. The file / segment encapsulation unit may perform the above-described operations on all point cloud data, or may perform the above-described operations on point cloud data indicated by the orientation information and / or viewport information. The sending unit may perform the above-described operations on all point cloud data, or may perform the above-described operations on point cloud data indicated by the orientation information and / or viewport information.

[0447] FIG. 22 illustrates an example of the configuration of a point cloud data receiving device according to an embodiment.

[0448] Figure 22 shows a point cloud system according to an embodiment, some / all of which may include some / all of the transceiver device of Figure 1, the encoding process of Figure 4, the 2D video / image encoder of Figure 15, the decoding process of Figure 16, the transmitting device of Figure 18, and / or the receiving device of Figure 19, etc. Also, some / all of which may be included in or correspond to some / all of the systems of Figures 20 and 21.

[0449] Each component of the receiving device may be a module, unit, component, hardware, software, processor, etc. A delivery client receives point cloud data, a point cloud bitstream, or a file / segment including the bitstream transmitted by the point cloud data transmitting device according to the embodiment. Depending on the transmitted channel, the receiving device receives point cloud data via a broadcast network or via broadband. Alternatively, the receiving device may receive point cloud data via a digital storage medium. The receiving device may include a process for decoding the received data and rendering it according to a user's viewport, etc. A receiving processor processes the received point cloud data in accordance with a transmission protocol. The receiving processor may be included in the receiving unit or may be configured as a separate component / module. Corresponding to the processing for transmission performed by the transmitting side, the receiving processor performs the reverse process of the above-mentioned transmitting processor. The receiving processor may transmit the acquired point cloud data to a file / segment decapsulating unit and transmit the acquired point cloud-related metadata to a metadata processing unit.

[0450] The sensing / tracking unit acquires orientation information and / or viewport information, and transmits the acquired orientation information and / or viewport information to the transmission client, file / segment decapsulation unit, and point cloud decoding unit.

[0451] The transmission client receives all point cloud data or receives point cloud data indicated by the orientation information and / or viewport information based on the orientation information and / or viewport information. The file / segment decapsulator decapsulates all point cloud data or decapsulates point cloud data indicated by the orientation information and / or viewport information based on the orientation information and / or viewport information. The point cloud decoder (video decoder and / or image decoder) decodes all point cloud data or decodes point cloud data indicated by the orientation information and / or viewport information based on the orientation information and / or viewport information. The point cloud processing unit processes all point cloud data or processes point cloud data indicated by the orientation information and / or viewport information.

[0452] The file / segment decapsulation unit 22000 performs video track decapsulation, metadata track decapsulation, and / or image decapsulation. The decapsulation processing unit decapsulates point cloud data in a file format transmitted from the receiving processing unit. The decapsulation processing unit decapsulates a file or segment in ISOBMFF format, etc., to obtain a point cloud bitstream and point cloud-related metadata (or another metadata bitstream). The obtained point cloud bitstream is transmitted to a point cloud decoding unit, and the obtained point cloud-related metadata (or metadata bitstream) is transmitted to a metadata processing unit. The point cloud bitstream may include metadata (metadata bitstream). The metadata processing unit may be included in the point cloud video decoder or may be configured as a separate component / module. The point cloud-related metadata obtained by the decapsulation processing unit may be in the form of a box or track in a file format. If necessary, the file / segment decapsulation unit may receive metadata required for decapsulation from the metadata processing unit. The point cloud related metadata may be transmitted to the point cloud decoding unit and used for point cloud decoding, or may be transmitted to the point cloud rendering unit and used for point cloud rendering. The file / segment decapsulating unit can generate metadata related to the point cloud data.

[0453] The Video Track Decapsulation module decapsulates the video tracks contained in files and / or segments, including the video streams containing geometry video, attribute video, occupancy maps, additional data, and / or mesh data.

[0454] Metadata Track Decapsulation decapsulates a bitstream that includes metadata and / or additional data associated with point cloud data.

[0455] Image Decapsulation decapsulates images including geometry images, feature images, occupancy maps, additional data and / or mesh data.

[0456] The decapsulation or decapsulation unit according to the embodiment divides and parses (decapsulates) a G-PCC / V-PCC bitstream based on one or more tracks in a file, and also decapsulates the corresponding signaling information. It also decapsulates an atlas stream included in the G-PCC / V-PCC bitstream based on the tracks in the file and parses the associated signaling information. It can also decapsulate an SEI message present in the G-PCC / V-PCC bitstream based on the tracks in the file, and obtain the associated signaling information.

[0457] The video decoding unit 22001 performs geometry video reconstruction, attribute video reconstruction, occupancy map reconstruction, additional data reconstruction, and / or mesh data reconstruction. The video decoding unit decodes the geometry video, attribute video, additional data, and / or mesh data, corresponding to the process of video encoding and additional data of the point cloud transmitting device according to the embodiment.

[0458] The image decoding unit 22002 performs geometry image restoration, feature image restoration, occupancy map restoration, additional data restoration, and / or mesh data restoration. The image decoding unit decodes the geometry image, feature image, additional data, and / or mesh data in accordance with the process performed by the image encoding unit of the point cloud transmitting device according to the embodiment.

[0459] The video decoding unit and image decoding unit according to the embodiment may be processed by a single video / image decoder as described above, or may be performed in separate paths as shown in the figure.

[0460] The video decoder and / or the image decoder generate metadata associated with the video data and / or the image data.

[0461] A point cloud video encoder and / or a point cloud image encoder according to the embodiment decodes a G-PCC / V-PCC bitstream according to the embodiment.

[0462] The Point Cloud Processing unit 22003 performs geometry reconstruction and / or attribute reconstruction.

[0463] Geometry reconstruction reconstructs geometry video and / or geometry images from the decoded video data and / or decoded image data based on the occupancy map, additional data and / or mesh data.

[0464] The feature reconstruction restores the feature video and / or feature image from the decoded feature video and / or the decoded feature image based on the occupancy map, the additional data, and / or the mesh data. According to an embodiment, for example, the feature is a texture. According to an embodiment, the feature means multiple feature information. When there are multiple features, the point cloud processing unit according to an embodiment performs multiple feature reconstructions.

[0465] The point cloud processing unit can receive metadata from the video decoder, image decoder and / or file / segment decapsulator and process the point cloud based on the metadata.

[0466] A Point Cloud Rendering unit renders the reconstructed point cloud. The Point Cloud Rendering unit receives metadata from the Video Decoder, Image Decoder, and / or File / Segment Decapsulator, and renders the point cloud based on the metadata.

[0467] The display displays the rendered results on a physical display device.

[0468] As shown in Figures 15 to 19, the method / apparatus according to the embodiment encodes / decodes point cloud data, and then encapsulates and / or decapsulates the bitstream containing the point cloud data into a file and / or segment format.

[0469] For example, a point cloud data transmission device according to an embodiment encapsulates point cloud data based on a file, where the file includes a V-PCC track containing parameters related to the point cloud, a geometry track containing geometry, a feature track containing features, and an occupancy track containing an occupancy map.

[0470] In addition, the point cloud data receiving device according to the embodiment decapsulates based on a point cloud data file, where the file includes a V-PCC track containing parameters related to the point cloud, a geometry track containing geometry, a feature track containing features, and an occupancy track containing an occupancy map.

[0471] The above-mentioned operations are performed by the file / segment encapsulation unit 20004 in FIG. 20, the file / segment encapsulation unit 21009 in FIG. 21, the file / segment decapsulation unit 22000 in FIG. 22, and the like.

[0472] FIG. 23 illustrates an example of a structure that can be linked with a method / apparatus for transmitting and receiving point cloud data according to an embodiment.

[0473] In the structure according to the embodiment, any of the server 2360, robot 2310, autonomous vehicle 2320, XR device 2330, smartphone 2340, home appliance 2350, and / or HMD 2370 connects to the cloud network 2310. Here, the robot 2310, autonomous vehicle 2320, XR device 2330, smartphone 2340, or home appliance 2350 is referred to as an apparatus. In addition, the XR device 1730 may correspond to or be linked to a point cloud compressed data (PCC) apparatus according to the embodiment.

[0474] Cloud network 2300 may refer to a network that constitutes a part of a cloud computing infrastructure or exists within a cloud computing infrastructure, where cloud network 2300 may be configured using a 3G network, a 4G or LTE (Long Term Evolution) network, a 5G network, or the like.

[0475] The AI ​​server 2360 is connected to at least one of the robot 2310, autonomous vehicle 2320, XR device 2330, smartphone 2340, home appliance 2350, and / or HMD 2370 via the cloud network 2300, and assists in at least part of the processing of the connected devices 2310-2370.

[0476] An HMD (Head-Mounted Display) 23700 represents one type of XR device and / or PCC device according to an embodiment. An HMD type device according to an embodiment includes a communication unit, a control unit, a memory unit, an I / O unit, a sensor unit, and a power supply unit.

[0477] Various embodiments of the devices 2310 to 2350 to which the above-described technology is applied will be described below. Here, the devices 2310 to 2350 shown in Fig. 23 can be linked / coupled with the point cloud data transmitting / receiving device according to the above-described embodiments.

[0478] <PCC+XR>

[0479] The XR / PCC device 2330 applies PCC and / or XR (AR+VR) technology and is embodied in an HMD (Head-Mounted Display), a HUD (Head-Up Display) installed in a vehicle, a TV, a mobile phone, a smartphone, a computer, a wearable device, a home appliance, a digital sign, a vehicle, a fixed robot, a mobile robot, etc.

[0480] The XR / PCC device 2330 can obtain information about the surrounding space or real objects by analyzing 3D point cloud data or image data acquired by various sensors or from external devices to generate position data and attribute data for 3D points, and can render and output the XR objects to be output. For example, the XR / PCC device 2330 can output XR objects including additional information about the recognized objects in correspondence with the recognized objects.

[0481] <PCC+XR+Mobile Phone>

[0482] The XR / PCC device 2330 is implemented as a mobile phone 2340 to which PCC technology is applied.

[0483] The mobile phone 2340 decodes and displays the point cloud content based on PCC technology.

[0484] [[ID=

[0485] The autonomous vehicle 2320 is realized as a mobile robot, vehicle, unmanned aerial vehicle, etc. by applying PCC technology and XR technology.

[0486] The autonomous vehicle 2320 to which XR / PCC technology is applied may refer to an autonomous vehicle equipped with a means for providing XR images, an autonomous vehicle that can be controlled / interacted with in the XR images, etc. In particular, the autonomous vehicle 2320 that can be controlled / interacted with in the XR images is separate from the XR device 2330 and can be linked to each other.

[0487] The autonomous vehicle 2320, which is equipped with a means for providing XR / PCC images, acquires sensor information from sensors including cameras and outputs XR / PCC images generated based on the acquired sensor information. For example, the autonomous vehicle may be equipped with a HUD and output XR / PCC images, thereby providing passengers with XR / PCC objects corresponding to real objects or objects on a screen.

[0488] In this case, when an XR / PCC object is output to a HUD, at least a portion of the XR / PCC object may be output so as to overlap with an actual object toward which the passenger's gaze is directed. On the other hand, when an XR / PCC object is output to a display provided in an autonomous vehicle, at least a portion of the XR / PCC object may be output so as to overlap with an object on the screen. For example, an autonomous vehicle may output XR / PCC objects corresponding to objects such as a roadway, other vehicles, traffic lights, traffic signs, motorcycles, pedestrians, buildings, etc.

[0489] The VR (Virtual Reality) technology, AR (Augmented Reality) technology, MR (Mixed Reality) technology, and / or PCC (Point Cloud Compression) technology according to the embodiments can be applied to various devices.

[0490] In other words, VR technology is a display technology that presents real objects and backgrounds only as CG images. On the other hand, AR technology is a technology that displays virtual CG images on top of images of real things. MR technology is similar to the above-mentioned AR technology in that it mixes virtual objects into the real world. However, AR technology clearly distinguishes between real objects and virtual objects made of CG images, and uses virtual objects in a way that complements real objects, while MR technology is different from AR technology in that virtual objects and real objects are considered to have the same characteristics. More specifically, for example, hologram services are an application of the above-mentioned MR technology.

[0491] However, recently, rather than clearly distinguishing between VR, AR, and MR technologies, they are referred to as XR (extended reality) technology. Therefore, embodiments of the present invention can be applied to any of VR, AR, MR, and XR technologies. These technologies are applied with encoding / decoding based on PCC, V-PCC, and G-PCC technologies.

[0492] The PCC method / apparatus according to the embodiment can be applied to vehicles that provide autonomous driving services.

[0493] The autonomous vehicles that provide the autonomous driving service are connected to the PCC device via wired or wireless communication.

[0494] When a point cloud compressed data (PCC) transceiver according to an embodiment is connected to an autonomous vehicle via wired or wireless communication, it can receive and process AR / VR / PCC service-related content data that can be provided along with an autonomous driving service and transmit the content data to the autonomous driving vehicle. Furthermore, when the point cloud data transceiver is mounted on the autonomous driving vehicle, the point cloud transceiver can receive and process AR / VR / PCC service-related content data according to a user input signal input through a user interface device and provide the content data to a user. The vehicle or user interface device according to an embodiment can receive a user input signal. The user input signal according to an embodiment may include a signal instructing an autonomous driving service.

[0495] Depending on the embodiment, the point cloud data transmitting method / apparatus, the encoding operation and encoder of the transmitting method / apparatus according to the embodiment, the encapsulation and encapsulation unit of the point cloud data transmitting method / apparatus, the point cloud data receiving method / apparatus, the decoding operation and decoder of the point cloud data receiving method / apparatus according to the embodiment, the decapsulation and decapsulation unit of the point cloud data receiving method / apparatus, etc. may be simply referred to as the method / apparatus according to the embodiment.

[0496] FIG. 24 illustrates an object and a bounding box for the object for point cloud data according to an embodiment.

[0497] The transmitting device 10000 of Fig. 1, the receiving device 10005 of Fig. 1, the encoder of Fig. 15, the decoding process of Fig. 16, the decoder of Fig. 17, the transmitting device of Fig. 18, the receiving device of Fig. 19, the system structures for processing point cloud data of Figs. 20 to 22, and the XR device 2330 of Fig. 23 process an object 24000 that is the target of point cloud data. The method / apparatus according to the embodiment represents the object 24000 as a point, encodes it, encapsulates it, transmits it, receives it, decapsulates it, decodes it, restores it, and renders it. In this process, the object 24000 is processed based on a bounding box 24010 that surrounds the object.

[0498] The bounding box 24010 according to the embodiment is a basic access unit for representing and processing the object 24000. That is, the bounding box 24010 according to the embodiment refers to a box that represents the object 24000, which is the target of the point cloud data, based on a hexahedron. The bounding box 24010 may be divided or changed over time 24020, 24030 according to the embodiment.

[0499] In addition, the bounding box 24010 according to the embodiment may be divided into sub-bounding boxes. When the bounding box is divided, the divided bounding boxes are called sub-bounding boxes, and the entire bounding box is also called an overall bounding box.

[0500] The object 24000 changes dynamically over time. For example, the object 24000 moves from position 24020 to position 24030. At this time, the bounding box 24010 moves along with the object 24000.

[0501] The method / apparatus according to the embodiment provides bounding box related metadata of V-PCC content to support partial or spatial access of the V-PCC content according to a user's viewport.

[0502] The method / apparatus according to the embodiment defines bounding box information for point cloud content in a point cloud bitstream.

[0503] The method / apparatus according to the embodiment provides a way to store and signal point cloud-related bounding box information associated with a video track within a file.

[0504] The method / apparatus according to the embodiment provides a way to store point cloud bounding box information associated with an image item in a file.

[0505] The method / apparatus according to the embodiment relates to a sending or receiving device for providing point cloud content services that efficiently stores V-PCC bitstreams in tracks within a file and provides signaling therefor.

[0506] The method / apparatus according to the embodiment provides an advantage of efficient access to the V-PCC bitstream by efficiently storing and signaling the V-PCC bitstream within a file track, by providing a file storage technique and / or a technique for splitting and storing the V-PCC bitstream across one or more tracks within a file.

[0507] FIG. 25 illustrates a bounding box and a global bounding box of a dynamic point cloud object according to an embodiment.

[0508] FIG. 25 corresponds to the bounding box 24010 in FIG.

[0509] The bounding boxes 24010, 25000, and 25010 of the point cloud objects change dynamically over time.

[0510] Box 25000 is called the overall bounding box.

[0511] Box 25010 is not the entire bounding box, but corresponds to a bounding box, a partial bounding box, a sub-bounding box, etc.

[0512] Box 25000 and box 25010 are transmitted and received by the file encapsulation units 20004, 21009 and / or file decapsulation units 20005, 22000, etc., based on the file structures shown in Figures 40 and 41. For example, box 25000 and box 25010 are each expressed as a sample entry. Since they change over time, dynamically changing boxes 25000 and 25010 can be generated and processed based on sample entry information. Box 25000 and box 25010 describe object 25020. Depending on the embodiment, an anchor point 25030 for object 25020 may be included.

[0513] Therefore, the receiving method / apparatus according to the embodiment or a player connected thereto can receive information about a bounding box or an overall bounding box that changes to enable spatial or partial access of the point cloud object / content depending on the user viewport, etc. To this end, the point cloud data transmitting method / apparatus according to the embodiment can include the information according to the embodiment in a V-PCC bitstream or in the form of signaling or metadata in a file.

[0514] FIG. 26 shows the structure of a bitstream containing point cloud data according to an embodiment.

[0515] Bitstream 26000 in Fig. 26 corresponds to bitstream 27000 in Fig. 27. The bitstreams in Fig. 26 and Fig. 27 are generated by transmitting device 10000 in Fig. 1, point cloud video encoder 10002, the encoder in Fig. 4, the encoder in Fig. 15, the transmitting device in Fig. 18, processor 20001 in Fig. 20, video / image encoder 20002, processors 21001 to 21006, and video / image encoders 21007 and 21008 in Fig. 21, etc.

[0516] The bitstreams of Figures 26 and 27 are stored in containers (files such as Figures 40 and 41) by the file / segment encapsulation unit of Figure 1, the file / segment encapsulation unit 20004 of Figure 20, and the file / segment encapsulation unit 21009 of Figure 21, etc.

[0517] The bitstreams in FIGS. 26 and 27 are transmitted by the transmitter 10004 in FIG. 1 or the like.

[0518] The receiving device 10005, receiver 10006, etc. in FIG. 1 receive a container (files such as those in FIGS. 40 and 41) including the bitstreams in FIGS.

[0519] The bitstreams of Figures 26 and 27 are parsed from the container by the file / segment decapsulating unit 10007 of Figure 1, the file / segment decapsulating unit 20005 of Figure 20, and the file / segment decapsulating unit 22000 of Figure 22, etc.

[0520] The bitstreams in Figures 26 and 27 are decoded and restored by the point cloud video decoder 10008 in Figure 1, the decoder in Figure 16, the decoder in Figure 17, the receiving device in Figure 19, the video / image decoder 20006 in Figure 20, the video / image decoders 22001 and 22002 in Figure 22, and the processor 22003, etc., and are provided to the user.

[0521] A sample stream V-PCC unit included in a bitstream 26000 for point cloud data according to an embodiment includes a V-PCC unit size 26010 and a V-PCC unit 26020.

[0522] The definitions of each abbreviation are as follows: VPS (V-PCC parameter set), аD (atlas data), OVD (occupancy video data), GVD (geometry video data), AVD (attribute video data).

[0523] Each V-PCC unit 26020 includes a V-PCC unit header 26030 and a V-PCC unit payload 26040. The V-PCC unit header 26030 describes the V-PCC unit type. The attribute video data V-PCC unit header describes the attribute type and its index, multiple instances of the same attribute type that are supported, etc.

[0524] The occupancy, geometry, and attribute video data unit payloads 26050, 26060, and 26070 correspond to video data units. For example, the occupancy video data, geometry video data, and attribute video data 26050, 26060, and 26070 are HEVC NAL units. Such video data is decoded by a video decoder according to an embodiment.

[0525] FIG. 27 shows the structure of a bitstream containing point cloud data according to an embodiment.

[0526] FIG. 27 shows the structure of a bitstream containing point cloud data according to an embodiment that is encoded or decoded as described in FIGS.

[0527] The method / apparatus according to the embodiment generates a bitstream for a dynamic point cloud object, proposes a file format for the bitstream, and provides a signaling scheme therefor.

[0528] The method / apparatus according to the embodiment is a transmitter, receiver and / or processor for providing a point cloud content service that efficiently stores V-PCC (=V3C) bitstreams in tracks of a file and provides signaling therefor.

[0529] The method / apparatus according to the embodiment provides a data format for storing a V-PCC bitstream containing point cloud data. Accordingly, the receiving method / apparatus according to the embodiment receives point cloud data and provides a data storage and signaling scheme that allows efficient access to the point cloud data. Therefore, a transmitter and / or receiver can provide a point cloud content service based on a file storage technique that includes efficiently accessible point cloud data.

[0530] The method / apparatus according to the embodiment efficiently stores a point cloud bitstream (V-PCC bitstream) in a track of a file. Signaling information regarding an efficient storage technique is generated and stored in the file. To support efficient access to the V-PCC bitstream stored in a file, a technique is provided that further (or further modifies / combines) the storage technique according to the file embodiment to split and store the V-PCC bitstream across one or more tracks in the file.

[0531] The definitions of terms used in this specification are as follows:

[0532] VPS: V-PCC parameter set. AD: Atlas data. OVD: Occupancy video data. GVD: Geometry video data. AVD: Attribute video data. ACL: Atlas Coding Layer. AAPS: Atlas Adaptation Parameter Set. ASPS: Atlas sequence parameter set. The syntax structure including syntax elements according to the embodiment applies to zero or more entire coded atlas sequences (CASs) and is determined by the content of the syntax elements of the ASPS referenced by the syntax elements in each tile group header.

[0533] AFPS: Atlas frame parameter set. A syntax structure containing syntax elements that apply to zero or more entire coded atlas frames, as determined by the content of the syntax elements in the tile group header.

[0534] SEI: Supplemental enhancement information

[0535] Atlas: A collection of 2D bounding boxes, i.e., patches projected onto a rectangular frame, that correspond to 3D bounding boxes in 3D space. An atlas represents a subset of a point cloud.

[0536] Atlas sub-bitstream: A sub-bitstream extracted from a V-PCC bitstream that contains an Atlas NAL bitstream portion.

[0537] V-PCC content: A point cloud encoded based on V-PCC (V3C).

[0538] V-PCC track: A volumetric visual track that carries the atlas bitstream of a V-PCC bitstream.

[0539] V-PCC component track: A video track that conveys 2D video coded data for the occupancy map, geometry, and attribute component video bitstream of a V-PCC bitstream.

[0540] An embodiment for supporting partial access to dynamic point cloud objects will be described. The embodiment includes atlas tile group information associated with partial data of V-PCC objects included in each spatial region at the file system level. The embodiment also includes an extended signaling scheme for label and / or patch information included in each atlas tile group.

[0541] Referring to FIG. 27, the structure of a point cloud bitstream included in data transmitted and received by a method / apparatus according to an embodiment is shown.

[0542] The point cloud data compression and decompression technique according to the embodiment illustrates volumetric encoding and decoding of point cloud visual information.

[0543] A point cloud bitstream (also referred to as a V-PCC bitstream or V3C bitstream, 27000) containing a coded point cloud sequence (CPCS) is composed of a sample stream V-PCC unit 27010. The sample stream V-PCC unit 27010 conveys V-PCC parameter set (VPS) data 27020, an atlas bitstream 27030, a 2D video coding occupancy map bitstream 27040, a 2D video coding geometry bitstream 27050, and zero or more 2D video coding attribute bitstreams 27060.

[0544] The point cloud bitstream 27000 includes a sample stream VPCC header 27070.

[0545] SSVH Unit Size Precision (ssvh_unit_size_precision_bytes_minus1): This value, plus 1, indicates the precision in bytes of the SSVU VPCC Unit Size (ssvu_vpcc_unit_size) element in all sample stream V-PCC units. ssvh_unit_size_precision_bytes_minus1 has a range of 0 to 7.

[0546] The syntax 27080 of sample stream V-PCC unit 27010 is as follows: Each sample stream V-PCC unit contains one of the following V-PCC unit types: VPS, AD, OVD, GVD, and AVD. The contents of each sample stream V-PCC unit are associated with the same access unit as the V-PCC unit contained within the sample stream V-PCC unit.

[0547] SSVU VPCC unit size (ssvu_vpcc_unit_size): Indicates the size in bytes of the subsequent V-PCC unit. The number of bits used to indicate ssvu_vpcc_unit_size is (ssvh_unit_size_precision_bytes_minus1+1)*8.

[0548] The method / apparatus according to the embodiment receives the bitstream of FIG. 27 including encoded point cloud data, and generates files such as those shown in FIGS. 40 and 41 using encapsulation units 20004, 21009, etc.

[0549] The method / apparatus according to the embodiment receives the files shown in FIGS. 40 and 41 and decodes the point cloud data using a decapsulator 22000 or the like.

[0550] The VPS 27020 and / or AD 27030 are encapsulated in the fourth track (V3C track) 40030.

[0551] The OVD 27040 is encapsulated in the second track (occupied track) 40010 .

[0552] The GVD 27050 is encapsulated in the third track (geometry track) 40020.

[0553] AVD27060 is encapsulated in the first track (attribute track) 40000.

[0554] FIG. 28 shows a V-PCC unit and a V-PCC unit header according to an embodiment.

[0555] FIG. 28 shows the syntax of the V-PCC unit 26020 and V-PCC unit header 26030 described in FIG.

[0556] A V-PCC bitstream according to an embodiment includes a series of V-PCC sequences.

[0557] A V-PCC unit type with a value of vuh_unit_type equal to VPCC_VPS is expected to be the first V-PCC unit type in a V-PCC sequence. All other V-PCC unit types follow this unit type without any additional restrictions in their coding order. The V-PCC unit payload of a V-PCC unit carrying occupancy video, attribute video, or geometry video is composed of one or more NAL units. (A V-PCC bitstream contains a series of V-PCC sequences. A vpcc unit type with a value of vuh_unit_type equal to VPCC_VPS is expected to be the first V-PCC unit type in a V-PCC sequence. All other V-PCC unit types follow this unit type without any additional restrictions in their coding order. A V-PCC unit payload of a V-PCC unit carrying occupancy video, attribute video, or geometry video is composed of one or more NAL units.)

[0558] The VPCC unit includes a header and a payload.

[0559] The VPCC unit header contains the following information based on the VUH unit type:

[0560] VUH unit type indicates the type of V-PCC unit 26020 as follows:

[0561] [Table 1]

[0562] If the VUH unit type (vuh_unit_type) indicates proprietary video data (VPCC_AVD), geometry video data (VPCC_GVD), proprietary video data (VPCC_OVD), or atlas data (VPCC_AD), the VUH VPCC parameter set ID (vuh_vpcc_parameter_set_id) and VUH atlas ID (vuh_atlas_id) are transmitted in the unit header. The parameter set ID and atlas ID associated with the V-PCC unit can be transmitted.

[0563] If the unit type is atlas video data, the header of the unit carries an attribute index (vuh_attribute_index), an attribute partition index (vuh_attribute_partition_index), a map index (vuh_map_index), and an auxiliary video flag (vuh_auxiliary_video_flag).

[0564] If the unit type is geometry video data, the map index (vuh_map_index) and the auxiliary video flag (vuh_auxiliary_video_flag) are transmitted.

[0565] If the unit type is proprietary video data or atlas data, the header of the unit contains additional reserved bits.

[0566] VUH VPCC Parameter Set ID (vuh_vpcc_parameter_set_id): Indicates the value of vps_vpcc_parameter_set_id for the active V-PCC VPS. The VPCC parameter set ID in the header of the current V-PCC unit identifies the ID of the VPS parameter set and indicates the relationship between the V-PCC unit and the V-PCC parameter set.

[0567] VUH Atlas ID (vuh_atlas_id): Indicates the index of the atlas corresponding to the current V-PCC unit. The atlas index can be determined by the atlas ID in the header of the current V-PCC unit, and the atlas corresponding to the V-PCC unit can be identified.

[0568] VUH attribute index (vuh_attribute_index): Indicates the index of the attribute data conveyed from the attribute video data unit.

[0569] VUH attribute partition index (vuh_attribute_partition_index): Indicates the index of the attribute dimension group conveyed from the attribute video data unit.

[0570] VUH Map Index (vuh_map_index): If present, this value indicates the map index of the current geometry or attribute stream.

[0571] VUH Auxiliary Video Flag (vuh_auxiliary_video_flag): A value of 1 indicates that the associated geometry or attribute video data unit is only RAW and / or EOM coded point video. A value of 0 indicates that the associated geometry or attribute video data unit can contain RAW and / or EOM coded points.

[0572] VUH Raw Video Flag (vuh_raw_video_flag): A value of 1 indicates that the associated geometry or attribute video data unit is just RAW coded point video. A value of 0 indicates that the associated geometry or attribute video data unit contains RAW coded points. If this flag is not present, the value is inferred to be 0.

[0573] FIG. 29 illustrates the payload of a V-PCC unit according to an embodiment.

[0574] Figure 29 shows the syntax of the payload 26040 of the V-PCC unit.

[0575] The V-PCC unit type (vuh_unit_type) is a V-PCC parameter set (VPCC_VPS). The payload of the V-PCC unit includes a parameter set (vpcc_parameter_set( )).

[0576] If the V-PCC unit type (vuh_unit_type) is V-PCC atlas data (VPCC_AD), the payload of the V-PCC unit contains the atlas sub-bitstream (atlas_sub_bitstream( )).

[0577] If the V-PCC unit type (vuh_unit_type) is V-PCC private video data (VPCC_OVD), geometry video data (VPCC_GVD), or attribute video data (VPCC_AVD), the payload of the V-PCC unit contains a video bitstream (video_sub_bitstream( )).

[0578] FIG. 30 shows a parameter set (V-PCC parameter set) according to an embodiment.

[0579] FIG. 30 shows the syntax of a parameter set when the payload 26040 of the unit 26020 of the bitstream according to the embodiment includes a parameter set, as in FIGS.

[0580] The VPS in Figure 30 includes the following elements:

[0581] Profile tier level (profile_tier_level()): Indicates restrictions on the bitstreams and hence limits on the capabilities needed to decode the bitstreams. Profiles, tiers, and levels may also be used to indicate interoperability points between individual decoder implementations.

[0582] Parameter Set ID (vps_vpcc_parameter_set_id): provides an identifier for the V-PCC VPS for reference by other syntax elements.

[0583] Bounding box presence flag (sps_bounding_box_present_flag): A flag indicating whether information about the overall bounding box (a bounding box that includes all bounding boxes that change over time) of the point cloud object / content in the bitstream is present (sps_bounding_box_present_flag equal to 1 indicates the overall bounding box offset and the size information of point cloud content carried in this bitstream).

[0584] When the bounding box present flag has a particular value, the following bounding box elements are included in the VPS:

[0585] Bounding Box Offset X (sps_bounding_box_offset_x): Indicates the x offset of the overall bounding box offset and the size information of point cloud content carried in this bitstream in the Cartesian coordinates. When not present, the value of sps_bounding_box_offset_x is inferred to be 0.

[0586] Bounding box offset Y (sps_bounding_box_offset_y): Indicates the y offset of the overall bounding box offset and the size information of point cloud content carried in this bitstream in the Cartesian coordinates. When not present, the value of sps_bounding_box_offset_y is inferred to be 0.

[0587] Bounding Box Offset Z (sps_bounding_box_offset_z): Indicates the z offset of the overall bounding box offset and the size information of point cloud content carried in this bitstream in the Cartesian coordinates. When not present, the value of sps_bounding_box_offset_z is inferred to be 0.

[0588] Bounding box size width (sps_bounding_box_size_width): Indicates the width of the overall bounding box offset and the size information of point cloud content carried in this bitstream in the Cartesian coordinates. When not present, the value of sps_bounding_box_size_width is inferred to be 1.

[0589] Bounding box size height (sps_bounding_box_size_height): Indicates the height of the overall bounding box offset and the size information of point cloud content carried in this bitstream in the Cartesian coordinates. When not present, the value of sps_bounding_box_size_height is inferred to be 1.

[0590] Bounding box size depth (sps_bounding_box_size_depth): Indicates the depth of the overall bounding box offset and the size information of point cloud content carried in this bitstream in the Cartesian coordinates. When not present, the value of sps_bounding_box_size_depth is inferred to be 1.

[0591] Atlas Count (vps_atlas_count_minus1): Adding 1 to this value indicates the total number of supported atlases in the current bitstream.

[0592] Depending on the number of atlases, the following parameters are further included in the parameter set:

[0593] Frame width (vps_frame_width[j]): Indicates the V-PCC frame width in terms of integer luma samples for the atlas with index j. This frame width is the nominal width that is associated with all V-PCC components for the atlas with index j.

[0594] Frame height (vps_frame_height[j]): Indicates the V-PCC frame height in terms of integer luma samples for the atlas with index j. This frame height is the nominal height that is associated with all V-PCC components for the atlas with index j.

[0595] MAP Count (vps_map_count_minus1[j]): This value plus 1 indicates the number of maps used for encoding the geometry and attribute data for the atlas with index j.

[0596] If the MAP count (vps_map_count_minus1[j]) is greater than 0, the following parameters are further included in the parameter set:

[0597] Depending on the value of the MAP count (vps_map_count_minus1[j]), the following parameters are further included in the parameter set:

[0598] Multi-map stream present flag (vps_multiple_map_streams_present_flag[j]): If this value is 0, it indicates that all geometry or attribute maps for index j are present in a single geometry or attribute video stream, respectively. If this value is 1, it indicates that all geometry or attribute maps for the atlas with index j are present in separate video streams (equal to 0 indicates that all geometry or attribute maps for the atlas with index j are placed in a single geometry or attribute video stream, respectively. vps_multiple_map_streams_present_flag[j] equal to 1 indicates that all geometry or attribute maps for the atlas with index j are placed in separate video streams).

[0599] If the multi map streams present flag (vps_multiple_map_streams_present_flag[j]) indicates 1, then vps_map_absolute_coding_enabled_flag[j][i] is further included in the parameter set; otherwise, vps_map_absolute_coding_enabled_flag[j][i] has 1.

[0600] Map absolute coding enabled flag (vps_map_absolute_coding_enabled_flag[j][i]): If this value is 1, it indicates that the geometry map with index I for the atlas with index j is coded without any form of map prediction. If this value is 1, it indicates that the geometry map with index I for the atlas with index j is predicted first, before any other coded map, before coding (equal to 1 indicates that the geometry map with index i for the atlas with index j is coded without any form of map prediction. vps_map_absolute_coding_enabled_flag[j][i]equal to 0 indicates that the geometry map with index i for the atlas with index j is first predicted from another, earlier coded map, prior to coding).

[0601] The map absolute coding enabled flag (vps_map_absolute_coding_enabled_flag[j][0]) being 1 indicates that the geometry map with index 0 is coded without map prediction.

[0602] If the map absolute coding enabled flag (vps_map_absolute_coding_enabled_flag[j][i]) is 0 and I is greater than 0, vps_map_predictor_index_diff[j][i] is further included in the parameter set. Otherwise, vps_map_predictor_index_diff[j][i] is 0.

[0603] Map predictor index difference (vps_map_predictor_index_diff[j][i]): This value is used to compute the predictor of the geometry map with index i for the atlas with index j when vps_map_absolute_coding_enabled_flag[j][i] is equal to 0.

[0604] Additional video present flag (vps_auxiliary_video_present_flag[j]): If this value is 1, it indicates that auxiliary information for the atlas with index j, such as RAW or EOM patch data, is stored in a separate video stream, referred to as the auxiliary video stream. If this value is 0, it indicates that auxiliary information for the atlas with index j, i.e., RAW or EOM patch data, may be stored in a separate video stream, referred to as the auxiliary video stream. vps_auxiliary_video_present_flag[j] equal to 0 indicates that auxiliary information for the atlas with index j is not stored in a separate video stream.

[0605] Row patch enabled flag (vps_raw_patch_enabled_flag[j]): Equal to 1 indicates that patches with RAW coded points for the atlas with index j may be present in the bitstream.

[0606] When the line patch enable flag has a specific value, the following elements are included in the VPS:

[0607] Row Separate Video Present Flag (vps_raw_separate_video_present_flag[j]): Equal to 1 indicates that RAW coded geometry and attribute information for the atlas with index j may be stored in a separate video stream.

[0608] Occupancy_information(): Contains a set of parameters related to the occupancy video.

[0609] Geometry_information(): Contains a set of parameters related to the geometry video.

[0610] Attribute_information(): Contains a set of parameters related to the attribute video.

[0611] Extension Present Flag (vps_extension_present_flag): Indicates whether the syntax element vps_extension_length is present in the parameter set (vpcc_parameter_set) syntax structure. If this value is 0, it indicates that the syntax element vps_extension_length is not present (equal to 1 specifies that the syntax element vps_extension_length is present in the vpcc_parameter_set syntax structure. vps_extension_present_flag equal to 0 specifies that the syntax element vps_extension_length is not present).

[0612] Extension Length (vps_extension_length_minus1): This value plus 1 specifies the number of vps_extension_data_byte elements that follow this syntax element.

[0613] The extension length (vps_extension_length_minus1) allows extension data to be further included in the parameter set.

[0614] Extension Data (vps_extension_data_byte): May contain any data contained by the extension (may have any value).

[0615] According to the method / apparatus of the embodiment, the transmitting device can generate point information included in the bounding box using the above-described bounding box-related information according to the embodiment and transmit the information to the receiving device. The receiving device can efficiently obtain and decode point cloud data related to the bounding box based on the point information included in the bounding box. Furthermore, according to the embodiment, the points related to the bounding box include points included in the bounding box, the origin point of the bounding box, etc.

[0616] FIG. 31 shows the structure of an atlas sub-bitstream according to an embodiment.

[0617] FIG. 31 shows an example in which payload 26040 of unit 26020 of bitstream 26000 of FIG. 26 carries atlas sub-bitstream 31000.

[0618] The V-PCC unit payload of a V-PCC unit carrying an atlas sub-bitstream includes one or more sample stream NAL units 31010.

[0619] The atlas sub-bitstream 31000 according to the embodiment includes a sample stream NAL header 31020 and a sample stream NAL unit 31010 .

[0620] The sample stream NAL header 31020 includes a unit size precision byte (ssnh_unit_size_precision_bytes_minus1), whose value, plus 1, specifies the precision, in bytes, of the ssnu_nal_unit_size element in all sample stream NAL units. ssnh_unit_size_precision_bytes_minus1 has a range of 0 to 7.

[0621] The sample stream NAL unit 31010 contains the NAL unit size (ssnu_nal_unit_size).

[0622] NAL unit size (ssnu_nal_unit_size) indicates the size in bytes of the subsequence NAL unit. The number of bits used to represent ssnu_nal_unit_size is equal to (ssnh_unit_size_precision_bytes_minus1+1)*8 (specifies the size, in bytes, of the subsequent NAL_unit. The number of bits used to represent ssnu_nal_unit_size is equal to (ssnh_unit_size_precision_bytes_minus1+1)*8).

[0623] Each sample stream NAL unit includes an atlas sequence parameter set (ASPS) 31030, an atlas frame parameter set (AFPS) 31040, one or more atlas tile group information 31050, and one or more supplemental enhancement information (SEI) 31060, each of which is described below. Depending on the embodiment, an atlas tile group may be equivalently referred to as an atlas tile.

[0624] FIG. 32 illustrates an atlas sequence parameter set according to an embodiment.

[0625] FIG. 32 shows the syntax of the RBSP data structure included in the NAL unit when the NAL unit type is atlas sequence parameters.

[0626] Each sample stream NAL unit includes an atlas parameter set, such as ASPS, AAPS, AFPS, one or more atlas tile group information, and SEI.

[0627] The ASPS contains syntax elements that apply to zero or more entire coded atlas sequences (CASs), as determined by the content of the syntax elements in the ASPS, which are referenced as syntax elements in each tile group (tile) header.

[0628] ASPS includes the following elements:

[0629] ASPS Atlas Sequence Parameter Set ID (asps_atlas_sequence_parameter_set_id): Provides an identifier for the atlas sequence parameter set for reference by other syntax elements.

[0630] ASPS frame width (asps_frame_width): indicates the atlas frame width in terms of integer luma samples for the current atlas.

[0631] ASPS frame height (asps_frame_height): indicates the atlas frame height in terms of integer luma samples for the current atlas.

[0632] ASPS Log Patch Packing Block Size (asps_log2_patch_packing_block_size): Specifies the value of the variable PatchPackingBlockSize, that is used for the horizontal and vertical placement of the patches within the atlas.

[0633] ASPS log max atlas frame order count lsb (asps_log2_max_atlas_frame_order_cnt_lsb_minus4): Specifies the value of the variable MaxAtlasFrmOrderCntLsb that is used in the decoding process for the atlas frame order count.

[0634] ASPS Max Decoded Atlas Frame Buffering (asps_max_dec_atlas_frame_buffering_minus1): This value plus 1 specifies the maximum required size of the decoded atlas frame buffer for the CAS in units of atlas frame storage buffers.

[0635] ASPS Long-Term Reference Atlas Frames Flag (asps_long_term_ref_atlas_frames_flag): A value of 0 indicates that no long-term reference atlas frames are used for inter-prediction of coded atlas frames in the CAS. A value of 1 indicates that long-term reference atlas frames can use inter-prediction of one or more coded atlas frames in the CAS (equal to 0 specifies that no long-term reference atlas frame is used for inter-prediction of any coded atlas frame in the CAS. asps_long_term_ref_atlas_frames_flag equal to 1 specifies that long-term reference atlas frames may be used for inter-prediction of one or more coded atlas frames in the CAS).

[0636] Number of ASPS Reference Atlas Frame Lists (asps_num_ref_atlas_frame_lists_in_asps): Specifies the number of the ref_list_struct(rlsIdx) syntax structures included in the atlas sequence parameter set.

[0637] The atlas sequence parameter set includes as many reference list structures (ref_list_struct(i)) as the number of ASPS reference atlas frame lists (asps_num_ref_atlas_frame_lists_in_asps).

[0638] ASPS Eight Orientations Flag (asps_use_eight_orientations_flag): A value of 0 indicates that the patch orientation index (pdu_orientation_index[i][j]) for a patch with index J in a frame with index I is in the range of 0 to 1, inclusive. A value of 1 indicates that the patch orientation index (pdu_orientation_index[i][j]) for a patch with index J in a frame with index I is in the range of 0 to 7, inclusive (equal to 0 specifies that the patch orientation index for a patch with index j in a frame with index i, pdu_orientation_index[i][j], is in the range of 0 to 1, inclusive. asps_use_eight_orientations_flag equal to 1 specifies that the patch orientation index for a patch with index j in a frame with index i, pdu_orientation_index[i][j], is in the range of 0 to 7, inclusive).

[0639] Projection Patch Present Flag (asps_45degree_projection_patch_present_flag): A value of 0 indicates that patch projection information is not currently signaled for the atlas tile group. A value of 1 indicates that patch projection information is currently signaled for the atlas tile group (Equal to 0 specifies that the patch projection information is not signaled for the current atlas tile group. asps_45degree_projection_present_flag equal to 1 specifies that the patch projection information is signaled for the current atlas tile group).

[0640] If the type (atgh_type) is not skip tile (SKIP_TILE_GRP), the following elements are included in the atgh_type group (or tile) header:

[0641] ASPS vertical axis limits quantization enabled flag (asps_normal_axis_limits_quantization_enabled_flag): A value of 1 indicates that quantization parameters shall be used and signaled for quantizing the normal axis related elements of a patch data unit, a merge patch data unit, or an inter patch data unit. A value of 0 indicates that quantization is not applied on the normal axis related elements of a patch data unit, a merge patch data unit, or an inter patch data unit. (equal to 1 specifies that quantization parameters shall be signaled and used for quantizing the normal axis related elements of a patch data unit, a merge patch data unit, or an inter patch data unit. If asps_normal_axis_limits_quantization_enabled_flag is equal to 0, then no quantization is applied on any normal axis related elements of a patch data unit, a merge patch data unit, or an inter patch data unit.)

[0642] If asps_normal_axis_limits_quantization_enabled_flag is 1, then atgh_pos_min_z_quantizer is included in the atlas tile group (or tile) header.

[0643] ASPS vertical axis max delta value enabled flag (asps_normal_axis_max_delta_value_enabled_flag): A value of 1 indicates that the maximum nominal shift value of the vertical axis that may be present in the geometry information of the patch with index I of the frame with index J is indicated in the bitstream for each patch data unit, integrated patch data unit, or inter patch data unit. If this value is 0, it indicates that the maximum nominal shift value of the normal axis that may be present in the geometry information of a patch with index i in a frame with index j will not be indicated in the bitstream for each patch data unit, a merge patch data unit, or an inter patch data unit. (If asps_normal_axis_max_delta_value_enabled_flag is equal to 0, the maximum nominal shift value of the normal axis that may be present in the geometry information of a patch with index i in a frame with index j shall not be indicated in the bitstream for each patch data unit, a merge patch data unit, or an inter patch data unit.)

[0644] If asps_normal_axis_max_delta_value_enabled_flag is 1, then atgh_pos_delta_max_z_quantizer is included in the atlas tile group (or tile) header.

[0645] ASPS Remove Duplicate Point Enabled Flag (asps_remove_duplicate_point_enabled_flag): If this value is 1, it indicates that duplicate points are not reconstructed for the current atlas, where a duplicate point is a point with the same 2D and 3D geometry coordinates as another point from a Lower index map. If this value is 0, it indicates that all points are reconstructed (equal to 1 indicates that duplicated points are not reconstructed for the current atlas, where a duplicated point is a point with the same 2D and 3D geometry coordinates as another point from a Lower index map. asps_remove_duplicate_point_enabled_flag equal to 0 indicates that all points are reconstructed).

[0646] ASPS Max Decoded Atlas Frame Buffering (asps_max_dec_atlas_frame_buffering_minus1): This value plus 1 specifies the maximum required size of the decoded atlas frame buffer for the CAS in units of atlas frame storage buffers.

[0647] ASPS pixel deinterleaving flag (asps_pixel_deinterleaving_flag) - A value of 1 indicates that the decoded geometry and attribute videos for the current atlas contain spatially interleaved pixels from two maps. A value of 0 indicates that the decoded geometry and attribute videos corresponding to the current atlas contain pixels from only a single map. (equal to 1 indicates that the decoded geometry and attribute videos for the current atlas contain spatially interleaved pixels from two maps. asps_pixel_deinterleaving_flag equal to 0 indicates that the decoded geometry and attribute videos corresponding to the current atlas contain pixels from only a single map.)

[0648] ASPS patch precedence order flag (asps_patch_precedence_order_flag): If this value is 1, it indicates that the patch precedence for the current atlas is the same as the decoding order. If this value is 0, it indicates that the patch precedence for the current atlas is the reverse of the decoding order.

[0649] ASPS Patch Size Quantizer Present Flag (asps_patch_size_quantizer_present_flag): If this value is 1, it indicates that the patch size quantization parameters are present in an atlas tile group header. If this value is 0, it indicates that the patch size quantization parameters are not present.

[0650] If asps_patch_size_quantizer_present_flag is 1, then atgh_patch_size_x_info_quantizer and atgh_patch_size_y_info_quantizer are included in the atlas tile group (or tile) header.

[0651] Enhanced occupancy map flag for depth (asps_enhanced_occupancy_map_for_depth_flag): A value of 1 indicates that the decoded occupancy map video for the current atlas contains information related to whether intermediate depth positions between two depth maps are occupied. A value of 0 indicates that the decoded occupancy map video does not contain information related to whether intermediate depth positions between two depth maps are occupied. asps_eom_patch_enabled_flag equal to 0 indicates that the decoded occupancy map video does not contain information related to whether intermediate depth positions between two depth maps are occupied.

[0652] If asps_enhanced_occupancy_map_for_depth_flag or asps_point_local_reconstruction_enabled_flag is 1, asps_map_count_minus1 is included in asps.

[0653] ASPS Point Local Reconstruction Enabled Flag (asps_point_local_reconstruction_enabled_flag): A value of 1 indicates that point local reconstruction mode information may be present in the bitstream for the current atlas. A value of 0 indicates that no information related to the point local reconstruction mode is present in the bitstream for the current atlas.

[0654] If the ASPS point local reconstruction enabled flag (asps_point_local_reconstruction_enabled_flag) is 1, the ASPS point local reconstruction information (asps_point_local_reconstruction_information) is transferred to the atlas sequence parameter set.

[0655] ASPS Map Count (asps_map_count_minus1): This value plus 1 indicates the number of maps that may be used for encoding the geometry and attribute data for the current atlas.

[0656] ASPS Enhanced Occupancy Map Fixed Bit Count (asps_enhanced_occupancy_map_fix_bit_count_minus1): This value plus 1 indicates the size in bits of the EOM code word.

[0657] When asps_enhanced_occupancy_map_for_depth_flag and asps_map_count_minus1 have 0, asps_enhanced_occupancy_map_fix_bit_count_minus1 is included in ASPS.

[0658] ASPS Representation Thickness (asps_surface_thickness_minus1): Adding 1 to this value specifies the maximum absolute difference between an explicitly coded depth value and an interpolated depth value when asps_pixel_deinterleaving_flag (or asps_pixel_interleaving_flag) or asps_point_local_reconstruction_flag is 1.

[0659] If asps_pixel_interleaving_flag or asps_point_local_reconstruction_enabled_flag is 1, the ASPS surface thickness is included in the ASPS.

[0660] asps_pixel_interleaving_flag corresponds to asps_map_pixel_deinterleaving_flag.

[0661] ASPS Map Pixel Deinterleaving Flag (asps_map_pixel_deinterleaving_flag[i]): If this value is 1, it indicates that the decoded geometry and attribute video corresponding to the map with index i in the current atlas contains spatially interleaved pixels corresponding to two maps. If this value is 0, it indicates that the decoded geometry and attribute video corresponding to the map with index i in the current atlas contains pixels corresponding to a single map. If not present, this value is inferred to 0 (Equal to 1 indicates that decoded geometry and attribute videos corresponding to map with index i in the current atlas contain spatially interleaved pixels corresponding to two maps. asps_map_pixel_deinterleaving_flag[i] equal to 0 indicates that decoded geometry and attribute videos corresponding to map index i in the current atlas contain pixels corresponding to a single map. When not present, the value of asps_map_pixel_deinterleaving_flag[i] is inferred to be 0).

[0662] ASPS Point Local Reconstruction Enabled Flag (asps_point_local_reconstruction_enabled_flag): A value of 1 indicates that point local reconstruction mode information may be present in the bitstream for the current atlas. A value of 0 indicates that no information related to the point local reconstruction mode is present in the bitstream for the current atlas.

[0663] ASPS vui parameters present flag (asps_vui_parameters_present_flag): A value of 1 indicates that the vui_parameters( ) syntax structure is present. A value of 0 indicates that the vui_parameters( ) syntax structure is not present (equal to 1 specifies that the vui_parameters( ) syntax structure is present. asps_vui_parameters_present_flag equal to 0 specifies that the vui_parameters( ) syntax structure is not present).

[0664] ASPS Extension Flag (asps_extension_flag): A value of 0 specifies that no asps_extension_data_flag syntax elements are present in the ASPS RBSP syntax structure.

[0665] ASPS Extension Data Flag (asps_extension_data_flag): Indicates that data for an extension is included within the ASPS RBSP syntax structure.

[0666] Trailing bits (rbsp_trailing_bits): Used to add a stop bit (1) to indicate the end of the RBSP data, and then fill the remaining bits with 0 for byte alignment.

[0667] FIG. 33 illustrates an atlas frame parameter set according to an embodiment.

[0668] Figure 33 shows the syntax of the atlas frame parameter set included in the NAL unit when the NAL unit type is NAL_AFPS.

[0669] An atlas frame parameter set (AFPS) comprises a syntax structure that includes syntax elements that apply to zero or more entire coded atlas frames.

[0670] AFPS Atlas Frame Parameter Set ID (afps_atlas_frame_parameter_set_id): Identifies the atlas frame parameter set for reference by other syntax elements. The AFPS atlas frame parameter set provides an identifier that can be referenced by syntax elements.

[0671] AFPS Atlas Sequence Parameter Set ID (afps_atlas_sequence_parameter_set_id): Specifies the value of the active atlas sequence parameter set.

[0672] Atlas frame tile information (atlas_frame_tile_information( ): This will be explained with reference to FIG. 37.

[0673] Number of AFPS Reference Indices (afps_num_ref_idx_default_active_minus1): This value plus 1 specifies the inferred value of the variable NumRefIdxActive for the tile group or tile with atgh_num_ref_idx_active_override_flag equal to 0.

[0674] AFPS Additional Variable (afps_additional_lt_afoc_lsb_len): Specifies the value of the variable MaxLtAtlasFrmOrderCntLsb that is used in the decoding process for reference atlas frame.

[0675] AFPS 2D Position X Bit Count (afps_2d_pos_x_bit_count_minus1): This value plus 1 specifies the number of bits in the fixed-length representation of pdu_2d_pos_x[j] of patch with index j in an atlas tile group that refers to afps_atlas_frame_parameter_set_id.

[0676] AFPS 2D Position Y Bit Count (afps_2d_pos_y_bit_count_minus1): This value plus 1 specifies the number of bits in the fixed-length representation of pdu_2d_pos_y[j] of patch with index j in an atlas tile group that refers to afps_atlas_frame_parameter_set_id.

[0677] AFPS 3D Position X Bit Count (afps_3d_pos_x_bit_count_minus1): This value plus 1 specifies the number of bits in the fixed-length representation of pdu_3d_pos_x[j] of patch with index j in an atlas tile group that refers to afps_atlas_frame_parameter_set_id.

[0678] AFPS 3D Position Y Bit Count (afps_3d_pos_y_bit_count_minus1): This value plus 1 specifies the number of bits in the fixed-length representation of pdu_3d_pos_y[j] of patch with index j in an atlas tile group that refers to afps_atlas_frame_parameter_set_id.

[0679] AFPS LOD Bit Count (afps_lod_bit_count): specifies the number of bits in the Fixed-length representation of pdu_lod[j] of patch with index j in an atlas tile group that refers to afps_atlas_frame_parameter_set_id.

[0680] AFPS override EOM flag (afps_override_eom_for_depth_flag): A value of 1 indicates that the values ​​of afps_eom_number_of_patch_bit_count_minus1 and afps_eom_max_bit_count_minus1 are explicitly present in the bitstream. A value of 0 indicates that the values ​​of afps_eom_number_of_patch_bit_count_minus1 and afps_eom_max_bit_count_minus1 are implicitly derived (equal to 1 indicates that the values ​​of afps_eom_number_of_patch_bit_count_minus1 and afps_eom_max_bit_count_minus1 are explicitly present in the bitstream. afps_override_eom_for_depth_flag equal to 0 indicates that the values ​​of afps_eom_number_of_patch_bit_count_minus1 and afps_eom_max_bit_count_minus1 are implicitly derived).

[0681] AFPS EOM Patch Bit Count (afps_eom_number_of_patch_bit_count_minus1): This value plus 1 specifies the number of bits used to represent the number of geometry patches associated with the current EOM attribute patch.

[0682] AFPS EOM Max Bit Count (afps_eom_max_bit_count_minus1): Specifies the number of bits used to represent the number of EOM points per geometry patch associated with the current EOM attribute patch (plus 1 specifies the number of bits used to represent the number of EOM points per geometry patch associated with the current EOM attribute patch).

[0683] AFPS RAW 3D Position Bit Count Explicit Mode Flag (afps_raw_3d_pos_bit_count_explicit_mode_flag): Equal to 1 indicates that the bit count for rpdu_3d_pos_x, rpdu_3d_pos_y, and rpdu_3d_pos_z is explicitly coded in an atlas tile group header that refers to afps_atlas_frame_parameter_set_id.

[0684] AFPS Extension Flag (afps_extension_flag): A value of 0 specifies that no afps_extension_data_flag syntax elements are present in the AFPS RBSP syntax structure.

[0685] AFPS Extension Data Flag (afps_extension_data_flag): Contains extension related data.

[0686] FIG. 34 shows atlas_frame_tile_information according to an embodiment.

[0687] FIG. 34 shows the syntax of the atlas frame tile information included in FIG.

[0688] Single tile in AFTI atlas frame flag (afti_single_tile_in_atlas_frame_flag): A value of 1 specifies that there is only one tile in each atlas frame referring to the AFPS. A value of 0 specifies that there is more than one tile in each atlas frame referring to the AFPS.

[0689] AFTI Uniform Tile Spacing Flag (afti_uniform_tile_spacing_flag): A value of 1 indicates that the tile column and row boundaries are uniformly distributed across the atlas frame and are signaled using the afti_tile_cols_width_minus1 and afti_tile_rows_height_minus1 syntax elements, respectively. A value of 0 indicates that the tile column and row boundaries are either uniformly distributed across the atlas frame or not, and are signaled using syntax elements such as afti_num_tile_columns_minus1, afti_num_tile_rows_minus1, a list of syntax element pairs afti_tile_column_width_minus1[i], afti_tile_row_height_minus1[i], etc.

[0690] AFTI Tile Column Width (afti_tile_cols_width_minus1): This value plus 1 specifies the width of the tile columns excluding the right-most tile column of the atlas frame in units of 64 samples when afti_uniform_tile_spacing_flag is equal to 1.

[0691] AFTI Tile Rows Height (afti_tile_rows_height_minus1): This value plus 1 specifies the height of the tile rows excluding the bottom tile row of the atlas frame in units of 64 samples when afti_uniform_tile_spacing_flag is equal to 1.

[0692] If afti_uniform_tile_spacing_flag is not 1, the following elements are included in the atlas frame tile information:

[0693] Number of AFTI tile columns (afti_num_tile_columns_minus1): This value plus 1 specifies the number of tile columns partitioning the atlas frame when afti_uniform_tile_spacing_flag is equal to 0.

[0694] Number of AFTI tile rows (afti_num_tile_rows_minus1): Adding 1 to this value specifies the number of tile rows partitioning the atlas frame when afti_uniform_tile_spacing_flag is equal to 0.

[0695] AFTI tile column width (afti_tile_column_width_minus1[i]): This value plus 1 specifies the width of the i-th tile column in units of 64 samples.

[0696] The width of the AFTI tile columns is included in the atlas frame tile information by the afti_num_tile_columns_minus1 value.

[0697] AFTI tile row height (afti_tile_row_height_minus1[i]): This value plus 1 specifies the height of the i-th tile row in units of 64 samples.

[0698] The height of AFTI tile rows equal to the afti_num_tile_rows_minus1 value is included in the atlas frame tile information.

[0699] AFTI Single Tile Per Tile Group Flag (afti_single_tile_per_tile_group_flag): If this value is 1, it indicates that each tile group (or tile) that refers to this AFPS contains one tile. If this value is 0, it indicates that the tile group (or tile) that refers to this AFPS contains more than one tile. If not present, this value is inferred to be equal to 1. (equal to 1 specifies that each tile group that refers to this AFPS includes one tile. afti_single_tile_per_tile_group_flag equal to 0 specifies that a tile group that refers to this AFPS may include more than one tile. When not present, the value of afti_single_tile_per_tile_group_flag is inferred to be equal to 1.)

[0700] Based on afti_num_tile_groups_in_atlas_frame_minus1, the AFTI tile index (afti_tile_idx[i]) is included in the atlas frame tile information.

[0701] If the single tile per AFTI tile group flag (afti_single_tile_per_tile_group_flag) is 0, then afti_num_tile_groups_in_atlas_frame_minus1 is conveyed in the atlas frame tile information.

[0702] Number of tile groups (or tiles) in AFTI atlas frame (afti_num_tile_groups_in_atlas_frame_minus1): Indicates the number of tile groups (or tiles) in each atlas frame that represents AFPS. The value of afti_num_tile_groups_in_atlas_frame_minus1 ranges from 0 to NumTilesInAtlasFrame-1 (inclusive). If not present and afti_single_tile_per_tile_group_flag is 1, then the value of afti_num_tile_groups_in_atlas_frame_minus1 shall be inferred to NumTilesInAtlasFrame-1 (plus 1 specifies the number of tile groups in each atlas frame referring to the afps, The value of afti_num_tile_groups_in_atlas_frame_minus1 shall be in the range of 0 to NumTilesInAtlasFrame - 1, inclusive. When not present and afti_single_tile_per_tile_group_flag is equal to 1, the value of afti_num_tile_groups_in_atlas_frame_minus1 is inferred to be equal to NumTilesInAtlasFrame-1).

[0703] The atlas frame tile information includes the following elements: afti_num_tile_groups_in_atlas_frame_minus1 value.

[0704] AFTI top left tile index (afti_top_left_tile_idx[i]): Indicates the tile index of the tile located at the top left of the I-th tile group (or tile). The value of afti_top_left_tile_idx[i] is not equal to the value of afti_top_left_tile_idx[j] for values ​​of i that are not equal to j. If not present, the value of afti_top_left_tile_idx[i] is inferred to be equal to i. The length of the afti_top_left_tile_idx[i] syntax element is Ceil(Log2(NumTilesInAtlasFrame) bits (specifies the tile index of the tile located at the top-left corner of the i-th tile group. The value of afti_top_left_tile_idx[i] is not be equal to the value of afti_top_left_tile_idx[j] for any i not equal to j. When not present, the value of afti_top_left_tile_idx[i] is inferred to be equal to. The length of the afti_top_left_tile_idx[i] syntax element is Ceil(Log2(NumTilesInAtlasFrame ) bits).

[0705] AFTI Bottom Right Tile Index Delta (afti_bottom_right_tile_idx_delta[i]): Indicates the difference between afti_top_left_tile_idx[i] and the tile index of the tile located at the bottom right of the I-th tile group (or tile). If afti_single_tile_per_tile_group_flag is 1, the value of afti_bottom_right_tile_idx_delta[i] is inferred to be equal to 0. The length of the afti_bottom_right_tile_idx_delta[i] syntax element is Ceil(Log2(NumTilesInAtlasFrame-afti_top_left_tile_idx[i])) bits. afti_single_tile_per_tile_group_flag is equal to 1, the value of afti_bottom_right_tile_idx_delta[i] is inferred to be equal to 0. The length of the afti_bottom_right_tile_idx_delta[i] syntax element is Ceil(Log2(NumTilesInAtlasFrame - afti_top_left_tile_idx[i])) bits).

[0706] AFTI signaled tile group ID flag (afti_signalled_tile_group_id_flag): A value of 1 indicates that the tile group ID or tile ID for each tile group or each tile is signaled (equal to 1 specifies that the tile group ID for each tile group is signaled).

[0707] If the AFTI signaled tile group ID flag (afti_signalled_tile_group_id_flag) is 1, then afti_signalled_tile_group_id_length_minus1 and afti_tile_group_id[i] are conveyed in the atlas frame tile information. If this value is 0, then the tile group ID may not be signaled.

[0708] AFTI signaled tile group ID length (afti_signalled_tile_group_id_length_minus1): This value plus 1 indicates the number of bits used to represent the syntax element afti_tile_group_id[i] when present, and the syntax element atgh_address in tile group headers.

[0709] AFTI Tile Group ID (afti_tile_group_id[i]): Specifies the ID of the i-th tile group (or tile). The length of the afti_tile_group_id[i] syntax element is afti_signalled_tile_group_id_length_minus1 + 1 bits.

[0710] The atlas frame tile information includes AFTI tile group IDs (afti_tile_group_id[i]) for the afti_num_tile_groupS_in_atlas_frame_minus1 value.

[0711] FIG. 35 shows supplemental enhancement information (SEI) according to an embodiment.

[0712] FIG. 35 shows the detailed syntax of the SEI information included in the bitstream according to the embodiment as in FIG.

[0713] The receiving method / apparatus, system, etc. according to the embodiment decodes, restores, and displays point cloud data based on the SEI message.

[0714] The SEI message indicates that the payload contains corresponding data based on each payload type (payloadType).

[0715] For example, if the payload type (payloadType) is 12, the payload contains 3D bounding box information (3d_bouding_box_info(payloadSize)).

[0716] When the unit type (psd_unit_type) is prefix (PSD_PREFIX_SEI), the SEI information according to the embodiment includes buffering_period(payloadSize), pic_timing(payloadSize), filler_payload(payloadSize), user_data_registered_itu_t_t35(payloadSize), user_data_unregistered(payloadSize), recovery_point(payloadSize), no_display(payloadSize), time_code(payloadSize), regional_nesting(payloadSize), sei_manifest(payloadSize), sei_prefix_indication(payloadSize), geometry_transformation_params(payloadSize), 3d_bounding_box_info(payloadSize) (see Figure 35, etc.), 3d_region_mapping(payloadSize) (see Figure 66, etc.), reserved_sei_message(payloadSize), etc.

[0717] When the unit type (psd_unit_type) is a suffix (PSD_SUFFIX_SEI), the SEI information according to the embodiment includes filler_payload(payloadSize), user_data_registered_itu_t_t35(payloadSize), user_data_unregistered(payloadSize), decoded_PCC_hash(payloadSize), reserved_sei_message(payloadSize), etc.

[0718] FIG. 36 shows a 3D bounding box SEI according to an embodiment.

[0719] FIG. 36 shows the detailed syntax of the SEI information included in the bitstream according to the embodiment as in FIG.

[0720] Cancellation flag (3dbi_cancel_flag): A value of 1 indicates that the 3D bounding box information SEI message cancels the presence of any previous 3D bounding box information SEI message in the output order.

[0721] Object ID (object_id): An identifier of the point cloud object / content transmitted in the bitstream.

[0722] Bounding Box X (3d_bounding_box_x): The X coordinate value of the origin of the object's 3D bounding box.

[0723] Bounding box Y (3d_bounding_box_y): The Y coordinate value of the origin of the object's 3D bounding box.

[0724] Bounding Box Z (3d_bounding_box_z): The Z coordinate value of the origin of the object's 3D bounding box.

[0725] Bounding Box Delta X (3d_bounding_box_z): Indicates the size of the object's bounding box on the X axis.

[0726] Bounding Box Delta Y (3d_bounding_box_delta_y): Indicates the size of the object's bounding box on the Y axis.

[0727] Bounding Box Delta Z (3d_bounding_box_delta_z): Indicates the size of the object's bounding box on the Z axis.

[0728] FIG. 37 shows volumetric tiling information according to an embodiment.

[0729] FIG. 37 shows the detailed syntax of the SEI information included in the bitstream according to the embodiment as in FIG.

[0730] Volumetric tiling information SEI message

[0731] This SEI message informs a V-PCC decoder according to this embodiment to avoid different characteristics of a decoded point cloud, including correspondence of areas within a 2D atlas and the 3D space, relationship and labeling of areasS and association with objects.

[0732] The persistence scope for this SEI message is the remainder of the bitstream or until a new volumetric tiling SEI message is encountered. Only the corresponding parameters described in this SEI message are updated. If not modified or if the value of vti_cancel_flag is not equal to 1, previously defined parameters from an earlier SEI message persist.

[0733] FIG. 38 shows a volumetric tiling information object according to an embodiment.

[0734] Figure 38 shows the detailed syntax of the volumetric tiling information objects (volumetric_tiling_info_objects) included in Figure 37.

[0735] Based on vtiObjectLabelPresentFlag, vti3DBoundingBoxPresentFlag, vtiObjectPriorityPresentFlag, tiObjectHiddenPresentFlag, vtiObjectCollisionShapePresentFlag, vtiObjectDependencyPresentFlag, etc., the volumetric tiling information object includes elements as shown in Figure 38.

[0736] FIG. 39 shows volumetric tiling information labels according to an embodiment.

[0737] Figure 39 shows the detailed syntax of the volumetric tiling information labels (volumetric_tiling_info_labels) included in Figure 37.

[0738] Cancel flag (vti_cancel_flag): A value of 1 indicates that the volumetric tiling information SEI message cancels the presence of any previous volumetric tiling information SEI message in the output order. If vti_cancel_flag is 0, the volumetric tiling information is as shown in Figure 37.

[0739] Object Label Present Flag (vti_object_label_present_flag): If this value is 1, it indicates that object label information is currently present in the Volumetric Tiling Information SEI message. If this value is 0, it indicates that object label information is not present.

[0740] 3D Bounding Box Present Flag (vti_3d_bounding_box_present_flag): A value of 1 indicates that 3D bounding box information is currently present in the Volumetric Tiling Information SEI message. A value of 0 indicates that 3D bounding box information is not present.

[0741] Object priority present flag (vti_object_priority_present_flag): If this value is 1, it indicates that object priority information is currently present in the Volumetric Tiling Information SEI message. If this value is 0, it indicates that object priority information is not present.

[0742] Hidden object present flag (vti_object_hidden_present_flag): If this value is 1, it indicates that hidden object information is currently present in the Volumetric Tiling Information SEI message. If this value is 0, it indicates that hidden object information does not exist.

[0743] Object Collision Shape Present Flag (vti_object_collision_shape_present_flag): If this value is 1, it indicates that object collision information is currently present in the Volumetric Tiling Information SEI message. If this value is 0, it indicates that object collision shape information is not present.

[0744] Object dependency present flag (vti_object_dependency_present_flag): If this value is 1, it indicates that object dependency information is currently present in the Volumetric Tiling Information SEI message. If this value is 0, it indicates that object dependency information is not present.

[0745] Object Label Language Present Flag (vti_object_label_language_present_flag): If this value is 1, it indicates that object label language information is currently present in the Volumetric Tiling Information SEI message. If this value is 0, it indicates that object label language information is not present.

[0746] Bit equal to zero (vti_bit_equal_to_zero): This value is equal to 0.

[0747] Object Label Language (vti_object_label_language): Contains a language tag followed by a null-termination byte equal to 0x00. The length of the vti_object_label_language syntax element is less than or equal to 255 bytes, excluding the null-termination byte.

[0748] Number of object labels (vti_num_object_label_updates): Indicates the number of object labels currently updated by SEI.

[0749] Label index (vti_label_idx[i]): Indicates the label index of the ith label to be updated.

[0750] Label Cancel Flag (vti_label_cancel_flag): If this value is 1, it indicates that the label with the same index as vti_label_idx[i] is canceled and set with an empty string as well. If this value is 0, it indicates that the label with the same index as vti_label_idx[i] is updated with the information according to this element.

[0751] Bit equal to zero (vti_bit_equal_to_zero): This value is equal to 0.

[0752] Label (vti_label[i]): Indicates the label of the i-th label. The length of the vti_label[i] syntax element is equal to or less than 255 bytes excluding the null termination byte.

[0753] Bounding Box Scale (vti_bounding_box_scale_log2): Indicates the scale applied to the 2D bounding box parameters described for the object.

[0754] 3D Bounding Box Scale (vti_3d_bounding_box_scale_log2): Indicates the scale applied to the 3D bounding box parameters described for the object.

[0755] 3D Bounding Box Precision (vti_3d_bounding_box_precision_minus8): This value plus 8 indicates the precision of the 3D bounding box parameters that may be specified for an object.

[0756] Number of objects (vti_num_object_updates): Indicates the number of objects currently being updated by SEI.

[0757] The volumetric tiling information object (see FIG. 38) contains object-related information equal to the number of objects (vti_num_object_updates).

[0758] Object index (vti_object_idx[i]): Indicates the object index of the i-th object to be updated.

[0759] Object Cancellation Flag (vti_object_cancel_flag[i]): If this value is 1, it indicates that the object with the same index as i is canceled and the variable ObjectTracked[i] is set to 0. The 2D and 3D bounding box parameters of the object are set to 0. If this value is 0, it indicates that the object with the same index as vti_object_idx[i] is updated with the information according to this element. Also, the variable ObjectTracked[i] is set to 1.

[0760] Bounding box update flag (vti_bounding_box_update_flag[i]): If this value is 1, it indicates that 2D bounding box information exists for the object with index i. If this value is 0, it indicates that 2D bounding box information does not exist.

[0761] If vti_bounding_box_update_flag is 1 for vti_object_idx[i], the following bounding box elements for vti_object_idx[i] are included in the volumetric tiling information object:

[0762] Bounding box top (vti_bounding_box_top[i]): Indicates the vertical coordinate value of the top left position of the bounding box of the object with index i in the current atlas frame.

[0763] Bounding box left (vti_bounding_box_left[i]): Indicates the horizontal coordinate value of the top left position of the bounding box of the object with index i in the current atlas frame.

[0764] Bounding box width (vti_bounding_box_width[i]): Indicates the width of the bounding box of the object with index i.

[0765] Bounding box height (vti_bounding_box_height[i]): Indicates the height of the bounding box of the object with index i.

[0766] If vti3dBoundingBoxPresentFlag is 1, the following bounding box elements are included in the volumetric tiling information object:

[0767] 3D bounding box update flag (vti_3d_bounding_box_update_flag[i]): If this value is 1, it indicates that 3D bounding box information exists for the object with index i. If this value is 0, it indicates that 3D bounding box information does not exist.

[0768] If vti_3d_bounding_box_update_flag is 1 for vti_object_idx[i], the following bounding box related elements are included in the volumetric tiling information object:

[0769] 3D bounding box X (vti_3d_bounding_box_x[i]): Indicates the X coordinate value of the origin position of the 3D bounding box of the object with index i.

[0770] 3D bounding box Y (vti_3d_bounding_box_y[i]): Indicates the Y coordinate value of the origin position of the 3D bounding box of the object with index i.

[0771] 3D bounding box Z (vti_3d_bounding_box_z[i]): Indicates the Z coordinate value of the origin position of the 3D bounding box of the object with index i.

[0772] 3D bounding box delta X (vti_3d_bounding_box_delta_x[i]): Indicates the size of the bounding box on the X axis of the object with index i.

[0773] 3D bounding box delta Y (vti_3d_bounding_box_delta_y[i]): Indicates the size of the bounding box on the Y axis of the object with index i.

[0774] 3D bounding box delta Z (vti_3d_bounding_box_delta_z[i]): Indicates the size of the bounding box on the Z axis of the object with index i.

[0775] If vtiObjectPriorityPresentFlag is 1, the following priority-related elements are included in the volumetric tiling information object:

[0776] Object priority update flag (vti_object_priority_update_flag[i]): If this value is 1, it indicates that object priority update information exists for the object with index i. If this value is 0, it indicates that object priority information does not exist.

[0777] Object priority value (vti_object_priority_value[i]): Indicates the priority of the object with index i. The lower the priority value, the higher the priority.

[0778] If vtiObjectHiddenPresentFlag is 1, the following hidden information for vti_object_idx[i] is included in the volumetric tiling information object:

[0779] Object Hidden Flag (vti_object_hidden_flag[i]): If this value is 1, it indicates that the object with index i is hidden. If this value is 0, it indicates that the object with index i exists.

[0780] If vtiObjectLabelPresentFlag is 1, label-related update flags are included in the volumetric tiling information object.

[0781] Object label update flag (vti_object_label_update_flag): If this value is 1, it indicates that object label update information exists for the object with index i. If this value is 0, it indicates that object label update information does not exist.

[0782] If vti_object_label_update_flag is 1 for vti_object_idx[i], the object label index for vti_object_idx[i] is included in the volumetric tiling information object.

[0783] Object Label Index (vti_object_label_idx[i]): Indicates the label index of the object with index i.

[0784] If vtiObjectCollisionShapePresentFlag is 1, object collision related elements are included in the volumetric tiling information object.

[0785] Object collision shape update flag (vti_object_collision_shape_update_flag[i]): If this value is 1, it indicates that object collision shape update information exists for the object with index i. If this value is 0, it indicates that object collision shape update information does not exist.

[0786] If vti_object_collision_shape_update_flag is 1 for vti_object_idx[i], the object collision shape ID for vti_object_idx[i] is included in the volumetric tiling information object.

[0787] Object collision shape ID (vti_object_collision_shape_id[i]): Indicates the collision shape ID of the object with index i.

[0788] If vtiObjectDependencyPresentFlag is 1, object dependency association elements are included in the volumetric tiling information object.

[0789] Object dependency update flag (vti_object_dependency_update_flag[i]): If this value is 1, it indicates that object dependency update information exists for the object with object index i. If this value is 0, it indicates that object dependency update information does not exist.

[0790] If vti_object_dependency_update_flag is 1 for vti_object_idx[i], the object dependency association element for vti_object_idx[i] is included in the volumetric tiling information object.

[0791] Number of object dependencies (vti_object_num_dependencies[i]): Indicates the number of object dependencies with index i.

[0792] The volumetric tiling information object contains as many object dependency indices as vti_object_num_dependencies.

[0793] Object dependency index (vti_object_dependency_idx[i][j]): Indicates the index of the jth object that has dependency on the object with index i.

[0794] FIG. 40 illustrates the structure of an encapsulated V-PCC data container according to an embodiment.

[0795] FIG. 41 illustrates an isolated V-PCC data container structure according to an embodiment.

[0796] The point cloud video encoder 10002 of the transmitting device 10000 of Figure 1, the encoders of Figures 4 and 15, the transmitting device of Figure 18, the video / image encoders 20002 and 20003 of Figure 29, the processor and encoders 21000 to 21008 of Figure 21, and the XR device 2330 of Figure 23, etc., generate a bitstream including point cloud data according to the embodiment.

[0797] The file / segment encapsulation unit 10003 in FIG. 1, the file / segment encapsulation unit 20004 in FIG. 20, the file / segment encapsulation unit 21009 in FIG. 21, and the XR device in FIG. 23 format the bitstream into the file structure in FIGS. 24 and 25.

[0798] Similarly, the file / segment decapsulator 10007 of the receiving device 10005 in Fig. 1, the file / segment decapsulators 20005, 21009, and 22000 in Fig. 20 to Fig. 23, and the XR device 2330 in Fig. 23 receive and decapsulate the file to parse the bitstream. The bitstream is decoded by the point cloud video decoder 101008 in Fig. 1, the decoders in Fig. 16 and Fig. 17, the receiving device in Fig. 19, the video / image decoders 20006, 21007, 21008, 22001, and 22002 in Fig. 20 to Fig. 23, and the XR device 2330 in Fig. 23, and the point cloud data is restored.

[0799] Figures 40 and 41 show the container structure of point cloud data in the ISOBMFF file format.

[0800] Figures 40 and 41 show the structure of a container that transmits a point cloud based on multi-track.

[0801] The method / apparatus according to the embodiment transmits and receives point cloud data and additional data related to the point cloud data in a container file based on multiple tracks.

[0802] The first track 40000 is a feature track and contains feature data 40040 encoded as in Figures 1, 4, 15, 18, etc.

[0803] The second track 40010 is an occupation track and contains geometry data 40050 encoded as in Figures 1, 4, 15, 18, etc.

[0804] The third track 40020 is a geometry track and contains occupancy data 40060 encoded as in Figures 1, 4, 15, 18, etc.

[0805] The fourth track 40030 is a v-pcc (v3c) track and includes an atlas bitstream 40070 that includes data related to the point cloud data.

[0806] Each track consists of sample entries and samples. A sample is a unit corresponding to a frame. To decode the Nth frame, the sample or sample entry corresponding to the Nth frame is required. A sample entry contains information describing the sample.

[0807] FIG. 41 is a detailed structural diagram of FIG.

[0808] The v3c track 41000 corresponds to the fourth track 40030. The data contained in the v3c track 41000 has a data container format called a box. The v3c track 41000 includes reference information regarding the V3C component tracks 41010 to 41030.

[0809] The receiving method / apparatus according to the embodiment receives a container (also called a file) containing point cloud data as shown in Figure 41, parses the V3C track, and decodes and restores the occupancy data, geometry data, and attribute data based on the reference information included in the V3C track.

[0810] The occupancy track 41010 corresponds to the second track 40010 and contains occupancy data, the geometry track 41020 corresponds to the third track 40020 and contains geometry data, and the attribute track 41030 corresponds to the first track 40000 and contains attribute data.

[0811] The syntax of the data structure contained in the files shown in FIGS. 40 and 41 will be explained in detail below.

[0812] Volumetric visual track

[0813] Each volumetric visual scene is represented by a unique volumetric visual track.

[0814] An ISOBMFF file contains multiple scenes, which results in multiple volumetric visual tracks being present within the file.

[0815] A volumetric visual track is identified by the volumetric visual media handler type 'volv' in the media box's handler box. The volumetric visual header is defined as follows:

[0816] Volumetric visual Media header

[0817] Box Type: 'vvhd'

[0818] Container:MediaInformationBox

[0819] Mandatory: Yes

[0820] Quantity: Exactly one

[0821] A volumetric visual track uses the VolumetricVisualMediaHeaderBox of the MediaInformationBox.

[0822] aligned(8) class VolumetricVisualMediaHeaderBox

[0823] extends FullBox('vvhd', version=0, 1){

[0824] }

[0825] Version is an integer that indicates the version of this box.

[0826] Volumetric visual sample entry

[0827] Volumetric visual tracks use volumetric visual sample entries (VolumetricVisualSampleEntry).

[0828] class VolumetricVisualSampleEntry(codingname)

[0829] extends SampleEntry(codingname){

[0830] unsigned int(8)

[32] compressor_name;

[0831] }

[0832] Compressor Name (compressor_name): A name for informative purposes. Forms a fixed 32-byte field. The first byte is set to the number of bytes to be displayed, followed by the number of bytes of displayable data encoded using UTF-8. The size byte is padded to complete 32 bytes. This field may be set to 0.

[0833] Volumetric visual samples

[0834] The format of the volumetric visual samples is defined by a coding system according to the embodiment.

[0835] V-PCC unit header box

[0836] This box is present in both the V-PCC track (in the sample entry) and in all video-coded V-PCC component tracks (in the scheme information). This box contains the V-PCC unit headers for the data carried by each track.

[0837] aligned(8) class VPCCUnitHeaderBox

[0838] extends FullBox('vunt', version=0, 0){

[0839] vpcc_unit_header() unit_header;

[0840] }

[0841] This box contains the V-PCC unit header (vpcc_unit_header()) as shown above.

[0842] V-PCC decoder configuration record

[0843] This record contains a version field. This version is version 1. Incompatible changes to the record are identified by a change in the version number. A reader / decoder according to an embodiment may not decode a record or stream where this version is an unrecognized version number.

[0844] The array for V-PCC parameter sets contains the V-PCC parameter sets as described above.

[0845] The atlas_setupUnit array contains the atlas parameter set that is constant for the stream named by the sample entry whose decoder configuration record is present with the atlas stream SEI message.

[0846] aligned(8) class VPCCDecoderConfigurationRecord{

[0847] unsigned int(8) configurationVersion=1;

[0848] unsigned int(3) sampleStreamSizeMinusOne;

[0849] unsigned int(5) numOfVPCCParameterSets;

[0850] for(i=0;i <numOfVPCCParameterSets;i++){

[0851] sample_stream_vpcc_unit VPCCParameterSet;

[0852] }

[0853] unsigned int(8) numOfAtlasSetupUnits;

[0854] for (i=0;i <numOfAtlasSetupUnits;i++){

[0855] sample_stream_vpcc_unit atlas_setupUnit;

[0856] }

[0857] }

[0858] ConfigurationVersion is a version field. Changes that are incompatible with this record are identified by changes in the version number.

[0859] Sample Stream Size (sampleStreamSizeMinusOne): This value plus one indicates the precision in bytes of the ssvu_vpcc_unit_size element in all sample stream V-PCC units in the V-PCC samples in this configuration record or the stream to which this configuration record applies.

[0860] Number of V-PCC Parameter Sets (numOfVPCCParameterSets): Indicates the number of V-PCC parameter sets (VPS) signaled in the decoder configuration record.

[0861] A V-PCC parameter set is a sample_stream_vpcc_unit() instance of a V-PCC unit of type VPCC_VPS. A V-PCC unit contains a V-PCC parameter set (vpcc_parameter_set()).

[0862] Number of Atlas Setup Units (numOfAtlasSetupUnits): Indicates the number of setup arrays for the Atlas stream signaled in this configuration record.

[0863] Atlas setup unit (Atlas_setupUnit): A sample_stream_vpcc_unit() instance that contains an Atlas sequence parameter set, an Atlas frame parameter set, or an SEI Atlas NAL unit. See, for example, the description in ISO / IEC 23090-5.

[0864] Also according to the embodiment, the V-PCC decoder configuration record is defined as follows:

[0865] aligned(8) class VPCCDecoderConfigurationRecord{

[0866] unsigned int(8) configurationVersion=1;

[0867] unsigned int(3) sampleStreamSizeMinusOne;

[0868] bit(2) reserved=1;

[0869] unsigned int(3) lengthSizeMinusOne;

[0870] unsigned int(5) numOVPCCParameterSets;

[0871] for(i=0;i<numOVPCCParameterSets;i++){

[0872] sample_stream_vpcc_unit VPCCParameterSet;

[0873] }

[0874] unsigned int(8) numOfSetupUnitArrays;

[0875] for(j=0;j<numOfSetupUnitArrays;j++){

[0876] bit(1) array_completeness;

[0877] bit(1) reserved=0;

[0878] unsigned int(6) NAL_unit_type;

[0879] unsigned int(8) numNALUnits;

[0880] for(i=0;i<numNALUnits;i++){

[0881] sample_stream_nal_unit setupUnit;

[0882] }

[0883] }

[0884] Configuration information (configurationVersion): A version field. Incompatible changes to this record are identified by changing the version number.

[0885] Length Size (lengthSizeMinusOne): This value plus one indicates the precision in bytes of the ssnu_nal_unit_size element in all sample stream NAL units in this configuration record or in the V-PCC samples in the stream to which this configuration record applies.

[0886] Sample Stream Size (sampleStreamSizeMinusOne): This value plus one indicates the precision in bytes of the ssvu_vpcc_unit_size element in all sample stream V-PCC units signaled in this configuration record.

[0887] Number of V-PCC Parameter Sets (numOfVPCCParameterSets): Indicates the number of V-PCC parameter sets (VPS) signaled in this configuration record.

[0888] A V-PCC parameter set is a sample_stream_vpcc_unit() instance for a V-PCC unit of type VPCC_VPS.

[0889] Number of Setup Unit Arrays (numOfSetupUnitArrays): The number of arrays of the atlas NAL unit of the indicated type.

[0890] Array completeness (array_completeness): A value of 1 indicates that all atlas NAL units of the given type are present in the following array and not in the stream. A value of 0 indicates that additional atlas NAL units of the indicated type may be present in the stream. The default and allowed values ​​are affected by the sample entry name.

[0891] NAL unit type (NAL_unit_type): Indicates the type of the atlas NAL unit in the following array. This type uses the values ​​defined in ISO / IEC 23090-5. This value indicates a NAL_ASPS, NAL_PREFIX_SEI, or NAL_SUFFIX_SEI atlas NAL unit.

[0892] Number of NAL units (numNALUnits): The number of Atlas NAL units of the indicated type contained in the configuration record for the stream to which this configuration record applies. The SEI array contains SEI messages that are purely descriptive in nature and provide information about the stream as a whole. An example of such an SEI is the user-data SEI.

[0893] Setup Unit (setupUnit): A sample_stream_nal_unit() instance that contains an atlas sequence parameter set or an atlas frame parameter set or a descriptive SEI atlas NAL unit.

[0894] V-PCC atlas parameter set sample group

[0895] The grouping type 'vaps' for sample grouping indicates that the atlas parameter set transmitted in the sample group of the sample in the V-PCC track is placed in the atlas parameter set. If a SampleToGroupBox with the same grouping type as 'vaps' exists, a SampleGroupDescriptionBox with the same grouping type exists and contains the ID of this group to which the sample belongs.

[0896] A V-PCC track contains at most one sample-to-group box with grouping type equal to 'vaps'.

[0897] aligned(8) class VPCCAtlasParamSampleGroupDescriptionEntry() extends SampleGroupDescriptionEntry('vaps'){

[0898] unsigned int(8) numOfAtlasParameterSets;

[0899] for(i=0;i <numOfAtlasParameterSets;i++){

[0900] sample_stream_vpcc_unit atlasParameterSet;

[0901] }

[0902] }

[0903] Number of Atlas Parameter Sets (numOfAtlasParameterSets): Indicates the number of Atlas Parameter Sets signaled within the sample group description.

[0904] The atlas parameter set is a sample_stream_vpcc_unit() instance that contains the atlas sequence parameter set and atlas frame parameter set associated with this group of samples.

[0905] The Atlas Parameter Sample Group Description Entry is as follows:

[0906] aligned(8) class VPCCAtlasParamSampleGroupDescriptionEntry() extends SampleGroupDescriptionEntry('vaps'){

[0907] unsigned int(3) lengthSizeMinusOne;

[0908] unsigned int(5) numOfAtlasParameterSets;

[0909] for(i=0;i <numOfAtlasParameterSets;i++){

[0910] sample_stream_nal_unit atlasParameterSetNALUnit;

[0911] }

[0912] }

[0913] Length Size (lengthSizeMinusOne): This value plus one indicates the precision in bytes of the ssnu_nal_unit_size element in all sample stream NAL units signaled in this sample group description.

[0914] Atlas Parameter Set NAL Unit (atlasParameterSetNALUnit): A sample_stream_nal_unit() instance that contains the atlas sequence parameter set and / or atlas frame parameter set associated with this group of samples.

[0915] V-PCC SEI sample group

[0916] The 'vsei' grouping type for a sample grouping indicates the placement of samples in a V-PCC track in the SEI information conveyed within this sample group. If a Sample To Group box with a grouping type equal to 'vsei' exists, a Sample Group Description Box with the same grouping type also exists and contains the ID of the group to which the sample belongs.

[0917] A V-PCC track contains at most one sample-to-group box with grouping type equal to 'vsei'.

[0918] aligned(8) class VPCCSEISampleGroupDescriptionEntry() extends SampleGroupDescriptionEntry('vsei'){

[0919] unsigned int(8) numOfSEIs;

[0920] for(i=0;i <numOfSEISets;i++){

[0921] sample_stream_vpcc_unit sei;

[0922] }

[0923] }

[0924] Number of SEIs (numOfSEIs): Indicates the number of V-PCC SEIs signaled within the sample group description.

[0925] SEI is a sample_stream_vpcc_unit() instance that contains the SEI information associated with this group of samples.

[0926] A V-PCC SEI sample group description entry is as follows:

[0927] aligned(8) class VPCCSEISampleGroupDescriptionEntry () extends SampleGroupDescriptionEntry('vsei'){

[0928] unsigned int(3) lengthSizeMinusOne;

[0929] unsigned int(5) numOfSEIs;

[0930] for(i=0;i <numOfSEIs;i++){

[0931] sample_stream_nal_unit seiNALUnit;

[0932] }

[0933] }

[0934] Length Size (lengthSizeMinusOne): This value plus one indicates the precision in bytes of the ssnu_nal_unit_size element in all sample stream NAL units signaled in this sample group description.

[0935] SEI NAL unit (seiNALUnit): A sample_stream_nal_unit() instance that contains the SEI information associated with this group of samples.

[0936] Multi track container of V-PCC bitstream

[0937] The general layout of a multi-track ISOBMFF V-PCC container. V-PCC units in a V-PCC elementary stream are mapped to individual tracks in the container file based on their type. There can be two types of tracks in a multi-track ISOBMFF V-PCC container: V-PCC tracks and V-PCC component tracks.

[0938] V-PCC tracks (or V3C tracks) 40030, 41000 are tracks that convey volumetric visual information in a V-PCC bitstream, including an atlas sub-bitstream and a sequence parameter set.

[0939] A V-PCC component track is a video scheme track that conveys 2D video coded data for the occupancy map, geometry, and attribute sub-streams of a V-PCC bitstream. Additionally, the following conditions are met for a V-PCC component track:

[0940] a) Within the sample entry, a new box is inserted that describes the role of the video stream contained in this track within the V-PCC system.

[0941] b) To generate the membership of V-PCC component tracks within a particular point cloud represented by a V-PCC track, track references are introduced from the V-PCC track to the V-PCC component tracks.

[0942] c) The track-header flag is set to 0 to indicate that this track does not directly contribute to the overall layup of the movie, but rather to the V-PCC system.

[0943] According to an embodiment, the atlas bitstream and signaling information (referred to as parameters, metadata, etc.) describing the point cloud data are contained in a data structure called a box. According to an embodiment, the method / apparatus transmits the atlas bitstream and parameter information in a v-pcc track (or v3c track) based on a multi-track basis. Furthermore, according to an embodiment, the method / apparatus transmits the atlas bitstream and parameter information in a sample entry of a v-pcc track (or v3c track).

[0944] In addition, the method / apparatus according to the embodiment transmits the atlas bitstream and parameter information according to the embodiment in a V-PCC elementary stream track based on a single track, and further transmits the atlas bitstream and parameter information in a sample entry or sample of the V-PCC elementary stream track.

[0945] Tracks belonging to the same V-PCC sequence are aligned by time. Samples and V-PCC tracks that contribute to the same point cloud frame from different video-coded V-PCC component tracks have the same presentation time. The V-PCC atlas sequence parameter set and atlas frame parameter set used for the samples have a decoding time equal to or earlier than the composition time of the point cloud frame. Furthermore, all tracks belonging to the same V-PCC sequence have the same implicit or explicit edit list.

[0946] Note: Synchronization between elementary streams within a component track is handled by the ISOBMFF track timing structures (stts, ctts, cslg) or by equivalent mechanisms within a movie fragment.

[0947] Based on this layout, a V-PCC ISOBMFF container contains the following (see Figure 24):

[0948] - A V-PCC track containing samples and V-PCC parameter sets in sample entries that carry the payload of V-PCC units (unit type VPCC_VPS) and atlas V-PCC units (unit type VPCC_AD). This track also contains track references to other tracks that carry the payload of video compressed V-PCC units, such as unit types VPCC_OVD, VPCC_GVD, and VPCC_AVD.

[0949] - A video scheme track with samples containing access units of video coded elementary streams for occupancy map data that are the payload of V-PCC units of type VPCC_OVD.

[0950] - One or more video scheme tracks with samples containing access units of video coded elementary streams of geometry data that are the payload of V-PCC units of type VPCC_GVD.

[0951] - Zero or more video scheme tracks with samples containing access units of a video coded elementary stream of attribute data that is the payload of a V-PCC unit of type VPCC_AVD.

[0952] V-PCC tracks:

[0953] V-PCC Track Sample Entry:

[0954] Sample Entry Type: 'vpc1', 'vpcg'

[0955] Container: SampleDescriptionBox

[0956] Mandatory: The 'vpc1' or 'vpcg' sample entry is mandatory.

[0957] Quantity: There is one or more sample entries.

[0958] V-PCC tracks use VPCCSampleEntry, which extends VolumetricVisualSampleEntry. The sample entry type is 'vpc1' or 'vpcg'.

[0959] The V-PCC sample entry contains a V-PCC configuration box (VPCCConfigurationBox), which contains a decoder configuration record (VPCCDecoderConfigurationRecord).

[0960] In the 'vpc1' sample entry, all atlas sequence parameter sets, atlas frame parameter sets or V-PCC SEIs are in the setupUnit array.

[0961] In the 'vpcg' sample entry, the atlas sequence parameter set, atlas frame parameter set, and V-PCC SEI are in this array or stream.

[0962] An optional BitRateBox is present in the V-PCC Volumetric Sample Entry to signal the bitrate information of the V-PCC track.

[0963] Volumetric Sequences:

[0964] class VPCCConfigurationBox extends box('vpcc'){

[0965] VPCCDecoderConfigurationRecord() VPCCConfig;

[0966] }

[0967] aligned(8) class VPCCSampleEntry() extends VolumetricVisualSampleEntry('vpc1'){

[0968] VPCCConfigurationBox config;

[0969] VPCCUnitHeaderBox unit_header;

[0970] VPCCBoundingInformationBox();

[0971] }

[0972] A point cloud data transmission method according to an embodiment includes the steps of: encoding the point cloud data; encapsulating the point cloud data; and transmitting the point cloud data.

[0973] Encapsulating the point cloud data according to an embodiment includes generating one or more tracks that include the point cloud data.

[0974] According to an embodiment, one or more tracks contain point cloud data, such as attribute data, occupancy data, geometry data, and / or parameters (metadata or signaling information) related thereto. Specifically, a track contains sample entries describing samples and / or samples. Multiple tracks may also be referred to as first tracks, second tracks, etc.

[0975] A point cloud data receiving apparatus according to an embodiment includes: a receiver for receiving point cloud data; a decapsulating unit for decapsulating the point cloud data; and a decoder for decoding the point cloud data.

[0976] The decapsulator, according to an embodiment, parses one or more tracks containing point cloud data.

[0977] FIG. 42 shows a V-PCC sample entry according to an embodiment.

[0978] FIG. 42 is a structural diagram of a sample entry included in the V-PCC track (or V3C track) 40030 in FIG. 40 and the V3C track 41000 in FIG.

[0979] Figure 42 shows an example of a V-PCC sample entry structure according to embodiments described herein. The sample entry includes a V-PCC parameter set (VPS) 42000, and optionally an atlas sequence parameter set (ASPS) 42010, an atlas frame parameter set (AFPS) 42020, and / or an SEI 42030.

[0980] The method / apparatus according to the embodiment stores point cloud data in a track of a file, and further stores parameters (or signaling information) related to the point cloud data in samples or sample entries of the track for transmission / reception.

[0981] The V-PCC bitstream of FIG. 42 is generated and parsed by the embodiment for generating and parsing a V-PCC bitstream of FIG.

[0982] The V-PCC bitstream includes a sample stream V-PCC header, a sample stream header, a V-PCC unit header box, and a sample stream V-PCC unit.

[0983] The V-PCC bitstream corresponds to the V-PCC bitstream described in Figures 26 and 27, or is an example of an additional extension.

[0984] V-PCC track sample format

[0985] Each sample in a V-PCC track corresponds to a single point cloud frame. Samples corresponding to this frame in the various component tracks have the same component duration as the V-PCC track sample. Each V-PCC sample contains one or more atlas NAL units.

[0986] aligned(8) class VPCCSample{

[0987] unsigned int PointCloudPictureLength=sample_size; / / Means the sample size from SampleSizeBox.

[0988] for(i=0;i <PointCloudPictureLength;){

[0989] sample_stream_nal_unit nalUnit

[0990] i+=(VPCCDecoderConfigurationRecord.lengthSizeMinusOne+1)+nalUnit.ssnu_nal_unit_size;

[0991] }

[0992] }

[0993] aligned(8) class VPCCSample

[0994] {

[0995] unsigned int PictureLength = sample_size; / / This means the size of the samples from the SampleSizeBox.

[0996] for (i = 0; i < PictureLength;) / / Signaling continues until the end of the picture

[0997] {

[0998] unsigned int((VPCCDecoderConfigurationRecord.LengthSizeMinusOne + 1) * 8)

[0999] NALUnitLength;

[1000] bit(NALUnitLength * 8) NALUnit;

[1001] i += (VPCCDecoderConfigurationRecord.LengthSizeMinusOne + 1) + NALUnitLength;

[1002] }

[1003] }

[1004] V-PCC Decoder Configuration Record (VPCCDecoderConfigurationRecord): Indicates the decoder configuration record within the matching V-PCC sample entry.

[1005] NAL Unit (nalUnit): Contains a single atlas NAL unit in the sample stream NAL unit format.

[1006] NAL Unit Length (NALUnitLength): Indicates the byte size within the following NAL unit.

[1007] NAL Unit (NALUnit): Contains a single atlas NAL unit.

[1008] V-PCC track sync sample:

[1009] The synchronization samples (arbitrary access points) within a V-PCC track are V-PCC IRAP coded patch data access units. The atlas parameter set is repeated at the synchronization sample for arbitrary access, if necessary.

[1010] Video-encoded V-PCC component tracks:

[1011] The transmission of video tracks coded using MPEG specific codecs follows the specifications of ISO BMFF. For example, the transmission of AVC and HEVC coded video can refer to ISO / IEC 14496-15. ISOBMFF can also provide an extension mechanism if other codec types are required.

[1012] Since it is not considered meaningful to display frames decoded from a feature, geometry, or occupancy map track without reconstructing the point cloud at the player, limited video scheme types can be defined for such video-coded tracks.

[1013] Restricted video scheme:

[1014] A V-PCC component video track is represented in the file as restricted video and is identified by the 'pccv' value in the scheme_type field of the SchemeTypeBox in the RestrictedSchemeInfoBox of the restricted video sample entry.

[1015] There is no restriction on the video codec used to encode the feature, geometry and occupancy map V-PCC components. Moreover, such components may be encoded using different video codecs.

[1016] Scheme information:

[1017] The SchemeInformationBox is present and contains a VPCCUnitHeaderBox.

[1018] Referencing V-PCC component tracks:

[1019] To link a V-PCC track to a component video track, three TrackReferenceTypeBoxes are added for each component to a TrackReferenceBox within the V-PCC track's TrackBox. The TrackReferenceTypeBox contains an array of track_IDs that specify the video tracks for the V-PCC track reference. The reference_type of the TrackReferenceTypeBox identifies the type of component, such as occupancy map, geometry, attribute, or occupancy map. The track reference types are:

[1020] 'pcco': The referenced track contains a video-coded occupancy map V-PCC component

[1021] 'pccg': The referenced track contains a video-coded geometry V-PCC component

[1022] 'pcca': The referenced track contains a video-coded attribute V-PCC component

[1023] The type of V-PCC component carried by the referenced restricted video track and signaled in the track's RestrictedSchemeInfoBox is matched to the reference type of the track reference from the V-PCC track.

[1024] FIG. 43 illustrates track substitution and grouping according to an embodiment.

[1025] 43 shows an example in which inter-track substitution or grouping of the ISOBMFF file structure is applied. Substitution or grouping is performed by an encapsulation unit 20004 or the like on the transmitting side, and parsing is performed by a decapsulation unit 20005 or the like on the receiving side.

[1026] Track alternatives and track grouping:

[1027] V-PCC component tracks with the same alternate_group value are different coded versions of the same V-PCC component. Volumetric visual scenes are coded alternatively. In this case, all V-PCC tracks that are interchangeable with each other have the same alternate_group value in the TrackHeaderBox.

[1028] Similarly, if a 2D video track representing one of the V-PCC components is coded into alternatives, there may be a track reference to one of the alternatives from such an alternative and alternative group.

[1029] Figure 43 shows V-PCC component tracks that make up V-PCC content based on the file structure. If they have the same atlas group ID, the ID may be 10, 11, or 12. The second and fifth tracks, which are feature videos, can be used interchangeably, the third and sixth tracks can be used interchangeably as geometry videos, and the fourth and seventh tracks can be used interchangeably as occupation videos.

[1030] Single track container of V-PCC Bitstream:

[1031] A single-track encapsulation of V-PCC data requires the V-PCC encoded elementary bitstream to be represented by a single-track declaration.

[1032] Single-track encapsulation of PCC data is used in the case of simple ISOBMFF encapsulation of V-PCC encoded bitstreams. Such bitstreams are immediately stored in a single track without further processing. The V-PCC unit header data structure is present in the bitstream. The single-track container for V-PCC data is provided to the media workflow for further processing (e.g., multi-track file generation, transcoding, DASH segmentation, etc.).

[1033] ISOBMFF files containing single-track encapsulated V-PCC data include 'pcst' in the compatible_brands[] list of the FileTypeBox.

[1034] V-PCC elementary stream track:

[1035] Sample Entry Type:'vpe1','vpeg'

[1036] Container:SampleDescriptionBox

[1037] Mandatory:a 'vpe1' or 'vpeg' sample entry is mandatory

[1038] Quantity:One or more sample entries may be present

[1039] A V-PCC elementary stream track uses a VolumetricVisualSampleEntry with a sample entry type 'vpe1' or 'vpeg'.

[1040] The V-PCC elementary stream sample entry includes a VPCCConfigurationBox.

[1041] In the 'vpe1' sample entry, all atlas sequence parameter sets, atlas frame parameter sets, and SEIs are in the setupUnit array. In the 'vpeg' sample entry, the atlas sequence parameter sets, atlas frame parameter sets, and SEIs are in this array or stream.

[1042] Volumetric Sequences:

[1043] class VPCCConfigurationBox extends box('vpcc'){

[1044] VPCCDecoderConfigurationRecord() VPCCConfig;

[1045] }

[1046] aligned(8) class VPCElementaryStreamSampleEntry() extends VolumetricVisualSampleEntry('vpe1'){

[1047] VPCCConfigurationBox config;

[1048] VPCCBoundingInformationBox 3d_bb;

[1049] }

[1050] V-PCC elementary stream sample format:

[1051] A V-PCC elementary stream sample is composed of one or more V-PCC units belonging to the same presentation time. Each sample has a unique presentation time, size, and duration. A sample may, for example, be a synchronization sample or be decoding dependent on other V-PCC elementary stream samples.

[1052] V-PCC elementary stream sync sample:

[1053] V-PCC elementary stream synchronization samples satisfy the following conditions:

[1054] -Independently decodable.

[1055] - Samples after the synchronization sample in decoding order have no decoding dependency on samples before the synchronization sample.

[1056] All samples after the synchronization sample in decoding order can be successfully decoded.

[1057] V-PCC elementary stream sub-sample:

[1058] A V-PCC elementary stream sub-sample is a V-PCC unit contained within a V-PCC elementary stream sample.

[1059] A V-PCC elementary stream track contains a SubSampleInformationBox in each SampleTableBox or TrackFragmentBox of MovieFragmentBoxes that lists the V-PCC elementary stream sub-samples.

[1060] The 32-bit unit header of the V-PCC unit representing the sub-sample is copied into the 32-bit codec_specific_parameters field of the sub-sample entry in the SubSampleInformationBox. The V-PCC unit type of each sub-sample is identified by parsing the codec_specific_parameters field of the sub-sample entry in the SubSampleInformationBox.

[1061] The information described below is transmitted as follows: For example, if the point cloud data is static, it is transmitted in the sample entry of the V3C track of a multi-track or the sample entry of the basic track of a single track. If the point cloud data is dynamic, the information is transmitted in a separate timed metadata track.

[1062] FIG. 44 shows a 3D bounding box information structure according to an embodiment.

[1063] Figure 44 shows a 3D bounding box included in the file according to Figures 40 and 41. For example, the 3D bounding box is transmitted by being included in multiple tracks and / or a single track of a file that is a container for a point cloud bitstream. The 3D bounding box is generated and parsed by the file / segment encapsulation units 20004 and 21009, file / segment decapsulation units 20005 and 22000, etc. according to the embodiment.

[1064] The 3D bounding box structure according to the embodiment provides a 3D bounding box for the point cloud data, including the X, Y, Z offsets of the 3D bounding box and the width, height, and depth of the 3D bounding box of the point cloud data.

[1065] Bounding Box X (bb_x), Bounding Box Y (bb_y), Bounding Box Z (bb_z): These represent the X, Y, and Z coordinate values ​​of the origin position of the 3D bounding box of the point cloud data in the coordinate system.

[1066] Bounding Box Delta X (bb_delta_x), Bounding Box Delta Y (bb_delta_y), Bounding Box Delta Z (bb_delta_z): Denote the extension of the 3D bounding box of the point cloud data in the coordinate system along the respective X, Y, and Z axes relative to the origin.

[1067] The file / segment encapsulation units 20004, 21009, file / segment decapsulation units 20005, 22000, etc. according to the embodiment generate and parse point cloud bounding boxes based on multi-track and / or single track of files that are containers of point cloud bitstreams.

[1068] According to an embodiment, the bounding information box (VPCCBoundingInformationBox) is present in a sample entry of a V-PCC track of a multi-track and / or a V-PCC elementary stream track of a single track. When present in a sample entry of a V-PCC track and / or a V-PCC elementary stream track, the bounding information box (VPCCBoundingInformationBox) provides overall bounding box information of the associated or transmitted point cloud data.

[1069] aligned(8) class VPCCBoundingInformationBox extends FullBox('vpbb',0,0) {

[1070] 3DBoundingBoxInfoStruct();

[1071] }

[1072] If a V-PCC track has an associated timed metadata track with sample entry type 'dybb', the timed metadata track contains dynamically changing 3D bounding box information for the point cloud data.

[1073] The associated timed metadata tracks include a 'cdsc' track that references a V-PCC track carrying the atlas stream.

[1074] aligned(8) class Dynamic3DBoundingBoxSampleEntry

[1075] extends MetaDataSampleEntry('dybb'){

[1076] VPCCBoundingInformationBox all_bb;

[1077] }

[1078] All Bounding Box (all_bb): Provides overall 3D bounding box information, including the extension of the overall 3D bounding box of the point cloud data in the coordinate system along each X, Y, and Z axis relative to the origin, and the X, Y, and Z coordinates of the origin position.

[1079] Sample syntax for sample entry type 'dybb' is as follows:

[1080] aligned(8) Dynamic3DBoundingBoxSample() {

[1081] VPCCBoundingInformationBox 3dBB;

[1082] }

[1083] 3D Bounding Box (3dBB): Provides 3D bounding box information signaled within samples.

[1084] The 3D spatial region structure (3DSpatialRegionStruct) provides information about the proximity of parts of volumetric visual data.

[1085] The 3D spatial region structure and 3D bounding box structure provide information for a spatial region of volumetric media and 3D bounding box information for the volumetric media, including the X, Y, and Z offsets of the spatial region and the width, height, and depth of the space within 3D space.

[1086] aligned(8) class 3DPoint() {

[1087] unsigned int(16) x;

[1088] unsigned int(16) y;

[1089] unsigned int(16) z;

[1090] }

[1091] aligned(8) class CuboidRegionStruct() {

[1092] unsigned int(16) cuboid_dx;

[1093] unsigned int(16) cuboid_dy;

[1094] unsigned int(16) cuboid_dz;

[1095] }

[1096] aligned(8) class 3DSpatialRegionStruct(dimensions_included_flag) {

[1097] unsigned int(16) 3D_region_id;

[1098] 3DPoint anchor;

[1099] if(dimensions_included_flag) {

[1100] CuboidRegionStruct();

[1101] }

[1102] }

[1103] aligned(8) class 3DBoundingBoxStruct() {

[1104] unsigned int(16) bb_dx;

[1105] unsigned int(16) bb_dy;

[1106] unsigned int(16) bb_dz;

[1107] }

[1108] 3D Region ID (3d_region_id): An identifier for the spatial region.

[1109] X, Y, Z: Denotes the X, Y, and Z coordinate values ​​of each of the 3D points in the coordinate system.

[1110] Cuboid Delta X, Cuboid Delta Y, Cuboid Delta Z (bb_dx, bb_dy, and bb_dz): Denote the extension of the 3D bounding box of the entire volumetric media in the coordinate system along the respective X, Y, and Z axes relative to the origin (0,0,0).

[1111] Dimensions included flag (dimensions_included_flag): Indicates whether the dimensions of the spatial domain are signaled.

[1112] A dimensions_included_flag of 0 indicates that the dimensions are not signaled, but that a previous instance of 3DSpatialRegionStruct with the same 3D region ID (3d_region_id) signaled the dimensions.

[1113] As described above, the transmitting method / apparatus according to the embodiment transmits 3D bounding box information included in a V3C track, and the receiving method / apparatus according to the embodiment can efficiently obtain the 3D bounding box information based on the V3C track.

[1114] That is, the embodiment proposes a file container for efficient transmission and decoding of a bitstream, and spatial or partial access to point cloud data is possible by transmitting 3D bounding box position and / or size information in a track within the file container. Also, there is no need to decode the entire bitstream, and desired spatial data can be quickly obtained by decapsulating and parsing the track within the file.

[1115] The method / apparatus according to the embodiment conveys 3D bounding box information according to the embodiment to sample entries and / or samples of a track according to the embodiment.

[1116] When a sample (e.g., a sample syntax with a sample entry type of 'dybb') according to the embodiment conveys 3D bounding box information, information about the bounding box that changes over time can be signaled. That is, by signaling the bounding box information at the time when the corresponding sample is decoded, it is possible to support fine spatial approximation at the time.

[1117] FIG. 45 shows an outline of a structure for encapsulating non-timed V-PCC data according to an embodiment.

[1118] FIG. 45 shows a structure for transmitting non-timed V-PCC data when the point cloud data related devices of FIGS. 20-22 process non-timed point cloud data.

[1119] The system included in the point cloud data transmitting / receiving method / apparatus and transmitting / receiving apparatus according to the embodiment encapsulates non-timed V-PCC data as shown in FIG. 45 and transmits / receives it.

[1120] When the point cloud data according to the embodiment is an image, the point cloud video encoder 10002 of FIG. 1 (or the encoder of FIG. 4, the encoder of FIG. 15, the transmitting device of FIG. 18, the processor 20001 of FIG. 20, the image encoder 20003, the processor of FIG. 21, and the image encoder 21008) encodes the image, the file / segment encapsulation unit 10003 (or the file / segment encapsulation unit 20004 of FIG. 20, the file / segment encapsulation unit 21009 of FIG. 21) stores the image and image-related information in a container (item) such as that of FIG. 45, and the transmitter 10004 transmits the container.

[1121] Similarly, the receiver in Fig. 1 receives the container in Fig. 45, and the file / segment decapsulator 10007 (or the file / segment decapsulator 20005 or the file / segment decapsulator 22000 in Fig. 20) parses the container. The point cloud video decoder 10008 in Fig. 1 (or the decoder in Fig. 16, the decoder in Fig. 17, the receiving device or the image decoder 20006 or the image decoder 22002 in Fig. 19) decodes the image included in the item and provides it to the user.

[1122] The image according to the embodiment is a still image. The method / apparatus according to the embodiment transmits and receives point cloud data for the image. The method / apparatus according to the embodiment stores the image in an item based on the data container structure shown in FIG. 45 and transmits and receives the image. In addition, attribute information related to the image can be stored in image properties.

[1123] Non-timed V-PCC data is stored within a file as an image item. Two new item types are defined to encapsulate non-timed V-PCC data: V-PCC Item and V-PCC Unit Item.

[1124] A new Handler Type 4CC code 'vpcc' is defined and stored in a HandlerBox of a MetaBox to indicate the presence of V-PCC items, V-PCC unit items, and other V-PCC-encoded content representation information.

[1125] V-PCC Items (V-PCC Items, 45000): A V-PCC item is an item that indicates an independently decodable V-PCC access unit. The item type 'vpci' is defined to identify a V-PCC item. A V-PCC item stores the V-PCC unit payload of an atlas sub-bitstream. If a Primary Item Box (PrimaryItemBox) exists, the item_id in this box is set to indicate the V-PCC item.

[1126] V-PCC Unit Item (45010): A V-PCC Unit Item is an item that represents V-PCC unit data. The V-PCC Unit Item of a video data unit stores the V-PCC unit payload for occupancy, geometry, and attributes. A V-PCC Unit Item stores only one V-PCC access unit related data.

[1127] The item type for a V-PCC unit item is set by the codec used to encode the corresponding video data unit. A V-PCC unit item is associated with the corresponding V-PCC unit header item property and codec specific configuration item property. Since it is not meaningful to display them independently, V-PCC unit items are displayed as hidden items.

[1128] The following three item reference types are used to indicate the relationship between V-PCC Items and V-PCC Unit Items: An item reference is defined from a V-PCC Item to its associated V-PCC Unit Item.

[1129] 'pcco': The referenced V-PCC unit item contains a proprietary video data unit.

[1130] 'pccg': The referenced V-PCC unit item contains a geometry video data unit.

[1131] 'pcca': The referenced V-PCC unit item contains a particular video data unit.

[1132] V-PCC configuration item property (45020)

[1133] box Types:'vpcp'

[1134] Property type:Descriptive item property

[1135] Container:ItemPropertyContainerBox

[1136] Mandatory(per item):Yes, for a V-PCC item of type 'vpci'

[1137] Quantity(per item):One or more for a V-PCC item of type 'vpci'

[1138] The V-PCC Configuration Item Property box type is 'vpcp' and the property type is Descriptive Item Property. The container is an Item Property Container Box. It is mandatory for V-PCC types of type 'vpci' per item. There may be one or more for V-PCC items of type 'vpci' per item.

[1139] The V-PCC parameter set is stored as this descriptive item property and is associated with the V-PCC item.

[1140] aligned(8) class vpcc_unit_payload_struct(){

[1141] unsigned int(16) vpcc_unit_payload_size;

[1142] vpcc_unit_payload();

[1143] }

[1144] V-PCC unit payload size (vpcc_unit_payload_size): Indicates the byte size of vpcc_unit_payload().

[1145] aligned(8) class VPCCConfigurationProperty extends ItemProperty('vpcc'){

[1146] vpcc_unit_payload_struct()[];

[1147] }

[1148] The V-PCC unit payload (vpcc_unit_payload()) contains a V-PCC unit of type VPCC_VPS.

[1149] V-PCC unit header item property (45030)

[1150] box Types:'vunt'

[1151] Property type:Descriptive item property

[1152] Container:ItemPropertyContainerBox

[1153] Mandatory (per item):Yes, for a V-PCC item of type 'vpcI' and for a V-PCC unit item

[1154] Quantity (per item): One

[1155] The V-PCC Unit Header Item Property box type is 'vunt', the property type is Descriptive Item Property, and the container is the Item Property Container Box. Required for V-PCC Unit Items per Item and V-PCC Items of type 'vpci'. There is one per item.

[1156] V-PCC unit headers are stored as descriptive item properties and are associated with V-PCC items and V-PCC unit items.

[1157] aligned(8) class VPCCUnitHeaderProperty() extends ItemFullProperty('vunt', version=0, 0){

[1158] vpcc_unit_header();

[1159] }

[1160] Based on the structure of FIG. 45, a method / apparatus / system according to an embodiment transmits non-time point cloud data.

[1161] Carriage of non-timed video-based point cloud compression data

[1162] Non-timed V-PCC data in a file is stored as an image item. A new handle, type 4 CC code 'vpcc', is defined and stored in the handler box of the meta box to indicate the presence of V-PCC items, V-PCC unit items and other V-PCC encoded content representation information.

[1163] In this embodiment, an item refers to an image, i.e., a piece of immobile data, such as an image.

[1164] The method / apparatus according to the embodiment generates and transmits data according to the embodiment based on a structure for encapsulating non-timed V-PCC data, as shown in FIG.

[1165] V-PCC Items

[1166] A V-PCC item is an item that represents an independently decodable V-PCC access unit. A new item type, 4CC code 'vpci', is defined to identify a V-PCC item. A V-PCC item stores the V-PCC unit payload of an atlas sub-bitstream.

[1167] If a PrimaryItemBox exists, the item_id in this box is set to indicate the V-PCC item.

[1168] V-PCC Unit Item

[1169] A V-PCC unit item is an item that indicates V-PCC unit data.

[1170] A V-PCC unit item stores the occupancy, geometry, and characteristics of a V-PCC unit payload of a video data unit. A V-PCC unit item contains data related to a single V-PCC access unit.

[1171] The Item Type 4CC code for a V-PCC Unit Item is set based on the codec used to encode the corresponding video data unit. A V-PCC Unit Item is associated with the corresponding V-PCC Unit Header Item properties and codec-specific Configuration Item properties.

[1172] V-PCC unit items are marked as hidden items because it may not make sense to display them separately.

[1173] To indicate the relationship between a V-PCC item and a V-PCC unit, three new item reference types, the 4CC codes 'pcco', 'pccg' and 'pcca', are defined. An item reference is defined from a V-PCC item to the associated V-PCC unit item. The 4CC codes for item reference times are as follows:

[1174] 'pcco': Referenced V-PCC unit item containing the dedicated video data unit

[1175] 'pccg': The referenced V-PCC unit item contains a geometry video data unit.

[1176] 'pcca': The referenced V-PCC unit item contains a particular video data unit.

[1177] V-PCC-related item properties

[1178] Descriptive item properties are defined to convey V-PCC parameter set information and V-PCC unit header information, respectively.

[1179] V-PCC configuration item property

[1180] Box Types:'vpcp'

[1181] Property type:Descriptive item property

[1182] Container:ItemPropertyContainerBox

[1183] Mandatory(per item):Yes, for a V-PCC item of type 'vpci'

[1184] Quantity(per item):One or more for a V-PCC item of type 'vpci'

[1185] V-PCC parameter sets are stored as descriptive item properties and are associated with V-PCC items.

[1186] essential is set to 1 for the 'vpcp' item property.

[1187] aligned(8) class vpcc_unit_payload_struct(){

[1188] unsigned int(16) vpcc_unit_payload_size;

[1189] vpcc_unit_payload();

[1190] }

[1191] aligned(8) class VPCCConfigurationProperty

[1192] extends ItemProperty('vpcc'){

[1193] vpcc_unit_payload_struct()[];

[1194] }

[1195] V-PCC unit payload size (vpcc_unit_payload_size): Indicates the size in bytes of vpcc_unit_payload().

[1196] V-PCC unit header item property

[1197] box Types:'vunt'

[1198] Property type:Descriptive item property

[1199] Container:ItemPropertyContainerBox

[1200] Mandatory(per item):Yes, for a V-PCC item of type 'vpci' and for a V-PCC unit item

[1201] Quantity (per item): One

[1202] V-PCC unit headers are stored as descriptive item properties and are associated with V-PCC items and V-PCC unit items.

[1203] Required is set to 1 for the 'vunt' item property.

[1204] aligned(8) class VPCCUnitHeaderProperty(){

[1205] extends ItemFullProperty('vunt', version=0, 0){

[1206] vpcc_unit_header();

[1207] }

[1208] Figure 46 shows V-PCC 3D bounding box item properties according to an embodiment.

[1209] Figure 46 shows the 3D bounding box item properties contained in the item in Figure 45.

[1210] V-PCC 3D bounding box item property

[1211] box Types:'v3dd'

[1212] Property type:Descriptive item property

[1213] Container:ItemPropertyContainerBox

[1214] Mandatory(per item):Yes, for a V-PCC item of type 'vpci' and for a V-PCC unit item

[1215] Quantity (per item): One

[1216] The 3D bounding information is stored as a descriptive item property and is associated with V-PCC items and V-PCC unit items.

[1217] aligned(8) class VPCC3DBoundingBoxInfoProperty(){

[1218] extends ItemFullProperty('V3DD', version=0, 0){

[1219] 3DBoundingBoxInfoStruct();

[1220] }

[1221] In the proposed method according to the embodiment, a transmitter or receiver for providing a point cloud content service constructs a V-PCC bitstream and stores the file as described above.

[1222] Transmits metadata in the bitstream for data processing and rendering within the V-PCC bitstream.

[1223] It allows players to access the space or parts of the point cloud object / content through a user viewport, etc. In other words, the above-described data representation method provides an effect of efficiently accessing the point cloud bitstream.

[1224] The point cloud data transmitting device according to the embodiment provides a bounding box and signaling information for partial and / or spatial access to point cloud content (e.g., V-PCC content), thereby enabling the point cloud data receiving device according to the embodiment to access point cloud content in a variety of ways taking into account the player or user environment.

[1225] The method / apparatus according to the embodiment stores and signals 3D bounding box information of point cloud videos / images in tracks or items, stores and signals information related to viewing information (viewing position, viewing orientation, viewing direction, etc.) of point cloud videos / images in tracks or items, and groups tracks / items associated with viewing information and 3D bounding box information and signals the grouping.

[1226] We propose a sender or receiver for providing point cloud content services that handles file storage techniques to support efficient access to stored V-PCC bitstreams.

[1227] The 3D bounding box of point cloud data (video or image) is static or dynamically changes over time. A method for constructing and signaling point cloud data based on the user's viewport and / or viewing information of the point cloud data (video or image) is required.

[1228] The definitions of terms used in this specification are as follows: VPS: V-PCC parameter set, AD: atlas data, OVD: occupancy video data, GVD: geometry video data, AVD: attribute video data, ACL: Atlas Coding Layer

[1229] AAPS: Atlas Adaptation Parameter Set, which contains camera parameters such as camera position, rotation, scale and camera model, as well as data associated with parts of the atlas sub-bitstream.

[1230] ASPS: atlas sequence parameter set. A syntax structure containing syntax elements that apply to zero or more entire coded atlas sequences (CASs) as determined by the content of the syntax elements in ASPS, which are called syntax elements in each tile group header.

[1231] AFPS:atlas frame parameter set: A syntax structure containing syntax elements that apply to zero or more entire coded atlas frames, as determined by the content of syntax elements in the tile group header.

[1232] SEI:Supplemental enhancement information

[1233] Atlas: A collection of 2D bounding boxes, i.e., patches, projected onto a rectangular frame corresponding to the 3D bounding box in 3D space. It represents a subset of a point cloud.

[1234] Atlas sub-bitstream: A sub-bitstream extracted from a V-PCC bitstream that contains parts of the Atlas NAL bitstream.

[1235] V-PCC content: Point clouds that are encoded based on video-coded point cloud compression (V-PCC).

[1236] V-PCC track: A volumetric visual track that carries the atlas bitstream for a V-PCC bitstream.

[1237] V-PCC component track: A video track that conveys 2D video coded data for the occupancy map, geometry, or attribute component video bitstream of a V-PCC bitstream.

[1238] The 3D bounding box of point cloud data (video or image) may be static or may dynamically change over time. The 3D bounding box information of the point cloud data may be used to select a track or item containing the point cloud data in a file, such as via a user viewport, or to parse / decode / render the data in the track or item. A method / apparatus according to an embodiment stores or signals the 3D bounding box information of the point cloud data within a track or item. Depending on the changing attributes of the 3D bounding box, the 3D bounding box information may be conveyed by including it in a sample entry, track group, sample group, or separate metadata track.

[1239] Viewing information (such as viewing position, viewing orientation, and viewing direction) of point cloud data (video or image) may be static or may change dynamically over time. To effectively provide point cloud video data to users, a method / apparatus according to an embodiment stores or signals the corresponding viewing information in a track or item. Depending on the changing attributes of the viewing information, the corresponding information may be included and transmitted in a sample entry, track group, sample group, or separate metadata track.

[1240] Referring to Figures 26 and 27, each V-PCC unit has a V-PCC unit header and a V-PCC unit payload. The V-PCC unit header describes the V-PCC unit type. The attribute video data V-PCC unit header describes the attribute type and its index. Multiple instances of the same attribute type can be supported. The occupancy, geometry, and attribute video data unit payloads correspond to video data units, e.g., HEVC NAL units. Such units are decoded by an appropriate video decoder.

[1241] A V-PCC bitstream contains a series of V-PCC sequences. The V-PCC unit type with a vuh_unit_type value equal to vpcc_VPS is the first V-PCC unit type in a V-PCC sequence. All other V-PCC unit types follow that unit type without any additional restrictions on their coding order. The V-PCC unit payload of a V-PCC unit carrying proprietary video, attribute video, or geometry video consists of one or more NAL units.

[1242] FIG. 47 shows a V-PCC unit according to an embodiment.

[1243] FIG. 47 shows the syntax of the V-PCC unit 26020 and V-PCC unit header 26030 described in FIG. 28 and FIG.

[1244] The definitions of the elements in Figure 47 refer to the definitions of the corresponding elements in Figure 28.

[1245] VUH attribute partition index (vuh_attribute_partition_index): Indicates the index of the attribute dimension group conveyed in the attribute video data unit.

[1246] VUH Auxiliary Video Flag (vuh_auxiliary_video_flag): A value of 1 indicates that the associated geometry or attribute video data unit is RAW and / or EOM coded point video only. A value of 0 indicates that the associated geometry or attribute video data unit contains RAW and / or EOM coded points.

[1247] FIG. 48 shows a V-PCC parameter set according to an embodiment.

[1248] FIG. 48 shows the syntax of a parameter set when the payload 26040 of a unit 26020 of a bitstream according to an embodiment includes a parameter set, as in FIG. 30, 26 to 29.

[1249] The definitions of the elements in FIG. 48 refer to the definitions of the corresponding elements in FIG.

[1250] Auxiliary video present flag (auxiliary_video_present_flag[j]): A value of 1 indicates that auxiliary information, i.e., RAW or EOM patch data, for the atlas with index J is stored in a separate video stream called the auxiliary video stream. A value of 0 indicates that auxiliary information for the atlas with index J is not stored in a separate video stream.

[1251] VPS Extension Present Flag (vps_extension_present_flag): A value of 1 indicates that the syntax element vps_extension_length is present in the vpcc_parameter_set syntax structure. A value of 0 indicates that the syntax element vps_extension_length is not present.

[1252] VPS Extension Length (vps_extension_length_minus1): This value plus 1 indicates the number of vps_extension_data_byte elements that follow this syntax element.

[1253] The VPS extension data byte (vps_extension_data_byte) has various values.

[1254] FIG. 49 shows an atlas frame according to an embodiment.

[1255] Figure 49 shows an atlas frame including tiles encoded by the encoder 10002 of Figure 1, the encoder of Figure 4, the encoder of Figure 15, the transmitting device of Figure 18, and the systems of Figures 20 and 21, etc., and shows an atlas frame including tiles decoded by the decoder 10008 of Figure 1, the decoders of Figures 16 and 17, the receiving device of Figure 19, and the system of Figure 23, etc.

[1256] An atlas frame is divided into one or more tile rows and one or more tile columns. A tile is a rectangular area of ​​the atlas frame. A tile group contains multiple tiles of the atlas frame. Tile and tile group are not distinguished from each other, and a tile group corresponds to one tile. Only rectangular tile groups are supported. In this mode, a tile group (or tile) is a rectangular area of ​​the atlas frame collectively containing multiple tiles of the atlas frame. Figure 49 shows tile or tile group division of an atlas frame according to an embodiment. 24 tiles (6 tile columns and 4 tile rows) and 9 rectangular tile groups of an atlas frame are shown in Figure 49. Without distinguishing between tile group and tile according to an embodiment, tile group is used as a term corresponding to tiles.

[1257] That is, in one embodiment, tile group 49000 corresponds to tile 49010 and is referred to as tile 49010. Tile 49010 corresponds to tile partition and is referred to as tile partition. The names of signaling information can also be changed and referred to according to such a mutually complementary correspondence.

[1258] Figure 50 shows the structure of an atlas bitstream according to an embodiment.

[1259] FIG. 50 shows an example in which the payload 26040 of the unit 26020 of the bitstream 26000 of FIGS. 31 and 26 carries an atlas sub-bitstream 31000.

[1260] The term 'sub' used in the embodiments can be interpreted as meaning a part, and a sub-bitstream can also be interpreted as a bitstream depending on the embodiment.

[1261] An atlas bitstream according to the embodiment includes a sample stream NAL header and a sample stream NAL unit.

[1262] Sample stream NAL units according to the embodiment each include an APSP, an AAPS, an AFPS, one or more atlas tile groups, one or more mandatory SEIs, and one or more non-mandatory SEIs.

[1263] The syntax of the information contained in the atlas bitstream of Figure 50 will be explained below.

[1264] Figure 51 shows a sample stream NAL unit header, sample stream NAL unit, NAL unit, and NAL unit header included in a bitstream containing point cloud data according to an embodiment.

[1265] Figure 51 shows the syntax of the data contained in the Atlas bitstream of Figure 50.

[1266] Sample Stream NAL Unit Header Unit Precision Byte (ssnh_unit_size_precision_bytes_minus1): This value, plus 1, indicates the accuracy, in bytes, of the sample stream NAL unit size (ssnu_nal_unit_size) element in all sample stream NAL units. The sample stream NAL unit header unit precision byte (ssnh_unit_size_precision_bytes_minus1) is in the range of 0 to 7.

[1267] Sample Stream NAL Unit NAL Unit Size (ssnu_nal_unit_size): Indicates the size in bytes of the following NAL unit. The number of bits used to indicate the sample stream NAL unit size is (ssnh_unit_size_precision_bytes_minus1+1)*8.

[1268] NAL unit number (NumBytesInNalUnit): Indicates the size of the NAL unit in bytes.

[1269] Number of bytes (NumBytesInRbsp): The number of bytes that belong to the payload of the NAL unit, initialized to 0.

[1270] Byte (rbsp_byte[i]): The i-th byte of the RBSP. The RBSP is represented as an ordered sequence of bytes.

[1271] NAL Forbidden Zero Bit (nal_forbidden_zero_bit): 0.

[1272] The NAL unit type (nal_unit_type) has values ​​as shown in Figure 52. The NAL unit type indicates the type of RBSP data structure included in the NAL unit.

[1273] NAL layer ID (nal_layer_id): Indicates the identifier of the layer to which the ACL NAL unit belongs or the identifier of the layer to which the NON-ACL NAL unit applies.

[1274] NAL Temporal ID (nal_temporal_id_plus1): This value minus 1 indicates the temporary identifier of the NAL unit.

[1275] Figure 52 shows the types of NAL units according to an embodiment.

[1276] FIG. 52 shows the types of NAL unit types included in the NAL unit headers of the sample stream NAL units of FIG.

[1277] NAL Trail (NAL_TRAIL): A coded tile group of a NON-TSA, NON-STSA trailing atlas frame is included in the NAL unit. The RBSP syntax structure of the NAL unit is atlas_tile_group_layer_rbsp( ) or atlas_tile_layer_rbsp( ). The type class of the NAL unit is ACL. Depending on the embodiment, a tile group corresponds to a tile.

[1278] NAL TSA: A coded tile group of a TSA atlas frame is contained in the NAL unit. The RBSP syntax structure of the NAL unit is atlas_tile_group_layer_rbsp( ) or atlas_tile_layer_rbsp( ). The type class of the NAL unit is ACL.

[1279] NAL_STSA: The NAL unit contains a coded tile group of an STSA atlas frame. The RBSP syntax structure of the NAL unit is atlas_tile_group_layer_rbsp( ) or atlas_tile_layer_rbsp( ). The type class of the NAL unit is ACL.

[1280] NAL_RADL: A coded tile group of a RADL atlas frame is contained in the NAL unit. The RBSP syntax structure of the NAL unit is atlas_tile_group_layer_rbsp( ) or atlas_tile_layer_rbsp( ). The type class of the NAL unit is ACL.

[1281] NAL_RASL: A coded tile group of an RASL atlas frame is contained in the NAL unit. The RBSP syntax structure of the NAL unit is atlas_tile_group_layer_rbsp( ) or atlas_tile_layer_rbsp( ). The type class of the NAL unit is ACL.

[1282] NAL_SKIP: A coded tile group of a skip atlas frame is included in the NAL unit. The RBSP syntax structure of the NAL unit is atlas_tile_group_layer_rbsp( ) or atlas_tile_layer_rbsp( ). The type class of the NAL unit is ACL.

[1283] NAL_RSV_ACL_6 to NAL_RSV_ACL_9: Reserved IRAP ACL NAL unit type is included in the NAL unit. The type class of the NAL unit is ACL.

[1284] NAL_BLA_W_LP, NAL_BLA_W_RADL, NAL_BLA_N_LP: The coded tile group of a BLA atlas frame is contained in the NAL unit. The RBSP syntax structure of the NAL unit is atlas_tile_group_layer_rbsp( ) or atlas_tile_layer (atlas_tile_layer_rbsp( )). The type class of the NAL unit is ACL.

[1285] NAL_GBLA_W_LP, NAL_GBLA_W_RADL, NAL_GBLA_N_LP: A coded tile group of a GBLA atlas frame is contained in the NAL unit. The RBSP syntax structure of the NAL unit is atlas_tile_group_layer_rbsp( ) or atlas_tile_layer_rbsp( ). The type class of the NAL unit is ACL.

[1286] NAL_IDR_W_RADL, NAL_IDR_N_LP: The coded tile group of an IDR atlas frame is contained in the NAL unit. The RBSP syntax structure of the NAL unit is atlas_tile_group_layer_rbsp( ) or atlas_tile_layer_rbsp( ). The type class of the NAL unit is ACL.

[1287] NAL_GIDR_W_RADL, NAL_GIDR_N_LP: The coded tile group of a GIDR atlas frame is contained in the NAL unit. The RBSP syntax structure of the NAL unit is atlas_tile_group_layer_rbsp( ) or atlas_tile_layer_rbsp( ). The type class of the NAL unit is ACL.

[1288] NAL_CRA: A coded tile group of a CRA atlas frame is contained in the NAL unit. The RBSP syntax structure of the NAL unit is atlas_tile_group_layer_rbsp( ) or atlas_tile_layer (atlas_tile_layer_rbsp( )). The type class of the NAL unit is ACL.

[1289] NAL_GCRA: A coded tile group of a GCRA atlas frame is contained in the NAL unit. The RBSP syntax structure of the NAL unit is atlas_tile_group_layer_rbsp( ) or atlas_tile_layer_rbsp( ). The type class of the NAL unit is ACL.

[1290] NAL_IRAP_ACL_22, NAL_IRAP_ACL_23: The reserved IRAP ACL NAL unit type is included in the NAL unit. The type class of the NAL unit is ACL.

[1291] NAL_RSV_ACL_24 to NAL_RSV_ACL_31: Reserved non-IRAP ACL NAL unit types are included in the NAL unit. The type class of the NAL unit is ACL.

[1292] NAL_ASPS: An atlas sequence parameter set is included in the NAL unit. The RBSP syntax structure of the NAL unit is atlas sequence parameter set (atlas_sequence_parameter_set_rbsp( )). The type class of the NAL unit is non-ACL.

[1293] NAL_AFPS: An atlas frame parameter set is included in the NAL unit. The RBSP syntax structure of the NAL unit is atlas frame parameter set (atlas_frame_parameter_set_rbsp( )). The type class of the NAL unit is non-ACL.

[1294] NAL_AUD: An access unit delimiter is included in the NAL unit. The RBSP syntax structure of the NAL unit is access_unit_delimiter_rbsp( ). The type class of the NAL unit is non-ACL.

[1295] NAL_VPCC_AUD: A V-PCC access unit delimiter is included in the NAL unit. The RBSP syntax structure of the NAL unit is access_unit_delimiter_rbsp( ). The type class of the NAL unit is non-ACL.

[1296] NAL_EOS: The NAL unit type is end of sequence. The RBSP syntax structure of the NAL unit is end of sequence (end_of_seq_rbsp( )). The type class of the NAL unit is non-ACL.

[1297] NAL_EOB: The NAL unit type is end of bitstream. The RBSP syntax structure of the NAL unit is end of atlas sub-bitstream (end_of_atlas_sub_bitstream_rbsp( )). The type class of the NAL unit is non-ACL.

[1298] NAL_FD Filler: The NAL unit type is filter data (filler_data_rbsp( )). The NAL unit type class is non-ACL.

[1299] NAL_PREFIX_NSEI, NAL_SUFFIX_NSEI: The NAL unit type is non-essential supplemental enhancement information. The RBSP syntax structure of the NAL unit is SEI(sei_rbsp( )). The type class of the NAL unit is non-ACL.

[1300] NAL_PREFIX_ESEI, NAL_SUFFIX_ESEI: The NAL unit type is Essential Supplemental Enhancement Information. The RBSP syntax structure of the NAL unit is SEI(sei_rbsp( )). The type class of the NAL unit is non-ACL.

[1301] NAL_AAPS: The NAL unit type is Atlas adaptation parameter set. The RBSP syntax structure of the NAL unit is atlas_adaptation_parameter_set_rbsp( ). The type class of the NAL unit is non-ACL.

[1302] NAL_RSV_NACL_44 to NAL_RSV_NACL_47: The NAL unit type is a reserved non-ACL NAL unit type. The type class of the NAL unit is non-ACL.

[1303] NAL_UNSPEC_48 to NAL_UNSPEC_63: The NAL unit type is an unspecified non-ACL NAL unit type. The type class of the NAL unit is non-ACL.

[1304] FIG. 53 shows an atlas sequence parameter set according to an embodiment.

[1305] FIG. 53 corresponds to the atlas sequence parameter set of FIG.

[1306] The definitions of the elements in FIG. 53 follow the definitions of the corresponding elements in FIG.

[1307] ASPS Extended Projection Enabled Flag (asps_extended_projection_enabled_flag): A value of 0 indicates that patch projection information is not signaled for the current atlas tile group or atlas tile. A value of 1 indicates that patch projection information is signaled for the current atlas tile group or atlas tile.

[1308] ASPS raw patch enabled flag (asps_raw_patch_enabled_flag): Indicates whether raw patch is enabled.

[1309] ASPS EOM patch enabled flag (asps_eom_patch_enabled_flag): A value of 1 indicates that the decoded occupancy map video for the current atlas contains information related to whether intermediate depth positions between two depth maps are occupied. A value of 0 indicates that the decoded occupancy map video does not contain information related to whether intermediate depth positions between two depth maps are occupied.

[1310] ASPS auxiliary video enabled flag (asps_auxiliary_video_enabled_flag): Indicates whether auxiliary video is enabled.

[1311] ASPS pixel deinterleaving enabled flag (asps_pixel_deinterleaving_enabled_flag): Indicates whether pixel deinterleaving is enabled.

[1312] ASPS pixel deinterleaving map flag [i] (asps_pixel_deinterleaving_map_flag[i]): A value of 1 indicates that the decoded geometry and feature video with index i in the current atlas contains spatially interleaved pixels corresponding to two maps. A value of 0 indicates that the decoded geometry and feature video corresponding to the map with index i in the current atlas contains pixels corresponding to a single map.

[1313] If asps_eom_patch_enabled_flag and asps_map_count_minus1 are zero, the asps_eom_fix_bit_count_minus1 element contains ASPS.

[1314] ASPS EOM Fixed Bit Count (asps_eom_fix_bit_count_minus1): This value plus 1 indicates the size in bits of the EOM codeword.

[1315] ASPS Extension Flag (asps_extension_flag): A value of zero indicates that the asps_extension_data_flag syntax element is not present in the ASPS RBSP syntax structure.

[1316] ASPS Extension Data Flag (asps_extension_data_flag): This element has various values.

[1317] FIG. 54 shows an atlas frame parameter set according to an embodiment.

[1318] Figure 54 corresponds to the atlas frame parameters in Figure 33. The definitions of the elements in Figure 54 follow the definitions of the corresponding elements in Figure 33.

[1319] AFPS Output Flag Present Flag (afps_output_flag_present_flag): A value of 1 indicates that the ATGH Frame Output Flag (atgh_frame_output_flag) syntax element is present in the associated tile group header. A value of 0 indicates that the...

Claims

1. 1. A method for encoding point cloud data, wherein an encoder comprises: encoding geometry data of the point cloud data; encoding characteristic data of the point cloud data; encapsulating the encoded geometry data and the encoded attribute data in a file; transmitting the file; the file includes a first track containing the geometry data and the attribute data; the file further comprises a second track comprising a timed metadata track; the timed metadata track includes a sample entry and a sample; a sample entry type of the sample entry indicating that a dynamic region of the point cloud data is being conveyed; the sample entry includes information for the dynamic region; The method, wherein the sample includes information for each dynamic region.

2. The method of claim 1 , wherein the second track includes a sample entry including a parameter set for the point cloud data and an atlas sub-bitstream parameter set for the point cloud data.

3. 1. A method for a decoder to decode point cloud data, comprising: receiving a file containing point cloud data; Decapsulating the file, the file includes a first track including geometry data of the point cloud data and attribute data of the point cloud data; the file further comprises a second track comprising a timed metadata track; the timed metadata track includes a sample entry and a sample; a sample entry type of the sample entry indicating that a dynamic region of the point cloud data is being conveyed; the sample entry includes information for the dynamic region; the sample includes information for each dynamic region; decoding the geometry data; and decoding said characteristic data.

4. The method of claim 3 , wherein the second track includes a sample entry including a parameter set for the point cloud data and an atlas sub-bitstream parameter set for the point cloud data.

5. 1. An apparatus for encoding point cloud data, comprising: Memory and at least one processor coupled to the memory; The at least one processor encoding geometry data of the point cloud data; encoding characteristic data of the point cloud data; encapsulating the encoded geometry data and the encoded attribute data in a file; transmitting the file; the file includes a first track containing the geometry data and the attribute data; the file further comprises a second track comprising a timed metadata track; the timed metadata track includes a sample entry and a sample; a sample entry type of the sample entry indicating that a dynamic region of the point cloud data is being conveyed; the sample entry includes information for the dynamic region; The sample includes information for each dynamic region.

6. 1. An apparatus for decoding point cloud data, comprising: Memory and at least one processor coupled to the memory; The at least one processor Decapsulate the file, the file includes a first track including geometry data of the point cloud data and attribute data of the point cloud data; the file further comprises a second track comprising a timed metadata track; the timed metadata track includes a sample entry and a sample; a sample entry type of the sample entry indicating that a dynamic region of the point cloud data is being conveyed; the sample entry includes information for the dynamic region; the sample includes information for each dynamic region; Decoding the geometry data; The device is configured to decode the characteristic data.

Citation Information

Patent Citations

  • 3D point cloud compression systems for delivery and access of a subset of a compressed 3D point cloud

    US20190318488A1

  • Method, apparatus and stream for volumetric video format

    WO2019079032A1

  • An apparatus, a method and a computer program for volumetric video

    WO2019135024A1