Point cloud data transmission device, point cloud data transmission method, point cloud data reception device, and point cloud data reception method
The point cloud data transmission method and devices address the inefficiencies in transmitting and receiving point cloud data by using V-PCC encoding and incorporating viewport information, resulting in efficient and optimized point cloud content delivery.
Patent Information
- Application Number
- JP2024140997
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-03-16
- Filing Date
- 2024-08-22
- Publication Date
- 2025-06-23
- Estimated Expiration
- 2040-10-07
AI Technical Summary
The existing technologies face challenges in efficiently transmitting and receiving point cloud data due to its large size, requiring significant processing power, and experiencing issues with latency and encoding/decoding complexity. Additionally, there is a need to optimize point cloud content for users by incorporating viewport information.
A point cloud data transmission method and devices that encode point cloud data using a video-based point cloud compression (V-PCC) method and transmit a bitstream including the encoded data and signaling information. The signaling information includes viewport information determined by the position and orientation of a camera or user, which is used to optimize the point cloud content for rendering.
The proposed solution enables efficient transmission and reception of high-quality point cloud data, reducing latency and encoding/decoding complexity. It also provides optimized point cloud content for users by utilizing viewport information, enhancing the rendering process and user experience.
Smart Images

Figure 0007697119000006 
Figure 0007697119000007 
Figure 0007697119000008
Abstract
Description
Technical Field
[0001] The embodiments provide a solution for providing point cloud content in order to provide various services such as VR (Virtual Reality), AR (Augmented Reality), MR (Mixed Reality), and autonomous driving services to users.
Background Art
[0002] A point cloud is a set of points in 3D space. There is a problem that the amount of points in 3D space is large and it is difficult to generate point cloud data.
[0003] For the transmission and reception of point cloud data, there is a problem that a large amount of processing power is required.
Summary of the Invention
Problems to be Solved by the Invention
[0004] The technical problem related to the embodiments is to provide a point cloud data transmission device, a transmission method, a point cloud data reception device, and a reception method for efficiently transmitting and receiving point clouds in order to solve the above-mentioned problems.
[0005] The technical problem related to the embodiments is to provide a point cloud data transmission device, a transmission method, a point cloud data reception device, and a reception method for solving latency and encoding / decoding complexity.
[0006] The technical problem related to the embodiments is to provide a point cloud data transmission device, a transmission method, a point cloud data reception device, and a reception method for providing point cloud content optimized for users by signaling information related to the viewport of point cloud data.
[0007] The technical problem related to the embodiment is to be able to transmit viewport information, recommended viewport information, and initial viewing orientation (i.e., viewpoint) information for data processing and rendering in a V-PCC bitstream into the bitstream, and to provide a point cloud data transmission device, a transmission method, a point cloud data reception device, and a reception method for providing point cloud content optimized for a user.
[0008] However, it is not limited only to the above-described technical problems, and the scope of rights of the embodiment can also be extended to other technical problems derived by those skilled in the art based on all the contents of this document.
Means for Solving the Problem
[0009] To achieve the above-described object and other advantages, the point cloud data transmission method according to the embodiment may include a step of encoding point cloud data and a step of transmitting a bitstream including the point cloud data and signaling information.
[0010] According to the embodiment, the point cloud data may include geometry data, texture data, and occupancy map data encoded by a video-based point cloud compression (V-PCC) method.
[0011] According to the embodiment, the signaling information may include information regarding a viewport for a viewport determined according to the position and orientation of a camera or a user.
[0012] According to the embodiment, the information regarding the viewport may include at least one of coordinate information of the camera or the user in 3D space, direction vector information indicating the direction in which the camera or the user is looking, up vector information indicating above the camera or the user, and right vector information indicating the right side of the camera or the user.
[0013] According to an embodiment, the information regarding the viewport may further include horizontal field of view (FOV) information and vertical FOV information for generating the viewport.
[0014] According to an embodiment, the point cloud data transmission device may include an encoder that encodes point cloud data and a transmitter that transmits a bitstream including the point cloud data and signaling information.
[0015] According to an embodiment, the point cloud data may include geometry data, texture data, and occupancy map data encoded by a video-based point cloud compression (V-PCC) method.
[0016] According to an embodiment, the signaling information may include information regarding a viewport for a viewport determined according to the position and orientation of a camera or a user.
[0017] According to an embodiment, the information regarding the viewport may include at least one of coordinate information of the camera or the user in 3D space, direction vector information indicating the direction in which the camera or the user is looking, up vector information indicating above the camera or the user, and right vector information indicating the right side of the camera or the user.
[0018] According to an embodiment, the information regarding the viewport may further include horizontal FOV information and vertical FOV information for generating the viewport.
[0019] According to an embodiment, the point cloud data reception method may include steps of receiving a bitstream including point cloud data and signaling information, decoding the point cloud data, and rendering the decoded point cloud data.
[0020] According to an embodiment, the signaling information may include information regarding a viewport for a viewport determined according to the position and orientation of a camera or a user.
[0021] According to an embodiment, the information regarding the viewport may include at least one of coordinate information of the camera or the user in 3D space, direction vector information indicating the direction the camera or the user is looking, up vector information indicating above the camera or the user, and right vector information indicating the right side of the camera or the user.
[0022] According to an embodiment, the information regarding the viewport may further include horizontal FOV information and vertical FOV information for generating the viewport.
[0023] According to an embodiment, the decoded point cloud data may be rendered based on the information regarding the viewport.
[0024] According to an embodiment, a point cloud data receiving device may include a receiver that receives a bitstream including point cloud data and signaling information, a decoder that decodes the point cloud data, and a renderer that renders the decoded point cloud data.
[0025] According to an embodiment, the signaling information may include information regarding a viewport for a viewport determined according to the position and orientation of a camera or a user.
[0026] According to an embodiment, the information regarding the viewport may include at least one of coordinate information of the camera or the user in 3D space, direction vector information indicating the direction the camera or the user is looking, up vector information indicating above the camera or the user, and right vector information indicating the right side of the camera or the user.
[0027] According to an embodiment, the information regarding the viewport may further include horizontal FOV information and vertical FOV information for generating the viewport.
[0028] According to an embodiment, the renderer may render the decoded point cloud data based on the information regarding the viewport.
Effect of the Invention
[0029] The point cloud data transmission method, transmission device, point cloud data reception method, and reception device according to the embodiment can provide a high-quality point cloud service.
[0030] The point cloud data transmission method, transmission device, point cloud data reception method, and reception device according to the embodiment can achieve various video codec methods.
[0031] The point cloud data transmission method, transmission device, point cloud data reception method, and reception device according to the embodiment can provide general-purpose point cloud contents such as autonomous driving services.
[0032] The point cloud data transmission method, transmission device, point cloud data reception method, and reception device according to the embodiment can provide an optimal point cloud content service by configuring a V-PCC bitstream and enabling the transmission, reception, and storage of files.
[0033] The point cloud data transmission method, transmission device, point cloud data reception method, and reception device according to the embodiment can provide an optimal point cloud content service by including metadata for data processing and rendering within the V-PCC bitstream and enabling its transmission and reception.
[0034] The point cloud data transmission method, transmission device, point cloud data reception method, and reception device according to the embodiments can provide the effect that the user can efficiently access and process the point cloud bitstream through the user's viewport by enabling access to the space or part of the point cloud object / content through the user viewport in a player or the like.
[0035] The point cloud data transmission method, transmission device, point cloud data reception method, and reception device according to the embodiments can provide the effect that the point cloud content can be accessed in various ways on the receiving side considering the player or user environment by providing a bounding box for partial access and / or spatial access to the point cloud content and signaling information therefor.
[0036] The point cloud data transmission method and transmission device according to the embodiments can provide 3D region information of the point cloud content for assisting in accessing the space or part of the point cloud content according to the user's viewport, and metadata related to the 2D region on the related video or atlas frame.
[0037] The point cloud data transmission method and transmission device according to the embodiments can process 3D region information of the point cloud in the point cloud bitstream, information signaling related to the 2D region on the related video or atlas frame, and the like.
[0038] The point cloud data reception method and reception device according to the embodiments can efficiently access the point cloud content based on storing and signaling information related to the 3D region information of the point cloud in the file and the 2D region on the related video or atlas frame.
[0039] The point cloud data reception method and reception apparatus according to the embodiment can provide point cloud content considering the user environment based on the 3D region information of the point cloud related to the in-file image item and the information related to the 2D region on the related video or atlas frame.
Brief Description of the Drawings
[0040] The drawings are attached to further understand the embodiments and illustrate the embodiments together with the descriptions of the embodiments.
[0041]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33
Figure 34
Figure 35
Figure 36
Figure 37
Figure 38
Figure 39
Figure 40
Figure 41
Figure 42
Figure 43
Figure 44
Figure 45
Figure 46
Figure 47
Figure 48
Figure 49
Figure 50
Figure 51
Figure 52
Figure 53
Figure 54
Figure 55
Figure 56
Figure 57
Figure 58
Figure 59
Figure 60
Figure 61
Figure 62
Figure 63
Best Mode for Carrying Out the Invention
[0042] Hereinafter, the preferred embodiments will be specifically described with reference to the accompanying drawings. The following detailed description with reference to the accompanying drawings is for explaining the preferred embodiments rather than showing only the embodiments that can be implemented by the examples. The following detailed description includes details to provide a thorough understanding of the embodiments. However, it is obvious to those skilled in the art that the embodiments can be implemented even without such details.
[0043] Most of the terms used in the embodiments are common and widely used in the relevant field, but some are arbitrarily selected by the applicant, and their meanings will be explained in detail below if necessary. Therefore, the embodiments should be understood based on the intended meanings of the terms rather than just their simple names and meanings.
[0044] FIG. 1 shows an example of the structure of a transmission / reception system for providing point cloud content according to an embodiment.
[0045] In this document, a solution for providing point cloud content is provided to offer various services such as VR (Virtual Reality), AR (Augmented Reality), MR (Mixed Reality), and autonomous driving services to users. The point cloud content according to the embodiment represents data in which an object is represented by points, and may be referred to as point cloud, point cloud data, point cloud video data, point cloud image data, etc.
[0046] The point cloud data transmission device 10000 according to the embodiment includes a point cloud video acquisition unit 10001, a point cloud video encoder 10002, a file / segment encapsulation unit 10003 and / or a transmitter (or communication module) 10004. The transmission device according to the embodiment can secure, process, and transmit point cloud video (or point cloud content). According to the embodiment, the transmission device may include a fixed station, a BTS (base transceiver system), a network, an AI (Artificial Intelligence) device and / or system, a robot, an AR / VR / XR device and / or a server, etc. Also, according to the embodiment, the transmission device 10000 may be a device that communicates with a base station and / or other wireless devices using a wireless connection technology (e.g., 5G NR (New RAT), LTE (Long Term Evolution)), a robot, a vehicle, an AR / VR / XR device, a portable device, a household appliance, an IoT (Internet of Thing) device, an AI device / server, etc.
[0047] The point cloud video acquisition unit 10001 according to the embodiment acquires point cloud video through processes such as capture, synthesis, or generation of point cloud video.
[0048] The point cloud video encoder 10002 according to the embodiment encodes the point cloud video data acquired by the point cloud video acquisition unit 10001. According to the embodiment, the point cloud video encoder 10002 is also called a point cloud encoder, a point cloud data encoder, an encoder, etc. Further, the point cloud compression coding (encoding) according to the embodiment is not limited to the above-described embodiment. The point cloud video encoder can output a bitstream including the encoded point cloud video data. The bitstream may include not only the encoded point cloud video data but also signaling information related to the encoding of the point cloud video data.
[0049] The point cloud video encoder 10002 according to the embodiment can support both the G-PCC (Geometry-based Point Cloud Compression) encoding method and / or the V-PCC (Video-based Point Cloud Compression) encoding method. Further, the point cloud video encoder 10002 can encode a point cloud (referring to all of the point cloud data or points) and / or signaling data related to the point cloud.
[0050] The file / segment encapsulation module 10003 according to the embodiment encapsulates the point cloud data in the form of a file and / or a segment. The point cloud data transmission method / apparatus according to the embodiment can transmit the point cloud data in the form of a file and / or a segment.
[0051] The transmitter (Transmitter (or Communication module)) 10004 according to the embodiment transmits the encoded point cloud video data in the form of a bitstream. According to the embodiment, a file or a segment is transmitted to a receiving device via a network or stored in a digital storage medium (e.g., USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.). The transmitter according to the embodiment can perform wired / wireless communication with a receiving device (or a receiver) via a network such as 4G, 5G, 6G, etc. Further, the transmitter can perform necessary data processing operations according to a network system (e.g., a communication network system such as 4G, 5G, 6G, etc.). Further, the transmitting device can also transmit data encapsulated in an On Demand manner.
[0052] The point cloud data receiving device (Reception device) 10005 according to the embodiment includes a receiver 10006, a file / segment decapsulation unit 10007, a point cloud video decoder 10008, and / or a renderer 10009. According to the embodiment, the receiving device may include a device that communicates with a base station and / or other wireless devices using a wireless connection technology (e.g., 5G NR (New RAT), LTE (Long Term Evolution)), a robot, a vehicle, an AR / VR / XR device, a portable device, a household appliance, an IoT (Internet of Thing) device, an AI device / server, etc.
[0053] The receiver 10006 according to the embodiment receives a bitstream including point cloud video data. According to the embodiment, the receiver 10006 can transmit feedback information to the point cloud data transmitting device 10000.
[0054] The File / Segment Decapsulation module 10007 decapsulates files and / or segments containing point cloud data.
[0055] The Point Cloud video Decoder 10008 decodes the received point cloud video data.
[0056] The Renderer 10009 renders the decoded point cloud video data. According to an embodiment, the Renderer 10009 can send the feedback information obtained at the receiving end to the Point Cloud video Decoder 10008. The point cloud video data according to an embodiment can send the feedback information to the Receiver 10006. According to an embodiment, the feedback information received by the point cloud transmitting device may be provided to the Point Cloud video Encoder 10002.
[0057] The arrow indicated by the dotted line in the drawing shows the transmission path of the feedback information obtained by the receiving device 10005. The feedback information is information for reflecting the interaction with the user who consumes the point cloud content, and includes the user's information (for example, head orientation information, viewport information, etc.). In particular, when the point cloud content is for a service that requires interaction with the user (such as an autonomous driving service, etc.), the feedback information can be transmitted to the content transmitting side (for example, the transmitting device 10000) and / or the service provider. According to an embodiment, the feedback information may or may not be used by the receiving device 10005 in addition to the transmitting device 10000.
[0058] The head orientation information according to the embodiment is information regarding the position, direction, angle, movement, etc. of the user's head. The receiving device 10005 according to the embodiment can calculate viewport information based on the head orientation information. The viewport information is information regarding the area of the point cloud video that the user is viewing. The viewpoint (viewpoint or orientation) is the point at which the user is viewing the point cloud video, and means the center of the viewport area. That is, the viewport is an area centered on the viewpoint, and the size, shape, etc. of the area are determined by the FOV (Field Of View). In other words, the viewport is determined according to the position of the virtual camera or the user and the viewpoint (viewpoint or orientation), and the point cloud data is rendered in the above-described viewport based on the viewport information. Therefore, in addition to the head orientation information, the receiving device 10005 can extract viewport information based on, for example, the vertical or horizontal FOV supported by the device. The receiving device 10005 also performs Gaze Analysis, etc. to confirm the user's point cloud consumption method, the point cloud video area that the user is gazing at, the gazing time, etc. According to the embodiment, the receiving device 10005 can transmit feedback information including the result of the gaze analysis to the transmitting device 10000. The feedback information according to the embodiment is obtained in the rendering and / or display process. The feedback information according to the embodiment is secured by one or more sensors included in the receiving device 10005. Also according to the embodiment, the feedback information is secured by the renderer 10009 or another external element (or device, component, etc.). The dotted line shown in FIG. 1 shows the transmission process of the feedback information secured by the renderer 10009. The point cloud content providing system processes (encodes / decodes) the point cloud data based on the feedback information. Therefore, the point cloud video data decoder 10008 can perform a decoding operation based on the feedback information.In addition, the receiving device 10005 can transmit feedback information to the transmitting device. The transmitting device (or the point cloud video data encoder 10002) can perform an encoding operation based on the feedback information. Therefore, the point cloud content providing system can efficiently process necessary data (for example, point cloud data corresponding to the user's head position) based on the feedback information without processing all the point cloud data (encoding / decoding), and provide point cloud content to the user.
[0059] In an embodiment, the transmitting device 10000 is called an encoder, a transmitting device, a transmitter, etc., and the receiving device 10005 is called a decoder, a receiving device, a receiver, etc.
[0060] The point cloud data processed (processed in a series of processes of acquisition / encoding / transmission / decoding / rendering) by the point cloud content providing system of FIG. 1 according to the embodiment is also called point cloud content data or point cloud video data. According to the embodiment, the point cloud content data can be used as a concept including metadata or signaling information related to the point cloud data.
[0061] The elements of the point cloud content providing system shown in FIG. 1 can be implemented by hardware, software, a processor, and / or a combination thereof.
[0062] The embodiment can provide point cloud content to provide various services such as VR (Virtual Reality), AR (Augmented Reality), MR (Mixed Reality), and autonomous driving services to the user.
[0063] To provide a point cloud content service, first, a point cloud video is acquired. The acquired point cloud video is transmitted to the receiving side through a series of processes, and the data received at the receiving side is processed back into the original point cloud video for rendering. In this way, the point cloud video can be provided to the user. The embodiments provide solutions necessary to effectively perform these series of processes.
[0064] The overall process (point cloud data transmission method and / or point cloud data reception method) for providing a point cloud content service includes an acquisition process, an encoding process, a transmission process, a decoding process, a rendering process, and / or a feedback process.
[0065] According to an embodiment, the process of providing point cloud content (or point cloud data) is also called a Point Cloud Compression process. According to an embodiment, the point cloud compression process means a Video-based Point Cloud Compression (hereinafter referred to as V-PCC) process.
[0066] Each element of the point cloud data transmission device and the point cloud data reception device according to an embodiment means hardware, software, a processor, and / or a combination thereof, etc.
[0067] The point cloud compression system can include a transmitting device and a receiving device. According to an embodiment, the transmitting device is called an encoder, a transmitting apparatus, a transmitter, a point cloud transmitting device, etc. According to an embodiment, the receiving device is called a decoder, a receiving apparatus, a receiver, a point cloud receiving device, etc. The transmitting device can encode a point cloud video to output a bitstream, which can be transmitted to the receiving device via a digital storage medium or a network in the form of a file or streaming (streaming segment). The digital storage medium includes various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.
[0068] As shown in FIG. 1, the transmitting device can include a point cloud video acquisition unit, a point cloud video encoder, a file / segment encapsulation unit, and a transmitting unit (or a transmitter). As shown in FIG. 1, the receiving device can generally include a receiving unit, a file / segment decapsulation unit, a point cloud video decoder, and a renderer. The encoder is also called a point cloud video / video / picture / frame encoding device, and the decoder is also called a point cloud video / video / picture / frame decoding device. The renderer may include a display unit, and the renderer and / or the display unit may be configured as another device or an external component. The transmitting device and the receiving device may further include another internal or external module / unit / component for the feedback process. Each element included in the transmitting device and the receiving device according to the embodiment can be composed of hardware, software, and / or a processor.
[0069] The operation of the receiving device according to the embodiment follows the reverse process of the operation of the transmitting device.
[0070] The point cloud video acquisition unit performs a process of acquiring a point cloud video through processes such as capturing, synthesizing, or generating a point cloud video. Through the acquisition process, 3D position (x, y, z) / attribute (color, reflectivity, transparency, etc.) data for a large number of points, for example, PLY (Polygon File format or the Stanford Triangle format) files, etc., are generated. In the case of a video having a plurality of frames, one or more files can be acquired. Metadata related to the point cloud (for example, metadata related to the capture, etc.) is generated in the capture process.
[0071] The point cloud data transmission device according to an embodiment may include an encoder that encodes point cloud data, and a transmitter that transmits the point cloud data (or a bitstream including the same).
[0072] The point cloud data reception device according to an embodiment includes a receiver that receives a bitstream including point cloud data, a decoder that decodes the point cloud data, and a renderer that renders the point cloud data.
[0073] The method / apparatus according to an embodiment shows a point cloud data transmission device and / or a point cloud data reception device.
[0074] Figure 2 shows an example of the capture of point cloud data according to an embodiment.
[0075] The point cloud data (or point cloud video data) according to an embodiment is acquired by a camera or the like. The capture method according to an embodiment includes, for example, an inward-facing method and / or an outward-facing method.
[0076] The inward method according to the embodiment is a capture method in which one or more cameras capture an object of point cloud data from the outside to the inside of the object.
[0077] The outward method according to the embodiment is a method in which one or more cameras capture an object of point cloud data from the inside to the outside of the object. For example, according to the embodiment, there are four cameras.
[0078] The point cloud data or point cloud content according to the embodiment is a video or still image of an object / environment represented on a 3D space in various forms. According to the embodiment, the point cloud content includes videos / audio / images, etc. for an object (such as an object).
[0079] As equipment for point cloud content capture, it is composed of a combination of a camera equipment capable of obtaining depth (a combination of an infrared pattern projector and an infrared camera) and an RGB camera capable of extracting color information corresponding to the depth information. Alternatively, the depth information can be extracted by a LiDAR that uses a radar system that emits laser pulses and measures the time it takes for them to reflect back to measure the position coordinates of the reflector. It is possible to extract the form of geometry composed of points in 3D space from the depth information, and extract the attributes representing the color / reflection of each point from the RGB information. The point cloud content may be composed of position (x, y, z), color (YCbCr or RGB), or reflectance (r) information for each point. There are an outward-facing method for capturing the external environment and an inward-facing method for capturing the central object in point cloud content. When configuring point cloud content in a VR / AR environment so that the user can freely view an object (e.g., a core object such as a character, athlete, object, actor, etc.) at 360°, the configuration of the capture camera uses the inward-facing method. Also, when configuring the current surrounding environment as point cloud content in a vehicle such as autonomous driving, the configuration of the capture camera uses the outward-facing method. Since point cloud content is captured by multiple cameras, a camera calibration process may be required before capturing the content to set the global coordinate system between the cameras.
[0080] Point cloud content is a video or still image of an object / environment shown in various forms in 3D space.
[0081] In addition, for the method of obtaining point cloud content, any point cloud video can be synthesized based on the captured point cloud video. Or, when attempting to provide a point cloud video for a virtual space generated by a computer, actual camera capture may not be performed. In this case, simply, the corresponding capture process can be replaced by the process of generating relevant data.
[0082] The captured point cloud video requires post-processing to improve the content quality. Although the maximum / minimum depth value can be adjusted within the range provided by the camera equipment in the video capture process, there may still be point data in undesired areas afterwards. Post-processing such as removing the undesired areas (e.g., the background) or recognizing the connected space and filling the spatial holes may be performed. Also, the point clouds extracted from cameras sharing a spatial coordinate system can be integrated into one content through the conversion process to the global coordinate system for each point, based on the position coordinates of each camera obtained by the calibration process. Thereby, it is also possible to generate one wide-range point cloud content or obtain point cloud content with a high point density.
[0083] The point cloud video encoder 10002 can encode an input point cloud video into one or more video streams. One point cloud video may include a plurality of frames, and one frame corresponds to a still video / picture. In this document, the point cloud video includes point cloud video / frame / picture / video / audio / image, etc., and the point cloud video can be used interchangeably with the point cloud video / frame / picture. The point cloud video encoder 10002 performs a video-based point cloud compression (V-PCC) procedure. The point cloud video encoder 10002 performs a series of procedures such as prediction, transformation, quantization, and entropy encoding for the sake of compression and coding efficiency. The encoded data (encoded video / video information) is output in the form of a bitstream. When based on the V-PCC procedure, the point cloud video encoder 10002 encodes the point cloud video into geometric video, attribute video, occupancy map video, and auxiliary information as described below. The geometric video includes geometric images, the attribute video includes attribute images, and the occupancy map video includes occupancy map images. The auxiliary information (or additional data) includes auxiliary patch information. The attribute video / image includes texture video / image.
[0084] The encapsulation module 10003 can encapsulate the encoded point cloud video data and / or point cloud video related metadata in a format such as a file. Here, the point cloud video related metadata may be transmitted from a metadata processing unit or the like. The metadata processing unit may be included in the point cloud video encoder 10002, or may be composed of another component / module. The encapsulation module 10003 may encapsulate the corresponding data in a file format such as ISOBMFF, or may process it in other formats such as DASH segments. According to an embodiment, the encapsulation module 10003 may include point cloud video related metadata on the file format. The point cloud video related metadata is included, for example, in boxes at various levels on the ISOBMFF file format, or in another track within the file. According to an embodiment, the encapsulation module 10003 can encapsulate the point cloud video related metadata itself in a file. The transmission processing unit may add processing for transmission to the point cloud video data encapsulated by the file format. The transmission processing unit may be included in the transmission unit 10004, or may be composed of another component / module. The transmission processing unit processes the point cloud video data according to any transmission protocol. The processing for transmission includes processing for transmission via a broadcast network and processing for transmission via broadband. The transmission processing unit according to an embodiment adds processing for transmission not only to the point cloud video data, but also to the point cloud video related metadata transmitted from the metadata processing unit.
[0085] The transmitting unit 10004 transmits the encoded video / video information or data output in the form of a bitstream to the receiver 10006 of the receiving device via a digital storage medium or a network in the form of a file or a stream. The digital storage medium includes USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit includes elements for generating a media file in a predetermined file format and elements for transmission via a broadcast / communication network. The receiving unit extracts the bitstream and transmits it to the decoding device.
[0086] The receiver 10006 receives the point cloud video data transmitted by the point cloud video transmitting device according to the present invention. Depending on the channel to be transmitted, the receiving unit may receive the point cloud video data via a broadcast network, or may receive the point cloud video data via a broadband, or may receive the point cloud video data via a digital storage medium.
[0087] The receiving processing unit performs processing according to the transmission protocol on the received point cloud video data. The receiving processing unit may be included in the receiver 10006, or may be composed of another component / module. Corresponding to the processing for transmission being performed on the transmitting side, the receiving processing unit performs the reverse process of the above-described transmitting processing unit. The receiving processing unit transmits the acquired point cloud video data to the decapsulation unit 10007 and transmits the acquired point cloud video-related metadata to a metadata processing unit (not shown). The point cloud video-related metadata acquired by the receiving processing unit may be in the form of a signaling table.
[0088] The decapsulation unit (file / segment decapsulation module) 10007 decapsulates the point cloud video data in file format transmitted from the reception processing unit. The decapsulation processing unit 10007 decapsulates a file such as ISOBMFF and obtains a point cloud video bitstream or point cloud video-related metadata (metadata bitstream). The obtained point cloud video bitstream is transmitted to the point cloud video decoder 10008, and the obtained point cloud video-related metadata (metadata bitstream) is transmitted to a metadata processing unit (not shown). The point cloud video bitstream may include metadata (metadata bitstream). The metadata processing unit may be included in the point cloud video decoder 10008, or may be composed of another component / module. The point cloud video-related metadata obtained by the decapsulation processing unit 10007 may be in the form of boxes or tracks within the file format. The decapsulation processing unit 10007 is transmitted the metadata necessary for decapsulation from the metadata processing unit if necessary. The point cloud video-related metadata may be transmitted to the point cloud video decoder 10008 and used in the procedure for decoding the point cloud video, or may be transmitted to the renderer 10009 and used in the procedure for rendering the point cloud video.
[0089] The point cloud video decoder 10008 receives a bitstream, performs operations corresponding to the operations of the point cloud video encoder, and can decode video / images. In this case, the point cloud video decoder 10008 decodes the point cloud video into a geometry video, an attribute video, an occupancy map video, and auxiliary information as described below. The geometry video includes a geometry image, the attribute video includes an attribute image, and the occupancy map video includes an occupancy map image. The auxiliary information includes auxiliary patch information. The attribute video / image includes a texture video / image.
[0090] Using the decoded geometry image, occupancy map, and auxiliary patch information, 3D geometry is restored and may then undergo a smoothing process. By applying texture images to the smoothed 3D geometry to assign color values, a color point cloud video / picture is restored. The renderer 10009 renders the restored geometry and color point cloud video / picture. The rendered video / image is displayed by a display unit (not shown). The user views all or part of the area of the rendered result through a VR / AR display or a general display, etc.
[0091] The feedback process may include a process of transmitting various feedback information obtainable in the rendering / display process to the transmitting side or to the decoder on the receiving side. Through the feedback process, interactivity is provided in the consumption of point cloud videos. According to an embodiment, in the feedback process, Head Orientation information, Viewport information indicating the area that the user is currently viewing, etc. are transmitted. According to an embodiment, the user can interact with what is embodied in the VR / AR / MR / autonomous driving environment, in which case information related to the interaction may be transmitted to the transmitting side or the service provider side in the feedback process. According to an embodiment, the feedback process may not be performed.
[0092] The Head Orientation information is information regarding the position, angle, movement, etc. of the user's head. Based on this information, information on the area that the user is currently viewing within the point cloud video, i.e., the Viewport information, is calculated.
[0093] The Viewport information is information on the area that the user is currently viewing in the point cloud video. Thereby, Gaze Analysis is performed, and it is also possible to confirm how the user consumes the point cloud video, how long the user gazes at which area of the point cloud video, etc. The Gaze Analysis is performed on the receiving side and transmitted to the transmitting side via the feedback channel. Devices such as VR / AR / MR displays extract the Viewport area based on the position / direction of the user's head, the vertical or horizontal FOV supported by the device, etc.
[0094] According to the embodiment, the above-described feedback information is not only transmitted to the transmission side, but may also be consumed at the reception side. That is, processes such as decoding and rendering at the reception side are performed using the above-described feedback information. For example, only the point cloud video for the area that the user is currently viewing is preferentially decoded and rendered using the head orientation information and / or the viewport information.
[0095] Here, the viewport or the viewport area is the area that the user is viewing in the point cloud video. The viewpoint is the point that the user is viewing in the point cloud video and means the center of the viewport area. That is, the viewport is an area centered on the viewpoint, and the size, shape, etc. occupied by the area are determined by the FOV (Field Of View).
[0096] This specification relates to point cloud video compression as described above. For example, the methods / embodiments disclosed in this specification are applicable to the PCC (point cloud compression or point cloud coding) standard of the MPEG (Moving Picture Experts Group) or the next-generation video / image coding standard.
[0097] In this specification, a picture / frame generally means a unit indicating one video in a specific time period.
[0098] A pixel or pel means the smallest unit constituting one picture (or video). Also, the term "sample" is used as a term corresponding to a pixel. A sample generally indicates a pixel or a pixel value, and may indicate only the pixel / pixel value of the luma component, only the pixel / pixel value of the chroma component, or only the pixel / pixel value of the depth component.
[0099] A "unit" represents the basic unit of video processing. A unit includes at least one of a specific region of a picture and information regarding that region. A unit may, in some cases, be used interchangeably with terms such as "block" or "area" or "module". In general, an MxN block includes a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.
[0100] Figure 3 shows an example of a point cloud, geometry, and texture image according to an embodiment.
[0101] The point cloud according to the embodiment is input into the V-PCC encoding process of Figure 4 described later, and a geometry image and a texture image are generated. According to the embodiment, the point cloud is used in the same sense as point cloud data.
[0102] In Figure 3, the left figure is a point cloud, showing a point cloud in which a point cloud object is located in 3D space and is represented by a bounding box or the like. The middle figure in Figure 3 shows a geometry image, and the right figure shows a texture image (non-padded). This specification also refers to the geometry image as a geometry patch frame / picture or a geometry frame / picture. The texture image is also referred to as a texture patch frame / picture or a texture frame / picture.
[0103] Video-based Point Cloud Compression (V-PCC) is a method for compressing 3D point cloud data based on 2D video codecs such as HEVC (High Efficiency Video Coding) and VVC (Versatile Video Coding). In the V-PCC compression process, the following data and information are generated.
[0104] Occupancy map: When dividing the points that make up a point cloud into patches and mapping them onto a 2D plane, it refers to a binary map that indicates the presence or absence of data at the corresponding position on the 2D plane with a value of 0 or 1. The occupancy map represents a 2D array corresponding to the atlas, and the value of the occupancy map indicates whether each sample position in the atlas corresponds to a 3D point. An atlas (ATLAS) means an object that contains information about 2D patches for each point cloud frame. For example, the atlas has the 2D arrangement and size of the patches, the position of the corresponding 3D region within the 3D points, the projection plan, LOD (Level of Detail) parameters, etc.
[0105] Patch: A set of points that make up a point cloud. Points belonging to the same patch are adjacent to each other in 3D space and are mapped in the same direction among the six-sided boundary box planes in the mapping process to a 2D image.
[0106] Geometry image: It refers to an image in the form of a depth map that represents the position information (geometry) of each point that makes up a point cloud in patch units. The geometry image is composed of pixel values of one channel. Geometry refers to a set of coordinates associated with the point cloud frame.
[0107] Texture image: It refers to an image that represents the color information of each point that makes up a point cloud in patch units. The texture image is composed of pixel values of multiple channels (e.g., 3 channels R, G, B). Texture is included in the characteristics. According to the example, texture and / or characteristics are interpreted as the same object and / or inclusion relationship.
[0108] Auxiliary patch info: Indicates the metadata necessary to reconstruct the point cloud from individual patches. The auxiliary patch info includes information regarding the position, size, etc. of the patch in 2D / 3D space.
[0109] Point cloud data according to an embodiment, e.g., a V-PCC component, includes an atlas, occupancy map, geometry, attributes, etc.
[0110] An atlas represents a collection of 2D bounding boxes, i.e., patches, projected into a rectangular frame that corresponds to a 3-dimensional bounding box in 3D space, which may represent a subset of a point cloud. In this case, the patch indicates a rectangular region within the atlas corresponding to a rectangular region in a planar projection. Also, the patch data indicates the data necessary for the transformation of the patches included in the atlas from 2D to 3D. In addition, the patch data group is also called an atlas.
[0111] An attribute indicates a scalar or vector associated with each point in the point cloud, e.g., color, reflectance, surface normal, time stamps, material ID, etc.
[0112] The point cloud data according to the embodiment shows PCC data by the V-PCC (Video-based Point Cloud Compression) method. The point cloud data includes a plurality of components. For example, it includes an occupancy map, patches, geometry, and / or texture, etc.
[0113] FIG. 4 shows an example of a point cloud video encoder according to an embodiment.
[0114] FIG. 4 shows a V-PCC encoding process for generating and compressing an occupancy map, a geometry image, a texture image, and auxiliary patch information. The V-PCC encoding process in FIG. 4 is processed by the point cloud video encoder 10002 in FIG. 1. Each component in FIG. 4 is performed by software, hardware, a processor, and / or a combination thereof.
[0115] The patch generation unit (patch generation, or patch generation section) 14000 receives a point cloud frame (which may be in the form of a bitstream including point cloud data). The patch generation unit 40000 generates patches from the point cloud data. Also, it generates patch information including information related to the generation of the patches.
[0116] The patch packing (patch packing, or patch packing section) 40001 packs one or more patches. Also, it generates an occupancy map including information related to patch packing.
[0117] The geometry image generation unit 40002 generates a geometry image based on point cloud data, patch information (or additional patch information), and / or occupancy map information. The geometry image refers to data including the geometry related to the point cloud data (i.e., the 3D coordinate values of the points), and is also called a geometry frame.
[0118] The texture image generation unit 40003 generates a texture image based on point cloud data, patches, packed patches, patch information (or additional patch information), and / or smoothed geometry. The texture image is also called a texture frame. Also, a texture image can be generated based on the smoothed geometry obtained by smoothing (numbering) the reconstructed geometry image based on the patch information and further smoothing the part that induces data errors.
[0119] The smoothing unit 40004 alleviates or removes the errors included in the image data. For example, a smoothed geometry can be generated by smoothing the reconstructed geometry image based on the patch information, that is, softly filtering the part that induces data errors. The smoothed geometry is output to the texture image generation unit 40003.
[0120] The auxiliary patch info compression unit 40005 compresses the auxiliary patch information related to the patch information generated in the patch generation process. Also, the auxiliary patch information compressed by the auxiliary patch info compression unit 40005 is transmitted to the multiple ksa 40013. The geometry image generation 40002 uses the auxiliary patch information when generating a geometry image.
[0121] The image pads (image padding or image padding portions) 40006 and 40007 pad the geometry image and the texture image respectively. That is, padding data is padded to the geometry image and the texture image.
[0122] The group dilation (group dilation or group dilation portion) 40008 adds data to the texture image in the same manner as the image pad. Additional patch information is inserted into the texture image.
[0123] The video compressions (video compression or video compression portions) 40009, 40010, and 40011 compress the padded geometry image, the padded texture image, and / or the occupancy map respectively. In other words, the video compression portions 40009, 40010, and 40011 compress the input geometry frame, texture frame, and / or occupancy map frame respectively, and output them to the video bitstream of the geometry, the video bitstream of the texture image, and the video bitstream of the occupancy map. Video compression encodes geometry information, texture information, occupancy information, etc.
[0124] The entropy compression (entropy compression or entropy compression portion) 40012 compresses the occupancy map based on the entropy method.
[0125] According to the embodiment, when the point cloud data is lossless and / or lossy, entropy compression and / or video compression is performed on the occupancy map frame.
[0126] The multiplexer 40013 multiplexes the video bitstreams of the geometry compressed in each compression unit, the video bitstream of the compressed texture image, the video bitstream of the compressed occupancy map, and the bitstream of the compressed additional patch information into one bitstream.
[0127] The above-described blocks may be omitted or may be replaced by blocks having similar or identical functions. Also, each block shown in FIG. 4 operates as at least one of a processor, software, and hardware.
[0128] Hereinafter, the detailed operations of each process in FIG. 4 according to the embodiments are shown.
[0129] Patch generation 40000
[0130] The process of patch generation means a process of dividing a point cloud into patches, which are units for performing mapping, in order to map the point cloud to a 2D image. The process of patch generation is divided into three steps: calculation of normal values, segmentation, and patch division, as follows.
[0131] Referring to FIG. 5, the process of calculating normal values will be specifically described.
[0132] FIG. 5 shows an example of the tangent plane and normal vector of a surface according to an embodiment.
[0133] The surface in FIG. 5 is used as follows in the process 40000 of patch generation in the V-PCC encoding process in FIG. 4.
[0134] Normal calculation related to patch generation
[0135] Each point (e.g., a point) that makes up the point cloud has a unique direction, which is represented by a 3D vector called the normal. Using the neighboring points (neighbors) of each point obtained using a K-D tree or the like, the tangent plane and the normal vector of each point forming the surface of the point cloud as shown in FIG. 5 are obtained. The search range in the process of searching for neighboring points is defined by the user.
[0136] Tangent plane: It refers to a plane that passes through a point on the surface and completely contains the tangent line to the curve on the surface.
[0137] FIG. 6 shows an example of a bounding box of a point cloud according to an embodiment.
[0138] The bounding box according to the embodiment is a unit box that divides point cloud data based on a hexahedron in 3D space.
[0139] The method / apparatus according to the embodiment, for example, patch generation 4000 uses the bounding box in the process of generating patches from point cloud data.
[0140] The bounding box is used in the process of projecting the target point cloud object of the point cloud data onto the planes of each hexahedron based on a hexahedron in 3D space. The bounding box is generated and processed by the point cloud video acquisition unit 10001 and the point cloud video encoder 10002 in FIG. 1. Also, based on the bounding box, patch generation 40000, patch packing 40001, geometry image generation 40002, and texture image generation 40003 of the V-PCC encoding process in FIG. 4 are performed.
[0141] Segmentation related to patch generation
[0142] Segmentation consists of two processes: initial segmentation and refine segmentation.
[0143] The point cloud video encoder 10002 according to the embodiment projects points onto one side of a bounding box. Specifically, each point forming the point cloud is projected onto one side of the six bounding box faces surrounding the point cloud as shown in FIG. 6. Initial segmentation is the process of determining one of the planes of the bounding box plane onto which each point is projected.
[0144] The normal values corresponding to each of the six planes are JPEG0007697119000001.jpg814 is defined as follows.
[0145] (1.0, 0.0, 0.0), (0.0, 1.0, 0.0), (0.0, 0.0, 1.0), (-1.0, 0.0, 0.0), (0.0, -1.0, 0.0), (0.0, 0.0, -1.0).
[0146] As in the following formula, the normal value of each point obtained from the above-described normal value calculation process ( JPEG0007697119000002.jpg89) and The plane with the largest dot product of JPEG0007697119000003.jpg914 is determined as the projection plane of that plane. That is, the plane having the normal in the direction most similar to the normal of the point is determined as the projection plane of that point.
[0147] JPEG0007697119000004.jpg1341
[0148] The determined plane is identified as a value of the cluster index in one of the index formats from 0 to 5.
[0149] Refine segmentation is a process of improving the projection plane of each point forming the point cloud determined in the above-described initial segmentation process in consideration of the projection planes of adjacent points. In this process, together with score normal representing the similarity between the normal of each point considered to determine the projection plane in the above-described initial segmentation process and the normal value of each plane of the bounding box, score smooth indicating the degree of coincidence between the projection plane of the current point and the projection planes of adjacent points is considered simultaneously.
[0150] Score smooth can be considered by giving a weight value to score normal, and at this time, the weight value is defined by the user. Refine segmentation is performed repeatedly, and the number of repetitions is also defined by the user.
[0151] Segment patches related to patch generation
[0152] Patch segmentation is a process of dividing the entire point cloud into patches, which are sets of adjacent points, based on the projection plane information of each point forming the point cloud obtained in the above-described initial / refine segmentation process. Patch segmentation consists of the following steps.
[0153] (1) Calculate the adjacent points of each point forming the point cloud using a K-D tree or the like. The maximum number of adjacent points is defined by the user.
[0154] (2) When the adjacent points are projected onto the same plane as the current point (when they have the same cluster index value), extract the current point and its adjacent points into one patch.
[0155] (3) Calculate the geometric values of the extracted patches.
[0156] (4) Repeat the processes of (2) to (3) until there are no unextracted points.
[0157] Through the patch splitting process, the size of each patch, the occupancy map of each patch, the geometry image, the texture image, etc. are determined.
[0158] FIG. 7 shows an example of the positioning of individual patches of an occupancy map according to an embodiment.
[0159] The point cloud encoder 10002 according to the embodiment can generate patch packing and an occupancy map.
[0160] Patch packing & Occupancy map generation 40001
[0161] This process is a process of determining the position of an individual patch within a 2D image of the patch in order to map the previously split patches to one 2D image. The occupancy map is one of the 2D images and is a binary map that indicates the presence or absence of data at that position with a value of 0 or 1. The occupancy map consists of blocks, and the resolution is determined according to the size of the blocks. For example, when the size of the block is 1*1, it has a resolution in units of pixels. The size of the occupancy packing block is determined by the user.
[0162] The process of determining the position of an individual patch within the occupancy map is as follows.
[0163] (1) Set all the values of the entire occupancy map to 0.
[0164] (2) Position the patch at the point (u, v) where the horizontal coordinate in the occupancy map plane is in the range [0, occupancySizeU - patch.sizeU0) and the vertical coordinate is in the range [0, occupancySizeV - patch.sizeV0).
[0165] (3) Set the point (x, y) where the horizontal coordinate in the patch plane is in the range [0, patch.sizeU0) and the vertical coordinate is in the range [0, patch.sizeV0) as the current point.
[0166] (4) For the point (x, y), if the (x, y) coordinate value of the patch occupancy map is 1 (data exists at the corresponding location within the patch) and the (u + x, v + y) coordinate value of the entire occupancy map is 1 (when the occupancy map is filled by the previous patch), change the (x, y) position in raster order and repeat the processes in (3) - (4). Otherwise, perform the process in (6).
[0167] (5) Change the (u, v) position in raster order and repeat the processes in (3) - (5).
[0168] (6) Determine (u, v) as the position of the corresponding patch, and assign (copy) the patch occupancy map data to the corresponding part of the entire occupancy map.
[0169] (7) Repeat the processes in (2) - (6) for the next patch.
[0170] Occupancy size U (occupancySizeU): Indicates the width of the occupancy map, with the unit being the occupancy packing block size.
[0171] Occupancy size V (occupancySizeV): Indicates the height of the occupancy map, with the unit being the occupancy packing block size.
[0172] Patch size U0 (patch.sizeU0): Indicates the width of the occupancy map, and the unit is the occupancy packing block size.
[0173] Patch size V0 (patch.sizeV0): Indicates the height of the occupancy map, and the unit is the occupancy packing block size.
[0174] For example, as shown in FIG. 7, there may be a box corresponding to a patch having an in-box patch size corresponding to the occupancy packing size block, and the in-box point (x, y) may be located.
[0175] FIG. 8 shows an example of the relationship of the normal, tangent, and bitangent axes according to an embodiment.
[0176] The point cloud video encoder 10002 according to the embodiment can generate a geometry image. The geometry image means image data including the geometry information of the point cloud. The generation process of the geometry image uses the three axes (normal, tangent, bitangent) of the patch in FIG. 8.
[0177] Geometry image generation 40002
[0178] In this process, the depth value constituting the geometry image of an individual patch is determined, and the overall geometry image is generated based on the position of the patch determined in the above-described patch packing process. The process of determining the depth value constituting the geometry image of an individual patch is configured as follows.
[0179] (1) Calculate the parameters regarding the position and size of an individual patch. The parameters include the following information. In one embodiment, the position of the patch is included in the patch information.
[0180] Index indicating the normal axis: The normal is obtained in the process of patch generation described above. The tangent axis is the axis that coincides with the horizontal (u) axis of the patch image among the axes perpendicular to the normal, and the bitangent axis is the axis that coincides with the vertical (v) axis of the patch image among the axes perpendicular to the normal. The three axes are shown as in Figure 8.
[0181] Figure 9 shows an example of the minimum mode and maximum mode configurations of the projection mode according to the embodiment.
[0182] The point cloud video encoder 10002 according to the embodiment performs patch-based projection to generate a geometry image, and the projection modes according to the embodiment are the minimum mode and the maximum mode.
[0183] 3D spatial coordinates of the patch: Calculated by the smallest bounding box surrounding the patch. For example, the 3D spatial coordinates of the patch include the minimum value in the tangent direction of the patch (patch 3D shift tangent axis), the minimum value in the bitangent direction of the patch (patch 3D shift bitangent axis), the minimum value in the normal direction of the patch (patch 3D shift normal axis), etc.
[0184] 2D size of the patch: Indicates the horizontal and vertical sizes when the patch is packed in a 2D image. The horizontal size (patch 2D size u) is the difference between the maximum value and the minimum value in the tangent direction of the bounding box, and the vertical size (patch 2D size v) is the difference between the maximum value and the minimum value in the bitangent direction of the bounding box.
[0185] (2) Determine the projection mode of the patch. The projection mode is either the min mode or the max mode. The geometric information of the patch is indicated by the depth value. When projecting each point forming the patch in the normal direction of the patch, two-layer images, i.e., an image composed of the maximum depth value and an image composed of the minimum depth value, are generated.
[0186] To generate the two-layer images d0 and d1, in the case of the min mode, as shown in FIG. 9, the minimum depth is configured as d0, and the maximum depth existing within the surface thickness from the minimum depth is configured as d1.
[0187] For example, when the point cloud is located in 2D as shown in FIG. 9, there may be a plurality of patches each including a plurality of points. As shown in FIG. 9, the points indicated by the same shading belong to the same patch. The process of projecting the patches of the points indicated by the blanks is shown.
[0188] When projecting the points indicated by the blanks to the left / right, with the left side as the reference, numbers for calculating the depth of the points are marked while increasing the depth one by one as 0, 1, 2,..6, 7, 8, 9 to the right.
[0189] The projection mode may be applied to all point clouds in the same way according to the user's definition, or different methods may be applied for each frame or patch. When different projection modes are applied for each frame or patch, a projection mode that can improve the compression efficiency or minimize the missed points is adaptively selected.
[0190] (3) Calculate the depth value of each individual point.
[0191] In the case of the minimum mode, a d0 image is constructed with a depth0 that is the value obtained by subtracting the minimum value of the normal direction of the patch (patch 3D shift normal axis) calculated in the process of (1) from the minimum value of the normal axis of each point. If there are other depth values within the range of the surface thickness from depth0 at the same position, this value is set as depth1. If not, the value of depth0 is also assigned to depth1. A d1 image is constructed with the value of depth1.
[0192] For example, the minimum value is calculated in determining the depth of the points of d0 (4 2 4 4 0 6 0 0 9 9 0 8 0). Also, when determining the depth of the points of d1, the larger value among two or more points is calculated, or if there is only one point, that value is calculated (4 4 4 4 6 6 6 8 9 9 8 8 9). Also, in the process where the points of the patch are encoded and reconstructed, some points are lost (for example, 8 points are lost in the figure).
[0193] In the case of the maximum mode, a d0 image is constructed with a depth0 that is the value obtained by subtracting the minimum value of the normal direction of the patch (patch 3D shift normal axis) calculated in the process of (1) from the maximum value of the normal axis of each point. If there are other depth values within the range of the surface thickness from depth0 at the same position, this value is set as depth1. If not, the value of depth0 is also assigned to depth1. A d1 image is constructed with the value of depth1.
[0194] For example, in determining the depth of the point of d0, the maximum value is calculated (4 4 4 4 6 6 6 8 9 9 8 8 9). Also, in determining the depth of the point of d1, the smaller value among two or more points is calculated, or if there is only one point, that value is calculated (4 2 4 4 5 6 0 6 9 9 0 8 0). Also, in the process where the points of the patch are encoded and reconstructed, some points are lost (for example, 6 points are lost in the figure).
[0195] The geometric image of the individual patch generated from the above-described process is arranged in the overall geometric image by using the position information of the individual patch generated through the above-described patch packing process, thereby generating the overall geometric image.
[0196] The d1 layer of the generated overall geometric image is encoded in various ways. The first is the method of directly encoding the depth value of the previously generated d1 image (absolute d1 encoding method). The second is the method of encoding the difference between the depth value of the previously generated d1 image and the depth value of the d0 image (differential encoding method).
[0197] Such a coding method using the depth values of the two layers of d0 and d1 loses the geometric information of that point in the process of encoding the geometric information of that point when there are other points between the two depths. Therefore, for lossless coding, an Enhanced-Delta-Depth (EDD) code may be used.
[0198] Referring to FIG. 10, the EDD code will be specifically described.
[0199] FIG. 10 shows an example of an EDD code according to an embodiment.
[0200] The point cloud video encoder 10002 and / or part / all of the processes of V-PCC encoding (e.g., video compression 40009) can encode the geometric information of points based on the EOD code.
[0201] The EDD code is a method of encoding the positions of all points within the surface thickness range including d1 in binary as shown in FIG. 10. As an example, for the points included in the second column from the left in FIG. 10, since there are points at the first and fourth positions above D0 and the second and third positions are empty, it is represented by the EDD code of 0b1001 (=9). When encoding and transmitting the EDD code together with D0, all the geometric information of the points can be restored without loss at the receiving end.
[0202] For example, if there is a point on the reference point, it is 1, and if there is no point, it is 0, and the code is represented based on four bits.
[0203] Smoothing 40004
[0204] Smoothing is an operation to remove the discontinuities that may occur at the patch boundaries due to the degradation of the image quality resulting from the compression process, and it is performed by the point cloud video encoder 10002 or the smoothing unit 40004 through the following process.
[0205] (1) Reconstruct the point cloud from the geometry image. This process can be said to be the reverse process of the above-mentioned geometry image generation. For example, the reverse process of encoding is reconstruction.
[0206] (2) Calculate the adjacent points of each point constituting the regenerated point cloud using a K-D tree or the like.
[0207] (3) For each point, determine whether the point is located on the patch boundary surface. As an example, if there is an adjacent point having a different projection plane (cluster index) from the current point, it can be determined that the point is located on the boundary surface of the patch.
[0208] (4) If the patch boundary surface exists, move the point to the centroid of the adjacent points (located at the average x, y, and z coordinates of the adjacent points). That is, change the geometry value. If it does not exist, maintain the previous geometry value.
[0209] FIG. 11 shows an example of recoloring using the color values of adjacent points according to the embodiment.
[0210] The point cloud video encoder 10002 or the texture image generation 40003 according to the embodiment can generate a texture image based on recoloring.
[0211] Texture image generation 40003
[0212] The process of texture image generation consists of a process of generating a texture image for the whole by generating a texture image for each individual patch and arranging them at determined positions, similar to the process of geometry image generation described above. However, in the process of generating a texture image for each individual patch, an image having the color values (e.g., R, G, B) of the points constituting the point cloud corresponding to that position is generated instead of the depth value for geometry generation.
[0213] In the process of obtaining the color values of each point constituting the point cloud, the geometry that has undergone the above-described smoothing process is used. Since the smoothed point cloud may be in a state where the positions of some points have moved in the original point cloud, a color restoration process for finding a color suitable for the changed position is required. Color restoration is performed using the color values of adjacent points. As an example, as shown in FIG. 11, the new color value can be calculated in consideration of the color values of the closest adjacent points and the color values of the adjacent points.
[0214] For example, referring to FIG. 11, color restoration calculates a color value suitable for the changed position based on the average of the characteristic information of the original point closest to the point and / or the average of the characteristic information of the original position closest to the point.
[0215] The texture image is also generated in two layers, t0 / t1, like the geometry image generated in two layers, d0 / d1.
[0216] Auxiliary patch info compression 40005
[0217] The point cloud video encoder 10002 or the additional patch information compression unit 40005 according to the embodiment can compress additional patch information (additional information related to the point cloud).
[0218] The additional patch information compression unit 40005 compresses the additional patch information generated in the processes of patch generation, patch packing, geometry generation, etc. described above. The additional patch information includes the following parameters:
[0219] An index (cluster index) for identifying the projection plane (normal)
[0220] 3D spatial position of the patch: minimum value of the tangent direction of the patch (patch 3D shift tangent axis), minimum value of the bitangent direction of the patch (patch 3D shift bitangent axis), minimum value of the normal direction of the patch (patch 3D shift normal axis)
[0221] 2D spatial position and size of the patch: horizontal size (patch 2D size u), vertical size (patch 2D size v), minimum value in the horizontal direction (patch 2D shift u), minimum value in the vertical direction (patch 2D shift u)
[0222] Mapping information for each block and patch: candidate index (when patches are sequentially positioned based on the 2D spatial position and size information of the patch described above, multiple patches may be mapped to one block overlappingly. At this time, the mapped patches form a candidate list, and the index indicating which patch data in this list exists in the corresponding block), local patch index (index indicating one of the entire patches existing in the frame). Table 1 is pseudo code showing the matching process of blocks and patches using the candidate list and local patch index).
[0223] The maximum number of the candidate list is defined by the user.
[0224]
Table 1
[0225] Figure 12 shows an example of push-pull background filling according to an embodiment.
[0226] Image padding and group dilation 40006, 40007, 40008
[0227] The image padding according to the embodiment can fill the space outside the patch region with meaningless additional data based on the push-pull background filling method.
[0228] The image paddings 40006 and 40007 are processes that fill the space outside the patch region with meaningless data for the purpose of improving the compression efficiency. For image padding, a method is used in which the pixel values of the columns or rows corresponding to the boundary surface side inside the patch are copied to fill the empty space. Alternatively, as shown in FIG. 12, in the process of gradually reducing the resolution of the unpadded image and then increasing the resolution again, a push-pull background filling method may be used to fill the empty space with pixel values from the low-resolution image.
[0229] Group dilation 40008 is a method of filling the empty space of the geometry and texture image consisting of two layers of d0 / d1 and t0 / t1, and is a process of filling the values of the empty space of the two layers calculated by the above-described image padding with the average value of the values for the same position of the two layers.
[0230] FIG. 13 shows an example of a possible traversal order for a 4*4 size block according to the embodiment.
[0231] Occupancy map compression 40012, 40011
[0232] The occupancy map compression according to the embodiment is a process of compressing the above-described occupancy map, and there are two methods: video compression for lossy compression and entropy compression for lossless compression. Video compression will be described later.
[0233] The process of entropy compression is performed as follows.
[0234] (1) For each block that makes up the occupancy map, if all blocks are filled, encode it as 1 and repeat the same process for the next block. Otherwise, encode it as 0 and perform the processes of (2) to (5).
[0235] (2) Determine the best traversal order for performing run - length coding on the filled pixels of the block. FIG. 13 shows, as an example, four possible traversal orders for a 4×4 - sized block.
[0236] FIG. 14 shows an example of the best traversal order according to an embodiment.
[0237] As described above, the entropy compression unit 40012 according to the embodiment can coat (encode) the blocks based on the traversal order method as shown in FIG. 14.
[0238] For example, among the possible traversal orders, select the best traversal order having the minimum number of runs and encode its index. As an example, in the case of selecting the third traversal order of FIG. 13 described above, in this case, the number of runs can be minimized to 2, and this is selected as the best traversal order.
[0239] Encode the number of runs. In the example of FIG. 14, since there are two runs, encode 2.
[0240] (4) Encode the occupancy of the first run. In the example of FIG. 14, since the first run corresponds to unfilled pixels, encode 0.
[0241] Encode the length for each individual run (as many lengths as the number of runs). In the example of FIG. 14, the lengths of the first run and the second run, which are 6 and 10, are encoded sequentially.
[0242] Video compression 40009, 40010, 40011
[0243] The video compression units (40009, 40010, 40011) according to the embodiments encode sequences such as the geometry image, texture image, occupancy map image, etc. generated by the above-described process using a 2D video codec such as HEVC, VVC.
[0244] FIG. 15 shows an example of a 2D video / image encoder according to an embodiment, which is also referred to as an encoding device.
[0245] FIG. 15 shows a schematic block diagram of a 2D video / image encoder 15000 to which the above-described video compression units 40009, 40010, 40011 are applied, where encoding of a video / video signal is performed. The 2D video / image encoder 15000 is included in the above-described point cloud video encoder 10002 or consists of internal / external components. Each component in FIG. 15 corresponds to software, hardware, a processor, and / or a combination thereof.
[0246] Here, the input image may be one of the geometric image, texture image (feature image), and occupancy map image described above. When the 2D video / image encoder in FIG. 15 is applied to the video compression unit 40009, the image input to the 2D video / image encoder 15000 is a padded geometric image, and the bitstream output from the 2D video / image encoder 15000 is a bitstream of the compressed geometric image. When the 2D video / image encoder in FIG. 15 is applied to the video compression unit 40010, the image input to the 2D video / image encoder 15000 is a padded texture image, and the bitstream output from the 2D video / image encoder 15000 is a bitstream of the compressed texture image. When the 2D video / image encoder in FIG. 15 is applied to the video compression unit 40011, the image input to the 2D video / image encoder 15000 is an occupancy map image, and the bitstream output from the 2D video / image encoder 15000 is a bitstream of the compressed occupancy map image.
[0247] The inter prediction unit 15090 and the intra prediction unit 15100 are collectively referred to as a prediction unit. That is, the prediction unit includes the inter prediction unit 15090 and the intra prediction unit 15100. The conversion unit 15030, the quantization unit 15040, the inverse quantization unit 15050, and the inverse conversion unit 15060 are also collectively referred to as a residual processing unit. The residual processing unit may further include a subtraction unit 15020. According to an embodiment, the video segmentation unit 15010, the subtraction unit 15020, the conversion unit 15030, the quantization unit 15040, the inverse quantization unit 15050, the inverse conversion unit 15060, the addition unit 155, the filtering unit 15070, the inter prediction unit 15090, the intra prediction unit 15100, and the entropy encoding unit 15110 in FIG. 15 are composed of one hardware component (for example, an encoder or a processor). Further, the memory 15080 includes a DPB (decoded picture buffer) and is composed of a digital storage medium.
[0248] The video segmentation unit 15010 divides the input video (or picture, frame) input to the encoding device 15000 into one or more processing units. As an example, the processing unit is also referred to as a coding unit (CU). In this case, the coding unit is recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) by a Quad-tree binary-tree (QTBT) structure. For example, based on one coding unit, a Quad-tree structure, and / or a binary-tree structure, it is divided into a plurality of coding units at a deeper depth. In this case, for example, the Quad-tree may be applied first, and then the binary-tree may be applied. Or the binary-tree may be applied first. Based on the final coding unit that cannot be further divided, the coding procedure according to this specification may be performed. In this case, based on the coding efficiency according to the characteristics of the video, etc., the largest coding unit may be used as the final encoding unit, or if necessary, the coding unit is recursively divided into coding units at a deeper depth, and the coding unit of the optimal size is used as the final coding unit. Here, the coding procedure includes procedures such as prediction, transformation, and restoration described later. As another example, the processing unit may further include a prediction unit (PU: Prediction Unit) or a transform unit (TU: Transform Unit). In this case, each of the prediction unit and the transform unit is divided or partitioned from the above-described final coding unit. The prediction unit is a unit of sample prediction, and the transform unit is a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0249] The unit is sometimes used interchangeably with terms such as block, area, or module. In general, an MxN block represents a set of samples or transform coefficients consisting of M columns and N rows. Samples generally represent pixels or pixel values, and may represent only the pixel / pixel values of the luma component, or only the pixel / pixel values of the chroma component. Samples are used as terms corresponding to pixels or pels in one picture (or video).
[0250] The subtraction unit 15020 of the encoding device 15000 subtracts the prediction signal (predicted block, predicted sample array) output from the inter prediction unit 15090 or the intra prediction unit 15100 in the input video signal (original block, original sample array) to generate a residual signal (residual signal, residual block, residual sample array), and the generated residual signal is transmitted to the conversion unit 15030. In this case, as shown in the figure, the unit that subtracts the prediction signal (predicted block, predicted sample array) in the input video signal (original block, original sample array) within the encoding device 15000 can be called the subtraction unit 15020. The prediction unit performs prediction on the processing target block (hereinafter referred to as the current block), and generates a predicted block including the prediction samples for the current block. The prediction unit determines whether to apply intra prediction or inter prediction in units of the current block or CU. The prediction unit generates various information related to prediction, such as prediction mode information, and transmits it to the entropy encoding unit 15110 as described later for each prediction mode. The information related to prediction is encoded by the entropy encoding unit 15110 and output in the form of a bitstream.
[0251] The intra prediction unit 15100 of the prediction unit predicts the current block by referring to samples within the current picture. The samples to be referred to are located adjacent to or away from the current block according to the prediction mode. In intra prediction, the prediction mode includes a plurality of non-directional modes and a plurality of directional modes. The non-directional modes include, for example, the DC mode and the Planar mode. The directional modes include, for example, 33 directional prediction modes or 65 directional prediction modes according to the precision of the prediction direction. However, this is an example, and more or fewer directional prediction modes may be used depending on the setting. The prediction mode applied to the current block may be determined using the prediction mode applied to the adjacent blocks of the intra prediction unit 15100.
[0252] The inter-prediction unit 15090 of the prediction unit derives a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter-prediction mode, motion information is predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between adjacent blocks and the current block. Motion information includes a motion vector and a reference picture index. The motion information further includes inter-prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-prediction, adjacent blocks include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. Temporal neighboring blocks are called collocated reference blocks, collocated CUs, etc., and the reference picture including the temporal neighboring block is also called a collocated picture (colPic). For example, the inter-prediction unit 15090 constructs a candidate list of motion information based on adjacent blocks and generates information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter-prediction is performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter-prediction unit 15090 uses the motion information of adjacent blocks as the motion information of the current block. In the case of the skip mode, unlike the merge mode, a residual signal may not be transmitted. In the case of the motion vector prediction (MVP) mode, the motion vector of an adjacent block is used as a motion vector predictor, and the motion vector of the current block is indicated by signaling a motion vector difference.
[0253] The prediction signal generated by the inter prediction unit 15090 or the intra prediction unit 15100 is used for generating a restored signal or for generating a residual signal.
[0254] The conversion unit 15030 applies a conversion method to the residual signal to generate transform coefficients. For example, the conversion method includes at least one of DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT means a transform obtained from a graph when representing the relationship information between pixels by a graph. CNT means a transform obtained based on generating a prediction signal using all previously reconstructed pixels. Also, the conversion process may be applied to a pixel block of the same size in a square shape or may be applied to a block of a variable size that is not square.
[0255] The quantization unit 15040 quantizes the transform coefficients and transmits them to the entropy encoding unit 15110, and the entropy encoding unit 15110 encodes the quantized signal (information regarding the quantized transform coefficients) and outputs it to a bit stream. The information regarding the quantized transform coefficients is called residual information. The quantization unit 15040 can also reorder the block-shaped quantized transform coefficients into a one-dimensional vector form based on the coefficient scan order, and generate information regarding the quantized transform coefficients based on the one-dimensional vector form of the quantized transform coefficients.
[0256] The entropy encoding unit 15110 performs various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit 15110 encodes, together or separately, information necessary for video / image restoration (such as values of syntax elements) in addition to the quantized transform coefficients. The encoded information (e.g., encoded video / video information) is transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units.
[0257] The bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network includes a broadcast network and / or a communication network, etc., and the digital storage medium includes various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) for transmitting and / or a storage unit (not shown) for storing the signal output from the entropy encoding unit 15110 may be configured as internal / external elements of the encoding device 15000, or the transmission unit may be included in the entropy encoding unit 15110.
[0258] The quantized transform coefficients output from the quantization unit 15040 are used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients by the inverse quantization unit 15040 and the inverse transformation unit 15060, a residual signal (residual block or residual sample) is restored. The addition unit 15200 adds the restored residual signal to the prediction signal output from the inter prediction unit 15090 or the intra prediction unit 15100 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). When there is no residual for the processing target block, as in the case where the skip mode is applied, the predicted block is used as the reconstructed block. The addition unit 15200 is called a restoration unit or a reconstructed block generation unit. The generated reconstructed signal may be used for intra prediction of the next processing target block within the current picture, or may be used for inter prediction of the next picture after passing through filtering as described later.
[0259] The filtering unit 15070 can apply filtering to the reconstructed signal output from the addition unit 15200 to improve subjective / objective image quality. For example, the filtering unit 15070 applies various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and stores the modified reconstructed picture in the memory 15080, specifically in the DPB of the memory 15080. Various filtering methods include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, and the like. The filtering unit 15070 generates various information related to filtering and transmits it to the entropy encoding unit 15110, like each filtering method described later. The information related to filtering is encoded by the entropy encoding unit 15110 and output in the form of a bitstream.
[0260] The modified restored picture stored in memory 15080 is used as a reference picture in the inter prediction unit 15090. When inter prediction is applied by the encoder in this way, prediction mismatches in the encoder 15000 and the decoder can be avoided, and the encoding efficiency can also be improved.
[0261] The DPB of memory 15080 stores the modified restored picture for use as a reference picture in the inter prediction unit 15090. Memory 15080 stores the motion information of the blocks for which the motion information in the current picture has been derived (or encoded) and / or the motion information of the blocks in the already restored pictures. The stored motion information is transmitted to the inter prediction unit 15090 for utilization as the motion information of spatially adjacent blocks or temporally adjacent blocks. Memory 15080 stores the restored samples of the blocks restored in the current picture and transmits them to the intra prediction unit 15100.
[0262] Note that at least one of the above-described prediction, transformation, and quantization procedures may be omitted. For example, for the blocks to which PCM (pulse code modulation) is applied, the prediction, transformation, and quantization procedures may be omitted, and the values of the original samples may be directly encoded and output to the bitstream.
[0263] FIG. 16 shows an example of a V-PCC decoding process according to an embodiment.
[0264] The V-PCC decoding process or the V-PCC decoder is the reverse process of the V-PCC encoding process (or encoder) in FIG. 4. Each component in FIG. 16 corresponds to software, hardware, a processor, and / or a combination thereof.
[0265] The demultiplexer 16000 demultiplexes the compressed bitstream and outputs a compressed texture image, a compressed geometry image, a compressed occupancy map image, and compressed additional patch information respectively.
[0266] The video decompression units 16001 and 16002 decompress the compressed texture image and the compressed geometry image respectively.
[0267] The occupancy map decompression unit 16003 decompresses the compressed occupancy map image.
[0268] The auxiliary patch information decompression unit 16004 decompresses the compressed additional patch information.
[0269] The geometry reconstruction unit 16005 restores (reconstructs) the geometry information based on the restored geometry image, the restored occupancy map, and / or the restored additional patch information. For example, it reconstructs the geometry changed in the encoding process.
[0270] The smoothing unit 16006 applies smoothing to the reconstructed geometry. For example, smoothing filtering is applied.
[0271] The texture reconstruction unit 16007 reconstructs the texture from the restored texture image and / or the smoothed geometry.
[0272] The color smoothing (color smoothing or color smoothing unit) 16008 smooths the color values from the reconstructed texture. For example, smoothing filtering is applied.
[0273] As a result, the reconstructed point cloud data is generated.
[0274] FIG. 16 shows the decoding process of V-PCC for reconstructing a point cloud by restoring (or decoding) a compressed occupancy map, a geometry image, a texture image, and additional patch information.
[0275] Each unit shown in FIG. 16 operates as at least one of a processor, software, and hardware. The detailed operations of each unit in FIG. 16 according to the embodiment are as follows.
[0276] Video decompression 16001, 16002
[0277] It is a reverse process of the above-described video compression, and is a process of decoding by reversing the process of video-compressing the bitstream of the geometry image, the bitstream of the compressed texture image, and / or the bitstream of the compressed occupancy map image generated in the above process using a 2D video codec such as HEVC or VVC.
[0278] FIG. 17 shows an example of a 2D video / image decoder according to an embodiment, which is also called a decoding device.
[0279] The 2D video / image decoder is a reverse process of the 2D video / image encoder in FIG. 15.
[0280] The 2D video / image decoder of FIG. 17 is an example of the video decompression units 16001 and 16002 of FIG. 16, and shows a schematic block diagram of the 2D video / image decoder 17000 in which decoding of video / video signals is performed. The 2D video / image decoder 17000 may be included in the above-described point cloud video decoder 10008, or may be configured as an internal / external component. Each component in FIG. 17 corresponds to software, hardware, a processor, and / or a combination thereof.
[0281] Here, the input bitstream is one of the bitstreams of the geometry image, the texture image (attribute(s) image), and the occupancy map image. When the 2D video / image decoder of FIG. 17 is applied to the video decompression unit 16001, the bitstream input to the 2D video / image decoder is the bitstream of the compressed texture image, and the restored image output from the 2D video / image decoder is the restored texture image. When the 2D video / image decoder of FIG. 17 is applied to the video decompression unit 16002, the bitstream input to the 2D video / image decoder is the bitstream of the compressed geometry image, and the restored image output from the 2D video / image decoder is the restored geometry image. The 2D video / image decoder of FIG. 17 is input with and restores the bitstream of the compressed occupancy map image. The restored video (or output video, decoded video) shows the restored video for the above-described geometry image, texture image (attribute(s) image), and occupancy map image.
[0282] Referring to FIG. 17, the inter prediction unit 17070 and the intra prediction unit 17080 are collectively referred to as a prediction unit. That is, the prediction unit includes the inter prediction unit 17070 and the intra prediction unit 17080. The inverse quantization unit 17020 and the inverse transform unit 17030 are collectively referred to as a residual processing unit. That is, the residual processing unit includes the inverse quantization unit 17020 and the inverse transform unit 17030. According to an embodiment, the entropy decoding unit 17010, the inverse quantization unit 17020, the inverse transform unit 17030, the addition unit 17040, the filtering unit 17050, the inter prediction unit 17070, and the intra prediction unit 17080 in FIG. 17 are constituted by one hardware component (for example, a decoder or a processor). Also, the memory 17060 may include a DPB (decoded picture buffer) or may be constituted by a digital storage medium.
[0283] When a bitstream including video / video information is input, the decoding apparatus 17000 restores the video corresponding to the process in which the video / video information was processed in the encoding apparatus of FIG. 15. For example, the decoding apparatus 17000 performs decoding using the processing units applied in the encoding apparatus. Thus, the decoding processing unit is, for example, a coding unit, and the coding unit is divided from a coding tree unit or a maximum coding unit by a Quad-tree structure and / or a binary-tree structure. Also, the restored video signal decoded and output by the decoding apparatus 17000 is reproduced by a reproducing apparatus.
[0284] The decoding device 17000 receives the signal output from the encoding device in the form of a bit stream, and the received signal is decoded by the entropy decoding unit 17010. For example, the entropy decoding unit 17010 parses the bit stream to derive information (e.g., video / video information) necessary for video restoration (or picture restoration). For example, the entropy decoding unit 17010 decodes the information in the bit stream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and outputs the value of the syntax element necessary for video restoration and the quantized value of the transform coefficient regarding the residue. More specifically, the CABAC entropy decoding method receives the bins corresponding to each syntax element in the bit stream, determines a context model using the information of the syntax element to be decoded, the adjacent information of the block to be decoded, and the decoded information of the block to be decoded or the symbol / bin information decoded in the previous step, predicts the occurrence probability of the bin according to the determined context model, performs arithmetic decoding of the bin, and generates a symbol corresponding to the value of each syntax element. At this time, after determining the context model, the CABAC entropy decoding method updates the context model using the symbol / bin information decoded for the context model of the next symbol / bin. Among the information decoded by the entropy decoding unit 17010, the information regarding prediction is provided to the prediction units (inter prediction unit 17070 and intra prediction unit 17080), and the residue value for which entropy decoding has been performed by the entropy decoding unit 17010, that is, the quantized transform coefficient and related parameter information, are input to the inverse quantization unit 17020. Also, among the information decoded by the entropy decoding unit 17010, the information regarding filtering is provided to the filtering unit 17050. On the other hand, a receiving unit (not shown) that receives the signal output from the encoding device may be further configured as an internal / external element of the decoding device 17000, and the receiving unit may be a component of the entropy decoding unit 17010.
[0285] In the inverse quantization unit 17020, the quantized transform coefficients are dequantized to output the transform coefficients. The inverse quantization unit 17020 rearranges the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement is performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit 17020 performs inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) to obtain the transform coefficients.
[0286] In the inverse transform unit 17030, the transform coefficients are inversely transformed to obtain the residual signal (residual block, residual sample array).
[0287] The prediction unit performs prediction on the current block and generates a predicted block including the predicted samples for the current block. The prediction unit determines whether intra prediction or inter prediction is applied to the current block based on the prediction-related information output from the entropy decoding unit 17010, and determines a specific intra / inter prediction mode.
[0288] The intra prediction unit 17080 of the prediction unit predicts the current block by referring to the samples within the current picture. The samples to be referred to may be located adjacent to or away from the current block depending on the prediction mode. In intra prediction, the prediction mode includes a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit 17080 determines the prediction mode to be applied to the current block using the prediction mode applied to the adjacent blocks.
[0289] The inter prediction unit 17070 of the prediction unit derives a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on the reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information is predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between adjacent blocks and the current block. The motion information includes a motion vector and a reference picture index. The motion information further includes information on an inter prediction method (such as L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, adjacent blocks include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit 17070 constructs a motion information candidate list based on adjacent blocks, and derives the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction is performed based on various prediction modes, and the information regarding the prediction includes information indicating the mode of inter prediction for the current block.
[0290] The addition unit 17040 generates a restored signal (restored picture, restored block, restored sample array) by adding the residual signal obtained by the inverse transform unit 17030 to the prediction signal (predicted block, predicted sample array) output from the inter prediction unit 17070 or the intra prediction unit 17080. When there is no residual for the processing target block, as in the case where the skip mode is applied, the predicted block is used as the restored block.
[0291] The addition unit 17040 is also called a restoration unit or a restored block generation unit. The generated restored signal may be used for intra prediction of the next processing target block within the current picture, and may also be used for inter prediction of the next picture after passing through filtering as described later.
[0292] Filtering unit 17050 applies filtering to the restored signal output from addition unit 17040 to improve subjective / objective image quality. For example, filtering unit 17050 applies various filtering methods to the restored picture to generate a modified restored picture, and transmits the modified restored picture to memory 17060, specifically to the DPB of memory 17060. Various filtering methods include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, and the like.
[0293] The (modified) restored picture stored in the DPB of memory 17060 is used as a reference picture in inter prediction unit 17070. Memory 17060 stores the motion information of the blocks for which the motion information in the current picture has been derived (or decoded) and / or the motion information of the blocks in the already restored picture. The stored motion information is transmitted to inter prediction unit 17070 for utilization as the motion information of spatially adjacent blocks or temporally adjacent blocks. Memory 17060 stores the restored samples of the restored blocks in the current picture and transmits them to intra prediction unit 17080.
[0294] In this specification, the embodiments described in filtering unit 15070, inter prediction unit 15090, and intra prediction unit 15100 of encoding apparatus 15000 in FIG. 15 can also be applied to filtering unit 17050, inter prediction unit 17070, and intra prediction unit 17080 of decoding apparatus 17000 in the same or corresponding manner.
[0295] On the other hand, at least one of the procedures of prediction, inverse transformation, and inverse quantization described above may be omitted. For example, for blocks to which PCM (pulse code modulation) is applied, the procedures of prediction, inverse transformation, and inverse quantization are omitted, and the values of the decoded samples are directly used as the samples of the restored video.
[0296] Occupancy map decompression 16003
[0297] It is the reverse process of the occupancy map compression described above, and it is a process of decoding the compressed occupancy map bit stream to restore the occupancy map.
[0298] Auxiliary patch info decompression 16004
[0299] It is the reverse process of the additional patch information compression described above, and it is a process of decoding the compressed additional patch information bit stream to restore the additional patch information.
[0300] Geometry reconstruction 16005
[0301] It is the reverse process of the geometry image generation described above. First, patches are extracted from the geometry image using the 2D position / size information of the patches and the mapping information between the blocks and the patches included in the restored occupancy map and the additional patch information. After that, using the geometry image of the extracted patches and the 3D position information of the patches included in the additional patch information, the point cloud is restored in the 3D space. If the geometry value corresponding to an arbitrary point (u, v) existing in one patch is denoted as g(u, v), and the coordinate values of the normal axis, tangent axis, and bitangent axis of the position of the patch in the 3D space are (δ0, s0, r0), then the coordinate values of the normal axis, tangent axis, and bitangent axis of the position in the 3D space mapped to the point (u, v), δ(u, v), s(u, v), r(u, v), are shown as follows.
[0302] δ(u,v) = δ0 + g(u, v)
[0303] s(u,v) = s0 + u
[0304] r(u,v) = r0 + v
[0305] Smoothing 16006
[0306] This is a process similar to the smoothing in the above-described encoding process, and is for removing discontinuities that may occur from the patch boundary surface due to image quality degradation occurring in the compression process.
[0307] Texture reconstruction 16007
[0308] This is a process of restoring a color point cloud by assigning a color value to each point constituting the smoothed point cloud. It is performed by using the mapping information between the geometry image and the point cloud reconstructed by the above-described diorama reconstruction process, and assigning the color value corresponding to the texture image pixel at the same position as the geometry image in the 2D space to the point of the point cloud corresponding to the same position in the 3D space.
[0309] Color smoothing 16008
[0310] This is a process similar to the above-described geometry smoothing process, and is for removing discontinuities of color values that may occur from the patch boundary surface due to image quality degradation generated from the compression process. Color smoothing is performed as follows.
[0311] (1) Calculate the adjacent points of each point constituting the restored color point cloud using a K-D tree or the like. The adjacent point information calculated in the above-described geometry smoothing process may be used as it is.
[0312] (2) Determine whether or not each point is located on the patch boundary surface. The boundary surface information calculated in the above-described geometry smoothing process may be used as it is.
[0313] (3) For adjacent points of the points existing on the boundary surface, examine the distribution of color values to determine whether to perform smoothing. As an example, when the entropy of the luminance value is less than or equal to the threshold local entry (when there are many similar luminance values), it is determined that it is not an edge part and smoothing is performed. As a method of smoothing, there is a method of replacing the color value of that point with the average value of adjacent points.
[0314] FIG. 18 shows an example of the flow of operations of a transmission device for compressing and transmitting V-PCC-based point cloud data according to an embodiment.
[0315] The transmission device according to the embodiment may correspond to the transmission device in FIG. 1, the encoding process in FIG. 4, the 2D video / image encoder in FIG. 15, or perform some / all of their operations. Each component of the transmission device corresponds to software, hardware, a processor, and / or a combination thereof.
[0316] The operations at the transmission end for compressing and transmitting point cloud data using V-PCC seem to be as shown in the figure.
[0317] The point cloud data transmission device according to the embodiment is called a transmission device, a transmission system, etc.
[0318] The patch generation unit 18000 receives point cloud data and generates patches for 2D image mapping of the point cloud. As a result of patch generation, patch information and / or additional patch information are generated, and the generated patch information and / or additional patch information are used in geometry image generation, texture image generation, smoothing, or the geometry restoration process for smoothing.
[0319] The patch packing unit 18001 performs the process of patch packing that maps the patches generated by the patch generation unit 18000 into the 2D image. For example, one or more patches are packed. An occupancy map is generated as a result of patch packing, and the occupancy map is used for the geometry image generation, geometry image partitioning, texture image partitioning, and / or geometry restoration process for smoothing.
[0320] The geometry image generation unit 18002 generates a geometry image using the point cloud data, patch information (or additional patch information), and / or occupancy map. The generated geometry image is preprocessed by the pre-encoding processing unit 18003 and then encoded into one bitstream by the video encoding unit 18006.
[0321] The pre-encoding processing unit 18003 includes image padding. That is, a part of the space of the generated geometry image and the generated texture image is padded with meaningless data. The pre-encoding processing unit 18003 may further include a process of group dilation for the generated texture image or the texture image with image padding performed.
[0322] The geometry restoration unit 18010 reconstructs a 3D geometry image using the geometry bitstream encoded by the video encoding unit 18006, additional patch information, and / or occupancy map.
[0323] The smoothing unit 18009 smooths the 3D geometry image reconstructed and output by the geometry restoration unit 18010 based on the additional patch information and outputs it to the texture image generation unit 18004.
[0324] The texture image generation unit 18004 generates a texture image using the smoothed 3D geometry, point cloud data, patches (or packed patches), patch information (or additional patch information), and / or occupancy map. The generated texture image is preprocessed by the pre-encoding processing unit 18003 and then encoded into one video bitstream by the video encoding unit 18006.
[0325] The metadata encoding unit 18005 encodes the additional patch information into one metadata bitstream.
[0326] The video encoding unit 18006 encodes the geometry image and the texture image output from the pre-encoding processing unit 18003 into respective video bitstreams, and encodes the occupancy map into one video bitstream. In one embodiment, the video encoding unit 18006 applies the 2D video / image encoder of FIG. 15 to each input image for encoding.
[0327] The multiplexing unit 18007 multiplexes the video bitstream of the geometry, the video bitstream of the texture image, the video bitstream of the occupancy map, and the metadata (including additional patch information) bitstream output from the metadata encoding unit 18005 into one bitstream.
[0328] The transmission unit 18008 transmits the bitstream output from the multiplexing unit 18007 to the receiving end. Alternatively, a file / segment encapsulation unit may be further provided between the multiplexing unit 18007 and the transmission unit 18008 to encapsulate the bitstream output from the multiplexing unit 18007 in the form of a file and / or segment and output it to the transmission unit 18008.
[0329] The patch generation unit 18000, patch packing unit 18001, geometry image generation unit 18002, texture image generation unit 18004, metadata encoding unit 18005, and smoothing unit 18009 in FIG. 18 respectively correspond to the patch generation unit 40000, patch packing unit 40001, geometry image generation unit 40002, texture image generation unit 40003, additional patch information compression unit 40005, and smoothing unit 40004 in FIG. 4. Further, the pre-encoding processing unit 18003 in FIG. 18 may include the image padding units 40006 and 40007 and the group expansion unit 40008 in FIG. 4, and the video encoding unit 18006 in FIG. 18 may include the video compression units 40009, 40010, 40011 and / or the entropy compression unit 40012 in FIG. 4. Therefore, for parts not described in FIG. 18, reference may be made to the descriptions in FIGS. 4 to 15. The above-described blocks may be omitted or may be replaced by blocks having similar or identical functions. Also, each block shown in FIG. 18 can operate as at least one of a processor, software, and hardware. Alternatively, the generated geometry, texture image, video bitstream of the occupancy map, and the additional patch information metadata bitstream are generated as one or more track data files or encapsulated in segments and transmitted from the transmitting unit to the receiving end.
[0330] Operation process of the receiving device
[0331] FIG. 19 shows an example of the operation flow of a receiving device for receiving and restoring V-PCC-based point cloud data according to an embodiment.
[0332] The receiving device according to the embodiment corresponds to the receiving device in FIG. 1, the decoding process in FIG. 16, and the 2D video / image encoder in FIG. 17, or performs some / all of their operations. Each component of the receiving device corresponds to software, hardware, a processor, and / or a combination thereof.
[0333] The operation process of the receiving end for receiving and restoring point cloud data using V-PCC follows the drawings. The operation of the V-PCC receiving end is the reverse process of the operation of the V-PCC transmitting end in FIG. 18.
[0334] The point cloud data receiving device according to the embodiment is called a receiving device, a receiving system, etc.
[0335] The receiving unit receives the bitstream of the point cloud (i.e., the compressed bitstream), and the demultiplexing unit 19000 demultiplexes the bitstream of the texture image, the bitstream of the geometry image, the bitstream of the occupancy map image, and the bitstream of the metadata (i.e., additional patch information) from the received point cloud bitstream. The demultiplexed bitstreams of the texture image, the geometry image, and the occupancy map image are output to the video decoding unit 19001, and the bitstream of the metadata is output to the metadata decoding unit 19002.
[0336] When the transmitting device in FIG. 18 is provided with a file / segment encapsulation unit, an example is to provide a file / segment decapsulation unit between the receiving unit of the receiving device in FIG. 19 and the demultiplexing unit 19000. In this case, in the transmitting device, the point cloud bitstream is encapsulated and transmitted in the form of a file and / or a segment, and in the receiving device, an example is to receive and decapsulate the file and / or segment including the point cloud bitstream.
[0337] The video decoding unit 19001 decodes the bitstreams of the geometry image, the texture image, and the occupancy map image into the geometry image, the texture image, and the occupancy map image, respectively. In one embodiment, the video decoding unit 19001 applies the 2D video / image decoder of FIG. 17 to each of the input bitstreams for decoding. The metadata decoding unit 19002 decodes the bitstream of the metadata into additional patch information and outputs it to the geometry restoration unit 19003.
[0338] Based on the geometry image, the occupancy map, and / or the additional patch information output from the video decoding unit 19001 and the metadata decoding unit 19002, the geometry restoration unit 19003 restores (reconstructs) the 3D geometry.
[0339] The smoothing unit 19004 smooths the 3D geometry reconstructed by the geometry restoration unit 19003.
[0340] The texture restoration unit 19005 restores the texture using the texture image and / or the smoothed 3D geometry output from the video decoding unit 19001. That is, the texture restoration unit 19005 assigns color values to the smoothed 3D geometry using the texture image to restore the color point cloud video / picture. Thereafter, in order to improve the objective / subjective visual quality, the color smoothing unit 19006 further performs color smoothing on the color point cloud video / picture. The modified point cloud video / picture derived thereby is shown to the user after the rendering process of the point cloud renderer (19007). Note that the color smoothing process may be omitted in some cases.
[0341] The above-described blocks may be omitted or replaced with blocks having similar or identical functions. Also, each block shown in FIG. 19 can operate as at least one of a processor, software, and hardware.
[0342] FIG. 20 shows an example of an architecture for storing and streaming V-PCC-based point cloud data according to an embodiment.
[0343] Part or all of the system of FIG. 20 includes part or all of the transmission / reception device of FIG. 1, the encoding process of FIG. 4, the 2D video / image encoder of FIG. 15, the decoding process of FIG. 16, the transmission device of FIG. 18, and / or the reception device of FIG. 19. Each component in the drawings corresponds to software, hardware, a processor, and combinations thereof.
[0344] FIG. 20 shows an overall architecture for storing or streaming point cloud data compressed based on video-based point cloud compression (V-PCC). The process of storing and streaming point cloud data can include an acquisition process, an encoding process, a transmission process, a decoding process, a rendering process, and / or a feedback process.
[0345] The embodiment proposes a method for efficiently providing point cloud media / content / data.
[0346] The point cloud acquisition unit 20000 first acquires a point cloud video in order to efficiently provide point cloud media / content / data. For example, point cloud data can be acquired by one or more cameras through processes such as capture, synthesis, or generation of the point cloud. Through this acquisition process, a point cloud video can be acquired that includes the 3D position of each point (indicated by x, y, z position values, etc., hereinafter referred to as geometry) and the characteristics of each point (color, reflectivity, transparency, etc.). Also, the acquired point cloud video can be generated in, for example, a PLY (Polygon File format or the Stanford Triangle format) file that includes this. In the case of point cloud data having a plurality of frames, one or more files can be acquired. Point cloud related metadata (for example, metadata related to capture, etc.) can be generated in this process.
[0347] The captured point cloud video may require post - processing to improve the quality of the content. In the process of video capture, the maximum / minimum depth values may be adjusted within the range provided by the camera equipment, but even after adjustment, point data in undesired areas may be included. Therefore, post - processing such as removing undesired areas (for example, the background) or recognizing connected spaces and filling holes (spatial holes) may be performed. Also, point clouds extracted from cameras sharing a spatial coordinate system may be integrated into one content by a conversion process to a global coordinate system for each point based on the position coordinates of each camera obtained by calibration. Thereby, a point cloud video with a high point density can be acquired.
[0348] The point cloud pre-processing unit 20001 can generate a point cloud video into one or more pictures / frames. Here, a picture / frame generally means a unit indicating one video in a specific time period. Also, when the point cloud pre-processing unit 20001 divides the points constituting the point cloud video into one or more patches and maps them onto a 2D plane, it can generate an occupancy map picture / frame which is a binary map that indicates whether data exists at that position on the 2D plane with a value of 0 or 1. Here, a patch is a set of points constituting the point cloud, and the points belonging to the same patch are adjacent to each other in 3D space and are a set of points mapped in the same direction among the planes of the six-sided bounding box in the mapping process to a 2D image. Also, the point cloud pre-processing unit 20001 can generate a geometry picture / frame which is a picture / frame in the form of a depth map representing the position information (geometry) of each point forming the point cloud video in patch units. Also, the point cloud pre-processing unit 20001 can generate a texture picture / frame which is a picture / frame representing the color information of each point forming the point cloud video in patch units. In this process, metadata necessary for reconstructing the point cloud from individual patches can be generated, and this metadata includes information about the patches (referred to as additional information or additional patch information) such as the position and size of each patch in 2D / 3D space. Such pictures / frames are continuously generated in chronological order and can constitute a video stream or a metadata stream.
[0349] The point cloud video encoder 20002 can encode one or more video streams related to point cloud video. One video includes a plurality of frames, and one frame corresponds to a still video / picture. In this specification, point cloud video includes point cloud video / frames / pictures, and point cloud video may be used interchangeably with point cloud video / frames / pictures. The point cloud video encoder 20002 performs a video-based point cloud compression (V-PCC) procedure. The point cloud video encoder 20002 can perform a series of procedures such as prediction, transformation, quantization, and entropy coding for the sake of compression and coding efficiency. The encoded data (encoded video / video information) is output in the form of a bitstream. When based on the V-PCC procedure, the point cloud video encoder 20002 can encode the point cloud video into geometric video, attribute video, occupancy map video, and also metadata, for example, information about patches, as described later. The geometric video may include a geometric image, the attribute video may include an attribute image, and the occupancy map video may include an occupancy map image. The patch data, which is additional information, may include information about the patch. The attribute video / image may include a texture video / image.
[0350] The point cloud image encoder 20003 can encode one or more images related to a point cloud video. The point cloud image encoder 20003 performs a video-based point cloud compression (V-PCC) procedure. The point cloud image encoder 20003 can perform a series of procedures such as prediction, transformation, quantization, and entropy coding for the sake of compression and coding efficiency. The encoded image is output in bitstream format. When based on the V-PCC procedure, the point cloud image encoder 20003 can encode the point cloud image by dividing it into a geometry image, an attribute image, an occupancy map image, and also metadata, for example, information about patches, as described below.
[0351] According to an embodiment, the point cloud video encoder 20002, the point cloud image encoder 20003, the point cloud video decoder 20006, and the point cloud image decoder 20008 may be performed by one encoder / decoder as described above, or may be performed by another path as shown in the drawings.
[0352] The encapsulation unit (file / segment encapsulation unit) 20004 can encapsulate the encoded point cloud data and / or the metadata related to the point cloud in a form such as a file or a segment for streaming. Here, the metadata related to the point cloud may be transmitted from a metadata processing unit (not shown). The metadata processing unit may be included in the point cloud video / image encoders 20002, 20003, or may be composed of other components / modules. The encapsulation unit 20004 encapsulates one bitstream or individual bitstreams including the video / image / metadata in a file format such as ISOBMFF, or processes them in a form such as a DASH segment. According to an embodiment, the encapsulation unit 20004 can include the metadata related to the point cloud on the file format. The point cloud metadata can be included, for example, in boxes at various levels on the ISOBMFF file format, or in the data in another track within the file. According to an embodiment, the encapsulation unit 20004 can encapsulate the point cloud related metadata itself in the file.
[0353] The encapsulation unit 20004 according to the embodiment divides and stores one bitstream or individual bitstreams in one or more tracks in the file, and also encapsulates the signaling information therefor. Further, the patch (or atlas) stream included on the bitstream may be stored in a track in the file, and the related signaling information may be stored. Furthermore, the SEI message existing on the bitstream may be stored in a track in the file, and the related signaling information may be stored.
[0354] The transmission processing unit (not shown) may perform processing for transmission on the point cloud data encapsulated according to the file format. The transmission processing unit may be included in the transmission unit (not shown) or may be composed of other components / modules. The transmission processing unit can process the point cloud data according to any transmission protocol. The processing for transmission may include processing for transmission via a broadcast network and processing for transmission via broadband. According to an embodiment, the transmission processing unit may transmit not only the point cloud data but also the point cloud-related metadata from the metadata processing unit and perform processing for transmission on this.
[0355] The transmission unit can transmit a point cloud bitstream or a file / segment including the bitstream to the receiving unit (not shown) of the receiving device via a digital storage medium or a network. For transmission, processing according to any transmission protocol may be performed. The data processed for transmission is transmitted via a broadcast network and / or broadband. This data is transmitted to the receiving side in an On Demand manner. The digital storage medium includes various ones such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmission unit may include elements for generating a media file in a predetermined file format and may include elements for transmission via a broadcast / communication network. The receiving unit extracts the bitstream and transmits it to the decoding device.
[0356] The receiving unit can receive the point cloud data transmitted by the point cloud data transmission device according to this specification. Depending on the channel to be transmitted, the receiving unit may receive the point cloud data via a broadcast network or may receive the point cloud data via broadband. Alternatively, the point cloud video data may be received by a digital storage medium. The receiving unit may decode the received data and render it according to, for example, the user's viewport.
[0357] The receiving processing unit (not shown) can perform processing according to the transmission protocol on the received point cloud video data. The receiving processing unit may be included in the receiving unit or may be composed of another component / module. In response to the processing for transmission being performed on the transmission side, the receiving processing unit performs the reverse process of the above-described transmission processing unit. The receiving processing unit transmits the acquired point cloud video to the decapsulation unit 20005 and transmits the metadata related to the acquired point cloud to a metadata processing unit (not shown).
[0358] The decapsulation unit (file / segment decapsulation unit) 20005 can decapsulate the file format point cloud data transmitted from the receiving processing unit. The decapsulation unit 20005 can decapsulate a file such as ISOBMFF and acquire a point cloud bitstream or point cloud related metadata (or another metadata bitstream). The acquired point cloud bitstream is transmitted to the point cloud video decoder 20006 and the point cloud image decoder 20008, and the acquired point cloud related metadata (or metadata bitstream) is transmitted to a metadata processing unit (not shown). The point cloud bitstream may include metadata (metadata bitstream). The metadata processing unit may be included in the point cloud video decoder 20006 or may be composed of another component / module. The point cloud related metadata acquired by the decapsulation unit 20005 may be in the form of a box or track in the file format. The decapsulation unit 20005 may receive metadata necessary for decapsulation from the metadata processing unit when necessary. The point cloud related metadata may be transmitted to the point cloud video decoder 20006 and / or the point cloud image decoder 20008 and used for point cloud decoding, or may be transmitted to the renderer 20009 and used for point cloud rendering.
[0359] The point cloud video decoder 20006 can decode video / images by receiving a bitstream and performing an inverse process corresponding to the operation of the point cloud video encoder 20002. In this case, as will be described later, the point cloud video decoder 20006 can decode the point cloud video into geometric video, attribute video, occupancy map video, and also auxiliary patch information. The geometric video may include a geometric image, the attribute video may include an attribute image, and the occupancy map video may include an occupancy map image. The additional information may include auxiliary patch information. The attribute video / image may include texture video / image.
[0360] The point cloud image decoder 20008 receives a bitstream and performs an inverse process corresponding to the operation of the point cloud image encoder 20003. In this case, the point cloud image decoder 20008 can decode the point cloud image into a geometric image, an attribute image, an occupancy map image, and also metadata, for example, auxiliary patch information.
[0361] The decoded geometry video / image, occupancy map, and additional patch information are used to restore the 3D geometry, and then smoothing processing is performed. By assigning color values to the smoothed 3D geometry using the texture video / image, a color point cloud video / picture is restored. The renderer 20009 can render the restored geometry and color point cloud video / picture. The rendered video / image is displayed on the display unit. The user can view all or part of the area of the rendered result through a VR / AR display or a general display, etc.
[0362] The sensing / tracking unit 20007 acquires orientation information and / or user viewport information from the user or the receiving side and transmits it to the receiving unit and / or the transmitting unit. The orientation information can indicate information regarding the position, angle, movement, etc. of the user's head, or information regarding the position, angle, movement, etc. of the device the user is looking at. Based on this information, information regarding the area that the user is currently viewing in the 3D space, i.e., viewport information, can be calculated.
[0363] The viewport information may be information regarding the area that the user is currently viewing in the 3D space through a device such as a device or HMD. A device such as a display can extract the viewport area based on the orientation information, the vertical or horizontal FOV supported by the device, etc. The orientation or viewport information can be extracted or calculated on the receiving side. The orientation or viewport information analyzed on the receiving side may be transmitted to the transmitting side via a feedback channel.
[0364] The receiving unit can efficiently extract or decrypt only the media data of a specific area, that is, the area indicated by the orientation information and / or the viewport information, using the orientation information acquired by the sensing / tracking unit 20007 and / or the viewport information indicating the area currently being viewed by the user from the file. Also, the transmitting unit can efficiently encode only the media data of a specific area, that is, the area indicated by the orientation information and / or the viewport information, using the orientation information acquired by the sensing / tracking unit 20007 and / or the viewport information, or generate and transmit a file.
[0365] The renderer 20009 can render the point cloud data decoded on the 3D space. The rendered video / image is displayed via the display unit. The user can view all or part of the area of the rendered result via a VR / AR display or a general display.
[0366] The feedback process may include transmitting various feedback information that can be obtained from the rendering / display process to the transmitting side or to the decoder on the receiving side. Through the feedback process, interactivity can be provided in the consumption of the point cloud data. According to an embodiment, in the feedback process, head orientation information, viewport information indicating the area currently being viewed by the user, etc. can be transmitted. According to an embodiment, the user can interact with what is embodied in a VR / AR / MR / autonomous driving environment, and in this case, information regarding the interaction can also be transmitted to the transmitting side and the service provider side in the feedback process. According to an embodiment, the feedback process may be omitted.
[0367] According to the embodiment, the above-described feedback information can not only be transmitted to the transmission side, but also be consumed on the reception side. That is, the above-described feedback information may be used to perform decapsulation processing, decoding, rendering processes, etc. on the reception side. For example, using the orientation information and / or viewport information, the point cloud data for the area currently viewed by the user may be preferentially decapsulated, decoded, and rendered.
[0368] FIG. 21 shows an example of the configuration of a point cloud data storage and transmission device according to an embodiment.
[0369] FIG. 21 shows a point cloud system according to an embodiment, and part / all of the system may include part / all of the transmission and reception device of FIG. 1, the encoding process of FIG. 4, the 2D video / image encoder of FIG. 15, the decoding process of FIG. 16, the transmission device of FIG. 18, and / or the reception device of FIG. 19. Also, it can be included in part / all of the system of FIG. 20 or can correspond thereto.
[0370] The point cloud data transmission device according to the embodiment is configured as shown in the drawings. Each component of the transmission device may be a module / unit / component / hardware / software / processor, etc.
[0371] The geometry, characteristics, additional data (or additional information), mesh data, etc. of the point cloud may be composed of independent streams respectively, or may be stored in different tracks in a file. Further, it may be included in another segment.
[0372] The Point Cloud Acquisition unit 21000 acquires a point cloud. For example, point cloud data can be acquired through processes such as capturing, synthesizing, or generating a point cloud via one or more cameras. Through such an acquisition process, point cloud data including the 3D position of each point (represented by x, y, z position values, etc., hereinafter referred to as geometry) and the characteristics of each point (such as color, reflectivity, transparency, etc.) can be acquired, and this can be generated, for example, in a PLY (Polygon File format or the Stanford Triangle format) file. In the case of point cloud data having a plurality of frames, one or more files can be acquired. In this process, point cloud-related metadata (such as metadata related to capture, etc.) can be generated. The Patch Generation unit 21001 generates patches from the point cloud data. The Patch Generation unit 21001 generates the point cloud data or point cloud video into one or more pictures / frames. Generally, a picture / frame may mean a unit indicating one video in a specific time period. When dividing the points constituting the point cloud video into one or more patches (a set of points constituting the point cloud, and the points belonging to the same patch are adjacent to each other in 3D space, and a set of points mapped in the same direction among the six-sided bounding box planes in the mapping process to a 2D image) and mapping them to a 2D plane, an occupancy map picture / frame, which is a binary map that indicates whether data exists at that position on the 2D plane with a value of 0 or 1, can be generated. Also, a geometry picture / frame, which is a depth map-formatted picture / frame representing the position information (geometry) of each point constituting the point cloud video in patch units, can be generated.A texture picture / frame can be generated that represents the color information of each point forming a point cloud video in terms of patches. In this process, metadata necessary for reconstructing the point cloud from individual patches can be generated, and this metadata may include information about the patches, such as the position and size of each patch in 2D / 3D space. Such pictures / frames are generated continuously in chronological order and can constitute a video stream or a metadata stream.
[0373] Also, the patches may be used for 2D image mapping. For example, the point cloud data may be projected onto each face of a cube. After patch generation, based on the generated patches, a geometry image, one or more attribute images, an occupancy map, auxiliary data, and / or mesh data can be generated.
[0374] Geometry Image Generation, Attribute Image Generation, Occupancy Map Generation, Auxiliary Data Generation, and / or Mesh Data Generation are performed by the point cloud preprocessing unit 20001 or a controller (not shown). In one embodiment, the point cloud preprocessing unit 20001 includes a patch generation unit 21001, a geometry image generation unit 21002, an attribute image generation unit 21003, an occupancy map generation unit 21004, an auxiliary data generation unit 21005, and a mesh data generation unit 21006.
[0375] The Geometry Image Generation unit 21002 generates a geometry image based on the result of patch generation. Geometry indicates points in 3D space. The geometry image is generated using an occupancy map, additional data (or additional information, including patch data), and / or mesh data that includes information related to 2D image packing of the patch based on the patch. The geometry image is related to information such as the depth (e.g., nearness, farness) of the patch generated after patch generation.
[0376] The Attribute Image Generation unit 21003 generates an attribute image. For example, the attribute can indicate a texture. The texture may be a color value corresponding to each point. According to an embodiment, a plurality (N) of attribute (such as color, reflectivity) images including a texture can be generated. The plurality of attributes can include a material (information related to the material), reflectivity, and the like. Also, according to an embodiment, the attribute may further include information in which the color changes depending on vision and light even with the same texture.
[0377] The Occupancy Map Generation unit 21004 generates an occupancy map from the patch. The occupancy map includes information indicating the presence or absence of data in pixels such as the geometry or attribute image.
[0378] The Auxiliary Data Generation unit 21005 generates auxiliary data (or auxiliary patch information) including information about the patch. That is, the auxiliary data indicates metadata regarding the patch of the point cloud object. For example, it can indicate information such as the normal vector for the patch. Specifically, according to an embodiment, the auxiliary data includes information necessary to reconstruct the point cloud from the patch (e.g., information regarding the position, size, etc. of the patch in 2D / 3D space, projection plane (normal) identification information, patch mapping information, etc.).
[0379] The Mesh Data Generation unit 21006 generates mesh data from the patch. The mesh indicates connection information between adjacent points. For example, it may indicate triangle data. For example, the mesh data according to an embodiment means the connection information between each point.
[0380] The point cloud preprocessing unit 20001 or the control unit generates metadata related to patch generation, geometry image generation, texture image generation, occupancy map generation, auxiliary data generation, and mesh data generation.
[0381] The point cloud transmission device performs video encoding and / or image encoding corresponding to the result generated by the point cloud preprocessing unit 20001. The point cloud transmission device can generate not only point cloud video data but also point cloud image data. According to an embodiment, the point cloud data may include only video data, only image data, and / or both video data and image data.
[0382] The video encoding unit 21007 performs geometric video compression, texture video compression, occupancy map video compression, additional data compression, and / or mesh data compression. The video encoding unit 21007 generates a video stream including each encoded video data.
[0383] Specifically, the geometric video compression encodes point cloud geometric video data. The texture video compression encodes point cloud texture video data. The additional data compression encodes additional data related to the point cloud video data. The mesh data compression encodes the mesh data of the point cloud video data. Each operation of the point cloud video encoding unit is performed in parallel.
[0384] The image encoding unit 21008 performs geometric image compression, texture image compression, occupancy map image compression, additional data compression, and / or mesh data compression. The image encoding unit generates an image including each encoded image data.
[0385] Specifically, the geometric image compression encodes point cloud geometric image data. The texture image compression encodes point cloud texture image data. The additional data compression encodes additional data related to the point cloud image data. The mesh data compression encodes the mesh data related to the point cloud image data. Each operation of the point cloud image encoding unit is performed in parallel.
[0386] The video encoding unit 21007 and / or the image encoding unit 21008 can receive metadata from the point cloud preprocessing unit 20001. The video encoding unit 21007 and / or the image encoding unit 21008 can perform each encoding process based on the metadata.
[0387] The File / Segment Encapsulation 21009 encapsulates video streams and / or images into file and / or segment formats. The File / Segment Encapsulation 21009 performs video track encapsulation, metadata track encapsulation and / or image encapsulation.
[0388] Video track encapsulation can encapsulate one or more video streams into one or more tracks.
[0389] Metadata track encapsulation can encapsulate metadata related to video streams and / or images into one or more tracks. The metadata includes data related to the content of the point cloud data. For example, it includes Initial Viewing Orientation Metadata. According to an embodiment, the metadata may be encapsulated into a metadata track, or may be encapsulated together with a video track or an image track.
[0390] Image encapsulation can encapsulate one or more images into one or more tracks or items.
[0391] For example, according to an embodiment, when four video streams and two images are input to the encapsulation unit, the four video streams and the two images are encapsulated into one file.
[0392] The File / Segment Encapsulation 21009 can receive metadata from the Point Cloud Preprocessing Unit 20001. The File / Segment Encapsulation 21009 can perform encapsulation based on the metadata.
[0393] The files and / or segments generated by file / segment encapsulation are transmitted by a point cloud transmission device or a transmission unit. For example, segments can be delivered based on a DASH-based protocol.
[0394] The Delivery unit can transmit a point cloud bitstream or a file / segment containing the bitstream to the receiving unit of the receiving device via a digital storage medium or a network. For transmission, processing is performed according to any transmission protocol. The data after the processing for transmission can be transmitted via a broadcast network and / or broadband. This data may be transmitted to the receiving side in an On Demand manner. The digital storage medium includes various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD.
[0395] The encapsulation unit 21009 according to the embodiment can divide and store one bitstream or individual bitstreams into one or more tracks in a file, and can also encapsulate the signaling information therefor together. In addition, the patch (or atlas) stream included on the bitstream can be stored in a track in the file, and the related signaling information can be stored. Further, the SEI message existing on the bitstream can be stored in a track in the file, and the related signaling information can be stored.
[0396] The transmission unit can include elements for generating media files in a predetermined file format and can include elements for transmission via a broadcast / communication network. The transmission unit receives orientation information and / or viewport information from the reception unit. The transmission unit can transmit the acquired orientation information and / or viewport information (or the information selected by the user) to the point cloud preprocessing unit 20001, the video encoding unit 21007, the image encoding unit 21008, the file / segment encapsulation unit 21009, and / or the point cloud encoding unit. Based on the orientation information and / or viewport information, the point cloud encoding unit can encode all the point cloud data or the point cloud data indicated by the orientation information and / or viewport information. Based on the orientation information and / or viewport information, the file / segment encapsulation unit can encapsulate all the point cloud data or the point cloud data indicated by the orientation information and / or viewport information. Based on the orientation information and / or viewport information, the transmission unit can transmit all the point cloud data or the point cloud data indicated by the orientation information and / or viewport information.
[0397] For example, the point cloud preprocessing unit 20001 may perform the above-described operations on all point cloud data, or may perform operations on the point cloud data indicated by the orientation information and / or the viewport information. The video encoding unit 21007 and / or the image encoding unit 21008 may perform the above-described operations on all point cloud data, or may perform the above-described operations on the point cloud data indicated by the orientation information and / or the viewport information. The file / segment encapsulation unit 21009 may perform the above-described operations on all point cloud data, or may perform the above-described operations on the point cloud data indicated by the orientation information and / or the viewport information. The transmission unit may perform the above-described operations on all point cloud data, or may perform the above-described operations on the point cloud data indicated by the orientation information and / or the viewport information.
[0398] FIG. 22 shows an example of the configuration of a point cloud data receiving apparatus according to an embodiment.
[0399] FIG. 22 shows a point cloud system according to an embodiment, and part or all of the system may include part or all of the transmission / reception apparatus of FIG. 1, the encoding process of FIG. 4, the 2D video / image encoder of FIG. 15, the decoding process of FIG. 16, the transmission apparatus of FIG. 18, and / or the reception apparatus of FIG. 19. Further, it may be included in or correspond to part or all of the systems of FIGS. 20 and 21.
[0400] Each component of the receiving device may be a module / unit / component / hardware / software / processor or the like. The Delivery Client 22006 can receive the point cloud data, point cloud bitstream, or a file / segment including the bitstream transmitted by the point cloud data transmission device according to the embodiment. Depending on the channel to be transmitted, the receiving device may receive the point cloud data via a broadcast network or via broadband. Alternatively, the point cloud data may be received by a digital storage medium. The receiving device may include a process of decrypting the received data and rendering it according to the user's viewport or the like. The Delivery Client 22006 (or the receiving processing unit) can perform processing according to the transmission protocol on the received point cloud data. The receiving processing unit may be included in the receiving unit or may be composed of another component / module. Corresponding to the processing for transmission performed on the transmission side, the receiving processing unit performs the reverse process of the above-described transmission processing unit. The receiving processing unit can transmit the acquired point cloud data to the File / Segment Decapsulation Unit 22000, and the acquired point cloud-related metadata can be transmitted to a metadata processing unit (not shown).
[0401] The Sensing / Tracking Unit 22005 acquires orientation information and / or viewport information. The Sensing / Tracking Unit 22005 can transmit the acquired orientation information and / or viewport information to the Delivery Client 22006, the File / Segment Decapsulation Unit 22000, the Point Cloud Decryption Units 22001, 22002, and the Point Cloud Processing Unit 22003.
[0402] The transmission client 22006 may receive all point cloud data based on the orientation information and / or viewport information, or may receive the point cloud data indicated by the orientation information and / or viewport information. The file / segment decapsulation unit 22000 can decapsulate all point cloud data or the point cloud data indicated by the orientation information and / or viewport information based on the orientation information and / or viewport information. The point cloud decoding unit (video decoding unit 22001 and / or image decoding unit 22002) can decode all point cloud data or the point cloud data indicated by the orientation information and / or viewport information based on the orientation information and / or viewport information. The point cloud processing unit 22003 can process all point cloud data or the point cloud data indicated by the orientation information and / or viewport information.
[0403] The File / Segment decapsulation unit 22000 performs Video Track Decapsulation, Metadata Track Decapsulation, and / or Image Decapsulation. The File / Segment decapsulation unit 22000 can decapsulate the point cloud data in file format transmitted by the reception processing unit. The File / Segment decapsulation unit 22000 can decapsulate a file or segment such as ISOBMFF to obtain a point cloud bitstream and point cloud related metadata (or another metadata bitstream). The obtained point cloud bitstream is transmitted to the point cloud decoding units 22001 and 22002, and the obtained point cloud related metadata (or metadata bitstream) can be transmitted to a metadata processing unit (not shown). The point cloud bitstream may include metadata (metadata bitstream). The metadata processing unit may be included in the point cloud video decoder or may be composed of another component / module. The point cloud related metadata obtained by the File / Segment decapsulation unit 22000 may be in the form of boxes or tracks within the file format. The File / Segment decapsulation unit 22000 may receive metadata necessary for decapsulation from the metadata processing unit if necessary. The point cloud related metadata may be transmitted to the point cloud decoding units 22001 and 22002 and used for point cloud decoding, or may be transmitted to the point cloud rendering unit 22004 and used for point cloud rendering. The File / Segment decapsulation unit 22000 can generate metadata related to the point cloud data.
[0404] Video Track Decapsulation within the File / Segment Decapsulation Unit 22000 decapsulates video tracks included in a file and / or segment. It decapsulates a video stream including geometric video, texture video, occupancy maps, additional data, and / or mesh data.
[0405] Metadata Track Decapsulation within the File / Segment Decapsulation Unit 22000 decapsulates a bitstream including metadata related to point cloud data and / or additional data, etc.
[0406] Image Decapsulation within the File / Segment Decapsulation Unit 22000 decapsulates an image including geometric images, texture images, occupancy maps, additional data, and / or mesh data.
[0407] The Decapsulation Unit 22000 according to the embodiment splits and parses (decapsulates) one bitstream or individual bitstreams based on one or more tracks in a file, and also decapsulates the signaling information therefor. Also, a patch (or atlas) stream included on the bitstream can be decapsulated based on tracks in the file, and the related signaling information can be parsed. Further, SEI messages present on the bitstream can be decapsulated based on tracks in the file, and the related signaling information can be obtained together.
[0408] The video decoding unit 22001 performs geometric video restoration, texture video restoration, occupancy map restoration, additional data restoration, and / or mesh data restoration. The video decoding unit decodes geometric video, texture video, additional data, and / or mesh data corresponding to the process of performing video encoding addition of the point cloud transmission device according to the embodiment.
[0409] The image decoding unit 22002 performs geometric image restoration, texture image restoration, occupancy map restoration, additional data restoration, and / or mesh data restoration. The image decoding unit decodes geometric images, texture images, additional data, and / or mesh data corresponding to the process performed by the image encoding unit of the point cloud transmission device according to the embodiment.
[0410] As described above, the video decoding unit 22001 and the image decoding unit 22002 according to the embodiment may be processed by one video / image decoder, or may be performed in separate paths as shown in the figure.
[0411] The video decoding unit 22001 and / or the image decoding unit 22002 can generate metadata related to video data and / or image data.
[0412] The point cloud processing unit 22003 performs geometry reconstruction and / or attribute reconstruction.
[0413] Geometry reconstruction restores geometric video and / or geometric images from the decoded video data and / or decoded image data based on the occupancy map, additional data, and / or mesh data.
[0414] Feature reconstruction restores the feature video and / or the feature image from the decoded feature video and / or the decoded feature image based on the occupancy map, additional data, and / or mesh data. According to an embodiment, for example, the feature may be a texture. According to an embodiment, the feature may mean a plurality of feature information. If there are a plurality of features, the point cloud processing unit 22003 according to the embodiment performs a plurality of feature reconstructions.
[0415] The point cloud processing unit 22003 can receive metadata from the video decoding unit 22001, the image decoding unit 22002, and / or the file / segment decapsulation unit 22000 and process the point cloud based on the metadata.
[0416] The Point Cloud Rendering unit 22004 renders the reconstructed point cloud. The Point Cloud Rendering unit 22004 can receive metadata from the video decoding unit 22001, the image decoding unit 22002, and / or the file / segment decapsulation unit 22000 and render the point cloud based on the metadata.
[0417] The display displays the rendered result on an actual display device.
[0418] According to the method / apparatus of the embodiment, as shown in FIGS. 20 to 22, on the transmitting side, point cloud data is encoded into a bit stream, encapsulated and transmitted in the form of a file and / or segment, and on the receiving side, the form of the file and / or segment is decapsulated into a bit stream including a point cloud, and can be decoded into point cloud data. For example, the point cloud data transmission apparatus according to the embodiment encapsulates point cloud data based on a file. At this time, the file may include a V-PCC track including parameters related to the point cloud, a geometry track including geometry, a property track including properties, and an occupancy track including an occupancy map.
[0419] Also, the point cloud data receiving apparatus according to the embodiment decapsulates point cloud data based on a file. At this time, the file may include a V-PCC track including parameters related to the point cloud, a geometry track including geometry, a property track including properties, and an occupancy track including an occupancy map.
[0420] The above-described encapsulation operation may be performed by the file / segment encapsulation unit 20004 in FIG. 20, the file / segment encapsulation unit 21009 in FIG. 21, etc., and the above-described decapsulation operation may be performed by the file / segment decapsulation unit 20005 in FIG. 20, the file / segment decapsulation unit 22000 in FIG. 22, etc.
[0421] FIG. 23 shows an example of a structure that can be interlocked with the method / apparatus for transmitting and receiving point cloud data according to the embodiment.
[0422] In the structure according to the embodiment, at least one or more of the AI (Artificial Intelligence) server 23600, robot 23100, autonomous vehicle 23200, XR device 23300, smartphone 23400, home appliance 23500, and / or HMD 23700 are connected to the cloud network 23000. Here, devices such as the robot 23100, autonomous vehicle 23200, XR device 23300, smartphone 23400, or home appliance 23500 can be referred to as devices. Also, the XR device 23300 may correspond to the point cloud compression data (PCC) device according to the embodiment or may operate in conjunction with the PCC device.
[0423] The cloud network 23000 may mean a network that constitutes part of the cloud computing infrastructure or exists within the cloud computing infrastructure. Here, the cloud network 23000 may be configured using a 3G network, 4G or LTE (Long Term Evolution) network, or 5G network, etc.
[0424] The AI server 23600 is connected to at least one or more of the robot 23100, autonomous vehicle 23200, XR device 23300, smartphone 23400, home appliance 23500, and / or HMD 23700 via the cloud network 23000 and can assist with at least a part of the processing of the connected devices 23100 to 23700.
[0425] The HMD (Head-Mount Display) 23700 represents one type that the XR device 23300 and / or PCC device according to the embodiment can implement. The device of the HMD type according to the embodiment includes a communication unit, a control unit, a memory unit, an I / O unit, a sensor unit, and a power supply unit, etc.
[0426] Hereinafter, various embodiments of apparatuses 23100 to 23500 to which the above technology is applied will be described. Here, the apparatuses 23100 to 23500 shown in FIG. 23 can be interlocked / connected with the point cloud data transmission / reception apparatus according to the above-described embodiment.
[0427] <PCC+XR>
[0428] The XR / PCC apparatus 23300 is applied with PCC and / or XR (AR+VR) technology and can be embodied in an HMD (Head-Mount Display), a HUD (Head-Up Display) provided in a vehicle, a TV, a mobile phone, a smartphone, a computer, a wearable device, a household electrical appliance, a digital signboard, a vehicle, a fixed robot, a mobile robot, and the like.
[0429] The XR / PCC apparatus 23300 can obtain information on the surrounding space or real object by analyzing 3D point cloud data or image data acquired from various sensors or an external device to generate position data and characteristic data for 3D points, and can render and output an XR object to be output. For example, the XR / PCC apparatus 23300 can output an XR object including additional information regarding the recognized object corresponding to the recognized object.
[0430] <PCC+Autonomous Driving+XR>
[0431] The autonomous driving vehicle 23200 is applied with PCC technology and XR technology and is embodied in a mobile robot, a vehicle, an unmanned aerial vehicle, and the like.
[0432] The autonomous driving vehicle 23200 to which XR / PCC technology is applied may mean an autonomous driving vehicle provided with means for providing an XR video or an autonomous driving vehicle that is a control / interaction target within the XR video. In particular, the autonomous driving vehicle 23200 that is a control / interaction target within the XR video is distinguished from the XR apparatus 23300 and can be interlocked with each other.
[0433] The autonomous vehicle 23200 equipped with means for providing an XR / PCC image acquires sensor information from sensors including a camera and outputs an XR / PCC image generated based on the acquired sensor information. For example, the autonomous vehicle 23200 can provide an XR / PCC object corresponding to a real object or an object within a screen to a passenger by outputting the XR / PCC image with a HUD.
[0434] At this time, when the XR / PCC object is output to the HUD, at least a part of the XR / PCC object may be output so as to overlap with an actual object towards which the passenger's line of sight is directed. On the other hand, when the XR / PCC object is output to a display provided within the autonomous vehicle 23200, at least a part of the XR / PCC object may be output so as to overlap with an object within the screen. For example, the autonomous vehicle 23200 can output an XR / PCC object corresponding to an object such as a road, other vehicles, traffic lights, traffic signs, motorcycles, pedestrians, buildings, etc.
[0435] The VR (Virtual Reality) technology, AR (Augmented Reality) technology, MR (Mixed Reality) technology, and / or PCC (Point Cloud Compression) technology according to the embodiment are applicable to various devices.
[0436] That is, the VR technology is a display technology that provides only CG images of real objects, backgrounds, etc. On the other hand, the AR technology is a technology that shows both a virtual CG image on the image of an actual thing. Also, the MR technology is similar to the above-described AR technology in that it shows a virtual object mixed in the real world. However, in the AR technology, the distinction between the real object and the virtual object composed of the CG image is clear, and while the virtual object is used in a form that complements the real object, the MR technology is distinguished from the AR technology in that the virtual object and the real object are regarded as having the same nature. More specifically, for example, the hologram service is an example where the above MR technology is applied.
[0437] However, recently, rather than clearly distinguishing VR, AR, and MR technologies, they are called XR (extended Reality) technologies. Therefore, the embodiments of the present invention are applicable to any of VR, AR, MR, and XR technologies. Such technologies are applied with coding / decoding based on PCC, V-PCC, and G-PCC technologies.
[0438] The PCC method / apparatus according to the embodiment can be applied to the vehicle 23200 that provides an autonomous driving service.
[0439] The autonomous driving vehicle 23200 that provides an autonomous driving service is connected to the PCC apparatus in a wired / wireless communication-capable manner.
[0440] When the point cloud compression data (PCC) transceiver according to the embodiment is connected to the autonomous driving vehicle 23200 in a wired / wireless communication-capable manner, it can receive / process AR / VR / PCC service-related content data provided together with the autonomous driving service and transmit it to the autonomous driving vehicle 23200. Further, when the point cloud data transceiver is mounted on the autonomous driving vehicle 23200, the point cloud transceiver can receive / process AR / VR / PCC service-related content data according to a user input signal input by the user interface device and provide it to the user. The vehicle or user interface device according to the embodiment can receive a user input signal. The user input signal according to the embodiment may include a signal for instructing an autonomous driving service.
[0441] As described above, the V-PCC-based point cloud video encoder shown in FIGS. 1, 4, 18, 20, or 21 projects 3D point cloud data (or content) onto a 2D space to generate patches. The patches generated in the 2D space are generated by being divided into a geometry image indicating position information (referred to as a geometry frame or a geometry patch frame) and a texture image indicating color information (referred to as a texture frame or a texture patch frame). The geometry image and the texture image are each video-compressed for each frame and output as a video bitstream of the geometry image (or a geometry bitstream) and a video bitstream of the texture image (or a texture bitstream). In addition, additional patch information (or patch information or metadata) including each patch projection plane information and patch size information necessary for decoding the 2D patch on the receiving side is also video-compressed and output to a bitstream of the additional patch information. In addition, an occupancy map indicating the presence or absence of a point for each pixel as 0 or 1 is entropy-compressed or video-compressed according to whether it is in a lossless mode or a lossy mode and output to a video bitstream of the occupancy map (or an occupancy map bitstream). The compressed geometry bitstream, the compressed texture bitstream, the compressed additional patch information bitstream, and the compressed occupancy map bitstream are multiplexed as the structure of the V-PCC bitstream.
[0442] According to an embodiment, the V-PCC bitstream may be directly transmitted to the receiving side, or may be encapsulated in the form of a file / segment in the file / segment encapsulation unit of FIGS. 1, 18, 20, or 21 and transmitted to the receiving device, or may be stored in a digital storage medium (for example, USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.). In this specification, it is taken as an example that the file is in the ISOBMFF file format in one embodiment.
[0443] According to an embodiment, the V-PCC bitstream may be transmitted via multiple tracks of a file or via a single track. Details will be described later.
[0444] FIG. 24 shows an example of a V-PCC bitstream structure according to an embodiment. The V-PCC bitstream in FIG. 24 is output from the V-PCC base point cloud video encoder of FIG. 1, FIG. 4, FIG. 18, FIG. 20, or FIG. 21 in one embodiment.
[0445] The V-PCC bitstream consists of one or more V-PCC units. That is, the V-PCC bitstream is a set of V-PCC units. Each V-PCC unit consists of a V-PCC unit header and a V-PCC unit payload. This specification classifies the data included in the V-PCC unit payload by the V-PCC unit header. For this purpose, the V-PCC unit header includes type information indicating the type of the V-PCC unit. Each V-PCC unit payload can include at least one of geometry video data (i.e., compressed geometry bitstream), texture video data (i.e., compressed texture bitstream), occupancy video data (i.e., compressed occupancy map bitstream), patch data group (PDG), and sequence parameter set (SPS). The patch data group is also called an atlas. In this specification, the atlas can be replaced with the patch data group.
[0446] The V-PCC unit payload, which includes at least geometry video data, texture video data, and occupancy map video data (or occupancy video data), corresponds to a video data unit (e.g., an HEVC NAL unit) that is decoded by an appropriate video decoder.
[0447] Geometry video data, texture video data, and occupancy map video data are referred to as 2D video encoding information for the geometry, texture, and occupancy map components of the encoded point cloud, and the patch data group (or atlas) is referred to as non-video encoded information. The patch data group includes additional patch information. The sequence parameter set includes the overall encoding information of the bitstream and may also be referred to as configuration and metadata information. The above sequence parameter set and patch data group may be referred to as signaling information, may be generated by a metadata processing unit in the point cloud video encoder, or may be generated by another component / module in the point cloud video encoder. Also, in this specification, the sequence parameter set and patch data group are referred to as initialization information, and geometry video data, texture video data, and occupancy map video data are referred to as point cloud data.
[0448] As an example, when the type information of the V-PCC unit header indicates a patch data group (VPCC-PDG), the V-PCC unit payload can include the patch data group. Conversely, when the V-PCC unit payload includes a patch data group, the type information of the V-PCC unit header can identify it. Since this is a matter of the designer's choice, this specification is not limited to the above example. Also, the same or similar applies to other V-PCC units.
[0449] The patch data group is included in the V-PCC unit payload in the patch data group unit () format. The patch data group can include information of one or more patch tile groups, and at least one of a patch sequence parameter set, a patch frame parameter set, a patch frame geometry parameter set, a geometry patch parameter set, a patch frame characteristic parameter set, and a characteristic patch parameter set.
[0450] On the other hand, a patch frame (or a point cloud object) targeted for point cloud data may be divided into one or multiple tiles. Tiles according to an embodiment may indicate a certain area in 3D space or a certain area in 2D plane. Also, a tile may be a rectangular cuboid or a sub-bounding box or a part of a patch frame within one bounding box. In this specification, dividing a patch frame (or a point cloud object) into one or more tiles may be performed by the point cloud video encoder in FIG. 1, the patch generation unit in FIG. 18, the point cloud preprocessing unit in FIG. 20, or the patch generation unit in FIG. 21, or may be performed by another component / module.
[0451] According to an embodiment, a V-PCC bitstream having a structure as shown in FIG. 24 may be directly transmitted to the receiving side, or may be encapsulated in a file / segment format and transmitted to the receiving side.
[0452] In this specification, as an embodiment, the V-PCC bitstream is encapsulated in a file format and transmitted. For example, the V-PCC bitstream can be encapsulated in a file format based on ISOBMFF (ISO Base Media File Format).
[0453] Encapsulating the V-PCC bitstream into a file is taken as an example to be performed by the file / segment encapsulation unit 10003 in FIG. 1, the transmission unit 18008 in FIG. 18, the file / segment encapsulation unit 20004 in FIG. 20, or the file / segment encapsulation unit 21009 in FIG. 21. Decapsulating a file into a V-PCC bitstream is taken as an example to be performed by the file / segment decapsulation unit 10007 in FIG. 1, the receiving unit in FIG. 19, the file / segment decapsulation unit 20005 in FIG. 20, or the file / segment decapsulation unit 22000 in FIG. 22.
[0454] FIG. 25 is a diagram visualizing a multiple-track V-PCC file structure according to an embodiment. That is, an example of the layout of an ISOBMFF-based file including multiple tracks is shown.
[0455] The ISOBMFF-based file according to the embodiment may be called a container, a container file, a media file, a V-PCC file, etc. Specifically, the file may be composed of boxes and / or information called ftyp, meta, moov, mdat, etc.
[0456] The ftyp box (file type box) can provide information related to the file type or file compatibility for that file. On the receiving side, the file can be classified by referring to the ftyp box.
[0457] The meta box can include vpcg{0, 1, 2, 3} boxes (V-PCC group boxes, which will be described in detail below).
[0458] The mdat box, also called the media data box, can include a video-encoded geometry bitstream, a video-encoded texture bitstream, a video-encoded occupancy map bitstream, and a patch data group bitstream.
[0459] The moov box, also known as the movie box, can contain metadata for the media data of its file (e.g., geometry bitstream, texture bitstream, occupancy map bitstream, etc.). For example, it can contain information necessary for decoding and playing back the media data, and can contain information regarding the samples of the file. The moov box can serve as a container for all metadata. The moov box may be the top - level box among the metadata - related boxes. According to an embodiment, there is only one moov box in the file.
[0460] In an ISOBMFF container structure such as that of FIG. 25, in one embodiment, the V - PCC units included in the V - PCC bitstream are mapped to individual tracks in the file based on their types.
[0461] Based on the layout as in FIG. 25, an ISOBMFF container for the V - PCC bitstream can include the following.
[0462] 1) It can include V - PCC tracks. A V - PCC track includes samples that transmit the payload of a V - PCC unit (e.g., type information in the V - PCC unit header indicating the type of the V - PCC unit, which indicates a sequence parameter set or a patch data group) that includes sequence parameter set and non - video coding information. The V - PCC track can also provide a track reference to other tracks that include samples that transmit the payload of a V - PCC unit (e.g., type information in the V - PCC unit header indicates geometry video data, texture video data, or occupancy map video data).
[0463] 2) It can include one or more restricted video scheme tracks for geometric video data. Samples included in this track include NAL units for a video-coded elementary stream for geometric video data. In this case, the type information in the V-PCC unit header indicates geometric video data, and the above NAL unit corresponds to the payload that carries the geometric video data.
[0464] 3) It can include one or more restricted video scheme tracks for texture video data. Samples included in this track include NAL units for a video-coded elementary stream for texture video data. In this case, the type information in the V-PCC unit header indicates texture video data, and the above NAL unit corresponds to the payload that carries the texture video data.
[0465] 4) It can include one restricted video scheme track for occupancy map video data. Samples included in this track include NAL units for a video-coded elementary stream for occupancy map video data. In this case, the type information in the V-PCC unit header indicates occupancy map video data, and the above NAL unit corresponds to the payload that carries the occupancy map video data.
[0466] According to the embodiment, a track including a video-coded elementary stream for geometric video data, texture video data, and occupancy map video data will be referred to as a component track.
[0467] Synchronization between elementary streams included in such a component track is, in one embodiment, handled by an ISOBMFF track timing structure (stts, ctts, cslg, or equivalent mechanisms within movie fragments). Samples contributing to the same point cloud frame across different video encoded component tracks and V-PCC tracks can have the same composition time.
[0468] The V-PCC parameter set used for this sample can be the same as the composition time of that frame or can have a previous decoding time.
[0469] An ISOBMFF file transmitting V-PCC content can be distinguished by a V-PCC definition brand. Tracks of V-PCC content can be grouped within a VPCCGroupBox, which is a file-level EntityToGroupBox having a V-PCC specific grouping 4CC value ('vpcg'). The VPCCGroupBox can be provided as an entry point for accessing V-PCC content within a container and can contain initial metadata identifying the V-PCC content. The VPCCGroupBox, which is an EntityToGroupBox, can be included within a MetaBox or Moov box.
[0470] An entity group is a group of items that group tracks. Entities within an entity group can share a particular characteristic or have a particular relationship as indicated by the grouping type.
[0471] The entity group is indicated within the GroupsListBox. The GroupsListBox is included in at least one of the file-level MetaBox box, Movie-level MetaBox box, and track-level MetaBox box. The entity group specified in the GroupsListBox of the file-level MetaBox refers to track or file-level items. The entity group specified in the GroupsListBox of the movie-level MetaBox refers to movie-level items. The entity group specified within the GroupsListBox of the track-level MetaBox refers to the track-level items of that track.
[0472] The GroupsListBox contains EntityToGroupBoxes as follows, each identifying one entity group.
[0473] Groups List box
[0474] Box Type: 'grpl'
[0475] Container: MetaBox that is not contained in AdditionalMetadataContainerBox
[0476] Mandatory: No
[0477] Quantity: Zero or One
[0478] The GroupsListBox contains entity groups specified for that file. This box contains a set of full boxes. Each is called an EntityToGroupBox and has a 4-character code indicating the defined grouping type.
[0479] The GroupsListBox does not exist within the AdditionalMetadataContainerBox.
[0480] If the GroupsListBox exists within the file-level MetaBox, there may be no item ID value in the ItemInfoBox within the same file-level MetaBox as the track ID value in the TrackHeaderBox, as follows.
[0481] aligned(8) class GroupsListBox extends Box('grpl') {
[0482] }
[0483] Box Type: As specified below with the grouping_type value for the EntityToGroupBox
[0484] Container: GroupsListBox
[0485] Mandatory: No
[0486] Quantity: One or more
[0487] The EntityToGroupBox specifies the entity group.
[0488] The box type (grouping_type) indicates the grouping type of an entity group. Each grouping_type code is related to the semantics that describe the grouping. The following is an explanation of the grouping_type value:
[0489] 'altr': The items and tracks mapped to this grouping can replace each other, and only one of them can be played (if the mapped items and tracks are part of a presentation, for example, a displayable item or track) or processed by other means (if the mapped item or track is not part of a presentation, for example, metadata). The player can select and process the first entity from the list of entity ID values (entity_id). For example, decrypt and play for mapped items and tracks that are part of a presentation. It also conforms to the needs of the application. The entity ID value is mapped to only one grouping of type 'altr'. The alternate group of entities consists of items and tracks mapped to the same entity group of type 'altr'.
[0490] Note: EntityToGroupBox contains specific extensions for grouping_type.
[0491] aligned(8) class EntityToGroupBox(grouping_type, version, flags) extends FullBox(grouping_type, version, flags) {
[0492] unsigned int(32) group_id;
[0493] unsigned int(32) num_entities_in_group;
[0494] for(i = 0; i < num_entities_in_group; i++)
[0495] unsigned int(32) entity_id;
[0496] }
[0497] The group ID (group_id) according to the embodiment is a non - negative integer assigned to a specific grouping that is not the same as the group ID (group_id) of other EntityToGroupBoxes, the item ID (item_ID) value of the hierarchical level (file, movie, or track) including the GroupsListBox, or the track ID (track_ID) value (when the GroupsListBox is included at the file level).
[0498] The num_entities_in_group according to the embodiment indicates the number of entity ID (entity_id) values mapped to this entity group.
[0499] The entity_id according to the embodiment is resolved to an item when an item with item_ID equal to entity_id is present in the hierarchy level (file, movie or track) that contains the GroupsListBox, or to a track when a track with track_ID equal to entity_id is present and the GroupsListBox is contained in the file level.
[0500] The V-PCC group box will be described below.
[0501] Box Type: 'vpcg'
[0502] Container: GroupListBox
[0503] Mandatory: Yes
[0504] Quantity: One or more
[0505] This box is included in GroupsListBox and provides a list of the tracks that comprise a V-PCC content.
[0506] V-PCC content specific information, such as information that maps the attribute types and layers to the related tracks, is listed in this box. This information provides a convenient way to have an initial understanding of the V-PCC content. For the flexible configuration of V-PCC content to support various client functions, multiple versions of the encoded V-PCC components are listed in this box. The profile, layer, and level information defined in V-PCC is also transmitted in this box as follows.
[0507] aligned(8) class VPCCGroupBox() extends EntityToGroupBox('vpcg', version, flags) {
[0508] for (i = 0; i < num_entities_in_group; i++) {
[0509] unsigned int(4) data_type;
[0510] unsigned int(4) attribute_type;
[0511] unsigned int(1) multiple_layer_present_flag;
[0512] unsigned int(4) layer_count_munus1;
[0513] for (i = 0; i < layer_count_minus1 + 1; i++) {
[0514] unsigned int(4) layer_id;
[0515] }
[0516] unsigned int(32) entity_id;
[0517] }
[0518] unsigned int(4) CC_layer_count_minus1;
[0519] vpcc_profile_tier_level()
[0520] }
[0521] The data_type according to the embodiment indicates the track type of the PCC data in the referenced track.
[0522] FIG. 26 is a table showing an example of the track type of the PCC data assigned to the data_type according to the embodiment. For example, when the value of data_type is 1, it can indicate a V-PCC track, when it is 2, a geometry video track, when it is 3, a texture video track, and when it is 4, an occupancy map video track.
[0523] The multiple_layer_present_flag according to the embodiment indicates whether a single geometry or a texture layer or a multiple geometry or a texture layer is transmitted to a related entity (or track). For example, when the multiple_layer_present_flag is 0, it indicates that a single geometry or a texture layer is transmitted to a related entity (or track), and when it is 1, it can indicate that a multiple geometry or a texture layer is transmitted to a related entity (or track). The V-PCC track according to the embodiment (i.e., the data_type indicates a V-PCC track) has a multiple_layer_present_flag with a value of 0. As another example, when the data_type indicates a geometry video track, if the value of the multiple_layer_present_flag is 0, the geometry video track transmits a single geometry layer, and if the value of the multiple_layer_present_flag is 1, it can indicate that the geometry video track transmits multiple geometry layers. As another example, when the data_type indicates a texture video track, if the value of the multiple_layer_present_flag is 0, the texture video track transmits a single texture layer, and if the value of the multiple_layer_present_flag is 1, it can indicate that the texture video track transmits multiple texture layers.
[0524] Layer_count_minus1 plus 1 according to the embodiment indicates the number of geometry and / or texture layers transmitted to the associated entity (or track). The V-PCC track according to the embodiment has layer_count_minus1 whose value is 0. For example, when the above data_type indicates a geometry video track, layer_count_minus1 plus 1 can indicate the number of geometry layers transmitted to the geometry video track. As another example, when the above data_type indicates a texture video track, layer_count_minus1 plus 1 can indicate the number of texture layers transmitted to the texture video track.
[0525] Layer_id according to the embodiment indicates the layer identifier of the geometry and / or texture layer within the associated entity (or track). The V-PCC track according to the embodiment has layer_id whose value is 0. The set of layer_id values for the V-PCC component track type is sorted in ascending order and is a continuous set of integers starting from 0. For example, when the above data_type indicates a geometry video track, layer_id can indicate the layer identifier of the geometry layer. As another example, when the above data_type indicates a texture video track, layer_id can indicate the layer identifier of the texture layer.
[0526] The pcc_layer_count_munus1 plus 1 according to the embodiment indicates the number of layers used to encode the geometric components and / or the characteristic components of the point cloud stream. For example, when the above data_type indicates a geometric video track, pcc_layer_count_munus1 plus 1 can indicate the number of layers used to encode the geometric components of the point cloud stream. As another example, when the above data_type indicates a characteristic video track, pcc_layer_count_munus1 plus 1 can indicate the number of layers used to encode the characteristic components of the point cloud stream.
[0527] The attribute_type according to the embodiment indicates the type of characteristic of the characteristic video data transmitted within the referenced entity (or track).
[0528] Figure 27 shows, according to the embodiment, an example of the type of characteristic assigned to attribute_type. For example, when the value of attribute_type is 0, it can indicate texture, when it is 1, it can indicate material ID, when it is 2, it can indicate transparency, when it is 3, it can indicate reflectivity, and when it is 4, it can indicate normal.
[0529] The entity_id according to the embodiment indicates an identifier for the related item or track. It is resolved by the item if an item with the same item_ID as entity_id exists within the hierarchical level (file, movie, or track) that includes the GroupsListBox. Or it is resolved by the track if a track with the same track_ID as entity_id exists and the GroupsListBox is included at the file level.
[0530] vpcc_profile_tier_level() is the same as profile_tier_level() specified in the sequence parameter set (sequence_parameter_set()).
[0531] Figure 28 shows an example of the syntax structure of profile_tier_level() according to an embodiment.
[0532] The ptl_tier_flag field indicates the codec profile tier used to encode V-PCC content.
[0533] The ptl_profile_idc field indicates the profile information that the encoded point cloud sequence follows.
[0534] The ptl_level_idc field indicates the level of the codec profile that the encoded point cloud sequence follows.
[0535] Hereinafter, the V-PCC track will be described.
[0536] The entry point of each V-PCC content can be represented by a unique V-PCC track. The ISOBMFF file can contain multiple V-PCC contents, so multiple V-PCC tracks exist in the above file. The V-PCC track is identified by the volume visual media bohandler type 'volm' handler type in the handler box of the media box.
[0537] Box Type: 'vohd'
[0538] Container: MediaInformationBox
[0539] Mandatory: Yes
[0540] Quantity: Exactly one
[0541] Volumetric tracks use the VolumetricMediaHeaderBox in the MediaInformationBox
[0542] aligned(8) class VolumetricMediaHeaderBox extends FullBox('vohd', version = 0, 1) {
[0543] / / if we don't need anything here, then use Null Media Header
[0544] }
[0545] / / random access point for the patch stream
[0546] The random access point for the V-PCC track may be a sample that transmits an I-patch frame in an empty decoder buffer. To indicate the random access point for the V-PCC track, there is a syncsampleBox. This sync sample indicates the sync sample that transmits the I-patch frame.
[0547] Some samples in a track that includes track fragments are non - sync samples, but the flag sample_is_non_sync_sample of samples in track fragments is valid and describes the samples. Even if the SyncSampleBox is not present, the SyncSampleBox should be present in the SampleTableBox. If the track is not fragmented and the SyncSampleBox is not present, then all samples in the track are sync samples.
[0548] Box Type: 'stbl'
[0549] Container: MediaInformationBox
[0550] Mandatory: Yes
[0551] Quantity: Exactly one
[0552] aligned(8) class SampleTableBox extends Box('stbl') {
[0553] }
[0554] Box Type: 'stss'
[0555] Container: SampleTableBox
[0556] Mandatory: No
[0557] Quantity: Zero or one
[0558] This box provides compact marking of in-stream sync samples. This table is strictly sorted in ascending order of sample number. If there is no SyncSampleBox, then all samples are sync samples.
[0559] aligned(8) class SyncSampleBox
[0560] extends FullBox('stss', version = 0, 0) {
[0561] unsigned int(32) entry_count;
[0562] int i;
[0563] for (i = 0; i < entry_count; i++) {
[0564] unsigned int(32) sample_number;
[0565] }
[0566] }
[0567] The version is an integer indicating the version of this box.
[0568] The entry_count is an integer and provides the number of entries in the following table. If the value of entry_count is 0, there are no in-stream sync samples and the following table is empty.
[0569] The sample_number provides the sample number for each sync sample in the stream. In particular, the sample_number gives, for each sync sample carrying an I-patch frame in the V-PCC track, its sample number.
[0570] The following describes the V-PCC track sample entry.
[0571] Sample Entry Type: 'vpc1'
[0572] Container: SampleDescriptionBox ('stsd')
[0573] Mandatory: A 'vpc1' sample entry is mandatory
[0574] Quantity: one or more
[0575] The track sample entry type 'vpc1' is used. The V-PCC track sample entry includes a VPCC Configuration Box as defined below. This includes a VPCCDecoderConfigurationRecord as defined. An optional BitRateBox may be present in the V-PCC track sample entry to signal the bitrate information of the V-PCC video stream.
[0576] aligned(8) class VPCCDecoderConfigurationRecord {
[0577] unsigned int(8) configurationVersion = 1;
[0578] unsigned int(8) numOfSequenceParameterSets;
[0579] for (i = 0; i < numOfSequenceParameterSets; i++) {
[0580] sequence_parameter_set();
[0581] }
[0582] / / additional fields
[0583] }
[0584] class VPCCConfigurationBox extends Box('vpcc') {
[0585] VPCCDecoderConfigurationRecord() VPCCConfig;
[0586] }
[0587] aligned(8) class VPCCSampleEntry() extends VolumetricSampleEntry ('vpc1') {
[0588] VPCCConfigurationBox config;
[0589] }
[0590] class VolumetricSampleEntry(codingname) extends SampleEntry (codingname){
[0591] }
[0592] The configurationVersion is a version field. Incompatible changes to the record are indicated by a change in the version number.
[0593] numOfSequenceParameterSets indicates the number of V-PCC sequence parameter sets signaled in the decoder Configuration record.
[0594] The compressor name within the base class VisualSampleEntry indicates the name of the compressor used in the recommended value "\012VPCC Coding".
[0595] The V-PCC specification allows multiple instances (within ids from 1 to 15) of the VPCC_SPS unit. Therefore, VPCCSampleEntry contains multiple sequence_parameter_set unit payloads.
[0596] The V-PCC sample format is described below.
[0597] Each sample within the V-PCC track corresponds to a single-point cloud frame.
[0598] Samples corresponding to this frame in various component tracks have the same composition time as the V-PCC track samples for the frame within the V-PCC track.
[0599] Each V-PCC sample contains only one V-PCC unit payload whose type information in its V-PCC unit header indicates a patch data group (PDG). This V-PCC unit payload contains one or more patch tile group unit payloads as shown in Figure 24.
[0600] aligned(8) class vpcc_unit_payload_struct {
[0601] vpcc_unit_payload();
[0602] }
[0603] aligned(8) class VPCCSample {
[0604] vpcc_unit_payload_struct();
[0605] }
[0606] vpcc_unit_payload() is the payload of a V-PCC unit whose type information in its V-PCC unit header indicates a patch data group (PDG), and contains one instance of patch_data_group_unit().
[0607] For a V-PCC track, a sample that transmits an I patch frame is defined as a sink sample.
[0608] The following describes the V-PCC track reference.
[0609] To link a V-PCC track to a component video track, the track reference tool of the ISOBMFF specification is used. Three TrackReferenceTypeBoxes are added once for each component in the TrackReferenceBox within the track box of the V-PCC track.
[0610] The TrackReferenceTypeBox contains an array of track_IDs that designate the video tracks that reference the V-PCC track.
[0611] The reference_type of the TrackReferenceTypeBox identifies the type of that component (i.e., geometry, texture, or occupancy map). The 4CCs of the new track reference types are 'pcca', 'pccg', 'pcco', etc.
[0612] In the 'pcca' type, the referenced track contains the video-encoded texture V-PCC component.
[0613] In the 'pccg' type, the referenced track contains the video-encoded geometry V-PCC component.
[0614] In the 'pcco' type, the referenced track contains the video-encoded occupancy map V-PCC component.
[0615] The following describes the video encoded V-PCC component tracks.
[0616] Since it does not make sense to display the frames decoded from the texture, geometry, or occupancy map tracks without reconstructing the point cloud on the player side, a restricted video scheme type is defined for that video-encoded track. The V-PCC video track contains the 4CC identifier 'pccv'. This 4CC identifier 'pccv' is included in the scheme_type field of the SchemeTypeBox of the RestrictedSchemeInfoBox of the restricted sample entry.
[0617] The use of the V-PCC video scheme for the restricted video sample entry type'resv' indicates that the decoded picture contains point cloud attributes, geometry, or occupancy map data.
[0618] The use of the V-PCC video scheme is represented by the same scheme_type as 'pccv' (video-based point cloud video) in the SchemeTypeBox of the RestrictedSchemeInfoBox.
[0619] This box is a container that contains boxes indicating V-PCC specific information for this track. The VPCCVideoBox provides PCC specific parameters applicable to all samples within the track.
[0620] Box Type: 'pccv'
[0621] Container: SchemeInformationBox
[0622] Mandatory: Yes, when scheme_type is equal to 'pccv'
[0623] Quantity: Zero or one
[0624] aligned(8) class VPCCVideoBox extends FullBox('vpcc', 0, 0) {
[0625] unsigned int(4) data_type;
[0626] unsigned int(4) attribute_count;
[0627] for(int i=0; i<attribute_count+1;i++) {
[0628] unsigned int(4) attribute_type;
[0629] }
[0630] unsigned int(1) multiple_layer_present_flag;
[0631] unsigned int(4) layer_count_minus1;
[0632] for (i = 0 ; i< layer_count_minus1+1 ; i++) {
[0633] unsigned int(4) layer_id;
[0634] }
[0635] }
[0636] The data_type indicates the data type included in the video samples contained in the corresponding track.
[0637] Figure 29 shows an example of the type of PCC data in the referenced track according to the embodiment. For example, when the value of data_type is 1, it indicates a V-PCC track; when it is 2, it indicates a geometry video track; when it is 3, it indicates a texture video track; and when it is 4, it indicates an occupancy map video track.
[0638] The multiple_layer_present_flag indicates whether a single geometry or characteristic layer or multiple geometries or characteristic layers are transmitted on this track. For example, if the multiple_layer_present_flag is 0, it indicates that a single geometry or characteristic layer is transmitted on this track, and if it is 1, it indicates that multiple geometries or characteristic layers are transmitted on this track. As another example, when the above data_type indicates a geometry video track, if the value of the multiple_layer_present_flag is 0, the above geometry video track transmits a single geometry layer, and if the value of the multiple_layer_present_flag is 1, it indicates that the above geometry video track transmits multiple geometry layers. As another example, when the above data_type indicates a characteristic video track, if the value of the multiple_layer_present_flag is 0, the above characteristic video track transmits a single characteristic layer, and if the value of the multiple_layer_present_flag is 1, it indicates that the above characteristic video track transmits multiple characteristic layers.
[0639] layer_count_minus1 plus 1 indicates the number of geometry and / or characteristic layers transmitted on this track. For example, when the above data_type indicates a geometry video track, layer_count_minus1 plus 1 indicates the number of geometry layers transmitted on the geometry video track. As another example, when the above data_type indicates a characteristic video track, layer_count_minus1 plus 1 indicates the number of characteristic layers transmitted on the characteristic video track.
[0640] The layer_id indicates the layer identifier of the geometry and / or property layer associated with the samples within this track. For example, when the above data_type indicates a geometry video track, the layer_id indicates the layer identifier of the geometry layer. As another example, when the above data_type indicates a property video track, the layer_id indicates the layer identifier of the property layer. The attribute_count can indicate the number of property data of the point cloud stream included in the corresponding track. One or more property data may be included in one track. In this case, in one embodiment, the above data_type indicates a property video track.
[0641] The attribute_type indicates the attribute type of the property video data transmitted to the corresponding track.
[0642] Figure 30 shows an example of an attribute type according to an embodiment. For example, if the value of the attribute_type field is 0, it indicates a texture, if it is 1, it indicates a material ID, if it is 2, it indicates transparency, if it is 3, it indicates reflectivity, and if it is 4, it indicates a normal.
[0643] When the field value existing in the VPCCVideoBox is included in the signaling related to the item, the above information regarding the corresponding item (image) can be indicated.
[0644] Hereinafter, attribute sample grouping will be described.
[0645] When the track includes property data of a point cloud, it can include data of one or more attribute types. In this case, as follows, it is possible to signal the data attribute types included in the samples included in the corresponding track.
[0646] class PCCAttributeSampleGroupEntry extends VisualSampleGroupEntry('pcca') {
[0647] unsigned int(4) attribute_type;
[0648] }
[0649] The attribute_type indicates the attribute type of the texture video data transmitted to the relevant samples within the corresponding track.
[0650] Figure 31 shows an example of the attribute type according to an embodiment. For example, if the value of the attribute_type field is 0, it indicates texture, if it is 1, it indicates material ID, if it is 2, it indicates transparency, if it is 3, it indicates reflectivity, and if it is 4, it indicates normal.
[0651] Next, layer sample grouping will be described.
[0652] It can include data related to one or more layers of the same data type (e.g., geometry data, attribute data) of the point cloud of the track. In this case, as follows, it is possible to signal the layer-related information of the data included in the samples contained in the corresponding track.
[0653] class PCCLayerSampleGroupEntry extends VisualSampleGroupEntry('pccl') {
[0654] unsigned int(4) pcc_layer_count_minus1;
[0655] unsigned int(4) layer_id;
[0656] }
[0657] pcc_layer_count_minus1 plus 1 indicates the number of layers used to encode the geometry and texture components of the point cloud stream.
[0658] layer_id indicates the layer identifier for the associated samples within the corresponding track.
[0659] The following describes layer tracking grouping.
[0660] From the perspective of V-PCC, in addition to the geometry, all other types of information (except occupancy and patch information) may have the same number of layers. Also, all information (geometry / texture) tagged with the same layer index is related to each other. That is, the information tagged with the same layer index refers to points that are similarly reconstructed in 3D space.
[0661] To reconstruct the texture with layer index M, the geometry layer with index M is also available. If for some reason that geometry layer is missing, since both are 'linked' to each other, that texture information is not needed.
[0662] From the perspective of rendering, the information is considered as 'augmenting' the reconstructed point cloud. This means that by discarding some layers, a reasonable reconstruction can be obtained. However, in one embodiment, layer 0 is not discarded.
[0663] In one embodiment, tracks belonging to the same layer have the same value of track_group_id for track_group_type 'pccl', and the track_group_id of tracks from one layer is different from the track_group_id of tracks from different layers.
[0664] Basically (by default), if this track grouping is not instructed for any track in the file, the file is considered to contain content for only one layer.
[0665] aligned(8) class PCCLayerTrackGroupBox extends TrackGroupTypeBox('pccl') {
[0666] unsigned int(4) pcc_layer_count_minus1;
[0667] unsigned int(4) layer_id;
[0668] }
[0669] pcc_layer_count_minus1 plus 1 indicates the number of layers used to encode the geometry and / or attribute components of the point cloud stream.
[0670] layer_id indicates the layer identifier for the associated pcc data (e.g., geometry and / or attribute) layer transmitted to the corresponding track.
[0671] This can inform the player in Figure 20 or the renderer in Figures 1, 19, 20, or 22 that the data of the tracks belonging to the same track group is used to restore the point cloud of the same layer.
[0672] As described above, this specification signals signaling information for the client / player to be able to select and use the attribute non-stream as needed to at least one of the V-PCC track and the attribute track.
[0673] In addition, this specification signals the signaling information related to the random access points among the samples in the V-PCC track to at least one of the V-PCC track and the texture track.
[0674] Therefore, the client / player can obtain the signaling information related to the random access points, and obtain and use the necessary texture bitstreams based on the signaling information, enabling efficient and faster data processing.
[0675] In addition, when the in-file geometry and / or texture video data consists of multiple layers, this specification signals the track grouping information used together to at least one of the V-PCC track and the corresponding video track (i.e., the geometry track or the texture track).
[0676] In addition, when the geometry and / or texture video data in one track is configured as multiple layers, this specification signals the sample grouping information for this to at least one of the V-PCC track and the corresponding video track (i.e., the geometry track or the texture track).
[0677] In addition, when one track contains multiple texture video data, this specification signals the sample grouping information for this to at least one of the V-PCC track and the texture track.
[0678] On one hand, a patch frame (or a point cloud object or an atlas frame) that is the target of point cloud data may be divided into one or multiple tiles as described above. A tile according to an embodiment may indicate a certain area in 3D space or a certain area on a 2D plane. Also, a tile may be a rectangular cuboid or a sub - bounding box or a part of a patch frame within one bounding box. In this specification, dividing a patch frame (or a point cloud object) into one or multiple tiles may be performed in the point cloud video encoder of FIG. 1, the patch generation unit of FIG. 18, the point cloud pre - processing unit of FIG. 20, or the patch generation unit of FIG. 21, or may be performed by another component / module.
[0679] FIG. 32 shows an example of dividing one or more patch frames in the row direction, dividing one or more in the column direction, and dividing one patch frame into multiple tiles. One tile is a rectangular region within one patch frame, and a tile group can include a number of tiles within the patch frame. This specification states that a tile group includes multiple tiles of the patch frame that collectively form a rectangular (or quadrilateral) region of the patch frame. In particular, the patch frame of FIG. 32 is shown divided into 24 tiles (= 6 tiles in the column direction * 4 tiles in the row direction) and divided into 9 rectangular (or quadrilateral) patch tile groups.
[0680] At this time, when the type information of the PCC unit header in FIG. 24 indicates a patch data group (VPCC - PDG) including at least one or more patches, the corresponding V - PCC unit payload includes the patch data group. Conversely, when the V - PCC unit payload includes a patch data group, the type information of the corresponding V - PCC unit header can identify this.
[0681] This patch data group may include information of one or more patch tile groups, and at least one of a patch sequence parameter set, a patch frame parameter set, a patch frame geometry parameter set, a geometry patch parameter set, a patch frame characteristic parameter set, and a characteristic patch parameter set.
[0682] FIG. 33 shows an example of the syntax structure of each V-PCC unit according to an embodiment. Each V-PCC unit consists of a V-PCC unit header and a V-PCC unit payload. The V-PCC unit in FIG. 33 can contain more data, and in this case, it can further include a trailing_zero_8bits field. The trailing_zero_8bits field according to the embodiment corresponds to bytes of 0x00.
[0683] FIG. 34 shows an example of the syntax structure of a V-PCC unit header according to an embodiment. In one embodiment, the V-PCC unit header (vpcc_unit_header()) in FIG. 34 includes a vpcc_unit_type field. The vpcc_unit_type field indicates the type of the corresponding V-PCC unit.
[0684] FIG. 35 shows an example of the type of V-PCC unit assigned to the vpcc_unit_type field according to an embodiment.
[0685] Referring to FIG. 35, in one embodiment, if the value of the vpcc_unit_type field is 0, it indicates that the data included in the V-PCC unit payload of the corresponding V-PCC unit is a sequence parameter set (VPCC_SPS); if it is 1, it indicates a patch data group (VPCC_PDG); if it is 2, it indicates occupied video data (VPCC_OVD); if it is 3, it indicates attribute video data (VPCC_AVD); and if it is 4, it indicates geometry video data (VPCC_GVD).
[0686] Since those skilled in the art can easily change the meaning, procedures, deletion, addition, etc. of the values assigned to the vpcc_unit_type field, the present invention is not limited to the above embodiments.
[0687] At this time, the V-PCC unit payload follows the format of an HEVC NAL unit. That is, the occupied, geometry, or attribute video data V-PCC unit payload according to the vpcc_unit_type field value corresponds to a video data unit (e.g., an HEVC NAL unit) that can be decoded by the video decoder specified in the corresponding occupied, geometry, or attribute parameter set V-PCC unit.
[0688] In one embodiment, when the vpcc_unit_type field indicates attribute video data (VPCC_AVD) or geometry video data (VPCC_GVD) or occupied video data (VPCC_OVD) or a patch data group (VPCC_PDG), the corresponding V-PCC unit header further includes a vpcc_sequence_parameter_set_id field.
[0689] This vpcc_sequence_parameter_set_id field specifies the identifier of the active sequence parameter set (VPCC SPS), i.e., sps_sequence_parameter_set_id. The value of the sps_sequence_parameter_set_id field may be in the range of 0 to 15.
[0690] In one embodiment, when the vpcc_unit_type field indicates video data with special characteristics (VPCC_AVD), the V-PCC unit header further includes a vpcc_attribute_index field and a vpcc_attribute_dimension_index field.
[0691] This vpcc_attribute_index field indicates the index of the video data with special characteristics transmitted to the video data unit with special characteristics. The value of this vpcc_attribute_index field may be in the range from 0 to (ai_attribute_count - 1).
[0692] This vpcc_attribute_dimension_index field indicates the index of the special dimension group transmitted to the video data unit with special characteristics. The value of this vpcc_attribute_dimension_index field may be in the range of 0 to 127.
[0693] The sps_multiple_layer_streams_present_flag field of the V-PCC unit header in Figure 34 indicates whether it includes a vpcc_layer_index field and a pcm_separate_video_data(11) field.
[0694] For example, if the value of the vpcc_unit_type field indicates texture video data (VPCC_AVD) and the value of the sps_multiple_layer_streams_present_flag field is true (e.g., 0), the corresponding V-PCC unit header further includes a vpcc_layer_index field and a pcm_separate_video_data(11) field. That is, if the value of the sps_multiple_layer_streams_present_flag field is true, it means that multiple layers for texture video data or geometry video data exist. In this case, a field (e.g., vpcc_layer_index) indicating the index of the current layer is required.
[0695] This vpcc_layer_index field indicates the index of the current layer of the texture video data. The vpcc_layer_index field has a value between 0 and 15.
[0696] For example, if the value of the vpcc_unit_type field indicates texture video data (VPCC_AVD) and the value of the sps_multiple_layer_streams_present_flag field is false (e.g., 1), the corresponding V-PCC unit header further includes a pcm_separate_video_data(15) field. That is, if the value of the sps_multiple_layer_streams_present_flag field is false, it means that no multiple layers for texture video data and / or geometry video data exist. In this case, a field indicating the index of the current layer is not required.
[0697] For example, if the value of the vpcc_unit_type field indicates geometric video data (VPCC_GVD) and the value of the sps_multiple_layer_streams_present_flag field is true (e.g., 0), the corresponding V-PCC unit header further includes a vpcc_layer_index field and a pcm_separate_video_data(18) field.
[0698] The vpcc_layer_index field indicates the index of the current layer of the geometric video data. The vpcc_layer_index field has a value between 0 and 15.
[0699] For example, if the value of the vpcc_unit_type field indicates geometric video data (VPCC_GVD) and the value of the sps_multiple_layer_streams_present_flag field is false (e.g., 1), the corresponding V-PCC unit header further includes a pcm_separate_video_data(22) field.
[0700] For example, if the value of the vpcc_unit_type field indicates occupied video data (VPCC_OVD) or indicates a patch data group (VPCC_PDG), the corresponding V-PCC unit header further includes a vpcc_reserved_zero_23bits field; otherwise, it further includes a vpcc_reserved_zero_27bits field.
[0701] On the other hand, the V-PCC unit header in FIG. 34 may further include a vpcc_pcm_video_flag field.
[0702] For example, if the value of this vpcc_pcm_video_flag field is 1, the related geometry video data unit or texture video data unit can indicate that it contains only Pulse Code Modulation (PCM) - coded points. As another example, if the value of the vpcc_pcm_video_flag field is 0, the related geometry video data unit or texture video data unit can indicate that it contains non - PCM - coded points. If the vpcc_pcm_video_flag field does not exist, it can be inferred that the field value is 0.
[0703] Figure 36 shows an example of the syntax structure of pcm_separate_video_data(bitCount) included in the V - PCC unit header according to the embodiment.
[0704] In Figure 36, the bitCount of pcm_separate_video_data is different according to the vpcc_unit_type field value in the V - PCC unit header of Figure 34 as described above.
[0705] pcm_separate_video_data(bitCount) can include the vpcc_pcm_video_flag field and the vpcc_reserved_zero_bitcount_bits field when the value of the sps_pcm_separate_video_present_flag field is true and it is not the vpcc_layer_index field, or can include the vpcc_reserved_zero_bitcountplus1_bits field otherwise.
[0706] If the value of this vpcc_pcm_video_flag field is 1, it can indicate that the related geometry or texture video data unit is only PCM-coded point video. Also, if the value of this vpcc_pcm_video_flag field is 0, it can indicate that the related geometry or texture video data unit contains PCM-coded points. If this vpcc_pcm_video_flag field does not exist, its value is considered to be 0.
[0707] Figure 37 shows an example of the syntax structure of a V-PCC unit payload according to an embodiment.
[0708] The V-PCC unit payload in Figure 37 contains one of a sequence parameter set (sequence_parameter_set()), a patch data group (patch_data_group()), and a video data unit (video_data_unit()) according to the vpcc_unit_type field value of the corresponding V-PCC unit header.
[0709] For example, when the vpcc_unit_type field indicates a sequence parameter set (VPCC_SPS), the V-PCC unit payload includes a sequence parameter set (sequence_parameter_set()), and when it indicates a patch data group (VPCC_PDG), it includes a patch data group (patch_data_group()). Also, when the vpcc_unit_type field indicates occupied video data (VPCC_OVD), the V-PCC unit payload includes a video data unit (video_data_unit()) for transmitting the occupied video data, when it indicates geometric video data (VPCC_GVD), it includes a geometric video data unit (video_data_unit()) for transmitting the geometric video data, and when it indicates attribute video data (VPCC_AVD), it includes an attribute video data unit (video_data_unit()) for transmitting the attribute video data, as an example of one embodiment.
[0710] FIG. 38 shows an example of the syntax structure of a sequence parameter set () included in a V-PCC unit payload according to an embodiment.
[0711] The sequence parameter set (SPS) in FIG. 38 can be applied to an encoded point cloud sequence including sequences of encoded geometric video data units, attribute video data units, and occupied video data units.
[0712] The sequence parameter set (SPS) in FIG. 38 can include profile_tier_level(), sps_sequence_parameter_set_id field, sps_frame_width field, sps_frame_height field, sps_avg_frame_rate_present_flag field.
[0713] This profile_tier_level() indicates codec information used to compress the sequence parameter set.
[0714] This sps_sequence_parameter_set_id field provides an identifier for the sequence parameter set for reference by other syntax elements.
[0715] This sps_frame_width field indicates the width of the nominal frame in terms of integer luma samples.
[0716] This sps_frame_height field indicates the height of the nominal frame in terms of integer luma samples.
[0717] This sps_avg_frame_rate_present_flag field indicates the presence or absence of average nominal frame rate information in this bitstream. For example, if the value of the sps_avg_frame_rate_present_flag field is 0, it indicates that there is no average nominal frame rate information in this bitstream. If the value of this sps_avg_frame_rate_present_flag field is 1, it can indicate that the average nominal frame rate information is present in this bitstream. For example, if the value of the sps_avg_frame_rate_present_flag field is true, i.e., 1, the sequence parameter set can further include the sps_avg_frame_rate field, the sps_enhanced_occupancy_map_for_depth_flag field, and the sps_geometry_attribute_different_layer_flag field.
[0718] This sps_avg_frame_rate field indicates the average nominal point cloud frame rate in units of point cloud frames per 256 seconds. If this sps_avg_frame_rate field does not exist, the value of that field shall be 0. During the reconstruction phase, the decoded occupancy, geometry, and texture videos are converted to nominal width, height, and frame rate using appropriate scaling.
[0719] The sps_enhanced_occupancy_map_for_depth_flag field indicates whether the decoded occupancy map video contains information on whether intermediate depth positions between two depth layers are occupied. For example, if the value of this sps_enhanced_occupancy_map_for_depth_flag field is 1, it can indicate the presence or absence of information regarding whether intermediate depth positions between two depth layers in the decoded occupancy map video are occupied. If the value of this sps_enhanced_occupancy_map_for_depth_flag field is 0, it can indicate that the decoded occupancy map video does not contain information regarding whether intermediate depth positions between two depth layers are occupied.
[0720] The sps_geometry_attribute_different_layer_flag field indicates whether the number of layers used to encode the geometry and texture video data is different. For example, if the value of the sps_geometry_attribute_different_layer_flag field is 1, it can indicate that the number of layers used to encode the geometry and texture video data is different. As an example, two layers may be used for encoding the geometry video data, and one layer may be used for encoding the texture video data. Also, if the value of the sps_geometry_attribute_different_layer_flag field is 1, it indicates whether the number of layers used to encode the geometry and texture video data is signaled in the patch sequence data unit.
[0721] The sps_geometry_attribute_different_layer_flag field indicates the presence or absence of the sps_layer_count_geometry_minus1 field and the sps_layer_count_minus1 field. For example, if the value of the sps_geometry_attribute_different_layer_flag field is true (e.g., 1), it may further include the sps_layer_count_geometry_minus1 field, and if it is false (e.g., 0), it may further include the sps_layer_count_minus1 field.
[0722] The sps_layer_count_geometry_minus1 field indicates the number of layers used to encode the geometry video data.
[0723] The sps_layer_count_minus1 field indicates the number of layers used to encode the geometry and texture video data.
[0724] If the value of the sps_layer_count_minus1 field is greater than 0, the sequence parameter set may further include the sps_multiple_layer_streams_present_flag field and the sps_layer_absolute_coding_enabled_flag [0]=1 field.
[0725] The sps_multiple_layer_streams_present_flag field indicates whether the geometry layers or texture layers are located in a single video stream or in separate video streams. For example, if the value of the sps_multiple_layer_streams_present_flag field is 0, it indicates that all geometry layers or texture layers are respectively placed in a single geometry video stream or a single texture video stream. If the value of the sps_multiple_layer_streams_present_flag field is 1, it can indicate that all geometry layers or texture layers are placed in separate video streams.
[0726] In addition, the sequence parameter set (SPS) includes a loop statement that is repeated only the number of times specified by the value of the sps_layer_count_minus1 field, and this loop statement includes the sps_layer_absolute_coding_enabled_flag field. At this time, i is initialized to 0, and is incremented by 1 each time the loop statement is executed. In one embodiment, the loop statement is repeated until the i value becomes equal to the value of the sps_layer_count_minus1 field. In addition, when the value of the sps_layer_absolute_coding_enabled_flag field is 0 and the i value is greater than 0, it further includes the sps_layer_predictor_index_diff field; otherwise, it does not include the sps_layer_predictor_index_diff field.
[0727] If the value of the sps_layer_absolute_coding_enabled_flag [i] field is 1, it can indicate that the geometry layer with index i is coded without any form of layer prediction. If the value of the sps_layer_absolute_coding_enabled_flag [i] field is 0, it can indicate that the geometry layer with index i is predicted first from another, earlier coded layer before coding.
[0728] The sps_layer_predictor_index_diff [i] field indicates that it is used for calculating the predictor of the geometry layer with index i if the value of the sps_layer_absolute_coding_enabled_flag [i] field is 0.
[0729] The sequence parameter set (SPS) according to this specification may further include the sps_pcm_patch_enabled_flag field. The sps_pcm_patch_enabled_flag field indicates the presence or absence of the sps_pcm_separate_video_present_flag field, the occupancy_parameter_set(), the geometry_parameter_set(), and the sps_attribute_count field. For example, if the sps_pcm_patch_enabled_flag field is 1, the sps_pcm_separate_video_present_flag field, the occupancy_parameter_set(), the geometry_parameter_set(), and the sps_attribute_count field may be further included. That is, if the sps_pcm_patch_enabled_flag field is 1, it indicates that there is a patch with PCM-coded points in the bitstream.
[0730] The sps_pcm_separate_video_present_flag field indicates whether the PCM-coded geometry video data and the texture video data are stored in separate video streams. For example, if the value of the sps_pcm_separate_video_present_flag field is 1, it indicates that the PCM-coded geometry video data and the texture video data are stored in a separate video stream.
[0731] The occupancy_parameter_set() includes information about the occupancy map.
[0732] The geometry_parameter_set() includes information about the geometry video data.
[0733] The sps_attribute_count field indicates the number of attributes associated with the point cloud.
[0734] Also, a sequence parameter set (SPS) according to this specification includes a repetition statement that is repeated only the value of the sps_attribute_count field. In one embodiment, this repetition statement includes the sps_layer_count_attribute_minus1 field and the attribute_parameter_set() if the sps_geometry_attribute_different_layer_flag field is 1. In this repetition statement, i is initialized to 0, incremented by 1 each time the repetition statement is executed, and repeated until the i value becomes the value of the sps_attribute_count field in one embodiment.
[0735] The sps_layer_count_attribute_minus1 [i] field indicates the number of layers used to encode the i-th attribute video data associated with the corresponding point cloud.
[0736] The attribute_parameter_set(i) includes information about the i-th attribute video data associated with the corresponding point cloud.
[0737] In one embodiment, a sequence parameter set (SPS) according to this specification further includes a sps_patch_sequence_orientation_enabled_flag field, a sps_patch_inter_prediction_enabled_flag field, a sps_pixel_deinterleaving_flag field, a sps_point_local_reconstruction_enabled_flag field, a sps_remove_duplicate_point_enabled_flag field, and a byte_alignment() field.
[0738] The sps_patch_sequence_orientation_enabled_flag field indicates whether flexible orientation is signaled in the patch sequence data unit. For example, if the value of the sps_patch_sequence_orientation_enabled_flag field is 1, it indicates that flexible orientation is signaled in the patch sequence data unit, and if it is 0, it indicates that it is not signaled.
[0739] If the value of the sps_patch_inter_prediction_enabled_flag field is 1, it indicates that inter prediction for patch information is used using patch information from previously encoded patch frames.
[0740] If the value of the sps_pixel_deinterleaving_flag field is 1, it indicates that the decoded geometry and texture video corresponding to a single stream contain pixels interleaved from two layers. If the value of the sps_pixel_deinterleaving_flag field is 0, it indicates that the decoded geometry and texture video corresponding to a single stream contain only pixels interleaved from a single layer.
[0741] If the value of the sps_point_local_reconstruction_enabled_flag field is 1, it indicates that the local reconstruction mode is used during the point cloud reconstruction process.
[0742] If the value of the sps_remove_duplicate_point_enabled_flag field is 1, it indicates that duplicated points are not reconstructed. Here, a duplicated point is a point that has the same 2D and 3D geometry coordinates as another point from the lower layer.
[0743] Figure 39 shows an example of the syntax structure of a patch_data_group() according to an embodiment.
[0744] As described above, when the value of the vpcc_unit_type field in the V-PCC unit header indicates a patch data group, the V-PCC unit payload in FIG. 37 includes the patch_data_group() in FIG. 39.
[0745] In one embodiment, a patch_data_group() includes a pdg_unit_type field, a patch_data_group_unit_payload(pdg_unit_type) in which the signaling information varies according to the value of the pdg_unit_type field, and a pdg_terminate_patch_data_group_flag field.
[0746] The pdg_unit_type field indicates the type of the patch data group.
[0747] The pdg_terminate_patch_data_group_flag field indicates the end of a patch data group. If the value of the pdg_terminate_patch_data_group_flag field is 0, it indicates that there is a further patch data group unit in the corresponding patch data group. Also, if the value of the pdg_terminate_patch_data_group_flag field is 1, it indicates that there is no further patch data group unit in the corresponding patch data group, and this is the end of the current patch data group unit.
[0748] Figure 40 shows an example of the type of patch data group assigned to the pdg_unit_type field of the patch data group in Figure 39.
[0749] For example, if the value of the pdg_unit_type field is 0, it indicates a Patch Sequence Parameter Set (PDG_PSPS), if it is 1, it indicates a Patch Frame Parameter Set (PDG_PFPS), if it is 2, it indicates a Patch Frame Geometry Parameter Set (PDG_PFGPS), if it is 3, it indicates a Patch Frame Attribute Parameter Set (PDG_PFAPS), if it is 4, it indicates a Geometry Patch Parameter Set (PDG_GPPS), if it is 5, it indicates an Attribute Patch Parameter Set (PDG_APPS), if it is 6, it indicates a Patch Tile Group Layer Unit (PDG_PTGLU), if it is 7, it indicates a Prefix SEI Message (PDG_PREFIX_SEI), and if it is 8, it indicates a Suffix SEI Message (PDG_SUFFIX_SEI).
[0750] The patch sequence parameter set (PDG_PSPS) includes sequence-level parameters, and the patch frame parameter set (PDG_PFPS) can include frame-level parameters. The patch frame geometry parameter set (PDG_PFGPS) includes frame-level geometry type parameters, and the patch frame attribute parameter set (PDG_PFAPS) can include frame-level attribute type parameters. The geometry patch parameter set (PDG_GPPS) includes patch-level geometry type parameters, and the attribute patch parameter set (PDG_APPS) can include patch-level attribute type parameters.
[0751] Figure 41 shows an example of the syntax structure of a patch data group unit payload (patch_data_group_unit_payload(pdg_unit_type)) according to an embodiment.
[0752] When the value of the pdg_unit_type field of the patch data group in FIG. 39 indicates the patch sequence parameter set (PDG_PSPS), the patch data group unit payload (patch_data_group_unit_payload()) can include the patch sequence parameter set (patch_sequence_parameter_set()).
[0753] When the value of the pdg_unit_type field indicates the geometry patch parameter set (PDG_GPPS), the patch data group unit payload (patch_data_group_unit_payload()) can include the geometry patch parameter set (geometry_patch_parameter_set()).
[0754] When the value of the pdg_unit_type field indicates an attribute patch parameter set (PDG_APPS), the patch data group unit payload (patch_data_group_unit_payload()) can include an attribute patch parameter set (attribute_patch_parameter_set()).
[0755] When the value of the pdg_unit_type field indicates a patch frame parameter set (PDG_PFPS), the patch data group unit payload (patch_data_group_unit_payload()) can include a patch frame parameter set (patch_frame_parameter_set()).
[0756] When the value of the pdg_unit_type field indicates a patch frame attribute parameter set (PDG_PFAPS), the patch data group unit payload (patch_data_group_unit_payload()) can include a patch frame attribute parameter set (patch_frame_attribute_parameter_set()).
[0757] When the value of the pdg_unit_type field indicates a patch frame geometry parameter set (PDG_PFGPS), the patch data group unit payload (patch_data_group_unit_payload()) can include a patch frame geometry parameter set (patch_frame_geometry_parameter_set()).
[0758] When the value of the pdg_unit_type field indicates a patch tile group layer unit (PDG_PTGLU), the patch data group unit payload (patch_data_group_unit_payload()) can include a patch tile group layer unit (patch_tile_group_layer_unit()).
[0759] When the value of the pdg_unit_type field indicates a prefix SEI message (PDG_PREFIX_SEI) or a suffix SEI message (PDG_SUFFIX_SEI), the patch data group unit payload (patch_data_group_unit_payload()) can include a sei_message().
[0760] Figure 42 shows an example of the syntax structure of a Supplemental Enhancement Information (SEI) message (sei_message()) according to an embodiment. That is, when the value of the pdg_unit_type field included in the patch data group of FIG. 39 indicates a prefix SEI message (PDG_PREFIX_SEI) or a suffix SEI message (PDG_SUFFIX_SEI), the patch data group unit payload of FIG. 41 includes an SEI message (sei_message()).
[0761] Each SEI message according to an embodiment consists of an SEI message header and an SEI message payload. The SEI message header includes an sm_payload_type_byte field and an sm_payload_size_byte field.
[0762] The sm_payload_type_byte field is the byte of the payload type of the corresponding SEI message. For example, based on the value of the sm_payload_type_byte field, it is possible to identify whether it is a prefix SEI message or a suffix SEI message.
[0763] The sm_payload_size_byte field is the byte of the payload size of the corresponding SEI message.
[0764] For the SEI message in Figure 42, after initializing the PayloadType value to 0, the value of the sm_payload_type_byte field in the loop is set to the PayloadType value, and if the value of the sm_payload_type_byte field is 0xFF, the loop ends.
[0765] Also, for the SEI message in Figure 42, after initializing the payload size value to 0, the value of the sm_payload_size_byte field in the loop is set to the payload size value, and if the value of the sm_payload_size_byte field is 0xFF, the loop ends.
[0766] Then, the information corresponding to the payload type and payload size set in the above two loops is signaled by the payload of the SEI message (sei_payload(payloadType, payloadSize)).
[0767] Figure 43 shows an example of the V-PCC bitstream structure according to another embodiment of this specification. In one embodiment, the V-PCC bitstream in Figure 43 is generated and output from the V-PCC based point cloud video encoder in Figures 1, 4, 18, 20 or 21.
[0768] The V-PCC bitstream according to the embodiment includes a coded point cloud sequence (CPCS) and consists of sample stream V-PCC units. This sample stream V-PCC unit transmits V-PCC parameter set (VPS) data, an atlas bitstream, a 2D video encoded occupancy map bitstream, a 2D video encoded geometry bitstream, and zero or more 2D video encoded attribute bitstreams.
[0769] In FIG. 43, the V-PCC bitstream can include one sample stream V-PCC header 40010 and one or more sample stream V-PCC units 40020. For the sake of convenience of explanation, one or more sample stream V-PCC units 40020 may also be referred to as the sample stream V-PCC payload. That is, the sample stream V-PCC payload may also be referred to as a set of sample stream V-PCC units.
[0770] Each sample stream V-PCC unit 40021 consists of V-PCC unit size information 40030 and a V-PCC unit 40040. The V-PCC unit size information 40030 indicates the size of the V-PCC unit 40040. For the sake of convenience of explanation, the V-PCC unit size information 40030 may also be referred to as the sample stream V-PCC unit header, and the V-PCC unit 40040 may also be referred to as the sample stream V-PCC unit payload.
[0771] Each V-PCC unit 40040 consists of a V-PCC unit header 40041 and a V-PCC unit payload 40042.
[0772] This specification classifies the data included in the corresponding V-PCC unit payload 40042 by the V-PCC unit header 40041. To this end, the V-PCC unit header 40041 includes type information indicating the type of the corresponding V-PCC unit. Each V-PCC unit payload 40042 can include at least one of geometric video data (i.e., a 2D video-encoded geometric bitstream), texture video data (i.e., a 2D video-encoded texture bitstream), occupancy video data (i.e., a 2D video-encoded occupancy map bitstream), atlas data, and a V-PCC parameter set (VPS) according to the type information of the corresponding V-PCC unit header 40041.
[0773] The V-PCC parameter set (VPS) according to the embodiment is also called a sequence parameter set (SPS) and may be used interchangeably.
[0774] The atlas data according to the embodiment may mean data composed of characteristics (e.g., texture (patch)) and / or depth of point cloud data, and may also be called a patch data group.
[0775] FIG. 44 shows an example of data transmitted by a sample stream V-PCC unit in a V-PCC bitstream according to an embodiment.
[0776] The V-PCC bitstream in FIG. 44 is an illustration including a sample stream V-PCC unit for transmitting a V-PCC parameter set (VPS), a sample stream V-PCC unit for transmitting atlas data (AD), a sample stream V-PCC unit for transmitting occupied video data (OVD), a sample stream V-PCC unit for transmitting geometry video data (GVD), and a sample stream V-PCC unit for transmitting attribute video data (AVD).
[0777] According to an embodiment, each sample stream V-PCC unit includes a V-PCC unit of one type among a V-PCC parameter set (VPS), atlas data (AD), occupied video data (OVD), geometry video data (GVD), and attribute video data (AVD).
[0778] FIG. 45 shows an example of the syntax structure of a sample stream V-PCC header included in a V-PCC bitstream according to an embodiment.
[0779] The sample stream V-PCC header () according to an embodiment can include an ssvh_unit_size_precision_bytes_minus1 field and an ssvh_reserved_zero_5bits field.
[0780] The ssvh_unit_size_precision_bytes_minus1 field can indicate the precision of the ssvu_vpcc_unit_size element in all sample stream V-PCC units in bytes by adding 1 to this field value. The value of this field is in the range of 0 to 7.
[0781] The ssvh_reserved_zero_5bits field is a reserved field for future use.
[0782] Figure 46 shows an example of the syntax structure of the sample stream V-PCC unit (sample_stream_vpcc_unit()) according to the embodiment.
[0783] The content of each sample stream V-PCC unit is associated with the same access unit as the V-PCC unit contained in the sample stream V-PCC unit.
[0784] The sample stream V-PCC unit (sample_stream_vpcc_unit()) according to the embodiment can include an ssvu_vpcc_unit_size field and vpcc_unit(ssvu_vpcc_unit_size).
[0785] The ssvu_vpcc_unit_size field corresponds to the V-PCC unit size information 40030 in FIG. 43 and specifies the size of the subsequent V-PCC unit 40040 in bytes. The number of bits used to indicate the ssvu_vpcc_unit_size field seems to be (ssvh_unit_size_precision_bytes_minus1 + 1) * 8.
[0786] vpcc_unit(ssvu_vpcc_unit_size) has a length corresponding to the value of the ssvu_vpcc_unit_size field and transmits one of the V-PCC parameter set (VPS), atlas data (AD), occupied video data (OVD), geometry video data (GVD), and attribute video data (AVD).
[0787] Figure 47 shows an example of the syntax structure of a V-PCC unit according to an embodiment. One V-PCC unit consists of a V-PCC unit header (vpcc_unit_header()) and a V-PCC unit payload (vpcc_unit_payload()). The V-PCC unit according to the embodiment can contain more data, and in this case, it can further contain a trailing_zero_8bits field. The trailing_zero_8bits field according to the embodiment is a byte corresponding to 0x00.
[0788] Figure 48 shows an example of the syntax structure of a V-PCC unit header according to an embodiment. In one embodiment, the V-PCC unit header (vpcc_unit_header()) in Figure 48 includes a vuh_unit_type field. The vuh_unit_type field indicates the type of the corresponding V-PCC unit. The vuh_unit_type field according to the embodiment is also referred to as the vpcc_unit_type field.
[0789] Figure 49 shows an example of the type of V-PCC unit assigned to the vuh_unit_type field according to an embodiment.
[0790] Referring to Figure 49, in one embodiment, if the value of the vuh_unit_type field is 0, it indicates that the data included in the V-PCC unit payload of the corresponding V-PCC unit is a V-PCC parameter set (VPCC_VPS); if it is 1, it indicates that it is atlas data (VPCC_AD); if it is 2, it indicates that it is occupied video data (VPCC_OVD); if it is 3, it indicates that it is geometric video data (VPCC_GVD); if it is 4, it indicates that it is texture video data (VPCC_AVD).
[0791] Since the meaning, procedure, deletion, addition, etc. of the values assigned to the vuh_unit_type field can be easily changed by those skilled in the art, the present invention is not limited to the above embodiments.
[0792] For the V-PCC unit header according to the embodiment, when the vuh_unit_type field indicates characteristic video data (VPCC_AVD) or geometric video data (VPCC_GVD) or occupancy video data (VPCC_OVD) or atlas data (VPCC_AD), it may further include a vuh_vpcc_parameter_set_id field and a vuh_atlas_id field.
[0793] The vuh_vpcc_parameter_set_id field specifies the identifier of the active V-PCC parameter set (VPCC VPS), that is, vuh_vpcc_parameter_set_id.
[0794] The vuh_atlas_id field specifies the index of the atlas corresponding to the current V-PCC unit.
[0795] For the V-PCC unit header according to the embodiment, when the vuh_unit_type field indicates characteristic video data (VPCC_AVD), it may further include a vuh_attribute_index field, a vuh_attribute_dimension_index field, a vuh_map_index field, and a vuh_raw_video_flag field.
[0796] The vuh_attribute_index field indicates the index of the characteristic video data carried in the characteristic video data unit.
[0797] The vuh_attribute_dimension_index field indicates the index of the characteristic dimension group carried in the characteristic video data unit.
[0798] The vuh_map_index field, if it exists, indicates the index of the current trait stream.
[0799] The vuh_raw_video_flag field can indicate the presence or absence of RAW-coded points. For example, if the value of the vuh_raw_video_flag field is 1, it can indicate that the related trait video data unit contains only RAW-coded points. As another example, if the value of the vuh_raw_video_flag field is 0, it can indicate that the related trait video data unit contains RAW-coded points. Note that if the vuh_raw_video_flag field does not exist, the value of that field can be inferred to be 0. According to an embodiment, RAW-coded points are also called Pulse Code Modulation (PCM)-coded points.
[0800] The V-PCC unit header according to an embodiment may further include a vuh_map_index field, a vuh_raw_video_flag field, and a vuh_reserved_zero_12bits field when the vuh_unit_type field indicates geometric video data (VPCC_GVD).
[0801] The vuh_map_index field, if it exists, indicates the index of the current geometry stream.
[0802] The vuh_raw_video_flag field can indicate the presence or absence of RAW-coded points. For example, if the value of the vuh_raw_video_flag field is 1, the related geometry video data unit can indicate that it contains only RAW-coded points. As another example, if the value of the vuh_raw_video_flag field is 0, the related geometry video data unit can indicate that it contains RAW-coded points. If the vuh_raw_video_flag field does not exist, it can be inferred that the value of the field is 0. According to an embodiment, RAW-coded points are also referred to as PCM (Pulse Code Modulation)-coded points.
[0803] The vuh_reserved_zero_12bits field is a reserved field for future use.
[0804] According to an embodiment, the V-PCC unit header can further include the vuh_reserved_zero_17bits field when the vuh_unit_type field indicates occupied video data (VPCC_OVD) or atlas data (VPCC_AD), and can further include the vuh_reserved_zero_27bits field otherwise.
[0805] The vuh_reserved_zero_17bits field and the vuh_reserved_zero_27bits field are reserved fields for future use.
[0806] FIG. 50 shows an example of the syntax structure of a V-PCC unit payload (vpcc_unit_payload()) according to an embodiment.
[0807] The V-PCC unit payload in Figure 50 may include one of a V-PCC parameter set (vpcc_parameter_set()), an atlas sub-bitstream (atlas_sub_bitstream()), and a video sub-bitstream (video_sub_bitstream()) according to the value of the vuh_unit_type field in the corresponding V-PCC unit header.
[0808] For example, when the vuh_unit_type field indicates a V-PCC parameter set (VPCC_VPS), the V-PCC unit payload includes a V-PCC parameter set (vpcc_parameter_set()) that contains the overall encoding information of the bitstream. When it indicates atlas data (VPCC_AD), it includes an atlas sub-bitstream (atlas_sub_bitstream()) that transmits the atlas data. Also, in one embodiment, when the vuh_unit_type field indicates occupied video data (VPCC_OVD), the V-PCC unit payload includes an occupied video sub-bitstream (video_sub_bitstream()) that transmits the occupied video data. When it indicates geometric video data (VPCC_GVD), it includes a geometric video sub-bitstream (video_sub_bitstream()) that transmits the geometric video data. When it indicates texture video data (VPCC_AVD), it includes a texture video sub-bitstream (video_sub_bitstream()) that transmits the texture video data.
[0809] According to an embodiment, the atlas sub-bitstream is also referred to as an atlas sub-stream, the occupied video sub-bitstream is also referred to as an occupied video sub-stream, the geometry video sub-bitstream is also referred to as a geometry video sub-stream, and the texture video sub-bitstream is also referred to as a texture video sub-stream. The V-PCC unit payload according to the embodiment follows the format of a HEVC (High Efficiency Video Coding) NAL (Network Abstraction Layer) unit.
[0810] FIG. 51 shows an example of an atlas sub-stream structure according to an embodiment. In one embodiment, the atlas sub-stream in FIG. 51 follows the format of a HEVC NAL unit.
[0811] The atlas sub-stream according to the embodiment consists of a sample stream NAL unit including an atlas sequence parameter set (ASPS), a sample stream NAL unit including an atlas frame parameter set (AFPS), one or more sample stream NAL units including one or more atlas style group information, and / or one or more sample stream NAL units including one or more SEI messages.
[0812] One or more SEI messages according to the embodiment can include a prefix SEI message and a suffix SEI message.
[0813] The atlas sub-stream according to the embodiment can further include a sample stream NAL header before one or more NAL units.
[0814] FIG. 52 shows an example of the syntax structure of a sample stream NAL header (sample_stream_nal_header()) included in the atlas sub-stream according to the embodiment.
[0815] The sample stream NAL header () according to the embodiment can include the ssnh_unit_size_precision_bytes_minus1 field and the ssnh_reserved_zero_5bits field.
[0816] The ssnh_unit_size_precision_bytes_minus1 field can indicate, in bytes, the precision of the ssnu_vpcc_unit_size element within all sample stream NAL units by adding 1 to this field value. The value of this field is within the range of 0 to 7.
[0817] The ssnh_reserved_zero_5bits field is a reserved field for future use.
[0818] Figure 53 shows an example of the syntax structure of the sample stream NAL unit (sample_stream_nal_unit ()) according to the embodiment.
[0819] The sample stream NAL unit (sample_stream_nal_unit ()) according to the embodiment can include the ssnu_nal_unit_size field and the nal_unit (ssnu_nal_unit_size).
[0820] The ssnu_nal_unit_size field specifies the size of the subsequent NAL unit in bytes. The number of bits used to indicate the ssnu_nal_unit_size field seems to be (ssnh_unit_size_precision_bytes_minus1 + 1) * 8.
[0821] The nal_unit(ssnu_nal_unit_size) has a length corresponding to the value of the ssnu_nal_unit_size field and transmits one of an atlas sequence parameter set (ASPS), an atlas frame parameter set (AFPS), atlas style group information, and an SEI message. That is, each sample stream NAL unit can include an atlas sequence parameter set (ASPS), an atlas frame parameter set (AFPS), atlas style group information, and an SEI message. In an embodiment, the atlas sequence parameter set (ASPS), the atlas frame parameter set (AFPS), the atlas style group information, and the SEI message are referred to as atlas data (or metadata for the atlas).
[0822] The SEI message according to the embodiment can assist processes related to decoding, reconstruction, display, or other purposes.
[0823] Each SEI message according to the embodiment consists of an SEI message header and an SEI message payload (sei_payload). The SEI message header may include payload type information (payloadType) and payload size information (payloadSize).
[0824] The payload type information (payloadType) indicates the payload type of the corresponding SEI message. For example, based on the payload type information (payloadType), it is possible to identify whether it is a prefix SEI message or a suffix SEI message.
[0825] The payload size information (payloadSize) indicates the payload size of the corresponding SEI message.
[0826] FIG. 54 shows an example of the syntax structure of the SEI message payload (sei_payload()) according to the embodiment.
[0827] The SEI message according to the embodiment may include a prefix SEI message or a suffix SEI message. Also, each SEI message payload signals information corresponding to the payload type information (payloadType) and the payload size information (payloadSize) by the SEI message payload (sei_payload(payloadType, payloadSize)).
[0828] The prefix SEI message according to the embodiment may include buffering_period(payloadSize) if the payload type information is 0, pic_timing(payloadSize) if it is 1, filler_payload(payloadSize) if it is 2, sei_prefix_indication(payloadSize) if it is 10, and 3D_region_mapping(payloadSize) if it is 13.
[0829] The suffix SEI message according to the embodiment may include filler_payload(payloadSize) if the payload type information is 2, user_data_registered_itu_t_t35(payloadSize) if it is 3, user_data_unregistered(payloadSize) if it is 4, and decoded_pcc_hash(payloadSize) if it is 11.
[0830] Note that as an embodiment, a V-PCC bitstream having the structure shown in FIG. 24 or FIG. 43 is encapsulated in the ISO BMFF file format in the file / segment encapsulation unit of FIG. 1, FIG. 18, FIG. 20, or FIG. 21.
[0831] At this time, the V-PCC stream may be transmitted via multiple tracks of a file or via a single track.
[0832] The ISOBMFF-based file according to the embodiment can be composed of boxes and / or information also called ftyp, meta, moov, mdat, etc.
[0833] The ftyp box (file type box) can provide information related to the file type or file compatibility for the corresponding file. On the receiving side, the corresponding file can be classified by referring to the ftyp box.
[0834] The meta box may include a vpcg{0, 1, 2, 3} box (V-PCC Group Box).
[0835] The mdat box, also called the media data box, may include a video-encoded geometry bitstream, a video-encoded texture bitstream, a video-encoded occupancy map bitstream, and / or an atlas data bitstream.
[0836] The moov box, also called the movie box, may include metadata for the media data (e.g., geometry bitstream, texture bitstream, occupancy map bitstream, etc.) of the corresponding file. For example, it may include information necessary for decoding and playing the corresponding media data, and may also include information about the samples of the corresponding file. The moov box can function as a container for all metadata. The moov box may be the top-layer box among the metadata-related boxes. According to the embodiment, there may be only one moov box in the file.
[0837] The box according to the embodiment includes a track (trak) box that provides information related to the track of the corresponding file, and the track (trak) box may include a media (mdia) box that provides media information of the corresponding track and a track reference container (tref) box for connecting (referencing) the samples of the corresponding track and the file corresponding to the corresponding track.
[0838] The media (mdia) box includes a media information container (minf) box that provides information on the corresponding media data, and the media information container (minf) box may include a sample table (stbl) box that provides metadata related to the samples of the mdat box.
[0839] The stbl box may include a sample description (stsd) box that provides information on the coding type used and the initialization information required for the corresponding coding type.
[0840] The stsd box may include a sample entry for a track that stores the V-PCC bitstream according to the embodiment.
[0841] According to the embodiment, in order to store the V-PCC bitstream in a single track or multiple tracks in a file, the Volumetric visual track, Volumetric visual media header, Volumetric sample entry, Volumetric samples, samples of the V-PCC track, and sample entry are defined as follows.
[0842] As used herein, the term V-PCC is similar to Visual Volumetric Video-based Coding (V3C) and can be referred to complementarily with each other.
[0843] According to an embodiment, V-PCC represents a volumetric encoding of point cloud visual information (video-based point cloud compression represents a volumetric encoding of point cloud visual information).
[0844] That is, the minf box within the track box of the moov box may further include a volumetric visual media header box. This volumetric visual media header box contains information regarding a volumetric visual track that includes a volumetric visual scene.
[0845] Each volumetric visual scene may be represented by a unique volumetric visual track. An ISOBMFF file may contain multiple scenes and therefore multiple volumetric visual tracks may be present in the ISOBMFF file (Each volumetric visual scene is represented by a unique volumetric visual track. An ISOBMFF file may contain multiple scenes and therefore multiple volumetric visual tracks may be present in the ISOBMFF file).
[0846] According to an embodiment, a volumetric visual track can be identified by the volumetric visual media handler type ’volv’ in the HandlerBox of the MediaBox (A volumetric visual track is identified by the volumetric visual media handler type ’volv’ in the HandlerBox of the MediaBox).
[0847] The syntax of the volumetric visual media header box according to an embodiment is as follows.
[0848] Box Type: 'vvhd'
[0849] Container: MediaInformationBox
[0850] Mandatory: Yes
[0851] Quantity: Exactly one
[0852] According to an implementation, a volumetric visual track can use a volumetric visual media header box (VolumetricVisualMediaHeaderBox) within a media information box (MediaInformationBox) as follows.
[0853] aligned(8) class VolumetricVisualMediaHeaderBox
[0854] extends FullBox('vvhd', version = 0, 1) {
[0855] }
[0856] The above "version" may be an integer indicating the version of this box.
[0857] The volume visual track according to the embodiment can use a volume visual sample entry as follows.
[0858] class VolumetricVisualSampleEntry(codingname)
[0859] extends SampleEntry (codingname) {
[0860] unsigned int(8)
[32] compressor_name;
[0861] }
[0862] compressor_name is a name, for informative purposes. It is formatted in a fixed 32 - byte field, with the first byte set to the number of bytes to be displayed, followed by that number of bytes of displayable data encoded using UTF - 8, and then padding to complete 32 bytes total (including the size byte). This field can be set to 0.
[0863] The format of the volumetric visual sample according to the embodiment can be defined by a coding system.
[0864] According to the embodiment, the V-PCC unit header box can be present in all video-coded V-PCC component tracks included in the V-PCC track and the scheme information included in the sample entry. The V-PCC unit header box can include a V-PCC unit header for the data transmitted by each track as follows.
[0865] aligned(8) class VPCCUnitHeaderBox
[0866] extends FullBox('vunt', version = 0, 0) {
[0867] vpcc_unit_header() unit_header;
[0868] }
[0869] That is, the VPCC unit header box may include vpcc_unit_header(). FIGS. 34 and 48 show examples of the syntax structure of the V-PCC unit header (vpcc_unit_header()).
[0870] According to the embodiment, the V-PCC track sample entry may include a VPCCConfigurationBox.
[0871] According to an embodiment, a VPCC Configuration Box may include a VPCC Decoder Configuration Record as follows.
[0872] aligned(8) class VPCCDecoderConfigurationRecord {
[0873] unsigned int(8) configurationVersion = 1;
[0874] unsigned int(3) sampleStreamSizeMinusOne;
[0875] unsigned int(5) numOfVPCCParameterSets;
[0876] for (i=0; i< numOfVPCCParameterSets; i++) {
[0877] sample_stream_vpcc_unit VPCCParameterSet;
[0878] }
[0879] unsigned int(8) numOfAtlasSetupUnits;
[0880] for (i=0; i< numOfAtlasSetupUnits; i++) {
[0881] sample_stream_vpcc_unit atlas_setupUnit;
[0882] }
[0883] }
[0884] The configurationVersion contained in the VPCC Decoder Configuration Record (VPCCDecoderConfigurationRecord) indicates the version field. Incompatible changes to the record are indicated by a change in the version number.
[0885] Adding 1 to the value of sampleStreamSizeMinusOne indicates the precision, in bytes, of the ssvu_vpcc_unit_size element in all sample stream V-PCC units in either this configuration record or a V-PCC sample in the stream to which this configuration record applies.
[0886] numOfVPCCParameterSets indicates the number of VPS (V-PCC parameter sets) signaled in the VPCC Decoder Configuration Record.
[0887] The VPCCParameterSet is an instance of the sample_stream_vpcc_unit() for the V-PCC unit of the VPCC_VPS type. The V-PCC unit may include vpcc_parameter_set(). That is, the VPCCParameterSet array may include vpcc_parameter_set(). Figure 46 shows an example of the syntax structure of the sample_stream_vpcc_unit().
[0888] numOfAtlasSetupUnits indicates the number of setup arrays for the atlas stream signaled in the VPCCDecoderConfigurationRecord.
[0889] An Atlas_setupUnit is an instance of sample_stream_vpcc_unit() that includes an atlas sequence parameter set, an atlas frame parameter set, or an SEI atlas NAL unit. Figure 46 shows an example of the syntax structure of the sample_stream_vpcc_unit().
[0890] That is, the atlas setup unit (atlas_setupUnit) array may include certain atlas parameter sets for the stream referenced by the sample entry where the VPCCDecoderConfigurationRecord exists and the atlas stream SEI message. According to an embodiment, the atlas setup unit may also be simply referred to as a setup unit.
[0891] According to other embodiments, the VPCC decoder configuration record (VPCCDecoderConfigurationRecord) can be shown as follows.
[0892] aligned(8) class VPCCDecoderConfigurationRecord {
[0893] unsigned int(8) configurationVersion = 1;
[0894] unsigned int(3) sampleStreamSizeMinusOne;
[0895] bit(2) reserved = 1;
[0896] unsigned int(3) lengthSizeMinusOne;
[0897] unsigned int(5) numOVPCCParameterSets;
[0898] for (i=0; i< numOVPCCParameterSets; i++) {
[0899] sample_stream_vpcc_unit VPCCParameterSet;
[0900] }
[0901] unsigned int(8) numOfSetupUnitArrays;
[0902] for (j=0; j<numOfSetupUnitArrays; j++) {
[0903] bit(1) array_completeness;
[0904] bit(1) reserved = 0;
[0905] unsigned int(6) NAL_unit_type;
[0906] unsigned int(8) numNALUnits;
[0907] for (i = 0; i < numNALUnits; i++) {
[0908] sample_stream_nal_unit setupUnit;
[0909] }
[0910] }
[0911] The configurationVersion is the version field. Incompatible changes to the record are indicated by a change in the version number.
[0912] Adding 1 to the value of lengthSizeMinusOne can indicate, in bytes, the precision of ssnu_nal_unit_size within all sample stream NAL units of the V-PCC samples included in the VPCCDecoderConfigurationRecord or the stream to which the VPCCDecoderConfigurationRecord applies. Figure 53 shows an example of the syntax structure of a sample stream NAL unit (sample_stream_nal_unit()) that includes the ssnu_nal_unit_size field.
[0913] numOfVPCCParameterSets indicates the number of VPSs (V-PCC parameter sets) signaled in the VPCCDecoderConfigurationRecord.
[0914] The VPCCParameterSet is an instance of the sample_stream_vpcc_unit() for a V-PCC unit of type VPCC_VPS. The V-PCC unit may contain vpcc_parameter_set(). That is, the VPCCParameterSet array may contain vpcc_parameter_set(). Figure 46 shows an example of the syntax structure of the sample_stream_vpcc_unit().
[0915] numOfSetupUnitArrays indicates the number of arrays of atlas NAL units of the indicated type.
[0916] The loop that is repeated only the value of numOfSetupUnitArrays may contain array_completeness.
[0917] If the value of array_completeness is 1, it indicates that all atlas NAL units of the given type are in the following array and none are in the stream; if the value of array_completeness is 0, it indicates that additional atlas NAL units of the indicated type may be in the stream (when equal to 1 indicates that all atlas NAL units of the given type are in the following array and none are in the stream; when equal to 0 indicates that additional atlas NAL units of the indicated type may be in the stream). The default and permitted values are constrained by the sample entry name (the default and permitted values are constrained by the sample entry name).
[0918] The NAL_unit_type indicates the type of the atlas NAL unit contained in the following array. The NAL_unit_type is restricted to take one of the values indicating a NAL_ASPS, NAL_PREFIX_SEI, or NAL_SUFFIX_SEI atlas NAL unit.
[0919] numNALUnits indicates the number of atlas NAL units of the indicated type for the stream to which the VPCCDecoderConfigurationRecord applies. That is, it provides information about the entire stream. The SEI array shall only contain SEI messages of a ‘declarative’ nature, that is, those that provide information about the stream as a whole. An example of such an SEI is the user-data SEI.
[0920] The setup unit is an instance of a sample_stream_nal_unit() that contains an atlas sequence parameter set, or an atlas frame parameter set or a declarative SEI atlas NAL unit.
[0921] According to an embodiment, the general layout of a multi-track container (also referred to as a multi-track ISOBMFF V-PCC container) of a V-PCC bitstream may be mapped to individual tracks within the container file based on their types (The general layout of a multi-track ISOBMFF V-PCC container, where V-PCC units in a V-PCC elementary stream are mapped to individual tracks within the container file based on their types). According to an embodiment, a multi-track ISOBMFF V-PCC container has two types of tracks. One of them is a V-PCC track, and the other is a V-PCC component track.
[0922] The V-PCC track according to an embodiment is a track that transmits volumetric visual information in a V-PCC bitstream including an atlas sub-bitstream and a sequence parameter set.
[0923] The V-PCC component track according to an embodiment is a restricted video scheme track that transmits 2D video-encoded data for an occupancy map, geometry, and texture sub-bitstream of a V-PCC bitstream. In addition to this, the V-PCC component track can satisfy the following conditions.
[0924] a) In the sample entry, a new box explaining the role of the video stream included in this track is inserted into the V-PCC system.
[0925] b) The track reference is introduced from the V-PCC track to the V-PCC component track. This is to establish the membership of the V-PCC component track included in the specific point-cloud represented by the V-PCC track.
[0926] c) The track-header flag is set to 0. This is to indicate that although the track contributes to the V-PCC system, it does not directly contribute to the overall layout of the movie.
[0927] Tracks belonging to the same V-PCC sequence are time-aligned. Samples that contribute to the same point cloud frame across the different video-encoded V-PCC component tracks and the V-PCC track have the same presentation time. The V-PCC atlas sequence parameter sets and atlas frame parameter sets used for such samples have a decoding time equal or prior to the composition time of the point cloud frame. In addition, all tracks belonging to the same V-PCC sequence have the same implied or explicit edit lists.
[0928] Note: The synchronization between elementary streams within a component track may be handled by the ISOBMFF track timing structures (stts, ctts, and cslg), or by an equivalent mechanism within a movie fragment.
[0929] Based on such a layout, the V-PCC ISOBMFF container may include the following.
[0930] - A V-PCC track containing V-PCC parameter sets included in samples and sample entries that transmit the payloads of V-PCC parameter set V-PCC units (unit type VPCC_VPS) and atlas V-PCC units (unit type VPCC_AD). Also, the track includes track references to other tracks that transmit the payloads of video-compressed V-PCC units such as unit types VPCC_OVD, VPCC_GVD, and VPCC_AVD.
[0931] - A restricted video scheme track with samples that include access units of a video-coded elementary stream for occupancy map data that is the payload of a V-PCC unit of type VPCC_OVD.
[0932] - One or more restricted video scheme tracks with samples that include access units of a video-coded elementary stream for geometry data that is the payload of a V-PCC unit of type VPCC_GVD.
[0933] - Zero or more restricted video scheme tracks with samples that include access units of a video-coded elementary stream for texture data that is the payload of a V-PCC unit of type VPCC_AVD.
[0934] The following describes V-PCC tracks.
[0935] According to an embodiment, the syntax structure of a V-PCC Track Sample Entry is as follows.
[0936] Sample Entry Type: 'vpc1', 'vpcg'
[0937] Container: SampleDescriptionBox
[0938] Mandatory: 'vpc1' or 'vpcg' sample entries are mandatory
[0939] Quantity: One or more sample entries may exist
[0940] The V-PCC track uses a VPCCSampleEntry that extends the volume visual sample entry. The sample entry type is 'vpc1' or 'vpcg'.
[0941] The V-PCC sample entry includes a V-PCC configuration box (VPCCConfigurationBox). This box contains a V-PCC decoder configuration record (VPCCDecoderConfigurationRecord).
[0942] Under the 'vpc1' sample entry, all atlas sequence parameter sets, atlas frame parameter sets, or V-PCC SEIs are within the setupUnit array.
[0943] Under the 'vpcg' sample entry, the atlas sequence parameter set, atlas frame parameter set, V-PCC SEI may be within this array or within the stream.
[0944] The optional BitRateBox may be present within the V-PCC volume sample entry to signal the bitrate information of the V-PCC track.
[0945] Volumetric Sequences:
[0946] class VPCCConfigurationBox extends Box('vpcC') {
[0947] VPCCDecoderConfigurationRecord() VPCCConfig;
[0948] }
[0949] aligned(8) class VPCCSampleEntry() extends VolumetricVisualSampleEntry ('vpc1') {
[0950] VPCCConfigurationBox config;
[0951] VPCCUnitHeaderBox unit_header;
[0952] }
[0953] FIG. 55 shows an example of a V-PCC sample entry structure according to an embodiment. In FIG. 55, the V-PCC sample entry includes one V-PCC parameter set and may include an optional atlas sequence parameter set (ASPS), an atlas frame parameter set (AFPS), or SEI.
[0954] The V-PCC bitstream according to the embodiment may further include a sample stream V-PCC header, a sample stream NAL header, and a V-PCC unit header box.
[0955] Hereinafter, the V-PCC track sample format will be described.
[0956] Each sample in the V-PCC track corresponds to a single-point cloud frame. Samples corresponding to this frame in the various component tracks have the same composition time as the V-PCC track samples. Each V-PCC sample contains one or more atlas NAL units as follows.
[0957] aligned(8) class VPCCSample {
[0958] unsigned int PointCloudPictureLength = sample_size; / / size of samble (e.g., from SampleSizeBox)
[0959] for (i=0; i<PointCloudPictureLength; ) {
[0960] sample_stream_nal_unit nalUnit
[0961] i += (VPCCDecoderConfigurationRecord.lengthSizeMinusOne+1) + nalUnit.ssnu_nal_unit_size;
[0962] }
[0963] }
[0964] aligned(8) class VPCCSample
[0965] {
[0966] unsigned int PictureLength = sample_size; / / size of samble (e.g., from SampleSizeBox)
[0967] for (i = 0; i < PictureLength; ) / / Signaled until the end of the picture
[0968] {
[0969] unsigned int((VPCCDecoderConfigurationRecord.LengthSizeMinusOne + 1) * 8)
[0970] NALUnitLength;
[0971] bit(NALUnitLength * 8) NALUnit;
[0972] i += (VPCCDecoderConfigurationRecord.LengthSizeMinusOne + 1) + NALUnitLength;
[0973] }
[0974] }
[0975] The sync sample (random access point) in the V-PCC track according to the embodiment is a V-PCC IRAP-coded patch data access unit. The atlas parameter set can be repeated at the sync sample, if necessary, to allow random access.
[0976] The following describes the video-encoded V-PCC component track.
[0977] The transmission of a video track encoded using an MPEG-specific codec follows the ISO BMFF standard. For example, the transmission of AVC- and HEVC-encoded video can refer to ISO / IEC 14496-15. ISO BMFF can further provide an extension mechanism if other codec types are required.
[0978] Since it makes no sense to display frames decoded from traits, geometry, or occupancy map tracks without the player side reconstructing the point cloud, restricted video scheme types can be defined for such video-encoded tracks.
[0979] The following describes a restricted video scheme.
[0980] The V-PCC component video track is represented in the file as restricted video. It can also be identified by the 'pccv' value in the scheme_type field of the SchemeTypeBox of the RestrictedSchemeInfoBox of the restricted video sample entry.
[0981] There are no restrictions on the video codec used to encode traits, geometry, and occupancy map V-PCC components. Furthermore, the components can be encoded using different video codecs.
[0982] According to an embodiment, scheme information (SchemeInformationBox) may exist and may include a VPCCUnitHeaderBox.
[0983] The following describes referencing V-PCC component tracks.
[0984] To link a V-PCC track to a component video track, three TrackReferenceTypeBoxes may be added to the track reference boxes within the track box of the V-PCC track for each component. The track reference type box contains an array of track_IDs that store the video tracks related to the V-PCC track reference. The reference_type of the TrackReferenceTypeBox identifies the type of component such as occupancy map, geometry, texture, or occupancy map. The track reference types are as follows:
[0985] In 'pcco', the referenced track contains a video-encoded occupancy map V-PCC component.
[0986] In 'pccg', the referenced track contains a video-encoded geometry V-PCC component.
[0987] In 'pcca', the referenced track contains a video-encoded texture V-PCC component.
[0988] The type of V-PCC component transmitted by the referenced restricted video track and signaled within the RestrictedSchemeInfoBox of the track matches the reference type of the track reference from the V-PCC track.
[0989] Next, the single track container of the V-PCC bitstream will be described.
[0990] A single-track encapsulation of V-PCC data requires the V-PCC encoded elementary bitstream to be represented by a single-track declaration.
[0991] Single-track encapsulation of PCC data is used for encapsulation of the simple ISOBMFF of the V-PCC encoded bitstream. Such a bitstream is immediately stored in a single track without additional processing. The V-PCC unit header data structure may be within the bitstream. The single-track container for V-PCC data is provided to the media workflow for additional processing (e.g., multi-track file generation, transcoding, DASH segmentation, etc.).
[0992] The ISOBMFF file containing the single-track encapsulated V-PCC data may contain 'pcst' in the compatible_brands[] list of the file type box.
[0993] V-PCC elementary stream track:
[0994] Sample Entry Type: 'vpe1', 'vpeg'
[0995] Container: SampleDescriptionBox
[0996] Mandatory: A 'vpe1' or 'vpeg' sample entry is mandatory
[0997] Quantity: One or more sample entries may be present
[0998] A V-PCC elementary stream track uses volume visual sample entries with sample entry type 'vpe1' or 'vpeg'.
[0999] The V-PCC elementary stream sample entry includes a VPCCConfigurationBox.
[1000] Under the 'vpe1' sample entry, all atlas sequence parameter sets, atlas frame parameter sets, and SEIs may be present in the setup unit array. Under the 'vpeg' sample entry, atlas sequence parameter sets, atlas frame parameter sets, and SEIs may be present in this array or stream.
[1001] Volumetric Sequences:
[1002] class VPCCConfigurationBox extends Box('vpcC') {
[1003] VPCCDecoderConfigurationRecord() VPCCConfig;
[1004] }
[1005] aligned(8) class VPCElementaryStreamSampleEntry() extends VolumetricVisualSampleEntry ('vpe1') {
[1006] VPCCConfigurationBox config;
[1007] VPCC Bounding Information Box 3D_bb;
[1008] }
[1009] The following describes the V-PCC elementary stream sample format.
[1010] A V-PCC elementary stream sample may be composed of one or more V-PCC units belonging to the same presentation time. Each sample has a unique presentation time, size, and duration. Samples are, for example, sink samples or are decoded depending on other V-PCC elementary stream samples.
[1011] The following describes the V-PCC elementary stream sync sample.
[1012] A V-PCC elementary stream sync sample satisfies the following conditions:
[1013] - It is independently decodable.
[1014] - Samples coming after the sync sample in decoding order have no decoding dependency on samples before the sync sample.
[1015] - All samples coming after the sync sample in decoding order are successfully decodable.
[1016] The following describes the V-PCC elementary stream sub-sample.
[1017] The V-PCC elementary stream sub-sample is a V-PCC unit included in the V-PCC elementary stream sample.
[1018] The V-PCC elementary stream track includes SubSampleInformationBoxes within the TrackFragmentBox of MovieFragmentBoxes that arrange V-PCC elementary stream sub-samples or within each SampleTableBox.
[1019] The 32-bit unit header of the V-PCC unit representing the sub-sample is copied to the 32-bit codec_specific_parameters field of the sub-sample entry within the SubSampleInformationBox. The V-PCC unit type of each sub-sample can be identified by parsing the codec_specific_parameters field of the sub-sample entry within the SubSampleInformationBox.
[1020] The rendering of point cloud data will be described below.
[1021] According to an embodiment, the rendering of point cloud data is performed by the renderer 10009 in FIG. 1, the point cloud renderer 19007 in FIG. 19, the renderer 20009 in FIG. 20, or the point cloud rendering section 22004 in FIG. 22. According to an embodiment, the point cloud data can be rendered in 3D space based on metadata. The user can view all or a part of the region of the result rendered by a VR / AR display or a general display, etc. In particular, the point cloud data is rendered by the user's viewport, etc.
[1022] Here, the viewport and the viewport area mean the area that the user is viewing in the point cloud video. The viewpoint is the point where the user is viewing in the point cloud video and means the center of the viewport area. That is, the viewport is an area centered on the viewpoint, and the size, shape, etc. of the area it occupies can be determined by the FOV (Field Of View).
[1023] The viewport information according to the embodiment is currently information regarding the area that the user is viewing on the 3D space by means of a device or an HMD, etc. Thereby, gaze analysis is performed, and it is also possible to confirm how the user consumes the point cloud video and how long the user gazes at which area of the point cloud video. The gaze analysis may be performed on the receiving side and transmitted to the transmitting side via a feedback channel. A display device such as VR / AR / MR can extract the viewport area based on orientation information, the position / direction of the user's value, the vertical or horizontal FOV supported by the device, etc. The orientation or viewport information is extracted or calculated by the receiving device. The orientation or viewport information analyzed by the receiving device may be transmitted to the transmitting device via a feedback channel.
[1024] The point cloud video decoder of the receiving device according to the embodiment efficiently extracts or decodes only the media data of a specific area, that is, the area indicated by the orientation information and / or the viewport information, using the orientation information and / or the viewport information indicating the area that the user is currently viewing. The point cloud video encoder of the transmitting device according to the embodiment may encode only the media data of a specific area, that is, the area indicated by the orientation information and / or the viewport information, using the orientation information and / or the viewport information feedbacked by the receiving device, and may encapsulate the encoded media data in a file and transmit it.
[1025] According to an embodiment, the file / segment encapsulation unit of the transmitting device may encapsulate all point cloud data into files / segments based on orientation information and / or viewport information, or may encapsulate the point cloud data indicated by the orientation information and / or viewport information into files / segments.
[1026] According to an embodiment, the file / segment decapsulation unit of the receiving device may decapsulate a file including all point cloud data based on orientation information and / or viewport information, or may decapsulate a file including the point cloud data indicated by the orientation information and / or viewport information.
[1027] According to an embodiment, viewport information is used in a similar or analogous sense to ViewInfoStruct information and view information.
[1028] The viewport information according to the embodiment is also referred to as information regarding the viewport or metadata related to the viewport information. The information regarding the viewport according to the embodiment may include at least one of viewport information, rendering parameters, and object rendering information.
[1029] According to an embodiment, the information regarding the viewport (viewport related information) is generated / encoded by the metadata encoding unit 18005 of the point cloud data transmitting device in FIG. 18, or the point cloud preprocessing unit 20001 and / or the video / image encoding units 21007, 21008 of the V-PCC system in FIG. 21, and can be acquired / decoded by the metadata decoding unit 19002 of the point cloud data receiving device in FIG. 19, or the video / image decoding units 22001, 22002 and / or the point cloud postprocessing unit 22003 of the V-PCC system in FIG. 22.
[1030] This specification defines metadata related to viewport (or view) information of point cloud data, and describes examples of storing and signaling metadata related to the corresponding viewport information in a file.
[1031] This specification describes an example of storing metadata related to viewport information of point cloud data that dynamically changes over time within a file.
[1032] This specification defines metadata related to rendering parameters of point cloud data, and describes examples of storing and signaling metadata related to the rendering parameters in a file.
[1033] This specification describes an example of storing metadata related to rendering parameters of point cloud data that dynamically changes over time within a file.
[1034] This specification defines metadata related to object rendering information of point cloud data within a file, and describes examples of storing and signaling metadata related to the object rendering information in the file.
[1035] This specification describes an example of storing metadata related to object rendering information of point cloud data that dynamically changes over time within a file.
[1036] FIG. 56 shows an example of generating a view using a virtual camera and ViewInfoStruct information according to an embodiment.
[1037] The ViewInfoStruct information according to the embodiment may include detailed information for generating a view to be rendered and provided. In particular, it includes the 3D position information of the virtual camera for generating the view, the vertical / horizontal FOV (field of view) of the virtual camera, the direction vector of the direction the virtual camera looks, the up vector information indicating the upward direction of the virtual camera, etc. The virtual camera is similar to the user's eye, i.e., the user's vision that looks at a part of the area on the 3D. Based on this information, a view frustum can be analogized. The view frustum means an area in 3D space that includes all or part of the point cloud data actually rendered and displayed. According to the embodiment, by projecting the analogized view frustum in the form of a 2D frame, a view (i.e., the 2D image / video frame actually displayed) is generated.
[1038] The following is the syntax showing an example of the information included in the ViewInfoStruct information.
[1039] aligned(8) class ViewInfoStruct(){
[1040] unsigned int(16) view_pos_x;
[1041] unsigned int(16) view_pos_y;
[1042] unsigned int(16) view_pos_z;
[1043] unsigned int(8) view_vfov;
[1044] unsigned int(8) view_hfov;
[1045] unsigned int(16) view_dir_x;
[1046] unsigned int(16) view_dir_y;
[1047] unsigned int(16) view_dir_z;
[1048] unsigned int(16) view_up_x;
[1049] unsigned int(16) view_up_y;
[1050] unsigned int(16) view_up_z;
[1051] }
[1052] view_pos_x, view_pos_y, and view_pos_z indicate the x, y, and z coordinate values in the 3D space of a virtual camera that can generate a view (e.g., a 2D image / video frame actually displayed).
[1053] view_vfov and view_hfov indicate the vertical field of view (FOV) and horizontal FOV information of a virtual camera that can generate a view.
[1054] view_dir_x, view_dir_y, and view_dir_z indicate the x, y, and z coordinate values in the 3D space for a direction vector indicating the direction in which the virtual camera is looking.
[1055] view_up_x, view_up_y, and view_up_z indicate the x, y, and z coordinate values in the 3D space for an up vector indicating the upward direction of the virtual camera.
[1056] The ViewInfoStruct() information according to the embodiment can be transmitted in a format such as SEI on the point cloud bitstream as follows.
[1057] V-PCC view information box
[1058] aligned(8) class VPCCViewInfoBox extends FullBox('vpvi',0,0) {
[1059] ViewInfoStruct();
[1060] }
[1061] The static V-PCC viewport information will be described below.
[1062] When the viewport information according to the embodiment does not change within the point cloud sequence, the VPCCViewInfoBox is included in the sample entry of the V-PCC track or the sample entry of the V-PCC elementary stream track as follows.
[1063] aligned(8) class VPCCSampleEntry() extends VolumetricVisualSampleEntry ('vpc1') {
[1064] VPCCConfigurationBox config;
[1065] VPCCUnitHeaderBox unit_header;
[1066] VPCCViewInfoBox view_info;
[1067] }
[1068] The detailed information for the VPCCViewInfoBox to generate a view where the point cloud data related to the atlas frames stored in the samples within the track is rendered and provided is as follows.
[1069] aligned(8) class VPCCElementaryStreamSampleEntry() extends VolumetricVisualSampleEntry ('vpe1') {
[1070] VPCCConfigurationBox config;
[1071] VPCCViewInfoBox view_info;
[1072] }
[1073] The VPCCViewInfoBox includes detailed information for generating a view where the point cloud data related to the atlas frames and video frames stored in the subsamples within the track is rendered and provided.
[1074] The following describes V-PCC view information sample grouping.
[1075] According to an embodiment, the 'vpvs' grouping_type for sample grouping indicates the assignment of samples within the V-PCC track to the view information transmitted to this sample group. If there is a SampleToGroupBox with the grouping_type 'vpvs', there is an accompanying SampleGroupDescriptionBox with the same grouping type, which includes the ID of this sample group.
[1076] aligned(8) class VPCCViewInfoSampleGroupDescriptionEntry() extends SampleGroupDescriptionEntry('vpvs') {
[1077] ViewInfoStruct();
[1078] }
[1079] The following describes the dynamic V-PCC view information.
[1080] If the V-PCC track has an associated time-domain metadata track with a sample entry type 'dyvi', the view information defined for the point cloud stream is transmitted by the V-PCC track considered as dynamic. That is, the view information can change dynamically over time.
[1081] The associated time-domain metadata track includes a 'cdsc' track reference for the V-PCC that transmits the atlas stream.
[1082] If the V-PCC element stream track has an associated time-domain metadata track with a sample entry type 'dyvi', the view information defined for the point cloud stream is transmitted by the V-PCC base track considered as dynamic. That is, the view information changes dynamically over time.
[1083] The associated time-domain metadata track includes a 'cdsc' track reference for the V-PCC base sample track.
[1084] aligned(8) class DynamicViewInfoSampleEntry extends MetaDataSampleEntry('dyvi') {
[1085] VPCCViewInfoBox init_view_info;
[1086] }
[1087] init_view_info may include view information (view information()) that generates an initial view of the point cloud data.
[1088] The sample syntax for this sample entry type 'dyvi' is as follows.
[1089] aligned(8) DynamicViewInfoSample() {
[1090] VPCCViewInfoBox view_info;
[1091] }
[1092] Each sample may include view information (view information()) that changes over time.
[1093] According to an embodiment, the RenderingParamStruct, which is a rendering parameter for rendering, may include parameter information applicable during the rendering of the point cloud data. This includes, for example, how the point size, point rendering type, and how to handle overlapping points, e.g., whether to display them or not, during the rendering of the point cloud data as follows.
[1094] aligned(8) class RenderingParamStruct(){
[1095] unsigned int(16) point_size;
[1096] unsigned int(7) point_type;
[1097] unsigned int(2) duplicated_point;
[1098] }
[1099] Point_size indicates the size of the points to be rendered / displayed.
[1100] Point_type can indicate the type of the points to be rendered / displayed. For example, if the value of Point_type is 0, it indicates a cuboid; if it is 1, it indicates a circle; if it is 2, it indicates a point. 【...
Claims
1. 1. A method for encoding point cloud data, comprising the steps of: encoding the point cloud data; encapsulating the bitstream containing the encoded point cloud data in a file; the encoded point cloud data includes encoded geometry data, encoded feature data, and encoded occupancy map data; The file contains multiple tracks, each of the multiple tracks includes a sample entry and a sample; the samples in at least one of the multiple tracks include one of the encoded geometry data, the encoded attribute data, and the encoded occupancy map data; the file further comprises signaling information; the signaling information includes information regarding a viewport for the point cloud data; The information about the viewport is classified into dynamic viewport information that changes over time and static viewport information that does not change over time; a track of the multiple tracks is a timed metadata track, a sample entry type for the timed metadata track having a value to indicate that the timed metadata track is used for the dynamic viewport information; The dynamic viewport information includes at least first viewport information including initial viewport position information, or second viewport information including viewport position information that dynamically changes over time; the first viewport information is conveyed via a sample entry of the timed metadata track; the second viewport information is conveyed via samples of the timed metadata track; A method according to claim 1, wherein the static viewport information is transmitted in a track of the multiple tracks that is different from the timed metadata track.
2. The method of claim 1 , wherein the point cloud data is encoded by a video based point cloud compression (V-PCC) scheme.
3. The method of claim 1 , wherein the information about the viewport further comprises camera orientation information.
4. The method of claim 1 , wherein the information about the viewport further comprises horizontal and vertical field of view (FOV) information for generating the viewport.
5. 1. An apparatus for encoding point cloud data, comprising: an encoder for encoding the point cloud data; an encapsulation unit for encapsulating a bitstream including the encoded point cloud data into a file, the encoded point cloud data includes encoded geometry data, encoded feature data, and encoded occupancy map data; The file contains multiple tracks, each of the multiple tracks includes a sample entry and a sample; the samples in at least one of the multiple tracks include one of the encoded geometry data, the encoded attribute data, and the encoded occupancy map data; the file further comprises signaling information; the signaling information includes information regarding a viewport for the point cloud data; The information about the viewport is classified into dynamic viewport information that changes over time and static viewport information that does not change over time; a track of the multiple tracks is a timed metadata track, a sample entry type for the timed metadata track having a value to indicate that the timed metadata track is used for the dynamic viewport information; The dynamic viewport information includes at least first viewport information including initial viewport position information, or second viewport information including viewport position information that dynamically changes over time; the first viewport information is conveyed via a sample entry of the timed metadata track; the second viewport information is conveyed via samples of the timed metadata track; An apparatus, wherein the static viewport information is transmitted in a track of the multiple tracks that is different from the timed metadata track.
6. The apparatus of claim 5 , wherein the point cloud data is encoded by a video based point cloud compression (V-PCC) scheme.
7. The apparatus of claim 5 , wherein the information about the viewport further comprises camera orientation information.
8. The apparatus of claim 5 , wherein the information about the viewport further comprises horizontal field of view (FOV) information and vertical FOV information for generating the viewport.
9. 1. A method for decoding point cloud data, comprising the steps of: Deencapsulating the file into a bitstream containing the point cloud data, the point cloud data includes geometry data, feature data, and occupancy map data; The file contains multiple tracks, each of the multiple tracks includes a sample entry and a sample; the samples in at least one of the multiple tracks include one of the geometry data, the attribute data, and the occupancy map data; the file further comprises signaling information; the signaling information including information regarding a viewport for the point cloud data; and decoding the point cloud data. The information about the viewport is classified into dynamic viewport information that changes over time and static viewport information that does not change over time; a track of the multiple tracks is a timed metadata track, a sample entry type for the timed metadata track having a value to indicate that the timed metadata track is used for the dynamic viewport information; The dynamic viewport information includes at least first viewport information including initial viewport position information, or second viewport information including viewport position information that dynamically changes over time; the first viewport information is conveyed via a sample entry of the timed metadata track; the second viewport information is conveyed via samples of the timed metadata track; A method according to claim 1, wherein the static viewport information is transmitted in a track of the multiple tracks that is different from the timed metadata track.
10. The method of claim 9 , wherein the information about the viewport further comprises camera orientation information.
11. The method of claim 9 , wherein the information about the viewport further comprises horizontal and vertical field of view (FOV) information for generating the viewport.
12. The method of claim 9 , further comprising: rendering the decoded point cloud data based on information about the viewport.
13. 1. An apparatus for decoding point cloud data, comprising: a decapsulator for decapsulating a file into a bitstream including the point cloud data, the point cloud data includes geometry data, feature data, and occupancy map data; The file contains multiple tracks, each of the multiple tracks includes a sample entry and a sample; the samples in at least one of the multiple tracks include one of the geometry data, the attribute data, and the occupancy map data; the file further comprises signaling information; a decapsulator, the signaling information including information regarding a viewport for the point cloud data; a decoder for decoding the point cloud data, The information about the viewport is classified into dynamic viewport information that changes over time and static viewport information that does not change over time; a track of the multiple tracks is a timed metadata track, a sample entry type for the timed metadata track having a value to indicate that the timed metadata track is used for the dynamic viewport information; The dynamic viewport information includes at least first viewport information including initial viewport position information, or second viewport information including viewport position information that dynamically changes over time; the first viewport information is conveyed via a sample entry of the timed metadata track; the second viewport information is conveyed via samples of the timed metadata track; An apparatus, wherein the static viewport information is transmitted in a track of the multiple tracks that is different from the timed metadata track.
14. The apparatus of claim 13 , wherein the information about the viewport further comprises camera orientation information.
15. The apparatus of claim 13 , wherein the information about the viewport further comprises horizontal field of view (FOV) information and vertical FOV information for generating the viewport.
16. The apparatus of claim 13 , further comprising a renderer for rendering the decoded point cloud data based on information about the viewport.
Citation Information
Patent Citations
Scalable point cloud compression with transform, and corresponding decompression
US20170347122A1
Methods and apparatus for signaling viewports and regions of interest
US20190297132A1
Multiple-viewpoints related metadata transmission and reception method and apparatus
US20190313081A1
Point cloud compression
WO2019055963A1
Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device
WO2020162495A1