Point cloud data transmission / reception apparatus and point cloud data transmission / reception method
By encoding and encapsulating point cloud data into a file, and using signaling data to provide partial access information, the efficiency and complexity problems of point cloud data transmission and reception in the prior art are solved, and efficient and low-latency point cloud content transmission is achieved.
Patent Information
- Application Number
- CN202510223891.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-22
- Filing Date
- 2021-06-21
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to efficiently transmit and receive point cloud data, and there are problems with delay and encoding/decoding complexity.
By encoding the point cloud data, the encoded data is encapsulated into a file, and sent and received by sending the file, signaling data is used to provide part of the access information to optimize the transmission of point cloud content.
It realizes efficient point cloud data transmission and reception, reduces latency and encoding/decoding complexity, and provides better point cloud content services by optimizing the transmission of signaling data.
Smart Images

Figure CN120075205A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with the original application number 202180044178.4 (International Application No.: PCT / KR2021 / 007763, filing date: June 21, 2021, invention title: Point Cloud Data Sending Device, Point Cloud Data Sending Method, Point Cloud Data Receiving Device, and Point Cloud Data Receiving Method). Technical Field
[0002] Embodiments provide a method of providing point cloud content to provide various services such as virtual reality (VR), augmented reality (AR), mixed reality (MR), and autonomous driving services to a user. Background Art
[0003] A point cloud is a set of points in a three-dimensional (3D) space. Since the number of points in the 3D space is large, it is difficult to generate point cloud data.
[0004] A large throughput is required to send and receive point cloud data. Summary of the Invention
[0005] Technical Problem
[0006] An object of the present disclosure is to provide a point cloud data sending device, a point cloud data sending method, a point cloud data receiving device, and a point cloud data receiving method for efficiently sending and receiving point clouds.
[0007] Another object of the present disclosure is to provide a point cloud data sending device, a point cloud data sending method, a point cloud data receiving device, and a point cloud data receiving method for solving latency and encoding / decoding complexity.
[0008] Another object of the present disclosure is to provide a point cloud data sending device, a point cloud data sending method, a point cloud data receiving device, and a point cloud data receiving method for providing optimized point cloud content to a user by signaling information about one or more atlases.
[0009] Another object of the present disclosure is to provide a point cloud data sending device, a point cloud data sending method, a point cloud data receiving device, and a point cloud data receiving method for providing optimized point cloud content to a user by signaling spatial region information and information related to an object and an atlas.
[0010] Another object of the present disclosure is to provide a point cloud data sending device, a point cloud data sending method, a point cloud data receiving device, and a point cloud data receiving method for providing optimized point cloud content to a user by signaling spatial region information and information related to an object and an atlas at a file format level.
[0011] Another object of the present disclosure is to provide a point cloud data transmission device, a point cloud data transmission method, a point cloud data reception device, and a point cloud data reception method for signaling information on a spatial area that is static or changes over time in an orbit or project and information related to an object and an atlas to efficiently provide point cloud data to a user.
[0012] Additional advantages, objects, and features of the present disclosure will be partly set forth in the following description, and partly will be obvious to those of ordinary skill in the art upon examination of the following, or may be learned from practice of the present disclosure. The objects and other advantages of the present disclosure may be realized and obtained by the structures particularly pointed out in the written description and claims hereof as well as the appended drawings.
[0013] Technical solutions
[0014] To achieve these objects and other advantages and in accordance with the purpose of the present disclosure, as implemented and broadly described herein, a method of transmitting point cloud data may include the steps of: encoding the point cloud data; encapsulating a bitstream including the encoded point cloud data into a file; and transmitting the file, the bitstream being stored in one or more tracks of the file, the file further including signaling data, the point cloud data being composed of one or more objects, and the signaling data including at least one parameter set and information for partial access to the point cloud data.
[0015] According to an embodiment, the point cloud data includes geometric data, attribute data, and occupancy map data encoded by a video-based encoding scheme.
[0016] According to an embodiment, the information for partial access is at least one of static information that does not change over time or dynamic information that changes dynamically over time.
[0017] According to an embodiment, the one or more objects are associated with the one or more atlas tiles, the one or more atlas tiles constituting an atlas frame, and the information for partial access includes mapping information between the one or more objects and the one or more atlas tiles.
[0018] According to an embodiment, the mapping information includes object index information for identifying each of the one or more objects, quantity information for identifying the number of atlas tiles associated with the object indicated by the object index information, and atlas identification information for identifying each of the atlas tiles associated with the object.
[0019] According to an embodiment, a point cloud data transmitting device may include: an encoder configured to encode point cloud data; a packager configured to package a bitstream including the encoded point cloud data into a file; and a transmitter configured to transmit the file, the bitstream being stored in one or more tracks of the file, the file further including signaling data, the point cloud data being composed of one or more objects; and the signaling data including at least one parameter set and information for partial access to the point cloud data.
[0020] According to an embodiment, the point cloud data includes geometric data, attribute data, and occupancy map data encoded by a video-based coding scheme.
[0021] According to an embodiment, the information for partial access is at least one of static information that does not change over time or dynamic information that changes dynamically over time.
[0022] According to an embodiment, the one or more objects are associated with the one or more atlas tiles, the one or more atlas tiles constituting an atlas frame, and the information for partial access includes mapping information between the one or more objects and the one or more atlas tiles.
[0023] According to an embodiment, the mapping information includes object index information for identifying each of the one or more objects, quantity information for identifying the number of atlas tiles associated with the object indicated by the object index information, and atlas identification information for identifying each of the atlas tiles associated with the object.
[0024] According to an embodiment, a method for receiving point cloud data may include the steps of: receiving a file; unpackaging the file into a bitstream including point cloud data, the bitstream being stored in one or more tracks of the file, and the file further including signaling data; decoding all or a part of the point cloud data based on the signaling data; and rendering all or a part of the decoded point cloud data based on the signaling data, the point cloud data being composed of one or more objects, and the signaling data including at least one parameter set and information for partial access to the point cloud data.
[0025] According to an embodiment, the point cloud data includes geometric data, attribute data, and occupancy map data decoded by a video-based coding scheme.
[0026] According to an embodiment, the information for partial access is at least one of static information that does not change over time or dynamic information that changes dynamically over time.
[0027] According to an embodiment, the one or more objects are associated with the one or more atlas tiles, the one or more atlas tiles form an atlas frame, and the information for partial access includes mapping information between the one or more objects and the one or more atlas tiles.
[0028] According to an embodiment, the mapping information includes object index information for identifying each of the one or more objects, quantity information for identifying the number of atlas tiles associated with the object indicated by the object index information, and atlas identification information for identifying each of the atlas tiles associated with the object.
[0029] According to an embodiment, a point cloud data receiving device may include: a receiver configured to receive a file; a de-encapsulator configured to de-encapsulate the file into a bitstream including point cloud data, the bitstream being stored in one or more tracks of the file, and the file further including signaling data; a decoder configured to decode all or part of the point cloud data based on the signaling data; and a renderer configured to render all or part of the decoded point cloud data based on the signaling data, the point cloud data being composed of one or more objects; and the signaling data includes at least one parameter set and information for partial access of the point cloud data.
[0030] According to an embodiment, the point cloud data includes geometric data, attribute data, and occupancy map data decoded by a video-based coding scheme.
[0031] According to an embodiment, the information for partial access is at least one of static information that does not change over time or dynamic information that changes dynamically over time.
[0032] According to an embodiment, the one or more objects are associated with the one or more atlas tiles, the one or more atlas tiles form an atlas frame, and the information for partial access includes mapping information between the one or more objects and the one or more atlas tiles.
[0033] According to an embodiment, the mapping information includes object index information for identifying each of the one or more objects, quantity information for identifying the number of atlas tiles associated with the object indicated by the object index information, and atlas identification information for identifying each of the atlas tiles associated with the object.
[0034] Advantageous Effects
[0035] The point cloud data sending method, sending device, point cloud data receiving method, and receiving device according to the embodiment can provide point cloud services of good quality.
[0036] The point cloud data sending method, sending device, point cloud data receiving method, and receiving device according to the embodiment can implement various video coding and decoding schemes.
[0037] The point cloud data sending method, sending device, point cloud data receiving method, and receiving device according to the embodiment can provide general point cloud content such as self-driving services.
[0038] The point cloud data sending method, sending device, point cloud data receiving method, and receiving device according to the embodiment can provide the best point cloud content service by configuring the V-PCC bitstream and allowing the file to be sent, received, and stored.
[0039] By using the point cloud data sending method, sending device, point cloud data receiving method, and receiving device according to the embodiment, the V-PCC bitstream can be efficiently accessed by multiplexing the V-PCC bitstream on the basis of the V-PCC unit. In addition, the atlas bitstream (or atlas sub-stream) of the V-PCC bitstream can be effectively stored in the track in the file, and thus can be sent and received.
[0040] By using the point cloud data sending method, sending device, point cloud data receiving method, and receiving device according to the embodiment, the V-PCC bitstream can be divided and stored in one or more tracks in the file, and the information indicating the relationship between the multiple tracks stored in the V-PCC bitstream can be signaled. Therefore, the file of the point cloud bitstream can be efficiently stored and sent.
[0041] By using the point cloud data sending method, point cloud data sending device, point cloud data receiving method, and point cloud data receiving device according to the embodiment, the metadata for data processing and rendering in the V-PCC bitstream can be sent and received in the V-PCC bitstream. Therefore, the best point cloud content service can be provided.
[0042] By using the point cloud data sending method, point cloud data sending device, point cloud data receiving method, and point cloud data receiving device according to the embodiment, the atlas parameter set can be stored and delivered in the track or item of the file for decoding and rendering the atlas sub-stream in the V-PCC bitstream. Therefore, the V-PCC decoder / player can operate effectively when decoding the V-PCC bitstream and the atlas sub-stream or parsing and processing the bitstream in the track / item. In addition, even when the atlas sub-bitstream is divided into one or more tracks and stored, the necessary atlas data and related video data can be effectively selected, extracted, and decoded.
[0043] Using a point cloud data transmission method, a point cloud data transmission device, a point cloud data reception method, and a point cloud data reception device according to an embodiment, point cloud data can be divided into a plurality of spatial regions to be processed for partial access and / or spatial access to point cloud content. Accordingly, encoding and transmission operations on the transmission side and decoding and rendering operations on the reception side can be performed in real time, and can be processed with low latency.
[0044] Using a point cloud data transmission method, a point cloud data transmission device, a point cloud data reception method, and a point cloud data reception device according to an embodiment, spatial region information about a spatial region segmented from point cloud content can be provided. Accordingly, point cloud content can be accessed in various ways in consideration of a player or user environment at the reception side.
[0045] Using a point cloud data transmission method, a point cloud data transmission device, a point cloud data reception method, and a point cloud data reception device according to an embodiment, spatial region information for data processing and rendering in a V-PCC bitstream can be transmitted and received at a file format level through a track. Accordingly, an optimal point cloud content service can be provided.
[0046] Using a point cloud data transmission method, a point cloud data transmission device, a point cloud data reception method, and a point cloud data reception device according to an embodiment, according to the degree of association between one or more objects segmented from point cloud content and one or more atlas tiles of an atlas frame, the association between the object and the atlas tile can be delivered in a sample, a sample group, a sample entry, a track group, on an entity group in a track, or in a timing metadata track. Accordingly, object-based access and access to atlas tiles associated with an object or associated point cloud data can be enabled.
[0047] Using a point cloud data transmission method, a point cloud data transmission device, a point cloud data reception method, and a point cloud data reception device according to an embodiment, according to the degree of association between one or more objects segmented from point cloud content and one or more atlases, the association between the object and the atlas can be delivered in a sample, a sample group, a sample entry, a track group, on an entity group in a track, or in a timing metadata track. Accordingly, an atlas associated with an object can be efficiently extracted.
[0048] Using a point cloud data transmission method, a point cloud data transmission device, a point cloud data reception method, and a point cloud data reception device according to an embodiment, an SEI message for processing and rendering data in a V-PCC bitstream can be stored and delivered to a track or an item in a file. Accordingly, a PCC decoder / player can operate efficiently in a step of decoding a V-PCC bitstream or in a step of parsing and processing a corresponding bitstream in a track / item. Description of the Drawings
[0049] The drawings are included to provide a further understanding of the present disclosure, and are incorporated into and constitute a part of this application. The drawings illustrate embodiments of the present disclosure and, together with the description, are used to explain the principles of the present disclosure. In the drawings:
[0050] Figure 1 Illustrates an exemplary structure of a transmission / reception system for providing point cloud content according to an embodiment.
[0051] Figure 2 Illustrates the capture of point cloud data according to an embodiment.
[0052] Figure 3 Illustrates exemplary point clouds, geometry, and texture images according to an embodiment.
[0053] Figure 4 Illustrates an exemplary V-PCC encoding process according to an embodiment.
[0054] Figure 5 Illustrates an example of a tangent plane and a normal vector of a surface according to an embodiment.
[0055] Figure 6 Illustrates an exemplary bounding box of a point cloud according to an embodiment.
[0056] Figure 7 Illustrates an example of the determination of respective patch positions on an occupancy map according to an embodiment.
[0057] Figure 8 Illustrates an exemplary relationship between a normal axis, a tangential axis, and a bitangential axis according to an embodiment.
[0058] Figure 9 Illustrates an exemplary configuration of a minimum mode and a maximum mode of a projection pattern according to an embodiment.
[0059] Figure 10 Illustrates an exemplary EDD code according to an embodiment.
[0060] Figure 11 Illustrates an example of recoloring based on color values of neighboring points according to an embodiment.
[0061] Figure 12 Illustrates an example of push-pull background filling according to an embodiment.
[0062] Figure 13 Illustrates an exemplary possible traversal order of a 4*4 block according to an embodiment.
[0063] Figure 14 Illustrates an exemplary optimal traversal order according to an embodiment.
[0064] Figure 15 Shows an exemplary 2D video / image encoder according to an embodiment.
[0065] Figure 16 Shows an exemplary V-PCC decoding process according to an embodiment.
[0066] Figure 17 Shows an exemplary 2D video / image decoder according to an embodiment.
[0067] Figure 18 Is a flowchart showing the operation of a transmitting device according to an embodiment of the present disclosure.
[0068] Figure 19 Is a flowchart showing the operation of a receiving device according to an embodiment.
[0069] Figure 20 Shows an exemplary architecture for V-PCC-based storage and streaming of point cloud data according to an embodiment.
[0070] Figure 21 Is an exemplary block diagram of a device for storing and transmitting point cloud data according to an embodiment.
[0071] Figure 22 Is an exemplary block diagram of a point cloud data receiving device according to an embodiment.
[0072] Figure 23 Shows an exemplary structure that can operate in conjunction with a point cloud data transmission / reception method / device according to an embodiment.
[0073] Figure 24 Shows an example of the relationship between an object and an atlas tile according to an embodiment.
[0074] Figure 25 Is a diagram showing an exemplary V-PCC bitstream structure according to an embodiment.
[0075] Figure 26 Shows an example of data carried by a sample stream V-PCC unit in a V-PCC bitstream according to an embodiment.
[0076] Figure 27 Shows an exemplary syntax structure of a sample stream V-PCC header included in a V-PCC bitstream according to an embodiment.
[0077] Figure 28 Shows an exemplary syntax structure of a sample stream V-PCC unit according to an embodiment.
[0078] Figure 29 Shows an exemplary syntax structure of a V-PCC unit according to an embodiment.
[0079] Figure 30 Shows an exemplary syntax structure of a V-PCC unit header according to an embodiment.
[0080] Figure 31 Shows an exemplary V-PCC unit type assigned to the vuh_unit_type field according to an embodiment.
[0081] Figure 32 Shows an exemplary syntax structure of a V-PCC unit payload (vpcc_unit_payload()) according to an embodiment.
[0082] Figure 33 Shows an exemplary syntax structure of a V-PCC parameter set included in a V-PCC unit payload according to an embodiment.
[0083] Figure 34 Shows an example of dividing an Atlas frame into multiple tiles according to an embodiment.
[0084] Figure 35 Is a diagram showing an exemplary Atlas sub-stream structure according to an embodiment.
[0085] Figure 36 Shows an exemplary syntax structure of a sample stream NAL header included in an Atlas sub-stream according to an embodiment.
[0086] Figure 37 Shows an exemplary syntax structure of a sample stream NAL unit according to an embodiment.
[0087] Figure 38 Shows an embodiment of the syntax structure of nal_unit(NumBytesInNalUnit) according to an embodiment.
[0088] Figure 39 Shows an embodiment of the syntax structure of an NAL unit header according to an embodiment.
[0089] Figure 40 Shows an example of the type of an RBSP data structure assigned to the nal_unit_type field according to an embodiment.
[0090] Figure 41 Shows the syntax structure of the syntax of an Atlas sequence parameter set according to an embodiment.
[0091] Figure 42 Shows the syntax structure of an Atlas frame parameter set according to an embodiment.
[0092] Figure 43Shows the syntax structure of the Atlas frame tile information according to an embodiment.
[0093] Figure 44 Shows the syntax structure of the Atlas adaptation parameter set according to an embodiment.
[0094] Figure 45 Shows the syntax structure of the camera parameters according to an embodiment.
[0095] Figure 46 Shows an example of the camera model assigned to the acp_camera_model field according to an embodiment.
[0096] Figure 47 Shows the syntax structure of the Atlas tile group layer according to an embodiment.
[0097] Figure 48 Shows the syntax structure of the Atlas tile group (or tile) header included in the Atlas tile group layer according to an embodiment.
[0098] Figure 49 Shows an example of the encoding type assigned to the atgh_type field according to an embodiment.
[0099] Figure 50 Shows an embodiment of the ref_list_struct() syntax structure according to an embodiment.
[0100] Figure 51 Shows the Atlas tile group (or tile) data unit according to an embodiment.
[0101] Figure 52 Shows an example of the patch mode type assigned to the atgdu_patch_mode field when the atgh_type field indicates I_TILE_GRP according to an embodiment.
[0102] Figure 53 Shows an example of the patch mode type assigned to the atgdu_patch_mode field when the atgh_type field indicates P_TILE_GRP according to an embodiment.
[0103] Figure 54 Shows an example of the patch mode type assigned to the atgdu_patch_mode field when the atgh_type field indicates SKIP_TILE_GRP according to an embodiment.
[0104] Figure 55 Shows the patch information data according to an embodiment.
[0105] Figure 56Shows the syntax structure of a patch data unit according to an embodiment.
[0106] Figure 57 Shows the rotation and offset regarding patch orientation according to an embodiment.
[0107] Figure 58 Shows the syntax structure of SEI information according to an embodiment.
[0108] Figure 59 Shows an exemplary syntax structure of an SEI message payload according to an embodiment.
[0109] Figure 60 Is a table showing examples of a sample entry structure and a sample format according to a sample entry type according to an embodiment.
[0110] Figure 61 Is a diagram showing an example of the transmission of one or more Atlas tiles through one V3C track according to an embodiment.
[0111] Figure 62 Is a diagram showing an example of the transmission of one or more Atlas tiles through multiple V3C tracks according to an embodiment.
[0112] Figure 63 Is a diagram showing an exemplary track structure for storing a single Atlas (or a single Atlas bitstream) with a single Atlas tile according to an embodiment.
[0113] Figure 64 Is a diagram showing an exemplary track structure for storing a single Atlas with multiple Atlas tiles according to an embodiment.
[0114] Figure 65 Is a diagram showing an exemplary track structure for storing multiple Atlases according to an embodiment.
[0115] Figure 66 Is a diagram showing an exemplary track structure for storing multiple Atlases according to an embodiment.
[0116] Figure 67 Is a table showing an example of the relationship between a sample entry type and a track reference according to an embodiment.
[0117] Figure 68 Is a diagram showing an exemplary structure for encapsulating non-temporal V-PCC data according to an embodiment.
[0118] Figure 69 Is a flowchart of a method for transmitting point cloud data according to an embodiment.
[0119] Figure 70It is a flowchart showing a method for receiving point cloud data according to an embodiment. Detailed Embodiment
[0120] Now, a preferred embodiment of the present disclosure will be described in detail, and examples thereof are shown in the accompanying drawings. The following detailed description with reference to the accompanying drawings is intended to illustrate the exemplary embodiments of the present disclosure, rather than showing the only embodiments that can be implemented according to the present disclosure. The following detailed description includes specific details to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure can be practiced without these specific details.
[0121] Although most of the terms used in the present disclosure are selected from the general terms widely used in the art, some terms are arbitrarily selected by the applicant and their meanings are detailed as needed in the following description. Therefore, the present disclosure should be understood based on the intended meanings of the terms rather than their simple names or meanings.
[0122] Figure 1 It shows an exemplary structure of a transmission / reception system for providing point cloud content according to an embodiment.
[0123] The present disclosure provides a method for providing point cloud content to provide various services such as virtual reality (VR), augmented reality (AR), mixed reality (MR), and autonomous driving to users. Point cloud data according to an embodiment represents data representing an object as points, and may be referred to as point cloud, point cloud data, point cloud video data, point cloud image data, etc.
[0124] The point cloud data transmission device 10000 according to an embodiment may include a point cloud video acquisition unit 10001, a point cloud video encoder 10002, a file / fragment encapsulation module (file / fragment encapsulator) 10003, and / or a transmitter (or communication module) 10004. The transmission device according to an embodiment may acquire and process point cloud video (or point cloud content) and transmit it. According to an embodiment, the transmission device may include a fixed station, a base transceiver system (BTS), a network, an artificial intelligence (AI) device and / or system, a robot, and an AR / VR / XR device and / or server. According to an embodiment, the transmission device 10000 may include a device configured to perform communication with a base station and / or other wireless devices using radio access technologies (e.g., 5G new radio access technology (NR), long term evolution (LTE)), a robot, a vehicle, an AR / VR / XR device, a portable device, a home appliance, an Internet of Things (IoT) device, and an AI device / server.
[0125] The point cloud video acquisition unit 10001 according to an embodiment acquires a point cloud video through a process of capturing, synthesizing, or generating the point cloud video.
[0126] The point cloud video encoder 10002 according to an embodiment encodes the point cloud video data obtained from the point cloud video acquisition unit 10001. According to an embodiment, the point cloud video encoder 10002 may be referred to as a point cloud encoder, a point cloud data encoder, an encoder, etc. The point cloud compression encoding (encoding) according to an embodiment is not limited to the above embodiment. The point cloud video encoder may output a bitstream including the encoded point cloud video data. The bitstream may include not only the encoded point cloud video data but also signaling information related to the encoding of the point cloud video data.
[0127] The point cloud video encoder 10002 according to an embodiment may support a geometry-based point cloud compression (G-PCC) encoding scheme and / or a video-based point cloud compression (V-PCC) encoding scheme. In addition, the point cloud video encoder 10002 may encode point clouds (referred to as point cloud data or points) and / or signaling data related to the point clouds.
[0128] The file / fragment encapsulation module 10003 according to an embodiment encapsulates the point cloud data in the form of a file and / or a fragment. The point cloud data sending method / device according to an embodiment may send the point cloud data in the form of a file and / or a fragment.
[0129] The transmitter (or communication module) 10004 according to an embodiment sends the encoded point cloud video data in the form of a bitstream. According to an embodiment, a file or a fragment may be sent to a receiving device via a network or stored in a digital storage medium (e.g., USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.). The transmitter according to an embodiment is capable of performing wired / wireless communication with a receiving device (or receiver) via a network such as 4G, 5G, 6G, etc. Additionally, the transmitter may perform necessary data processing operations according to a network system (e.g., a 4G, 5G, or 6G communication network system). The sending device may send the encapsulated data in an on-demand manner.
[0130] The point cloud data receiving device 10005 according to an embodiment may include a receiver 10006, a file / fragment de-encapsulator (or file / fragment de-encapsulation module) 10007, a point cloud video decoder 10008, and / or a renderer 10009. According to an embodiment, the receiving device may include a device configured to perform communication with a base station and / or other wireless devices using radio access technologies (e.g., 5G new radio access technology (NR), long-term evolution (LTE)), a robot, a vehicle, an AR / VR / XR device, a portable device, a household appliance, an Internet of Things (IoT) device, and an AI device / server.
[0131] The receiver 10006 according to an embodiment receives a bitstream containing the point cloud video data. According to an embodiment, the receiver 10006 may send feedback information to the point cloud data sending device 10000.
[0132] The file / fragment decompression module 10007 decompresses the file and / or fragment containing the point cloud data.
[0133] The point cloud video decoder 10008 decodes the received point cloud video data.
[0134] The renderer 10009 renders the decoded point cloud video data. According to an embodiment, the renderer 10009 may send the feedback information obtained on the receiving side to the point cloud video decoder 10008. According to an embodiment of the point cloud video data, the feedback information may be carried to the receiver 10006. According to an embodiment, the feedback information received by the point cloud transmitting device may be provided to the point cloud video encoder 10002.
[0135] The arrow indicated by the dashed line in the figure represents the transmission path of the feedback information acquired by the receiving device 10005. The feedback information is information reflecting the interactivity with the user consuming the point cloud content and includes user information (e.g., head orientation information, viewport information, etc.). Specifically, when the point cloud content is content for a service that requires interaction with the user (e.g., autonomous driving service, etc.), the feedback information may be provided to the content sender (e.g., the transmitting device 10000) and / or the service provider. According to an embodiment, the feedback information may be used in the receiving device 10005 and the transmitting device 10000 and may not be provided.
[0136] The head orientation information according to an embodiment is information about the user's head position, orientation, angle, movement, etc. The receiving device 10005 according to an embodiment may calculate the viewport information based on the head orientation information. The viewport information may be information about the area of the point cloud video that the user is watching. The viewpoint (or orientation) is the point where the user watches the point cloud video and may refer to the center point of the viewport area. That is, the viewport is an area centered on the viewpoint, and the size and shape of the area may be determined by the field of view (FOV). In other words, the viewport is determined according to the position and viewpoint (or orientation) of the visual camera or the user, and the point cloud data is rendered in the viewport based on the viewport information. Therefore, in addition to the head orientation information, the receiving device 10005 may also extract the viewport information based on the vertical or horizontal FOV supported by the device. In addition, the receiving device 10005 performs gaze analysis to check how the user consumes the point cloud, the area in the point cloud video that the user gazes at, the gaze time, etc. According to an embodiment, the receiving device 10005 may send the feedback information including the gaze analysis result to the transmitting device 10000. The feedback information according to an embodiment may be acquired during the rendering and / or display process. The feedback information according to an embodiment may be obtained by one or more sensors included in the receiving device 10005. In addition, according to an embodiment, the feedback information may be obtained by the renderer 10009 or a separate external element (or device, component, etc.).Figure 1 The dashed line in Figure 1 indicates the process of obtaining the feedback information by the transmitting renderer 10009. The point cloud content providing system can process (encode / decode) the point cloud data based on the feedback information. Therefore, the point cloud video decoder 10008 can perform a decoding operation based on the feedback information. The receiving device 10005 can send the feedback information to the transmitting device. The transmitting device (or the point cloud video encoder 10002) can perform an encoding operation based on the feedback information. Therefore, the point cloud content providing system can efficiently process the necessary data (e.g., the point cloud data corresponding to the user's head position) based on the feedback information instead of processing (encoding / decoding) all the point cloud data, and provide the point cloud content to the user.
[0137] According to an embodiment, the transmitting device 10000 may be referred to as an encoder, a transmitting device, a transmitter, etc., and the receiving device 10005 may be referred to as a decoder, a receiving device, a receiver, etc.
[0138] According to an embodiment of Figure 1 the point cloud data processed in the point cloud content providing system (through a series of processes of acquisition / encoding / transmission / decoding / rendering) may be referred to as point cloud content data or point cloud video data. According to an embodiment, the point cloud content data can be used as a concept covering metadata or signaling information related to the point cloud data.
[0139] Figure 1 The components of the point cloud content providing system shown can be implemented by hardware, software, a processor, and / or a combination thereof.
[0140] An embodiment can provide a method for providing point cloud content to provide various services such as virtual reality (VR), augmented reality (AR), mixed reality (MR), and autonomous driving to a user.
[0141] To provide a point cloud content service, a point cloud video can be first acquired. The acquired point cloud video can be sent to the receiving side through a series of processes, and the receiving side can process the received data back to the original point cloud video and render the processed point cloud video. Thus, the point cloud video can be provided to the user. An embodiment provides a method for effectively performing this series of processes.
[0142] All the processes for providing the point cloud content service (the point cloud data sending method and / or the point cloud data receiving method) may include an acquisition process, an encoding process, a transmission process, a decoding process, a rendering process, and / or a feedback process.
[0143] According to an embodiment, the process of providing the point cloud content (or the point cloud data) may be referred to as a point cloud compression process. According to an embodiment, the point cloud compression process may represent a video-based point cloud compression (V-PCC) process.
[0144] Each element of the point cloud data transmitting device and the point cloud data receiving device according to the embodiment may be hardware, software, a processor, and / or a combination thereof.
[0145] The point cloud compression system may include a transmitting device and a receiving device. According to an embodiment, the transmitting device may be referred to as an encoder, a transmitting device, a transmitter, a point cloud data transmitting device, etc. According to an embodiment, the receiving device may be referred to as a decoder, a receiving device, a receiver, a point cloud data receiving device, etc. The transmitting device may output a bitstream by encoding a point cloud video and transmit it to the receiving device in the form of a file or a stream (streaming segment) through a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD.
[0146] The transmitting device may include a point cloud video acquisition unit, a point cloud video encoder, a file / segment encapsulator, and a transmitting unit (transmitter), as Figure 1 shown. The receiving device may include a receiver, a file / segment de-encapsulator, a point cloud video decoder, and a renderer, as Figure 1 shown. The encoder may be referred to as a point cloud video / picture / picture / frame encoder, and the decoder may be referred to as a point cloud video / picture / picture / frame decoding device. The renderer may include a display. The renderer and / or the display may be configured as a separate device or an external component. The transmitting device and the receiving device may also include separate internal or external modules / units / components for feedback processing. According to an embodiment, each element in the transmitting device and the receiving device may be configured by hardware, software, and / or a processor.
[0147] According to an embodiment, the operation of the receiving device may be a reverse process of the operation of the transmitting device.
[0148] The point cloud video acquirer may perform the process of acquiring a point cloud video through processes such as capturing, arranging, or generating a point cloud video. In the acquisition process, data of the 3D positions (x, y, z) / attributes (color, reflectivity, transparency, etc.) of multiple points may be generated, such as a polygon file format (PLY) (or Stanford triangle format) file. For a video having multiple frames, one or more files may be acquired. During the capture process, point cloud-related metadata (e.g., capture-related metadata) may be generated.
[0149] The point cloud data transmitting device according to an embodiment may include an encoder configured to encode point cloud data and a transmitter configured to transmit the point cloud data or a bitstream including the point cloud data.
[0150] The point cloud data receiving device according to an embodiment may include a receiver configured to receive a bitstream including point cloud data, a decoder configured to decode the point cloud data, and a renderer configured to render the point cloud data.
[0151] The method / device according to an embodiment represents a point cloud data transmitting device and / or a point cloud data receiving device.
[0152] Figure 2 Shows the capture of point cloud data according to an embodiment.
[0153] The point cloud data (point cloud video data) according to an embodiment may be obtained through a camera or the like. The capture technique according to an embodiment may include, for example, inward-facing and / or outward-facing.
[0154] In the inward-facing according to an embodiment, one or more cameras facing an object of point cloud data inward may capture the object from outside the object.
[0155] In the outward-facing according to an embodiment, one or more cameras facing an object of point cloud data outward may capture the object. For example, according to an embodiment, there may be four cameras.
[0156] The point cloud data or point cloud content according to an embodiment may be a video or still image of an object / environment represented in various types of 3D spaces. According to an embodiment, the point cloud content may include video / audio / image of an object.
[0157] As a device for capturing point cloud content, a combination of a camera device capable of acquiring depth (a combination of an infrared pattern projector and an infrared camera) and an RGB camera capable of extracting color information corresponding to depth information may be configured. Alternatively, depth information may be extracted by using LiDAR of a radar system that measures the position coordinates of a reflector by emitting laser pulses and measuring the return time. Geometries composed of points in 3D space may be extracted from the depth information, and attributes representing the color / reflection rate of each point may be extracted from the RGB information. The point cloud content may include information about the position (x, y, z) and the color (YCbCr or RGB) or reflection rate (r) of the points. For point cloud content, an outward-facing technique for capturing the external environment and an inward-facing technique for capturing the central object may be used. In a VR / AR environment, when an object (e.g., a core object such as a character, player, thing, or actor) is configured in the point cloud content that a user can view in any direction (360 degrees), the configuration of the capture camera may be based on the inward-facing technique. When the current surrounding environment is configured in the point cloud content in a vehicle mode such as autonomous driving, the configuration of the capture camera may be based on the outward-facing technique. Since point cloud content may be captured by multiple cameras, it may be necessary to perform camera calibration processing before capturing the content to configure a global coordinate system for the cameras.
[0158] The point cloud content may be a video or a still image of an object / environment existing in various types of 3D spaces.
[0159] In addition, in the point cloud content acquisition method, any point cloud video may be arranged based on the captured point cloud video. Alternatively, when a point cloud video of a computer-generated virtual space is to be provided, the capture using an actual camera may not be performed. In this case, the capture process may simply be replaced by a process of generating relevant data.
[0160] Post-processing of the captured point cloud video may be required to improve the quality of the content. In the video capture process, the maximum / minimum depth may be adjusted within the range provided by the camera device. Even after the adjustment, there may still be point data in unwanted areas. Therefore, post-processing such as removing unwanted areas (e.g., the background) or identifying connected spaces and filling space holes may be performed. In addition, the point clouds extracted from cameras sharing a spatial coordinate system may be integrated into one piece of content by a process of transforming each point into a global coordinate system based on the position coordinates of each camera obtained through a calibration process. Thus, one piece of point cloud content with a wide range or point cloud content with a high density of points may be generated.
[0161] The point cloud video encoder 10002 may encode an input point cloud video into one or more video streams. A point cloud video may include multiple frames, and each frame may correspond to a still image / frame. In this specification, a point cloud video may include point cloud images / frames / frames / videos / audio. In addition, the term "point cloud video" may be used interchangeably with point cloud images / frames / frames. The point cloud video encoder 10002 may perform video-based point cloud compression (V-PCC) processing. For compression and encoding efficiency, the point cloud video encoder may perform a series of processes such as prediction, transformation, quantization, and entropy encoding. The encoded data (encoded video / image information) may be output in the form of a bitstream. Based on the V-PCC processing, the point cloud video encoder may encode the point cloud video by dividing it into a geometry video, an attribute video, an occupancy map video, and auxiliary information (or auxiliary data) (which will be described later). The geometry video may include geometry images, the attribute video may include attribute images, and the occupancy map video may include occupancy map images. The auxiliary information may include auxiliary patch information. The attribute video / image may include texture video / images.
[0162] A file / fragment encapsulator (file / fragment encapsulation module) 10003 may encapsulate, for example, in the form of a file, the encoded point cloud video data and / or the metadata related to the point cloud video. Here, the metadata related to the point cloud video may be received from a metadata processor. The metadata processor may be included in the point cloud video encoder 10002 or may be configured as a separate component / module. The file / fragment encapsulator 10003 may encapsulate the data in a file format such as ISOBMFF or process the data in the form of DASH fragments, etc. According to an embodiment, the file / fragment encapsulator 10003 may include the point cloud video related metadata in the file format. The point cloud video metadata may be included in boxes at various levels of, for example, the ISOBMFF file format, or as data in a separate track within the file. According to an embodiment, the file / fragment encapsulator 10003 may encapsulate the point cloud video related metadata into the file. A transmission processor may perform a transmission process on the point cloud video data encapsulated according to the file format. The transmission processor may be included in the transmitter 10004 or may be configured as a separate component / module. The transmission processor may process the point cloud video data according to a transmission protocol. The transmission process may include a process of transmitting via a broadcast network and a process of transmitting via broadband. According to an embodiment, the transmission processor may receive the point cloud video related metadata together with the point cloud video data from the metadata processor and perform the process of the point cloud video data for transmission.
[0163] The transmitter 10004 may send the encoded video / image information or data output in the form of a bitstream to the receiver 10006 of the receiving device in the form of a file or a stream through a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitter may include an element for generating a media file in a predetermined file format and may include an element for transmitting via a broadcast / communication network. The receiver may extract the bitstream and send the extracted bitstream to the decoding device.
[0164] The receiver 10006 may receive the point cloud video data transmitted by the point cloud video transmitting device according to the present disclosure. According to the transmission channel, the receiver may receive the point cloud video data via a broadcast network or via broadband. Alternatively, the point cloud video data may be received through a digital storage medium.
[0165] The receiving processor can process the received point cloud video data according to the transmission protocol. The receiving processor can be included in the receiver 10006 or can be configured as a separate component / module. The receiving processor can perform the above processing of the sending processor in reverse, such that the processing corresponds to the transmission processing performed on the sending side. The receiving processor can transmit the acquired point cloud video data to the file / fragment de-packager 10007 and transmit the acquired point cloud video-related metadata to the metadata processor (not shown). The point cloud video-related metadata acquired by the receiving processor can take the form of a signaling table.
[0166] The file / fragment de-packager (file / fragment de-packaging module) 10007 can de-package the point cloud video data received from the receiving processor in the form of a file. The file / fragment de-packager 10007 can de-package the file according to ISOBMFF, etc., and can acquire the point cloud video bitstream or the point cloud video-related metadata (metadata bitstream). The acquired point cloud video bitstream can be transmitted to the point cloud video decoder 10008, and the acquired point cloud video-related metadata (metadata bitstream) can be transmitted to the metadata processor (not shown). The point cloud video bitstream can include metadata (metadata bitstream). The metadata processor can be included in the point cloud video decoder 10008 or can be configured as a separate component / module. The point cloud video-related metadata acquired by the file / fragment de-packager 10007 can take the form of a box or track in the file format. When needed, the file / fragment de-packager 10007 can receive the metadata required for de-packaging from the metadata processor. The point cloud video-related metadata can be transmitted to the point cloud video decoder 10008 and used in the point cloud video decoding process, or can be transmitted to the renderer 10009 and used in the point cloud video rendering process.
[0167] The point cloud video decoder 10008 can receive the bitstream and decode the video / image by performing operations corresponding to the operations of the point cloud video encoder. In this case, the point cloud video decoder 10008 can decode the point cloud video by dividing the point cloud video into a geometry video, an attribute video, an occupancy map video, and auxiliary information as described below. The geometry video can include geometry images, the attribute video can include attribute images. The occupancy map video can include occupancy map images. The auxiliary information can include auxiliary patch information. The attribute video / image can include texture video / images.
[0168] The 3D geometry can be reconstructed based on the decoded geometry image, occupancy map, and auxiliary patch information, and then can be subjected to smoothing processing. The color point cloud image / frame can be reconstructed by assigning color values to the smoothed 3D geometry based on the texture image. The renderer 10009 can render the reconstructed geometry and color point cloud image / frame. The rendered video / image can be displayed through a display (not shown). The user can view all or part of the rendering result through a VR / AR display or a typical display.
[0169] The feedback processing can include transmitting various types of feedback information that can be obtained in the rendering / display processing to the decoder on the transmitting side or the receiving side. Interactivity can be provided through the feedback processing when consuming the point cloud video. According to an embodiment, the head orientation information, viewport information indicating the area that the user is currently viewing, etc. can be transmitted to the transmitting side in the feedback processing. According to an embodiment, the user can interact with things implemented in the VR / AR / MR / autonomous driving environment. In this case, the information related to the interaction can be transmitted to the transmitting side or the service provider during the feedback processing. According to an embodiment, the feedback processing can be skipped.
[0170] The head orientation information can represent the information of the position, angle, and movement of the user's head. Based on this information, the information about the area of the point cloud video that the user is currently viewing (i.e., the viewport information) can be calculated.
[0171] The viewport information can be the information about the area of the point cloud video that the user is currently viewing. The viewport information can be used to perform gaze analysis to examine the way the user consumes the point cloud video, the area of the point cloud video that the user gazes at, and how long the user gazes at that area. The gaze analysis can be performed on the receiving side, and the analysis result can be transmitted to the transmitting side on the feedback channel. Devices such as VR / AR / MR displays can extract the viewport area based on the position / direction of the user's head, the vertical or horizontal FOV supported by the device, etc.
[0172] According to an embodiment, the above feedback information can be not only transmitted to the transmitting side, but also consumed on the receiving side. That is, the decoding and rendering processing on the receiving side can be performed based on the above feedback information. For example, only the point cloud video in the area that the user is currently viewing can be preferentially decoded and rendered based on the head orientation information and / or the viewport information.
[0173] Here, the viewport or the viewport area can represent the area of the point cloud video that the user is currently viewing. The viewing point is the point in the point cloud video that the user views, and can represent the center point of the viewport area. That is, the viewport is the area around the viewing point, and the size and form of the area can be determined by the field of view (FOV).
[0174] The present disclosure relates to point cloud video compression as described above. For example, the methods / embodiments disclosed in the present disclosure can be applied to the point cloud compression or point cloud coding (PCC) standard of the Moving Picture Experts Group (MPEG) or the next-generation video / image coding standard.
[0175] As used herein, a picture / frame generally can represent a unit representing an image in a specific time interval.
[0176] A pixel or picture element can be the smallest unit constituting a picture (or image). Additionally, "sample" can be used as a term corresponding to a pixel. A sample generally can represent a pixel or a pixel value. It can represent only the pixel / pixel value of the luminance component, only the pixel / pixel value of the chrominance component, or only the pixel / pixel value of the depth component.
[0177] A unit can represent a basic unit of image processing. A unit can include at least one of a specific area of a picture and information related to the area. In some cases, a unit can be used interchangeably with terms such as a block or a region or a module. Generally, an M×N block can include a sample (or an array of samples) or a set (or an array) of transform coefficients configured in M columns and N rows.
[0178] Figure 3 Examples of a point cloud, a geometry image, and a texture image according to an embodiment are shown.
[0179] A point cloud according to an embodiment can be input to the Figure 4 V-PCC encoding process described below to generate a geometry image and a texture image. According to an embodiment, a point cloud can have the same meaning as point cloud data.
[0180] As Figure 3 shown, the left part shows a point cloud, where the point cloud object is located in a 3D space and can be represented by a bounding box or the like. Figure 3 The middle part in Figure 3 shows a geometry image, and
[0181] The right part in shows a texture image (non-filled image). In the present disclosure, a geometry image can be referred to as a geometry patch frame / picture or a geometry frame / picture, and a texture image can be referred to as an attribute patch frame / picture or an attribute frame / picture.
[0182] Occupancy Map: This is a binary map that uses values 0 or 1 to indicate whether there is data at the corresponding position in the 2D plane when dividing the points that make up the point cloud into patches and mapping them to the 2D plane. The occupancy map can represent a 2D array corresponding to an atlas, and the values of the occupancy map can indicate whether each sample position in the atlas corresponds to a 3D point. An atlas refers to an object that includes information about 2D patches of each point cloud frame. For example, the atlas can include the 2D layout and the size of the patch, the position of the corresponding 3D region within the 3D point, the projection plan, and the level-of-detail parameters.
[0183] Patch: A set of points that make up a point cloud, indicating that the points belonging to the same patch are adjacent to each other in 3D space and are mapped in the same direction between the six-sided bounding box planes during the process of mapping to a 2D image.
[0184] Geometry Image: This is an image in the form of a depth map that presents the position information (geometry) of each point that makes up the point cloud patch by patch. The geometry image can consist of pixel values of one channel. The geometry represents a set of coordinates associated with the point cloud frame.
[0185] Texture Image: This is an image that represents the color information of each point that makes up the point cloud patch by patch. The texture image can consist of pixel values of multiple channels (e.g., three channels of R, G, and B). The texture is included in the attributes. According to an embodiment, the texture and / or the attributes can be interpreted as the same object and / or have an inclusion relationship.
[0186] Auxiliary Patch Information: This indicates the metadata required to reconstruct the point cloud using each patch. The auxiliary patch information can include information about the position, size, etc. of the patch in the 2D / 3D space.
[0187] Point cloud data according to an embodiment, for example, a V-PCC component, can include an atlas, an occupancy map, geometry, and attributes.
[0188] The atlas represents a set of 2D bounding boxes. It can be a set of patches, for example, patches projected into a rectangular frame corresponding to a 3D bounding box in 3D space, and the 3D space can represent a subset of the point cloud. In this case, the patch can represent a rectangular region in the atlas corresponding to a rectangular region in the plane projection. In addition, the patch data can represent the data required to perform the transformation of the patches included in the atlas from 2D to 3D. In addition, the set of patch data is also called an atlas.
[0189] Attributes can represent a scalar or vector associated with each point in the point cloud. For example, the attributes can include color, reflectivity, surface normal, timestamp, material ID.
[0190] The point cloud data according to the embodiment represents PCC data according to a video-based point cloud compression (V-PCC) scheme. The point cloud data may include multiple components. For example, it may include an occupancy map, patches, geometry, and / or texture.
[0191] Figure 4 An example of a point cloud video encoder according to an embodiment is shown.
[0192] Figure 4 A V-PCC encoding process for generating and compressing an occupancy map, a geometry image, a texture image, and auxiliary patch information is shown. Figure 4 The V-PCC encoding process of can be processed by Figure 1 the point cloud video encoder 10002 of. Figure 4 Each element of can be executed by software, hardware, a processor, and / or a combination thereof.
[0193] The patch generation or patch generator 14000 receives a point cloud frame (which may be in the form of a bitstream containing point cloud data). The patch generator 14000 generates patches according to the point cloud data. In addition, patch information including information about patch generation is generated.
[0194] The patch packing or patch packer 14001 packs one or more patches. Additionally, the patch packer 14001 generates an occupancy map containing information about patch packing.
[0195] The geometry image generation or geometry image generator 14002 generates a geometry image based on the point cloud data, patch information (or auxiliary information), and / or occupancy map information. The geometry image refers to data containing geometry related to the point cloud data (i.e., 3D coordinate values of points), and refers to a geometry framework.
[0196] The texture image generation or texture image generator 14003 generates a texture image based on the point cloud data, patches, packed patches, patch information (or auxiliary information), and / or smoothed geometry. The texture image refers to an attribute frame. That is, the texture image can also be generated based on the smoothed geometry generated by smoothing processing based on patch information.
[0197] The smoothing or smoother 14004 can reduce or eliminate errors contained in the image data. For example, the reconstructed geometry image is smoothed based on patch information. That is, parts that may cause errors between data can be smoothly filtered out to generate smoothed geometry.
[0198] Auxiliary patch information compression or auxiliary patch information compressor 14005 may compress auxiliary patch information related to the patch information generated in patch generation. In addition, the auxiliary patch information compressed in the auxiliary patch information compressor 14005 may be sent to the multiplexer 14013. The auxiliary patch information may be used in the geometric image generator 14002. According to an embodiment, the compressed auxiliary patch information may be referred to as a bitstream of the compressed auxiliary patch information, an auxiliary patch information bitstream, a bitstream of the compressed atlas, or an atlas bitstream, etc.
[0199] Image filling or image fillers 14006 and 14007 may fill the geometric image and the texture image, respectively. The filling data may be filled into the geometric image and the texture image.
[0200] Group dilation or group dilator 14008 may add data to the texture image in a manner similar to image filling. The auxiliary patch information may be inserted into the texture image.
[0201] Video compression or video compressors 14009, 14010, and 14011 may compress the filled geometric image, the filled texture image, and / or the occupancy map, respectively. In other words, the video compressors 14009, 14010, and 14011 may compress the input geometric frame, attribute frame, and / or occupancy map frame, respectively, to output a video bitstream of the geometric image, a video bitstream of the texture image, and a video bitstream of the occupancy map. Video compression may encode geometric information, texture information, and occupancy information. According to an embodiment, the video bitstream of the compressed geometry may be referred to as a 2D video-encoded geometry bitstream, a compressed geometry bitstream, a video-encoded geometry bitstream, or geometric video data, etc. According to an embodiment, the video bitstream of the compressed texture image may be referred to as a 2D video-encoded attribute bitstream, a compressed attribute bitstream, a video-encoded attribute bitstream, or attribute video data, etc.
[0202] Entropy compression or entropy compressor 14012 may compress the occupancy map based on an entropy scheme.
[0203] According to an embodiment, entropy compression and / or video compression may be performed on the occupancy map frame according to whether the point cloud data is lossless and / or lossy. According to an embodiment, the entropy and / or video-compressed map may be referred to as a video bitstream of the compressed occupancy map, a 2D video-encoded occupancy map bitstream, an occupancy map bitstream, a compressed occupancy map bitstream, a video-encoded occupancy map bitstream, or occupancy map video data, etc.
[0204] The multiplexer 14013 multiplexes the video bitstream of the compressed geometry, the video bitstream of the compressed texture image, the video bitstream of the compressed occupancy map, and the bitstream of the compressed auxiliary patch information from the respective compressors into one bitstream.
[0205] The above blocks may be omitted or replaced by blocks having similar or identical functions. In addition, Figure 4 each of the blocks shown may be used as at least one of a processor, software, and hardware.
[0206] According to an embodiment Figure 4 a detailed operation description of each process is as follows.
[0207] Patch Generation (14000)
[0208] The patch generation process refers to a process of dividing a point cloud into patches (mapping units) in order to map the point cloud to a 2D image. The patch generation process can be divided into three steps: normal value calculation, segmentation, and patch segmentation.
[0209] The normal value calculation process will be described in detail with reference to Figure 5 below.
[0210] Figure 5 An example of a tangent plane and a normal vector of a surface according to an embodiment is shown.
[0211] In Figure 4 the patch generator 14000 of the V-PCC encoding process Figure 5 the following surface is used.
[0212] Normal Calculation Related to Patch Generation
[0213] Each point of the point cloud has its own direction, which is represented by a 3D vector called a normal vector. Using the neighbors of each point obtained by using a K-D tree or the like, the tangent plane and the normal vector of each point constituting the surface of the point cloud can be obtained as Figure 5 shown. The search range applied to the process of searching for neighbors can be defined by the user.
[0214] A tangent plane refers to a plane that passes through a point on the surface and completely includes the tangent of a curve on the surface.
[0215] Figure 6 An exemplary bounding box of a point cloud according to an embodiment is shown.
[0216] A bounding box according to an embodiment refers to a box that is a unit for dividing point cloud data based on a hexahedron in a 3D space.
[0217] A method / apparatus (e.g., the patch generator 14000) according to an embodiment may use a bounding box in the process of generating patches from point cloud data.
[0218] The bounding box can be used in the process of projecting an object of interest in point cloud data onto the planes of the respective flat faces of a hexahedron in a 3D space. The bounding box may be defined by Figure 1The point cloud video acquisition unit 10001 and the point cloud video encoder 10002 generate and process. In addition, based on the bounding box, the patch generation 14000, patch packing 14001, geometric image generation 14002, and texture image generation 14003 for V-PCC encoding processing can be performed. Figure 4 The segmentation is divided into two processes: initial segmentation and refinement segmentation.
[0219] Segmentation Related to Patch Generation
[0220] The segmentation is divided into two processes: initial segmentation and refinement segmentation.
[0221] According to an embodiment, the point cloud video encoder 10002 projects points onto one face of the bounding box. Specifically, as Figure 6 shown, each point constituting the point cloud is projected onto one of the six faces of the bounding box surrounding the point cloud. The initial segmentation is a process of determining one of the flat faces of the bounding box to which each point is to be projected.
[0222] These are the normal values corresponding to each of the six flat faces, defined as follows:
[0223] (1.0, 0.0, 0.0), (0.0, 1.0, 0.0), (0.0, 0.0, 1.0), (-1.0, 0.0, 0.0), (0.0, -1.0, 0.0), (0.0, 0.0, -1.0).
[0224] As shown in the following formula, the normal vector of each point obtained in the normal value calculation process and The plane with the largest value is determined as the projection plane of the corresponding point. That is, the plane with the normal vector most similar to the direction of the normal vector of the point is determined as the projection plane of the point.
[0225]
[0226] The determined plane can be identified by a cluster index (one of 0 to 5).
[0227] The refinement segmentation is a process of enhancing the projection plane of each point constituting the point cloud determined in the initial segmentation process by considering the projection planes of neighboring points. In this process, the score normal and the score smoothness can be considered together. The score normal represents the similarity between the normal vector of each point considered when determining the projection plane in the initial segmentation process and the normal of each flat face of the bounding box, and the score smoothness indicates the similarity between the projection plane of the current point and the projection planes of neighboring points.
[0228] The score smoothness can be considered by assigning a weight to the score normal. In this case, the weight value can be defined by the user. The refinement segmentation can be repeatedly executed, and the number of repetitions can also be defined by the user.
[0229] Patch Segmentation Related to Patch Generation
[0230] Patch segmentation is a process of dividing the entire point cloud into patches (sets of neighboring points) based on the projection plane information of each point constituting the point cloud obtained in the initial / refinement segmentation process. Patch segmentation may include the following steps:
[0231] Calculate the neighboring points of each point constituting the point cloud using a K-D tree or the like. The maximum number of neighbors can be defined by the user;
[0232] ① When the neighboring points are projected onto the same plane as the current point (when they have the same cluster index), ② extract
[0233] the current point and the neighboring points as a patch;
[0234] ③ Calculate the geometric values of the extracted patch.
[0235] ④ Repeat operations ② to ③ until there are no unextracted points.
[0236] The occupancy map, geometric image, and texture image of each patch, as well as the size of each patch, are determined through the patch segmentation process.
[0237] Figure 7 An example of determining the position of each patch on the occupancy map according to an embodiment is shown.
[0238] The point cloud video encoder 10002 according to an embodiment can perform patch packing and generate an occupancy map.
[0239] Patch Packing and Occupancy Map Generation (14001)
[0240] This is a process of determining the position of each patch in a 2D image to map the segmented patches to the 2D image. As a 2D image, the occupancy map is a binary map using values 0 or 1 to indicate whether there is data at the corresponding position. The occupancy map consists of blocks, and its resolution can be determined by the size of the blocks. For example, when the block is a 1*1 block, pixel-level resolution is obtained. The occupancy packing block size can be determined by the user.
[0241] The process of determining the position of each patch on the occupancy map can be configured as follows:
[0242] ① Set all positions on the occupancy map to 0;
[0243] ② Place the patch at the point (u, v) in the occupancy map plane where the horizontal coordinate is in the range (0, occupancySizeU - patch.sizeU0) and the vertical coordinate is in the range (0, occupancySizeV - patch.sizeV0);
[0244] ③Set the point (x, y) in the patch plane where the horizontal coordinate is within the range (0, patch.sizeU0) and the vertical coordinate is within the range (0, patch.sizeV0) as the current point;
[0245] ④Change the position of the point (x, y) in raster order, and if the value of the coordinate (x, y) on the occupancy map occupied by the patch is 1 (there is data at the point in the patch) and the value of the coordinate (u + x, v + y) on the global occupancy map is 1 (the occupancy map is filled with the previous patch), then repeat operations ③ and ④. Otherwise, proceed to operation ⑥;
[0246] ⑤Change the position of (u, v) in raster order and repeat operations ③ to ⑤;
[0247] ⑥Determine (u, v) as the position of the patch and copy the occupancy map data regarding the patch to the corresponding part on the global occupancy map; and
[0248] ⑦Repeat operations ② to ⑥ for the next patch.
[0249] occupancySizeU: Indicates the width of the occupancy map. Its unit is the occupancy packing block size.
[0250] occupancySizeV: Indicates the height of the occupancy map. Its unit is the occupancy packing block size.
[0251] patch.sizeU0: Indicates the width of the occupancy map. Its unit is the occupancy packing block size.
[0252] patch.sizeV0: Indicates the height of the occupancy map. Its unit is the occupancy packing block size.
[0253] For example, as Figure 7 shown, in the box corresponding to the occupancy packing size block, there is a box corresponding to the patch with the patch size, and the point (x, y) can be located in this box.
[0254] Figure 8 Shows an exemplary relationship between the normal axis, tangential axis, and bitangential axis according to an embodiment.
[0255] The point cloud video encoder 10002 according to an embodiment can generate a geometry image. The geometry image refers to image data including geometric information about the point cloud. The geometry image generation process can adopt the three axes (normal, tangential, and bitangential) of the patch in FIG. 8.
[0256] Geometric Image Generation (14002)
[0257] In this process, depth values of the geometric images constituting each patch are determined, and the entire geometric image is generated based on the positions of the patches determined in the above patch packing process. The process of determining the depth values of the geometric images constituting each patch can be configured as follows.
[0258] ① Calculate parameters related to the positions and sizes of each patch. The parameters may include the following information. According to an embodiment, the position of the patch is included in the patch information.
[0259] The normal index indicating the normal axis is obtained in the previous patch generation process. The tangential axis is the axis among the axes perpendicular to the normal axis that coincides with the horizontal axis u of the patch image, and the bi - tangential axis is the axis among the axes perpendicular to the normal axis that coincides with the vertical axis v of the patch image. The three axes can be as Figure 8 shown.
[0260] Figure 9 Illustrates exemplary configurations of the minimum mode and the maximum mode of the projection mode according to an embodiment.
[0261] The point cloud video encoder 10002 according to an embodiment can perform patch - based projection to generate a geometric image, and the projection mode according to an embodiment includes a minimum mode and a maximum mode.
[0262] The 3D spatial coordinates of the patch can be calculated based on the bounding box around the minimum size of the patch. For example, the 3D spatial coordinates may include the minimum tangential value of the patch (on the patch 3d - shifted tangential axis), the minimum bi - tangential value of the patch (on the patch 3d - shifted bi - tangential axis), and the minimum normal value of the patch (on the patch 3d - shifted normal axis).
[0263] The 2D size of the patch indicates the horizontal size and the vertical size of the patch when the patch is packed into a 2D image. The horizontal size (patch 2d size u) can be obtained as the difference between the maximum tangential value and the minimum tangential value of the bounding box, and the vertical size (patch 2d size v) can be obtained as the difference between the maximum bi - tangential value and the minimum bi - tangential value of the bounding box.
[0264] ② Determine the projection mode of the patch. The projection mode can be the minimum mode or the maximum mode. The geometric information about the patch is represented using depth values. When each point constituting the patch is projected in the normal direction of the patch, two layers of images can be generated, an image constructed using the maximum depth value and an image constructed using the minimum depth value.
[0265] In the minimum mode, when generating the two layers of images d0 and d1, the minimum depth can be configured for d0, and the maximum depth within the surface thickness from the minimum depth can be configured for d1, as Figure 9 shown.
[0266] For example, when the point cloud is in 2D (as Figure 9When in the state shown (as shown), there are multiple patches including multiple points. As shown in the figure, the points marked with the same style of shading can belong to the same patch. The figure shows the processing of the patch that projects the points marked with blanks.
[0267] When left / right projecting the points marked with blanks, relative to the left side, the depth can be incremented by 1, being 0, 1, 2, …, 6, 7, 8, 9, and the numbers used to calculate the depth of the points can be marked on the right side.
[0268] The same projection mode can be applied to all point clouds, or different projection modes can be applied to each frame or patch according to user definition. When different projection modes are applied to each frame or patch, the projection mode that can enhance the compression efficiency or minimize the missing points can be adaptively selected.
[0269] ③ Calculate the depth values of each point.
[0270] In the minimum mode, the image d0 is constructed using depth0, where depth0 is the value obtained by subtracting the minimum normal value of the patch (on the 3D shifted normal axis of the patch) calculated in operation 1) from the minimum normal value of each point. If there is another depth value within the range between depth0 and the surface thickness at the same position, that value is set as depth1. Otherwise, the value of depth0 is assigned to depth1. The image d1 is constructed using the value of depth1.
[0271] For example, when calculating the depth of the points in the image d0, the minimum value (4 2 4 4 0 6 0 0 9 9 0 80) can be calculated. When calculating the depth of the points in the image d1, the greater value among two or more points can be calculated. When there is only one point, its value can be calculated (4 4 4 4 6 6 6 8 9 9 8 8 9). In the processing of encoding and reconstructing the points of the patch, some points may be missing (for example, in the figure, eight points are missing).
[0272] In the maximum mode, the image d0 is constructed using depth0, where depth0 is the value obtained by subtracting the minimum normal value of the patch (on the 3D shifted normal axis of the patch) calculated in operation 1) from the maximum normal value of each point. If there is another depth value within the range between depth0 and the surface thickness at the same position, that value is set as depth1. Otherwise, the value of depth0 is assigned to depth1. The image d1 is constructed using the value of depth1.
[0273] For example, when calculating the depth of a point in image d0, the minimum value (4 4 4 4 6 6 6 8 9 9 8 8 9) can be calculated. When calculating the depth of a point in image d1, the smaller value among two or more points can be calculated. When there is only one point, its value can be calculated (4 2 4 4 5 6 0 6 9 9 0 8 0). In the processing of encoding and reconstructing the points of a patch, some points may be missing (for example, in the figure, six points are missing).
[0274] The entire geometric image can be generated by placing the geometric images of the respective patches generated by the above processing onto the entire geometric image based on the patch position information determined in the patch packing process.
[0275] The layer d1 of the generated entire geometric image can be encoded using various methods. The first method (absolute d1 encoding method) is to encode the depth values of the previously generated image d1. The second method (differential encoding method) is to encode the difference between the depth values of the previously generated image d1 and the depth values of image d0.
[0276] In the encoding method using the depth values of the two layers d0 and d1 as described above, if there is another point between the two depths, the geometric information about that point is lost in the encoding process. Therefore, the enhanced - delta - depth (EDD) code can be used for lossless encoding.
[0277] Hereinafter, with reference to Figure 10 EDD code will be described in detail.
[0278] Figure 10 An exemplary EDD code according to an embodiment is shown.
[0279] In some / all processes of the point cloud video encoder 10002 and / or V - PCC encoding (such as video compression 14009), the geometric information about points can be encoded based on the EOD code.
[0280] As Figure 10 shown, the EDD code is used for binary encoding of the positions of all points within the surface thickness range including d1. For example, in Figure 10 , since there are points at the first and fourth positions on D0 and the second and third positions are empty, the points included in the second column from the left can be represented by the EDD code 0b1001 (=9). When the EDD code is encoded and transmitted together with D0, the receiving terminal can recover the geometric information about all points losslessly.
[0281] For example, when there is a point above the reference point, the value is 1. When there is no point, the value is 0. Therefore, the code can be represented based on 4 bits.
[0282] Smoothing (14004)
[0283] Smoothing is an operation for eliminating discontinuities that may appear on the patch boundary due to image quality degradation occurring during the compression process. Smoothing can be performed by the point cloud video encoder 10002 or the smoother 14004:
[0284] ① Reconstruct the point cloud from the geometric image. This operation can be the reverse of the geometric image generation described above. For example, the reverse process of encoding can be reconstructed;
[0285] ② Use a K-D tree or the like to calculate the neighboring points of each point constituting the reconstructed point cloud;
[0286] ③ Determine whether each point is located on the patch boundary. For example, when there are neighboring points with a different projection plane (cluster index) from the current point, it can be determined that the point is located on the patch boundary;
[0287] ④ If there is a point on the patch boundary, move the point to the centroid of the neighboring points (located at the average x, y, and z coordinates of the neighboring points). That is, change the geometric value. Otherwise, maintain the previous geometric value.
[0288] Figure 11 An example of recoloring based on the color values of neighboring points according to an embodiment is shown.
[0289] The point cloud video encoder 10002 or the texture image generator 14003 according to an embodiment can generate a texture image based on the recoloring.
[0290] Texture Image Generation (14003)
[0291] Similar to the above geometric image generation process, the texture image generation process includes generating the texture image of each patch and generating the entire texture image by arranging the texture images at the determined positions. However, in the operation of generating the texture image of each patch, instead of the depth value used for geometric generation, an image having the color values (e.g., R, G, and B values) of the points constituting the point cloud corresponding to the position is generated.
[0292] When estimating the color values of each point constituting the point cloud, the geometry obtained previously through the smoothing process can be used. In the smoothed point cloud, the positions of some points may have shifted relative to the original point cloud, so a recoloring process for finding the color suitable for the changed positions may be required. Recoloring can be performed using the color values of neighboring points. For example, as Figure 11 shown, the color values of the nearest neighboring points and the neighboring points can be considered to calculate the new color value.
[0293] For example, referring to Figure 11, in recoloring, a suitable color value for the changed position can be calculated based on the average of the attribute information of the nearest original points with respect to the points and / or the average of the attribute information of the nearest original points with respect to the points.
[0294] Similar to the geometric image generated with two layers d0 and d1, the texture image can also be generated with two layers t0 and t1.
[0295] Auxiliary Patch Information Compression (14005)
[0296] According to an embodiment, the point cloud video encoder 10002 or the auxiliary patch information compressor 14005 can compress the auxiliary patch information (auxiliary information about the point cloud).
[0297] The auxiliary patch information compressor 14005 compresses the auxiliary patch information generated in the above patch generation, patch packing, and geometry generation processes. The auxiliary patch information may include the following parameters:
[0298] An index (cluster index) for identifying the projection plane (normal plane);
[0299] The 3D spatial position of the patch, i.e., the minimum tangential value of the patch (on the patch 3d shifted tangential axis), the minimum bitangential value of the patch (on the patch 3d shifted bitangential axis), and the minimum normal value of the patch (on the patch 3d shifted normal axis);
[0300] The 2D spatial position and size of the patch, i.e., the horizontal size (patch 2d size u), the vertical size (patch 2d size v), the minimum horizontal value (patch 2d shifted u), and the minimum vertical value (patch 2d shifted u); and
[0301] Mapping information about each block and patch, i.e., the candidate index (when patches are sequentially set based on the 2D spatial position and size information of the patch, multiple patches can be mapped to one block in an overlapping manner. In this case, the mapped patches form a candidate list, and the candidate index indicates the sequential position of the patch whose data exists in the block) and the local patch index (an index indicating a patch existing in a frame). Table 1 shows the pseudocode representing the process of matching between the block and the patch based on the candidate list and the local patch index.
[0302] The maximum number of candidate lists can be defined by the user.
[0303] [Table 1]
[0304]
[0305] Figure 12 Shows the push - pull type background filling according to an embodiment.
[0306] Image Filling and Group Expansion (14006, 14007, 14008)
[0307] An image filler according to an embodiment can fill a space other than the patch area with meaningless supplementary data based on a push-pull background filling technique.
[0308] Image fillings 14006 and 14007 are processes of filling a space other than the patch area with meaningless data to improve compression efficiency. For image filling, pixel values in columns or rows near the boundary in the patch can be copied to fill the blank space. Alternatively, as Figure 12 shown, a push-pull background filling method can be used. According to this method, in the process of gradually reducing the resolution of the non-filled image and then increasing the resolution again, pixel values from the low-resolution image are used to fill the blank space.
[0309] Group dilation 14008 is a process of filling blank spaces in a geometric image and a texture image respectively configured with two layers d0 / d1 and t0 / t1. In this process, the blank spaces in the two layers calculated by image filling are filled with the average of the values at the same position.
[0310] Figure 13 An exemplary possible traversal order of a 4×4 block according to an embodiment is shown.
[0311] Occupancy Map Compression (14012, 14011)
[0312] An occupancy map compressor according to an embodiment can compress a previously generated occupancy map. Specifically, two methods can be used, namely: video compression for lossy compression and entropy compression for lossless compression. Video compression is described below.
[0313] Entropy compression can be performed by the following operations.
[0314] ① If a block constituting the occupancy map is fully occupied, encode 1 and repeat the same operation for the next block of the occupancy map. Otherwise, encode 0 and perform operations 2) to 5).
[0315] ② Determine the best traversal order for performing run-length encoding on the occupied pixels of the block. Figure 13 Four possible traversal orders of a 4×4 block are shown.
[0316] Figure 14 An exemplary best traversal order according to an embodiment is shown.
[0317] As described above, the entropy compressor according to an embodiment can encode (encode) the block based on Figure 14 the traversal order scheme shown.
[0318] For example, select the best traversal order with the minimum number of runs from the possible traversal orders, and encode its index. The selection is shown in the figureFigure 13 The case of the third traversal order in Figure 13 . In the shown case, the number of runs can be minimized to 2, so the third traversal order can be selected as the optimal traversal order.
[0319] ③ Encode the number of runs. In Figure 14 the example of Figure 14 , there are two runs, so 2 is encoded.
[0320] ④ Encode the occupancy of the first run. In Figure 14 the example of Figure 14 , 0 is encoded because the first run corresponds to unoccupied pixels.
[0321] ⑤ Encode the length of each run (as many as the number of runs). In Figure 14 the example of Figure 14 , the lengths 6 and 10 of the first run and the second run are encoded in sequence.
[0322] Video Compression (14009, 14010, 14011)
[0323] Video compressors 14009, 14010, 14011 according to the embodiment use a 2D video codec such as HEVC or VVC to encode a sequence of geometric images, texture images, occupancy map images, etc. generated in the above operations.
[0324] Figure 15 An exemplary 2D video / image encoder according to the embodiment is shown. According to the embodiment, the 2D video / image encoder may be referred to as an encoding device.
[0325] Denote an embodiment applying the above video compressors 14009, 14010, and 14011 Figure 15 is a schematic block diagram of a 2D video / image encoder 15000 configured to encode a video / image signal. The 2D video / image encoder 15000 may be included in the above point cloud video encoder 10002, or may be configured as an internal / external component. Figure 15 Each component of Figure 15 may correspond to software, hardware, a processor, and / or a combination thereof.
[0326] Here, the input image may include one of the above geometric image, texture image (attribute image), and occupancy map image. When Figure 15 the 2D video / image encoder of Figure 15 is applied to video compressor 14009, the image input to the 2D video / image encoder 15000 is a filled geometric image, and the bitstream output from the 2D video / image encoder 15000 is the bitstream of the compressed geometric image. When Figure 15When the 2D video / image encoder is applied to the video compressor 14010, the image input to the 2D video / image encoder 15000 is a padded texture image, and the bitstream output from the 2D video / image encoder 15000 is the bitstream of the compressed texture image. When Figure 15 the 2D video / image encoder is applied to the video compressor 14011, the image input to the 2D video / image encoder 15000 is an occupancy map image, and the bitstream output from the 2D video / image encoder 15000 is the bitstream of the compressed occupancy map image.
[0327] The inter-frame predictor 15090 and the intra-frame predictor 15100 may be collectively referred to as a predictor. That is, the predictor may include the inter-frame predictor 15090 and the intra-frame predictor 15100. The transformer 15030, the quantizer 15040, the inverse quantizer 15050, and the inverse transformer 15060 may be collectively referred to as a residual processor. The residual processor may further include a subtractor 15020. According to an embodiment, Figure 15 the image splitter 15010, the subtractor 15020, the transformer 15030, the quantizer 15040, the inverse quantizer 15050, the inverse transformer 15060, the adder 15200, the filter 15070, the inter-frame predictor 15090, the intra-frame predictor 15100, and the entropy encoder 15110 may be configured by one hardware component (e.g., an encoder or a processor). Additionally, the memory 15080 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium.
[0328] The image splitter 15010 can split an image (or a picture or a frame) input to the encoder 15000 into one or more processing units. For example, the processing unit can be referred to as a coding unit (CU). In this case, the CU can be recursively split from a coding tree unit (CTU) or a largest coding unit (LCU) according to a quadtree binary tree (QTBT) structure. For example, a CU can be split into multiple CUs of a lower depth based on a quadtree structure and / or a binary tree structure. In this case, for example, the quadtree structure can be applied first, and the binary tree structure can be applied later. Alternatively, the binary tree structure can be applied first. The encoding process according to the present disclosure can be performed based on the final CU that is no longer split. In this case, based on the characteristics of the image and the encoding efficiency, the LCU can be used as the final CU. If necessary, the CU can be recursively split into CUs of a lower depth, and the CU of the optimal size can be used as the final CU. Here, the encoding process can include prediction, transformation, and reconstruction (which will be described later). As another example, the processing unit can also include a prediction unit (PU) or a transformation unit (TU). In this case, the PU and the TU can be split or divided from the above-mentioned final CU. The PU can be a unit for sample prediction, and the TU can be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0329] The term "unit" can be used interchangeably with terms such as block or region or module. In general, an M×N block can represent a set of samples or transform coefficients configured in M columns and N rows. Samples can generally represent pixels or pixel values, and can indicate only the pixels / pixel values of the luminance component, or only the pixels / pixel values of the chrominance component. "Sample" can be used as a term corresponding to a pixel or a pel in a picture (or an image).
[0330] The subtractor 15020 of the encoder 15000 can generate a residual signal (residual block or residual sample array) by subtracting the prediction signal (prediction block or prediction sample array) output from the inter-frame predictor 15090 or the intra-frame predictor 15100 from the input image signal (original block or original sample array), and the generated residual signal is sent to the transformer 15030. In this case, as shown in the figure, the unit that subtracts the prediction signal (prediction block or prediction sample array) from the input image signal (original block or original sample array) in the encoder 15000 can be referred to as the subtractor 15020. The predictor can perform prediction on the processing target block (hereinafter referred to as the current block) and generate a prediction block including the prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or the CU. As will be described later in the description of each prediction mode, the predictor can generate various types of information about the prediction (for example, prediction mode information), and transmit the generated information to the entropy encoder 15110. The information about the prediction can be encoded by the entropy encoder 15110 and output in the form of a bitstream.
[0331] The intra-predictor 15100 of the predictor may predict the current block with reference to samples in the current picture. Depending on the prediction mode, the samples may be adjacent to or distant from the current block. In intra-prediction, the prediction mode may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, the DC mode and the planar mode. Depending on the fineness of the prediction direction, the directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is merely an example, and more or fewer directional prediction modes may be used according to the settings. The intra-predictor 15100 may determine the prediction mode to be applied to the current block based on the prediction mode applied to adjacent blocks.
[0332] The inter-predictor 15090 of the predictor may derive the prediction block of the current block based on the reference block (reference sample array) specified by the motion vector on the reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter-prediction mode, the motion information may be predicted based on the correlation of the motion information between adjacent blocks on a per-block, sub-block, or sample basis. The motion information may include a motion vector and a reference picture index. The motion information may also include information regarding the inter-prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-prediction, the adjacent blocks may include spatial adjacent blocks present in the current picture and temporal adjacent blocks present in the reference picture. The reference picture including the reference block may be the same as or different from the reference picture including the temporal adjacent blocks. The temporal adjacent blocks may be referred to as co-located reference blocks or co-located CUs (colCUs), and the reference picture including the temporal adjacent blocks may be referred to as a co-located picture (colPic). For example, the inter-predictor 15090 may configure a motion information candidate list based on adjacent blocks and generate information indicating candidates for the motion vector and / or reference picture index to be used to derive the current block. Inter-prediction may be performed based on various prediction modes. For example, in the skip mode and the merge mode, the inter-predictor 15090 may use the motion information regarding adjacent blocks as the motion information regarding the current block. In the skip mode, unlike the merge mode, the residual signal may not be transmitted. In the motion vector prediction (MVP) mode, the motion vector of the adjacent block may be used as the motion vector predictor, and the motion vector difference may be signaled to indicate the motion vector of the current block.
[0333] The prediction signal generated by the inter-predictor 15090 or the intra-predictor 15100 may be used to generate a reconstructed signal or a residual signal.
[0334] The transformer 15030 can generate transform coefficients by applying a transformation technique to the residual signal. For example, the transformation technique can include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen–Loève transform (KLT), a graph-based transform (GBT), or a conditional non-linear transform (CNT). Here, GBT refers to a transform obtained from a graph depicting the relationships between pixels. CNT refers to a transform obtained based on a prediction signal generated according to all previously reconstructed pixels. Additionally, the transform operation can be applied to square pixel blocks of the same size, or can be applied to blocks of variable size other than square.
[0335] The quantizer 15040 can quantize the transform coefficients and send them to the entropy encoder 15110. The entropy encoder 15110 can encode the quantized signal (information about the quantized transform coefficients) and output a bitstream of the encoded signal. The information about the quantized transform coefficients can be referred to as residual information. The quantizer 15040 can rearrange the quantized transform coefficients in block form into a one-dimensional vector based on the coefficient scan order, and generate information about the quantized transform coefficients based on the quantized transform coefficients in one-dimensional vector form.
[0336] The entropy encoder 15110 can employ various coding techniques such as, for example, exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoder 15110 can encode information required for video / image reconstruction (e.g., the values of syntax elements) together with or separately from the quantized transform coefficients. The encoded information (e.g., encoded video / image information) can be sent or stored in the form of a bitstream based on network abstraction layer (NAL) units.
[0337] The bitstream can be sent via a network or can be stored in a digital storage medium. Here, the network can include a broadcast network and / or a communication network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. A transmitter (not shown) for sending the signal output from the entropy encoder 15110 and / or a storage unit (not shown) for storing the signal can be configured as internal / external elements of the encoder 15000. Alternatively, the transmitter can be included in the entropy encoder 15110.
[0338] The quantized transform coefficients output from the quantizer 15040 can be used to generate a prediction signal. For example, the inverse quantizer 15050 and the inverse transformer 15060 can apply inverse quantization and inverse transformation to the quantized transform coefficients to reconstruct the residual signal (residual block or residual sample). The adder 15200 adds the reconstructed residual signal to the prediction signal output from the inter-frame predictor 15090 or the intra-frame predictor 15100. Thus, a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) can be generated. When there is no residual signal for the processing target block as in the case of applying the skip mode, the prediction block can be used as the reconstructed block. The adder 15200 can be referred to as a reconstructor or a reconstructed block generator. As described below, the generated reconstructed signal can be used for intra-frame prediction of the next processing target block in the current picture, or can be used for inter-frame prediction of the next picture through filtering.
[0339] The filter 15070 can improve the subjective / objective image quality by applying filtering to the reconstructed signal output from the adder 15200. For example, the filter 15070 can generate a modified reconstructed picture by applying various filtering techniques to the reconstructed picture, and the modified reconstructed picture can be stored in the memory 15080 (specifically, the DPB of the memory 15080). For example, various filtering techniques can include deblocking filtering, sample adaptive offset, adaptive loop filtering, and bilateral filtering. As described below in the description of the filtering technique, the filter 15070 can generate various types of information about the filtering and transmit the generated information to the entropy encoder 15110. The information about the filtering can be encoded by the entropy encoder 15110 and output in the form of a bitstream.
[0340] The modified reconstructed picture stored in the memory 15080 can be used by the inter-frame predictor 15090 as a reference picture. Therefore, when applying inter-frame prediction, the encoder can avoid prediction mismatch between the encoder 15000 and the decoder and improve the coding efficiency.
[0341] The DPB of the memory 15080 can store the modified reconstructed picture to be used by the inter-frame predictor 15090 as a reference picture. The memory 15080 can store the motion information of the block for deriving (or encoding) the motion information in the current picture and / or the motion information of the blocks that have been reconstructed in the picture. The stored motion information can be transmitted to the inter-frame predictor 15090 to be used as the motion information about spatially neighboring blocks or temporally neighboring blocks. The memory 15080 can store the reconstructed samples of the reconstructed blocks in the current picture and transmit the reconstructed samples to the intra-frame predictor 15100.
[0342] At least one of the above prediction, transformation, and quantization processes can be skipped. For example, for a block to which pulse code modulation (PCM) is applied, the prediction, transformation, and quantization processes can be skipped, and the values of the original samples can be encoded and output in the form of a bitstream.
[0343] Figure 16 Shows an exemplary V-PCC decoding process according to an embodiment.
[0344] The V-PCC decoding process or V-PCC decoder can follow Figure 4 the reverse process of the V-PCC encoding process (or encoder). Figure 16 Each component of
[0345] The demultiplexer 16000 demultiplexes the compressed bitstream to output the compressed texture image, compressed geometry image, compressed occupancy map, and compressed auxiliary patch information, respectively.
[0346] The video decompressor or video decompressors 16001, 16002 decompress each of the compressed texture image and the compressed geometry image.
[0347] The occupancy map decompressor or occupancy map decompressor 16003 decompresses the compressed occupancy map image.
[0348] The auxiliary patch information decompressor or auxiliary patch information decompressor 16004 decompresses the compressed auxiliary patch information.
[0349] The geometry reconstructor or geometry reconstructor 16005 restores (reconstructs) the geometry information based on the decompressed geometry image, decompressed occupancy map, and / or decompressed auxiliary patch information. For example, the geometry changed in the encoding process can be reconstructed.
[0350] The smoother or smoother 16006 can apply smoothing to the reconstructed geometry. For example, smoothing filtering can be applied.
[0351] The texture reconstructor or texture reconstructor 16007 reconstructs the texture from the decompressed texture image and / or the smoothed geometry.
[0352] The color smoother or color smoother 16008 smooths the color values from the reconstructed texture. For example, smoothing filtering can be applied.
[0353] As a result, the reconstructed point cloud data can be generated.
[0354] Figure 16 Shows the decoding process of V-PCC for reconstructing the point cloud by decompressing (decoding) the compressed occupancy map, geometry image, texture image, and auxiliary patch information.
[0355] Figure 16 Each of the units shown can be used as at least one of a processor, software, and hardware. According to an embodiment, Figure 16 a detailed operation description of each unit is as follows.
[0356] Video Decompression (16001, 16002)
[0357] Video decompression is the reverse process of the above video compression. It is a process of using a 2D video codec such as HEVC or VVC to decode the bitstream of the geometric image, the bitstream of the compressed texture image, and / or the bitstream of the compressed occupancy map image generated in the above process.
[0358] Figure 17 An exemplary 2D video / image decoder according to an embodiment is also referred to as a decoding device.
[0359] The 2D video / image decoder can follow Figure 15 the reverse process of the operation of the 2D video / image encoder of
[0360] Figure 17 The 2D video / image decoder of Figure 16 is an embodiment of the video decompressors 16001 and 16002 of Figure 17 is a schematic block diagram of a 2D video / image decoder 17000 that decodes a video / image signal. The 2D video / image decoder 17000 can be included in the above point cloud video decoder 10008, or can be configured as an internal / external component. Figure 17 Each component of
[0361] Here, the input bitstream can be one of the bitstream of the geometric image, the bitstream of the texture image (attribute image), and the bitstream of the occupancy map image. When Figure 17 the 2D video / image decoder of Figure 17 is applied to the video decompressor 16001, the bitstream input to the 2D video / image decoder is the bitstream of the compressed texture image, and the reconstructed image output from the 2D video / image decoder is the decompressed texture image. When Figure 17 the 2D video / image decoder of
[0362] Reference Figure 17 The inter-frame predictor 17070 and the intra-frame predictor 17080 may be collectively referred to as a predictor. That is, the predictor may include the inter-frame predictor 17070 and the intra-frame predictor 17080. The inverse quantizer 17020 and the inverse transformer 17030 may be collectively referred to as a residual processor. That is, according to an embodiment, the residual processor may include the inverse quantizer 17020 and the inverse transformer 17030. Figure 17 The entropy decoder 17010, the inverse quantizer 17020, the inverse transformer 17030, the adder 17040, the filter 17050, the inter-frame predictor 17070, and the intra-frame predictor 17080 of Figure 17 may be configured by one hardware component (e.g., a decoder or a processor). Additionally, the memory 17060 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium.
[0363] When a bitstream containing video / image information is input, the decoder 17000 may reconstruct an image in a process corresponding to the process in which the encoder of Figure 15 processes the video / image information. For example, the decoder 17000 may use the processing units applied in the encoder to perform decoding. Thus, the decoding processing unit may be, for example, a CU. The CU may be divided from a CTU or an LCU along a quadtree structure and / or a binary tree structure. Then, the reconstructed video signal decoded and output by the decoder 17000 may be played by a player.
[0364] The decoder 17000 may receive a signal output from the encoder in the form of a bitstream, and the received signal may be decoded by the entropy decoder 17010. For example, the entropy decoder 17010 may parse the bitstream to derive information (e.g., video / image information) required for image reconstruction (or picture reconstruction). For example, the entropy decoder 17010 may decode the information in the bitstream based on an encoding technique such as exponential Golomb coding, CAVLC, or CABAC, and output the syntax elements required for image reconstruction and the quantization values of the transform coefficients of the residuals. More specifically, in CABAC entropy decoding, bins corresponding to each syntax element in the bitstream may be received, and a context model may be determined based on the decoding target syntax element information and the decoding information about the neighboring decoding target blocks or the information about the symbols / bins decoded in the previous step. Then, the probability of the bin occurrence may be predicted according to the determined context model, and arithmetic decoding of the bin may be performed to generate a symbol corresponding to the value of each syntax element. According to CABAC entropy decoding, after the context model is determined, the context model may be updated based on the information about the symbols / bins decoded for the context model of the next symbol / bin. The information about prediction in the information decoded by the entropy decoder 17010 may be provided to the predictors (inter-frame predictor 17070 and intra-frame predictor 17080), and the residual values (i.e., the quantized transform coefficients and related parameter information) that have been entropy decoded by the entropy decoder 17010 may be input to the inverse quantizer 17020. In addition, the information about filtering in the information decoded by the entropy decoder 17010 may be provided to the filter 17050. A receiver (not shown) configured to receive the signal output from the encoder may also be configured as an internal / external component of the decoder 17000. Alternatively, the receiver may be a component of the entropy decoder 17010.
[0365] The inverse quantizer 17020 may output transform coefficients by inverse quantizing the quantized transform coefficients. The inverse quantizer 17020 may rearrange the quantized transform coefficients in the form of two-dimensional blocks. In this case, the rearrangement may be performed based on the coefficient scan order implemented by the encoder. The inverse quantizer 17020 may perform inverse quantization on the quantized transform coefficients using a quantization parameter (e.g., quantization step information) and obtain the transform coefficients.
[0366] The inverse transformer 17030 obtains a residual signal (residual block and residual sample array) by transforming the transform coefficients.
[0367] The predictor may perform prediction on the current block and generate a prediction block including the prediction samples of the current block. The predictor may determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on the information about prediction output from the entropy decoder 17010, and may determine a specific intra-frame / inter-frame prediction mode.
[0368] The intra predictor 17080 of the predictor may predict the current block by referring to samples in the current picture. Depending on the prediction mode, the samples may be adjacent to or far from the current block. In intra prediction, the prediction mode may include a plurality of non - directional modes and a plurality of directional modes. The intra predictor 17080 may use the prediction mode applied to adjacent blocks to determine the prediction mode applied to the current block.
[0369] The inter predictor 17070 of the predictor may derive the prediction block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter - prediction mode, the motion information may be predicted based on the correlation of the motion information between adjacent blocks and the current block on a per - block, sub - block, or sample basis. The motion information may include a motion vector and a reference picture index. The motion information may also include information about the inter - prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter - prediction, adjacent blocks may include spatial adjacent blocks present in the current picture and temporal adjacent blocks present in the reference picture. For example, the inter predictor 17070 may configure a motion information candidate list based on adjacent blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. The inter - prediction may be performed based on various prediction modes. Information about the prediction may include information indicating the inter - prediction mode of the current block.
[0370] The adder 17040 may add the residual signal obtained in the inverse transformer 17030 to the prediction signal (prediction block or prediction sample array) output from the inter predictor 17070 or the intra predictor 17080, thereby generating a reconstructed signal (reconstructed picture, reconstructed block, or reconstructed sample array). When there is no residual signal for processing the target block as in the case of applying the skip mode, the prediction block may be used as the reconstructed block.
[0371] The adder 17040 may be referred to as a reconstructor or a reconstructed - block generator. The generated reconstructed signal may be used for intra - prediction of the next processing target block in the current picture, or may be used for inter - prediction of the next picture through filtering as described below.
[0372] The filter 17050 may improve the subjective / objective image quality by applying filtering to the reconstructed signal output from the adder 17040. For example, the filter 17050 may generate a modified reconstructed picture by applying various filtering techniques to the reconstructed picture and send the modified reconstructed picture to the memory 17060 (specifically, the DPB of the memory 17060). For example, various filtering methods may include de - blocking filtering, sample - adaptive offset, adaptive loop filtering, and bilateral filtering.
[0373] The reconstructed picture stored in the DPB of the memory 17060 can be used as a reference picture in the inter - predictor 17070. The memory 17060 can store motion information of blocks for which motion information is derived (or decoded) in the current picture and / or motion information of blocks in the already - reconstructed pictures. The stored motion information can be transmitted to the inter - predictor 17070 to be used as motion information for spatially - adjacent blocks or motion information for temporally - adjacent blocks. The memory 17060 can store the reconstructed samples of the reconstructed blocks in the current picture and transmit the reconstructed samples to the intra - predictor 17080.
[0374] In the present disclosure, the embodiments described with respect to Figure 15 the filter 15070, the inter - predictor 15090, and the intra - predictor 15100 of the encoder 15000 can be applied to the filter 17050, the inter - predictor 17070, and the intra - predictor 17080 of the decoder 17000, respectively, in the same or corresponding manner.
[0375] At least one of the above - mentioned prediction, inverse - transform, and inverse - quantization processes can be skipped. For example, for a block to which pulse - code modulation (PCM) is applied, the prediction, inverse - transform, and inverse - quantization processes can be skipped, and the values of the decoded samples can be used as the samples of the reconstructed image.
[0376] Occupancy Map Decompression (16003)
[0377] This is the inverse process of the above - mentioned occupancy - map compression. Occupancy - map decompression is a process of reconstructing an occupancy map by decompressing an occupancy - map bitstream.
[0378] Auxiliary Patch Information Decompression (16004)
[0379] The auxiliary patch information can be reconstructed by performing the inverse process of the above - mentioned auxiliary - patch - information compression and decoding the compressed auxiliary - patch - information bitstream.
[0380] Geometric Reconstruction (16005)
[0381] This is the inverse process of the above - mentioned geometric - image generation. Initially, patches are extracted from the geometric image using the reconstructed occupancy map, the 2D position / size information of the patches included in the auxiliary patch information, and the information on the mapping between the blocks and the patches. Then, a point cloud is reconstructed in 3D space based on the geometric image of the extracted patches and the 3D position information of the patches included in the auxiliary patch information. When the geometric value corresponding to the point (u, v) within the patch is g(u, v), and the position coordinates of the patch on the normal axis, tangential axis, and binormal axis in 3D space are (δ0, s0, r0), the normal, tangential, and binormal coordinates δ(u, v), s(u, v), and r(u, v) of the position in 3D space mapped to the point (u, v) can be expressed as follows.
[0382] δ(u, v) = δ0 + g(u, v)
[0383] s(u, v) = s0 + u
[0384] r(u, v) = r0 + v
[0385] Smoothing (16006)
[0386] Similar to the smoothing in the above encoding process, smoothing is a process for eliminating discontinuities that may appear at patch boundaries due to image quality degradation occurring during the compression process.
[0387] Texture Reconstruction (16007)
[0388] Texture reconstruction is a process of reconstructing a color point cloud by assigning color values to each point constituting the smoothed point cloud. This can be performed by assigning color values corresponding to texture image pixels at the same position in the 2D space geometric image to the points in the point cloud corresponding to the same position in the 3D space based on the mapping information between the geometric image and the point cloud reconstructed in the above geometric reconstruction process.
[0389] Color Smoothing (16008)
[0390] Color smoothing is similar to the above geometric smoothing process. Color smoothing is a process for eliminating discontinuities that may appear at patch boundaries due to image quality degradation occurring during the compression process. Color smoothing can be performed by the following operations:
[0391] ① Calculate the neighboring points of each point constituting the reconstructed point cloud using a K - D tree or the like. The neighboring point information calculated in the above geometric smoothing process can be used.
[0392] ② Determine whether each point is located on the patch boundary. These operations can be performed based on the boundary information calculated in the above geometric smoothing process.
[0393] ③ Check the distribution of the color values of the neighboring points of the points existing on the boundary and determine whether to perform smoothing. For example, when the entropy of the luminance value is less than or equal to a threshold local entry (there are many similar luminance values), it can be determined that the corresponding part is not an edge part, and smoothing can be performed. As a smoothing method, the color value of the point can be replaced with the average of the color values of the neighboring points.
[0394] Figure 18 It is a flowchart showing the operations of a transmitting device for compressing and transmitting V - PCC - based point cloud data according to an embodiment of the present disclosure.
[0395] The transmitting device according to the embodiment can correspond to Figure 1 the transmitting device of Figure 4encoding process, and Figure 15 a 2D video / image encoder, or some / all of its operations. Each component of the transmitting device may correspond to software, hardware, a processor, and / or a combination thereof.
[0396] The operation of the transmitting terminal using V-PCC to compress and transmit point cloud data can be performed as shown in the figure.
[0397] The point cloud data transmitting device according to the embodiment may be referred to as a transmitting device or a transmitting system.
[0398] Regarding the patch generator 18000, patches for 2D image mapping of the point cloud are generated based on the input point cloud data. As a result of patch generation, patch information and / or auxiliary patch information are generated. The generated patch information and / or auxiliary patch information can be used in geometric image generation, texture image generation, smoothing, and processing for geometric reconstruction for smoothing.
[0399] The patch packer 18001 performs a patch packing process of mapping the patches generated by the patch generator 18000 into a 2D image. For example, one or more patches can be packed. As a result of patch packing, an occupancy map can be generated. The occupancy map can be used in geometric image generation, geometric image filling, texture image filling, and / or processing for geometric reconstruction for smoothing.
[0400] The geometric image generator 18002 generates a geometric image based on the point cloud data, patch information (or auxiliary patch information), and / or the occupancy map. The generated geometric image is preprocessed by the encoding preprocessor 18003 and then encoded into a bitstream by the video encoder 18006.
[0401] The encoding preprocessor 18003 may include an image filling process. In other words, some spaces in the generated geometric image and the generated texture image can be filled with meaningless data. The encoding preprocessor 18003 may also include a group expansion process for the generated texture image or the texture image for which image filling has been performed.
[0402] The geometric reconstructor 18010 reconstructs a 3D geometric image based on the geometric bitstream encoded by the video encoder 18006, auxiliary patch information, and / or the occupancy map.
[0403] The smoother 18009 smooths the 3D geometric image reconstructed and output by the geometric reconstructor 18010 based on the auxiliary patch information and outputs the smoothed 3D geometric image to the texture image generator 18004.
[0404] The texture image generator 18004 can generate a texture image based on smooth 3D geometry, point cloud data, patches (or packed patches), patch information (or auxiliary patch information), and / or occupancy maps. The generated texture image can be preprocessed by the encoding preprocessor 18003 and then encoded into a video bitstream by the video encoder 18006.
[0405] The metadata encoder 18005 can encode the auxiliary patch information into a metadata bitstream.
[0406] The video encoder 18006 can encode the geometry image and the texture image output from the encoding preprocessor 18003 into corresponding video bitstreams, and can encode the occupancy map into a video bitstream. According to an embodiment, the video encoder 18006 encodes each input image by applying Figure 15 a 2D video / image encoder.
[0407] The multiplexer 18007 multiplexes the video bitstream of the geometry, the video bitstream of the texture image, the video bitstream of the occupancy map, and the bitstream of the metadata (including auxiliary patch information) output from the metadata encoder 18005 into one bitstream.
[0408] The transmitter 18008 sends the bitstream output from the multiplexer 18007 to the receiving side. Alternatively, a file / fragment encapsulator may also be provided between the multiplexer 18007 and the transmitter 18008, and the bitstream output from the multiplexer 18007 can be encapsulated in the form of a file and / or a fragment and output to the transmitter 18008.
[0409] Figure 18 The patch generator 18000, the patch packer 18001, the geometry image generator 18002, the texture image generator 18004, the metadata encoder 18005, and the smoother 18009 of can respectively correspond to patch generation 14000, patch packing 14001, geometry image generation 14002, texture image generation 14003, auxiliary patch information compression 14005, and smoothing 14004. Figure 18 The encoding preprocessor 18003 of can include Figure 4 image fillers 14006 and 14007 and a group expander 14008 of, and Figure 18 the video encoder 18006 of can include Figure 4 video compressors 14009, 14010, and 14011 and / or an entropy compressor 14012 of. For parts not described with reference to Figure 18 refer to the description of Figures 4 to 15 The above blocks may be omitted or may be replaced by blocks with similar or identical functions. In addition,Figure 18 Each of the blocks shown can be used as at least one of a processor, software, or hardware. Alternatively, the generated geometric, texture image, occupancy map video bitstream, and metadata bitstream of the auxiliary patch information can be formed into one or more track data in a file, or encapsulated into segments, and sent to the receiving side through a transmitter.
[0410] Process of Operating the Receiving Device
[0411] Figure 19 is a flowchart showing the operation of a receiving device for receiving and recovering V-PCC-based point cloud data according to an embodiment.
[0412] The receiving device according to an embodiment may correspond to Figure 1 the receiving device of Figure 16 the decoding process of Figure 17 the 2D video / image encoder of , or perform some / all of its operations. Each component of the receiving device may correspond to software, hardware, a processor, and / or a combination thereof.
[0413] The operation of the receiving terminal for receiving and reconstructing point cloud data using V-PCC can be performed as shown in the figure. The operation of the V-PCC receiving terminal can follow Figure 18 the reverse process of the operation of the V-PCC transmitting terminal of .
[0414] The point cloud data receiving device according to an embodiment may be referred to as a receiving device, a receiving system, etc.
[0415] The receiver receives the bitstream of the point cloud (i.e., the compressed bitstream), and the demultiplexer 19000 demultiplexes the bitstream of the texture image, the bitstream of the geometric image, the bitstream of the occupancy map image, and the bitstream of the metadata (i.e., the auxiliary patch information) from the received point cloud bitstream. The demultiplexed bitstreams of the texture image, the geometric image, and the occupancy map image are output to the video decoder 19001, and the bitstream of the metadata is output to the metadata decoder 19002.
[0416] According to an embodiment, when Figure 18 the transmitting device of is provided with a file / fragment encapsulator, the file / fragment de-encapsulator is provided between Figure 19 the receiver of the receiving device of and the demultiplexer 19000. In this case, the transmitting device encapsulates and sends the point cloud bitstream in the form of a file and / or a segment, and the receiving device receives and de-encapsulates the file and / or the segment containing the point cloud bitstream.
[0417] Video decoder 19001 decodes the bitstreams of the geometry image, the texture image, and the occupancy map image into a geometry image, a texture image, and an occupancy map image respectively. According to an embodiment, video decoder 19001 performs a decoding operation on each input bitstream by applying Figure 17 's 2D video / image decoder. Metadata decoder 19002 decodes the bitstream of metadata into auxiliary patch information and outputs this information to geometry reconstructor 19003.
[0418] Geometry reconstructor 19003 reconstructs 3D geometry based on the geometry image, occupancy map, and / or auxiliary patch information output from video decoder 19001 and metadata decoder 19002.
[0419] Smoother 19004 smooths the 3D geometry reconstructed by geometry reconstructor 19003.
[0420] Texture reconstructor 19005 reconstructs a texture using the texture image output from video decoder 19001 and / or the smoothed 3D geometry. That is, texture reconstructor 19005 reconstructs a color point cloud image / view by assigning color values to the smoothed 3D geometry using the texture image. Thereafter, to improve the objective / subjective visual quality, additional color smoothing processing can be performed on the color point cloud image / view by color smoother 19006. After the rendering process in point cloud renderer 19007, the modified point cloud image / view derived through the above operations is displayed to the user. In some cases, the color smoothing process can be omitted.
[0421] The above blocks can be omitted or can be replaced by blocks with similar or identical functions. In addition, Figure 19 each of the blocks shown can be used as at least one of a processor, software, and hardware.
[0422] Figure 20 An exemplary architecture for V-PCC-based storage and streaming of point cloud data according to an embodiment is shown.
[0423] Figure 20 Some or all of the system of Figure 1 can include a transmitting device and a receiving device of Figure 4 's encoding process, Figure 15 's 2D video / image encoder, Figure 16 's decoding process, Figure 18 's transmitting device and / or Figure 19 's receiving device. Each component in the figure can correspond to software, hardware, a processor, and / or a combination thereof.
[0424] Figure 20Shows an overall architecture for storing or streaming point cloud data compressed according to video-based point cloud compression (V-PCC). The processing of storing and streaming point cloud data may include acquisition processing, encoding processing, transmission processing, decoding processing, rendering processing, and / or feedback processing.
[0425] Embodiments propose a method for effectively providing point cloud media / content / data.
[0426] To effectively provide point cloud media / content / data, a point cloud acquirer 20000 may acquire a point cloud video. For example, one or more cameras may acquire point cloud data by capturing, arranging, or generating point clouds. Through this acquisition processing, a point cloud video including the 3D positions of individual points (which may be represented by x, y, and z position values, etc.) (hereinafter referred to as geometry) and the attributes of individual points (color, reflectivity, transparency, etc.) may be acquired. For example, a file in the Polygon file format (PLY) (or Stanford triangle format) containing the point cloud video may be generated. For point cloud data having multiple frames, one or more files may be acquired. Herein, point cloud-related metadata (e.g., metadata related to capture, etc.) may be generated.
[0427] Post-processing may be required for the captured point cloud video to improve the quality of the content. In the video capture processing, the maximum / minimum depth may be adjusted within the range provided by the camera device. Even after the adjustment, there may still be point data in unwanted regions. Therefore, post-processing for removing unwanted regions (e.g., background) or identifying connected spaces and filling spatial holes may be performed. Additionally, point clouds extracted from cameras sharing a spatial coordinate system may be integrated into a single piece of content through a process of transforming individual points to a global coordinate system based on the position coordinates of each camera obtained through calibration processing. Thus, a point cloud video with a high density of points may be acquired.
[0428] The point cloud pre-processor 20001 may generate one or more frames of a point cloud video. Generally, a frame may be a unit representing an image at a specific time interval. In addition, when dividing the points constituting the point cloud video into one or more patches and mapping them to a 2D plane, the point cloud pre-processor 20001 may generate an occupancy map frame with values of 0 or 1, which is a binary map indicating the presence or absence of data at corresponding positions in the 2D plane. Here, a patch is a set of points constituting the point cloud video. In this document, points belonging to the same patch are adjacent to each other in 3D space and are mapped to the same face among the flat faces of a 6-sided bounding box when mapped to a 2D image. In addition, the point cloud pre-processor 20001 may generate a geometric frame in the form of a depth map that represents, patch by patch, information about the positions (geometry) of the individual points constituting the point cloud video. The point cloud pre-processor 20001 may also generate a texture frame that represents, patch by patch, color information about the individual points constituting the point cloud video. Herein, metadata required to reconstruct the point cloud from each patch may be generated. The metadata may include information about the patches (auxiliary information or auxiliary patch information), such as the positions and sizes of the individual patches in 2D / 3D space. These frames may be continuously generated in chronological order to construct a video stream or a metadata stream.
[0429] The point cloud video encoder 20002 may encode one or more video streams related to a point cloud video. A video may include multiple frames, and one frame may correspond to a still image / frame. In the present disclosure, a point cloud video may include point cloud images / frames / frames, and the term "point cloud video" may be used interchangeably with point cloud video / frame / frame. The point cloud video encoder 20002 may perform video-based point cloud compression (V-PCC) processing. For compression and encoding efficiency, the point cloud video encoder 20002 may perform a series of processes such as prediction, transformation, quantization, and entropy encoding. The encoded data (encoded video / image information) may be output in the form of a bitstream. Based on the V-PCC processing, as described below, the point cloud video encoder 20002 may encode a point cloud video by dividing the point cloud video into a geometric video, an attribute video, an occupancy map video, and metadata (e.g., information about patches). The geometric video may include geometric images, the attribute video may include attribute images, and the occupancy map video may include occupancy map images. Patch data as auxiliary information may include patch-related information. The attribute video / image may include texture video / image.
[0430] The point cloud image encoder 20003 can encode one or more images related to a point cloud video. The point cloud image encoder 20003 can perform video-based point cloud compression (V-PCC) processing. For compression and encoding efficiency, the point cloud image encoder 20003 can perform a series of processes such as prediction, transformation, quantization, and entropy encoding. The encoded images can be output in the form of a bitstream. Based on the V-PCC processing, as described below, the point cloud image encoder 20003 can encode a point cloud image by dividing the point cloud image into a geometry image, an attribute image, an occupancy map image, and metadata (e.g., information about patches).
[0431] According to an embodiment, the point cloud video encoder 20002, the point cloud image encoder 20003, the point cloud video decoder 20006, and the point cloud image decoder 20008 can be executed by one encoder / decoder as described above and can be executed along separate paths as shown in the figure.
[0432] In the file / fragment encapsulator 20004, the encoded point cloud data and / or point cloud-related metadata can be encapsulated as a file or a fragment for streaming. Here, the point cloud-related metadata can be received from a metadata processor (not shown) or the like. The metadata processor can be included in the point cloud video / image encoder 20002 / 20003 or can be configured as a separate component / module. The file / fragment encapsulator 20004 can encapsulate the corresponding video / image / metadata in a file format such as ISOBMFF or in the form of DASH fragments. According to an embodiment, the file / fragment encapsulator 20004 can include point cloud metadata in the file format. The point cloud-related metadata can be included in various levels of boxes in, for example, the ISOBMFF file format or as data in a separate track within the file. According to an embodiment, the file / fragment encapsulator 20004 can encapsulate the point cloud-related metadata into a file.
[0433] The file / fragment encapsulator 20004 according to an embodiment can store one bitstream or separate bitstreams into one or more tracks in a file and can also encapsulate signaling information for this operation. In addition, an atlas stream (or patch stream) included in the bitstream can be stored as a track in the file, and related signaling information can be stored. In addition, SEI messages present in the bitstream can be stored in a track of the file, and related signaling information can be stored.
[0434] A sending processor (not shown) may perform transmission processing of the encapsulated point cloud data according to the file format. The sending processor may be included in a transmitter (not shown) or may be configured as a separate component / module. The sending processor may process the point cloud data according to the transmission protocol. The transmission processing may include processing for transmission via a broadcast network and processing for transmission via broadband. According to an embodiment, the sending processor may receive point cloud related metadata and point cloud data from the metadata processor and perform transmission processing of the point cloud video data.
[0435] The transmitter may send a point cloud bitstream or a file / fragment including the bitstream to a receiver (not shown) of the receiving device via a digital storage medium or a network. For transmission, processing according to any transmission protocol may be performed. The data processed for transmission may be transmitted via a broadcast network and / or via broadband. The data may be transmitted to the receiving side on a demand basis. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitter may include elements for generating a media file in a predetermined file format and may include elements for transmission via a broadcast / communication network. The receiver may extract the bitstream and send the extracted bitstream to the decoder.
[0436] The receiver may receive the point cloud data sent by the point cloud data sending device according to the present disclosure. According to the transmission channel, the receiver may receive the point cloud data via a broadcast network or via broadband. Alternatively, the point cloud data may be received via a digital storage medium. The receiver may include processing for decoding the received data and rendering the data according to the user's viewport.
[0437] A receiving processor (not shown) may perform processing on the received point cloud video data according to the transmission protocol. The receiving processor may be included in the receiver or may be configured as a separate component / module. The receiving processor may conversely perform the processing of the above-mentioned sending processor to correspond to the transmission processing performed on the sending side. The receiving processor may send the acquired point cloud video to the file / fragment de-encapsulator 20005 and send the acquired point cloud related metadata to the metadata parser.
[0438] The file / fragment de-encapsulator 20005 can de-encapsulate the point cloud data received from the receiving processor in the form of a file. The file / fragment de-encapsulator 20005 can de-encapsulate the file according to ISOBMFF, etc., and can obtain the point cloud bitstream or point cloud-related metadata (or a separate metadata bitstream). The obtained point cloud bitstream can be transmitted to the point cloud video decoder 20006 and the point cloud image decoder 20008, and the obtained point cloud video-related metadata (metadata bitstream) can be transmitted to the metadata processor (not shown). The point cloud bitstream can include metadata (metadata bitstream). The metadata processor can be included in the point cloud video decoder 20006 or can be configured as a separate component / module. The point cloud video-related metadata obtained by the file / fragment de-encapsulator 20005 can take the form of boxes or tracks in the file format. When necessary, the file / fragment de-encapsulator 20005 can receive the metadata required for de-encapsulation from the metadata processor. The point cloud-related metadata can be transmitted to the point cloud video decoder 20006 and / or the point cloud image decoder 20008 and used in the point cloud decoding process, or can be transmitted to the renderer 20009 and used in the point cloud rendering process.
[0439] The point cloud video decoder 20006 can receive the bitstream and decode the video / image by performing operations corresponding to those of the point cloud video encoder 20002. In this case, as described below, the point cloud video decoder 20006 can decode the point cloud video by dividing the point cloud video into geometric video, attribute video, occupancy map video, and auxiliary patch information. The geometric video can include geometric images, the attribute video can include attribute images, and the occupancy map video can include occupancy map images. The auxiliary information can include auxiliary patch information. The attribute video / image can include texture video / image.
[0440] The point cloud image decoder 20008 can receive the bitstream and perform the inverse process corresponding to the operations of the point cloud image encoder 20003. In this case, the point cloud image decoder 20008 can divide the point cloud image into geometric images, attribute images, occupancy map images, and metadata (which is, for example, auxiliary patch information) to decode it.
[0441] The 3D geometry can be reconstructed based on the decoded geometric video / image, occupancy map, and auxiliary patch information, and then can be subjected to smoothing processing. The color point cloud image / frame can be reconstructed by assigning color values to the smoothed 3D geometry based on the texture video / image. The renderer 20009 can render the reconstructed geometry and color point cloud image / frame. The rendered video / image can be displayed through a display. All or part of the rendering result can be displayed to the user through a VR / AR display or a typical display.
[0442] The sensor / tracker (sensing / tracking) 20007 obtains orientation information and / or user viewport information from the user or the receiving side and transmits the orientation information and / or user viewport information to the receiver and / or transmitter. The orientation information may represent information about the position, angle, movement, etc. of the user's head, or information about the position, angle, movement, etc. of the device through which the user is viewing the video / image. Based on this information, information about the area currently viewed by the user in the 3D space (i.e., viewport information) can be calculated.
[0443] The viewport information may be information about the area currently viewed by the user in the 3D space through the device or the HMD. A device such as a display can extract the viewport area based on the orientation information, the vertical or horizontal FOV supported by the device, etc. The orientation or viewport information can be extracted or calculated on the receiving side. The orientation or viewport information analyzed on the receiving side can be sent to the transmitting side on the feedback channel.
[0444] Based on the orientation information obtained by the sensor / tracker 20007 and / or the viewport information indicating the area currently viewed by the user, the receiver can effectively extract or decode only the media data of a specific area (i.e., the area indicated by the orientation information and / or viewport information) from the file. In addition, based on the orientation information and / or viewport information obtained by the sensor / tracker 20007, the transmitter can effectively encode only the media data of a specific area (i.e., the area indicated by the orientation information and / or viewport information), or generate and send its file.
[0445] The renderer 20009 can render the decoded point cloud data in the 3D space. The rendered video / image can be displayed through the display. The user can view all or part of the rendering result through the VR / AR display or a typical display.
[0446] The feedback processing may include transmitting various feedback information that can be obtained in the rendering / display processing to the decoder on the transmitting side or the receiving side. Through the feedback processing, interactivity can be provided when consuming the point cloud data. According to an embodiment, the head orientation information, the viewport information indicating the area currently viewed by the user, etc. can be transmitted to the transmitting side in the feedback processing. According to an embodiment, the user can interact with the content implemented in the VR / AR / MR / autonomous driving environment. In this case, the information related to the interaction can be transmitted to the transmitting side or the service provider in the feedback processing. According to an embodiment, the feedback processing can be skipped.
[0447] According to an embodiment, the above feedback information can be not only sent to the transmitting side, but also consumed at the receiving side. That is, the unpacking processing, decoding, and rendering processing at the receiving side can be performed based on the above feedback information. For example, the point cloud data about the area currently viewed by the user can be preferentially unpacked, decoded, and rendered based on the orientation information and / or viewport information.
[0448] Figure 21 is an exemplary block diagram of an apparatus for storing and transmitting point cloud data according to an embodiment.
[0449] Figure 21 Shows a point cloud system according to an embodiment. Figure 21 Some or all of the system of Figure 1 the transmitting device and the receiving device of Figure 4 the encoding process of Figure 15 the 2D video / image encoder of Figure 16 the decoding process of Figure 18 the transmitting device and / or Figure 19 some or all of the receiving device of Figure 20 Some or all of the system of
[0450] The point cloud data transmitting device according to an embodiment may be configured as shown. Each element of the transmitting device may be a module / unit / component / hardware / software / processor.
[0451] The geometry, attributes, occupancy map, auxiliary data (auxiliary information), and mesh data of the point cloud may each be configured as separate streams or stored in different tracks in a file. In addition, they may be included in separate segments.
[0452] The point cloud acquirer 21000 acquires a point cloud. For example, one or more cameras may acquire point cloud data by capturing, arranging, or generating a point cloud. Through this acquisition process, point cloud data including the 3D positions of individual points (which may be represented by x, y, and z position values, etc.) (hereinafter referred to as geometry) and the attributes of individual points (color, reflectivity, transparency, etc.) can be acquired. For example, a Polygon File Format (PLY) (or Stanford Triangle Format) file including the point cloud data may be generated. For point cloud data having multiple frames, one or more files may be acquired. Herein, metadata related to the point cloud (for example, metadata related to capture, etc.) may be generated. The patch generator 21001 generates patches from the point cloud data. The patch generator 21001 generates the point cloud data or a point cloud video into one or more frames. A frame is generally a unit that represents an image at a specific time interval. When the points constituting the point cloud video are divided into one or more patches (a set of points constituting the point cloud video, where the points belonging to the same patch are adjacent to each other in 3D space and are mapped in the same direction between the flat surfaces of a six-sided bounding box when mapped to a 2D image), an occupancy map frame of a binary image can be generated, which indicates whether there is data at the corresponding position in the 2D plane with 0 or 1. In addition, a geometry frame in the form of a depth map representing information about the positions (geometry) of the individual points constituting the point cloud video can be generated for each patch. A texture frame representing color information about the individual points constituting the point cloud video can be generated for each patch. Herein, metadata required to reconstruct the point cloud from each patch can be generated. The metadata may include information about the patches, such as the positions and sizes of the individual patches in 2D / 3D space. These frames can be continuously generated in chronological order to construct a video stream or a metadata stream.
[0453] In addition, the patches can be used for 2D image mapping. For example, the point cloud data can be projected onto the respective faces of a cube. After patch generation, a geometry image, one or more attribute images, an occupancy map, auxiliary data, and / or mesh data can be generated based on the generated patches.
[0454] The generation of the geometry image, the attribute image, the occupancy map, the auxiliary data, and / or the mesh data is performed by the point cloud pre-processor 20001 or a controller (not shown). The point cloud pre-processor 20001 may include a patch generator 21001, a geometry image generator 21002, an attribute image generator 21003, an occupancy map generator 21004, an auxiliary data generator 21005, and a mesh data generator 21006.
[0455] The geometric image generator 21002 generates a geometric image based on the result of patch generation. The geometric representation is points in 3D space. An occupancy map is used to generate the geometric image, which includes information related to the 2D image packing of the patches, auxiliary data (including patch data), and / or mesh data based on the patches. The geometric image is related to information such as the depth of the patches (e.g., near, far) generated after patch generation.
[0456] The attribute image generator 21003 generates an attribute image. For example, the attribute can represent a texture. The texture can be a color value that matches each point. According to an embodiment, an image including multiple attributes (e.g., color and reflectivity) (N attributes) of the texture can be generated. The multiple attributes can include material information and reflectivity. According to an embodiment, the attribute can additionally include information indicating color, which can vary according to the viewing angle and light even for the same texture.
[0457] The occupancy map generator 21004 generates an occupancy map from the patches. The occupancy map includes information indicating whether there is data in the pixels (e.g., corresponding to the geometric or attribute image).
[0458] The auxiliary data generator 21005 generates auxiliary data (or auxiliary information) including information about the patches. That is, the auxiliary data represents the metadata of the patches of the point cloud object. For example, it can represent information such as the normal vector of the patches. Specifically, the auxiliary data can include information required to reconstruct the point cloud from the patches (e.g., information about the position, size, etc. of the patches in 2D / 3D space, projection (normal) plane identification information, patch mapping information, etc.).
[0459] The mesh data generator 21006 generates mesh data from the patches. The mesh represents the connection between adjacent points. For example, it can represent data in a triangular shape. For example, the mesh data refers to the connectivity between points.
[0460] The point cloud preprocessor 20001 or the controller generates metadata related to patch generation, geometric image generation, attribute image generation, occupancy map generation, auxiliary data generation, and mesh data generation.
[0461] The point cloud transmitting device performs video encoding and / or image encoding in response to the result generated by the point cloud preprocessor 20001. The point cloud transmitting device can generate point cloud image data and point cloud video data. According to an embodiment, the point cloud data can have only video data, only image data, and / or have both video data and image data.
[0462] The video encoder 21007 performs geometric video compression, attribute video compression, occupancy map video compression, auxiliary data compression, and / or mesh data compression. The video encoder 21007 generates a video stream containing the encoded video data.
[0463] Specifically, in geometric video compression, point cloud geometric video data is encoded. In attribute video compression, attribute video data of the point cloud is encoded. In auxiliary data compression, auxiliary data associated with the point cloud video data is encoded. In mesh data compression, mesh data of the point cloud video data is encoded. Each operation of the point cloud video encoder can be executed in parallel.
[0464] The image encoder 21008 performs geometric image compression, attribute image compression, occupancy map image compression, auxiliary data compression, and / or mesh data compression. The image encoder generates an image containing the encoded image data.
[0465] Specifically, in geometric image compression, point cloud geometric image data is encoded. In attribute image compression, attribute image data of the point cloud is encoded. In auxiliary data compression, auxiliary data associated with the point cloud image data is encoded. In mesh data compression, mesh data associated with the point cloud image data is encoded. Each operation of the point cloud image encoder can be executed in parallel.
[0466] The video encoder 21007 and / or the image encoder 21008 can receive metadata from the point cloud preprocessor 21001. The video encoder 21007 and / or the image encoder 21008 can perform respective encoding processes based on the metadata.
[0467] The file / fragment encapsulator 21009 encapsulates the video stream and / or the image in the form of a file and / or a fragment. The file / fragment encapsulator 21009 performs video track encapsulation, metadata track encapsulation, and / or image encapsulation.
[0468] In video track encapsulation, one or more video streams can be encapsulated into one or more tracks.
[0469] In metadata track encapsulation, metadata related to the video stream and / or the image can be encapsulated in one or more tracks. The metadata includes data related to the content of the point cloud data. For example, it can include initial viewing orientation metadata. According to an embodiment, the metadata can be encapsulated into a metadata track, or can be encapsulated together in a video track or an image track.
[0470] In image encapsulation, one or more images can be encapsulated into one or more tracks or items.
[0471] For example, according to an embodiment, when four video streams and two images are input to the encapsulator, the four video streams and the two images can be encapsulated in one file.
[0472] The file / fragment encapsulator 21009 can receive metadata from the point cloud preprocessor 21001. The file / fragment encapsulator 21009 can perform encapsulation based on the metadata.
[0473] The files and / or segments generated by the file / fragment encapsulator 21009 are sent by the point cloud transmitting device or transmitter. For example, the segments may be transmitted according to a DASH-based protocol.
[0474] The transmitter may send the point cloud bitstream or the file / segment including the bitstream to the receiver of the receiving device via a digital storage medium or a network. For transmission, processing according to any transmission protocol may be performed. The processed data for transmission may be sent via a broadcast network and / or by broadband. The data may be transmitted to the receiving side on a demand basis. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD.
[0475] The file / fragment encapsulator 21009 according to an embodiment may divide and store one bitstream or individual bitstreams into one or more tracks in a file, and may encapsulate the signaling information therefor. In addition, the patch (or atlas) stream included in the bitstream may be stored as a track in the file, and the related signaling information may be stored. In addition, the SEI messages present in the bitstream may be stored in the tracks of the file, and the related signaling information may be stored.
[0476] The transmitter may include elements for generating a media file in a predetermined file format, and may include elements for transmission via a broadcast / communication network. The transmitter receives the orientation information and / or viewport information from the receiver. The transmitter may transmit the acquired orientation information and / or viewport information (or the information selected by the user) to the point cloud pre-processor 21001, the video encoder 21007, the image encoder 21008, the file / fragment encapsulator 21009, and / or the point cloud encoder. Based on the orientation information and / or viewport information, the point cloud encoder may encode all the point cloud data or the point cloud data indicated by the orientation information and / or viewport information. Based on the orientation information and / or viewport information, the file / fragment encapsulator may encapsulate all the point cloud data or the point cloud data indicated by the orientation information and / or viewport information. Based on the orientation information and / or viewport information, the transmitter may transmit all the point cloud data or the point cloud data indicated by the orientation information and / or viewport information.
[0477] For example, the point cloud pre-processor 21001 may perform the above operations on all point cloud data or on the point cloud data indicated by the orientation information and / or viewport information. The video encoder 21007 and / or the image encoder 21008 may perform the above operations on all point cloud data or on the point cloud data indicated by the orientation information and / or viewport information. The file / fragment encapsulator 21009 may perform the above operations on all point cloud data or on the point cloud data indicated by the orientation information and / or viewport information. The transmitter may perform the above operations on all point cloud data or on the point cloud data indicated by the orientation information and / or viewport information.
[0478] Figure 22 is an exemplary block diagram of a point cloud data receiving device according to an embodiment.
[0479] Figure 22 illustrates a point cloud system according to an embodiment. Figure 22 Part / All of the system may include Figure 1 a transmitting device and a receiving device of Figure 4 the encoding process of Figure 15 a 2D video / image encoder of Figure 16 the decoding process of Figure 18 a transmitting device of and / or Figure 19 some or all of the receiving device of. In addition, it may be included or correspond to Figure 20 and Figure 21 Part / All of the system of.
[0480] Each component of the receiving device may be a module / unit / component / hardware / software / processor. The transport client may receive point cloud data, a point cloud bitstream, or a file / fragment including the bitstream transmitted by the point cloud data transmitting device according to an embodiment. Depending on the channel for transmission, the receiver may receive the point cloud data via a broadcast network or over broadband. Alternatively, the point cloud data may be received via a digital storage medium. The receiver may include processing to decode the received data and render the received data according to the user viewport. The transport client (receiving processor) 22006 may perform processing on the received point cloud data according to the transmission protocol. The receiving processor may be included in the receiver or configured as a separate component / module. Instead, the receiving processor may perform the processing of the above-mentioned transmitting processor to correspond to the transmission processing performed on the transmitting side. The receiving processor may transmit the acquired point cloud data to the file / fragment de-encapsulator 22000 and transmit the acquired point cloud-related metadata to a metadata processor (not shown).
[0481] The sensor / tracker 22005 obtains orientation information and / or viewport information. The sensor / tracker 22005 may transmit the obtained orientation information and / or viewport information to the transmission client 22006, the file / fragment de-packager 22000, the point cloud decoders 22001 and 22002, and the point cloud processor 22003.
[0482] The transmission client 22006 may receive all point cloud data or the point cloud data indicated by the orientation information and / or viewport information based on the orientation information and / or viewport information. The file / fragment de-packager 22000 may de-package all point cloud data or the point cloud data indicated by the orientation information and / or viewport information based on the orientation information and / or viewport information. The point cloud decoders (video decoder 22001 and / or image decoder 22002) may decode all point cloud data or the point cloud data indicated by the orientation information and / or viewport information based on the orientation information and / or viewport information. The point cloud processor 22003 may process all point cloud data or the point cloud data indicated by the orientation information and / or viewport information based on the orientation information and / or viewport information.
[0483] The file / fragment de-packager 22000 performs video track de-packaging, metadata track de-packaging, and / or image de-packaging. The file / fragment de-packager 22000 may de-package the point cloud data received in the form of a file from the receiving processor. The file / fragment de-packager 22000 may de-package the file or fragment according to ISOBMFF, etc. to obtain the point cloud bitstream or point cloud-related metadata (or a separate metadata bitstream). The obtained point cloud bitstream may be transmitted to the point cloud decoders 22001 and 22002, and the obtained point cloud-related metadata (or metadata bitstream) may be transmitted to the metadata processor (not shown). The point cloud bitstream may include metadata (metadata bitstream). The metadata processor may be included in the point cloud video decoder or may be configured as a separate component / module. The point cloud-related metadata obtained by the file / fragment de-packager 22000 may take the form of boxes or tracks in the file format. If necessary, the file / fragment de-packager 22000 may receive the metadata required for de-packaging from the metadata processor. The point cloud-related metadata may be transmitted to the point cloud decoders 22001 and 22002 and used in the point cloud decoding process, or may be transmitted to the renderer 22004 and used in the point cloud rendering process. The file / fragment de-packager 22000 may generate metadata related to the point cloud data.
[0484] In the video track de-packaging performed by the file / fragment de-packager 22000, the video tracks contained in the file and / or fragment are de-packaged. Video streams including geometric video, attribute video, occupancy maps, auxiliary data, and / or mesh data are de-packaged.
[0485] In the metadata track demultiplexing performed by the file / fragment demultiplexer 22000, a bitstream including metadata related to point cloud data and / or auxiliary data is demultiplexed.
[0486] In the image demultiplexing performed by the file / fragment demultiplexer 22000, an image including a geometric image, an attribute image, an occupancy map, auxiliary data, and / or mesh data is demultiplexed.
[0487] The file / fragment demultiplexer 22000 according to an embodiment may store one bitstream or separate bitstreams into one or more tracks in a file and may also demultiplex signaling information therefor. In addition, a bitstream included in the bitstream or an atlas (patch) stream may be demultiplexed based on the tracks in the file, and the related signaling information may be parsed. In addition, SEI messages present in the bitstream may be demultiplexed based on the tracks in the file, and the related signaling information may also be obtained.
[0488] The video decoder 22001 performs geometric video decompression, attribute video decompression, occupancy map decompression, auxiliary data decompression, and / or mesh data decompression. The video decoder 22001 decodes geometric video, attribute video, auxiliary data, and / or mesh data in a process corresponding to the process performed by the video encoder of the point cloud transmission device according to an embodiment.
[0489] The image decoder 22002 performs geometric image decompression, attribute image decompression, occupancy map decompression, auxiliary data decompression, and / or mesh data decompression. The image decoder 22002 decodes geometric images, attribute images, auxiliary data, and / or mesh data in a process corresponding to the process performed by the image encoder of the point cloud transmission device according to an embodiment.
[0490] The video decoder 22001 and the video decoder 22002 according to an embodiment may be processed by one video / image decoder as described above and may be executed along separate paths as shown.
[0491] The video decoder 22001 and / or the image decoder 22002 may generate metadata related to video data and / or image data.
[0492] In the point cloud processor 22003, geometric reconstruction and / or attribute reconstruction is performed.
[0493] In geometric reconstruction, geometric video and / or geometric images are reconstructed from decoded video data and / or decoded image data based on an occupancy map, auxiliary data, and / or mesh data.
[0494] In property reconstruction, the property video and / or property image are reconstructed from the decoded property video and / or decoded property image based on the occupancy map, auxiliary data, and / or grid data. According to an embodiment, for example, the property can be texture. According to an embodiment, the property can represent multiple pieces of property information. When there are multiple properties, the point cloud processor 22003 according to the embodiment performs multiple property reconstructions.
[0495] The point cloud processor 22003 can receive metadata from the video decoder 22001, the image decoder 22002, and / or the file / fragment demultiplexer 22000, and process the point cloud based on the metadata.
[0496] The point cloud renderer 22004 renders the reconstructed point cloud. The point cloud renderer 22004 can receive metadata from the video decoder 22001, the image decoder 22002, and / or the file / fragment demultiplexer 22000, and render the point cloud based on the metadata.
[0497] The display shows the rendering result on an actual display device.
[0498] According to the method / apparatus according to the embodiment, as Figures 20 to 22 shown, the transmitting side can encode the point cloud data into a bitstream, encapsulate the bitstream into the form of a file and / or fragment, and send it. The receiving side can demultiplex the file and / or fragment into a bitstream containing the point cloud, and can decode the bitstream into point cloud data. For example, the point cloud data device according to the embodiment can encapsulate the point cloud data based on a file. The file can include a V-PCC track containing the parameters of the point cloud, a geometry track containing the geometry, an attribute track containing the attributes, and an occupancy track containing the occupancy map.
[0499] In addition, the point cloud data receiving device according to the embodiment demultiplexes the point cloud data based on a file. The file can include a V-PCC track containing the parameters of the point cloud, a geometry track containing the geometry, an attribute track containing the attributes, and an occupancy track containing the occupancy map.
[0500] The above encapsulation operation can be performed by Figure 20 the file / fragment encapsulator 20004 of Figure 21 or the file / fragment encapsulator 21009 of Figure 20 The above demultiplexing operation can be performed by Figure 22 the file / fragment demultiplexer 20005 of
[0501] Figure 23 or the file / fragment demultiplexer 22000 of
[0502] In the structure according to an embodiment, at least one of a server 23600, a robot 23100, a self-driving vehicle 23200, an XR device 23300, a smart phone 23400, a household appliance 23500, and / or a head-mounted display (HMD) 23700 is connected to a cloud network 23000. Here, the robot 23100, the self-driving vehicle 23200, the XR device 23300, the smart phone 23400, or the household appliance 23500 may be referred to as a device. Additionally, the XR device 23300 may correspond to a point cloud data (PCC) device according to an embodiment, or may be operatively connected to a PCC device.
[0503] The cloud network 23000 may represent a network that forms part of or exists in a cloud computing infrastructure. Here, the cloud network 23000 may be configured using a 3G network, a 4G or Long Term Evolution (LTE) network, or a 5G network.
[0504] The server 23600 may be connected via the cloud network 23000 to at least one of the robot 23100, the self-driving vehicle 23200, the XR device 23300, the smart phone 23400, the household appliance 23500, and / or the HMD 23700, and may assist in at least a part of the processing of the connected devices 23100 to 23700.
[0505] The HMD 23700 represents one of the implementation types of an XR device and / or a PCC device according to an embodiment. The HMD-type device according to an embodiment includes a communication unit, a control unit, a memory, an I / O unit, a sensor unit, and a power unit.
[0506] Hereinafter, various embodiments of the devices 23100 to 23500 to which the above technology is applied will be described. Figure 23 The illustrated devices 23100 to 23500 may be operatively connected / linked to a point cloud data sending and receiving device according to the above embodiment.
[0507] <PCC+XR>
[0508] The XR / PCC device 23300 may adopt PCC technology and / or XR (AR+VR) technology, and may be implemented as an HMD, a head-up display (HUD) provided in a vehicle, a television, a mobile phone, a smart phone, a computer, a wearable device, a household appliance, a digital signage, a vehicle, a stationary robot, or a mobile robot.
[0509] The XR / PCC device 23300 can analyze 3D point cloud data or image data obtained through various sensors or from external devices and generate position data and attribute data about 3D points. Thus, the XR / PCC device 23300 can obtain information about the surrounding space or real objects, and render and output XR objects. For example, the XR / PCC device 23300 can match an XR object including auxiliary information about the identified object with the identified object and output the matched XR object.
[0510] <PCC + Autonomous Driving + XR>
[0511] The autonomous driving vehicle 23200 can be implemented as a mobile robot, a vehicle, an unmanned aerial vehicle, etc. by applying PCC technology and XR technology.
[0512] The autonomous driving vehicle 23200 applying the XR / PCC technology can represent an autonomous vehicle provided with means for providing an XR image, or an autonomous vehicle that is a control / interaction target in the XR image. Specifically, the autonomous driving vehicle 23200 that is a control / interaction target in the XR image can be distinguished from and operably connected to the XR device 23300.
[0513] The autonomous driving vehicle 23200 having means for providing an XR / PCC image can obtain sensor information from sensors including a camera and output the generated XR / PCC image based on the obtained sensor information. For example, the autonomous driving vehicle can have a HUD and output the XR / PCC image thereto to provide an XR / PCC object corresponding to a real object or an object present on the screen to passengers.
[0514] In this case, when the XR / PCC object is output to the HUD, at least a part of the XR / PCC object can be output to overlap with the real object pointed to by the passenger's eyes. On the other hand, when the XR / PCC object is output on a display provided inside the autonomous driving vehicle, at least a part of the XR / PCC object can be output to overlap with the object on the screen. For example, the autonomous driving vehicle can output an XR / PCC object corresponding to objects such as roads, other vehicles, traffic lights, traffic signs, two-wheeled vehicles, pedestrians, and buildings.
[0515] According to an embodiment, virtual reality (VR) technology, augmented reality (AR) technology, mixed reality (MR) technology, and / or point cloud compression (PCC) technology are applicable to various devices.
[0516] In other words, VR technology is a display technology that only provides real-world objects, backgrounds, etc. as CG images. On the other hand, AR technology refers to a technology that displays CG images virtually created on real object images. MR technology is similar to the above AR technology in that the virtual objects to be displayed are mixed and combined with the real world. However, MR technology is different from AR technology. AR technology clearly distinguishes between real objects and virtual objects created as CG images and uses virtual objects as supplementary objects for real objects, while MR technology treats virtual objects as objects having the same characteristics as real objects. More specifically, an example of the application of MR technology is holographic services.
[0517] Recently, VR, AR, and MR technologies are generally referred to as extended reality (XR) technologies, rather than being clearly distinguished from each other. Therefore, the embodiments of the present disclosure are applicable to all VR, AR, MR, and XR technologies. For these technologies, encoding / decoding based on PCC, V-PCC, and G-PCC technologies can be applied.
[0518] The PCC method / device according to the embodiment can be applied to the self-driving vehicle 23200 that provides self-driving services.
[0519] The self-driving vehicle 23200 that provides self-driving services is connected to the PCC device for wired / wireless communication.
[0520] When the point cloud data compression transmission and reception device (PCC device) according to the embodiment is connected to the self-driving vehicle 23200 for wired / wireless communication, the device can receive and process content data related to AR / VR / PCC services that can be provided together with the self-driving service and send the processed content data to the self-driving vehicle 23200. In the case where the point cloud data transmission and reception device is installed on the vehicle, the point cloud transmission and reception device can receive and process content data related to AR / VR / PCC services according to a user input signal input through the user interface device and provide the processed content data to the user. The self-driving vehicle 23200 or the user interface device according to the embodiment can receive a user input signal. The user input signal according to the embodiment can include a signal indicating the self-driving service.
[0521] As described above, Figure 1 、 Figure 4 、 Figure 18 、 Figure 20 or Figure 21The V-PCC based point cloud video encoder projects 3D point cloud data (or content) into a 2D space to generate patches. Patches are generated in the 2D space by dividing the data into a geometric image representing position information (referred to as a geometric frame or geometric patch frame) and a texture image representing color information (referred to as an attribute frame or attribute patch frame). The geometric image and the texture image are video compressed for each frame, and a video bitstream of the geometric image (referred to as a geometric bitstream) and a video bitstream of the texture image (referred to as an attribute bitstream) are output. In addition, auxiliary patch information (also referred to as patch information or metadata or atlas data), including projection plane information and patch size information for each patch (which are required for decoding 2D patches on the receiving side), is also video compressed and a bitstream of the auxiliary patch information is output. In addition, an occupancy map indicating the presence / absence of a point for each pixel as 0 or 1 is entropy compressed or video compressed depending on whether it is in a lossless mode or a lossy mode, and a video bitstream of the occupancy map (or referred to as an occupancy map bitstream) is output. The compressed geometric bitstream, the compressed attribute bitstream, the compressed auxiliary patch information bitstream (also referred to as the compressed atlas bitstream), and the compressed occupancy map bitstream are multiplexed into the structure of a V-PCC bitstream.
[0522] According to an embodiment, the V-PCC bitstream can be sent to the receiving side as it is, or can be encapsulated in the form of a file / fragment by Figure 1 , Figure 18 , Figure 20 or Figure 21 's file / fragment encapsulator, and sent to the receiving device or stored in a digital storage medium (e.g., USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.). According to an embodiment of the present disclosure, the file is in the file format ISOBMFF.
[0523] According to an embodiment, the V-PCC bitstream can be sent through multiple tracks in a file, or can be sent through a single track. Details will be described later.
[0524] In the present disclosure, point cloud data (i.e., V-PCC data) represents volumetric encoding of a point cloud composed of a sequence of point cloud frames. In a point cloud sequence (which is a sequence of point cloud frames), each point cloud frame includes a set of points. Each point can have a 3D position (i.e., geometric information) and multiple attributes such as color, reflectivity, surface normal, etc. That is, each point cloud frame represents a set of 3D points specified by the Cartesian coordinates (x, y, z) (i.e., position) of 3D points and zero or more attributes at a specific time instance.
[0525] As used herein, the term video-based point cloud compression (V-PCC) is the same as vision-based volumetric video coding (V3C). According to an embodiment, the terms V-PCC and V3C may have the same meaning and may be used interchangeably.
[0526] According to an embodiment, point cloud content (or referred to as V-PCC content or V3C content) represents volumetric media (or point cloud data) encoded using V-PCC.
[0527] According to an embodiment, point cloud content (or referred to as a volumetric scene) represents 3D data and may be divided into one or more objects (or composed of one or more objects).
[0528] That is, a volumetric scene is a region or unit composed of one or more objects that make up the volumetric media. In addition, when encapsulating the V-PCC bitstream in a file format, the region obtained by dividing the bounding box of the entire volumetric media according to a spatial standard is called a 3D spatial region. According to an embodiment, the 3D spatial region may be referred to as a 3D region or a spatial region.
[0529] According to an embodiment, an object may represent a piece of point cloud data or volumetric media or V3C content or V-PCC content. Additionally, an object may be divided into multiple objects, etc., according to spatial criteria, etc. In the present disclosure, each of the divided objects is referred to as a sub-object or simply as an object. According to an embodiment, a 3D bounding box may represent information indicating the position information of an object in 3D space, and a 2D bounding box may represent a rectangular region surrounding a patch corresponding to an object in a 2D frame. That is, the data generated after the process of projecting an object onto a 2D plane (which is one of the processes of encoding an object) is a patch, and the box surrounding the patch may be referred to as a 2D bounding box. That is, since an object may be composed of multiple patches during the encoding process, an object is related to patches. In addition, an object in 3D space is represented by 3D bounding box information surrounding the object, and an atlas frame includes patch information corresponding to each object and 3D bounding box information in 3D space. Therefore, an object is related to a 3D bounding box, an atlas tile, or a 3D spatial region.
[0530] According to an embodiment, in the step of encapsulating the V-PCC bitstream in a file format, the 3D bounding box of the point cloud data can be divided into one or more 3D regions, and each divided 3D region can include one or more objects. Further, in the step of encapsulating patches in an atlas frame, the patches are collected on an object-by-object basis and encapsulated (mapped) into one or more atlas patch regions in the atlas frame. That is, one 3D region (i.e., file level) can be associated with one or more objects (i.e., bitstream level), and one object can be associated with one or more 3D regions. And since each object is associated with one or more atlas patches, the 3D region can be associated with one or more atlas patches. An atlas patch (or patch) according to an embodiment represents an independently decodable rectangular region of the atlas frame.
[0531] Figure 24 Shows an example of the relationship between an object and an atlas patch (or patch) according to an embodiment.
[0532] According to an embodiment, assume that the point cloud content (or referred to as point cloud data or V3C content or V-PCC content) is divided into three objects (Object #1, Object #2, Object #3), and the atlas frame is composed of five atlas patches (Atlas Patch #1, Atlas Patch #2, Atlas Patch #3, Atlas Patch #4, Atlas Patch #5).
[0533] According to an embodiment, the 3D regions can overlap each other. As an example, the 3D (spatial) region #1 can include Object #1, and the 3D region #2 can include Object #2 and Object #3. As another example, the 3D region #1 can include Object #1 and Object #2, and the 3D region #2 can include Object #2 and Object #3. In other words, Object #2 can be associated with both 3D region #1 and 3D region #2. Additionally, since the same object (e.g., Object #2) can be included in different 3D regions (e.g., 3D region #1 and 3D region #2), the patches corresponding to Object #2 can be assigned to (included in) different 3D regions (3D region #1 and 3D region #2).
[0534] According to an embodiment, the atlas frame is composed of one or more atlas patches (or referred to as patches). The objects in the point cloud content can be associated with one or more atlas patches in the atlas frame. For example, as shown in Figure 24In this example, it is assumed that the point cloud content (or point cloud data) is divided into three objects (Object #1, Object #2, Object #3), and the atlas frame is composed of five atlas patches (Atlas Patch #1, Atlas Patch #2, Atlas Patch #3, Atlas Patch #4, Atlas Patch #5). Object #1 can be related to Atlas Patch #1 and Atlas Patch #2, Object #2 can be related to Atlas Patch #3, and Object #3 can be related to Atlas Patch #4 and Atlas Patch #5. This is merely an example for better understanding of the present disclosure, and the number of objects into which the point cloud content is divided and the atlas patches related to each object can vary.
[0535] The present disclosure proposes a method for extracting, decoding, and rendering a V-PCC sub-bitstream related to a necessary object only from a V-PCC bitstream.
[0536] According to an embodiment, an object can be related to one or more atlas sub-bitstreams (or referred to as atlas sub-streams). According to an embodiment, an atlas sub-bitstream can be defined as a sub-bitstream extracted from a V-PCC bitstream including a part of atlas NAL units. In an embodiment, an atlas sub-bitstream can correspond to one atlas. If there are multiple atlases, multiple atlas sub-bitstreams can be generated. Additionally, a video sub-bitstream can include occupancy, geometry, and attribute components of each atlas or each atlas patch.
[0537] For example, one or more atlases (or atlas sub-bitstreams) can be included in a V-PCC bitstream to support 3DOF+ video. The present disclosure proposes a method for efficiently using one or more atlases received by a receiver. Specifically, the present disclosure proposes a method for extracting, decoding, and rendering a necessary sub-bitstream only from a V-PCC bitstream even when there are multiple atlas sub-bitstreams.
[0538] According to an embodiment, an atlas sub-bitstream carries a part or all of the atlas data.
[0539] According to an embodiment, Atlas data is signaling information, which includes an Atlas Sequence Parameter Set (ASPS), an Atlas Frame Parameter Set (AFPS), an Atlas Adaptation Parameter Set (AAPS), Atlas tile group information (also referred to as Atlas tile information), and SEI messages, and can be referred to as metadata about Atlas. According to an embodiment, the ASPS is a syntax structure containing syntax elements that are applied to zero or more complete coded Atlas sequences (CASs) as determined by the content of the syntax elements in the ASPS (referred to as the syntax elements in each tile group (or tile) header). According to an embodiment, the AFPS is a syntax structure including syntax elements that are applied to zero or more complete coded Atlas frames as determined by the content of the syntax elements in each tile group (or tile). According to an embodiment, the AAPS may include camera parameters related to a part of an Atlas sub-bitstream, e.g., camera position, rotation, scale, and camera model. In the present disclosure, for simplicity, the ASPS, AFPS, and AAPS are referred to as Atlas parameter sets.
[0540] According to an embodiment, as used herein, the term syntax element may have the same meaning as a field or parameter.
[0541] According to an embodiment, an Atlas represents a set of 2D bounding boxes and can be a patch projected onto a rectangular frame.
[0542] According to an embodiment, an Atlas frame is a 2D rectangular array of Atlas samples onto which patches are projected, and additional information corresponding to a volume frame related to the patches. An Atlas sample is the position of the rectangular frame onto which a patch related to the Atlas is projected.
[0543] According to an embodiment, an Atlas frame can be divided into one or more rectangular partitions, which may be referred to as tile partitions or tiles. Alternatively, two or more tile partitions can be grouped and referred to as a tile. In other words, one or more tile partitions can constitute a tile. In the present disclosure, a tile has the same meaning as an Atlas tile. A tile is a partitioning unit of the signaling information of the point cloud data called Atlas. According to an embodiment, the tiles in an Atlas frame do not overlap with each other, and an Atlas frame may include regions unrelated to the tiles (i.e., one or more tile partitions). Additionally, the height and width of each tile included in an Atlas frame may vary between tiles.
[0544] According to an embodiment, some of the point cloud data corresponding to a specific 3D spatial region among all the point cloud data may be associated with one or more 2D regions. Thus, a 3D region may correspond to an atlas frame and may be associated with multiple 2D regions. According to an embodiment, the 2D region represents one or more video frames or atlas frames containing data related to the point cloud data in the 3D region.
[0545] According to an embodiment, a patch is a set of points that make up a point cloud, which indicates that the points belonging to the same patch are adjacent to each other in 3D space and are mapped in the same direction among the six-sided bounding box planes during the process of mapping to a 2D image. The patch is signaling information regarding the configuration of the point cloud data.
[0546] The receiving device according to an embodiment may reconstruct attribute video data, geometric video data, and occupancy video data, which are actual video data based on atlases (tiles, patches) having the same presentation time.
[0547] For spatial access or partial access to the point cloud data, it is necessary to access a part of the point cloud data according to a 3D (spatial) region or object. To this end, in the present disclosure, the mapping information between the 3D region and the tile, the mapping information between the 3D region and the object, the mapping information between the object and the atlas, or the mapping information between the object and the tile is signaled.
[0548] That is, in the present disclosure, in order to support the decoding and / or rendering / display of a specific object in the PCC decoder / player of the receiving device, the transmitting device may generate signaling information for identifying the relationship between the object and the atlas, the relationship between the object and the atlas tile, or the relationship between the object and the atlas sub-bitstream and transmit it to the receiving side. The signaling information may be static or vary over time. In the present disclosure, the signaling information may be transmitted through the V-PCC bitstream or through the sample entries and / or samples of the file carrying the V-PCC bitstream or in the form of metadata. According to an embodiment, the signaling information may be stored in a sample, a sample entry, a sample group, or an individual metadata track in an track group in a track or in a file. Specifically, a part of the signaling information may be stored in the sample entry of the track in the form of a box or a full box. Details of the signaling information related to the object and the atlas and the method for storing / transmitting the signaling information will be described later.
[0549] When the user zooms in or the user changes the viewport, a part of the entire point cloud object / data can be rendered or displayed on the changed viewport. In this case, it may be more efficient for the PCC decoder / player to decode or process the video or atlas data associated with the part of the point cloud data that is rendered or displayed on the user viewport and skip decoding or processing the video or atlas data associated with the part / region of the point cloud data that is not rendered or displayed. To this end, support for partial access to the point cloud data is required.
[0550] To achieve spatial access or partial access to the rendered / displayed point cloud data when the PCC decoder / player of the receiving device renders / displays the point cloud data, the transmitting device may send 3D region information about the point cloud (which may be static or change over time) through the V-PCC bitstream, signal the 3D region information about the point cloud in the sample entry and / or sample of the file carrying the V-PCC bitstream, or send the 3D region information about the point cloud in the form of metadata. In the present disclosure, the information signaled in the V-PCC bitstream or file or in the form of metadata is referred to as signaling information.
[0551] According to an embodiment, the signaling information may be stored in a sample, a sample entry, a sample group, or an orbit group in an orbit or in a separate metadata orbit in a file. Specifically, a part of the signaling information may be stored in the sample entry in the form of a box or a full box. Details of the signaling information related to the 3D spatial region and the method for storing / transmitting the signaling information will be described later.
[0552] According to an embodiment, the signaling information may be generated by a metadata generator (e.g., Figure 18 the metadata encoder 18005) of the transmitting device, and then signaled in a sample, a sample entry, a sample group, or an orbit group in an orbit or in a separate metadata orbit by a file / fragment encapsulator, or may be generated by the file / fragment encapsulator and signaled by the file / fragment encapsulator in a sample, a sample entry, a sample group, or an orbit group in an orbit or in a separate metadata orbit. In the present disclosure, the signaling information may include metadata related to the point cloud data (e.g., set values, etc.). Depending on the application, the signaling information may be defined at the system level (such as file format, HTTP Dynamic Adaptive Streaming over HTTP (DASH), or MPEG Media Transport (MMT)) or at the wired interface level (such as High-Definition Multimedia Interface (HDMI), DisplayPort, Video Electronics Standards Association (VESA), or CTA).
[0553] Figure 25 An example of the V-PCC bitstream structure according to other embodiments of the present disclosure is shown. In an embodiment, Figure 25 the V-PCC bitstream ofFigure 1 , Figure 4 , Figure 18 , Figure 20 or Figure 21 generation and output of a V-PCC based point cloud video encoder
[0554] A V-PCC bitstream containing an encoded point cloud sequence (CPCS) according to an embodiment may consist of a sample stream V-PCC unit or V-PCC units. The sample stream V-PCC unit or V-PCC units carry V-PCC parameter set (VPS) data, an atlas bitstream, a 2D video coded occupancy map bitstream, a 2D video coded geometry bitstream, and zero or more 2D video coded attribute bitstreams.
[0555] In Figure 25 , the V-PCC bitstream may include a sample stream V-PCC header 40010 and one or more sample stream V-PCC units 40020. For simplicity, one or more sample stream V-PCC units 40020 may be referred to as the sample stream V-PCC payload. That is, the sample stream V-PCC payload may be referred to as a set of sample stream V-PCC units. A detailed description of the sample stream V-PCC header 40010 will be described in Figure 27 .
[0556] Each sample stream V-PCC unit 40021 may include V-PCC unit size information 40030 and a V-PCC unit 40040. The V-PCC unit size information 40030 indicates the size of the V-PCC unit 40040. For simplicity, the V-PCC unit size information 40030 may be referred to as the sample stream V-PCC unit header, and the V-PCC unit 40040 may be referred to as the sample stream V-PCC unit payload.
[0557] Each V-PCC unit 40040 may include a V-PCC unit header 40041 and a V-PCC unit payload 40042.
[0558] In the present disclosure, the data contained in the V-PCC unit payload 40042 is distinguished by the V-PCC unit header 40041. To this end, the V-PCC unit header 40041 contains type information indicating the type of the V-PCC unit. According to the type information in the V-PCC unit header 40041, each V-PCC unit payload 40042 may contain at least one of geometric video data (i.e., a 2D video coded geometry bitstream), attribute video data (i.e., a 2D video coded attribute bitstream), occupancy video data (i.e., a 2D video coded occupancy map bitstream), atlas data, or a V-PCC parameter set (VPS).
[0559] The VPS according to an embodiment is also referred to as a Sequence Parameter Set (SPS). These two terms can be used interchangeably.
[0560] According to an embodiment, Atlas data is signaling information including an Atlas Sequence Parameter Set (ASPS), an Atlas Frame Parameter Set (AFPS), an Atlas Adaptation Parameter Set (AAPS), Atlas tile group information (or referred to as Atlas tile information), and SEI messages, and is referred to as an Atlas bitstream or a patch data group. Additionally, ASPS, AFPS, and AAPS are also referred to as Atlas parameter sets.
[0561] Figure 26 An example of data carried by a sample stream V-PCC unit in a V-PCC bitstream according to an embodiment is shown.
[0562] In Figure 26 the example, the V-PCC bitstream includes a sample stream V-PCC unit carrying a V-PCC Parameter Set (VPS), a sample stream V-PCC unit carrying Atlas data (AD), a sample stream V-PCC unit carrying Occupied Video Data (OVD), a sample stream V-PCC unit carrying Geometric Video Data (GVD), and a sample stream V-PCC unit carrying Attribute Video Data (AVD).
[0563] According to an embodiment, each sample stream V-PCC unit includes one type of V-PCC unit among VPS, AD, OVD, GVD, and AVD.
[0564] A field as a term used in the syntax of the present disclosure described below may have the same meaning as a parameter or an element (or a syntax element).
[0565] Figure 27 An example of the syntax structure of a sample stream V-PCC header 40010 included in a V-PCC bitstream according to an embodiment is shown.
[0566] The sample_stream_v-pcc_header() according to an embodiment may include an ssvh_unit_size_precision_bytes_minus1 field and an ssvh_reserved_zero_5bits field.
[0567] Adding 1 to the value of the ssvh_unit_size_precision_bytes_minus1 field may specify the precision (in bytes) of the ssvu_vpcc_unit_size element in all sample stream V-PCC units. The value of this field may be in the range of 0 to 7.
[0568] The ssvh_reserved_zero_5bits field is a reserved field for future use.
[0569] Figure 28 An example of the syntax structure of a sample_stream_vpcc_unit() according to an embodiment is shown.
[0570] The content of each sample_stream_vpcc_unit is associated with an access unit that is the same as the V-PCC unit contained in the sample_stream_vpcc_unit.
[0571] The sample_stream_vpcc_unit() according to an embodiment may include an ssvu_vpcc_unit_size field and a vpcc_unit(ssvu_vpcc_unit_size).
[0572] The ssvu_vpcc_unit_size field corresponds to Figure 25 the V-PCC unit size information 40030 and specifies the size (in bytes) of the subsequent vpcc_unit. The number of bits used to represent the ssvu_vpcc_unit_size field is equal to (ssvh_unit_size_precision_bytes_minus1 + 1)*8.
[0573] The vpcc_unit(ssvu_vpcc_unit_size) has a length corresponding to the value of the ssvu_vpcc_unit_size field and carries one of VPS, AD, OVD, GVD, and AVD.
[0574] Figure 29 An example of the syntax structure of a V-PCC unit according to an embodiment is shown. The V-PCC unit consists of a V-PCC unit header (vpcc_unit_header()) 40041 and a V-PCC unit payload (vpcc_unit_payload()) 40042. The V-PCC unit according to an embodiment may contain more data. In this case, it may also include a trailing_zero_8bits field. The trailing_zero_8bits field according to an embodiment is a byte corresponding to 0x00.
[0575] Figure 30 An example of the syntax structure of the V-PCC unit header 40041 according to an embodiment is shown. In the embodiment, Figure 30The vpcc_unit_header() includes a vuh_unit_type field. The vuh_unit_type field indicates the type of the corresponding V-PCC unit. The vuh_unit_type field according to the embodiment is also referred to as the vpcc_unit_type field.
[0576] Figure 31 An example of the V-PCC unit type assigned to the vuh_unit_type field according to the embodiment is shown.
[0577] Referring to Figure 31 , according to the embodiment, the vuh_unit_type field set to 0 indicates that the data included in the V-PCC unit payload of the V-PCC unit is a V-PCC parameter set (VPCC_VPS). The vuh_unit_type field set to 1 indicates that the data is Atlas data (VPCC_AD). The vuh_unit_type field set to 2 indicates that the data is occupied video data (VPCC_OVD). The vuh_unit_type field set to 3 indicates that the data is geometric video data (VPCC_GVD). The vuh_unit_type field set to 4 indicates that the data is attribute video data (VPCC_AVD).
[0578] Those skilled in the art can easily change the meaning, order, deletion, addition, etc. of the values assigned to the vuh_unit_type field. Therefore, the present disclosure will not be limited to the above embodiments.
[0579] When the vuh_unit_type field indicates VPCC_AVD, VPCC_GVD, VPCC_OVD, or VPCC_AD, the V-PCC unit header according to the embodiment may further include a vuh_vpcc_parameter_set_id field and a vuh_atlas_id field.
[0580] The vuh_vpcc_parameter_set_id field specifies the value of vps_vpcc_parameter_set_id of the active V-PCC VPS.
[0581] The vuh_atlas_id field specifies the index of the atlas corresponding to the current V-PCC unit.
[0582] When the vuh_unit_type field indicates VPCC_AVD, the V-PCC unit header according to the embodiment may further include a vuh_attribute_index field, a vuh_attribute_partition_index field, a vuh_map_index field, and a vuh_auxiliary_video_flag field.
[0583] The vuh_attribute_index field indicates the index of the attribute data carried in the attribute video data unit.
[0584] The vuh_attribute_partition_index field indicates the index of the attribute dimension group carried in the attribute video data unit.
[0585] When present, the vuh_map_index field may indicate the map index of the current geometry or attribute stream.
[0586] When the vuh_auxiliary_video_flag field is set to 1, it may indicate that the associated attribute video data unit contains only RAW and / or EOM (Enhanced Occupancy Mode) coding points. As another example, when the vuh_auxiliary_video_flag field is set to 0, it may indicate that the associated attribute video data unit may contain RAW and / or EOM coding points. When the vuh_auxiliary_video_flag field does not exist, its value may be inferred to be equal to 0. According to the embodiment, the RAW and / or EOM coding points are also referred to as Pulse Code Modulation (PCM) coding points.
[0587] When the vuh_unit_type field indicates VPCC_GVD, the V-PCC unit header according to the embodiment may further include a vuh_map_index field, a vuh_auxiliary_video_flag field, and a vuh_reserved_zero_12bits field.
[0588] When present, the vuh_map_index field indicates the index of the current geometry stream.
[0589] When the vuh_auxiliary_video_flag field is set to 1, it can indicate that the relevant geometric video data unit only contains RAW and / or EOM encoded points. As another example, when the vuh_auxiliary_video_flag field is set to 0, it can indicate that the associated geometric video data unit can contain RAW and / or EOM encoded points. When the vuh_auxiliary_video_flag field does not exist, its value can be inferred to be equal to 0. According to an embodiment, RAW and / or EOM encoded points are also referred to as PCM encoded points.
[0590] The vuh_reserved_zero_12bits field is a reserved field for future use.
[0591] If the vuh_unit_type field indicates VPCC_OVD or VPCC_AD, the V-PCC unit header according to an embodiment may further include the vuh_reserved_zero_17bits field. Otherwise, the V-PCC unit header may further include the vuh_reserved_zero_27bits field.
[0592] The vuh_reserved_zero_17bits field and the vuh_reserved_zero_27bits field are reserved fields for future use.
[0593] Figure 32 An example of the syntax structure of the V-PCC unit payload (vpcc_unit_payload()) according to an embodiment is shown.
[0594] Figure 32 The V-PCC unit payload can contain one of the V-PCC parameter set (vpcc_parameter_set()), the atlas sub-bitstream (atlas_sub_bitstream()), and the video sub-bitstream (video_sub_bitstream()) according to the value of the vuh_unit_type field in the V-PCC unit header.
[0595] For example, when the vuh_unit_type field indicates VPCC_VPS, the V-PCC unit payload contains vpcc_parameter_set(), which contains the overall coding information about the bitstream. When the vuh_unit_type field indicates VPCC_AD, the V-PCC unit payload contains atlas_sub_bitstream() that carries Atlas data. Additionally, according to an embodiment, when the vuh_unit_type field indicates VPCC_OVD, the V-PCC unit payload contains an occupancy video sub-bitstream (video_sub_bitstream()) that carries occupancy video data. When the vuh_unit_type field indicates VPCC_GVD, the V-PCC unit payload contains a geometry video sub-bitstream (video_sub_bitstream()) that carries geometry video data. When the vuh_unit_type field indicates VPCC_AVD, the V-PCC unit payload contains an attribute video sub-bitstream (video_sub_bitstream()) that carries attribute video data.
[0596] According to an embodiment, the Atlas sub-bitstream may be referred to as the Atlas sub-stream, and the occupancy video sub-bitstream may be referred to as the occupancy video sub-stream. The geometry video sub-bitstream may be referred to as the geometry video sub-stream, and the attribute video sub-bitstream may be referred to as the attribute video sub-stream. The V-PCC unit payload according to an embodiment may conform to the format of a High Efficiency Video Coding (HEVC) Network Abstraction Layer (NAL) unit.
[0597] Figure 33 An example of the syntax structure of the V-PCC parameter set included in the V-PCC unit payload according to an embodiment is shown.
[0598] In Figure 33 profile_tier_level() contains V-PCC codec profile-related information and specifies the constraints on the bitstream. It represents the limitations on the capabilities required to decode the bitstream. Profiles, tiers, and levels can be used to indicate interoperability points between individual decoder implementations.
[0599] The vps_vpcc_parameter_set_id field provides an identifier for the V-PCC VPS for other syntax elements to reference.
[0600] The value of the vps_atlas_count_minus1 field plus 1 indicates the total number of supported Atlases in the current bitstream.
[0601] The iterative statement is also included in the V-PCC parameter set, and this iterative statement repeats as many times as the value of the vps_atlas_count_minus1 field (i.e., the total number of atlases). In an embodiment, in the iterative statement, j can be initialized to 0 and incremented by 1 each time the iterative statement is executed until j reaches the value of the vps_atlas_count_minus1 field + 1.
[0602] In an embodiment, the iterative statement includes the following fields. In addition to the following fields (not shown), the iterative statement may further include an atlas identifier for identifying the atlas with index j. In one embodiment, the index j may be an identifier for identifying the j-th atlas.
[0603] The vps_frame_width[j] field indicates the V-PCC frame width in terms of integer luminance samples of the atlas with index j. This frame width is the nominal width associated with all V-PCC components of the atlas with index j.
[0604] The vps_frame_height[j] field indicates the V-PCC frame height in terms of integer luminance samples of the atlas with index j. This frame height is the nominal height associated with all V-PCC components of the atlas with index j.
[0605] Adding 1 to the vps_map_count_minus1[j] field indicates the number of maps used to encode the geometry and attribute data of the atlas with index j.
[0606] When the vps_map_count_minus1[j] field is greater than 0, the following parameters may also be included in the parameter set.
[0607] According to the value of the vps_map_count_minus1[j] field, the following parameters may also be included in the parameter set.
[0608] The vps_multiple_map_streams_present_flag[j] field being equal to 0 indicates that all geometry or attribute maps of the atlas with index j are placed in a single geometry or attribute video stream respectively. The vps_multiple_map_streams_present_flag[j] field being equal to 1 indicates that all geometry or attribute maps of the atlas with index j are placed in separate video streams.
[0609] If the vps_multiple_map_streams_present_flag[j] field is equal to 1, the vps_map_absolute_coding_enabled_flag[j][i] field may also be included in the parameter set. Otherwise, the vps_map_absolute_coding_enabled_flag[j][i] field may be 1.
[0610] The vps_map_absolute_coding_enabled_flag[j][i] field being equal to 1 indicates that the geometry map with index i of the atlas with index j is coded without any form of map prediction. The vps_map_absolute_coding_enabled_flag[j][i] field being equal to 0 indicates that the geometry map with index i of the atlas with index j is first predicted from another earlier coded map before coding.
[0611] The vps_map_absolute_coding_enabled_flag[j][0] field being equal to 1 indicates that the geometry map with index 0 is coded without map prediction.
[0612] If the vps_map_absolute_coding_enabled_flag[j][i] field is 0 and i is greater than 0, the vps_map_predictor_index_diff[j][i] field may also be included in the parameter set. Otherwise, the vps_map_predictor_index_diff[j][i] field may be 0.
[0613] When the vps_map_absolute_coding_enabled_flag[j][i] field is equal to 0, the vps_map_predictor_index_diff[j][i] field is used to calculate the predictor for the geometry map with index i of the atlas with index j.
[0614] The vps_auxiliary_video_present_flag[j] field being equal to 1 indicates that the auxiliary information of the atlas with index j, i.e., RAW or EOM patch data, can be stored in a separate video stream, called the auxiliary video stream. The vps_auxiliary_video_present_flag[j] field being equal to 0 indicates that the auxiliary information of the atlas with index j is not stored in a separate video stream.
[0615] occupancy_information() includes occupancy video related information.
[0616] geometry_information() includes related geometry video information.
[0617] attribute_information() includes attribute video related information.
[0618] That is to say, the V-PCC parameter set can include occupancy_information(), geometry_information() and attribute_information() for each atlas.
[0619] The V-PCC parameter set can also include the vps_extension_present_flag field.
[0620] The vps_extension_present_flag field equal to 1 specifies that the vps_extension_length field exists in the vpcc_parameter_set. The vps_extension_present_flag field equal to 0 specifies that the vps_extension_length field does not exist.
[0621] The vps_extension_length_minus1 field plus 1 specifies the number of vps_extension_data_byte elements that follow this syntax element.
[0622] According to the vps_extension_length_minus1 field, the extended data (vps_extension_data_byte) can also be included in the parameter set..
[0623] The vps_extension_data_byte field can have any value.
[0624] An Atlas frame (or point cloud object or patch frame) that is the target of point cloud data can be divided (or segmented) into one or more tiles or one or more Atlas tiles. According to an embodiment, a tile can represent a specific region in 3D space or a specific region in a 2D plane. Additionally, a tile can be a rectangular cuboid, a sub-boundary box, or a part of an Atlas frame within a bounding box. According to other embodiments, an Atlas frame can be divided into one or more rectangular partitions, which can be referred to as tile partitions or tiles. Alternatively, two or more tile partitions can be grouped and referred to as a tile. In the present disclosure, dividing an Atlas frame (or point cloud object) into one or more tiles can be performed by Figure 1 a point cloud video encoder, Figure 18 a patch generator, Figure 20 a point cloud preprocessor, or Figure 21 a patch generator, or can be performed by a separate component / module.
[0625] Figure 34 FIG. shows an example of dividing an Atlas frame (or patch frame) 43010 into multiple tiles by dividing the Atlas frame (or patch frame) 43010 into one or more tile rows and one or more tile columns. A tile 43030 is a rectangular region of the Atlas frame, and a tile group 43050 can contain multiple tiles in the Atlas frame. In the present disclosure, the tile group 43050 contains multiple tiles of the Atlas frame, which together form a rectangular (quadrilateral) region of the Atlas frame. In the present disclosure, tiles and tile groups may not be distinguished from each other, and one tile group can correspond to one tile. For example, Figure 34 FIG. shows an example in which an Atlas frame (or patch frame) is divided into 24 tiles (i.e., 6 tile columns and 4 tile rows) and 9 rectangular (quadrilateral) tile groups.
[0626] According to other embodiments, in Figure 34 , the tile group 43050 can be referred to as a tile, and the tile 43030 can be referred to as a tile partition. The term signaling information can also be varied and referred to according to the above complementary relationship.
[0627] Figure 35 is a diagram showing an example of the structure of an Atlas sub-stream as described above. In an embodiment, Figure 35 the Atlas sub-stream conforms to the format of an HEVC NAL unit.
[0628] An Atlas sub-stream according to an embodiment can include a sample stream NAL header and one or more sample stream NAL units. In Figure 35Among them, one or more sample stream NAL units may be referred to as a sample stream NAL payload. That is to say, a sample stream NAL payload may be referred to as a set of sample stream NAL units.
[0629] One or more sample stream NAL units according to an embodiment may consist of a sample stream NAL unit containing an atlas sequence parameter set (ASPS), a sample stream NAL unit containing an atlas frame parameter set (AFPS), and one or more sample stream NAL units containing information about one or more atlas tile groups (or atlas tiles) and / or one or more sample stream NAL units containing one or more SEI messages.
[0630] One or more SEI messages according to an embodiment may include a prefix SEI message and a suffix SEI message.
[0631] An example of the syntax structure of a sample_stream_nal_header() included in an atlas substream according to an embodiment is shown.
[0632] According to an embodiment, sample_stream_nal_header() may include an ssnh_unit_size_precision_bytes_minus1 field and an ssnh_reserved_zero_5bits field.
[0633] Adding 1 to the value of the ssnh_unit_size_precision_bytes_minus1 field may specify the precision (in bytes) of the ssnu_nal_unit_size element in all sample stream NAL units. The value of this field may be in the range of 0 to 7.
[0634] The ssnh_reserved_zero_5bits field is a reserved field for future use.
[0635] An example of the syntax structure of a sample_stream_nal_unit() according to an embodiment is shown.
[0636] According to an embodiment, sample_stream_nal_unit() may include an ssnu_nal_unit_size field and nal_unit(ssnu_nal_unit_size).
[0637] The ssnu_nal_unit_size field specifies the size of the subsequent NAL unit in bytes. The number of bits used to represent the ssnu_nal_unit_size field is equal to (ssnh_unit_size_precision_bytes_minus1 + 1) * 8.
[0638] The nal_unit (ssnu_nal_unit_size) has a length corresponding to the value of the ssnu_nal_unit_size field and carries one of an atlas sequence parameter set (ASPS), an atlas frame parameter set (AFPS), an atlas adaptation parameter set (AAPS), atlas tile group information (or atlas tile information), and an SEI message. That is, a sample stream NAL unit can contain one of an ASPS, an AFPS, an AAPS, atlas tile group information (or atlas tile information), or an SEI message. According to an embodiment, the ASPS, AFPS, AAPS, atlas tile group information (or atlas tile information), and SEI message are referred to as atlas data (or metadata related to atlas).
[0639] The SEI message according to an embodiment can assist in processing related to decoding, reconstruction, display, or other purposes.
[0640] Shown is an example of the syntax structure of the nal_unit (NumBytesInNalUnit).
[0641] In , NumBytesInNalUnit indicates the size of the NAL unit in bytes. NumBytesInNalUnit indicates the value of the ssnu_nal_unit_size field.
[0642] According to an embodiment, the NAL unit can include a NAL unit header (nal_unit_header()) and a NumBytesInRbsp field. The NumBytesInRbsp field is initialized to 0 and indicates the number of bytes of the payload belonging to the NAL unit.
[0643] The NAL unit includes an iterative statement with the same number of iterations as the value of NumBytesInNalUnit. In an embodiment, the iterative statement includes rbsp_byte[NumBytesInRbsp++]. According to an embodiment, in the iterative statement, i is initialized to 2 and incremented by 1 each time the iterative statement is executed. The iterative statement is repeated until i reaches the value of NumBytesInNalUnit.
[0644] rbsp_byte[NumBytesInRbsp++] is the i-th byte of the raw byte sequence payload (RBSP) that carries Atlas data. The RBSP is specified as a contiguous sequence of bytes. That is, rbsp_byte[NumBytesInRbsp++] carries one of the Atlas sequence parameter set (ASPS), Atlas frame parameter set (AFPS), Atlas adaptation parameter set (AAPS), Atlas tile group information (Atlas tile information), and SEI message.
[0645] Shown Figure 38 An example of the syntax structure of the NAL unit header. The NAL unit header may include a nal_forbidden_zero_bit field, a nal_unit_type field, a nal_layer_id field, and a nal_temporal_id_plus1 field.
[0646] The nal_forbidden_zero_bit is used for error detection in the NAL unit and must be 0.
[0647] The nal_unit_type specifies the type of the RBSP data structure contained in the NAL unit. Examples of the RBSP data structure according to the value of the nal_unit_type field will be described with reference to Figure 40 Examples of the RBSP data structure according to the value of the nal_unit_type field will be described.
[0648] The nal_layer_id specifies the identifier of the layer to which the Atlas coded layer (ACL) NAL unit belongs or the identifier of the layer to which a non-ACL NAL unit is applied.
[0649] One less than the nal_temporal_id_plus1 specifies the temporal identifier of the NAL unit.
[0650] Figure 40 Examples of the types of RBSP data structures assigned to the nal_unit_type field are shown. That is, the figure shows the types of the nal_unit_type field in the NAL unit header of the NAL unit included in the sample stream NAL unit.
[0651] In Figure 40 , NAL_TRAIL indicates that the coded patch group of a non-TSA and non-STSA ending Atlas frame is included in the NAL unit. The RBSP syntax structure of the NAL unit is atlas_tile_group_layer_rbsp() or atlas_tile_layer_rbsp(). The type class of the NAL unit is ACL. According to the embodiment, the patch group may be referred to as a patch.
[0652] NAL TSA indicates that the coded patch group of a TSA Atlas frame is included in the NAL unit. The RBSP syntax structure of the NAL unit is atlas_tile_group_layer_rbsp() or atlas_tile_layer_rbsp(). The type class of the NAL unit is ACL.
[0653] NAL_STSA indicates that the coded patch group of an STSA Atlas frame is included in the NAL unit. The RBSP syntax structure of the NAL unit is atlas_tile_group_layer_rbsp() or atlas_tile_layer_rbsp(). The type class of the NAL unit is ACL.
[0654] NAL_RADL indicates that the coded patch group of a RADL Atlas frame is included in the NAL unit. The RBSP syntax structure of the NAL unit is atlas_tile_group_layer_rbsp() or atlas_tile_layer_rbsp(). The type class of the NAL unit is ACL.
[0655] NAL_RASL indicates that the coded patch group of a RASL Atlas frame is included in the NAL unit. The RBSP syntax structure of the NAL unit is atlas_tile_group_layer_rbsp() or aatlas_tile_layer_rbsp(). The type class of the NAL unit is ACL.
[0656] NAL_SKIP indicates that the coded patch group of a skipped Atlas frame is included in the NAL unit. The RBSP syntax structure of the NAL unit is atlas_tile_group_layer_rbsp() or atlas_tile_layer_rbsp(). The type class of the NAL unit is ACL.
[0657] NAL_RSV_ACL_6 to NAL_RSV_ACL_9 indicate that reserved non-IRAP ACL NAL unit types are included in the NAL unit. The type class of the NAL unit is ACL.
[0658] NAL_BLA_W_LP, NAL_BLA_W_RADL, and NAL_BLA_N_LP indicate that the coded tile group of the BLA atlas frame is included in the NAL unit. The RBSP syntax structure of the NAL unit is atlas_tile_group_layer_rbsp() or atlas_tile_layer_rbsp(). The type class of the NAL unit is ACL.
[0659] NAL_GBLA_W_LP, NAL_GBLA_W_RADL, and NAL_GBLA_N_LP indicate that the coded tile group of the GBLA atlas frame may be included in the NAL unit. The RBSP syntax structure of the NAL unit is atlas_tile_group_layer_rbsp() or atlas_tile_layer_rbsp(). The type class of the NAL unit is ACL.
[0660] NAL_IDR_W_RADL and NAL_IDR_N_LP indicate that the coded tile group of the IDR atlas frame is included in the NAL unit. The RBSP syntax structure of the NAL unit is atlas_tile_group_layer_rbsp() or atlas_tile_layer_rbsp(). The type class of the NAL unit is ACL.
[0661] NAL_GIDR_W_RADL and NAL_GIDR_N_LP indicate that the coded tile group of the GIDR atlas frame is included in the NAL unit. The RBSP syntax structure of the NAL unit is atlas_tile_group_layer_rbsp() or atlas_tile_layer_rbsp(). The type class of the NAL unit is ACL.
[0662] NAL_CRA indicates that the coded tile group of the CRA atlas frame is included in the NAL unit. The RBSP syntax structure of the NAL unit is atlas_tile_group_layer_rbsp() or atlas_tile_layer_rbsp(). The type class of the NAL unit is ACL.
[0663] NAL_GCRA indicates that the coded tile group of the GCRA atlas frame is included in the NAL unit. The RBSP syntax structure of the NAL unit is atlas_tile_group_layer_rbsp() or atlas_tile_layer_rbsp(). The type class of the NAL unit is ACL.
[0664] NAL_IRAP_ACL_22 and NAL_IRAP_ACL_23 indicate that the reserved IRAP ACL NAL unit types are included in the NAL unit. The type class of the NAL unit is ACL.
[0665] NAL_RSV_ACL_24 to NAL_RSV_ACL_31 indicate that the reserved non-IRAP ACL NAL unit types are included in the NAL unit. The type class of the NAL unit is ACL.
[0666] NAL_ASPS indicates that ASPS is included in the NAL unit. The RBSP syntax structure of the NAL unit is atlas_sequence_parameter_set_rbsp(). The type class of the NAL unit is non-ACL.
[0667] NAL_AFPS indicates that AFPS is included in the NAL unit. The RBSP syntax structure of the NAL unit is atlas_frame_parameter_set_rbsp(). The type class of the NAL unit is non-ACL.
[0668] NAL_AUD indicates that an access unit delimiter is included in the NAL unit. The RBSP syntax structure of the NAL unit is access_unit_delimiter_rbsp(). The type class of the NAL unit is non-ACL.
[0669] NAL_VPCC_AUD indicates that the V-PCC access unit delimiter is included in the NAL unit. The RBSP syntax structure of the NAL unit is access_unit_delimiter_rbsp(). The type class of the NAL unit is non-ACL.
[0670] NAL_EOS indicates that the NAL unit type can be the end of the sequence. The RBSP syntax structure of the NAL unit is end_of_seq_rbsp(). The type class of the NAL unit is non-ACL.
[0671] NAL_EOB indicates that the NAL unit type can be the end of the bitstream. The RBSP syntax structure of the NAL unit is end_of_atlas_sub_bitstream_rbsp(). The type class of the NAL unit is non-ACL.
[0672] NAL_FD fill indicates that fill_data_rbsp() is included in the NAL unit. The type class of the NAL unit is non-ACL.
[0673] NAL_PREFIX_NSEI and NAL_SUFFIX_NSEI indicate that non-essential supplementary enhancement information is included in the NAL unit. The RBSP syntax structure of the NAL unit is sei_rbsp(). The type class of the NAL unit is non-ACL.
[0674] NAL_PREFIX_ESEI and NAL_SUFFIX_ESEI indicate that essential supplementary enhancement information is included in the NAL unit. The RBSP syntax structure of the NAL unit is sei_rbsp(). The type class of the NAL unit is non-ACL.
[0675] NAL_AAPS indicates that the Atlas adaptation parameter set is included in the NAL unit. The RBSP syntax structure of the NAL unit is atlas_adaptation_parameter_set_rbsp(). The type class of the NAL unit is non-ACL.
[0676] NAL_RSV_NACL_44 to NAL_RSV_NACL_47 indicate that the NAL unit type can be a reserved non-ACL NAL unit type. The type class of the NAL unit is non-ACL.
[0677] NAL_UNSPEC_48 to NAL_UNSPEC_63 indicate that the NAL unit type can be an unspecified non-ACL NAL unit type. The type class of the NAL unit is non-ACL.
[0678] Figure 41 Shows the syntax structure of the Atlas sequence parameter set according to an embodiment.
[0679] Figure 41 Shows the syntax of the RBSP data structure included in the NAL unit when the NAL unit type is an Atlas sequence parameter.
[0680] Each sample stream NAL unit can contain an Atlas parameter set (e.g., ASPS, AAPS, or AFPS), one or more Atlas tile groups (or Atlas tiles), and one of the SEIs.
[0681] ASPS can contain zero or more syntax elements of the complete coded Atlas sequence (CAS) applied to the content determined by the syntax elements (referred to as the syntax elements in each tile group (or tile) header) found in the ASPS. According to an embodiment, the syntax elements can have the same meaning as fields or parameters.
[0682] In Figure 41In it, the asps_atlas_sequence_parameter_set_id field can provide an identifier used to identify the atlas sequence parameter set for other syntax elements to reference.
[0683] The asps_frame_width field indicates the atlas frame width in terms of an integer number of samples, where the samples correspond to the luminance samples of the video component.
[0684] The asps_frame_height field indicates the atlas frame height in terms of an integer number of samples, where the samples correspond to the luminance samples of the video component.
[0685] The asps_log2_patch_packing_block_size field specifies the value of the variable PatchPackingBlockSize, which is used for the horizontal and vertical placement of patches within the atlas.
[0686] The asps_log2_max_atlas_frame_order_cnt_lsb_minus4 field specifies the value of the variable MaxAtlasFrmOrderCntLsb in the decoding process for the atlas frame order count.
[0687] Adding 1 to the asps_max_dec_atlas_frame_buffering_minus1 field specifies the maximum size required for the decoding atlas frame buffer of the CAS, in units of the atlas frame storage buffer.
[0688] The asps_long_term_ref_atlas_frames_flag field being equal to 0 specifies that long-term reference atlas frames are not used for inter prediction of any coded atlas frames in the CAS. The asps_long_term_ref_atlas_frames_flag field being equal to 1 specifies that long-term reference atlas frames can be used for inter prediction of one or more coded atlas frames in the CAS.
[0689] The asps_num_ref_atlas_frame_lists_in_asps field specifies the number of ref_list_struct(rlsIdx) syntax structures included in the atlas sequence parameter set.
[0690] ASAP includes as many iteration statements as the value of the asps_num_ref_atlas_frame_lists_in_asps field. In an implementation, the iteration statement includes ref_list_struct(i). According to an implementation, in the iteration statement, i is initialized to 0 and incremented by 1 each time the iteration statement is executed. The iteration statement is repeated until i reaches the value of the asps_num_ref_atlas_frame_lists_in_asps field.
[0691] will be referred to Figure 50 Describe ref_list_struct(i) in detail.
[0692] When the asps_num_ref_atlas_frame_lists_in_asps field is greater than 0, the atgh_ref_atlas_frame_list_sps_flag field may be included in the Atlas tile group (tile) header. When the asps_num_ref_atlas_frame_lists_in_asps field is greater than 1, the atgh_ref_atlas_frame_list_idx field may be included in the Atlas tile group (tile) header.
[0693] The asps_use_eight_orientations_flag field being equal to 0 specifies that the patch orientation index pdu_orientation_index[i][j] field for the patch with index j in the frame with index i is in the range from 0 to 1 (inclusive). The asps_use_eight_orientations_flag field being equal to 1 specifies that the patch orientation index pdu_orientation_index[i][j] field for the patch with index j in the frame with index i is in the range from 0 to 7 (inclusive).
[0694] The asps_extended_projection_enabled_flag field being equal to 0 specifies that patch projection information is not signaled for the current Atlas tile group. The asps_extended_projection_enabled_flag field being equal to 1 specifies that patch projection information is signaled for the current Atlas tile group.
[0695] The asps_normal_axis_limits_quantization_enabled_flag field being equal to 1 specifies that quantization parameters shall be signaled and used for quantizing the normal axis related elements of a patch data unit, a merged patch data unit, or an inter-patch data unit. If the asps_normal_axis_limits_quantization_enabled_flag field is equal to 0, quantization shall not be applied to any normal axis related elements of a patch data unit, a merged patch data unit, or an inter-patch data unit.
[0696] When the asps_normal_axis_limits_quantization_enabled_flag field is 1, the atgh_pos_min_z_quantizer field may be included in the atlas tile group (or tile) header.
[0697] The asps_normal_axis_max_delta_value_enabled_flag field being equal to 1 specifies that the maximum nominal shift value of the normal axis that may be present in the geometry information of a patch with index i in a frame with index j shall be indicated in the bitstream of each patch data unit, merged patch data unit, or inter-patch data unit. If the asps_normal_axis_max_delta_value_enabled_flag field is equal to 0, the maximum nominal shift value of the normal axis that may be present in the geometry information of a patch with index i in a frame with index j shall not be indicated in the bitstream of each patch data unit, merged patch data unit, or inter-patch data unit.
[0698] When the asps_normal_axis_max_delta_value_enabled_flag field is 1, the atgh_pos_delta_max_z_quantizer field may be included in the atlas tile group (or tile) header.
[0699] The asps_remove_duplicate_point_enabled_flag field being equal to 1 indicates that duplicate points are not constructed for the current atlas, where a duplicate point is a point having the same 2D and 3D geometric coordinates as another point from a lower index map. The asps_remove_duplicate_point_enabled_flag field being equal to 0 indicates that all points are reconstructed.
[0700] Incrementing the asps_max_dec_atlas_frame_buffering_minus1 field by 1 specifies the maximum size required for decoding the atlas frame buffer of the CAS, in terms of the atlas frame storage buffer.
[0701] The asps_pixel_deinterleaving_flag field being equal to 1 indicates that the decoded geometry and attribute video of the current atlas contains spatially interleaved pixels from two maps. The asps_pixel_deinterleaving_flag field being equal to 0 indicates that the decoded geometry and attribute video corresponding to the current atlas contains pixels from only a single map.
[0702] The asps_patch_precedence_order_flag field being equal to 1 indicates that the patch precedence of the current atlas is the same as the decoding order. The asps_patch_precedence_order_flag field being equal to 0 indicates that the patch precedence of the current atlas is opposite to the decoding order.
[0703] The asps_patch_size_quantizer_present_flag field being equal to 1 indicates that the patch size quantization parameter is present in the atlas tile group header. The asps_patch_size_quantizer_present_flag field being equal to 0 indicates that the patch size quantization parameter is not present.
[0704] When the asps_patch_size_quantizer_present_flag field is equal to 1, the atgh_patch_size_x_info_quantizer field and the atgh_patch_size_y_info_quantizer field may be included in the atlas tile group (or tile) header.
[0705] The asps_eom_patch_enabled_flag field being equal to 1 indicates that the decoded occupancy map video of the current atlas contains information regarding whether the intermediate depth position between two depth maps is occupied. The asps_eom_patch_enabled_flag field being equal to 0 indicates that the decoded occupancy map video does not contain information regarding whether the intermediate depth position between two depth maps is occupied.
[0706] The asps_raw_patch_enabled_flag field being equal to 1 indicates that the decoded geometry and attribute video of the current atlas contains information related to RAW coding points. The asps_raw_patch_enabled_flag field being equal to 0 indicates that the decoded geometry and attribute video does not contain information related to RAW coding points.
[0707] When the asps_eom_patch_enabled_flag field or the asps_raw_patch_enabled_flag field is equal to 1, the asps_auxiliary_video_enabled_flag field can be included in the atlas sequence parameter set syntax.
[0708] The asps_auxiliary_video_enabled_flag field being equal to 1 indicates that information related to RAW and EOM patch types can be placed in the auxiliary video sub-bitstream. The asps_auxiliary_video_enabled_flag field being equal to 0 indicates that information related to RAW and EOM patch types can only be placed in the main video sub-bitstream.
[0709] The asps_point_local_reconstruction_enabled_flag field being equal to 1 indicates that point local reconstruction mode information can be present in the bitstream of the current atlas. The asps_point_local_reconstruction_enabled_flag field being equal to 0 indicates that there is no information related to the point local reconstruction mode in the bitstream of the current atlas.
[0710] The asps_map_count_minus1 field plus 1 indicates the number of maps that can be used to encode the geometry and attribute data of the current atlas.
[0711] When the asps_pixel_deinterleaving_enabled_flag field is equal to 1, for each value of the asps_map_count_minus1 field, the asps_pixel_deinterleaving_map_flag[j] field can be included in the ASPS.
[0712] The asps_pixel_deinterleaving_map_flag[i] field being equal to 1 indicates that the decoded geometry and attribute video corresponding to the picture with index i in the current atlas contains spatially interleaved pixels corresponding to two pictures. The asps_pixel_deinterleaving_map_flag[i] field being equal to 0 indicates that the decoded geometry and attribute video corresponding to the picture with index i in the current atlas contains pixels corresponding to a single picture.
[0713] When the asps_eom_patch_enabled_flag field and the asps_map_count_minus1 field are equal to 0, the ASPS may also include the asps_eom_fix_bit_count_minus1 field.
[0714] The asps_eom_fix_bit_count_minus1 field plus 1 indicates the bit size of the EOM codeword.
[0715] When the asps_point_local_reconstruction_enabled_flag field is equal to 1, the asps_point_local_reconstruction_information(asps_map_count_minus1) may be included in the ASPS and sent.
[0716] When the asps_pixel_deinterleaving_enabled_flag field or the asps_point_local_reconstruction_enabled_flag field is equal to 1, the ASPS may also include the asps_surface_thickness_minus1 field.
[0717] The asps_surface_thickness_minus1 field plus 1 specifies the maximum absolute difference between the explicitly encoded depth value and the interpolated depth value when the asps_pixel_deinterleaving_enabled_flag field or the asps_point_local_reconstruction_enabled_flag field is equal to 1.
[0718] The asps_vui_parameters_present_flag field being equal to 1 indicates the presence of the vui_parameters() syntax structure. The asps_vui_parameters_present_flag field being equal to 0 indicates the absence of the vui_parameters() syntax structure.
[0719] The asps_extension_flag field being equal to 0 indicates the absence of the asps_extension_data_flag field in the ASPS RBSP syntax structure.
[0720] The asps_extension_data_flag field indicates that data for extension is included in the ASPS RBSP syntax structure.
[0721] The rbsp_trailing_bits field is used to pad the remaining bits with 0s after adding 1 (which is the stop bit) for byte alignment to indicate the end of the RBSP data.
[0722] Figure 42 Shows an Atlas frame parameter set according to an embodiment.
[0723] Figure 42 Shows the syntax structure of the Atlas frame parameter set included in the NAL unit when the NAL unit type (nal_unit_type) is NAL_AFPS as Figure 40 shown.
[0724] In Figure 42 the Atlas frame parameter set (AFPS) contains a syntax structure that contains syntax elements applied to all zero or more complete coded Atlas frames.
[0725] The afps_atlas_frame_parameter_set_id field can provide an identifier for identifying the Atlas frame parameter set for reference by other syntax elements. That is, an identifier that can be referenced by other syntax elements can be provided through the AFPS Atlas frame parameter set.
[0726] The afps_atlas_sequence_parameter_set_id field specifies the value of asps_atlas_sequence_parameter_set_id of the active Atlas sequence parameter set.
[0727] Will refer to Figure 43 to describe atlas_frame_tile_information().
[0728] The afps_output_flag_present_flag field being equal to 1 indicates that the atgh_atlas_output_flag or ath_atlas_output_flag field is present in the associated tile group (or tile) header. The afps_output_flag_present_flag field being equal to 0 indicates that the atgh_atlas_output_flag field or ath_atlas_output_flag field is not present in the associated tile group (or tile) header.
[0729] Adding 1 to the afps_num_ref_idx_default_active_minus1 field specifies the inferred value of the variable NumRefIdxActive for tile groups where the atgh_num_ref_idx_active_override_flag field is equal to 0.
[0730] The afps_additional_lt_afoc_lsb_len field specifies the value of the variable MaxLtAtlasFrmOrderCntLsb used during the decoding of the reference atlas frame.
[0731] Adding 1 to the afps_3d_pos_x_bit_count_minus1 field specifies the number of bits in the fixed-length representation of the pdu_3d_pos_x[j] field for the patch with index j in the atlas tile group that references the afps_atlas_frame_parameter_set_id field.
[0732] Adding 1 to the afps_3d_pos_y_bit_count_minus1 field specifies the number of bits in the fixed-length representation of the pdu_3d_pos_y[j] field for the patch with index j in the atlas tile group that references the afps_atlas_frame_parameter_set_id field.
[0733] The afps_lod_mode_enabled_flag field being equal to 1 indicates that LOD parameters may be present in the patch. The afps_lod_mode_enabled_flag field being equal to 0 indicates that LOD parameters are not present in the patch.
[0734] The afps_override_eom_for_depth_flag field being equal to 1 indicates that the values of the afps_eom_number_of_patch_bit_count_minus1 field and the afps_eom_max_bit_count_minus1 field are explicitly present in the bitstream. The afps_override_eom_for_depth_flag field being equal to 0 indicates that the values of the afps_eom_number_of_patch_bit_count_minus1 field and the afps_eom_max_bit_count_minus1 field are implicitly derived.
[0735] One plus the afps_eom_number_of_patch_bit_count_minus1 field specifies the number of bits used to represent the number of geometric patches associated with the EOM attribute patches in the atlas frame associated with this atlas frame parameter set.
[0736] One plus the afps_eom_max_bit_count_minus1 field specifies the number of bits used to represent the number of EOM points for each geometric patch associated with the EOM attribute patches in the atlas frame associated with this atlas frame parameter set.
[0737] The afps_raw_3d_pos_bit_count_explicit_mode_flag field being equal to 1 indicates that the number of bits in the fixed-length representation of the rpdu_3d_pos_x field, the rpdu_3d_pos_y field, and the rpdu_3d_pos_z field is explicitly encoded by the atgh_raw_3d_pos_axis_bit_count_minus1 field in the atlas tile group header that references the afps_atlas_frame_parameter_set_id field. The afps_raw_3d_pos_bit_count_explicit_mode_flag field being equal to 0 indicates that the value of the atgh_raw_3d_pos_axis_bit_count_minus1 field is implicitly derived.
[0738] When the afps_raw_3d_pos_bit_count_explicit_mode_flag field is equal to 1, the atgh_raw_3d_pos_axis_bit_count_minus1 field may be included in the atlas tile group (or tile) header.
[0739] The afps_fixed_camera_model_flag indicates whether there is a fixed camera model.
[0740] The afps_extension_flag field being equal to 0 indicates that the afps_extension_data_flag field does not exist in the AFPS RBSP syntax structure.
[0741] The afps_extension_data_flag field may contain extension-related data.
[0742] Figure 43 Shows the syntax structure of atlas_frame_tile_information according to an embodiment.
[0743] Figure 43 Shows Figure 42 the syntax of atlas_frame_tile_information contained in
[0744] The afti_single_tile_in_atlas_frame_flag field being equal to 1 specifies that there is only one tile in each atlas frame that references AFPS. The afti_single_tile_in_atlas_frame_flag field being equal to 0 indicates that there is more than one tile in each atlas frame that references AFPS.
[0745] The afti_uniform_tile_spacing_flag field being equal to 1 specifies that the tile column and row boundaries are evenly distributed over the atlas frame and are signaled using the syntax elements afti_tile_cols_width_minus1 field and afti_tile_rows_height_minus1 field respectively. The afti_uniform_tile_spacing_flag field being equal to 0 specifies that the tile column and row boundaries may or may not be evenly distributed over the atlas frame and are signaled using the afti_num_tile_columns_minus1 field and afti_num_tile_rows_minus1 field and a list of syntax element pairs afti_tile_column_width_minus1[i] and afti_tile_row_height_minus1[i].
[0746] Adding 1 to the afti_tile_cols_width_minus1 field specifies the width of the tile column of the rightmost tile column excluding the atlas frame in units of 64 samples when the afti_uniform_tile_spacing_flag field is equal to 1.
[0747] Adding 1 to the afti_tile_rows_height_minus1 field specifies the height of the tile row of the bottom tile row excluding the atlas frame in units of 64 samples when the afti_uniform_tile_spacing_flag field is equal to 1.
[0748] Adding 1 to the afti_num_tile_columns_minus1 field specifies the number of tile columns for partitioning the atlas frame when the afti_uniform_tile_spacing_flag field is equal to 0.
[0749] Adding 1 to the afti_num_tile_rows_minus1 field specifies the number of tile rows for partitioning the atlas frame when the pti_uniform_tile_spacing_flag field is equal to 0.
[0750] Adding 1 to the afti_tile_column_width_minus1[i] field specifies the width of the i-th tile column in units of 64 samples.
[0751] Adding 1 to the afti_tile_row_height_minus1[i] field specifies the height of the i-th tile row in units of 64 samples.
[0752] The afti_single_tile_per_tile_group_flag field being equal to 1 specifies that each tile group referring to this AFPS includes one tile (or tile partition). The afti_single_tile_per_tile_group_flag field being equal to 0 specifies that the tile group referring to this AFPS can include more than one tile (or tile partition).
[0753] When the afti_single_tile_per_tile_group_flag field is equal to 0, the afti_num_tile_groups_in_atlas_frame_minus1 field is carried in the atlas frame tile information.
[0754] Incrementing the afti_num_tile_groups_in_atlas_frame_minus1 field by 1 specifies the number of tile groups (or tiles) in each atlas frame that reference the AFPS.
[0755] For each value of the afti_num_tile_groups_in_atlas_frame_minus1 field, the afti_top_left_tile_idx[i] field and the afti_bottom_right_tile_idx_delta[i] field may also be included in the AFPS.
[0756] The afti_top_left_tile_idx[i] field specifies the tile index of the tile located at the top left corner of the i-th tile group.
[0757] The afti_bottom_right_tile_idx_delta[i] field specifies the difference between the tile index of the tile located at the bottom right corner of the i-th tile group and the afti_top_left_tile_idx[i] field.
[0758] Setting the afti_signalled_tile_group_id_flag field equal to 1 signals the tile group ID for each tile group or the tile ID for each tile.
[0759] When the afti_signalled_tile_group_id_flag field is 1, the afti_signalled_tile_group_id_length_minus1 field and the afti_tile_group_id[i] field may be carried in the atlas frame tile information.
[0760] Incrementing the afti_signalled_tile_group_id_length_minus1 field by 1 specifies the number of bits used to represent the afti_tile_group_id[i] field (when present) and the atgh_address syntax element in the tile group header.
[0761] The afti_tile_group_id[i] field specifies the tile group ID of the i-th tile group. The length of the afti_tile_group_id[i] field is afti_signalled_tile_group_id_length_minus1 + 1 bits.
[0762] In Figure 43In this case, the tile group can be referred to as a tile, and the tile can be referred to as a partition (or tile partition). In this case, afti_tile_group_id[i] indicates the ID of the i-th tile.
[0763] Figure 44 Shows an atlas adaptation parameter set (atlas_adaptation_parameter_set_rbsp()) according to an embodiment.
[0764] Figure 44 Shows the syntax structure of an atlas adaptation parameter set (AAPS) carried by an NAL unit when the NAL unit type (nal_unit_type) is NAL_AAPS.
[0765] The AAPS RBSP includes parameters that can be referenced by one or more coded tile group (or tile) NAL units encoding atlas frames. At any given moment during the operation of the decoding process, at most one AAPS RBSP is considered active, and the activation of any particular AAPS RBSP causes the deactivation of the previously active AAPS RBSP.
[0766] In Figure 44 the aaps_atlas_adaptation_parameter_set_id field provides an identifier for identifying the atlas adaptation parameter set for reference by other syntax elements.
[0767] The aaps_atlas_sequence_parameter_set_id field specifies the value of the asps_atlas_sequence_parameter_set_id field of the active atlas sequence parameter set.
[0768] The aaps_camera_parameters_present_flag field being equal to 1 specifies that camera parameters (atlas_camera_parameters) are present in the current atlas adaptation parameter set. The aaps_camera_parameters_present_flag field being equal to 0 specifies that the camera parameters of the current adaptation parameter set are not present. atlas_camera_parameters will be described with reference to Figure 45 below.
[0769] The aaps_extension_flag field being equal to 0 indicates that the aaps_extension_data_flag field does not exist in the AAPS RBSP syntax structure.
[0770] The aaps_extension_data_flag field can indicate whether the AAPS contains extension-related data.
[0771] Figure 45 Shows atlas_camera_parameters according to an embodiment.
[0772] Figure 45 Shows Figure 44 the detailed syntax of atlas_camera_parameters.
[0773] In Figure 45 it, the acp_camera_model field indicates the camera model of the point cloud frame associated with the current atlas adaptation parameter set.
[0774] Figure 46 Is a table showing examples of camera models assigned to the acp_camera_model field according to an embodiment.
[0775] For example, the acp_camera_model field being equal to 0 indicates that the camera model is not specified.
[0776] The acp_camera_model field being equal to 1 indicates that the camera model is an orthographic camera model.
[0777] When the acp_camera_model field is 2 - 255, the camera model can be reserved.
[0778] According to an embodiment, when the value of the acp_camera_model field is equal to 1, the camera parameters can also include the acp_scale_enabled_flag field, the acp_offset_enabled_flag field, and / or the acp_rotation_enabled_flag field related to scaling, offset, and rotation.
[0779] The acp_scale_enabled_flag field being equal to 1 indicates the existence of the scaling parameter of the current camera model. The acp_scale_enabled_flag field being equal to 0 indicates the non-existence of the scaling parameter of the current camera model.
[0780] When the acp_scale_enabled_flag field is equal to 1, the acp_scale_on_axis[d] field can be included in the atlas camera parameters for each d value.
[0781] The `acp_scale_on_axis[d]` field specifies the scaling value `Scale[d]` along axis `d` for the current camera model. The value of `d` ranges from 0 to 2 (inclusive), where the values 0, 1, and 2 correspond to the X, Y, and Z axes, respectively.
[0782] The `acp_offset_enabled_flag` field being equal to 1 indicates the existence of the offset parameter for the current camera model. The `acp_offset_enabled_flag` field being equal to 0 indicates the non-existence of the offset parameter for the current camera model.
[0783] When the `acp_offset_enabled_flag` field is equal to 1, for each value of `d`, the `acp_offset_on_axis[d]` field may be included in the Atlas camera parameters.
[0784] The `acp_offset_on_axis[d]` field indicates the value of the offset `Offset[d]` along axis `d` for the current camera model, where `d` is in the range from 0 to 2 (inclusive). The values of `d` equal to 0, 1, and 2 correspond to the X, Y, and Z axes, respectively.
[0785] The `acp_rotation_enabled_flag` field being equal to 1 indicates the existence of the rotation parameter for the current camera model. The `acp_rotation_enabled_flag` field being equal to 0 indicates the non-existence of the rotation parameter for the current camera model.
[0786] When the `acp_rotation_enabled_flag` field is equal to 1, the Atlas camera parameters may also include the `acp_rotation_qx` field, the `acp_rotation_qy` field, and the `acp_rotation_qz` field.
[0787] The `acp_rotation_qx` field specifies the x-component `qX` of the rotation for the current camera model in quaternion representation.
[0788] The `acp_rotation_qy` field specifies the y-component `qY` of the rotation for the current camera model in quaternion representation.
[0789] The `acp_rotation_qz` field specifies the z-component `qZ` of the rotation for the current camera model in quaternion representation.
[0790] The `atlas_camera_parameters()` as described above may be included in at least one SEI message and sent.
[0791] Figure 47Shows an atlas_tile_group_layer according to an embodiment.
[0792] Figure 47 Shows according to Figure 40 The syntax structure of an atlas_tile_group_layer or atlas_tile_layer carried in a NAL unit with the NAL unit type shown.
[0793] According to an embodiment, a tile group may correspond to a tile. In the present disclosure, the term "tile group" may be referred to as the term "tile". Similarly, the term "atgh" may be interpreted as the term "ath".
[0794] The atlas_tile_group_layer field or atlas_tile_layer field may contain an atlas_tile_group_header (atlas_tile_group_header) or an atlas_tile_header (atlas_tile_header). The atlas tile group (or tile) header (atlas_tile_group_header or atlas_tile_header) will be described with reference to Figure 48 Describe the atlas tile group (or tile) header (atlas_tile_group_header or atlas_tile_header).
[0795] When the atgh_type field of the atlas tile group (or tile) is not SKIP_TILE_GRP, the atlas tile group (or tile) data may be included in the atlas_tile_group_layer or atlas_tile_layer.
[0796] Figure 48 Shows an example of the syntax structure of an atlas tile group (or tile) header (atlas_tile_group_header() or atlas_tile_header()) included in an atlas tile group layer according to an embodiment.
[0797] In Figure 48 the atgh_atlas_frame_parameter_set_id field specifies the value of the afps_atlas_frame_parameter_set_id field of the active atlas frame parameter set of the current atlas block group.
[0798] The value of the atgh_atlas_adaptation_parameter_set_id field specifies the value of the aaps_atlas_adaptation_parameter_set_id field for the active atlas adaptation parameter set for the current atlas tile group.
[0799] The atgh_address field specifies the tile group (or tile) address of the tile group (or tile). When it does not exist, the value of the atgh_address field is inferred to be equal to 0. The tile group (or tile) address is the tile group ID (or tile ID) of the tile group (or tile). The length of the atgh_address field is the afti_signalled_tile_group_id_length_minus1 field + 1 bit. If the afti_signalled_tile_group_id_flag field is equal to 0, the value of the atgh_address field is in the range from 0 to the afti_num_tile_groups_in_atlas_frame_minus1 field (inclusive). Otherwise, the value of the atgh_address field is in the range from 0 to 2(afti_signalled_tile_group_id_length_minus1 field + 1) - 1 (inclusive). The afti_signalled_tile_group_id_length_minus1 field and the afti_signalled_tile_group_id_flag field are included in AFTI.
[0800] The atgh_type field specifies the coding type of the current atlas tile group (tile).
[0801] Figure 49 Examples of the coding types assigned to the atgh_type field according to an embodiment are shown.
[0802] For example, when the value of the atgh_type field is 0, the coding type of the atlas tile group (or tile) is P_TILE_GRP (inter-frame atlas tile group (or tile)).
[0803] When the value of the atgh_type field is 1, the coding type of the atlas tile group (or tile) is I_TILE_GRP (intra-frame atlas tile group (or tile)).
[0804] When the value of the atgh_type field is 2, the coding type of the atlas tile group (or tile) is SKIP_TILE_GRP (skip atlas tile group (or tile)).
[0805] When the value of the afps_output_flag_present_flag field included in AFTI is equal to 1, the Atlas tile group (or tile) header may further include an atgh_atlas_output_flag field.
[0806] The atgh_atlas_output_flag field affects the decoded Atlas output and removal process.
[0807] The atgh_atlas_frm_order_cnt_lsb field specifies the Atlas frame order count modulo MaxAtlasFrmOrderCntLsb for the current Atlas tile group.
[0808] According to an embodiment, when the value of the asps_num_ref_atlas_frame_lists_in_asps field included in the Atlas Sequence Parameter Set (ASPS) is greater than 1, the Atlas tile group (or tile) header may further include an atgh_ref_atlas_frame_list_sps_flag field. asps_num_ref_atlas_frame_lists_in_asps specifies the number of ref_list_struct (rlsIdx) syntax structures included in the ASPS.
[0809] The atgh_ref_atlas_frame_list_sps_flag field being equal to 1 specifies that the reference Atlas frame list for the current Atlas tile group (or tile) is derived based on one of the ref_list_struct (rlsIdx) syntax structures in the active ASPS. The atgh_ref_atlas_frame_list_sps_flag field being equal to 0 specifies that the reference Atlas frame list for the current atlastile list is derived based on the ref_list_struct (rlsIdx) syntax structure directly included in the tile group header of the current Atlas tile group.
[0810] According to an embodiment, when the atgh_ref_atlas_frame_list_sps_flag field is equal to 0, the Atlas tile group (or tile) header includes ref_list_struct (asps_num_ref_atlas_frame_lists_in_asps), and when the atgh_ref_atlas_frame_list_sps_flag is greater than 1, the Atlas tile group (or tile) header includes the atgh_ref_atlas_frame_list_sps_flag field.
[0811] The atgh_ref_atlas_frame_list_idx field specifies the index of the ref_list_struct (rlsIdx) syntax structure, which is used to derive the reference Atlas frame list of the current Atlas tile group (or tile). The reference Atlas frame list is a list of ref_list_struct (rlsIdx) syntax structures included in the active ASPS.
[0812] According to an embodiment, according to the value of the NumLtrAtlasFrmEntries field, the Atlas tile group (or tile) header further includes the atgh_additional_afoc_lsb_present_flag[j] field, and if the atgh_additional_afoc_lsb_present_flag[j] is equal to 1, the atgh_additional_afoc_lsb_val[j] field may also be included.
[0813] The atgh_additional_afoc_lsb_val[j] specifies the value of the FullAtlasFrmOrderCntLsbLt[RlsIdx][j] of the current Atlas tile group (or tile).
[0814] According to an embodiment, if the atgh_type field does not indicate SKIP_TILE_GRP, the Atlas tile group (or tile) header may further include an atgh_pos_min_z_quantizer field, an atgh_pos_delta_max_z_quantizer field, an atgh_patch_size_x_info_quantizer field, an atgh_patch_size_y_info_quantizer field, an atgh_raw_3d_pos_axis_bit_count_minus1 field, and / or an atgh_num_ref_idx_active_minus1 field, depending on the information included in the ASPS or AFPS.
[0815] According to an embodiment, when the value of the asps_normal_axis_limits_quantization_enabled_flag field included in the ASPS is 1, the atgh_pos_min_z_quantizer field is included. When the values of both the asps_normal_axis_limits_quantization_enabled_flag and asps_axis_max_enabled_flag fields included in the ASPS are 1, the atgh_pos_delta_max_z_quantizer field is included.
[0816] According to an embodiment, when the value of the asps_patch_size_quantizer_present_flag field included in the ASPS is 1, the atgh_patch_size_x_info_quantizer field and the atgh_patch_size_y_info_quantizer field are included. When the value of the afps_raw_3d_pos_bit_count_explicit_mode_flag field included in the AFPS is 1, the atgh_raw_3d_pos_axis_bit_explicit_minus1 field is included.
[0817] According to an embodiment, when the atgh_type field indicates P_TILE_GRP and num_ref_entries[RlsIdx] is greater than 1, the Atlas tile group (or tile) header further includes the atgh_num_ref_idx_active_override_flag field. When the value of the atgh_num_ref_idx_active_override_flag field is 1, the atgh_num_ref_idx_active_minus1 field is included in the Atlas tile group (or tile) header.
[0818] The atgh_pos_min_z_quantizer specifies the quantizer to be applied to pdu_3d_pos_min_z[p] of the patch with index p. If the atgh_pos_min_z_quantizer field does not exist, its value can be inferred to be equal to 0.
[0819] The atgh_pos_delta_max_z_quantizer field specifies the quantizer to be applied to the value of the pdu_3d_pos_delta_max_z[p] field of the patch with index p. If the atgh_pos_delta_max_z_quantizer field does not exist, its value can be inferred to be equal to 0.
[0820] The atgh_patch_size_x_info_quantizer field specifies the value of the quantizer PatchSizeXQuantizer to be applied to the variables pdu_2d_size_x_minus1[p], mpdu_2d_delta_size_x[p], ipdu_2d_delta_size_x[p], rpdu_2d_size_x_minus1[p], and epdu_2d_size_x.minus1[p] of the patch with index p. If the atgh_patch_size_x_info_quantizer field does not exist, its value can be inferred to be equal to the asps_log2_patch_packing_block_size field.
[0821] The atgh_patch_size_y_info_quantizer field specifies the value of the variable PatchSizeYQuantizer of the variables pdu_2d_size_y_minus1[p], mpdu_2d_delta_size_y[p], ipdu_2d_delta_size_y[p], rpdu_2d_size_y_minus1[p], and epdu_2d_size_y_minus1[p] to be applied to the patch with index p. If the atgh_patch_size_y_info_quantizer field does not exist, its value can be inferred to be equal to the asps_log2_patch_packing_block_size field.
[0822] The atgh_raw_3d_pos_axis_bit_count_minus1 field plus 1 specifies the number of bits in the fixed-length representation of rpdu_3d_pos_x, rpdu_3d_pos_y, and rpdu_3d_pos_z.
[0823] The atgh_num_ref_idx_active_override_flag field being equal to 1 specifies that the atgh_num_ref_idx_active_minus1 field of the syntax element exists in the current atlas tile group (or tile). The atgh_num_ref_idx_active_override_flag field being equal to 0 specifies that the atgh_num_ref_idx_active_minus1 field of the syntax element does not exist. If the atgh_num_ref_idx_active_override_flag field does not exist, its value can be inferred to be equal to 0.
[0824] The atgh_num_ref_idx_active_minus1 field plus 1 specifies the maximum reference index that can be used for the reference atlas frame list for decoding the current atlas tile group. When the value of the atgh_num_ref_idx_active_minus1 field is equal to 0, the reference index of the reference atlas frame list can be not used for decoding the current atlas tile group (or tile).
[0825] Byte alignment is used to fill the remaining bits with 0s after adding 1 (which is the stop bit) for byte alignment to indicate the end of the data.
[0826] As described above, one or more ref_list_struct(rlsIdx) syntax structures may be included in the ASPS and / or may be directly included in the atlas tile group (or tile) header.
[0827] Figure 50 An example of the syntax structure of ref_list_struct() according to an embodiment is shown.
[0828] In Figure 50 the num_ref_entries[rlsIdx] field specifies the number of entries in the ref_list_struct(rlsIdx) syntax structure.
[0829] As many of the following elements as the value of the num_ref_entries[rlsIdx] field may be included in the reference list structure.
[0830] When the asps_long_term_ref_atlas_frames_flag field is equal to 1, the reference atlas frame flag (st_ref_atlas_frame_flag[rlsIdx][i]) may be included in the reference list structure.
[0831] When the i-th entry is the first short-term reference atlas frame entry in the ref_list_struct(rlsIdx) syntax structure, the abs_delta_afoc_st[rlsIdx][i] field specifies the absolute difference between the current atlas tile group and the atlas frame order count value of the atlas frame referenced by the i-th entry. When the i-th entry is a short-term reference atlas frame entry but not the first short-term reference atlas frame entry in ref_list_struct(rlsIdx), this field specifies the absolute difference between the atlas frame order count values of the atlas frames referenced by the i-th entry and the previous short-term reference atlas frame entry in the ref_list_struct(rlsIdx) syntax structure.
[0832] When the st_ref_atlas_frame_flag[rlsIdx][i] field is equal to 1, the abs_delta_afoc_st[rlsIdx][i] field may be included in the reference list structure.
[0833] The abs_delta_afoc_st[rlsIdx][i] field specifies the absolute difference between the current atlas tile group and the atlas frame sequence count value of the atlas frame referenced by the i-th entry when the i-th entry is the first short-term reference atlas frame entry in the ref_list_struct(rlsIdx) syntax structure, or specifies the absolute difference between the i-th entry in the ref_list_struct(rlsIdx) syntax structure and the atlas frame sequence count value of the atlas frame referenced by the previous short-term reference atlas frame entry when the i-th entry is a short-term reference atlas frame entry but not the first short-term reference atlas frame entry in the ref_list_struct(rlsIdx) syntax structure.
[0834] When the value of the abs_delta_afoc_st[rlsIdx][i] field is greater than 0, the entry symbol flag (strpf_entry_sign_flag[rlsIdx][i]) field may be included in the reference list structure.
[0835] The strpf_entry_sign_flag[rlsIdx][i] field being equal to 1 specifies that the i-th entry in the ref_list_struct(rlsIdx) field has a value greater than or equal to 0. The strpf_entry_sign_flag[rlsIdx][i] field being equal to 0 specifies that the i-th entry in the ref_list_struct(rlsIdx) has a value less than 0. When not present, it can be inferred that the value of the strpf_entry_sign_flag[rlsIdx][i] field is equal to 1.
[0836] When the asps_long_term_ref_atlas_frames_flag included in ASPS is equal to 0, the afoc_lsb_lt[rlsIdx][i] field may be included in the reference list structure.
[0837] The afoc_lsb_lt[rlsIdx][i] field specifies the value of the atlas frame sequence count modulo MaxAtlasFrmOrderCntLsb of the atlas frame referenced by the i-th entry in the ref_list_struct(rlsIdx) field. The length of the afoc_lsb_lt[rlsIdx][i] field is asps_log2_max_atlas_frame_order_cnt_lsb_minus4 + 4 bits.
[0838] Figure 51Shows the atlas tile group data (atlas_tile_group_data_unit) according to an embodiment.
[0839] Figure 51 Shows the syntax of the atlas tile group data (atlas_tile_group_data_unit()) included in the Figure 47 atlas tile group (or tile) layer. The atlas tile group data may correspond to the atlas tile data, and the tile group may be referred to as a tile.
[0840] In Figure 51 , as p increments by 1 from 0, the atlas-related elements (or fields) according to the index p may be included in the atlas tile group (or tile) data.
[0841] The atgdu_patch_mode[p] field indicates the patch mode of the patch with index p in the current atlas tile group. When the atgh_type field included in the atlas tile group (or tile) header indicates SKIP_TILE_GRP, this indicates that the entire tile group (or tile) information is directly copied from the tile group (or tile) having the same atgh_address as the current tile group (or tile) corresponding to the first reference atlas frame.
[0842] When the atgdu_patch_mode[p] field is not I_END and the atgdu_patch_mode[p] is not P_END, patch_information_data and atgdu_patch_mode[p] may be included in the atlas tile group data (or atlas tile data) for each index p.
[0843] Figure 52 Shows an example of the patch mode types assigned to the atgdu_patch_mode field when the atgh_type field indicates I_TILE_GRP according to an embodiment.
[0844] For example, the atgdu_patch_mode field equal to 0 indicates a non-predictive patch mode with the identifier I_INTRA.
[0845] The atgdu_patch_mode field equal to 1 indicates a RAW point patch mode with the identifier I_RAW.
[0846] The atgdu_patch_mode field equal to 2 indicates an EOM point patch mode with the identifier I_EOM.
[0847] The atgdu_patch_mode field being equal to 14 indicates the patch termination mode with the identifier I_END.
[0848] Figure 53 An example of the patch mode type assigned to the atgdu_patch_mode field when the atgh_type field indicates P_TILE_GRP according to an embodiment is shown.
[0849] For example, the atgdu_patch_mode field being equal to 0 indicates the patch skip mode with the identifier P_SKIP.
[0850] The atgdu_patch_mode field being equal to 1 indicates the patch merge mode with the identifier P_MERGE.
[0851] The atgdu_patch_mode field being equal to 2 indicates the inter-frame prediction patch mode with the identifier P_INTER.
[0852] The atgdu_patch_mode field being equal to 3 indicates the non-prediction patch mode with the identifier P_INTRA.
[0853] The atgdu_patch_mode field being equal to 4 indicates the RAW point patch mode with the identifier P_RAW.
[0854] The atgdu_patch_mode field being equal to 5 indicates the EOM point patch mode with the identifier P_EOM.
[0855] The atgdu_patch_mode field being equal to 14 indicates the patch termination mode with the identifier P_END.
[0856] Figure 54 An example of the patch mode type assigned to the atgdu_patch_mode field when the atgh_type field indicates SKIP_TILE_GRP according to an embodiment is shown.
[0857] For example, atgdu_patch_mode being equal to 0 indicates the patch skip mode with the identifier P_SKIP.
[0858] According to an embodiment, the Atlas tile group (or tile) data unit may further include the AtgduTotalNumberOfPatches field. The AtgduTotalNumberOfPatches field indicates the number of patches and may be set to the final value of p.
[0859] Figure 55Shows patch information data (patch_information_data(patchIdx, patchMode)) according to an embodiment.
[0860] Figure 55 Shows an example of the syntax structure of patch information data (patch_information_data(p, atgdu_patch_mode[p])) included in the Figure 51 patch group (or patch) data unit. In the patch_information_data(p, atgdu_patch_mode[p]) of Figure 51 , p corresponds to the Figure 55 patchIdx, and atgdu_patch_mode[p] corresponds to the Figure 55 patchMode.
[0861] For example, when the atgh_type field indicates SKIP_TILE_GR, skip_patch_data_unit(patchIdx) is included as patch information data.
[0862] When the atgh_type field indicates P_TILE_GR, one of skip_patch_data_unit(patchIdx), merge_patch_data_unit(patchIdx), patch_data_unit(patchIdx), inter_patch_data_unit(patchIdx), raw_patch_data_unit(patchIdx), and eom_patch_data_unit(patchIdx) can be included as patch information data according to the patchMode.
[0863] For example, when patchMode indicates the patch skip mode (P_SKIP), skip_patch_data_unit(patchIdx) is included. When patchMode indicates the patch merge mode (P_MERGE), merge_patch_data_unit(patchIdx) is included. When patchMode indicates P_INTRA, patch_data_unit(patchIdx) is included. When patchMode indicates P_INTER, inter_patch_data_unit(patchIdx) is included. When patchMode indicates the RAW dot patch mode (P_RAW), raw_patch_data_unit(patchIdx) is included. When patchMode is the EOM dot patch mode (P_EOM), eom_patch_data_unit(patchIdx) is included.
[0864] When the atgh_type field indicates I_TILE_GR, one of patch_data_unit(patchIdx), raw_patch_data_unit(patchIdx), and eom_patch_data_unit(patchIdx) can be included as patch information data according to patchMode.
[0865] For example, when patchMode indicates I_INTRA, patch_data_unit(patchIdx) is included. When patchMode indicates the RAW dot patch mode (I_RAW), raw_patch_data_unit(patchIdx) is included. When patchMode indicates the EOM dot patch mode (I_EOM), eom_patch_data_unit(patchIdx) is included.
[0866] Figure 56 The syntax structure of the patch data unit (patch_data_unit(patchIdx)) according to an embodiment is shown. As described above, when the atgh_type field indicates P_TILE_GR and patchMode indicates P_INTRA, or when the atgh_type field indicates I_TILE_GR and patchMode indicates I_INTRA, patch_data_unit(patchIdx) can be included as patch information data.
[0867] In Figure 56Among them, the pdu_2d_pos_x[p] field indicates the x coordinate (or left offset) of the upper left corner of the patch bounding box of the patch with index p in the current Atlas patch group (or patch) (tileGroupIdx). The Atlas patch group (or patch) can have a patch group (or patch) index (tileGroupIdx). tileGroupIdx can be expressed as a multiple of PatchPackingBlockSize.
[0868] The pdu_2d_pos_y[p] field indicates the y coordinate (or upper offset) of the upper left corner of the patch bounding box of the patch with index p in the current Atlas patch group (or patch) (tileGroupIdx). tileGroupIdx can be expressed as a multiple of PatchPackingBlockSize.
[0869] Adding 1 to the pdu_2d_size_x_minus1[p] field specifies the quantized width value of the patch with index p in the current Atlas patch group (or patch) tileGroupIdx.
[0870] Adding 1 to the pdu_2d_size_y_minus1[p] field specifies the quantized height value of the patch with index p in the current Atlas patch group (or patch) tileGroupIdx.
[0871] The pdu_3d_pos_x[p] field specifies the offset along the tangent axis of the reconstructed patch point in the patch with index p in the current Atlas patch group (or patch).
[0872] The pdu_3d_pos_y[p] field specifies the offset along the bitangent axis of the reconstructed patch point in the patch with index p in the current Atlas patch group (or patch).
[0873] The pdu_3d_pos_min_z[p] field specifies the offset along the normal axis of the reconstructed patch point in the patch with index p in the current Atlas patch group (or patch).
[0874] When the value of the asps_normal_axis_max_delta_value_enabled_flag field included in ASPS is equal to 1, the pdu_3d_pos_delta_max_z[patchIdx] field can be included in the patch data unit.
[0875] If it exists, the pdu_3d_pos_delta_max_z[p] field specifies the nominal maximum of the shift along the normal axis expected to be present in the reconstructed bitdepth patch geometry samples in the patch of index p of the current atlas tile group (or tile) after conversion to their nominal representation.
[0876] The pdu_projection_id[p] field specifies the value of the projection mode and the index of the normal of the projection plane for the patch of index p of the current atlas tile group (or tile).
[0877] The pdu_orientation_index[p] field indicates the patch orientation index for the patch of index p of the current atlas tile group (or tile). The pdu_orientation_index[p] field will be described with reference to Figure 57 below.
[0878] When the value of the afps_lod_mode_enabled_flag field included in the AFPS is equal to 1, the pdu_lod_enabled_flag[patchIndex] can be included in the patch data unit.
[0879] When the pdu_lod_enabled_flag[patchIndex] field is greater than 0, the pdu_lod_scale_x_minus1[patchIndex] field and the pdu_lod_scale_y[patchIndex] field can be included in the patch data unit.
[0880] When the pdu_lod_enabled_flag[patchIndex] field is equal to 1 and patchIndex is p, it specifies that the current patch with index p has LOD parameters. If the pdu_lod_enabled_flag[p] field is equal to 0, the current patch does not have LOD parameters.
[0881] The pdu_lod_scale_x_minus1[p] field plus 1 specifies the LOD scaling factor to be applied to the local x coordinate of the points in the patch of index p of the current atlas tile group (or tile) before adding it to the patch coordinate Patch3dPosX[p].
[0882] The pdu_lod_scale_y[p] field specifies the LOD scaling factor to be applied to the local y coordinate of the points in the patch of index p of the current atlas tile group (or tile) before adding it to the patch coordinate Patch3dPosY[p].
[0883] When the value of the asps_point_local_reconstruction_enabled_flag field included in the ASPS is equal to 1, point_local_reconstruction_data(patchIdx) can be included in the patch data unit.
[0884] According to an embodiment, point_local_reconstruction_data(patchIdx) can contain information that allows the decoder to recover points lost due to compression loss and the like.
[0885] Figure 57 Shows rotation and offset relative to the patch orientation according to an embodiment.
[0886] Figure 57 Shows the Figure 56 rotation matrix and offset assigned to the patch orientation index (pdu_orientation_index[p] field).
[0887] A method / apparatus according to an embodiment can perform an orientation operation on point cloud data. This operation can be performed using Figure 57 the identifiers, rotation, and offset shown.
[0888] According to an embodiment, the NAL unit can include SEI information. For example, non-essential supplementary enhancement information or essential supplementary enhancement information can be included in the NAL unit according to the nal_unit_type.
[0889] Figure 58 Exemplarily shows the syntax structure of SEI information including sei_message() according to an embodiment.
[0890] SEI messages assist in processes related to decoding, reconstruction, display, or other purposes. According to an embodiment, there can be two types of SEI messages: essential and non-essential.
[0891] The decoding process may not require non-essential SEI messages. A conforming decoder may not be required to process this information for output order conformance.
[0892] Essential SEI messages can be an integral part of the V-PCC bitstream and should not be removed from the bitstream. Essential SEI messages can be classified into the following two types:
[0893] Type A necessary SEI messages contain the information required to check bitstream consistency and output timing decoder consistency. Each V-PCC decoder compliant with Point A may not discard any Type A necessary SEI messages and consider them for bitstream consistency and output timing decoder consistency.
[0894] Type B necessary SEI messages: A V-PCC decoder compliant with a specific reconstruction profile may not discard any Type B necessary SEI messages and consider them for 3D point cloud reconstruction and consistency purposes.
[0895] According to an embodiment, an SEI message consists of an SEI message header and an SEI message payload. The SEI message header includes an sm_payload_type_byte field and an sm_payload_size_byte field.
[0896] The sm_payload_type_byte field indicates the payload type of the SEI message. For example, it can be based on the value of the sm_payload_type_byte field to identify whether the SEI message is a prefix SEI message or a suffix SEI message.
[0897] The sm_payload_size_byte field indicates the payload size of the SEI message.
[0898] According to an embodiment, the sm_payload_type_byte field is set to the value of PayloadType in the SEI message payload, and the sm_payload_size_byte field is set to the value of PayloadSize in the SEI message payload.
[0899] Figure 59 An example of the syntax structure of an SEI message payload (sei_payload(payloadType, payloadSize)) according to an embodiment is shown.
[0900] In an embodiment, when nal_unit_type is NAL_PREFIX_NSEI or NAL_PREFIX_ESEI, the SEI message payload may include sei(payloadSize) according to PayloadType.
[0901] In another embodiment, when nal_unit_type is NAL_SUFFIX_NSEI or NAL_SUFFIX_ESEI, the SEI message payload may include sei(payloadSize) according to PayloadType.
[0902] In addition, a V-PCC bitstream (also referred to as a V3C bitstream) having the Figure 25 shown structure can be sent to the receiving side as it is, or can be encapsulated into an ISOBMFF file format by a Figure 1 , Figure 20 or Figure 21 file / fragment encapsulator and sent to the receiving side.
[0903] In the latter case, the V-PCC bitstream can be sent through multiple tracks in the file or through a single track. In this case, the file can be de-encapsulated into a V-PCC bitstream by a Figure 1 , Figure 20 or Figure 22 file / fragment de-encapsulator of the receiving device.
[0904] For example, a V-PCC bitstream carrying a V-PCC parameter set, a geometry bitstream, an occupancy map bitstream, an attribute bitstream, and / or an atlas bitstream can be encapsulated into a file format based on ISOBMFF (ISO Base Media File Format) by a Figure 1 , Figure 20 or Figure 21 file / fragment encapsulator. In this case, according to the embodiment, the V-PCC bitstream can be stored in a single track or multiple tracks in the ISOBMFF-based file.
[0905] According to the embodiment, the ISOBMFF-based file can be referred to as a container, a container file, a media file, a V-PCC file, etc. Specifically, the file can be composed of boxes and / or information that can be referred to as ftyp, meta, moov, or mdat.
[0906] The ftyp box (file type box) can provide information related to the file compatibility or file type of the file. The receiving side can refer to the ftyp box to identify the file.
[0907] The meta box can include a vpcg{0,1,2,3} box (V-PCC group box).
[0908] The mdat box, also referred to as the media data box, includes the actual media data. According to the embodiment, the video-encoded geometry bitstream, the video-encoded attribute bitstream, the video-encoded occupancy map bitstream, and the atlas bitstream are included in the samples of the mdat box of the file. According to the embodiment, the samples can be referred to as V-PCC samples.
[0909] A moov box, also known as a movie box, can contain metadata about the media data of a file (e.g., geometric bitstream, attribute bitstream, occupancy map bitstream, etc.). For example, it can contain information required to decode and play the media data, as well as information about the samples of the file. The moov box can be used as a container for all metadata. The moov box can be the highest-level box among the metadata-related boxes. According to an embodiment, there can be only one moov box in a file.
[0910] A box according to an embodiment can include a track (trak) box that provides information related to the tracks of a file. The track box can include a media (mdia) box that provides media information about the track and a track reference container (tref) box for referencing the track and the samples of the file corresponding to the track.
[0911] The mdia box can include a media information container (minf) box that provides information about the corresponding media data and a handler (hdlr) box (HandlerBox) that indicates the type of the stream.
[0912] The minf box can include a sample table (stbl) box that provides metadata related to the samples of the mdat box.
[0913] The stbl box can include a sample description (stsd) box that provides information about the encoding type employed and the initialization information required for that encoding type.
[0914] According to an embodiment, the stsd box can include sample entries for a track that stores a V-PCC bitstream.
[0915] The term V-PCC used herein has the same meaning as the term video coding based on visual volume (V3C). These two terms can be used complementarily to each other.
[0916] In the present disclosure, in order to store a V-PCC bitstream according to an embodiment in a single track or multiple tracks, volumetric visual tracks, volumetric visual media headers, volumetric visual sample entries, volumetric visual samples, samples and sample entries of a V-PCC track (or referred to as a V3C track) in a file, samples and sample entries of a V-PCC video component track (or referred to as a V3C video component track), etc. can be defined as follows.
[0917] A volumetric visual track (or referred to as a volumetric track) is a track that has a handler type reserved for describing a volumetric visual track. That is, a volumetric visual track can be identified by a volumetric visual media handler type 'volv' included in the HandlerBox of the MediaBox and / or a volumetric visual media header (vvhd) in the minf box of the media box (MediaBox).
[0918] The V3C track refers to the V3C bitstream track, the V3C Atlas track, and the V3C Atlas tile track.
[0919] In the case of a single-track container, the V3C bitstream track is a volumetric visual track containing the V3C bitstream.
[0920] In the case of a multi-track container, the V3C Atlas track is a volumetric visual track containing the V3C Atlas bitstream.
[0921] In the case of a multi-track container, the V3C Atlas tile track is a volumetric visual track containing a part of the V3C Atlas bitstream corresponding to one or more tiles.
[0922] The V3C video component track is a video track carrying 2D video coding data corresponding to one of the occupancy video bitstream, geometric video bitstream, and attribute video bitstream in the V3C bitstream.
[0923] According to an embodiment, video-based point cloud compression (V-PCC) represents volumetric encoding of point cloud visual information.
[0924] That is to say, the minf box in the trak box of the moov box may further include a volumetric visual media header box. The volumetric visual media header box contains information about the volumetric visual track containing the volumetric visual scene.
[0925] Each volumetric visual scene can be represented by a unique volumetric visual track. The ISOBMFF file may contain multiple scenes, so there may be multiple volumetric visual tracks in the ISOBMFF file.
[0926] According to an embodiment, the volumetric visual track can be identified by the volumetric visual media handler type 'volv' in the HandlerBox of the MediaBox and / or the volumetric visual media header (vvhd) in the minf box of the mdia box (MediaBox). The minf box may be referred to as a media information container or a media information box. The minf box is included in the mdia box. The mdia box is included in the trak box. The trak box is included in the moov box of the file. There may be a single volumetric visual track or multiple volumetric visual tracks in the file.
[0927] According to an embodiment, the volumetric visual track may use the VolumetricVisualMediaHeaderBox in the MediaInformationBox. The MediaInformationBox is referred to as the minf box, and the VolumetricVisualMediaHeaderBox is referred to as the vvhd box.
[0928] According to an embodiment, the syntax of the volumetric visual media header (vvhd) box may be defined as follows.
[0929] Box type: 'vvhd'
[0930] Container: MediaInformationBox
[0931] Mandatory: Yes
[0932] Quantity: Exactly one
[0933] The syntax of the volumetric visual media header box (i.e., the vvhd type box) according to an embodiment is as follows.
[0934] aligned(8) class VolumetricVisualMediaHeaderBox
[0935] extends FullBox('vvhd', version = 0, 1) {
[0936] }
[0937] The version may be an integer indicating the version of the box.
[0938] According to an embodiment, the volumetric visual track may use the VolumetricVisualSampleEntry to send signaling information and may use the VolumetricVisualSample to send actual data.
[0939] According to an embodiment, the volumetric visual sample entry may be referred to as the sample entry or the V-PCC sample entry, and the volumetric visual sample may be referred to as the sample or the V-PCC sample.
[0940] According to an embodiment, a single volumetric visual track or multiple volumetric visual tracks may exist in the file. According to an embodiment, a single volumetric visual track may be referred to as a single track or a V-PCC single track, and multiple volumetric visual tracks may be referred to as multiple tracks or multiple V-PCC tracks.
[0941] An example of the syntax structure of the VolumetricVisualSampleEntry is as follows.
[0942] Class VolumetricVisualSampleEntry(codingname) extends SampleEntry(codingname) {
[0943] unsigned int(8)
[32] compressor_name;
[0944] }
[0945] The compressor_name field is the name of the compressor for informational purposes. It is formatted as a fixed 32 - byte field, where the first byte is set to the number of bytes to be displayed, followed by the bytes of the displayable data encoded using UTF - 8, and then padded to complete a total of 32 bytes (including the size byte). This field can be set to 0.
[0946] According to an embodiment, the format of the volumetric visual sample can be defined by the coding system.
[0947] According to an embodiment, a V - PCC unit header box including a V - PCC unit header can exist in the sample entry of a V - PCC track and in the sample entries of all video - encoded V - PCC component tracks included in the program information. The V - PCC unit header box can contain a V - PCC unit header for the data carried by each track as follows.
[0948] aligned(8) class VPCCUnitHeaderBox extends FullBox('vunt', version = 0, 0) {
[0949] vpcc_unit_header() unit_header;
[0950] }
[0951] That is, the VPCCUnitHeaderBox can include vpcc_unit_header().
[0952] Figure 30 An example showing the syntax structure of vpcc_unit_header() is presented.
[0953] According to an embodiment, the sample entry from which VolumetricVisualSampleEntry inherits (i.e., the superclass of VolumetricVisualSampleEntry) includes a VPCC decoder configuration box (VPCCConfigurationBox).
[0954] According to an embodiment, the VPCCConfigurationBox may include a VPCCDecoderConfigurationRecord as shown below.
[0955] class VPCCConfigurationBox extends Box('VpcC'){
[0956] VPCCDecoderConfigurationRecord() VPCCConfig;
[0957] }
[0958] According to an embodiment, the syntax of VPCCDDecoderConfigurationRecord() may be defined as follows.
[0959] aligned(8) class VPCCDDecoderConfigurationRecord{
[0960] unsigned int(8) configurationVersion = 1;
[0961] unsigned int(2) lengthSizeMinusOne;
[0962] bit(1) reserved = 1;
[0963] unsigned int(5) numOfVPCCParameterSets;
[0964] for(i = 0; i < numOfVPCCParameterSets; i++){
[0965] unsigned int(16) VPCCParameterSetLength;
[0966] vpcc_unit(VPCCParameterSetLength) vpccParameterSet;
[0967] }
[0968] unsigned int(8) numOfSetupUnitArrays;
[0969] for(j = 0; j < numOfSetupUnitArrays; j++){
[0970] bit(1) array_completeness;
[0971] bit(1) reserved = 0;
[0972] unsigned int(6) NAL_unit_type;
[0973] unsigned int(8) numNALUnits;
[0974] for(i = 0; i < numNALUnits; i++){
[0975] unsigned int(16) SetupUnitLength;
[0976] nal_unit(SetupUnitLength) setupUnit;
[0977] }
[0978] }
[0979] }
[0980] The configurationVersion is the version field. Incompatible changes to the record are indicated by a change in the version number.
[0981] When the value of the lengthSizeMinusOne field is incremented by 1, it can indicate the length (in bytes) of the NALUnitLenght field included in the VPCCDecoderConfigurationRecord or in the V-PCC samples of the stream to which the VPCCDecoderConfigurationRecord is applied. For example, a size of 1 byte is indicated by "0". The value of this field is the same as the value of the ssnh_unit_size_precision_bytes_minus1 field in the sample_stream_nal_header() of the Atlas substream. Figure 36 A syntax structure example of the sample_stream_nal_header() that shows the ssnh_unit_size_precision_bytes_minus1 field is presented. Incrementing the value of the ssnh_unit_size_precision_bytes_minus1 field by 1 can indicate the precision of the ssnu_vpcc_unit_size element in all sample stream NAL units, in bytes.
[0982] numOfVPCCParameterSets specifies the number of V-PCC parameter sets (VPSs) signaled in the VPCCDecoderConfigurationRecord.
[0983] A VPCC parameter set is an example of sample_stream_vpcc_unit() of a V-PCC unit of type VPCC_VPS. A V-PCC unit may include vpcc_parameter_set(). That is, the VPCCParameterSet array may include vpcc_parameter_set(). Figure 28 Shows an example of the syntax structure of sample_stream_vpcc_unit().
[0984] numOfSetupUnitArrays indicates the number of arrays of Atlas NAL units of the specified type.
[0985] An n-iteration statement with the same number of repetitions as the value of numOfSetupUnitArrays may include array_completeness.
[0986] array_completeness equal to 1 indicates that all Atlas NAL units of the given type are in the following array and none are in the stream. array_completeness equal to 0 indicates that additional Atlas NAL units of the specified type may be in the stream. The default value and allowed values are constrained by the sample entry name.
[0987] NAL_unit_type indicates the type of Atlas NAL units in the following array. NAL_unit_type is constrained to take one of the values indicating NAL_ASPS, NAL_AFPS, NAL_AAPS, NAL_PREFIX_ESEI, NAL_SUFFIX_ESEI, NAL_PREFIX_NSEI, or NAL_SUFFIX_NSEI Atlas NAL units.
[0988] The numNALUnits field indicates the number of Atlas NAL units of the specified type included in the VPCCDDecoderConfigurationRecord applied to the stream of the VPCCDDecoderConfigurationRecord. The SEI array should contain only SEI messages.
[0989] The SetupUnitLength field indicates the size of the setupUnit field in bytes. This field includes the size of both the NAL unit header and the NAL unit payload, but does not include the length field itself.
[0990] A setupUnit is an example of sample_stream_nal_unit() that contains a NAL unit of type NAL_ASPS, NAL_AFPS, NAL_AAPS, NAL_PREFIX_ESEI, NAL_PREFIX_NSEI, NAL_SUFFIX_ESEI, or NAL_SUFFIX_NSEI.
[0991] According to an embodiment, the SetupUnit array may include an Atlas parameter set that is constant for a stream referenced by a sample entry in which a VPCCDDecoderConfigurationRecord exists, and includes essential or non-essential SEI messages for the Atlas sub-stream. According to an embodiment, the Atlas setup unit may be abbreviated as the setup unit.
[0992] According to an embodiment, the file / fragment encapsulator of the present disclosure may perform grouping of samples, grouping of tracks, single-track encapsulation of a V-PCC bitstream, or multi-track encapsulation of a V-PCC bitstream. In addition, the file / fragment encapsulator may add signaling information including SEI messages, signaling information related to tiles and objects, and / or signaling information for supporting spatial access in the form of boxes or FullBoxes to sample entries or separate metadata tracks. Each box will be described in detail below.
[0993] Next, a description will be given of information related to 3D regions, signaling information including SEI messages, and / or information related to tiles and objects in the signaling information signaled in samples, sample entries, sample groups, or track groups included in at least one track of a file or signaled in a separate metadata track.
[0994] SEI information structure
[0995] According to an embodiment, the SEI information structure (referred to as VPCCSEIInfoStruct or V3CSEIInfoStruct) includes essential SEI Atlas NAL units and / or non-essential SEI Atlas NAL units as follows.
[0996] aligned(8) class VPCCSEIInfoStruct() {
[0997] unsigned int(16) numEssentialSEIs;
[0998] for(i = 0; i < numEssentialSEIs; i++){
[0999] unsigned int(16) ESEI_type;
[1000] unsigned int(16) ESEI_length
[1001] nal_unit(ESEI_length) ESEI_byte;
[1002] }
[1003] unsigned int(16) numNonEssentialSEIs;
[1004] for(i = 0; i < numNonEssentialSEIs; i++){
[1005] unsigned int(16) NSEI_type;
[1006] unsigned int(16) NSEI_length
[1007] nal_unit(NSEI_length) NSEI_byte;
[1008] }
[1009] }
[1010] The numEss...
Claims
1. A method for transmitting point cloud data, the method for transmitting point cloud data comprises the following steps: Encoding the point cloud data; Encapsulating the bitstream including the encoded point cloud data into a file; And Transmitting the file, wherein the bitstream is stored in one or more tracks of the file, wherein the file further includes signaling data, wherein the signaling data includes at least one parameter set and information for partial access to the point cloud data, wherein the point cloud data consists of one or more regions, wherein the one or more regions are associated with one or more objects, wherein the one or more objects are associated with one or more atlas patches, wherein the one or more atlas patches constitute an atlas frame, wherein the information for partial access includes region identification information for identifying each of the one or more regions and mapping information between the one or more objects and the one or more atlas patches, and wherein the mapping information includes object index information for identifying each of the one or more objects, quantity information for identifying the number of atlas patches associated with the object, and atlas identification information for identifying each of the atlas patches associated with the object.
2. The method for transmitting point cloud data according to claim 1, wherein, the point cloud data includes geometric data, attribute data, and occupancy map data encoded by a video-based encoding scheme.
3. The method for transmitting point cloud data according to claim 1, wherein, the information for partial access is at least one of static information that does not change over time or dynamic information that changes dynamically over time.
4. A point cloud data transmission device, the point cloud data transmission device comprises: An encoder that encodes the point cloud data; An encapsulator that encapsulates the bitstream including the encoded point cloud data into a file; And A transmitter that transmits the file, wherein the bitstream is stored in one or more tracks of the file, wherein the file further includes signaling data, wherein the signaling data includes at least one parameter set and information for partial access to the point cloud data, wherein the point cloud data consists of one or more regions, wherein the one or more regions are associated with one or more objects, wherein the one or more objects are associated with one or more atlas patches, wherein the one or more atlas patches constitute an atlas frame, wherein the information for partial access includes region identification information for identifying each of the one or more regions and mapping information between the one or more objects and the one or more atlas patches, and Wherein, the mapping information includes object index information for identifying each of the one or more objects, quantity information for identifying the number of atlas patches associated with the object, and atlas identification information for identifying each of the atlas patches associated with the object.
5. The point cloud data sending device according to claim 4, Wherein, The point cloud data includes geometric data, attribute data, and occupancy map data encoded by a video-based coding scheme.
6. The point cloud data sending device according to claim 4, Wherein, The information for partial access is at least one of static information that does not change over time or dynamic information that changes dynamically over time.
7. A method for receiving point cloud data, the method for receiving point cloud data comprises the following steps: Receiving a file; Unencapsulating the file into a bitstream including point cloud data, wherein the bitstream is stored in one or more tracks of the file, and wherein the file further includes signaling data; and Decoding all or part of the point cloud data based on the signaling data, Wherein the signaling data includes at least one parameter set and information for partial access to the point cloud data, Wherein the point cloud data is composed of one or more regions, Wherein the one or more regions are associated with one or more objects, Wherein the one or more objects are associated with one or more atlas patches, Wherein the one or more atlas patches constitute an atlas frame, Wherein the information for partial access includes region identification information for identifying each of the one or more regions and mapping information between the one or more objects and the one or more atlas patches, and Wherein the mapping information includes object index information for identifying each of the one or more objects, quantity information for identifying the number of atlas patches associated with the object, and atlas identification information for identifying each of the atlas patches associated with the object.
8. The method for receiving point cloud data according to claim 7, Wherein, The point cloud data includes geometric data, attribute data, and occupancy map data decoded by a video-based coding scheme.
9. The method for receiving point cloud data according to claim 7, Wherein, The information for partial access is at least one of static information that does not change over time or dynamic information that changes dynamically over time.
10. A point cloud data receiving device, the point cloud data receiving device comprises: A receiver that receives a file; An unencapsulator that unencapsulates the file into a bitstream including point cloud data, wherein the bitstream is stored in one or more tracks of the file, and wherein the file further includes signaling data; and A decoder that decodes all or part of the point cloud data based on the signaling data, Wherein the signaling data includes at least one parameter set and information for partial access to the point cloud data, Among them, the point cloud data consists of one or more regions, wherein the one or more regions are associated with one or more objects, wherein the one or more objects are associated with one or more atlas patches, wherein the one or more atlas patches constitute an atlas frame, wherein the information for partial access includes region identification information for identifying each of the one or more regions and mapping information between the one or more objects and the one or more atlas patches, and wherein the mapping information includes object index information for identifying each of the one or more objects, quantity information for identifying the number of atlas patches associated with the object, and atlas identification information for identifying each of the atlas patches associated with the object.
11. The point cloud data receiving device according to claim 10, wherein, the point cloud data includes geometric data, attribute data, and occupancy map data decoded by a video-based coding scheme.
12. The point cloud data receiving device according to claim 10, wherein, the information for partial access is at least one of static information that does not change over time or dynamic information that changes dynamically over time.