Point cloud data transmission method, point cloud data transmission device, point cloud data reception method, and point cloud data reception device

The method and device for spatially adaptive G-PCC transmission in MPEG-DASH address the challenge of streaming 3D multimedia content by encoding and delivering point cloud data efficiently, adapting to user location and environment changes for optimized bandwidth usage and reduced parsing time.

WO2025143933A1PCT designated stage expired Publication Date: 2025-07-03FOUND FOR RES & BUSINESS SEOUL NAT UNIV OF SCI & TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/021399
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-10-07
Filing Date
2024-12-30
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Current technologies lack efficient methods for transmitting 3D multimedia content produced using point clouds in a streaming manner, particularly in adapting to user location and environment changes.

Method used

A method and device for encoding and transmitting point cloud data using Geometry-based Point Cloud Compression (G-PCC) that includes spatially adaptive encoding and transmission, utilizing MPEG-DASH for efficient delivery of 3D multimedia content based on user location and environment, with hierarchical MPD generation and object grouping for optimized bandwidth usage.

Benefits of technology

Enables efficient and adaptive transmission of 3D multimedia content, ensuring seamless streaming of 3D multimedia content by adapting to user location and environment changes, optimizing bandwidth usage and reducing parsing time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024021399_03072025_PF_FP_ABST
    Figure KR2024021399_03072025_PF_FP_ABST
Patent Text Reader

Abstract

A point cloud data transmission method according to embodiments may comprise the steps of: encoding point cloud data; encapsulating the point cloud data into a file; and transmitting the file. A point cloud data reception method according to embodiments may comprise the steps of: receiving a file including point cloud data; decapsulating the file; and decoding the point cloud data.
Need to check novelty before this filing date? Find Prior Art

Description

Point cloud data transmission method, point cloud data transmission device, point cloud data reception method, and point cloud data reception device

[0001] The embodiments relate to a technology for transmitting three-dimensional multimedia content produced based on G-PCC (Geometry-based Point Cloud Compression) in a streaming manner.

[0002] The concepts of digital twins and the metaverse are being introduced, and related technologies are growing and establishing themselves at the heart of cutting-edge industries. Accordingly, media services are expected to evolve into streaming formats, such as 6DoF 360VR video or 3D multimedia content.

[0003] Meanwhile, one of the most widely used methods for describing the digital world, such as the metaverse, is to create 3D multimedia content based on Geometry-based Point Cloud Compression (G-PCC). G-PCC is used because, in the digital world, the relationship between the user and other objects is not a fixed value but constantly changing, making it efficient to describe objects using G-PCC.

[0004] As previously explained, although it is expected that 3D multimedia content will evolve into a form that provides streaming services, there is currently no technology for transmitting 3D multimedia content produced as point clouds in a streaming manner.

[0005] The embodiments were created against this technical background, and the purpose is to design signaling that includes spatiotemporal information to adaptively transmit 3D multimedia content according to the user's location, environment, etc., thereby enabling transmission of 3D multimedia content in a streaming manner.

[0006] A method for transmitting point cloud data according to embodiments may include a step of encoding point cloud data; a step of encapsulating point cloud data into a file; and a step of transmitting the file. A method for receiving point cloud data according to embodiments may include a step of receiving a file including point cloud data; a step of decapsulating the file; and a step of decoding the point cloud data.

[0007] The method and device according to the embodiments can efficiently encode and transmit point cloud data.

[0008] The method and device according to the embodiments can spatially adaptively encode and transmit point cloud data.

[0009] The drawings are included to further understand the embodiments, and the drawings illustrate the embodiments together with the description related to the embodiments. For a better understanding of the various embodiments described below, reference should be made to the following description of the embodiments in conjunction with the following drawings, in which like reference numerals correspond to corresponding parts throughout the drawings.

[0010] FIG. 1 is a flowchart schematically illustrating a method for streaming G-PCC-based content via MPEG-DASH according to embodiments.

[0011] Figure 2 shows objects within a user's field of view according to embodiments.

[0012] Figure 3 shows a group of objects according to embodiments.

[0013] Figure 4 illustrates a group of objects within a user's field of view according to embodiments.

[0014] Figure 5 shows a structure for transmitting geometry-based point cloud compression data according to embodiments.

[0015] FIG. 6 illustrates multi-track encapsulation of G-PCC data according to ISOBMFF (ISO Base Media File Format) according to embodiments.

[0016] Figure 7 illustrates sample entries and samples of a track of a G-PCC file according to embodiments.

[0017] FIG. 8 illustrates grouping of GPCC components within MPD (MPEG-DASH Media Presentation Description) according to embodiments.

[0018] Figure 9 shows an IDD (Init Description Document) according to embodiments.

[0019] Figure 10 illustrates a method for describing IDD-based spatial information.

[0020] Fig. 11 illustrates a point cloud data transmission method according to embodiments.

[0021] Figure 12 illustrates a method for receiving point cloud data according to embodiments.

[0022] Preferred embodiments of the embodiments are described in detail, examples of which are illustrated in the accompanying drawings. The following detailed description, with reference to the accompanying drawings, is intended to illustrate preferred embodiments of the embodiments, rather than merely show embodiments that can be implemented according to the embodiments. The following detailed description includes details to provide a thorough understanding of the embodiments. However, it will be apparent to those skilled in the art that the embodiments may be practiced without these details.

[0023] While most of the terms used in the examples are commonly used in the field, some terms were arbitrarily selected by the applicant, and their meanings are described in detail in the following descriptions as needed. Therefore, the examples should be understood based on the intended meaning of the terms, not simply their names or meanings.

[0024] FIG. 1 is a flowchart schematically illustrating a method for streaming G-PCC-based content via MPEG-DASH according to embodiments.

[0025] At step S10, when a user (10) requests data transmission from a server (20) to view specific G-PCC-based content, the server (20) transmits an IDD (Init Description Document) containing information about media presentation descriptions (MPDs) that divide and describe the entire space to the user (10).

[0026] This is to reduce the time required for the user (10) to parse the file, as when providing spatial information based on MPEG-DASH, a single MPD contains too many objects, resulting in a large file size. This step may be selectively applied depending on the system environment.

[0027] At step S20, the user (10) selects an MPD that describes the spatial information he or she needs based on his or her location and IDD, and then requests the related MPD while transmitting his or her location information to the server (20).

[0028] At step S30, the server (20) generates an MPD file describing the relationship between the user and objects placed in space based on the user's location. This MPD file can be generated in advance based on 3D spatial information, or can be generated in real time according to the user's request at step S20.

[0029] At step S40, the server (20) transmits the MPD file generated according to the user's request to the user (10).

[0030] At step S50, the user (10) requests the server (20) for the necessary amount of G-PCC-based content based on the received MPD, taking into account the user's network environment and the relative distance and angle from the object.

[0031] At step S60, the server (20) that received the request transmits the requested segment data.

[0032] In the above operation process, since MPEG-DASH stipulates that an MPD describing the relationship between a user and content be created and distributed, the server (20) basically creates and distributes an MPD.

[0033] To avoid filling the MPD with too much information, the entire space can be divided into multiple sections and an MPD can be created that describes only the sections based on the user's location information.

[0034] In this case, since the generated MPD cannot represent the entire space, the server (20) additionally generates an IDD containing information about which space within the entire space each MPD describes. For this purpose, the generated IDD information includes the size of the identified space and a URI that can receive the corresponding MPD. Table 1 shows an example of an IDD.

[0035] Space URI(0,0,0 / 180,259,372)https: / / www.IDD.kr / first_MPD(180,0,0 / 375,259,372)https: / / www.IDD.kr / second_MP D(0,259,0 / 180,511,372)https: / / www.IDD.kr / third_MPD(180,259,0 / 375,511,372)https: / / www.IDD.kr / fourth_MPD

[0036] In accordance with this, step S20 is explained again. After the user (10) selects an MPD that describes the partitioned space he / she needs from the entire space based on his / her location and IDD, he / she requests the related MPD while transmitting his / her location information to the server (20). In step S30, the server (20) generates an MPD corresponding to the space. The reason for generating the MPD after the user request is that various relationships, such as the Occlusion Area and the proportion of the Object on the screen, change depending on the user's location within the space.

[0037] Meanwhile, MPD is updated based on factors such as the length of the content and the user's movement speed. If the user moves quickly, more frequent updates are required because the relationship between objects and the user can change frequently.

[0038] Therefore, in this embodiment, the TTL (time to live) of the MPD is set as in mathematical formula 1 by considering the segment length, the movement speed of the user in the space, etc. Here,

[0039] T mpd is the holding time, T s is the segment length, V user is the user's movement speed.

[0040]

[0041] At this time, the segment length is the unit of time in which the server (20) divides the content. The content is maintained for the segment length, and the space occupied during the segment length time is defined as the object size.

[0042] Content is created based on G-PCC. This geometry-based point cloud compression method relies on a three-dimensional tree structure called an octree. Users should be able to receive objects individually based on their relative position and distance from each other, or receive multiple objects simultaneously.

[0043] The resolution quality of compressed objects using node-based compression algorithms is determined by the relative distance between the user and the object. Therefore, a standard is needed when dividing or grouping objects, or when specifying content quality.

[0044] Therefore, below, we present the weight values ​​used to specify these two, and explain how to use these values ​​to group and determine quality.

[0045] ND: Node Depth

[0046] ND stands for Node Depth Weighting, and it represents the required resolution required for an object to be consumed at a certain quality. To determine ND, the object first needs the Required Node Depth (RND), which is the minimum Node Depth value required at a reference distance (e.g., 1 m). This RND is stored in the MPD as additional information for each object.

[0047] According to RND, ND is defined as in Equation 2, where r is the distance between the object and the user.

[0048]

[0049] TL: Tile Level

[0050] TL stands for Tile Level Weight. TL is a measure of whether an object should be segmented and transmitted based on the user's viewport (FOV) when transmitting an object to the user, or transmitted as a group with other objects. Here, a tile is the same concept as a unit for dividing and storing space, and a single object can be segmented and expressed as multiple tiles depending on the segmentation level. TL is a level value, with values ​​ranging from [1, 2, 3, 4, 5], as shown in Table 2.

[0051] TL Meaning 1. Transmit a single object by dividing it into smaller pieces of two or more levels 2. Transmit a single object after dividing it into one level 3. Transmit a single object 4. Transmit a group containing one or more objects 5. Transmit a space containing one or more groups

[0052] The TL value of each object is determined by a set of reference values. The reference value set is an element for distinguishing TL and expresses the degree of user field of view range using latitude and longitude. Each TL reference value set is configured as a reference value for passing the TL stage, and distinguishes between TL n and (n+1). Each value Defined as . Defined as . Longitude standard ( ) and latitude standard ( )silver class 2 shows an object within the user's field of view according to embodiments.

[0053] As illustrated in Fig. 2, the proportions θ and φ occupied by objects within the user's field of view are compared with a set of reference values ​​to determine the TL value. The proportion occupied by an object within the field of view is calculated as the difference between the largest θ and φ and the smallest θ and φ among the eight polar coordinate representations (r, θ, φ) of the eight points constituting the object.

[0054] The TL value is determined by comparing the θ, φ obtained through this with the TL reference value set θ, φ. For example, if the θ of the object is greater than θ1, the TL is 1. If θ is less than θ4, the TL is 5. Through this method, the values ​​of θ, φ of the object are confirmed, and the TL values ​​of each of the two axes are determined. It is designated as . When proceeding or dividing the group thereafter, each direction is divided / grouped separately.

[0055] Meanwhile, in one embodiment, objects having the same ND and TL based on the above-described ND and TL can be grouped into the same group and reconstructed into a single file.

[0056] For example, objects with a TL value greater than a threshold (e.g., 4 or greater) can be grouped together. Objects in this group are objects with a TL of 4 or 5 within the user's field of view (FoV).

[0057] Figure 3 shows a group of objects according to embodiments.

[0058] Objects grouped together are reconstructed into a single file by creating an additional node called a group root node above the root node of the existing node structure, as shown in Figure 3, and decoded with the same node depth. During this process, if the ND of individual objects differs, the quality of objects on the user's screen may show a very large gap.

[0059] Therefore, when grouping objects by TL, ND must also be considered. Consequently, objects within a group share the same ND and TL. This means that all objects within a group appear to the user with the same quality at the same node depth.

[0060] Furthermore, groups should be presented in various combinations to ensure adaptive transmission based on the user's viewing direction. This is because fixed group combinations waste user bandwidth. As shown in Figure 4, when grouping objects, the number of objects grouped is adjusted based on the user's viewing direction, grouping objects with the same ND TL into the same group.

[0061] This will be explained in more detail with reference to Fig. 4.

[0062] Figure 4 illustrates a group of objects within a user's field of view according to embodiments.

[0063] As described above, the server generates object groups of various combinations to provide an object group suitable for the user's head direction. The object groups are combined within the user's field of view, and various object groups can be generated depending on the user's head direction. Fig. 4 shows object groups that are variously combined within the same TL and ND depending on the user's head direction. The user's head direction can be expressed as latitude and longitude values ​​θ and φ among the polar coordinate representations of the head direction. In addition, the user's field of view (FoV) is expressed as (q, j) by the service providing terminal. In order to generate an object group, it is necessary to determine whether the two axes of latitude and longitude are included within the user's field of view. However, for the convenience of understanding, an example based on the longitude axis is described, as in Fig. 4.

[0064] When the user's head direction is expressed as θ and the field of view is defined as q, the user's field of view is expressed as (θ-q / 2, θ+q / 2). At this time, when θ is 0, an object group is created by combining objects included in the field of view considering TL and ND, and a new object group is created whenever the composition of objects within the field of view changes as θ increases.

[0065] Through the above explanation, the following three elements are required to describe objects reconstructed to fit the user's location in MPD.

[0066] 1. Adaptation Set

[0067] 2. Representation (resolution)

[0068] 3. SRD(Spartial Relationship Description)

[0069] 1. Adaptation Set

[0070] An Adaptation Set is typically a content element that can be individually decoded. This Adaptation Set is defined as an individual Adaptation Set of objects segmented / combined through the previously described TL and ND.

[0071] 2. Representation (playback quality)

[0072] Representation refers to the quality of content and is adjusted based on the user's network environment. G-PCC-based content has a variety of node depths, and providing all node depths as representations can lead to object-specific gaps. Therefore, the number of levels (m) and depth gap (n) provided systematically are defined. Here, the number of levels refers to the number of playback quality levels provided by the system, and the depth gap refers to the gap between quality levels.

[0073] m,n set the Node Depth of each object specified in Representation in conjunction with ND. ND is the Depth Level of the lowest Representation, and a maximum of m levels are created with a difference of at least n levels up to the total Node Depth of the object.

[0074] For example, if the ND of an object with a depth of 16 at the leaf node is 6 and m = 5, the number of representations of the object is 5, which is at most m. The lowest of these is

[0075] Representation is Node Depth 6. The highest Representation is Node Depth 16. Node depths 9 and 12 are each Representation to allow for an interval of n or more between them. As a result, the system can provide 5 resolutions, but only 4 Representations are provided due to the relationship between the object ND and the maximum depth.

[0076] 3. SRD(Spartial Relationship Description)

[0077] SRD (spartial relation desecription) is an item that describes the relationship between users and objects and the objects. It is written as a sub-item called Essential Property within the Adaptation Set and is an item that writes additional descriptions about the Adaptation Set. The SRD of the system is Write in the form.

[0078] is a polar coordinate system that expresses the positional relationship from the user to the center of an individual object or a group of objects. is the Width, Depth, and Height of the object. is a quaternion value that describes the rotation of the object. Since the three-axis-based rotation values ​​represented by yaw, pitch, and roll have the problem that the final direction changes depending on the rotation order, the rotation of the object is expressed as a quaternion.

[0079] Based on the SRD value, the user can check information such as the object's location, size, and rotation to place the object in space.

[0080] FIG. 5 illustrates a structure for transmitting geometry-based point cloud compression data according to embodiments.

[0081] A real-world visual scene (A) is captured by a camera set or a camera device with multiple lenses and sensors. The resulting acquisition generates source point cloud data (B). One or more point cloud frames are encoded into a coded G-PCC bitstream, which includes a coded geometry bitstream and an attribute bitstream (E). The one or more coded bitstreams consist of a media file (F) for file playback or a sequence of initialization segments and media segments (Fs) for streaming, according to a specific media container file format. The media container file format according to embodiments is the ISO base media file format specified in ISO / IEC 14496-12 [ISOBMFF]. The file encapsulator also includes metadata in the file or segment. The segment F is delivered to the player using a delivery mechanism.

[0082] The file (F) output by the file encapsulator is identical to the file (F') input to the file decapsulator. The file decapsulator processes the file (F') or the received segment (F'), extracts the coded bitstream (E'), and parses the metadata. The G-PCC bitstream is then decoded into a decoded signal (D'), and point cloud data is generated from the decoded signal (D'). If applicable, the point cloud data is rendered and displayed on the screen of a head-mounted display or other display device based on the current viewing position, viewing direction, or viewport determined by various types of sensors, such as the head. Tracking and position or eye-tracking sensors are also possible. In addition to being used to help the player access the appropriate portion of the decoded point cloud data, the current viewing position or viewing direction may also be used to optimize decoding. In viewport-dependent delivery, the current viewing position and viewing direction are also passed to the strategy module to determine which tracks to receive.

[0083] The process described above is applicable to both live and on-demand use cases.

[0084] The definition of each interface in Figure 5 is as follows:

[0085] E / E',: coded G-PCC bitstream

[0086] F / F': A media file containing a specification of a track format, which may include constraints on the elementary streams contained within the track samples.

[0087] Although not shown in Figure 5, geometry-based point cloud compression data can be transmitted in DASH format.

[0088] FIG. 6 illustrates multi-track encapsulation of G-PCC data according to ISOBMFF (ISO Base Media File Format) according to embodiments.

[0089] When a G-PCC bitstream is carried on multiple tracks per G-PCC component, each G-PCC component bitstream is mapped to a separate track. There are two types of G-PCC component tracks: G-PCC geometry tracks and G-PCC attribute tracks. Each sample in a track contains at least one G-PCC unit carrying a single G-PCC component data unit, and does not contain both geometry and attribute data units, or multiplexing of different attribute data units.

[0090] The conditions for multi-track according to the embodiments are as follows.

[0091] a) There is one G-PCC geometry track and it becomes the entry point.

[0092] b) Zero or more G-PCC attribute tracks may appear. The track header flag of a G-PCC attribute track is set to 0.

[0093] c) A sample entry has one GPCCComponentInfoBox to indicate the role of the stream contained in the track.

[0094] d) Track references are introduced from the G-PCC geometry track to the G-PCC attribute track.

[0095] Tracks belonging to the same G-PCC sequence are time-aligned. Samples contributing to the same point cloud frame across different tracks have the same presentation time. If a parameter set or tile list is present in a sample, the decoding time of the parameter set or tile list used for that sample is equal to or earlier than the decoding time of the corresponding G-PCC component data unit sample. If all parameter sets are present in samples across multiple tracks, the decoding time of a sample containing an SPS is equal to or earlier than the decoding time of a sample containing a GPS, APS, or tile list. Additionally, all tracks belonging to the same G-PCC sequence have the same implicit or explicit edit list.

[0096] Figure 7 illustrates sample entries and samples of a track of a G-PCC file according to embodiments.

[0097] Figure 7(a) shows a sample entry of a track of a G-PCC file. Figure 7(b) shows a sample of a track of a G-PCC file.

[0098] Referring to Figure 7(a), the compressorname in the base class VolumetricVisualSampleEntry indicates the name of the compressor to be used with the recommended " / 013GPCC Coding" value. The first byte is the remaining byte count, represented here as / 013 (octal 13). This is 11 (decimal), the number of bytes in the remaining string.

[0099] config contains the 7 G-PCC decoder configuration record information.

[0100] type indicates the type of G-PCC component passed to each track.

[0101] Referring to Fig. 7(b), each sample corresponds to a single point cloud frame, and samples contributing to the same point cloud frame in multiple tracks have the same presentation time.

[0102] Each sample consists of one or more G-PCC units of the G-PCC component indicated in the GPCCComponentInfoBox of the sample entry, and zero or more G-PCC units carrying either a parameter set or a tile list. If a G-PCC unit containing a parameter set or a tile list exists in the sample, it appears before the G-PCC unit of the G-PCC component. The sample structure of the G-PCC geometry track is as shown in Fig. 7(b).

[0103] FIG. 8 illustrates grouping of GPCC components within MPD (MPEG-DASH Media Presentation Description) according to embodiments.

[0104] Figure 8 is an exemplary DASH configuration for grouping G-PCC components belonging to a single point cloud media within an MPEG-DASH MPD file.

[0105] Period 1 includes PreSelection 1, and PreSelection 1 may include a GPCC descriptor. The GPCC descriptor may include parameter information regarding compression of G-PCC data according to embodiments. PreSelection 1 may be referenced by an AdaptationSet. Depending on the configuration of point cloud data, there may be multiple AdaptationSets, such as geometry, attribute (color), attribute (reflectivity), attribute (material), etc., and each AdaptationSet may include a GPCC component descriptor together with each Representation. The GPCC component descriptor may include parameter information for each component.

[0106] Figure 9 shows an IDD (Init Description Document) according to embodiments.

[0107] A point cloud data transmission device according to embodiments may include a point cloud data acquisition unit, a point cloud data encoder, and / or a file / segment encapsulator, as shown in FIG. 5. A point cloud data reception device according to embodiments may include a file / segment decapsulator, a point cloud data decoder, and / or a point cloud data renderer.

[0108] A point cloud data transmitting device according to embodiments can adaptively transmit G-PCC data, and a point cloud data receiving device according to embodiments can adaptively receive G-PCC data.

[0109] The embodiments enable the efficient use of descriptor documents (Descriptor Documents) when transmitting data for large spaces based on a hierarchical MPD. Object information can be adaptively transmitted through MPDs. Transmission efficiency can be increased by varying the object composition by group or tile. A user information feedback channel can be defined for the server to push optimal content based on user information throughout the entire operation.

[0110] Referring to FIG. 9, a transmitting device according to embodiments may transmit additional IDD information within a file. When selecting an MPD on the receiving side, the necessary information may be selected based on the user's location, motion vector, field of view, etc. The decoder on the receiving side may receive the MPD based on the user's selection. The decoder may request segments from the server. The decoder may receive segment data.

[0111] Referring to Figures 3 and 4, multiple objects can be grouped together, and information about them can be added to the MPD and / or file. The object grouping can vary depending on the user's position or field of view. Using the user's position or field of view information, the decoder can obtain the necessary information from the MPD and / or file to decode and render the desired G-PCC data.

[0112] The transmission method according to the embodiments can perform a scalable encoding method for adaptive transmission.

[0113] The transmission method according to the embodiments can encode an object by dividing it into stages based on the user's location. For example, the units of division can include groups and / or tiles.

[0114] The transmission method according to the embodiments can scalably encode objects divided into steps.

[0115] The transmission method according to the embodiments can generate and transmit hierarchical MPD and / or file tracks for wide space transmission implementation.

[0116] The transmission method according to the embodiments can divide the space into a certain size and express it as an individual MPD, and generate a descriptor document that explains which position multiple MPDs are responsible for.

[0117] The transmission method according to the embodiments can generate Object information within the MPD for adaptive selection based on user location. For example, the information can include location information, size information, and / or rotation information of the Object. In addition, the information can include dynamic-related information.

[0118] The transmission method according to the embodiments may form a feedback channel to support push operations for users. For example, a feedback channel that allows a server to push content based on user location, direction, speed, etc. may be used to transmit track information of the relevant MPD and / or file to a decoder so that the receiving decoder can decode G-PCC content optimized for the user.

[0119] An IDD according to embodiments refers to an initial description document and may be abbreviated as initial information, etc. An IDD may include the following information:

[0120] Information that distinguishes the extent of the entire space, background and / or foreground, etc.: For example, information that distinguishes the background and foreground may be included to separately convey background elements that should be consistently visualized relative to the user's position in an open space.

[0121] Information that distinguishes whether a space is open or closed when defining a space: For example, this information can be included to determine whether elements not included in the corresponding MPD are visible when the user looks in a particular direction. Additionally, if the space is open, additional objects may need to be downloaded via the background or additional MPD for that direction.

[0122] The URL of the MPD containing the segmented space definition and the corresponding spatial information: For example, the MPD can describe the space as accurately as possible by distinguishing between Points and Faces, allowing for the inclusion of curved spatial elements rather than simple shapes like rectangular solids. Furthermore, the Face Index can be used to indicate whether a face is open or not. The Face Index can have the same meaning as the Face ID.

[0123] An IDD according to embodiments may be configured as follows:

[0124] IDD's Global Structure contains size information for the entire space. Specifically, IDD's bounding box contains extent information for the entire space. For example, it may include bounding box location and bounding box size.

[0125] IDD's Spatial Partition contains information about subspaces. It can contain information about subspaces within the entire space, as follows. For example, Space ID indicates the ID value of the space. Name indicates the name of the space as a string. Bounding Box provides information about the vertices needed to implement the space. Face provides information about the surface (face) created by connecting the vertices. Properties provide additional information about the space. State indicates whether the space is open. Face indicates the Face Index, which is the same for both open and closed states. Type indicates whether the space is the background or foreground. Reference indicates the URL address of the MPD. MPD contains the URL address of the MPD.

[0126] A bounding box element can contain vertex IDs that constitute the bounding box and position information of points corresponding to the vertex IDs.

[0127] The Face element contains information about the face ID and the vertices that make up the face identified by the face ID.

[0128] The Properties element indicates whether the space is open or closed. The State element can be used to indicate whether the space is open or closed.

[0129] Figure 10 illustrates a method for describing IDD-based spatial information.

[0130] As shown in Figure 10, the entire space exists, and subspaces of the entire space can exist. The entire space can be expressed based on a background and a certain range. The space can include open and / or closed spaces. By representing this space with various pieces of information, the decoder can effectively achieve spatial access.

[0131] The composition of the IDD described above can be expressed as follows:

[0132] <idd>

[0133] <globalstructure>

[0134] <boundingbox>

[0135] <minx> 0< / minx> <miny> 0< / miny> <minz> 0< / minz>

[0136] <maxx> 1000< / maxx> <maxy> 1000< / maxy> <maxz> 1000< / maxz>

[0137] < / boundingbox>

[0138] < / globalstructure>

[0139] <spatialpartition>

[0140] <space id="S1">

[0141] <name>MainRoom < / name>

[0142] <boundingbox>

[0143] <vertices>

[0144] <vertex id="A" x="0" y="0" z="0" / >

[0145] <vertex id="B" x="2" y="0" z="0" / >

[0146] <vertex id="C" x="2" y="1" z="0" / >

[0147] <vertex id="D" x="3" y="1" z="0" / >

[0148] <vertex id="E" x="3" y="3" z="0" / >

[0149] …< / vertices>

[0150] < / boundingbox>

[0151] <faces>

[0152] <Face id=" <vertexref> ABCDEFGHIJ < / vertexref>

[0153] <Face id="” <vertexref> ABBA' < / vertexref> <face>…

[0154] <Properties = "”

[0155] <state> 1,2< / state>

[0156]

[0157] <Properties state = "”

[0158] <state> 3,4< / state>

[0159]

[0160] <type> Background < / type>

[0161] <reference>

[0162] <mpd> http: / example.com / mainroom.mpd< / mpd>

[0163] < / reference>

[0164] < / face> < / faces> < / space>

[0165] < / spatialpartition>

[0166] < / idd>

[0167] That is, as described above, IDD information can include global structural information. Global structural information can include bounding box information. Bounding box information can include minimum and maximum values ​​for each axis of the bounding box.

[0168] IDD information may include spatial partition information. The spatial partition information may include identification information of a space corresponding to the spatial partition. The spatial partition information may include name information of the space identified by the space identification information. The spatial partition information may include identification information of vertices of a bounding box, which is a space identified by the space identification information, and position information of the vertices.

[0169] Spatial partition information may include face information. The face information may include information identifying the face. The face information may include reference information of a vertex included in the face identified by the face identification information. Spatial partition information may include state information of the face. The state information may indicate at least one state of closed or open through attribute information. If the attribute (property) is open, ID information of faces having the same open state may be indicated through state information. For example, if the IDs of faces whose property state is closed are 3 and 4, the state information may include the IDs of faces having the same closed state. Spatial partition information may include type information. The type information may indicate attributes of the space, such as whether the space of the spatial partition is a background or a foreground.

[0170] Spatial partition information may include MPD reference information. The MPD reference information may include address information from which the MPD can be obtained.

[0171] Referring to FIG. 8, a point cloud data transmission method according to embodiments can transmit G-PCC configuration information to an MPD-based descriptor based on the DASH method.

[0172] You can indicate that you want to pass information for GPCC via schemeIDUri in EssentialProperty in MPD: "urn:mpeg:mpegI:gpcc:2020:component".

[0173] To convey a G-PCC, each item in the MPD can contain a GPCC Component Descriptor. The descriptor contains information about the G-PCC. For example, the descriptor contains a group identifier, such as @component@group_id. The descriptor indicates a G-PCC data type, such as @component@component_type. 'geom' indicates the coordinates of the geometry, and 'attr' indicates the color of the attribute. Regarding @component@geomtry_type, it contains the location information of the anchor point of the area through Anchor Point (X,Y,Z), contains the coordinate information of the geometry, such as Position (X,Y,Z), and contains the location information of the geometry that changes over time, such as Dynamic Position (X,Y,Z). In addition, it can convey the final location information of a point in a dynamic area, such as Final Position.

[0174] The method for transmitting point cloud data according to embodiments may generate and transmit a GPCC Component Descriptor within an MPD. The GPCC Component Descriptor may include a group ID that identifies a group of geometry data in the point cloud data. Additionally, the GPCC Component Descriptor may include a group ID that identifies a group of attribute data.

[0175] A method for receiving point cloud data according to embodiments may receive and parse a GPCC Component Descriptor within an MPD. The GPCC Component Descriptor may include a group ID that identifies a group of geometry data of the point cloud data. Additionally, the GPCC Component Descriptor may include a group ID that identifies a group of attribute data. Through the group ID within the MPD, point cloud data associated with a specific object and / or a specific space may be efficiently partially decoded.

[0176] The method / device according to the embodiments can adaptively transmit and receive point cloud data through recording position information in the IDD and / or MPD.

[0177] The MPD according to the embodiments may further include additional elements that define and verify the relationship between the IDD and the MPD. For example, the MPD may further include IDD index information, etc.

[0178] The method / device according to the embodiments can check the IDD in the MPD using the following method. For example, information indicating the ID value can be added to the MPD item itself. <IDD ID = “” / <MPD IDD = “”와 같은 형태로 IDD의 ID를 나타낼 수 있다. MPD 하위 항목에 IDD의 ID를 추가하여 동일성 여부 나타낼 수 있다.

[0179] Fig. 11 illustrates a point cloud data transmission method according to embodiments.

[0180] A method for transmitting point cloud data according to embodiments may include a step (S1100) of encoding point cloud data.

[0181] The point cloud data transmission method according to the embodiments may further include a step (S1110) of encapsulating the point cloud data into a file.

[0182] The point cloud data transmission method according to the embodiments may further include a step of transmitting a file (S1120).

[0183] Referring to FIG. 6, with respect to multi-tracks, the file may include a first track including geometry of the point cloud data, and a second track including attributes of the point cloud data.

[0184] Referring to FIG. 7, with respect to the sample entries and samples of the tracks, each track of the file includes a sample entry including configuration information regarding the point cloud data and a sample including the point cloud data, and the sample may further include a parameter set regarding the point cloud data.

[0185] Referring to FIG. 8, with respect to MPD, the method further includes a step of transmitting MPD information for point cloud data, wherein the MPD includes first adaptation set information for a geometry of the point cloud data and second adaptation set information for an attribute of the point cloud data, wherein the first adaptation set information includes a component descriptor for the geometry, and the second adaptation set information includes a component descriptor for the attribute.

[0186] Referring to FIG. 9, with respect to IDD, the method further includes: a step of transmitting initial information indicating information about an entire space including point cloud data, wherein the initial information includes at least one of size information about the entire space or information about a subspace of the entire space, and the information about the subspace may include at least one of an identifier for the subspace, a name for the subspace, or position information of a bounding box for the subspace.

[0187] Referring to FIG. 9, with respect to the face, properties, state, face, type, MPD URL, etc. of the IDD, the initial information may further include at least one of a face ID that identifies a face generated based on points of the point cloud data, vertex information included in the face identified by the face ID, information indicating whether the subspace is an open space or a closed space, index information for the face, information indicating whether the subspace is a background or a foreground, or address information of an MPD regarding the point cloud data.

[0188] A method for transmitting point cloud data can be performed by a transmitting device of FIG. 5. The point cloud data transmitting device can include an encoder for encoding point cloud data; an encapsulator for encapsulating point cloud data into a file; and a transmitter for transmitting the file.

[0189] Figure 12 illustrates a method for receiving point cloud data according to embodiments.

[0190] The receiving method of Fig. 12 can follow the reverse process of the transmitting method of Fig. 11.

[0191] A method for receiving point cloud data according to embodiments may include a step (S1200) of receiving a file including point cloud data.

[0192] The method for receiving point cloud data according to embodiments may further include a step of decapsulating a file (S1210).

[0193] The method for receiving point cloud data according to embodiments may further include a step (S1220) of decoding point cloud data.

[0194] The received file may include a first track containing the geometry of the point cloud data, and a second track containing attributes of the point cloud data.

[0195] Each track of the received file includes a sample entry containing configuration information about the point cloud data and a sample containing the point cloud data, wherein the sample may further include a set of parameters about the point cloud data.

[0196] Referring to the MPD of FIG. 8, the method further includes: receiving initial information indicating information about an entire space including point cloud data, wherein the initial information includes at least one of size information about the entire space or information about a subspace of the entire space, and the information about the subspace may include at least one of an identifier for the subspace, a name for the subspace, or position information of a bounding box for the subspace.

[0197] With respect to the IDD of FIG. 9, the method further includes: receiving initial information indicating information about the entire space including the point cloud data, wherein the initial information includes size information about the entire space, the initial information includes information about a subspace of the entire space, and the information about the subspace may include an identifier for the subspace, a name for the subspace, and position information of a bounding box for the subspace.

[0198] Additionally, the initial information may further include at least one of a face ID that identifies a face generated based on points of the point cloud data, vertex information included in the face identified by the face ID, information indicating whether the subspace is an open space or a closed space, index information for the face, information indicating whether the subspace is a background or a foreground, or address information of an MPD regarding the point cloud data.

[0199] A method for receiving point cloud data can be performed by a receiving device of FIG. 5. The point cloud data receiving device can include a receiving unit that receives a file containing point cloud data; a decapsulator that decapsulates the file; and a decoder that decodes point cloud data.

[0200] The embodiments have been described in terms of methods and / or devices, and the descriptions of methods and devices may be applied complementarily.

[0201] For the convenience of explanation, each drawing has been described separately, but it is also possible to design a new embodiment by combining the embodiments described in each drawing. In addition, designing a computer-readable recording medium having a program recorded thereon for executing the previously described embodiments, as needed by a person skilled in the art, also falls within the scope of the embodiments. The devices and methods according to the embodiments are not limited to the configurations and methods of the embodiments described above, but the embodiments may be configured by selectively combining all or part of the embodiments so that various modifications can be made. Although preferred embodiments of the embodiments have been illustrated and described, the embodiments are not limited to the specific embodiments described above, and various modifications can be made by a person skilled in the art to which the present invention pertains without departing from the gist of the embodiments claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the embodiments.

[0202] The various components of the devices of the embodiments may be implemented by hardware, software, firmware, or a combination thereof. The various components of the embodiments may be implemented by a single chip, for example, a single hardware circuit. According to embodiments, the components according to the embodiments may be implemented by separate chips. According to embodiments, at least one of the components of the devices of the embodiments may be configured with one or more processors capable of executing one or more programs, and the one or more programs may perform, or include instructions for performing, one or more of the operations / methods according to the embodiments. The executable instructions for performing the methods / operations of the devices of the embodiments may be stored in non-transitory CRMs or other computer program products configured to be executed by one or more processors, or may be stored in temporary CRMs or other computer program products configured to be executed by one or more processors. In addition, the memory according to the embodiments may be used as a concept including not only volatile memory (e.g., RAM, etc.), but also non-volatile memory, flash memory, PROM, etc. Additionally, it may include implementations in the form of carrier waves, such as transmissions via the Internet. Furthermore, processor-readable recording media may be distributed across network-connected computer systems, allowing processor-readable code to be stored and executed in a distributed manner.

[0203] In this document, “ / ” and “,” are interpreted as “and / or”. For example, “A / B” is interpreted as “A and / or B”, and “A, B” is interpreted as “A and / or B”. Additionally, “A / B / C” means “at least one of A, B, and / or C”. Also, “A, B, C” means “at least one of A, B, and / or C”. Additionally, “or” in this document is interpreted as “and / or”. For example, “A or B” can mean 1) “A” only, 2) “B” only, or 3) “A and B”. In other words, “or” in this document can mean “additionally or alternatively”.

[0204] Terms such as "first" and "second" may be used to describe various components of the embodiments. However, the various components according to the embodiments should not be interpreted as limited by these terms. These terms are merely used to distinguish one component from another. For example, a first user input signal may be referred to as a "second user input signal." Similarly, a second user input signal may be referred to as a "first user input signal." The use of these terms should be interpreted as not departing from the scope of the various embodiments. Although "first user input signal" and "second user input signal" are both user input signals, they do not mean the same user input signals unless the context clearly indicates otherwise.

[0205] The terminology used to describe the embodiments is for the purpose of describing particular embodiments and is not intended to be limiting of the embodiments. As used in the description of the embodiments and in the claims, the singular is intended to include the plural unless the context clearly dictates otherwise. The expressions “and / or” are used to mean all possible combinations of terms. The expression “includes” describes the presence of features, numbers, steps, elements, and / or components, but does not mean that additional features, numbers, steps, elements, and / or components are not included. Conditional expressions such as “if” or “when” used to describe the embodiments are not intended to be limited to only optional cases. When a specific condition is satisfied, a related action is performed in response to a specific condition, or a related definition is intended to be interpreted.

[0206] Additionally, the operations according to the embodiments described in this document may be performed by a transceiver device including a memory and / or a processor according to the embodiments. The memory may store programs for processing / controlling the operations according to the embodiments, and the processor may control various operations described in this document. The processor may be referred to as a controller, etc. The operations according to the embodiments may be performed by firmware, software, and / or a combination thereof, and the firmware, software, and / or a combination thereof may be stored in the processor or in the memory.

[0207] Meanwhile, the operations according to the embodiments described above may be performed by a transmitting device and / or a receiving device according to the embodiments. The transmitting / receiving device may include a transmitting / receiving unit for transmitting and receiving media data, a memory for storing instructions (program code, algorithm, flowchart, and / or data) for a process according to the embodiments, and a processor for controlling the operations of the transmitting / receiving device.

[0208] The processor may be referred to as a controller or the like, and may correspond to, for example, hardware, software, and / or a combination thereof. The operations according to the above-described embodiments may be performed by the processor. Furthermore, the processor may be implemented as an encoder / decoder or the like for the operations of the above-described embodiments.

[0209] As described above, the relevant contents have been described in the best form for carrying out the embodiments.

[0210] As described above, the embodiments may be applied in whole or in part to a point cloud data transmission and reception device and system.

[0211] Those skilled in the art may make various changes or modifications to the embodiments within the scope of the embodiments.

[0212] Embodiments may include modifications / changes, which do not depart from the scope of the claims and their equivalents.

Claims

1. Step of encoding point cloud data; A step of encapsulating the above point cloud data into a file; and comprising the step of transmitting the above file; Method for transmitting point cloud data.

2. In paragraph 1, The file comprises a first track including the geometry of the point cloud data, and a second track including attributes of the point cloud data. Method for transmitting point cloud data.

3. In paragraph 1, Each track of the above file includes a sample entry containing configuration information about the point cloud data and a sample containing the point cloud data, The above sample further comprises a set of parameters relating to the point cloud data, Method for transmitting point cloud data.

4. In paragraph 1, the method: Further comprising a step of transmitting MPD information for the above point cloud data, The above MPD includes first adaptation set information for the geometry of the point cloud data and second adaptation set information for the attributes of the point cloud data, The above first adaptation set information includes a component descriptor for the geometry, The above second adaptation set information includes a component descriptor for the above attribute, The above MPD further includes information identifying initial information representing information about the entire space including the point cloud data. Method for transmitting point cloud data.

5. In paragraph 1, the method: Further comprising a step of transmitting initial information representing information about the entire space including the above point cloud data, The above initial information includes at least one of size information of the entire space or information about a subspace of the entire space, The information about the subspace includes at least one of an identifier for the subspace, a name for the subspace, or location information of a bounding box for the subspace. Method for transmitting point cloud data.

6. In paragraph 5, The initial information further includes at least one of a face ID for identifying a face generated based on points of the point cloud data, vertex information included in the face identified by the face ID, information indicating whether the subspace is an open space or a closed space, index information for the face, information indicating whether the subspace is a background or a foreground, or address information of an MPD regarding the point cloud data. Method for transmitting point cloud data.

7. An encoder that encodes point cloud data; An encapsulator for encapsulating the above point cloud data into a file; and a transmitter for transmitting the above file; comprising; Point cloud data transmission device.

8. Step of receiving a file containing point cloud data; a step of decapsulating the above file; and A step of decoding the above point cloud data; comprising: How to receive point cloud data.

9. In paragraph 8, The file comprises a first track including the geometry of the point cloud data, and a second track including attributes of the point cloud data. How to receive point cloud data.

10. In paragraph 8, Each track of the above file contains a sample entry containing configuration information about the point cloud data and a sample containing the point cloud data, The above sample further comprises a set of parameters relating to the point cloud data, How to receive point cloud data.

11. In paragraph 8, the method: Further comprising a step of transmitting MPD information for the above point cloud data, The above MPD includes first adaptation set information for the geometry of the point cloud data and second adaptation set information for the attributes of the point cloud data, The above first adaptation set information includes a component descriptor for the geometry, The above second adaptation set information includes a component descriptor for the above attribute, The above MPD further includes information identifying initial information representing information about the entire space including the point cloud data. How to receive point cloud data.

12. In paragraph 8, the method: Further comprising a step of receiving initial information representing information about the entire space including the above point cloud data, The above initial information includes the size information of the entire space, The above initial information contains information about a subspace of the entire space, Information about the subspace includes an identifier for the subspace, a name for the subspace, and location information of a bounding box for the subspace. How to receive point cloud data.

13. In paragraph 12, The initial information further includes at least one of a face ID for identifying a face generated based on points of the point cloud data, vertex information included in the face identified by the face ID, information indicating whether the subspace is an open space or a closed space, index information for the face, information indicating whether the subspace is a background or a foreground, or address information of an MPD regarding the point cloud data. How to receive point cloud data.

14. A receiving unit for receiving a file containing point cloud data; A decapsulator for decapsulating the above file; and A decoder for decoding the above point cloud data; comprising: Point cloud data receiving device.

Citation Information

Patent Citations

  • Apparatus and method for regerssion test

    KR1020220102490A

  • The high perfomance antiviral sterilization system securing the antibacterial effect and UV-C sterilization time of the silver fiber HEPA filter

    KR1020230009023A

  • Illumination device for skin care and operating method thereof

    KR102777467B1

  • KR20230028792A

  • KR20230141897A