Point cloud data transmission method, point cloud data transmission device, point cloud data reception method, and point cloud data reception device

CN122460084APending Publication Date: 2026-07-24FOUND FOR RES & BUSINESS SEOUL NAT UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FOUND FOR RES & BUSINESS SEOUL NAT UNIV OF SCI & TECH
Filing Date
2024-12-30
Publication Date
2026-07-24

Smart Images

  • Figure CN122460084A_ABST
    Figure CN122460084A_ABST
Patent Text Reader

Abstract

A method for transmitting point cloud data includes encoding point cloud data, encapsulating the point cloud data into a file, and transmitting the file. A method for receiving point cloud data includes receiving a file containing point cloud data, decapsulating the file, and decoding the point cloud data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to a technique for streaming three-dimensional multimedia content generated by geometry-based point cloud compression (G-PCC). Background Technology

[0002] With the introduction of the concepts of digital twins and metaverses, related technologies are becoming central to advanced industries and continue to expand. Accordingly, media services are expected to gradually evolve into the delivery of six degrees of freedom (6DoF) 360-degree virtual reality (VR) video or 3D multimedia content via streaming.

[0003] On the other hand, one widely used method for representing digital scenes, such as the metaverse, is generating 3D multimedia content using geometry-based point cloud compression (G-PCC). The reason for using G-PCC is that in digital scenes, the relationships between users and other objects are not fixed values ​​but change in real time; therefore, describing objects using G-PCC is efficient.

[0004] As mentioned above, although 3D multimedia content is expected to evolve into a streaming format, there is currently no technology for streaming 3D multimedia content generated in point cloud format. Summary of the Invention

[0005] Technical issues Based on the above-mentioned technical background, this invention aims to design signaling that includes spatial and temporal information to adaptively transmit three-dimensional multimedia content according to user location, environment, etc., thereby enabling the transmission of three-dimensional multimedia content in a streaming manner.

[0006] Technical solution According to embodiments of the present invention, a method for sending point cloud data may include: encoding the point cloud data; encapsulating the point cloud data in a file; and sending the file. According to embodiments of the present invention, a method for receiving point cloud data may include: receiving a file containing the point cloud data; decapsulating the file; and decoding the point cloud data.

[0007] Beneficial technical effects The method and apparatus of the present invention can efficiently encode and transmit point cloud data.

[0008] The method and apparatus of the present invention can encode and transmit point cloud data in a spatially adaptive manner. Attached Figure Description

[0009] To facilitate a better understanding of the present invention, embodiments of the invention will be described below in conjunction with the accompanying drawings. For a better understanding of the various embodiments described below, the following drawings and descriptions are provided, wherein the same reference numerals denote corresponding parts in all drawings, wherein: Figure 1 The schematic diagram illustrates the process of the method of the present invention for streaming geometry-based point cloud compression (G-PCC) content via HTTP-based Moving Picture Experts Group-Dynamic Adaptive Streaming (MPEG-DASH); Figure 2 A schematic diagram of objects within the user's field of view according to the present invention is shown; Figure 3 A schematic diagram of the object group of the present invention is shown; Figure 4 A schematic diagram of the object group within the user's field of view according to the present invention is shown; Figure 5 A schematic diagram of the architecture for transmitting G-PCC data according to the present invention is shown; Figure 6 A schematic diagram of multitrack encapsulation of G-PCC data based on the ISO Basic Media File Format (ISOBMFF) of the present invention is shown. Figure 7 The sample entries and sample diagrams of the tracks in the G-PCC file of the present invention are shown; Figure 8 A schematic diagram illustrating the grouping of G-PCC components within the MPEG-DASH Media Presentation Description (MPD) of the present invention is shown; Figure 9 A schematic diagram of the initial description document (IDD) of the present invention is shown; Figure 10 A schematic diagram of the method for describing spatial information based on IDD according to the present invention is shown; Figure 11 A schematic diagram of the method for sending point cloud data according to the present invention is shown; Figure 12 A schematic diagram of the method for receiving point cloud data according to the present invention is shown. Detailed Implementation

[0010] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings. The following detailed description, taken in conjunction with the drawings, is intended to explain the preferred embodiments and not to limit the embodiments of the invention to those described herein. The detailed description includes specific details that provide a full understanding of the embodiments. However, it will be apparent to those skilled in the art that the invention can be practiced without these specific details.

[0011] Most of the terminology used in these embodiments is derived from terms widely used in the relevant technical field. However, some terms are defined by the applicant, and their definitions will be explained in detail below. Therefore, these embodiments should be understood based on the intended meaning of the terms rather than on their literal names or superficial meanings.

[0012] Figure 1 According to an embodiment of the present invention, a flowchart of a method for streaming geometry-based point cloud compression (G-PCC) content via HTTP-based Moving Picture Experts Group-Dynamic Adaptive Streaming (MPEG-DASH) is schematically shown.

[0013] In step S10, when user 10 initiates a data transmission request to server 20 to view specified G-PCC-based content, server 20 sends user 10 an initial description document (IDD) including information related to the Media Presentation Description (MPD) describing the global spatial division.

[0014] This is because when spatial information is provided based on MPEG-DASH, a single MPD may contain a massive number of objects, resulting in an excessively large file size. Therefore, the above operation can reduce the time required for user parsing. This step can be applied selectively depending on the system operating environment.

[0015] In step S20, user 10 selects MPD describing the spatial information required by user 10 based on user 10's location and IDD, and requests the corresponding MPD from server 20 while transmitting user 10's location information to server 20.

[0016] In step S30, server 20 generates an MPD file describing the relationship between objects arranged in space and user 10 based on user 10's location information. The MPD file can be pre-generated based on three-dimensional spatial information, or it can be generated in real time according to the content requested by user 10 when a request is received in step S20.

[0017] In step S40, server 20 transmits the MPD file generated according to user 10's request to user 10.

[0018] In step S50, based on the received MPD, combined with the user's network environment and the relative distance and angle between the user and the object, the user 10 requests the required amount of G-PCC-based content from the server 20.

[0019] In step S60, after receiving the request, server 20 transmits the requested segmented data to user 10.

[0020] During the above operation, since the MPEG-DASH specification generates and distributes MPDs describing the relationship between users and content, server 20 generates and distributes MPDs by default.

[0021] To avoid excessive information in the MPD, the entire space can be divided into multiple regions, and an MPD that only describes the region corresponding to the location information of user 10 within the divided region can be generated.

[0022] In this case, since the generated MPD does not represent the entire domain space, server 20 also needs to generate an Idd that includes information indicating which part of the entire domain space each MPD describes. The generated IDD information includes the size of each partitioned region and the Uniform Resource Identifier (URI) that can receive the corresponding MPD. Table 1 shows a sample of IDDs.

[0023] [Table 1]

[0024] Therefore, step S20 can be described as follows. Based on its location and IDD, user 10 selects an MPD from the global space that describes the area to be divided by user 10, and requests the corresponding MPD from server 20 while transmitting user 10's location information. In step S30, server 20 generates the MPD corresponding to that area. The reason for generating the MPD after user 10's request is that various relationships, such as occlusion areas and the proportion of objects within the image, change depending on user 10's location in the space.

[0025] MPD will be updated based on the content length and the movement speed of user 10. If user 10 moves quickly, the relationship between objects and user 10 may change frequently, thus requiring more frequent updates.

[0026] Therefore, in this embodiment, considering factors such as segment length and the user 10's movement speed in space, the time to live (TTL) of the MPD is set as shown in Formula 1.

[0027] T mpd Indicates the duration, T s V represents the segment length. 用户 This indicates the movement speed of user 10.

[0028] [Formula 1]

[0029] The segment length is a time unit for dividing content into 20 segments on the server. Content is maintained in segments, and the space occupied by each segment is defined as the object size.

[0030] The content is generated based on G-PCC. This point cloud compression scheme utilizing geometric information is performed based on a three-dimensional tree structure called an octree. User 10 can receive objects individually or multiple objects at once, depending on the relative position and distance between user 10 and the objects.

[0031] The parsing quality of an object compressed by a node-based compression algorithm is determined based on the relative distance between the user 10 and the object. Therefore, a reference standard is needed when partitioning objects, grouping multiple objects, or specifying content quality.

[0032] The following section provides weight values ​​for specifying the two aspects mentioned above, and describes a method for grouping objects and determining quality using these weight values.

[0033] ND: Node Depth ND stands for Node Depth Weight, which corresponds to the resolution required to render an object at a specified quality. To determine ND, the minimum node depth value required for an object at a reference distance (e.g., 1 meter) is first defined, known as the Required Node Depth (RND). The RND is stored in the MPD as additional information for each object.

[0034] According to RND, the ND is as shown in Formula 2. Here, r represents the distance between the object and user 10.

[0035] [Formula 2]

[0036] TL: Tile Level TL represents the tile hierarchy weight. TL is a metric characterizing whether, when transmitting an object to user 10, the object needs to be partitioned and transmitted according to user 10's viewport (i.e., field of view (FOV)), or whether the object is transmitted as a group with other objects. Here, a tile corresponds to a unit used to store the partitioned space, and a single object can be partitioned and represented as multiple tiles according to the partitioning level. TL is represented by a hierarchy value, for example, having values ​​in the range [1,2,3,4,5], and each value is shown in Table 2.

[0037] [Table 2]

[0038] The TL value for each object is determined using a set of reference values. These reference values ​​represent the object's proportion within the user's FOV (Field of View) in terms of longitude and latitude, and are used as factors in determining the TL. The set of TL reference values ​​is configured as a threshold for upgrading to the TL level, and the boundary value θ, φ distinguishing TLn from TL(n+1) is defined as θ n , φ nThe longitude-based threshold (θ) and the latitude-based threshold (φ) can be represented as θ(θ1, θ2, θ3, θ4) and φ(φ1, φ2, φ3, φ4), respectively. Figure 2 The objects within the user viewport of the present invention are shown.

[0039] like Figure 2 As shown, the TL value is determined by comparing the proportions θ and φ of the object within the user's viewport with a set of reference values. The proportion of the object within the viewport is calculated as the difference between the maximum and minimum values ​​of θ and φ in the eight sets of polar coordinates (r, θ, φ) of the eight points constituting the object.

[0040] The θ and φ values ​​obtained as described above are compared with those in the TL reference value set to determine the TL value. For example, if the object's θ is greater than θ1, then the TL is 1. If θ is less than θ4, then the TL is 5. Based on this method, the object's θ and φ values ​​are identified, and the TL values ​​for the two axes are assigned as follows: When performing grouping or partitioning subsequently, grouping or partitioning is performed separately for each direction.

[0041] In one implementation, based on the aforementioned ND and TL, objects with the same ND and TL can be grouped into the same group and reconfigured as a single file.

[0042] For example, objects with a TL value equal to or greater than a reference value (e.g., 4) can be grouped into the same group. Objects belonging to this group correspond to objects within the user's field of view (FOV) with a TL value of 4 or 5.

[0043] Figure 3 According to an embodiment of the present invention, a group of objects is shown.

[0044] like Figure 3 As shown, by generating an additional node called the group root node above the root node of the existing node structure, objects divided into a single group are reconfigured as a single file, and these objects are decoded at the same node depth. During this process, if the ND values ​​of the individual objects are different, the quality of the objects displayed on the user's screen may vary significantly.

[0045] Therefore, when performing grouping based on TL, ND should also be considered within the group. As a result, objects within a single group have the same ND and TL. This means that objects within a group are decoded at the same node depth and presented to the user with the same quality.

[0046] Furthermore, to support adaptive transmission based on the user's viewing direction, various combinations of groups should be configured. Groups with fixed combinations may lead to inefficient use of user bandwidth. For example... Figure 4As shown, when configuring group combinations, the number of objects to be grouped is adjusted according to the user's viewing direction, and objects with the same ND and TL are grouped into the same group.

[0047] Reference Figure 4 To describe the details.

[0048] Figure 4 According to an embodiment of the present invention, a group of objects within the user viewport is shown.

[0049] As described above, the server generates various combinations of object groups to provide an object group suitable for the user's head orientation. These object groups are configured within the user's field of view (FOV) and can generate various combinations of object groups based on the user's head orientation. Figure 4 Based on the user's head orientation, different combinations of object groups are shown within the same TL and ND. The user's head orientation can be represented as θ and φ corresponding to latitude and longitude in polar coordinates. Furthermore, the user's FOV is represented by the service-providing terminal as (q, j). To generate object groups, it is necessary to determine whether all objects are included within the user's FOV relative to the latitude and longitude axes. However, for ease of explanation, references are now provided. Figure 4 Describe an implementation based on a longitude axis.

[0050] When the user's head direction is represented by θ and the FOV is defined as q, the user's FOV is represented as (θ - q / 2, θ + q / 2). In this case, when θ is 0, considering TL and ND, an object group is generated by combining objects included within the FOV. As θ increases, a new object group is generated each time the composition of objects within the FOV changes.

[0051] Based on the above description, the following three elements are needed to describe reconfiguring objects based on the user's location in the MPD.

[0052] 1. Adaptive Set 2. Characterization (resolution) 3. SRD (Spatial Relationship Description) 1. Adaptive Set An adaptive set typically corresponds to a content element that can be decoded individually. The adaptive set is defined as a separate adaptive set for objects partitioned / combined based on the aforementioned TL and ND.

[0053] 2. Characterization (playback quality) The representation indicates the quality of the content and can be adjusted according to the user's network environment. G-PCC-based content has various node depths, and providing all node depths as representation levels may lead to parallax between objects. Therefore, the number of levels (m) and depth interval (n) provided by the system are defined. Here, the number of levels refers to the number of playback quality levels of the content provided by the system, while the depth interval refers to the interval between quality levels.

[0054] Furthermore, m and n determine the node depth of the corresponding object assigned to the representation associated with ND. ND corresponds to the lowest depth level of the representation and generates at most m levels until the total node depth of the object has a difference of at least n levels.

[0055] For example, when an object with a leaf node depth of 16 has an ND of 6 and m=5, the object's representation number is 5, which corresponds to the maximum value of m.

[0056] The lowest representation corresponds to a node depth of 6. The highest representation corresponds to a node depth of 16. Between these values, node depths of 9 and 12 become representations, thus maintaining a difference of at least n levels. As a result, although the system can provide 5 resolution levels, only 4 representations are provided due to the relationship between the object's ND and the object's total node depth.

[0057] 3. SRD (Spatial Relationship Description) SRDs (Search Engine Registries) are elements that describe the relationships between users and objects, as well as the objects themselves. These elements are written as sub-elements of basic attributes within an adaptive set and are used to provide additional information about the adaptive set. The system's SRDs are... Write it in the form of .

[0058] r, θ, and φ represent the positional relationship from the user to the center of a single object or group of objects in polar coordinates. W, D, and H represent the width, depth, and height of the object. Q represents the quaternion value describing the object's rotation. Because the three-axis rotation values ​​represented by yaw, pitch, and roll will result in different final orientations depending on the order of rotation, quaternion values ​​are used to represent the object's rotation.

[0059] Based on the SRD value, users can place objects in space by recognizing information such as the object's position, size, and rotation.

[0060] Figure 5 According to an embodiment of the present invention, a schematic diagram of an architecture for transmitting G-PCC data is shown.

[0061] The actual visual scene A is captured by a camera group or camera device including multiple lenses and sensors. As a result of the acquisition, source point cloud data B is generated. One or more point cloud frames are encoded into an encoded G-PCC bitstream including an encoded geometric bitstream and an attribute bitstream E. Depending on a specific media container file format, one or more encoded bitstreams are configured as a media file F for file playback or as an initialization segment sequence and media segments Fs for streaming. According to an embodiment of the invention, the media container file format is the ISO Basic Media File Format [ISOBMFF] as specified in ISO / IEC 14496-12. The file container also includes metadata from the file or segment. Segment F is transmitted to the player using a transport mechanism.

[0062] The file F output by the file encapsulator is the same as the file F′ input to the file decapsulator. The file decapsulator processes file F′ or the received segment F′, extracts the encoded bitstream E′, and parses metadata. Subsequently, the G-PCC bitstream is decoded into a decoded signal D′, and point cloud data is generated from the decoded signal D′. If applicable, the point cloud data is rendered and displayed on a head-mounted display or other display device based on the current observation position, observation direction, or viewport determined by various sensors, such as head-tracking sensors. Tracking and positioning sensors or eye-tracking sensors may also be used. In addition to allowing the player to access appropriate portions of the decoded point cloud data, the current observation position or observation direction can also be used for decoding optimization. In viewport-related transmissions, the current observation position and observation direction are also provided to the strategy module to determine the trajectory to be received.

[0063] The above process also applies to live streaming and video-on-demand scenarios.

[0064] Figure 5 The definitions of each interface are as follows: E / E′: Encoded G-PCC bitstream F / F′: A media file that includes a track format specification, which may include constraints on the basic data stream contained within the track sample. although Figure 5 Although not shown in the diagram, G-PCC data can also be transmitted using the DASH scheme.

[0065] Figure 6 According to an embodiment of the present invention, a multitrack encapsulation of G-PCC data based on the ISO Basic Media File Format (ISOBMFF) is shown.

[0066] When a G-PCC bitstream is transmitted as multiple tracks for each G-PCC component, each G-PCC component bitstream is mapped to a separate track. There are two types of G-PCC component tracks: G-PCC geometry tracks and G-PCC attribute tracks. Each sample of a track includes at least one G-PCC unit that transmits a single G-PCC component data unit, and may not simultaneously include geometry and attribute data units or multiplexing of different attribute data units.

[0067] According to an embodiment of the present invention, the conditions for multi-track configuration are as follows.

[0068] a) There exists a G-PCC geometric orbit and it is used as the entry point.

[0069] b) There may be zero or more G-PCC attribute tracks. The track header identifier for each G-PCC attribute track is set to 0.

[0070] c) The sample entry includes a GPCC component information box to indicate the role of the stream contained in the track.

[0071] d) Introduce a track reference from the G-PCC geometric track to the G-PCC property track.

[0072] Tracks belonging to the same G-PCC sequence are time-aligned. Samples of the same point cloud frame contributed to different tracks have the same rendering time. If a sample contains a parameter set or tile list, the decoding time of the parameter set or tile list used in that sample is equal to or earlier than the decoding time of the corresponding G-PCC component data unit sample. If all parameter sets exist in samples across multiple tracks, the decoding time of samples including the Sequence Parameter Set (SPS) is equal to or earlier than the decoding time of samples including the Geometry Parameter Set (GPS), Attribute Parameter Set (APS), or tile list. Furthermore, all tracks belonging to the same G-PCC sequence have the same implicit or explicit edit list.

[0073] Figure 7 According to an embodiment of the present invention, sample entries and sample diagrams of tracks in a G-PCC file are shown.

[0074] Figure 7 (a) shows a sample entry for the G-PCC file track. Figure 7 (b) shows a track sample from the G-PCC file.

[0075] refer to Figure 7(a) The compressor name in the base class of the volumetric visual sample entry indicates the name of the compressor used with the recommended value " / 013GPCC encoding". The first byte represents the count of the remaining bytes. In this case, the value is represented as " / 013 (octal 13)". This value corresponds to 11 (decimal), which represents the number of bytes of the remaining string.

[0076] In addition, the configuration (config field) includes 7 sets of G-PCC decoder configuration record information.

[0077] In addition, the type field represents the type of G-PCC component delivered in each track.

[0078] refer to Figure 7 (b) Each sample corresponds to a single point cloud frame, and samples that contribute to the same point cloud frame across multiple tracks have the same presentation time.

[0079] Each sample includes one or more G-PCC cells of the G-PCC component indicated in the GPCC component information box that transmits the sample entry, and zero or more G-PCC cells that transmit parameter sets or tile lists. If a sample contains G-PCC cells that include parameter sets or tile lists, that cell appears before the G-PCC cells that transmit the G-PCC component. The sample structure of the G-PCC geometry track is as follows: Figure 7 As shown in (b).

[0080] Figure 8 According to an embodiment of the present invention, a schematic diagram is shown of grouping G-PCC components within an MPEG-DASH Media Presentation Description (MPD).

[0081] Figure 8 An exemplary DASH configuration is shown for grouping G-PCC components belonging to a single point cloud media within an MPD file.

[0082] Period 1 includes a preselection set 1, which may include a G-PCC descriptor. In this embodiment, the G-PCC descriptor may include parameter information related to the compression of G-PCC data. Preselection set 1 may be referenced by an adaptive set. Depending on the configuration of the point cloud data, multiple adaptive sets may exist, such as geometry, attribute (color), attribute (reflectivity), attribute (material), etc. Each adaptive set may include a G-PCC component descriptor and a corresponding representation. The G-PCC component descriptor may include parameter information for each component.

[0083] Figure 9 A schematic diagram of IDD is shown according to an embodiment of the present invention.

[0084] like Figure 5As shown, according to an embodiment, the point cloud data transmitting device may include a point cloud data acquisition unit, a point cloud data encoder, and / or a file / segment encapsulator. According to an embodiment, the point cloud data receiving device may include a file / segment decapsulator, a point cloud data decoder, and / or a point cloud data renderer.

[0085] According to an embodiment, the point cloud data transmitting device can adaptively transmit G-PCC data, and according to an embodiment, the point cloud data receiving device can adaptively receive G-PCC data.

[0086] According to the implementation method, the efficiency of descriptor documents can be improved when transmitting wide-area data based on hierarchical MPD. Objects can be transmitted adaptively using MPD that includes object information. Transmission efficiency can be improved by changing the configuration of objects on a group or tile basis. In overall operation, a user information feedback channel can be defined so that the server can push optimal content based on user information.

[0087] refer to Figure 9 The transmitting device of the present invention can transmit a file that also includes IDD information. When the receiver selects MPD, the receiver can select necessary information based on the user's location, the user's motion vector, the user's viewport, etc. The receiver's decoder can receive MPD based on the user's selection. The decoder can request segments from the server. The decoder can receive segmented data.

[0088] refer to Figures 3 to 4 Multiple objects can be grouped into a single group, and related information can also be included in the MPD and / or file. Object groups can vary depending on the user's location or viewport. Using the user's location or viewport information, the decoder can obtain the necessary information from the MPD and / or file for decoding and rendering the required G-PCC data.

[0089] The transmission method of the present invention can execute a scalable coding method for adaptive transmission.

[0090] The transmission method of the present invention can encode objects by hierarchically dividing them based on the user's location. For example, the division units may include groups and / or tiles.

[0091] The sending method of the present invention can encode objects after hierarchical division in a scalable manner.

[0092] The transmission method of the present invention can generate and transmit layered MPDs and / or file tracks for implementing wide-area transmission.

[0093] The transmission method of the present invention can divide space into regions of a predetermined size, represent each region as a separate MPD, and generate a descriptor document describing which location each MPD covers.

[0094] The transmission method of the present invention can generate object information in the MPD for user location-based adaptive selection. For example, the object information may include the object's position information, size information, and / or rotation information. Furthermore, the object information may include dynamic related information.

[0095] The transmission method of the present invention can establish a feedback channel to support the user's push operation. For example, by utilizing a feedback channel that enables the server to consider the user's location, orientation, and speed when pushing content, the corresponding MPD and / or file track information can be provided to the receiver's decoder to decode optimized G-PCC content for the user.

[0096] In this invention, IDD refers to an initial description document, and may be simply referred to as initial information, etc. The IDD may include the following information: Information used to distinguish the extent, background, and / or foreground elements of the open space: For example, IDD may include information to distinguish between background and foreground so that background elements that should be visualized are transmitted separately, regardless of the user's location in the open space.

[0097] Information used to distinguish between open and / or closed spaces when defining them: For example, the IDD may include information to determine whether elements not included in the corresponding MPD become visible when a user observes from a particular direction. Furthermore, in the case of open spaces, it may be necessary to additionally download objects corresponding to that direction via background data or an additional MPD.

[0098] The definition of the partitioned space and the URL of the MPD containing that space information: For example, to include spatial elements with curved shapes rather than simple shapes such as simple rectangles, the space described by the MPD is defined as precisely as possible by distinguishing between points and faces. Furthermore, face indexes can be used to indicate whether a corresponding face corresponds to open space. The face index can have the same meaning as a face ID.

[0099] The IDD of this invention can be configured as follows: The global structure of an IDD includes size information for the entire space. Specifically, the bounding box of an IDD includes the extent information of the entire space. For example, the bounding box may include the bounding box position and the bounding box size.

[0100] The spatial partitioning of an IDD includes information related to the subspace. Information related to the subspace of the global space may include the following: for example, a space ID representing the space's identifier value; a name representing the space's name in string format; a bounding box providing vertex information for implementing the corresponding space; faces providing information related to the faces formed by connecting vertices; attributes providing additional information about the corresponding space; a state providing information about whether the space is open or closed; a face indicating the index of faces with the same open / closed state; a type indicating whether the space corresponds to the background or foreground; and a reference indicating the URL address of the MPD. The MPD includes the URL address of the MPD.

[0101] The bounding box element may include the vertex IDs that make up the bounding box and the position information of the points corresponding to the vertex IDs.

[0102] A face element includes a face ID and vertex information that constitutes the face identified by the face ID.

[0103] The attribute element indicates whether the space is open or closed. In this case, the open / closed information of the space can be indicated by the status element.

[0104] Figure 10 A method for describing spatial information based on IDD is shown.

[0105] like Figure 10 As shown, a global space and subspaces of the global space may exist. The global space can be represented based on a background and a predetermined range. The space may include open space and / or closed space. By utilizing various types of information to represent these spaces, the decoder can perform spatial access.

[0106] The configuration of the above IDD can be restated as follows: <idd> <globalstructure> <boundingbox> <minx> 0< / minx> <miny> 0 <miny> <minz> 0 <minz> <maxx> 1000< / maxx> <maxy> 1000 <maxy> <maxz> 1000< / maxz> < / maxy> < / maxy> < / minz> < / minz> < / miny> < / miny> < / boundingbox> <spatialpartition> <Spaceid="S1"> <name> MainRoom< / name> <boundingbox> <vertices> <vertex id="A" x="0" y="0" z="0" / > <vertex id="B" x="2" y="0" z="0" / > <vertex id="C" x="2" y="1" z="0" / > <vertex id="D" x="3" y="1" z="0" / > <vertex id="E" x="3" y="3" z="0" / > ..< / vertices> < / boundingbox> <faces> <Face id=" <vertexref> ABCDEFGHIJ< / vertexref> <Face id="” <vertexrcf>ABBA' <face>... <Properties ="” <state> l,2 <state> <Properties state ="” <state> 3,4< / state> <type> Background< / type> <reference> <mpd> http: / example.com / mainroom.mpd< / mpd> < / reference> < / state> < / state> < / face> < / vertexrcf> < / faces> < / spatialpartition> < / globalstructure> < / idd> Therefore, as described above, IDD information may include global structure information. The global structure information may include bounding box information. The bounding box information may include the minimum and maximum values ​​of the position information for each axis of the bounding box.

[0107] The IDD information may include spatial partitioning information. The spatial partitioning information may include identification information corresponding to the spaces within the spatial partitioning. The spatial partitioning information may include the name information of the spaces identified by the spatial identification information. The spatial partitioning information may include the identification information of the vertices of the bounding box corresponding to the spaces identified by the spatial identification information, as well as the position information of the vertices.

[0108] The spatial partitioning information may include face information. The face information may include information identifying the faces. The face information may include reference information for the vertices included in the face identified by the face identification information. The spatial partitioning information may include face state information. The state information may indicate at least one state, either closed or open, through attribute information. When the attribute is open, the state information may indicate face IDs with the same open state. For example, when face IDs with the closed attribute state are 3 and 4, the state information may include face IDs with the same closed state. The spatial partitioning information may include type information. The type information may indicate the attributes of the space corresponding to the spatial partitioning, such as whether the space is background or foreground.

[0109] The spatial partitioning information may include MPD reference information. The MPD reference information may include address information used to obtain the MPD.

[0110] refer to Figure 8 The method for sending point cloud data of the present invention can transmit G-PCC configuration information to an MPD-based descriptor based on the DASH scheme.

[0111] The instruction to provide GPCC information can be sent via signaling through the scheme ID Uri in the basic attributes of MPD, for example, "urn:mpeg:mpegI:gpcc:2020:component".

[0112] To transmit G-PCCs, each item in the MPD may include a G-PCC component descriptor. This descriptor may include information associated with the corresponding G-PCC. For example, in `@component@group_id`, the descriptor may include a group identifier. In `@component@component_type`, the descriptor may indicate the G-PCC data type. When the value is "geom", the descriptor may indicate geometric coordinates, while when the value is "attr", the descriptor may indicate attribute color. For `@component@geometry_type`, the descriptor may include anchor point location information for the region via anchor points (X, Y, Z), geometric coordinate information such as position (X, Y, Z), and time-varying geometric position information such as dynamic position (X, Y, Z). Furthermore, the descriptor may transmit the final position information of points in the dynamic region via the final position.

[0113] The method for transmitting point cloud data according to the present invention can generate and transmit G-PCC component descriptors in an MPD (Multi-Level Device). The G-PCC component descriptor may include a group ID for identifying geometric data groups of point cloud data. The G-PCC component descriptor may also include a group ID for identifying attribute data groups.

[0114] The method for receiving point cloud data according to the present invention can receive and parse G-PCC component descriptors in an MPD (Multi-Party Descriptor). The G-PCC component descriptor may include a group ID for identifying geometric data groups of point cloud data. The G-PCC component descriptor may also include a group ID for identifying attribute data groups. Using the group IDs in the MPD, point cloud data associated with specific objects and / or specific spaces can be decoded efficiently and partially.

[0115] The method / apparatus of the present invention can adaptively send and receive point cloud data by recording location information in IDD and / or MPD.

[0116] The MPD of the present invention may further include additional elements for defining and verifying the relationship between the IDD and the MPD. For example, the MPD may further include IDD index information.

[0117] The method / apparatus of the present invention can verify the IDD in the MPD according to the following method. For example, information indicating the ID value can be added to the MPD item. The ID of the IDD can be indicated as, for example... <idd id="">or <mpd idd="">The form of the IDD is used. The ID of the IDD can be added to the sub-items of the MPD to indicate whether the IDDs are the same.

[0118] Figure 11 According to an embodiment of the present invention, a method for transmitting point cloud data is shown.

[0119] The method for transmitting point cloud data according to the present invention may include encoding the point cloud data (S1100).

[0120] The method for sending point cloud data according to the present invention may further include encapsulating the point cloud data in a file (S1110).

[0121] The method for sending point cloud data according to the present invention may further include transmitting the file (S1120).

[0122] refer to Figure 6 In a multi-track configuration, the file may include a first track containing geometric data of the point cloud data and a second track containing attributes of the point cloud data.

[0123] refer to Figure 7 For the sample entries and samples of the track, each track of the file may include a sample entry containing configuration information related to the point cloud data and a sample containing the point cloud data, and the sample may further include a set of parameters related to the point cloud data.

[0124] refer to Figure 8 For MPD, the method may further include sending MPD information for point cloud data. The MPD may include first adaptive set information for geometric data of the point cloud data and second adaptive set information for attributes of the point cloud data. The first adaptive set information may include component descriptors for the geometric data, while the second adaptive set information may include component descriptors for the attributes.

[0125] refer to Figure 9 For IDD, the method may further include sending initial information indicating global spatial information related to the point cloud data. The initial information may include at least one of global spatial size information or partial spatial information related to the global space. The partial spatial information may include at least one of a partial space identifier, a partial space name, or partial space bounding box location information.

[0126] refer to Figure 9 For the face, attributes, status, type, and MPD URL of the IDD, the initial information may further include at least one of the following: face ID identifying the face generated from the points based on the point cloud data; vertex information included in the face identified by the face ID; information indicating whether the partial space is open or closed; face index information; information indicating whether the partial space is background or foreground; or address information of the MPD associated with the point cloud data.

[0127] can be Figure 5 The transmitting apparatus shown performs a method for transmitting point cloud data. The point cloud data transmitting apparatus may include: an encoder configured to encode point cloud data; a wrapper configured to encapsulate the point cloud data in a file; and a transmitter configured to transmit the file.

[0128] Figure 12 According to an embodiment of the present invention, a method for receiving point cloud data is shown.

[0129] Figure 12 The receiving method shown can correspond to Figure 11 The reverse process of the sending method.

[0130] The method for receiving point cloud data according to the present invention may include receiving a file containing point cloud data (S1200).

[0131] The method for receiving point cloud data according to the present invention may further include decapsulating the file (S1210).

[0132] The method for receiving point cloud data of the present invention may further include decoding the point cloud data (S1220).

[0133] The received file may include a first track containing geometric data of point cloud data and a second track containing attributes of point cloud data.

[0134] Each track of the received file may include a sample entry containing configuration information related to the point cloud data and a sample containing the point cloud data, and the sample may further include a set of parameters related to the point cloud data.

[0135] refer to Figure 8 The MPD method may further include receiving initial information indicating global spatial information related to the point cloud data. The initial information may include at least one of the size information of the global space or partial spatial information related to the global space. The partial spatial information may include at least one of the identifier of the partial space, the name of the partial space, or the location information of the bounding box of the partial space.

[0136] refer to Figure 9 The method, in addition to the IDD (Integrated Data Generation) method, may further include receiving initial information indicating global spatial information including point cloud data. The initial information may include size information of the global space. The initial information may also include partial spatial information related to the global space. The partial spatial information may include an identifier of the partial space, a name of the partial space, and location information of the bounding box of the partial space.

[0137] The initial information may also include at least one of the following: a face ID that identifies a face generated from points based on point cloud data; vertex information contained in the face identified by the face ID; information indicating whether the partial space is open or closed; face index information; information indicating whether the partial space is background or foreground; or address information of MPD associated with the point cloud data.

[0138] can be Figure 5 The receiving device performs a method for receiving point cloud data. The data receiving device may include: a receiver configured to receive a file containing point cloud data, a decapsulator configured to decapsulate the file, and a decoder configured to decode the point cloud data.

[0139] The embodiments of the present invention have been described from the perspective of methods and / or apparatus, and the descriptions of the methods and apparatus are complementary and applicable to each other.

[0140] For ease of understanding, the accompanying drawings have been described separately. However, new embodiments can be designed by combining the embodiments described in the drawings. Furthermore, computer-readable recording media on which programs for performing the above embodiments are recorded, according to the actual needs of those skilled in the art, also fall within the scope of protection of this invention. The apparatus and method of this invention are not limited to the above configurations and methods. New embodiments can be configured by selectively combining all or part of the embodiments to allow for various modifications. Although preferred embodiments of the invention have been shown and described, the embodiments of the invention are not limited to the specific embodiments described above. Those skilled in the art can make various modifications without departing from the spirit of the invention as claimed in the claims. These modifications should not be understood separately from the core technology or scope of the invention.

[0141] Various components of the apparatus of the present invention can be implemented by hardware, software, firmware, or a combination thereof. Various components can be implemented on a single chip, for example, as a single hardware circuit. Alternatively, the components can be implemented on separate chips. At least one component of the apparatus can be configured as one or more processors capable of executing one or more programs, and the one or more programs can include instructions for causing the one or more processors to perform one or more operations / methods of the present invention. Executable instructions for performing the methods / operations of the apparatus of the present invention can be stored in a non-transitory computer-readable medium (CRM) or other computer program products configured to be executed by one or more processors, or can be stored in a transient CRM or other computer program products configured to be executed by one or more processors. Furthermore, the concept of memory in the present invention includes not only volatile memory (e.g., RAM) but also non-volatile memory, flash memory, PROM, etc. Furthermore, implementations in the form of carrier waves, for example, transmitted via the Internet, can also be included. Furthermore, the processor-readable recording medium can be deployed on a network-connected computer system, enabling processor-readable code to be stored and executed in a distributed manner. Furthermore, the memory of the present invention may include not only volatile memory (e.g., random access memory (RAM)), but also non-volatile memory, flash memory, programmable read-only memory (PROM), etc. Additionally, implementations in the form of carrier waves transmitted via, for example, the Internet may also be included. Furthermore, the processor-readable recording medium can be deployed on a network-connected computer system, enabling processor-readable code to be stored and executed in a distributed manner.

[0142] In this document, " / " and "," should be interpreted as "and / or". For example, "A / B" should be understood as "A and / or B", and "A, B" should be understood as "A and / or B". Furthermore, "A / B / C" means at least one of A, B, and / or C. Similarly, "A, B, C" means at least one of A, B, and / or C. Additionally, in this document, "or" should be understood as "and / or". For example, "A or B" could mean (1) only A, (2) only B, or (3) both A and B. In other words, "or" in this document can be interpreted as additionally or alternatively.

[0143] Terms such as "first" and "second" are intended to describe the components in the embodiments. However, the various components of the present invention are not limited by these terms. These terms are only used to distinguish one component from another. For example, a first user input signal may be referred to as a second user input signal, and similarly, a second user input signal may be referred to as a first user input signal. The use of these terms should be interpreted as not departing from the scope of protection of the present invention. Although both the first user input signal and the second user input signal are user input signals, they do not refer to the same user input signal unless the context clearly indicates otherwise.

[0144] The terminology used to describe embodiments is intended to describe particular examples and not to limit the scope of any particular embodiment. As used in the description of embodiments and claims, the singular form is intended to cover its plural reference unless the context clearly indicates otherwise. The expressions "and / or" are intended to cover all possible combinations of the related terms. The term "including" is intended to describe the presence of features, values, steps, elements, and / or components, but does not exclude the presence or addition of one or more other features, values, steps, elements, and / or components. Conditional expressions used in describing embodiments, such as "when..." or "in the case of...", should not be construed as limiting to optional cases. When a particular condition is met, it is intended to perform a related operation in response to that particular condition, or to interpret the related definitions accordingly.

[0145] Furthermore, the operations of the embodiments described in this invention can be performed by the transmitting and receiving apparatus of this invention, including a memory and / or a processor. The memory may store programs for processing and / or controlling the operations of this invention, and the processor may control the various operations described in this invention. The processor may also be referred to as a controller. The operations of this invention can be performed by firmware, software, and / or combinations thereof, and the firmware, software, and / or combinations thereof may be stored in a processor or memory.

[0146] The operation of the present invention can be performed by the transmitting and / or receiving apparatus of the present invention. The transmitting and receiving apparatus may include: a transmitting and receiving unit for transmitting and receiving media data, a memory for storing instructions (program code, algorithms, flowcharts, and / or data) for the process of the present invention, and a processor for controlling the operation of the transmitting and receiving apparatus.

[0147] The processor may be referred to as a controller, and may correspond, for example, to hardware, software, and / or a combination thereof. Operations of the present invention may be performed by the processor. Furthermore, the processor may be implemented as an encoder and / or decoder for performing operations of the present invention.

[0148] Methods of implementing the present invention As described above, the relevant contents of the present invention have been explained with reference to the best embodiment.

[0149] Industrial applicability As described above, the present invention can be applied, in whole or in part, to point cloud data transmission and reception devices and systems.

[0150] Those skilled in the art can make various modifications or changes to this invention within the scope of this invention.

[0151] This invention includes various modifications and alterations, and these modifications and alterations are intended to fall within the scope of the claims of this invention and their equivalents.< / mpd> < / idd>

Claims

1. A method for transmitting point cloud data, characterized in that, The method includes: The point cloud data is encoded; The point cloud data is encapsulated in a file; and Send the file.

2. The method according to claim 1, characterized in that: The file includes a first track containing geometric data of the point cloud data and a second track containing attributes of the point cloud data.

3. The method according to claim 1, characterized in that: Each track of the file includes a sample entry containing configuration information related to the point cloud data and a sample containing the point cloud data; and The sample also includes a set of parameters related to the point cloud data.

4. The method according to claim 1, characterized in that: It also includes sending information for the Media Presentation Description (MPD) of the point cloud data; The MPD includes a first adaptive set of information for the geometric data of the point cloud data and a second adaptive set of information for the attributes of the point cloud data. The first adaptive set information includes component descriptors for the geometric data; The second adaptive set information includes a component descriptor for the attribute; and The MPD also includes information identifying initial information related to global spatial information including the point cloud data.

5. The method according to claim 1, characterized in that: It also includes sending initial information representing global spatial information related to the point cloud data; The initial information includes at least one of the size information of the global space or information related to the subspaces of the global space; and The information related to the subspace includes at least one of the subspace's identifier, the subspace's name, or the location information of the subspace's bounding box.

6. The method according to claim 5, characterized in that: The initial information also includes at least one of the following: A face identifier (ID) that identifies the face generated from points based on the point cloud data; Including vertex information in the face identified by the face ID; Information indicating whether the subspace is an open space or a closed space; Index information for the face; Information indicating whether the subspace corresponds to the background or foreground; or Address information for the Media Presentation Description (MPD) used for the point cloud data.

7. An apparatus configured to transmit point cloud data, the apparatus comprising: An encoder, configured to encode the point cloud data; A wrapper, specifically configured to wrap the point cloud data in a file; as well as A transmitter configured to send the file.

8. A method for receiving point cloud data, characterized in that, The method includes: Receive a file containing the point cloud data; Decapsulate the file; and The point cloud data is decoded.

9. The method according to claim 8, characterized in that: The file includes a first track containing geometric data of the point cloud data and a second track containing attributes of the point cloud data.

10. The method according to claim 8, characterized in that: Each track of the file includes a sample entry containing configuration information related to the point cloud data and a sample containing the point cloud data, and The sample also includes a set of parameters related to the point cloud data.

11. The method according to claim 8, characterized in that: It also includes sending information for the Media Presentation Description (MPD) of the point cloud data; The MPD includes a first adaptive set of information for the geometric data of the point cloud data and a second adaptive set of information for the attributes of the point cloud data. The first adaptive set information includes component descriptors for the geometric data; The second adaptive set information includes a component descriptor for the attribute; and The MPD also includes information identifying initial information related to global spatial information including the point cloud data.

12. The method according to claim 8, characterized in that: It also includes receiving initial information representing global spatial information related to the point cloud data; The initial information includes the size information of the global space; The initial information includes information related to subspaces of the global space; and The information related to the subspace includes at least one of the subspace's identifier, the subspace's name, or the location information of the subspace's bounding box.

13. The method according to claim 12, characterized in that: The initial information also includes at least one of the following: A face identifier (ID) that identifies the face generated from points based on the point cloud data; Including vertex information in the face identified by the face ID; Information indicating whether the subspace is an open space or a closed space; The index information of the face; Information indicating whether the subspace corresponds to the background or foreground; or Address information for the Media Presentation Description (MPD) used for the point cloud data.

14. An apparatus configured to receive point cloud data, characterized in that, The device includes: A receiver configured to receive a file including the point cloud data; A decapsulator, configured to decapsulate the file; and A decoder, configured to decode the point cloud data.