Method for streaming g-PCC-based three-dimensional multimedia content

The method addresses the lack of streaming technologies for G-PCC-based 3D content by adapting MPD selection and data segmentation based on user location and network conditions, enabling efficient streaming via MPEG-DASH.

WO2025143589A1PCT designated stage expired Publication Date: 2025-07-03FOUND FOR RES & BUSINESS SEOUL NAT UNIV OF SCI & TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/019287
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-28
Filing Date
2024-11-29
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

There is currently no technology for transmitting three-dimensional multimedia content produced using point clouds in a streaming manner, particularly in adapting to user location, environment, and network conditions.

Method used

A method for streaming G-PCC-based content involves selecting an MPD based on user location, generating an MPD file describing spatial relationships, requesting and transmitting necessary content considering network environment and relative distance/angle, and segmenting data transmission using TTL, ND, and TL weights to ensure adaptive streaming.

Benefits of technology

Enables efficient streaming of three-dimensional multimedia content based on G-PCC, adapting to user location and network conditions, using MPEG-DASH standard, and ensuring consistent playback quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024019287_03072025_PF_FP_ABST
    Figure KR2024019287_03072025_PF_FP_ABST
Patent Text Reader

Abstract

This method for streaming G-PCC-based content comprises: a first step of selecting a media presentation description (MPD) that describes spatial information required for a user on the basis of a user location, and then requesting a related MPD while transferring information about the user location to a server; a second step of generating the MPD file that describes, on the basis of the received user location, the relationship between an object arranged in a space and the user, and transmitting same to the user; a third step of making, on the basis of the received MPD, a request to the server for necessary G-PCC-based content in consideration of the network environment of the user, and the relative distance and angle to the object, and the like; a fourth step in which the server transmits, to the user, the MPD file generated according to the user request; a fifth step in which the user makes, on the basis of the received MPD, a request to the server for necessary G-PCC-based content in consideration of variables such as the network environment of the user, and the relative distance and angle to the object; and a sixth step in which the server transmits, to the user, segment data of the requested G-PCC-based content.
Need to check novelty before this filing date? Find Prior Art

Description

Streaming method of 3D multimedia content based on G-PCC

[0001] The present invention relates to a technology for transmitting three-dimensional multimedia content produced based on G-PCC (Geometry-based Point Cloud Compression) in a streaming manner.

[0002]

[0003] The concepts of digital twins and the metaverse are being introduced, and related technologies are growing and establishing themselves at the heart of cutting-edge industries. Accordingly, media services are expected to evolve into streaming formats, such as 6DoF 360VR video or 3D multimedia content.

[0004] Meanwhile, one of the most widely used methods for describing the digital world, such as the metaverse, is to create 3D multimedia content based on Geometry-based Point Cloud Compression (G-PCC). G-PCC is used because, in the digital world, the relationship between the user and other objects is not a fixed value but constantly changing, making it efficient to describe objects using G-PCC.

[0005] As previously explained, although it is expected that 3D multimedia content will evolve into a form that provides streaming services, there is currently no technology for transmitting 3D multimedia content produced as point clouds in a streaming manner.

[0006]

[0007] The present invention was created against this technical background, and aims to design signaling including spatiotemporal information to adaptively transmit 3D multimedia content according to the user's location, environment, etc., thereby enabling transmission of 3D multimedia content in a streaming manner.

[0008]

[0009] In order to solve the above technical problem, a method for streaming G-PCC-based content includes a first step of selecting an MPD (media presentation description) that describes spatial information required by a user based on the user's location and then requesting a related MPD while transmitting the user's location information to a server, a second step of generating the MPD file that describes a relationship between an object arranged in space and the user based on the received user's location and transmitting the MPD file to the user, a third step of requesting a necessary amount of G-PCC-based content from the server based on the received MPD, taking into account the user's network environment and the relative distance and angle from the object, a fourth step of transmitting the MPD file generated by the server according to the user's request to the user, a fifth step of transmitting the necessary amount of G-PCC-based content from the server based on the received MPD, taking into account variables such as the user's network environment and the relative distance and angle from the object, and a sixth step of transmitting segment data of the requested G-PCC-based content to the user.

[0010] Prior to the first step, the present invention further includes a step of transmitting an IDD (Init Description Document) containing information on MPDs (media presentation descriptions) that divide and explain the entire space to the user, and in the first step, the MPD is selected based on the received IDD.

[0011] The TTL (time to live) of the above MPD is is defined as follows, is the retention time, is the segment length, is the user's movement speed.

[0012] The above MPD includes the content of grouping objects having the same TL (Tile Level) and ND (Node Depth) and reconstructing them into a single file, and when the user's head direction is θ and the field of view is q, the user's field of view is (θ-q / 2, θ+q / 2), and when θ=0, the object group is created by combining objects included in the user's field of view by considering the TL (Tile Level) and ND (Node Depth), and a new object group is created each time the composition of objects in the user's field of view changes while increasing θ.

[0013] The above same group includes at least two subgroups based on the user's field of view (FoV), and at least one object is identical between two adjacent subgroups among the subgroups.

[0014] The playback quality of the above streaming is set based on the number of levels (m) and the depth gap (n) provided by the system, and the number of levels is adjusted by the value obtained by dividing the difference between the ND (Node Depth) of the object and the ND of the Leaf Node by the depth gap.

[0015]

[0016] According to the present invention, 3D multimedia content produced based on G-PCC can be transmitted in accordance with the MPEG-DASH standard, thereby enabling streaming service.

[0017]

[0018] FIG. 1 is a flowchart schematically showing a method for streaming G-PCC-based content via MPEG-DASH in one embodiment of the present invention.

[0019] Figure 2 is a schematic diagram explaining TL.

[0020] Figures 3 and 4 are drawings illustrating how objects are grouped into the same group.

[0021]

[0022] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. However, detailed descriptions of well-known functions or components that may obscure the gist of the present invention will be omitted in the following description and the attached drawings. Additionally, throughout the specification, the term "including" a component does not exclude other components, unless specifically stated otherwise, but rather implies the inclusion of other components.

[0023] Additionally, while terms such as "first" and "second" may be used to describe various components, these components should not be limited by these terms. These terms may be used to distinguish one component from another. For example, without departing from the scope of the present invention, a first component may be referred to as a second component, and similarly, a second component may also be referred to as a first component.

[0024] The terminology used herein is merely used to describe specific embodiments and is not intended to limit the present invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, it should be understood that the terms "comprise" or "have" indicate the presence of a described feature, number, step, operation, component, part, or combination thereof, but do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0025] Unless specifically defined otherwise, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by those of ordinary skill in the art to which this invention pertains. Terms defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning within the context of the relevant technology, and shall not be construed in an idealized or overly formal sense unless explicitly defined herein.

[0026]

[0027] The present invention relates to a method for streaming G-PCC-based 3D multimedia content. In the following description, for convenience, 3D multimedia content produced based on G-PCC is abbreviated as "G-PCC-based content."

[0028] In one embodiment of the present invention, G-PCC-based content is streamed via MPEG-DASH (Dynamic Adaptive Streaming over HTTP).

[0029] FIG. 1 is a flowchart schematically showing a method for streaming G-PCC-based content via MPEG-DASH in one embodiment of the present invention.

[0030] In step 10, when a user (10) requests data transmission from a server (20) to view specific G-PCC-based content, the server (20) transmits an IDD (Init Description Document) containing information about media presentation descriptions (MPDs) that divide and describe the entire space to the user (10).

[0031] This is to reduce the time required for the user (10) to parse the file, as when providing spatial information based on MPEG-DASH, a single MPD contains too many objects, resulting in a large file size. This step may be selectively applied depending on the system environment.

[0032] At step S20, the user (10) selects an MPD that describes the spatial information he or she needs based on his or her location and IDD, and then requests the related MPD while transmitting his or her location information to the server (20).

[0033] At step S30, the server (20) generates an MPD file describing the relationship between the user and objects placed in space based on the user's location. This MPD file can be generated in advance based on 3D spatial information, or can be generated in real time according to the user's request at step S20.

[0034] At step S40, the server (20) transmits the MPD file generated according to the user's request to the user (10).

[0035] At step S50, the user (10) requests the server (20) for the necessary amount of G-PCC-based content based on the received MPD, taking into account the user's network environment and the relative distance and angle from the object.

[0036] At step S60, the server (20) that received the request transmits the requested segment data.

[0037]

[0038] In the above operation process, since MPEG-DASH stipulates that an MPD describing the relationship between a user and content be created and distributed, the server (20) basically creates and distributes an MPD.

[0039] In one example, to avoid filling the MPD with too much information, the entire space can be divided into multiple sections, and an MPD can be created that describes only the sections based on the user's location information.

[0040] In this case, since the generated MPD cannot represent the entire space, the server (20) additionally generates an IDD containing information about which space each MPD describes within the entire space. For this purpose, the generated IDD information includes the size of the identified space and a URI that can retrieve the corresponding MPD. Table 1 provides an example of an IDD.

[0041] Space URI(0,0,0 / 180,259,372)https: / www.IDD.kr / first_MPD(180,0,0 / 375,259,372)https: / www.IDD.kr / second_MPD(0,259,0 / 180,511,372)https: / www.IDD.kr / third_MPD(180,259,0 / 375,511,372)https: / www.IDD.kr / fourth_MPD

[0042]

[0043] To explain step S20 again in this regard, the user (10) selects an MPD that describes the partitioned space he or she needs from the entire space based on his or her location and IDD, and then requests the related MPD while transmitting his or her location information to the server (20). Then, in step S30, the server (20) generates an MPD corresponding to the space. The reason for generating the MPD after the user request is that various relationships, such as the Occlusion Area and the proportion of the Object on the screen, change depending on the user's location within the space.

[0044]

[0045] Meanwhile, MPD is updated based on factors such as the length of the content and the user's movement speed. If the user moves quickly, more frequent updates are required because the relationship between objects and the user can change frequently.

[0046] Therefore, in this embodiment, the TTL (time to live) of the MPD is set as in mathematical formula 1 by considering the segment length, the movement speed of the user in the space, etc. Here, is the retention time, is the segment length, is the user's movement speed.

[0047]

[0048] At this time, the segment length is the unit of time in which the server (20) divides the content. The content is maintained for the segment length, and the space occupied during the segment length time is defined as the object size.

[0049]

[0050] In the present invention, content is produced based on G-PCC. The Point Cloud compression method utilizing this geometry is based on a three-dimensional tree structure called an Octree. Users should be able to receive objects separately based on their relative position and distance from each other, or receive multiple objects at once.

[0051] In other words, the resolution quality of objects compressed using a node-based compression algorithm is determined by the relative distance between the user and the object. Therefore, a standard is needed when dividing or grouping objects, or specifying content quality.

[0052] Therefore, below, we present the weight values ​​used to specify these two, and explain how to use these values ​​to group and determine quality.

[0053] ND: Node Depth

[0054] ND stands for Node Depth Weighting, and it represents the required resolution required for an object to be consumed at a certain quality. To determine ND, the object first needs the Required Node Depth (RND), which is the minimum Node Depth value required at a reference distance (e.g., 1 m). This RND is stored in the MPD as additional information for each object.

[0055] According to RND, ND is defined as in Equation 2, where r is the distance between the object and the user.

[0056]

[0057] TL: Tile Level

[0058] TL stands for Tile Level Weight. TL is a measure of whether an object should be segmented and transmitted based on the user's viewport (FOV) when transmitting an object to the user, or transmitted as a group with other objects. Here, a tile is the same concept as a unit for dividing and storing space, and a single object can be segmented and expressed as multiple tiles depending on the segmentation level. TL is a level value, with values ​​ranging from [1, 2, 3, 4, 5], as shown in Table 2.

[0059] TL Meaning 1. Transmit a single object by dividing it into smaller pieces of two or more levels 2. Transmit a single object after dividing it into one level 3. Transmit a single object 4. Transmit a group containing one or more objects 5. Transmit a space containing one or more groups

[0060] The TL value of each object is determined through a set of reference values. The reference value set is an element for distinguishing TL and expresses the degree of user field of view range using latitude and longitude. Each TL reference value set is configured as a reference value for passing the TL stage, and the θ and φ values ​​that distinguish TL n and (n+1) are θn and φ, respectively. n It is defined as follows. The longitude reference (θ) and the latitude reference (φ) are class It can be expressed as. As shown in Fig. 2, the proportions θ, φ occupied by objects in the user's field of view are compared with a set of reference values ​​to determine the TL value. The proportion occupied by an object in the field of view is calculated as the difference between the largest θ, φ and the smallest θ, φ among the eight values ​​of the polar coordinate system representation (r, θ, φ) of the eight points constituting the object.

[0061] The TL value is determined by comparing the θ, φ obtained through this with the TL reference value set θ, φ. For example, if the θ of the object is greater than θ1, the TL is 1. If θ is less than θ4, the TL is 5. Through this method, the values ​​of θ, φ of the object are confirmed, and the TL values ​​of each of the two axes are determined. , It is designated as . When proceeding or dividing the group thereafter, each direction is divided / grouped separately.

[0062]

[0063] Meanwhile, in one embodiment, objects having the same ND and TL based on the above-described ND and TL can be grouped into the same group and reconstructed into a single file.

[0064] For example, objects with a TL value greater than a threshold (e.g., 4 or greater) can be grouped together. Objects in this group are objects with a TL of 4 or 5 within the user's field of view (FoV).

[0065] Objects grouped together are reconstructed into a single file by creating an additional node called a group root node above the root node of the existing node structure, as shown in Figure 3, and decoded with the same node depth. During this process, if the ND of individual objects differs, the quality of objects on the user's screen may show a very large gap.

[0066] Therefore, when grouping objects by TL, ND must also be considered. Consequently, objects within a group share the same ND and TL. This means that all objects within a group appear to the user with the same quality at the same node depth.

[0067] Furthermore, groups should be presented in various combinations to ensure adaptive transmission based on the user's viewing direction. This is because fixed group combinations waste user bandwidth. As shown in Figure 4, when grouping objects, the number of objects grouped is adjusted based on the user's viewing direction, grouping objects with the same ND TL into the same group.

[0068] This will be explained in more detail with reference to Fig. 4.

[0069] As described above, the server generates object groups of various combinations to provide an object group suitable for the user's head direction. The object groups are combined within the user's field of view range, and various object groups can be generated depending on the user's head direction. Fig. 4 shows object groups that are variously combined within the same TL and ND depending on the user's head direction. The user's head direction can be expressed as latitude and longitude values ​​θ and φ among the polar coordinate representations of the head direction. In addition, the user's field of view (FoV) is expressed as (q, j) by the service providing terminal. In order to generate an object group, it is necessary to determine whether the two axes of latitude and longitude are included within the user's field of view range. However, for the sake of convenience of understanding, an example based on the longitude axis is described, as in Fig. 4.

[0070] When the user's head direction is expressed as θ and the field of view is defined as q, the user's field of view is expressed as (θ-q / 2, θ+q / 2). At this time, when θ is 0, an object group is created by combining objects included in the field of view considering TL and ND, and a new object group is created whenever the composition of objects within the field of view changes as θ increases.

[0071]

[0072] Through the above explanation, the following three elements are required to describe objects reconstructed to fit the user's location in MPD.

[0073] 1. Adaptation Set

[0074] 2. Representation (resolution)

[0075] 3. SRD(Spartial Relationship Description)

[0076]

[0077] 1. Adaptation Set

[0078] An Adaptation Set is typically a content element that can be individually decoded. This Adaptation Set is defined as an individual Adaptation Set of objects segmented / combined through the previously described TL and ND.

[0079]

[0080] 2. Representation (playback quality)

[0081] Representation refers to the quality of content and is adjusted based on the user's network environment. G-PCC-based content has a variety of node depths, and providing all node depths as representations can lead to object-specific gaps. Therefore, the number of levels (m) and depth gap (n) provided systematically are defined. Here, the number of levels refers to the number of playback quality levels provided by the system, and the depth gap refers to the gap between quality levels.

[0082] m,n set the Node Depth of each object specified in Representation in conjunction with ND. ND is the Depth Level of the lowest Representation, and a maximum of m levels are created with a difference of at least n levels up to the total Node Depth of the object.

[0083] For example, if the ND of an object with a Leaf Node Depth of 16 is 6 and m = 5, the object has a maximum of m Representations, which are 5. The lowest Representation among them is Node Depth 6. The highest Representation is Node Depth 16. To allow an interval of at least n between them, node depths 9 and 12 are each Representation. As a result, the system can provide 5 resolutions, but only 4 Representations are provided due to the relationship between the object ND and the maximum Depth.

[0084]

[0085] 3. SRD(Spartial Relationship Description)

[0086] SRD(spartial relation desecription) is an item that describes the relationship between users and objects and the objects. It is written as a sub-item called Essential Property within the Adaptation Set and is an item that writes additional description about the Adaptation Set. The SRD of the system is ( Write in the form.

[0087] It is a polar coordinate system that expresses the positional relationship from the user to the center of an individual object or a group of objects. is the Width, Depth, and Height of the object. is a quaternion value that describes the rotation of the object. Since the three-axis-based rotation values ​​represented by yaw, pitch, and roll have a problem in that the final direction changes depending on the rotation order, the rotation of the object is expressed as a quaternion.

[0088] Based on the SRD value, the user can check information such as the object's location, size, and rotation to place the object in space.

[0089]

[0090] Meanwhile, the above-described G-PCC-based content streaming method of the present invention can be implemented as computer-readable code on a computer-readable recording medium. A computer-readable recording medium includes any type of recording device that stores data that can be read by a computer system.

[0091] Examples of computer-readable recording media include ROM, RAM, CD-ROM, magnetic tape, floppy disks, and optical data storage devices. Furthermore, computer-readable recording media can be distributed across network-connected computer systems, allowing computer-readable code to be stored and executed in a distributed manner. Furthermore, functional programs, codes, and code segments for implementing the present invention can be readily inferred by programmers in the technical field to which the present invention pertains.

[0092] The present invention has been described above, focusing on various embodiments thereof. Those skilled in the art will appreciate that the present invention can be implemented in modified forms without departing from its essential characteristics. Therefore, the disclosed embodiments should be considered illustrative rather than limiting. The scope of the present invention is set forth in the claims, not the foregoing description, and all differences within the scope equivalent thereto should be construed as being encompassed by the present invention.

Claims

1. In a method for streaming G-PCC-based content, The first step is to select an MPD (media presentation description) that describes the spatial information required for the user based on his / her location, and then request the related MPD while transmitting his / her location information to the server; A second step of generating an MPD file describing a relationship between an object placed in space and the user based on the location of the received user and transmitting the MPD file to the user; A third step of requesting the server for the necessary amount of G-PCC-based content based on the received MPD, taking into account the network environment of the user and the relative distance and angle from the object; Step 4: The server sends the MPD file generated according to the user's request to the user; Step 5: The user requests the server to provide as much G-PCC-based content as needed, taking into account variables such as the user's network environment and the relative distance and angle from the object based on the received MPD; Step 6: The server transmits segment data of the requested G-PCC-based content to the user; A method for streaming content based on G-PCC including:

2. In paragraph 1, Prior to the first step, the step of transmitting an IDD (Init Description Document) containing information about MPDs (media presentation descriptions) that divide and describe the entire space to the user is further included. A G-PCC based content streaming method, wherein the MPD is selected based on the received IDD in the first step.

3. In paragraph 1, The TTL (time to Live) of the above MPD is is defined as follows, is the retention time, is the segment length, A G-PCC based content streaming method, which is a user's movement speed.

4. In paragraph 1, The above MPD includes the content of grouping objects with the same TL (Tile Level) and ND (Node Depth) into an object group and reconstructing them into a single file. When the user's head direction is θ and the field of view is q, the user's field of view is (θ-q / 2, θ+q / 2). When θ=0, the object group is generated by combining objects included in the user's field of view by considering TL (Tile Level) and ND (Node Depth), and a new object group is generated each time the composition of objects within the user's field of view changes while increasing θ. Method for streaming content based on G-PCC.

5. In paragraph 1, A G-PCC-based content streaming method, wherein the playback quality of the above streaming is set based on the number of levels (m) and the depth gap (n) provided by the system, and the number of levels is adjusted by a value obtained by dividing the difference between the ND (Node Depth) of an object and the ND of a Leaf Node by the depth gap.

Citation Information

Patent Citations

  • Methods for switching between a MBMS download and an http-based delivery of dash formatted content over an IMS network

    KR1020140041912A

  • Wire cutter

    KR1020230018274A

  • System and method for predicting effectiveness of advertising based scored contents channel positive percentage

    KR1020240053882A

  • Noise diagnosis method of motor

    KR102822500B1

  • KR20230028792A