System for asset exchange

A metadata framework for immersive media conversion and rendering addresses the challenge of distributing immersive media across diverse devices by preserving scene information, enhancing efficiency and adaptability in network resource utilization.

CN120323032APending Publication Date: 2025-07-15TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480005282.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-11-14
Filing Date
2024-11-15
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Existing commercial networks are difficult to effectively distribute heterogeneous immersive media to diversified client devices. The lack of a standardized metadata framework leads to inefficient media format conversion and cannot meet the diversified needs of heterogeneous clients.

Method used

Adopting the independent mapped space (IMS) metadata and ITMF specification in ISO/IEC 23090 Part 28, combined with the glTF2.0 extension, a standardized metadata framework is provided for conversion and adaptation between scene graph formats, and supporting immersive media distribution of heterogeneous client devices.

Benefits of technology

It realizes immersive media format conversion across multiple client devices, keeps media features not lost, adapts to the diversified needs of heterogeneous clients, and improves distribution efficiency and quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120323032A_ABST
    Figure CN120323032A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a method for processing immersive media. The method comprises the following steps: receiving first scene information to be converted into a second scene graph format in a first scene graph format; the method can further comprise the steps that a metadata frame is obtained so that scene information stored in the scene graph can be reserved in the scene graph conversion process, and the metadata frame comprises a plurality of subsystems; converting the first scene into the second scene graph format using a metadata framework; and rendering the first scene in the second scene graph format based on the conversion. The plurality of subsystems may include: a first subsystem including information associated with geometric assets of the first scene; a second subsystem comprising information associated with animations of one or more assets in the first scene; and a third subsystem including information associated with a logical sequence of data in the first scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Cross - Reference to Related Applications

[0001] This application claims priority to U.S. Provisional Application Nos. 63 / 599,409, 63 / 599,426, and 63 / 599,480, filed on November 15, 2023, and U.S. Application No. 18 / 947,615, filed on November 14, 2024, the disclosures of which are incorporated herein by reference in their entireties. Technical Field

[0002] This disclosure describes embodiments of architectures, structures, and components generally related to systems and networks for distributing media, including video, audio, geometric (3D) objects, haptics, associated metadata, or other content for client devices. Specific embodiments are directed to systems, structures, and architectures for distributing media content to heterogeneous immersive and interactive client devices. Background Art

[0003] “Immersive media” generally refers to media that stimulates any or all of the human sensory systems (visual, auditory, somatosensory, olfactory, and possibly gustatory) to create or enhance the perception of the user's physical presence in the media experience, i.e., beyond the content distributed via existing (e.g., “traditional”) commercial networks for temporal two - dimensional (2D) video and corresponding audio; such temporal media is also referred to as “traditional media”.

[0004] Another definition of “immersive media” is media that attempts to create or mimic the physical world through digital simulations of dynamics and physical laws, thereby stimulating any or all of the human sensory systems in order to create the perception of the user's physical presence within a scene depicting a real or virtual world.

[0005] A rendering device with immersive media capabilities may refer to a device equipped with sufficient resources and capabilities to access, interpret, and render immersive media. In terms of media provided by a network, such devices are heterogeneous in terms of the number and format of media they can support. Similarly, media is heterogeneous in terms of the amount and type of network resources required for large - scale distribution of such media. “Large - scale” may refer to the distribution of media by service providers that achieve a distribution comparable to that of traditional video and audio media via networks (e.g., Netflix, Hulu, Comcast subscriptions, and Spectrum subscriptions).

[0006] In contrast, traditional rendering devices such as laptop displays, televisions, and mobile handheld device displays are homogeneous in capabilities because these devices currently all include rectangular displays that use frame-based 2D rectangular videos or still images as their primary visual media format. Some of the commonly used frame-based visual media formats in traditional rendering devices may include High Efficiency Video Coding / H.265, Advanced Video Coding / H.264, and Versatile Video Coding / H.266 for video media.

[0007] The term "frame-based" media refers to the characteristic that visual media includes one or more consecutive rectangular frames of images. In contrast, "scene-based" media refers to visual media organized by "scenes", where each scene refers to the various assets that jointly describe a visual scene.

[0008] Taking a forest as an example in visual media, a contrast example between frame-based visual media and scene-based visual media is illustrated. In the frame-based representation, a camera device (such as those provided on mobile phones) is used to capture the forest. The user enables the camera to focus on the forest, and the frame-based media captured by the mobile phone is the same as what the user observes through the viewfinder provided on the mobile phone, including the changes in the picture caused by the user actively moving the device. The resulting frame-based representation of the forest is a series of 2D images recorded by the camera at a standard rate (usually 30 frames per second or 60 frames per second). Each image is composed of a set of pixels, and the information types stored in each pixel are completely uniform and arranged continuously.

[0009] In contrast, a scene - based representation of a forest consists of individual assets that describe the various objects in the forest, along with a human - readable scene graph description that presents a large amount of metadata describing the assets or how to render these assets. For example, this scene - based representation can include individual objects called "trees", where each tree consists of a collection of smaller assets called "trunk", "branches", and "leaves". Each trunk can be described individually by a mesh that describes the complete 3D geometry of the trunk and a texture applied to the trunk mesh to capture the color and radiative properties of the trunk. Additionally, the trunk can be accompanied by additional information that describes the smoothness or roughness of the surface of the trunk or its ability to reflect light. The corresponding human - readable scene graph description can provide information on where to place the trunk relative to the viewport of a virtual camera focused on the forest scene. Additionally, the human - readable description can include information on how many branches to generate from a single branch asset called "branches" and where to place them in the scene. Similarly, the description can include how many leaves to generate and the position of these leaves relative to the branches and the trunk. Additionally, transformation matrices can provide information on how to scale or rotate the leaves so that they do not appear homogeneous. Overall, the individual assets that make up the scene differ in terms of the type and amount of information stored in each asset. Each asset is typically stored in its own file, but typically these assets are used to create multiple instances of the objects they are designed to create, for example, the branches and leaves of each tree.

[0010] Those skilled in the art will understand that the human - readable part of the scene graph has a large amount of metadata to not only describe the relationship of the assets to their position within the scene but also contain instructions on how to render the objects, for example, using various types of light sources, or using surface properties (to indicate that the object has a shiny metal versus a matte surface) or other materials (porous or smooth texture). Other information typically stored in the human - readable part of the graph is the relationship of an asset to other assets, for example, to form groups of assets that are rendered or processed as a single entity, such as a trunk with branches and leaves.

[0011] Examples of scene graphs with human - readable components include glTF 2.0, where the node tree component is provided in JavaScript Object Notation (JSON), which is a human - readable notation for describing objects. Another example of a scene graph with human - readable components is the Immersive Technology Media Format, where an OCS file is generated using another human - readable representation format, XML.

[0012] Another difference between scene-based media and frame-based media is that in frame-based media, the view created for a scene is the same as the view captured by the user via a camera, i.e., when the media is created. When frame-based media is presented by a client, the view of the presented media is the same as the view captured by, for example, the camera used to record the video in the media. However, with scene-based media, the user can use a variety of virtual cameras (e.g., thin lens cameras and panoramic cameras) in multiple ways to view the scene.

[0013] The distribution of any media over a network can employ media delivery systems and architectures that reformat the media from an input or network "ingested" media format into a delivery media format that is not only suitable for ingestion by the target client device and its applications but also conducive to "streaming" over the network. Thus, there can be two processes performed on the ingested media by the network: 1) converting the media from format A to format B that is suitable for ingestion by the target client, i.e., based on the ability of the client to ingest certain media formats, and 2) preparing the media for streaming.

[0014] The "streaming" of media generally refers to segmenting and / or packaging the media so that it can be transmitted over the network in successive smaller-sized "chunks" that are logically organized and ordered according to either or both of the temporal or spatial structure of the media. The "transformation" (sometimes referred to as "transcoding") of the media from format A to format B can be a process typically performed by the network or service provider before distributing the media to the client. Such transcoding can involve converting the media from format A to format B based on the prior knowledge that format B is a preferred or the only format that can be ingested by the target client or is more suitable for distribution over a constrained resource such as a commercial network. In many cases, but not all, both the steps of transforming the media and preparing the media for streaming are necessary before the client can receive and process the media from the network.

[0015] The one - step or two - step process (i.e., before distributing the media to the client) that acts on the network - ingested media results in a media format called the "distribution media format" or simply the "distribution format". Typically, if the network can obtain information to indicate that the client will need the transformed and / or streamed media object on multiple occasions (otherwise these multiple occasions would trigger the transformation and streaming of such media multiple times), then for a given media data object, these steps should be performed only once (if at all). That is, the data processing and transmission for the transformation and streaming of media are generally regarded as a source of latency and consume potentially significant network and / or computing resources. Therefore, a network design that cannot obtain information to indicate when the client has potentially stored in its cache or locally stored relative to the client a particular media data object will perform poorly compared to a network that can obtain such information.

[0016] For traditional rendering devices, the distribution format can be equivalent to or sufficiently equivalent to the "rendering format" that the client - side rendering device ultimately uses to create the rendering. That is, the rendered media format is a media format whose characteristics (resolution, frame rate, bit depth, color gamut, etc.) are closely related to the capabilities of the client - side rendering device. Some examples of the distribution format and the rendering format include: an HD video signal (1920 pixel columns × 1080 pixel rows) distributed by the network to a UHD client device with a resolution of (3840 pixel columns × 2160 pixel rows). In this scenario, the UHD client will apply a process called "super - resolution" to the HD distribution format to increase the resolution of the video signal from HD to UHD. Thus, the final signal format rendered by the client device is the "rendering format", which in this example is the UHD signal, while the HD signal includes the distribution format. In this example, the HD signal distribution format and the UHD signal rendering format are quite similar because both signals are linear video formats, and the process of converting the HD format to the UHD format is relatively simple and easy on most traditional client devices.

[0017] Alternatively, the preferred rendering format of the target client device may be significantly different from the ingested format received by the network. However, the client can access sufficient computing, storage, and bandwidth resources to transform the media from the ingested format into the necessary rendering format suitable for client - side rendering. In this scenario, the network can bypass the step of re - formatting the ingested media (e.g., "transcoding" the media from format A to format B) simply because the client can access sufficient resources to perform all media transformations without the network having to do so first. However, the network can still perform the steps of segmenting and packaging the ingested media so that the media can be streamed to the client.

[0018] Another alternative is that the ingested media received by the network is significantly different from the preferred presentation format of the client, and the client does not have access to sufficient computing, storage, and / or bandwidth resources to convert the media into the preferred presentation format. In such a scenario, the network can assist the client by performing part or all of the transformation from the ingested format to a format that is equivalent or nearly equivalent to the client's preferred presentation format on behalf of the client. In some architectural designs, this assistance provided by the network on behalf of the client is commonly referred to as "split rendering" or "adaptation" of the media.

[0019] Regarding the goal of converting a scene graph format X to another scene graph format Y, there are multiple problems to be solved as follows. The first problem is to define a general conversion between two representations of the same type of media object, media attribute, or rendering function to be performed.

[0020] The second problem is that for a specific instance of a scene graph (e.g., the scene graph representation using format X), each individual object and other parts in the scene graph need to be annotated with metadata including IMS. That is, the metadata used to annotate a specific instance of the scene graph should be directly related to the corresponding individual media objects, media attributes, and scene graph rendering characteristics represented in format X. Summary of the Invention

[0021] A method for processing an immersive media stream, the method being executed by at least one processor, and the method includes: obtaining a metadata framework to preserve the scene information stored in the scene graph during the process of scene graph conversion, the metadata framework including multiple subsystems; receiving first scene information to be converted from a first scene graph format to a second scene graph format; using the metadata framework to convert the first scene into the second scene graph format, and the multiple subsystems from the metadata framework used include one or more of the following: a first subsystem including information associated with the geometric assets of the first scene; a second subsystem including information associated with the animation of one or more assets in the first scene; and a third subsystem including information associated with the logical sequence of the data in the first scene; rendering the first scene in the second scene graph format based on the conversion.

[0022] A non - volatile computer - readable medium stores instructions for processing an immersive media stream. The instructions include one or more instructions that, when executed by one or more processors of a device, cause the one or more processors to: obtain a metadata framework to preserve scene information stored in a scene graph during the transformation of the scene graph. The metadata framework includes a first subsystem that includes information associated with a logical sequence of data in a first scene. Parameters of the first subsystem include one or more of the following: a first binary data container for storing various types of data; a second binary data container formed by a GL transmission - format binary file; and a third binary data container for vertex data during subdivision surface evaluation stored in an OpenSubdiv library; receive first scene information to be transformed into a second scene - graph format in a first scene - graph format; use the metadata framework to transform the first scene into the second scene - graph format; and render the first scene in the second scene - graph format based on the transformation.

[0023] A device for processing an immersive media stream includes at least one memory configured to store program code; and at least one processor configured to read the program code and operate according to the instructions of the program code. The program code may include: an acquisition code configured to cause the at least one processor to obtain a metadata framework to preserve scene information stored in a scene graph during the transformation of the scene graph. The metadata framework includes a first subsystem that includes metadata information associated with the animation of one or more assets in a first scene. The first subsystem includes animation parameters, and the animation parameters include one or more of the following: a data type for indicating the data type provided to an animation generator or renderer; a period for indicating the time pattern of the animation; a mode for specifying the input / output time or key frames of the animation in the form of an array of time samples; and an end time for indicating the time when the animation stops; a reception code configured to cause the at least one processor to receive first scene information to be transformed into a second scene - graph format in a first scene - graph format; a transformation code configured to cause the at least one processor to use the metadata framework to transform the first scene into the second scene - graph format; and a rendering code configured to cause the at least one processor to render the first scene in the second scene - graph format based on the transformation. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 is a schematic diagram of a process for distributing immersive media to a client via a network according to an embodiment.

[0025] Figure 2Schematic diagram of the process of immersive media over a network before distributing the media to a client according to an embodiment.

[0026] Figure 3 Exemplary embodiment of a data model for representing and streaming temporal immersive media according to an embodiment.

[0027] Figure 4 Exemplary embodiment of a data model for representing and streaming non - temporal immersive media according to an embodiment.

[0028] Figure 5 Schematic diagram of the process of capturing a natural scene and converting it into an immersive representation in an ingestion format that can be used for a network according to an embodiment.

[0029] Figure 6 Schematic diagram of the process of using 3D modeling tools and formats to create an immersive representation of a synthetic scene in an ingestion format that can be used for a network according to an embodiment.

[0030] Figure 7 System diagram of a computer system according to an embodiment.

[0031] Figure 8 Schematic diagram of a network serving multiple heterogeneous client endpoints.

[0032] Figure 9 Schematic diagram of a network providing adaptation information about a specific media represented in a media ingestion format according to an embodiment.

[0033] Figure 10 System diagram of a media adaptation process consisting of a media renderer - converter that converts a source media from its ingestion format into a suitable specific format according to an embodiment.

[0034] Figure 11 Schematic diagram of a network formatting an adapted source media into a data model suitable for representation and streaming according to an embodiment.

[0035] Figure 12 System diagram of a media streaming process that segments a data model into the payloads of network protocol packets according to an embodiment.

[0036] Figure 13 Sequence diagram of a network adapting a specific immersive media in an ingestion format into a streamable and suitable distribution format for a specific immersive media client endpoint according to an embodiment.

[0037] Figure 14A Depicts an exemplary architecture for a scene graph.

[0038] Figure 14BDepicts an extended example of the architecture depicted in FIG. 14 according to an embodiment.

[0039] Figure 15 Depicts an example of an annotated scenario diagram according to an embodiment.

[0040] Figure 16 Depicts an example of an annotated scenario diagram according to an embodiment.

[0041] Figure 17 Depicts mapping an IMS subsystem identifier to one or more nodes, pins, or attributes according to an embodiment.

[0042] Figure 18 Depicts an example of an IMS subsystem for organizing IMS metadata according to an embodiment.

[0043] Figure 19 Depicts exemplary information items corresponding to a buffer subsystem of metadata according to an embodiment.

[0044] Figure 20 Depicts exemplary information items corresponding to a scenario subsystem of metadata according to an embodiment.

[0045] Figure 21 Depicts exemplary information items corresponding to an animation subsystem of metadata according to an embodiment. Detailed Description

[0046] Although the present disclosure has described several exemplary embodiments, there are changes, permutations, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, it should be understood that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are thus within the spirit and scope of the present disclosure.

[0047] Figure 1 Illustrates a media streaming process 100 including a general sequence of steps that may be performed by a network cloud or edge device 104. At step 101, the network receives media stored in an ingest media format A from a content provider. The network processing step 102 prepares the media for distribution to the client by formatting the media into format B and / or by preparing to stream the media to the client 108. The media is streamed from 104 to the client via the network connection 105. The client 108 receives or ingests the distributed media and optionally prepares the media for presentation via the rendering process 106. The output of the rendering process 106 is the presented media in yet another possibly different format C at 107.

[0048] Figure 2Depicts a media transformation decision process 200, which illustrates the network logic flow for processing ingested media through manual or automated processes within a network. At 201, the media is ingested by the network from a content provider. If the attributes of the target client are not yet known, the attributes of the target client are obtained at 202. If needed, decision step 203 determines whether the network should assist in the transformation of the media. Only when the decision step determines that the network must or should assist in the transformation, the ingested media is transformed through process 204 to convert the media from format A to format B, resulting in transformed media 205. At 206, preparation is made for streaming the transformed or original-form media. At 207, the media is streamed to the client or media storage.

[0049] Figure 2 An important aspect of the logic in is decision process 203, which can be performed by a person or by an automated process. This decision step must determine whether the media can be streamed in its original ingested format A, or whether it must be transformed into a different format B to facilitate presentation of the media by the client.

[0050] Such a decision process 203 may need access to information that describes aspects or characteristics of the ingested media in such a way as to help process 203 make the best choice, that is, to determine whether the ingested media needs to be transformed before streaming it to the client, or whether the media should be streamed directly to the client in its original ingested format A.

[0051] For each of the above scenarios where the transformation of media from format A to another format can be done entirely by the network, entirely by the client, or jointly by the network and the client, e.g., for split rendering, it is clear that a terminology library of attributes of the media format may be needed so that both the client and the network have complete information to characterize the media and the work that must be done. Additionally, a terminology library of attributes of the client capabilities may be equally needed, e.g., in terms of available computing resources, available storage resources, and access to bandwidth. Further, a mechanism is needed to characterize the level of computational, storage, or bandwidth complexity of the ingestion format so that the network and the client can jointly or separately determine whether or when the network can adopt a split rendering step to distribute the media to the client. Additionally, if the transformation and / or streaming of a particular media object that the client has completed or will need to complete the rendering has already been done as part of the work of a previous scenario that processed the rendering, the network can completely skip the step of transforming and / or streaming the ingested media, provided that the client can still access or use the media that was previously streamed to the client. Finally, if the transformation from format A to another format is determined to be a necessary step to be performed by or on behalf of the client, a priority scheme for ordering the transformation processes of the various assets within the scenario can benefit an intelligent and efficient network architecture.

[0052] An example of such a terminology library for characterizing media is the so-called Independent Mapping Space (IMS) nomenclature, which is designed to assist in the conversion from one scenegraph format to another that may be completely different. The Independent Mapping Space will be defined in Part 28 of the ISO / IEC 23090 series of standards; this series of standards is informally referred to as "MPEG-I". According to the scope of Part 28, the IMS includes metadata and other information that describe common aspects of scene-based media formats. For example, scene-based media typically can provide a mechanism for describing the geometry of a visual scene. One aspect of the IMS in ISO / IEC 23090 Part 28 is to provide standards-based metadata for annotating the human-readable part of the scenegraph, thus guiding the conversion from one format to another, i.e., from one scene geometry description to another. This annotation can also be attached to the scenegraph as a separate binary component. The same guiding conversion may apply to cameras; i.e., many scenegraph formats provide means for describing the characteristics of a virtual camera, and the characteristics of the virtual camera can be used as part of the rendering process to create a viewport into the scene. The IMS in Part 28 is also intended to provide metadata for describing common camera types. The purpose of the IMS is to provide a nomenclature that can be used to describe common aspects across multiple scenegraph formats, thus enabling the conversion from one format to another guided by the IMS. This conversion enables asset exchange across multiple clients.

[0053] Another important aspect of ISO / IEC 23090 Part 28 is that it intentionally does not specify how to perform the conversion from one format to another. Instead, the IMS only provides guidance on how to characterize the common features of all scene graphs. In addition to the geometric and camera features of the scene graph, other common features of the scene include lighting and object surface properties, such as albedo, material, roughness, and smoothness.

[0054] Regarding the goal of converting a scene graph format X to another scene graph format Y, there are at least two potential problems to be solved as follows. The first problem is to define a general conversion between two representations of the same type of media object, media attribute, or rendering function to be performed. For example, the IMS metadata for a static mesh object can be expressed using a general code (such as: IMS_STATIC_MESH). The scene graph represented by the syntax of format X can use an identifier such as: FORMAT_X_STATIC_MESH to refer to the static mesh, while the scene graph represented by the syntax of format Y can use an identifier such as: FORMAT_Y_STATIC_MESH to refer to the static mesh. Defining the general conversion using the IMS in ISO / IEC 23090 Part 28 can include the mapping of FORMAT_X_STATIC_MESH to IMS_STATIC_MESH and FORMAT_Y_STATIC_MESH to IMS_STATIC_MESH. Therefore, the general conversion from format X static mesh to format Y static mesh can be facilitated by using the metadata IMS_STATIC_MESH of the IMS from ISO / IEC 23090 Part 28.

[0055] It should be noted that at the time of this disclosure, ISO / IEC JTC1 SC29 / WG7 (the 7th working group of MPEG) was still developing the first version of Part 28. The latest version of the specification released by WG7 is ISO / IEC JTC1 / SC29 WG7N00657, which was released by WG7 on July 22, 2023. Document N00657 does not provide a complete specification of the Independent Mapping Space (IMS), especially regarding the goal of establishing a set of standards-based metadata for the exchange of scene graphs.

[0056] Regarding the problem of defining metadata to facilitate the conversion from one scene graph format to another, one approach is to utilize the availability of unique tags and metadata defined in the ITMF series of specifications to create an Independent Mapping Space such as that planned in ISO / IEC 23090 Part 28 which is under development. The role of such a space is to facilitate media exchange from one format to another while preserving or closely preserving the information represented by different media formats.

[0057] In the ITMF specification, most of the nodes, node pins, and node attributes crucial for encoding ITMF scenarios are organized into node systems related to the functions of their encodings used in ITMF. However, ITMF does not define sufficient metadata to describe how to construct, organize, or access media data within a buffer for the purpose of animation. That is, within ITMF, there are many nodes and node groups related to the descriptions of geometry, materials, textures, etc. These are organized into specific groups according to the purposes they serve, and such groups and their constituent nodes are also specified in Part 28. For example, nodes related to geometry descriptions are defined within the set of "geometry nodes" in ITMF; nodes related to texture descriptions are defined within the set of nodes called "textures".

[0058] Although ITMF defines many nodes, pins, and attributes to describe the logical and physical relationships between scene assets (such as geometry, textures, materials, etc.), ITMF does not provide detailed metadata to precisely define how to organize the binary data associated with such assets within computer memory for the purpose of animation, nor does it define the precise mechanism for animating objects. This information is helpful for use cases where an application attempts to animate scene assets, which is a common use case for glTF players and other renderers. Since glTF provides a description of how to organize buffers for the purpose of animation, as well as the precise mechanism for animating the assets stored in such buffers, the IMS in Part 28 should also do this by specifying metadata that can assist in converting between the glTF media format (or other formats that define how animation should be performed) and other media formats.

[0059] Some definitions known to those skilled in the art are mentioned below.

[0060] Scene graph: A common data structure used in vector-based graphics editing applications and modern computer games, which arranges the logical and usually (but not necessarily) spatial representation of a graphical scene; a collection of nodes and vertices in a graph structure.

[0061] Scene: In the context of computer graphics, a scene is a collection of objects (e.g., 3D assets), object attributes, and other metadata, which include visual, acoustic, and physical characteristics that describe a specific setting bounded by space or time and involve the interaction of objects within that setting.

[0062] Node: The basic element of a scene graph, including information related to the logical or spatial or temporal representation of visual, audio, tactile, olfactory, gustatory, or related processed information; each node should have at most one output edge, zero or more input edges, and at least one edge (input or output) connected to it.

[0063] Base layer: The nominal representation of an asset, typically formulated to minimize the computational resources or time required to render the asset, or the time required to transmit the asset over a network.

[0064] Enhancement layer: A set of information that, when applied to the base layer representation of an asset, enhances the base layer to include features or capabilities not supported in the base layer.

[0065] Attribute: Metadata associated with a node, used to describe a specific characteristic or feature of the node in a canonical or more complex form (e.g., in terms of another node).

[0066] Binding LUT: A logical structure that associates metadata from the IMS of ISO / IEC 23090 Part 28 with metadata or other mechanisms used to describe the features or capabilities of a specific scene graph format (e.g., ITMF, glTF, Universal Scene Description).

[0067] Container: A serialized format for storing and exchanging information to represent all natural scenes, all synthetic scenes, or a mixture of synthetic and natural scenes, including all media resources required for the scene graph and the rendered scene.

[0068] Serialization: The process of converting a data structure or object state into a format that can be stored (e.g., in a file or memory buffer) or transmitted (e.g., over a network connection link) and later reconstructed (possibly in a different computer environment). When the resulting series of bits is reread according to the serialization format, this process can be used to create a clone that is semantically identical to the original object.

[0069] Renderer: An (usually software-based) application or process that, based on a selective mix of disciplines related to acoustic physics, optical physics, visual perception, audio perception, mathematics, and software development, given an input scene graph and an asset container, emits a typical visual and / or audio signal that is suitable for presentation on a target device or conforms to the desired characteristics specified by the properties of the rendering target node in the scene graph. For visual-based media assets, the renderer can emit a visual signal suitable for the target display or suitable for storage as an intermediate asset (e.g., repackaged into another container, i.e., used in a series of rendering processes in the graphics pipeline); for audio-based media assets, the renderer can emit an audio signal for presentation in multi-channel speakers and / or binaural headphones, or for repackaging into another (output) container. Popular examples of renderers include the real-time rendering features of game engines Unity and Unreal Engine.

[0070] Evaluation: Produces a result (e.g., an evaluation similar to a document object model of a web page) that turns the output from abstract to a concrete result.

[0071] Scripting language: An interpreted programming language that can be executed by a renderer at runtime to handle dynamic inputs and variable state changes of scene graph nodes that affect the rendering and evaluation of spatial and temporal object topologies (including physical forces, constraints, inverse kinematics, deformations, collisions), as well as energy propagation and transmission (light, sound).

[0072] Shader: A computer program that was originally used for shading (producing appropriate levels of brightness, darkness, and color within an image), but now performs various specialized functions in various fields of computer graphics special effects, or performs video post-processing unrelated to shading, or even functions completely unrelated to graphics.

[0073] Path tracing: A computer graphics method for rendering three-dimensional scenes such that the lighting of the scene is faithful to reality.

[0074] Temporal media: Media that is ordered by time; e.g., having start and end times according to a specific clock.

[0075] Non-temporal media: Media that is organized by spatial, logical, or temporal relationships; e.g., as in an interactive experience that is realized based on actions taken by one or more users.

[0076] Neural network model: A collection of parameters and tensors (e.g., matrices) that define the weights (i.e., numerical values) used in well-defined mathematical operations applied to visual signals to achieve an improved visual output, which may include interpolating new views of visual signals not explicitly provided in the original signal.

[0077] OCS: The human-readable part of the ITMF scene graph, which uses a unique identifier represented as "id=nnn", where "nnn" is an integer value.

[0078] IMS: Independent mapping space metadata standardized in ISO / IEC 23090, Part 28.

[0079] Pin: Input and output parameters for the nodes of a scene graph.

[0080] Attribute: A characteristic of a given node that cannot be changed by other nodes.

[0081] In the past decade, many devices with immersive media capabilities (including head-mounted displays, augmented reality glasses, handheld controllers, multi-view displays, haptic gloves, and gaming consoles) have been introduced to the consumer market. Similarly, holographic displays and other forms of volumetric displays will enter the consumer market in the next three to five years. Despite the immediate or upcoming availability of these devices, a coherent end-to-end ecosystem for distributing immersive media over commercial networks has not been realized for several reasons.

[0082] One of the obstacles to realizing a coherent end-to-end ecosystem for distributing immersive media over commercial networks is that the client devices that act as endpoints of such a distribution network for immersive displays are highly diverse. Some of these client devices support certain immersive media formats, while others do not. Some of these client devices are capable of creating immersive experiences based on traditional raster-based formats, while others are not. Different from a network designed only for distributing traditional media, a network that must support multiple display clients requires a large amount of information regarding the details of each of the client capabilities and the formats of the media to be distributed, and then such a network can adopt an adaptation process to convert the media into a format suitable for each target display and the corresponding application. Such a network will at least need to obtain information directly describing the characteristics of each target display and the media itself in order to determine the exchange of the media. That is, media information can be represented in different ways depending on how the media is organized according to various media formats; a network that supports heterogeneous clients and immersive media formats will need to obtain information that enables it to identify when one or more media representations (according to the specifications of the media format) essentially represent the same media information. Therefore, a major challenge in distributing heterogeneous media to heterogeneous client endpoints is to achieve media "exchange".

[0083] Media exchange can be considered to preserve the characteristics of the media after the media conversion is completed (or adapted in the conversion from format A to format B as described above). That is, the information represented by format A is either not lost or is very close to the information represented by format B.

[0084] Immersive media can be organized into "scenes" described by a scene graph, which are also called scene descriptions. To date, there are many popular scene-based media formats, including: FBX, USD, Alembic, and glTF.

[0085] Such scenes refer to the scene-based media as described above. The scope of the scene graph is to describe visual, audio, and other forms of immersive assets, which include specific settings as part of the presentation, such as actors and events occurring at specific locations in a building as part of the presentation (e.g., a movie). A list of all the scenes that make up a single presentation can be formulated into a scene inventory.

[0086] The techniques provided herein describe a collection of metadata to create a set of standardized metadata for representing or describing how media assets are stored and managed in a computer storage device (i.e., a "buffer").

[0087] The techniques provided herein describe a collection of metadata to create a set of standardized metadata for representing or describing how media assets are stored for animation and how to animate media assets.

[0088] The techniques provided herein describe a collection of metadata to create a set of standardized metadata for representing or describing media assets formatted according to various specifications being used as geometric objects for a specific scene. That is, a "superset" scene can include geometric assets formatted according to the specifications of Alembic (ABC), Universal Scene Description (USD), Filebox (FBX), and Graphics Language Transmission Format (glTF).

[0089] Figure 3 The temporal media representation 300 is depicted as an example representation of a streamable format for heterogeneous immersive media for time series. Figure 4 The non-temporal media representation 400 is depicted as an example representation of a streamable format for heterogeneous immersive media for non-time series. Both figures refer to scenes; Figure 3 refers to scene 301 for temporal media, and Figure 4 refers to scene 401 for non-temporal media. For both cases, the scene can be embodied by various scene representations or scene descriptions.

[0090] For example, in some immersive media designs, a scene can be represented by a scene graph, or as a multi-planar image (MPI), or as a multi-spherical image (MSI). Both MPI technology and MSI technology are examples of technologies that help create a display-independent scene representation for natural content (i.e., real-world images captured simultaneously from one or more cameras). On the other hand, scene graph technology can be used to represent natural images and computer-generated images in the form of a synthetic representation. However, for the case where content is captured by one or more cameras as a natural scene, creating such a representation requires a particularly large amount of computation. That is, creating a scene graph representation of naturally captured content is both time-consuming and computationally intensive, requiring complex analysis of natural images using photogrammetry or deep learning or both techniques in order to generate a synthetic representation that can subsequently be used to interpolate a sufficient number of views to meet the frustum filling requirements of a target immersive client display. Therefore, it is currently impractical to consider such synthetic representations as candidates for representing natural content because, given use cases that require real-time distribution, they cannot actually be created in real time. However, currently, the best candidate representation for computer-generated images is to use a scene graph with a synthetic model because computer-generated images are created using 3D modeling processes and tools.

[0091] This polarization of the best representations for natural and computer-generated content suggests that the optimal ingestion format for naturally captured content is different from the optimal ingestion format for computer-generated content, or from the best ingestion formats for natural content that are not necessary for real-time distribution applications. Therefore, the goal of the disclosed subject matter is to be robust enough to support multiple ingestion formats for visually immersive media, whether they are created naturally using physical cameras or created by computers.

[0092] The following are example techniques for representing a scene graph as a format that is suitable for representing visually immersive media created using computer-generated techniques, or natural captured content that employs deep learning or photogrammetry techniques to create a corresponding synthetic representation, i.e., for real-time distribution applications that are not necessary.

[0093] 1. OTOY's

[0094] OTOY's ORBX is one of several scenegraph technologies that can support any type of temporal or non-temporal visual media, including ray-traceable, traditional (frame-based), volumetric, and other types of synthetic or vector-based visual formats. ORBX is unique compared to other scenegraphs because ORBX provides native support for free and / or open-source formats for meshes, point clouds, and textures. ORBX is a scenegraph that is deliberately designed to facilitate exchange between multiple vendor technologies that operate on the scenegraph. Additionally, ORBX provides a rich material system, support for open shader languages, a robust camera system, and support for Lua scripting. ORBX is also the basis for an immersive technology media format that is released for licensing under royalty-free terms by the Immersive Digital Experience Alliance (IDEA). In the context of real-time distribution of media, the ability to create and distribute an ORBX representation of a natural scene depends on the availability of computing resources to perform complex analysis of the data captured by a camera and synthesize that data into a synthetic representation. To date, it has not been practical but is not impossible to provide sufficient computing for real-time distribution.

[0095] 2. Pixar's Universal Scene Description

[0096] Pixar's Universal Scene Description (USD) is another well-known and mature scenegraph that is popular in the VFX and professional content production communities. USD is integrated into Nvidia's Omniverse platform, which is a suite of tools for developers to create and render 3D models using Nvidia's GPUs. A subset of USD is released as USDZ by Apple and Pixar. USDZ is supported by Apple's ARKit.

[0097] 3. Khronos' glTF2.0

[0098] glTF2.0 is the latest version of the "Graphics Language Transmission Format" specification written by the Khronos 3D group. This format supports a simple scenegraph format that generally can support static (non-temporal) objects in a scene, including "png" and "jpeg" image formats. glTF2.0 supports simple animations, including support for translation, rotation, and scaling of basic shapes described by glTF primitives, i.e., for geometric objects. glTF2.0 does not support temporal media and therefore does not support video or audio.

[0099] 4. ISO / IEC 23090 Part 14 Scene Description is an extension of glTF2.0 that adds support for temporal media such as video and audio.

[0100] These known designs for scene representations of immersive visual media are provided only as examples and do not limit the ability of the disclosed subject matter to specify a process for adapting an input immersive media source to a format with specific characteristics suitable for a client endpoint device.

[0101] In addition, any or all of the above example media representations are currently adopting or may adopt deep learning techniques in the future to train and create a neural network model that can implement or facilitate the selection of a specific view to fill the frustum of a specific display based on the specific dimensions of the frustum. The views selected for the frustum of a specific display can be interpolated from existing views explicitly provided in the scene representation (e.g., from MSI technology or MPI technology), or they can be rendered directly from a rendering engine based on a specific virtual camera position, filter, or description of the virtual camera for these rendering engines.

[0102] Accordingly, the disclosed subject matter is robust enough to consider that there exists a relatively small but well-known set of immersive media ingestion formats that can adequately meet the requirements for real-time or "on-demand" (e.g., non-real-time) distribution of media captured naturally (e.g., using one or more cameras) or created using computer-generated techniques.

[0103] With the deployment of advanced network technologies such as 5G for mobile networks and fiber optic cables for fixed networks, the interpolation of views of immersive media ingestion formats using neural network models or network-based rendering engines has been further facilitated. That is, these advanced network technologies increase the capacity and capabilities of commercial networks because such advanced network infrastructures can support the transmission and delivery of an increasing amount of visual information. Network infrastructure management technologies such as multi-access edge computing (MEC), software-defined networking (SDN), and network function virtualization (NFV) enable commercial network service providers to flexibly configure their network infrastructure to adapt to changes in the demand for certain network resources, e.g., in response to dynamic increases or decreases in the demand for network throughput, network speed, round-trip latency, and computing resources. In addition, this inherent ability to adapt to dynamic network requirements also facilitates the network's ability to adapt immersive media ingestion formats to suitable distribution formats in order to support various immersive media applications with potentially heterogeneous visual media formats for heterogeneous client endpoints.

[0104] Immersive media applications themselves may also have different requirements for network resources, including gaming applications (which require significantly lower network latency to respond to real-time updates of game states), telepresence applications (which have symmetric throughput requirements for both the upstream and downstream portions of the network), and passive viewing applications (the demand for downstream link resources of which may increase depending on the type of client endpoint display consuming the data). Generally, any consumer-facing application may be supported by a variety of client endpoints, which have various on-board client capabilities for storage, computing, and power, and also have various requirements for specific media presentations.

[0105] Accordingly, the disclosed subject matter enables a well-equipped network (i.e., a network that incorporates some or all of the features of modern networks) to support multiple legacy and immersive media-capable devices simultaneously according to the characteristics specified therein:

[0106] 1. Flexibly utilize media ingestion formats that are practical for both real-time use cases and "on-demand" use cases of media distribution.

[0107] 2. Flexibly support natural and computer-generated content for both legacy client endpoints and immersive media-capable client endpoints.

[0108] 3. Support temporal and non-temporal media.

[0109] 4. Provide a process for dynamically adapting source media ingestion formats to suitable distribution formats based on the characteristics and capabilities of the client endpoint and the requirements of the application.

[0110] 5. Ensure that the distribution format can be streamed over an IP-based network.

[0111] 6. Enable the network to serve multiple heterogeneous client endpoints simultaneously, which may include legacy and immersive media-capable devices and applications.

[0112] 7. Provide an exemplary media presentation framework that facilitates organizing distributed media along scene boundaries.

[0113] The improved end-to-end embodiments enabled by the disclosed subject matter are implemented according to the processes and components described in the following Figures 3 to 16 detailed description.

[0114] Figure 3 and Figure 4 both employ a single exemplary global distribution format that has been adapted from the ingestion source format to match the capabilities of a specific client endpoint. As described above, Figure 3 the media shown in Figure 4The media shown is non-sequential. A particular global format is robust enough in its structure to accommodate a variety of media attributes, where each media attribute can be layered based on the amount of significant information contributed by each layer to the media presentation. Note that this layering process is already well-known state-of-the-art technology, as demonstrated by progressive JPEG and scalable video architectures (such as those specified in ISO / IEC 14496-10 (Scalable Advanced Video Coding)).

[0115] 1. Media streamed according to a global media format is not limited to traditional visual and audio media, but can also include any type of media information capable of generating signals that interact with a machine to stimulate human vision, hearing, taste, touch, and smell.

[0116] 2. Media streamed according to a global media format can be sequential or non-sequential media, or a mixture of both.

[0117] 3. By using a base layer and enhancement layer architecture to achieve a layered representation of media objects, this global media format can be further streamed. In one example, separate base and enhancement layers are calculated by applying multi-resolution or multi-tessellation analysis techniques to media objects in each scene. This is similar to the progressive rendering image formats specified in ISO / IEC 10918-1 (JPEG) and ISO / IEC 15444-1 (JPEG2000), but is not limited to raster-based visual formats. In an example embodiment, the progressive representation of a geometric object can be a multi-resolution representation of the object calculated using wavelet analysis.

[0118] In another example of the layered representation of a media format, the enhancement layer applies different attributes to the base layer, such as modifying the material properties of the surface of a visual object represented by the base layer. In yet another example, these attributes can modify the texture of the surface of the base layer object, such as changing the surface from a smooth texture to a porous texture, or from a matte surface to a glossy surface.

[0119] In yet another example of the layered representation, the surface of one or more visual objects in a scene can be changed from a Lambertian surface to a ray-traced surface.

[0120] In yet another example of the layered representation, the network will distribute the base layer representation to the client so that the client can create a nominal representation of the scene while waiting at the client for the transmission of additional enhancement layers to modify the resolution or other characteristics of the base representation.

[0121] 4. The resolution of the attributes or correction information in the enhancement layer is not explicitly coupled to the resolution of the objects in the base layer as in existing MPEG video and JPEG image standards today.

[0122] 5. This global media format support enables any type of information media that can be presented or driven by a rendering device or machine, thus enabling heterogeneous media format support for heterogeneous client endpoints. In one embodiment of a network that distributes media formats, the network will first query the client endpoint to determine the capabilities of the client, and if the client cannot meaningfully ingest the media representation, the network will remove the attribute layers that the client does not support, or adapt the media from its current format to a format suitable for the client endpoint. In one example of such adaptation, the network will convert volumetric visual media assets into a 2D representation of the same visual asset by using a network-based media processing protocol. In another example of such adaptation, the network can employ a neural network process to reformat the media into an appropriate format, or alternatively synthesize the views required by the client endpoint.

[0123] 6. A manifest for a complete or partially complete immersive experience (live event, game, or playback of an on-demand asset), organized as scenes, which is the minimum amount of information that current rendering and game engines can ingest to generate rendered content. The manifest includes a list of the individual scenes to be rendered for the entire immersive experience requested by the client. Associated with each scene is one or more representations of the geometric objects within the scene, which correspond to a streamable version of the scene geometry. One embodiment of a scene representation refers to a low-resolution version of the geometric objects for the scene. Another embodiment of the same scene refers to an enhancement layer for the low-resolution representation of the scene to add additional detail or increase tessellation to the geometric objects of the same scene. As described above, each scene can have more than one enhancement layer to incrementally increase the detail of the geometric objects of the scene in a progressive manner.

[0124] 7. Each layer of a media object referenced within a scene is associated with a token (e.g., a URI) that points to an address where the resource can be accessed within the network. Such resources are similar to a CDN and the client can fetch content from the CDN.

[0125] 8. The token for the representation of a geometric object can point to a location within the network or a location within the client. That is, the client can signal to the network that its resources are available for network-based media processing.

[0126] Figure 3 Depicted is a temporal media representation 300, which includes an embodiment of a global media format for temporal media as follows. The temporal scene manifest includes a list of scenes 301. Scene 301 refers to a list of components 302 that individually describe the processing information and the type of media asset that includes scene 301. Component 302 refers to asset 303, which further refers to a base layer 304 and an attribute enhancement layer 305. A list of unique assets not previously used in other scenes is provided in 307.

[0127] Figure 4 Depicts an atemporal media representation 400, which includes an embodiment of a global media format for atemporal media as shown below. Information for a scene 401 is not associated with clock-based start and end durations. Scene 401 refers to a list of components 402 that separately describe processing information and the types of media assets that include scene 401. Component 402 refers to an asset 403, which further refers to a base layer 404 and property enhancement layers 405 and 406. Additionally, scene 401 refers to other scenes 401 for atemporal media. Scene 401 also refers to a scene 407 for a temporal media scene. List 406 identifies unique assets associated with a particular scene that have not been previously used in a higher-order (e.g., parent) scene.

[0128] Figure 5 Illustrates an example embodiment of a natural media synthesis process 500 for synthesizing an ingestion format from natural content. Camera unit 501 uses a single camera lens to capture a scene of a person. Camera unit 502 captures a scene with five divergent fields of view by mounting five camera lenses around an annular object. The arrangement in 502 is an exemplary arrangement commonly used to capture omnidirectional content for VR applications. Camera unit 503 captures a scene with seven convergent fields of view by mounting seven camera lenses on the inner diameter portion of a sphere. Arrangement 503 is an exemplary arrangement commonly used to capture a light field for a light field or holographic immersive display. Natural image content 509 is provided as an input to a synthesis process 504, which can optionally employ a neural network training process 505 using a set of training images 506 to produce an optional capture neural network model 508. Another process commonly used in place of training process 505 is photogrammetry. If model 508 is created during the process 500 depicted in Figure 5 then model 508 becomes one of the assets in an ingestion format 510 for natural content. An annotation process 507 can optionally be performed to annotate scene-based media with IMS metadata. Exemplary embodiments of ingestion format 510 include MPI and MSI.

[0129] Figure 6An embodiment of a synthetic media ingestion creation process 600 is illustrated to create an ingestion format for synthetic media (e.g., computer-generated imagery). A LIDAR camera 601 captures a point cloud 602 of a scene. CGI tools, 3D modeling tools, or another animation process for creating synthetic content are employed on a computer 603 to create 604 CGI assets over a network. An action capture suit with sensors 605A is worn by an actor 605 to capture a digital record of the actions of the actor 605 to produce animated MoCap data 606. Data 602, 604, and 606 are provided as inputs to a synthesis process 607 that outputs a synthetic media ingestion format 608. The format 608 can then be input into an optional IMS annotation process 609, the output of which is an IMS-annotated synthetic media ingestion format 610.

[0130] Techniques for representing and streaming the above-described heterogeneous immersive media can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 7 A computer system 700 is depicted that is suitable for implementing certain embodiments of the disclosed subject matter.

[0131] The computer software can be encoded using any suitable machine code or computer language that can be subject to assembly, compilation, linking, or similar mechanisms to create code that includes instructions that can be executed directly by a computer central processing unit (CPU), a graphics processing unit (GPU), etc., or executed through interpretation, microcode execution, etc.

[0132] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0133] Figure 7 The components shown for the computer system 700 are exemplary in nature and are not intended to impose any limitations on the scope of use or functionality of the computer software implementing embodiments of the present disclosure. The configuration of the components should also not be construed as having any dependencies or requirements related to any one or combination of the components illustrated in the exemplary embodiments of the computer system 700.

[0134] The computer system 700 may include certain human-machine interface input devices. Such human-machine interface input devices may respond to inputs generated by one or more human users through, for example, tactile inputs (such as keystrokes, swipes, data glove movements), audio inputs (such as speech, clapping), visual inputs (such as gestures), and olfactory inputs (not depicted). The human-machine interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (such as voice, music, ambient sound), images (such as scanned images, photographic images obtained from a still-image camera), and video (such as two-dimensional video, three-dimensional video including stereoscopic video).

[0135] The input human-machine interface devices may include one or more of the following (only one of each is depicted): keyboard 701, mouse 702, touchpad 703, touch screen 710, data glove (not depicted), joystick 705, microphone 706, scanner 707, camera 708.

[0136] The computer system 700 may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback through the touch screen 710, data glove (not depicted), or joystick 705, but there may also be tactile feedback devices that do not function as input devices), audio output devices (such as speakers 709, headphones (not depicted)), visual output devices (such as screen 710, including CRT screens, LCD screens, plasma screens, OLED screens (each screen having or not having touch screen input capabilities, each screen having or not having tactile feedback capabilities), some of which are capable of outputting two-dimensional visual output or more than three-dimensional output through means such as stereoscopic output; virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted)), and printers (not depicted).

[0137] The computer system 700 may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW 720 with CD / DVD or similar media 721, thumb drives 722, removable hard disk drives or solid-state drives 723, traditional magnetic media such as tapes and floppy disks (not depicted), devices based on dedicated ROM / ASIC / PLD such as secure dongles (not depicted), etc.

[0138] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other volatile signals.

[0139] The computer system 700 may also include an interface to one or more communication networks. The network can be, for example, wireless, wired, or optical. The network can further be a local area network, wide area network, metropolitan area network, in-vehicle network, and industrial network, real-time network, delay-tolerant network, and so on. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, in-vehicle and industrial networks including CANBus, and so on. Some networks typically require an external network interface adapter attached to certain general-purpose data ports or peripheral buses (749) (such as, for example, the USB port of the computer system 700); other networks are typically integrated into the core of the computer system 700 by attaching to the system bus as described below (for example, an Ethernet interface is integrated into a PC computer system or a cellular network interface is integrated into a smartphone computer system). Using any of these networks, the computer system 700 can communicate with other entities. Such communication can be unidirectional receive-only (such as broadcast TV), unidirectional transmit-only (such as CANbus to certain CANbus devices), or bidirectional (such as communicating with other computer systems using a local area digital network or a wide area digital network). Certain protocols and protocol stacks can be used on each of the networks and network interfaces described above.

[0140] The above-mentioned human-machine interface devices, human-accessible storage devices, and network interfaces can be attached to the core 740 of the computer system 700.

[0141] The core 740 may include one or more central processing units (CPUs) 741, a graphics processing unit (GPU) 742, a dedicated programmable processing unit 743 in the form of a field-programmable gate array (FPGA), a hardware accelerator 744 for certain tasks, and so on. These devices, together with read-only memory (ROM) 745, random access memory 746, and internal mass storage devices (such as internal non-user-accessible hard disk drives, SSDs, etc.) 747, can be connected via a system bus (748). In some computer systems, the system bus 748 can be accessed in the form of one or more physical plugs to enable expansion by adding additional CPUs, GPUs, etc. Peripheral devices can be attached directly to the system bus 748 of the core or attached to the system bus 748 of the core via a peripheral bus 749. Architectures for peripheral buses include PCI, USB, and so on.

[0142] The CPU 741, GPU 742, FPGA 743, and accelerator 744 can execute certain instructions that, when combined, can constitute the aforementioned computer code. The computer code can be stored in the ROM 745 or RAM 746. Transitional data can also be stored in the RAM 746, while permanent data can be stored in, for example, the internal mass storage device 747. Fast storage and retrieval of any memory device can be achieved by using a cache memory that can be closely associated with one or more of the CPU 741, GPU 742, mass storage device 747, ROM 745, RAM 746, etc.

[0143] A computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and the computer code can be media and computer code that are specially designed and constructed for the purposes of this disclosure, or they can be of the type well-known and available to those skilled in the computer software arts.

[0144] By way of example and not limitation, a computer system having architecture 700 (and particularly core 740) can provide functionality due to one or more processors (including CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible, computer-readable media. Such computer-readable media can be media associated with a user-accessible mass storage device as introduced above, as well as certain storage devices of core 740 having a non-volatile nature, such as core-internal mass storage device 747 or ROM 745. The software implementing various embodiments of this disclosure can be stored in such devices and executed by core 740. Depending on specific requirements, the computer-readable media can include one or more memory devices or chips. The software can cause core 740 (and particularly the processors therein (including CPU, GPU, FPGA, etc.)) to perform specific processes or specific portions of specific processes described herein, including defining data structures stored in RAM 746 and modifying such data structures according to processes defined by the software. Additionally or alternatively, the computer system can provide functionality through logic hard-wired or otherwise embodied in circuitry (e.g., accelerator 744) that can replace the software or operate in conjunction with the software to perform specific processes or specific portions of specific processes described herein. In appropriate instances, references to software can encompass logic and vice versa. In appropriate instances, references to computer-readable media can encompass circuitry (such as an integrated circuit (IC)) storing software for execution, circuitry embodying logic for execution, or both. This disclosure encompasses any suitable combination of hardware and software.

[0145] Figure 8Illustrated is an exemplary network media distribution system 800 that supports various traditional and heterogeneous displays with immersive media capabilities as client endpoints. The content acquisition process 801 uses Figure 6 or Figure 5 to capture or create media using the example embodiments therein. In the content preparation process 802, ingestion formats are created and then transferred to the network media distribution system using the transfer process 803. The gateway 804 can serve customer premise equipment to provide network access to various client endpoints of the network. The set-top box 805 can also be used as customer premise equipment to provide access to aggregated content through a network service provider. The wireless demodulator 806 can be used as a mobile network access point for mobile devices, e.g., as shown by the mobile handset and display 813. In this particular embodiment of the system 800, a traditional 2D television 807 is shown connected directly to the gateway 804, the set-top box 805, or the WiFi router 808. A laptop computer with a traditional 2D display 809 is illustrated as a client endpoint connected to the WiFi router 808. A head-mounted 2D (raster-based) display 810 is also connected to the router 808. The lenticular light field display 811 is shown connected to the gateway 804. The display 811 includes a local computing GPU 811A, a storage device 811B, and a visual rendering unit 811C that creates multiple views using ray-based lenticular optics technology. The holographic display 812 is shown connected to the set-top box 805. The display 812 includes a local computing CPU 812A, a GPU 812B, a storage device 812C, and a Fresnel mode, wave-based holographic visualization unit 812D. The augmented reality headset 814 is shown connected to the wireless demodulator 806. The headset 814 includes a GPU 814A, a storage device 814B, a battery 814C, and a volumetric visual rendering component 814D. The dense light field display 815 is shown connected to the WiFi router 808. The display 815 includes multiple GPUs 815A, a CPU 815B, a storage device 815C, an eye tracking device 815D, a camera 815E, and a dense ray-based light field panel 815F.

[0146] Figure 9 Illustrated is an embodiment of immersive media distribution with a scene analyzer for a default viewport process 900 that is capable of serving traditional and heterogeneous displays with immersive media capabilities as previously depicted in Figure 8 . Content is created or acquired in process 901, which is further embodied in Figure 5 and Figure 6 for natural content and CGI content respectively. The content 901 is then converted into an ingestion format using the create network ingestion format process 902. Similarly, in Figure 5 andFigure 6 Process 902 is further embodied in , respectively for natural content and CGI content. The ingested media is optionally annotated with IMS metadata by a scene analyzer 911 having an optional IMS representation. The ingested media format is transmitted to the network and stored on a storage device 903. Optionally, the storage device may reside in the immersive media content producer's network and be remotely accessed by an immersive media network distribution process (unnumbered), as depicted by the dashed line bisecting 903. Client and application specific information is optionally available on a remote storage device 904, which may optionally exist remotely in an alternative "cloud" network.

[0147] like Figure 9 As depicted in FIG. 1 , the network orchestration process 905 serves as the primary source and sink of information to perform the important tasks of the distribution network. In this particular embodiment, the process 905 may be implemented in a format that is unified with the other components of the network. However, Figure 9 The tasks depicted in process 905 in form the basic elements of the disclosed subject matter. The orchestration process 905 can further employ a two-way message protocol with the client to facilitate all processing and distribution of the media according to the characteristics of the client. In addition, the two-way protocol can be implemented across different transport channels (i.e., control plane channels and data plane channels).

[0148] Process 905 receives information about the characteristics and properties of client 908 and further collects requirements about the application currently running on 908. This information may be obtained from device 904, or in an alternative embodiment, may be obtained by directly querying client 908. In the case of direct querying of client 908, assuming a two-way protocol ( Figure 9 ) exists and is operable so that the client can communicate directly with the orchestration process 905.

[0149] The arrangement process 905 also initiates Figure 10 908B, or the client 908 itself may initiate a "pull" request for the media 906 from the storage device 909. The orchestration process 905 may employ a bidirectional messaging interface ( s ) to communicate with and process 910 for media adaptation and segmentation as described in . As the process 910 adapts and segments the ingested media, the media is optionally transferred to an intermediate storage device depicted as media prepared for distribution storage device 909. When the distribution media is prepared and stored in device 909, the orchestration process 905 ensures that the immersive client 908 receives the distribution media and corresponding descriptive information 906 via its network interface 908B through a "push" request, or the client 908 itself may initiate a "pull" request for the media 906 from the storage device 909. The orchestration process 905 may employ a bidirectional messaging interface ( Figure 9to perform a "push" request (not shown in the figure) or initiate a "pull" request through the client 908. The immersive client 908 may optionally employ a GPU (or a CPU not shown) 908C. The distribution format of the media is stored in the storage device or storage cache 908D of the client 908. Finally, the client 908 visually presents the media via its visualization component 908A.

[0150] Throughout the process of streaming immersive media to the client 908, the orchestration process 905 will detect the status of the client progress via the client progress and status feedback channel 907. The detection of the status can be performed by means of a two-way communication message interface ( Figure 9 (not shown in the figure).

[0151] Figure 10 depicts a specific embodiment of a scene analyzer for the media adaptation process 1000, such that the ingested source media can be appropriately adapted to meet the requirements of the client 908. The media adaptation process 1001 includes multiple components that facilitate the adaptation of the ingested media into an appropriate distribution format for the client 908. These components should be considered exemplary. In Figure 10 it, the adaptation process 1001 receives an input network state 1005 to follow the current traffic load on the network; client 908 information, which includes property and feature descriptions, application features and descriptions, and the current state of the application, as well as the client neural network model (if available) to assist in mapping the geometry of the client's frustum to the interpolation capabilities of the ingested immersive media. This information can be obtained by means of a two-way message interface ( Figure 10 (not shown in the figure). The adaptation process 1001 ensures that when creating the adapted output, it is stored in the client-adapted media storage device 1006. The scene analyzer 1007 with an optional IMS notation process is Figure 10 depicted in it as an optional process that can be executed preferentially or as part of a network automation process for media distribution.

[0152] The adaptation process 1001 is controlled by the logic controller 1001F. The adaptation process 1001 also employs a renderer 1001B or a neural network processor 1001C to adapt a specific ingestion source media into a format suitable for the client. The neural network processor 1001C uses the neural network model in 1001A. Examples of such neural network processors 1001C include the Deepview neural network model generators described in MPI and MSI. If the media is in 2D format but the client must have 3D format, the neural network processor 1001C can invoke a process to use highly correlated images from the 2D video signal to derive a volumetric representation of the scene depicted in the video. An example of a suitable renderer 1001B can be a modified version of the OTOY Octane renderer (not shown), which will be modified to interact directly with the adaptation process 1001. The adaptation process 1001 can optionally employ a media compressor 1001D and a media decompressor 1001E, depending on the need for these tools regarding the format of the ingestion media and the format required by the client 908.

[0153] Figure 11 Depicts the distribution format creation process 1100. The adapted media packaging process 1103 packages the media that now resides on the client-adapted media storage device 1102 from the media adaptation process 1101 (depicted as process 1000 in Figure 10 ). The packaging process 1103 formats the adapted media from process 1101 into a robust distribution format 1104, for example, Figure 3 or Figure 4 the exemplary format shown in. The manifest information 1104A provides the client 908 with a list of the scene data assets 1104B it can expect to receive. The list 1104B depicts a list of visual assets, audio assets, and tactile assets, each with its corresponding metadata.

[0154] Figure 12 Depicts the packetization processing system 1200. The packetization process 1202 divides the adapted media 1201 into independent packets 1203 suitable for streaming to the client 908.

[0155] Figure 13The components and communications shown in sequence diagram 1300 are explained as follows: The client endpoint 1301 initiates a media request 1308 to the network distribution interface 1302. This request 1308 includes information for identifying the media requested by the client, which identification can be achieved through a URN or other standard nomenclature. The network distribution interface (also referred to as client 1302) responds to request 1308 with a profile request 1309, which requests the client 1301 to provide information about its currently available resources (including computing, storage, battery charge percentage, and other information characterizing the current operating state of the client). The profile request 1309 also requests the client to provide one or more neural network models, if such models are available at the client, which the network can use for neural network inference to extract the correct media view or perform correct media view interpolation to match the characteristics of the client's rendering system. The response 1311 from the client 1301 to the interface 1302 provides a client token, an application token, and one or more neural network model tokens (if such neural network model tokens are available at the client). The interface 1302 then provides the client 1301 with a session ID token 1311. The interface 1302 then requests the ingest media server 1303 with an ingest media request 1312, which includes the URN or other standard name of the media identified in request 1308. The server 1303 replies to request 1312 with a response 1313 that includes an ingest media token. The interface 1302 then provides the media token from response 1313 to the client 1301 in an invocation 1314. Then, the interface 1302 initiates an adaptation process for the media requested in 1308 by providing the ingest media token, the client token, the application token, and the neural network model token to the adaptation interface 1304. The interface 1304 requests access to the ingest media by providing the ingest media token to the server 1303 at an invocation 1316 to request access to the ingest media asset. The server 1303 provides the ingest media access token to the interface 1304 in a response 1317 in response to request 1316. The interface 1304 then requests the media adaptation process 1305 to adapt the ingest media located at the ingest media access token for the client, application, and neural network inference model corresponding to the session ID token created at 1313. The request 1318 from the interface 1304 to the process 1305 contains the required tokens and the session ID. The process 1305 provides the adapted media access token and the session ID to the interface 1302 in an update 1319. The interface 1302 provides the adapted media access token and the session ID to the packaging process 1306 in an interface invocation 1320. The packaging process 1306 provides the interface 1302 with a response 1321, which has a packaged media access token and the session ID in the response 1321.Process 1306 provides the packaged asset, URN, and the packaged media access token for the session ID to the packaged media server 1307 in response 1322. The client 1301 executes request 1323 to initiate the streaming of the media asset corresponding to the packaged media access token received in message 1321. The client 1301 executes other requests and provides a status update to interface 1302 in message 1324.

[0156] Figure 14A Illustrates an exemplary scenario graph architecture 1400. The human-readable scenario graph description 1401 serves as the part of the scenario graph that stores the spatial, logical, physical, and temporal aspects of the attached assets. The description 1401 also contains references to binary assets that further include the scenario. Associated with the description 1401 is the binary asset 1402. Figure 14 illustrates that there are four binary assets for the exemplary graph, including: binary asset A 1402, binary asset B 1402, binary asset C 1402, and binary asset D 1402. The references 1403 from the description 1401 are also illustrated as: reference 1403 to binary asset A, reference 1403 to binary asset B, reference 1403 to binary asset C, and reference 1403 to binary asset D. Figure 14B Illustrates an example of an extended scenario graph architecture.

[0157] Figure 15An exemplary annotated scenario graph architecture 1500 is provided, where IMS subsystem metadata 1503* (where * represents a character in the figure) is written directly into the human-readable part 1501 of the scenario graph architecture 1500. In this example, the IMS subsystem metadata 1503* includes multiple subsystems of metadata: 1503A, 1503B, 1503C, 1503D, 1503E, 1503F, 1503G, and 1503H, where each subsystem is associated with its own unique IMS subsystem identifier tag, and the unique IMS subsystem identifier tag corresponds to what is depicted for item 1503 in the figure. Mapping 1504* (where * represents a character in the figure) further provides additional information for a unique ITMF tag (obtained from the ITMF series specifications), and the unique ITMF tag fully or partially characterizes the information contained in each section of the human-readable part 1501. Such mappings 1504* depicted in the figure include: 1504A, 1504B, 1504C, 1504D, 1504E, 1504F, and 1504G. 1504H does not have a mapping to a unique ITMF tag because there is no such node group in the ITMF. In this case, the metadata for 1504H is fully defined within the IMS (rather than from the ITMF). The IMS metadata written into the human-readable part 1501 includes the information depicted in the mappings 1504* as described above. The scenario graph architecture 1500 further includes scenario assets 1502.

[0158] Figure 16 An exemplary annotated scenario graph architecture 1600 is provided, where IMS subsystem metadata 1606* (where * represents a character in the figure) is written directly into the binary part 1603 of the architecture, rather than storing such metadata in the human-readable part 1601 of the scenario graph architecture 1600 or otherwise storing such metadata in the human-readable part 1601 of the scenario graph architecture 1600 (as Figure 15In (depicted in). In this example, the IMS subsystem metadata 1606* includes multiple subsystems of metadata: 1606A, 1606B, 1606C, 1606D, 1606E, 1606F, 1606G, and 1606H, where each subsystem is associated with its own unique IMS system identifier label, and the unique IMS system identifier label corresponds to that depicted for item 1606 in the figure. The mapping 1604* (where * represents a character in the figure) further provides additional information of a unique ITMF label (obtained from the ITMF series specifications), and the unique ITMF label fully or partially characterizes the information contained in the human-readable part 1601. Such mappings 1604* depicted in the figure include: 1604A, 1604B, 1604C, 1604D, 1604E, 1604F, and 1604G. 1604H does not have a mapping to a unique ITMF label because there is no such node group in the ITMF. In this case, the metadata of 1604H is fully defined within the IMS (rather than from the ITMF). The IMS metadata written to the binary part 1603 includes the information depicted in the mapping 1604* as described above. The scene graph architecture 1600 further includes scene assets 1602.

[0159] Figure 17Depicts an example mapping 1700 of an IMS subsystem identifier 1702* (where * represents a character in the figure) to one or more unique tags 1701 from the ITMF series specification version 2.0. The IMS subsystem identifier 1702* includes: IMS_ID_1702A, IMS_ID_1702B, IMS_ID_1702C, IMS_ID_1702D, IMS_ID_1702E, IMS_ID_1702F, IMS_ID_1702G, IMS_ID_1702H, IMS_ID_1702I, IMS_ID_1702J, IMS_ID_1702K, IMS_ID_1702L, IMS_ID_1702M, IMS_ID_1702N, IMS_ID_1702O, IMS_ID_1702P, IMS_ID_1702Q, IMS_ID_1702R, and IMS_ID_1702S. The mapping 1700 illustrates (for illustrative purposes): IMS_ID_1702A is mapped to an ITMF tag for a numerical node; IMS_ID_1702B is mapped to ITMF tags for a render target node, a film setup node, an animation setup node, a core node, and a render AOV node; IMS_ID_1702C is mapped to an ITMF tag for a render target node; IMS_ID_1702D is mapped to an ITMF tag for a camera node; IMS_ID_1702E is mapped to an ITMF tag for a lighting node; IMS_ID_1702F is mapped to an ITMF tag for an object layer node; IMS_ID_1702G is mapped to an ITMF tag for a material node; IMS_ID_1702H is mapped to an ITMF tag for a medium node; IMS_ID_1702I is mapped to an ITMF tag for a texture node; IMS_ID_1702J is mapped to an ITMF tag for a transform node; IMS_ID_1702K is mapped to an ITMF tag for a render layer node; IMS_ID_1702L is mapped to an ITMF tag for a render channel node; IMS_ID_1702M is mapped to an ITMF tag for a camera imager node; IMS_ID_1702N is mapped to an ITMF tag for a custom lookup table node; IMS_ID_1702O is mapped to an ITMF tag for a post-processor node; IMS_ID_1702P is mapped to an ITMF tag for an unknown node; IMS_ID_1702Q is mapped to an ITMF tag for a node graph node; IMS_ID_1702R is mapped to an ITMF tag for a node pin; and IMS_ID_1702S is mapped to an ITMF tag for a node attribute.

[0160] Figure 18 Illustrates an exemplary system architecture 1800 for organizing the IMS subsystems described in the disclosed subject matter. In this example, the following IMS subsystems are defined: 1801A is an independent mapped space numerical node subsystem; 1801B is an independent mapped space rendering node subsystem; 1801C is an independent mapped space camera node subsystem; 1801D is an independent mapped space geometry node subsystem; 1801E is an independent mapped space object layer subsystem; 1801F is an independent mapped space material node subsystem; 1801G is an independent mapped space media node subsystem; 1801H is an independent mapped space texture node subsystem; 1801I is an independent mapped space file settings node subsystem; 1801X is an independent mapped space node graph subsystem; 1801Y is an independent mapped space node pin subsystem; and 1801Z is an independent mapped space node attribute subsystem.

[0161] Figure 19 Illustrates an example 1900 of a metadata tag list that forms a buffer subsystem 1901 for the disclosed IMS metadata framework. For subsystem 1901, the following metadata tags are included: BinaryBlob, BufferSpecification, GLBBuffer, OpenSubDiv buffer, Shading Buffer, Asset Buffer, Accessor, AccessorSparse, AccessorSparseIndices, AccessorSparseValues, and CircularBuffer.

[0162] This subsystem 1901 can be included as a stream node object indicating a logical sequence of data bytes that may be organized into one or more chunks. Subsystem 1901 can direct a processor, importer, or renderer by indicating the organization of binary data into a stream.

[0163] In an embodiment, subsystem 1901 may include one or more of the following parameters. The binaryBlob parameter, which describes a binary data container for storing various types of data such as geometry, animation, textures, and shaders. The bufferSpecification parameter, which describes the organization of the raw data stored in a buffer. This can be part of the local properties of a stream. The GLBBuffer parameter, which describes the binary buffer component of a GL Transmission Format Binary file (GLB). The openSubDiv buffer parameter, which describes the buffer in the OpenSubdiv library used to store and manipulate vertex data during subdivision surface evaluation. The shading buffer parameter, which describes a type of data buffer in computer graphics used to store information about the shading of objects in a scene. The asset buffer parameter, which describes a data structure for storing and managing various types of assets such as geometry, textures, and other resources required to render a 3D scene. The accessor parameter, which describes a data structure that describes the organization and one or more types of data within a buffer such that the content of the buffer can be retrieved (efficiently) according to the accessor. The accessorSparse parameter, which describes a way to optimize the storage and transmission of geometry data by storing only the necessary vertex positions that differ between objects. The accessorSparse parameter can be organized into two parts: sparse indices and sparse values. The accessorSparseIndices can describe the positions and data types of the values to be replaced in the sparse accessor. The accessorSparseValues can describe the values used to replace the default values of the sparse accessor.

[0164] Figure 20 Example 2000 depicts an example of a list of metadata tags that form a scene object subsystem 2001 for the disclosed IMS metadata framework. For subsystem 2001, the following metadata tags are included: ABCScene, FBXScene, glTFScene, USDScene.

[0165] The subsystem 2001 may be included as a scene object node that describes a geometric object that may be animated, was created using digital content creation tools, and is included in a composite scene. As described above, it may represent various geometric assets of a larger scene using Alembic, Universal Scene Description, glTF, and Filmbox.

[0166] Figure 21Illustrates an example 2100 of a metadata tag list that forms an animation subsystem 2101 for the disclosed IMS metadata framework. For this subsystem 2101, the following metadata tags are included: Data Type, Period, Pattern, Animation Type, End Time, Node Target, Input Accessor, Output Accessor, Interpolation, Channel, Animation Settings.

[0167] The subsystem 2101 can be included as an animation node object that describes how an asset will be animated. Animating the asset by the renderer can be guided by the animation parameters of the asset.

[0168] In some embodiments, the subsystem 2101 can include parameters from one or more of the following. A data type parameter that indicates the type of data provided to the animator, such as a string (for a file name), an integer value, a floating-point value. A period parameter that indicates the time pattern of the animation in seconds. An input mode parameter that defines the input time or keyframes of the animation as an array of time samples (e.g., in seconds). An output mode parameter that defines the output time or keyframes of the animation as an array of time samples (e.g., in seconds). An animation type parameter that specifies how to interpret data values when there are more samples defined by the time sampling than there are data values. For example, the animation can be looped, played back and forth, or played only once. An end time parameter that indicates the time at which the animation should stop. A target parameter that is an indicator of the location of the data value to be animated. A property descriptor of the property to be animated (e.g., translation, rotation, scaling, or deformation). An interpolation parameter that provides a description of the type of interpolation to be used for the animation. A shutter alignment parameter that describes how the shutter interval aligns with the current time, such as "before" the current time, "symmetric" with the current time, and "after" the current time. A shutter open time parameter that indicates the amount of time the shutter remains open (as a percentage of the single-frame duration). A subframe start parameter that indicates the minimum start time at which the shutter can be opened without having to reconstruct the geometry (as a percentage of the single-frame duration). A subframe end parameter that indicates the maximum end time at which the shutter can remain open without having to reconstruct the geometry (as a percentage of the single-frame duration). A stacksAvailable parameter that describes the list of animation stacks available to the end user. A stack selection parameter that indicates the animation stack selected by the end user.

[0169] The above disclosure also covers the features mentioned below. These features can be combined in various ways and are not limited to the combinations mentioned below.

[0170] (1) A method for processing immersive media, which is executed by at least one processor. The method includes obtaining a metadata framework to preserve the scene information stored in the scene graph during the process of scene graph conversion. The metadata framework includes multiple subsystems; receiving first scene information to be converted into a second scene graph format in a first scene graph format; using the metadata framework to convert the first scene into the second scene graph format. The multiple subsystems from the metadata framework used include one or more of the following: a first subsystem, including information associated with the geometric assets of the first scene; a second subsystem, including information associated with the animation of one or more assets in the first scene; and a third subsystem, including information associated with the logical sequence of data in the first scene; rendering the first scene in the second scene graph format based on the conversion.

[0171] (2) The method according to feature (1), wherein the first subsystem includes parameters indicating whether the objects in the first scene include one or more of Alembic objects, USD objects, gITF objects, and Filmbox objects.

[0172] (3) The method according to features (1) to (2), wherein the second subsystem includes animation parameters, and these animation parameters include one or more of the following: data type, used to indicate the data type provided to the animation generator or renderer; period, used to indicate the time pattern of the animation; mode, used to specify the input / output time or key frames of the animation in the form of an array of time samples; and end time, used to indicate the time when the animation stops.

[0173] (4) The method according to features (1) to (3), wherein the animation parameters of the second subsystem further include: target, used to indicate the position of the data value of the animation; a type description of the interpolation used in the animation; and animation type, indicating the method of interpreting the data value when the number of time samples is greater than the number of data values.

[0174] (5) The method according to features (1) to (4), wherein the third subsystem includes buffer parameters, and these buffer parameters include one or more of the following: a first binary data container, used to store various types of data; a second binary data container, composed of a GL transmission format binary file; and a third binary data container, used to store vertex data during the evaluation of subdivision surfaces in the OpenSubdiv library.

[0175] (6) The method according to features (1) to (5), wherein the buffer parameters further include: an asset buffer data structure for storing one or more types of assets required for rendering the first scene; an accessor data structure for describing the organization and type of data in the buffer; and an accessor sparse data structure for storing different necessary vertex positions between objects, the accessor sparse data structure including an accessor sparse index parameter and an accessor sparse value parameter.

[0176] (7) An apparatus for processing immersive media, the apparatus including a memory storing program code; and at least one processor configured to read the program code and operate in accordance with the instructions of the program code, the program code being configured to execute the method according to any one of features (1) to (6).

[0177] (8) A non - volatile computer - readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to execute the method according to any one of features (1) to (6).

[0178] Although the present disclosure has described several exemplary embodiments, there are changes, permutations, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, it should be understood that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are thus within the spirit and scope of the present disclosure.

Claims

1. A method for processing immersive media, characterized in that, The method is executed by at least one processor, and the method includes: Obtaining a metadata framework to retain the scene information stored in the scene graph during the conversion of the scene graph, the metadata framework including a plurality of subsystems; Receiving first scene information to be converted into a second scene graph format in a first scene graph format; Using the metadata framework to convert the first scene into the second scene graph format, and the plurality of subsystems from the metadata framework used include one or more of the following: A first subsystem including information associated with the geometric assets of the first scene; A second subsystem including information associated with the animation of one or more assets in the first scene; and A third subsystem including information associated with the logical sequence of data in the first scene; Rendering the first scene in the second scene graph format based on the conversion.

2. The method according to claim 1, wherein The first subsystem includes parameters indicating whether the objects in the first scene include one or more of Alembic objects, USD objects, gITF objects, and Filmbox objects.

3. The method according to claim 1, characterized in that, The second subsystem includes animation parameters, and the animation parameters include one or more of the following: A data type for indicating the data type provided to the animation generator or renderer; A period for indicating the time mode of the animation; A mode for specifying the input / output time or key frames of the animation in the form of an array of time samples; And An end time for indicating the time when the animation stops.

4. The method according to claim 3, characterized in that The animation parameters of the second subsystem further include: A target for indicating the position of the data value of the animation; A type description of the interpolation used in the animation; and An animation type indicating the method of interpreting data values when the number of time samples is greater than the number of data values.

5. The method according to claim 1, characterized in that, The third subsystem includes buffer parameters, and the buffer parameters include one or more of the following: A first binary data container for storing various types of data; A second binary data container composed of a GL transmission format binary file; and A third binary data container for storing vertex data during the subdivision surface evaluation in the OpenSubdiv library.

6. The method according to claim 5, characterized in that, The buffer parameters further include: An asset buffer data structure for storing one or more types of assets required for rendering the first scene; An accessor data structure for describing the organization and type of data in the buffer; and An accessor sparse data structure for storing different necessary vertex positions between objects, and the accessor sparse data structure includes an accessor sparse index parameter and an accessor sparse value parameter.

7. An apparatus for processing immersive media, characterized in that, The device includes: A memory storing program code; and At least one processor configured to read the program code and operate according to the instructions of the program code, and the program code includes: Fetch code, the fetch code being configured to cause the at least one processor to fetch a metadata framework to preserve scene information stored in a scene graph during the conversion of the scene graph, the metadata framework including a first subsystem, the first subsystem including metadata information associated with the animation of one or more assets in a first scene, wherein the first subsystem includes animation parameters, the animation parameters including one or more of the following: Data type, for indicating the data type provided to an animation generator or renderer; Period, for indicating the time pattern of the animation; Mode, for specifying the input / output time or keyframes of the animation in the form of an array of time samples; and End time, for indicating the time at which the animation stops; Receive code, the receive code being configured to cause the at least one processor to receive first scene information to be converted to a second scene graph format in a first scene graph format; Conversion code, the conversion code being configured to cause the at least one processor to convert the first scene into the second scene graph format using the metadata framework; and Render code, the render code being configured to cause the at least one processor to render the first scene in the second scene graph format based on the conversion.

8. The device according to claim 7, characterized in that, The animation parameters further include: Target, for indicating the location of the data value of the animation; A type description of the interpolation used in the animation; and Animation type, indicating the method of interpreting data values when the number of time samples is greater than the number of data values.

9. The device according to claim 7, characterized in that, The metadata framework further includes a second subsystem, wherein the second subsystem includes information associated with the geometric assets of the first scene, and the parameters of the second subsystem indicate whether the objects in the first scene include one or more of Alembic objects, USD objects, gITF objects, and Filmbox objects.

10. The device according to claim 7, characterized in that, The metadata framework further includes a third subsystem, the third subsystem including information associated with the logical sequence of data in the first scene.

11. The device according to claim 10, characterized in that, The third subsystem includes buffer parameters, the buffer parameters including: A first binary data container, for storing various types of data; A second binary data container, composed of a GL transmission format binary file; and A third binary data container, for storing vertex data during the evaluation of subdivision surfaces in the OpenSubdiv library.

12. The device according to claim 11, characterized in that, The buffer parameters further include: An asset buffer data structure, for storing one or more types of assets required for rendering the first scene; An accessor data structure, for describing the organization and type of data in the buffer; and An accessor sparse data structure, for storing different necessary vertex positions between objects, the accessor sparse data structure including an accessor sparse index parameter and an accessor sparse value parameter.

13. A non-volatile computer-readable medium, characterized in that, The non - volatile computer - readable medium stores one or more instructions for processing immersive media, the one or more instructions including: Obtain a metadata framework to preserve the scene information stored in the scene graph during the conversion of the scene graph. The metadata framework includes a first subsystem, and the first subsystem includes information associated with the logical sequence of data in the first scene. The parameters of the first subsystem include one or more of the following: A first binary data container for storing various types of data; A second binary data container composed of a GL transmission format binary file; and A third binary data container for storing vertex data during the evaluation of subdivision surfaces in the OpenSubdiv library; Receive first scene information to be converted to a second scene graph format in a first scene graph format; Use the metadata framework to convert the first scene into the second scene graph format; and Render the first scene in the second scene graph format based on the conversion.

14. The non-volatile computer-readable medium according to claim 13, wherein The parameters of the first subsystem further include: An asset buffer data structure for storing one or more types of assets required to render the first scene; An accessor data structure for describing the organization and type of data in the buffer; and An accessor sparse data structure for storing different necessary vertex positions between objects. The accessor sparse data structure includes an accessor sparse index parameter and an accessor sparse value parameter.

15. The non-volatile computer-readable medium according to claim 13, wherein The metadata framework further includes a second subsystem and a third subsystem. The second subsystem includes information associated with the geometric assets of the first scene, and the third subsystem includes information associated with the animation of one or more assets in the first scene.

16. The non-volatile computer-readable medium according to claim 15, wherein The second subsystem includes parameters indicating whether the objects in the first scene include one or more of Alembic objects, USD objects, gITF objects, and Filmbox objects.

17. The non-volatile computer-readable medium according to claim 15, wherein The third subsystem includes animation parameters, and the animation parameters include one or more of the following: A data type for indicating the data type provided to the animation generator or renderer; A period for indicating the time mode of the animation; A mode for specifying the input / output time or key frames of the animation in the form of an array of time samples; And An end time for indicating the time when the animation stops.

18. The non-volatile computer-readable medium according to claim 17, wherein The animation parameters further include: A target for indicating the position of the data value of the animation; A type description of the interpolation used in the animation; and An animation type indicating the method of interpreting data values when the number of time samples is greater than the number of data values.