Method, apparatus, and storage medium for optimizing media distribution

By introducing immersive media optimization procedures and caching mechanisms into the network, analyzing media asset characteristics and generating metadata, the diversity problem of heterogeneous client devices is solved, the efficiency of immersive media distribution is improved, and resource waste is reduced.

CN116686275BActive Publication Date: 2026-07-31TENCENT AMERICA LLC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT AMERICA LLC
Filing Date
2022-10-25
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In the current technology, an end-to-end ecosystem for distributing immersive media on commercial networks has not yet been realized. The diversity of heterogeneous client devices leads to inefficient media distribution, and an effective method is needed to represent and stream heterogeneous immersive media.

Method used

By introducing an immersive media optimization process into the network, the characteristics of media assets are analyzed, metadata information that uniquely identifies the corresponding media assets is generated, and a caching mechanism is used to store adapted assets, thereby optimizing the media distribution process and reducing redundant conversion and streaming steps.

Benefits of technology

It improves media distribution efficiency, reduces network resource waste, optimizes media conversion and streaming processes, and adapts to the diversity of heterogeneous client devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116686275B_ABST
    Figure CN116686275B_ABST
Patent Text Reader

Abstract

A media distribution scheme for an optimized media stream, executed by at least one processor, is provided, comprising: receiving immersive media data for immersive presentation from a content source; acquiring asset information corresponding to media assets in a scene set included in the immersive media data; analyzing the characteristics of the media assets used in the scene set from the asset information to determine whether the corresponding media asset is unique among the media assets; and generating metadata information that uniquely identifies the corresponding media asset based on the determination that the corresponding media asset is not unique among the media assets in the scene set included in the immersive media data for immersive presentation, for reusing the corresponding media asset in the scene set.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application is based on and claims priority to U.S. Provisional Patent Application No. 63 / 276,540, filed November 5, 2021, and U.S. Patent Application No. 17 / 970,712, filed October 21, 2022, the disclosures of which are incorporated herein by reference in their entirety. Technical Field

[0003] This disclosure generally describes embodiments relating to the architecture, structure, and components of systems and networks for distributing media, including video, audio, geometric (3D) objects, haptic, associated metadata, or other content for client-side rendering devices. Some embodiments relate to systems, structures, and architectures for distributing media content to heterogeneous immersive and interactive client-side rendering devices. Background Technology

[0004] Immersive media generally refers to media that stimulates any or all human sensory systems (e.g., vision, hearing, tactile sensation, smell, and possibly taste) to create or enhance the user's perception of physical presence in the media experience; that is, media that goes beyond timed two-dimensional (2D) video and corresponding audio distributed on existing (e.g., "traditional") commercial networks; such timed media are also referred to as "traditional media." Immersive media can also be defined as media that attempts to create or mimic the physical world through digital simulations of dynamics and physical laws, thereby stimulating any or all human sensory systems so that the user can create a perception of physical presence in a scene depicting a real or virtual world.

[0005] Many devices with immersive media capabilities have been introduced (or are about to appear) in the consumer market, including head-mounted displays, augmented reality glasses, handheld controllers, multi-view displays, haptic gloves, game consoles, holographic displays, and other forms of volumetric displays. Despite these devices, a coherent end-to-end ecosystem for distributing immersive media on commercial networks has yet to be realized.

[0006] One of the obstacles to realizing a coherent end-to-end ecosystem for distributing immersive media over commercial networks in related technologies is the highly diverse range of client devices used as endpoints in such distribution networks for immersive displays. Unlike networks designed solely for distributing traditional media, networks that must support a wide variety of display clients (i.e., heterogeneous clients) require a wealth of information related to the characteristics of each client's capabilities and the format of the media to be distributed before they can employ adaptation processing to convert media into a format suitable for each target display and corresponding application. Such a network needs at least access to information describing the characteristics of each target display and the complexity of the ingested media so that the network can determine how to purposefully adapt the input media source to a format suitable for the target display and application.

[0007] Therefore, there is a need for methods to effectively represent and stream heterogeneous immersive media to different clients. Summary of the Invention

[0008] According to an embodiment, a method is provided for optimizing a system for distributing media assets to at least one set of scenes including immersive media presentation, wherein a separate media asset that is used more than once in at least one set of scenes is distributed only once to a client, the client being equipped with sufficient storage resources or having access to sufficient storage resources to store a copy of the media asset received by the client.

[0009] According to one aspect of this disclosure, a method for optimizing media distribution, executed by at least one processor, is provided. The method includes: receiving immersive media data for immersive presentation from a content source; acquiring asset information corresponding to media assets in a scene set included in the immersive media data; analyzing characteristics of the media assets used in the scene set from the asset information to determine whether the corresponding media asset is unique among the media assets; and generating metadata information uniquely identifying the corresponding media asset for reusing the corresponding media asset in the scene set, based on the determination that the corresponding media asset is not unique among the media assets in the scene set included in the immersive media data for immersive presentation.

[0010] According to another aspect of this disclosure, an apparatus (or device) for optimizing media distribution is provided, comprising at least one memory configured to store computer program code and at least one processor configured to read the computer program code and operate according to the instructions of the computer program code. The computer program code includes receiving code, acquiring code, analyzing code, and generating code. The receiving code is configured to cause at least one processor to receive immersive media data for immersive presentation from a content source. The acquiring code is configured to cause at least one processor to acquire asset information corresponding to media assets in a scene set included in the immersive media data. The analyzing code is configured to cause at least one processor to analyze characteristics of media assets used in the scene set from the asset information to determine whether the corresponding media asset is unique among the media assets. The generating code is configured to cause at least one processor to generate metadata information uniquely identifying the corresponding media asset for reusing the corresponding media asset in the scene set, based on the determination that the corresponding media asset is not unique among the media assets in the scene set included in the immersive media data for immersive presentation.

[0011] According to another aspect of this disclosure, a non-transitory computer-readable medium is provided storing instructions executable by at least one processor of a device for optimizing media distribution. The instructions cause at least one processor to: receive immersive media data for immersive presentation from a content source; acquire asset information corresponding to media assets in a scene set included in the immersive media data; analyze the characteristics of the media assets used in the scene set from the asset information to determine whether the corresponding media asset is unique among the media assets; and, based on the determination that the corresponding media asset is not unique among the media assets in the scene set included in the immersive media data for immersive presentation, generate metadata information uniquely identifying the corresponding media asset for reusing the corresponding media asset in the scene set.

[0012] Additional embodiments will be set forth in the following description, and some embodiments will be apparent from the description and / or may be implemented by practice of the embodiments presented in this disclosure. Attached Figure Description

[0013] Figure 1A This is a schematic diagram of a media stream distributed to a client over a network according to an embodiment; Figure 1B This demonstrates the workflow for reusing logical decision processing; Figure 2A This is a schematic diagram of media streaming over a network according to an embodiment, in which decision processing is employed to determine whether the network should convert the media before distributing it to the client; Figure 2BThis is an illustration of a workflow for media conversion decision processing with asset reuse logic, based on an embodiment. Figure 3 This is a schematic diagram of a data model for the representation and streaming of timed immersive media according to an embodiment; Figure 4 This is a schematic diagram of a data model for representation and streaming of intermittently immersive media according to an embodiment; Figure 5 This is a schematic diagram of the natural media synthesis process according to an embodiment; Figure 6 This is a schematic diagram illustrating an example of a synthetic media ingestion and creation process according to an embodiment; Figure 7 This is a schematic diagram of a computer system according to an embodiment; Figure 8 This is a schematic diagram of a network media distribution system according to an embodiment; Figure 9 This is a schematic diagram of an exemplary workflow for immersive media distribution processing according to an embodiment; Figure 10 This is a system diagram of a media adaptation processing system according to an embodiment; Figure 11 This is a schematic diagram illustrating the creation process of an exemplary distribution format according to an embodiment; Figure 12 This is a schematic diagram of an exemplary grouping process according to an embodiment; Figure 13 This is a sequence diagram illustrating an example of communication flow between components according to an embodiment; Figure 14A This is a workflow illustrating a method for analyzing the uniqueness of objects / assets in a scene using a media reuse analyzer according to an embodiment; Figure 14B This is an example of a list of unique assets used for rendering a scene, according to an embodiment; Figure 15 This is a block diagram of an example of computer code for a media reuse analyzer used to analyze the uniqueness of objects / assets in a scene, according to an embodiment. Detailed Implementation

[0014] The following detailed description of exemplary embodiments is with reference to the accompanying drawings. The same reference numerals in different drawings may identify the same or similar elements.

[0015] The foregoing disclosure provides illustrations and descriptions, but is not intended to be exhaustive or to limit the embodiments to the precise forms disclosed. Modifications and variations are possible based on the foregoing disclosure, or such modifications and variations may be obtained from practice of the embodiments. Furthermore, one or more features or components of one embodiment may be incorporated into or combined with another embodiment (or one or more features of another embodiment). Moreover, in the flowcharts and descriptions of operations provided below, it should be understood that one or more operations may be omitted, one or more operations may be added, one or more operations may be performed simultaneously (at least partially), and the order of one or more operations may be switched.

[0016] It is evident that the systems and / or methods described herein can be implemented in various forms of hardware, software, or combinations of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods is not limiting to specific embodiments. Therefore, the operation and behavior of the systems and / or methods described herein do not refer to any particular software code. It should be understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.

[0017] Even if a particular combination of features is detailed in the claims and / or disclosed in the specification, such combinations are not intended to limit the disclosure of possible implementations. In fact, many of these features can be combined in ways not specifically listed in the claims and / or not disclosed in the specification. Although each dependent claim listed below may depend directly on only one claim, the disclosure of possible implementations includes the combination of each dependent claim with every other claim in the claim set.

[0018] The suggested features discussed below can be used individually or in any combination in any order. Furthermore, embodiments can be implemented using processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-transitory computer-readable medium.

[0019] Unless explicitly stated otherwise, no element, action, or instruction used herein should be construed as critical or necessary. Furthermore, as used herein, the articles “a” and “an” are intended to include one or more items and may be used interchangeably with “one or more.” If only one item is intended to be used, the term “one” or similar wording is used. Additionally, as used herein, the terms “have,” “possess,” “include,” etc., are intended to be open-ended terms. Furthermore, unless explicitly stated otherwise, the word “based on” means “at least partially based on.” Furthermore, expressions such as “at least one of [A] and [B]” or “at least one of [A] or [B]” should be understood to include only A, only B, or both A and B.

[0020] In the implementation of this application, the collection and processing of relevant data should be strictly in accordance with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.

[0021] This disclosure provides an example embodiment of a method and apparatus for analyzing and converting media assets for distribution to presentation devices with immersive media capabilities. A presentation device with immersive media capabilities can refer to one equipped with sufficient resources and capabilities to access, interpret, and present immersive media. Such devices are heterogeneous in terms of the quantity and format of media (provided by the network) they can support. Similarly, the media is heterogeneous in terms of the quantity and type of network resources required for large-scale distribution of such media. "Large-scale" can refer to media distribution by service providers that achieve distribution equivalent to traditional video and audio media distribution via networks (e.g., Netflix, Hulu, Comcast subscriptions, Spectrum subscriptions, etc.). In contrast, traditional presentation devices such as laptops, televisions, and mobile handheld displays are homogeneous in capability because all of these devices consist of rectangular displays that use 2D rectangular video or still images as their primary visual media format. Some visual media formats commonly used in traditional presentation devices can include, for example, High Efficiency Video Coding / H.265, Advanced Video Coding / H.264, and Universal Video Coding / H.266.

[0022] As mentioned above, client devices that serve as endpoints for distributing immersive media over a network are highly diverse. Some client devices support certain immersive media formats, while others do not. Some client devices are capable of creating immersive experiences based on traditional raster-based formats, while others are not. To address this issue, the distribution of any media over a network can employ a media delivery system and architecture that reformats the media from the input or network-ingested media format into a distribution media format that is not only suitable for ingestion by the target client device and its applications but also conducive to streaming over the network. Therefore, there can be two processes performed by the network using the ingested media: 1) converting the media from format A to format B suitable for ingestion by the target client device, i.e., ingesting a specific media format based on the client device's capabilities, and 2) preparing the media for streaming.

[0023] In embodiments, streaming media broadly refers to the segmentation and / or grouping of media such that it can be transmitted over a network in logically organized and ordered blocks of smaller sizes, based on one or both of the media's temporal or spatial structure. Converting media from format A (sometimes called "transcoding") to format B can be a process typically performed by a network or service provider before distributing the media to client devices. This transcoding may include converting media from format A to format B based on prior knowledge that format B is the preferred or unique format, which may be ingested by the target client device or is more suitable for distribution on limited resources such as commercial networks. In many, but not all, steps of converting media and preparing media for streaming are necessary before the target client device can receive and process the media from the network.

[0024] Media conversion (or transformation) and media preparation for streaming are steps in the processing of ingested media by the network before it is distributed to client devices. The result of this processing (i.e., conversion and preparation for streaming) is a media format known as the distribution media format, or simply the distribution format. If all these steps are performed for a given media data object, these steps should only be performed once. However, if the network has access to information instructing the client to use the media object requiring conversion and / or streaming for multiple scenarios, such media conversion and streaming will be triggered multiple times. That is, the processing and transmission of data for media conversion and streaming are generally considered sources of latency, requiring potentially significant network and / or computational resources. Therefore, a network designed to be unable to access information indicating when a client may have stored a particular media data object in its cache or relative to the client's local storage will be suboptimal compared to a network that does have access to such information.

[0025] An ideal network supporting heterogeneous clients should leverage the fact that assets adapted from input media formats to a specific target format can be reused across a set of similar display targets. That is, once assets are converted to a format suitable for the target display, they can be reused across multiple such displays with similar adaptation requirements. Therefore, according to embodiments, such an ideal network can employ a caching mechanism to store adapted assets in an area similar to, for example, a Content Distribution Network (CDN) used in traditional networks.

[0026] Immersive media can be organized into scenes described by scene graphs, also known as scene descriptions. In embodiments, a scene (in the context of computer graphics) is a collection of objects (e.g., 3D assets), object attributes, and other metadata, including visual, acoustic, and physical features describing a particular setting whose interaction with respect to objects within that setting is spatially or temporally constrained. A scene graph is a general data structure commonly used in vector-based graphics editing applications and modern computer games, representing the logical and, in many cases, (but not necessarily) spatial representation of a graphical scene. A scene graph can consist of a collection of nodes and vertices in a graphical structure. Nodes can include information related to a logical, spatial, or temporal representation of visual, auditory, tactile, olfactory, gustatory, or related processing information. Each node should have at most one output edge, zero or more input edges, and at least one edge (input or output) connected to it. Attributes or object attributes refer to metadata associated with a node that describes a particular characteristic or feature of that node in a canonical or more complex form (e.g., in relation to another node). A scene graph describes visual, audio, and other forms of immersive assets that include specific settings as part of a presentation, such as events and actors occurring in a specific location within a building that is part of a presentation (e.g., a film). A list of all scenes that include a single presentation can be formulated as a list of scenes.

[0027] An additional benefit of using a caching mechanism to store adapted assets is that a bill of materials can be created for content that must be prepared before distribution. This bill of materials identifies all assets that will be used throughout the rendering, and the frequency with which each asset is used in various scenarios within the rendering. Ideally, the network should be aware of the existence of cached resources that can be used to meet the asset needs of a particular rendering. Similarly, a client device rendering a series of scenarios may want to know the frequency with which any given asset will be used across multiple scenarios. For example, if a media asset (also called a media object) is referenced multiple times in multiple scenarios being or to be processed by a client device, the client device should avoid discarding that asset from its cached resources until the last scenario in which that particular asset was needed has been rendered. In embodiments of this disclosure, the terms media “object” and media “asset” are used interchangeably, both referring to a specific instance of media data in a specific format.

[0028] For traditional media presentation devices, the distribution format can be equivalent to or fully equivalent to the "presentation format" that the client presentation device ultimately uses to create the presentation. That is, the presentation media format is the media format whose attributes (e.g., resolution, frame rate, bit depth, color gamut, etc.) are closely adapted to the capabilities of the client presentation device. An example of a distribution relative to a presentation format includes a high-definition (HD) video signal (1920 pixels x 1080 pixels) distributed over a network to an Ultra-high-definition (UHD) client device with a resolution of (3840 pixels in columns x 2160 pixels in rows). In the aforementioned example, the UHD client device would apply super-resolution processing to the HD distribution format to upscale the video signal from HD to UHD. Therefore, the final signal format presented by the client device is the "presentation format," which in this example is the UHD signal, while the HD signal includes the distribution format. In this example, the HD signal distribution format is very similar to the UHD signal presentation format because both signals are linear video formats, and the process of converting HD to UHD is relatively simple and easy to perform on most traditional media client devices.

[0029] In some embodiments, the preferred presentation format of the client device may differ significantly from the ingested format received by the network. However, the client device has access to sufficient computing, storage, and bandwidth resources to convert the media from the ingested format to the necessary presentation format suitable for presentation by the client device. In this case, the network can bypass the step of reformatting or transcoding the ingested media from format A to format B simply because the client device has sufficient resources to perform all media conversions, which the network need not do beforehand. However, the network can still perform the steps of segmenting and packaging the ingested media so that the media can be streamed over the network to the client device.

[0030] In some embodiments, the ingested media may differ significantly from the client's preferred rendering format, and the client device may not have access to sufficient computational, storage, and / or bandwidth resources to convert the media from the ingested format to the preferred rendering format. In such cases, the network can assist the client by performing some or all of the conversions on behalf of the client device from the ingested format to a format that is equivalent to or nearly equivalent to the client's preferred rendering format. In some architectural designs, this assistance provided by the network on behalf of the client device is often referred to as segmented rendering.

[0031] Figure 1A This is a schematic diagram of media stream processing 100 distributed to a client over a network according to some embodiments. Figure 1A An exemplary processing of media in format A (hereinafter referred to as "ingesting media format A") is illustrated. This processing (i.e., media stream processing 100) can be performed or implemented by a network cloud or edge device (hereinafter referred to as "network device 104") and distributed to a client, such as client device 108. In some embodiments, the same processing can be performed manually or by the client device. Network device 104 may include an ingesting media module 101, a network processing module 102 (hereinafter referred to as "processing module 102"), and a distribution module 103. Client device 108 may include a rendering module 106 and a presentation module 107.

[0032] First, network device 104 receives ingested media from a content provider or similar source. Ingestion media module 101 obtains the ingested media stored in ingestion media format A. Network processing module 102 performs any necessary transformations or adjustments on the ingested media to create potential alternative representations of the media. That is, processing module 102 creates a distribution format for media objects in the ingested media. As described above, the distribution format is a media format that can be distributed to clients by formatting the media into distribution format B. Distribution format B is a format prepared for streaming to client device 108. Processing module 102 may include optimization reuse logic 121 to perform decision processing to determine whether a specific media object has been streamed to client device 108. (See reference...) Figure 1BThe operation of processing module 102 and optimization reuse logic 121 is described in detail.

[0033] Media formats A and B may or may not be representations of the same syntax that follow a specific media format specification; however, format B may be limited to a scheme that facilitates media distribution via a network protocol. The network protocol may be, for example, a connection-oriented protocol (TCP) or a connectionless protocol (UDP). Distribution module 103 streams the streamable media (i.e., media format B) from network device 104 to client device 108 via network connection 105.

[0034] Client device 108 receives the distribution media and optionally prepares the media for presentation via rendering module 106. Rendering module 106 may have access to rendering capabilities, depending on the target client device 108, which may be basic or, similarly, advanced. Rendering module 106 creates the presentation media in presentation format C. Presentation format C may or may not be represented according to a third format specification. Therefore, presentation format C may be the same as or different from media formats A and / or B. Rendering module 106 outputs presentation format C to presentation module 107, which can then present the presentation media on the display (or similar device) of client device 108.

[0035] Embodiments of this disclosure facilitate decision processing adopted by the network and / or client to determine whether the network should convert some or all ingested media from format A to format B, thereby further facilitating the client's ability to produce media presentation in a potential third format C. To aid such decision processing, embodiments describe an immersive media optimization procedure for reusing scene assets. The immersive media optimization procedure introduced herein serves as a mechanism that analyzes one or more media objects across a set of scenes constituting the presentation, including a portion or the entire immersive media scene. The immersive media optimization procedure creates informational metadata associated with the uniqueness of each object in each scene constituting the presentation. The metadata may contain information that uniquely identifies the object based on its attributes. The immersive media optimization procedure further determines the number of times a unique object is used in at least one scene. This immersive media optimization procedure can be implemented independently or in conjunction with an immersive media analyzer that determines the complexity of each object in each scene. Therefore, once all such metadata about some or all parts of an immersive media scene is available, decision processing is better equipped with information indicating how frequently a particular object is used for full presentation in the scene, thus allowing the decision to skip the transformation and / or streaming of the original media if the client has previously received a copy of such media from the network.

[0036] The embodiments address the need for mechanisms or processes to analyze immersive media scenarios to obtain sufficient information to support decision-making processes that, when adopted by a network or client, provide indications as to whether the conversion of media objects from format A to format B should be performed entirely by the network, entirely by the client, or via a combination of both (and indications of which assets should be converted by the client or network). This immersive media data complexity analyzer can be used by the client or network in an automated context, or manually by a person, such as by an operating system or device.

[0037] According to embodiments, the process of adapting an input immersive media source to a specific endpoint client device can be the same as or similar to the process of adapting the same input immersive media source to a specific application running on a specific client endpoint device. Therefore, the problem of adapting an input media source to the characteristics of an endpoint device has the same complexity as the problem of adapting a specific input media source to the characteristics of a specific application.

[0038] Figure 1B This is the workflow of processing module 102 according to the embodiment. More specifically, Figure 1B The workflow is a reuse logic decision process executed by the optimization reuse logic 121 to help the decision process determine whether a specific media object has been streamed to the client device 108.

[0039] In S101, media creation processing is performed. Therefore, reuse logic decision processing also begins. In S102, conditional logic is executed to determine whether the current media object has previously been streamed to client device 108. Optimization reuse logic 121 can access a list of unique assets used for presentation to determine whether the media object has previously been streamed to the client. If the current media object has previously been streamed, processing proceeds to S103. In S103, an indicator (hereinafter referred to as a "proxy") is created to identify that the client has received the current media object and should use a local copy of the media object. Subsequently, processing of the current media object ends (S104).

[0040] If it is determined that the media object has not been streamed before, the process proceeds to S105. In S105, the processing module 102 continues to prepare the media object for conversion and / or distribution, and creates the distribution format for the media object. Subsequently, the process proceeds to S104, ending the processing of the current media object.

[0041] Figure 2A It is a logical workflow used to process ingested media over a network. Figure 2AThe illustrated workflow describes a media conversion decision process 200 according to an embodiment. The media conversion decision process 200 determines whether the network should convert media before distributing it to client devices. The media conversion decision process 200 can be processed manually or automatically within the network.

[0042] The ingested media, represented in format A, is provided to the network by the content provider. In S201, the network ingests the media from the content provider. Then, in S202, if the attributes of the target client are not yet known, the attributes of the target client are obtained. These attributes describe the processing capabilities of the target client.

[0043] In S203, it is determined whether the network (or client) should assist in converting the ingested media. Specifically, before the media is streamed to the target client, it is determined whether any format conversion is required for any media assets contained within the ingested media (e.g., conversion of one or more media objects from format A to format B). The decision processing at S203 can be performed manually (i.e., by a device operator, etc.) or can be automated. The decision processing at S203 can be based on the determination that the media can be streamed in its original ingestion format A, or whether it must be converted to a different format B to facilitate the client's presentation of the media. Such a decision may require access to information describing aspects or characteristics of the ingested media in a way that helps the decision processing make the best choice (i.e., determining whether the ingested media needs to be converted before being streamed to the client, or whether the media should be streamed directly to the client in its original ingestion format A).

[0044] If it is determined that the network (or client) should assist in the conversion of any media assets (the judgment result at S203 is "yes"), then process 200 proceeds to S204.

[0045] In S204, the ingested media is converted from format A to format B, resulting in converted media 205. The converted media 205 is output, and processing proceeds to S206. In S206, the input media undergoes preparation processing for streaming to the client. In this case, the converted media 205 (i.e., the input media) is prepared for streaming.

[0046] The conversion of media from format A to another format (e.g., format B) can be done entirely by the network, entirely by the client, or jointly by both the network and the client. For split rendering, it is clear that a dictionary of attributes describing the media format may be needed so that both the client and the network have complete information to characterize the work that must be done. Furthermore, a dictionary of attributes providing client capabilities may also be needed, such as attributes regarding available computational resources, available storage resources, and access to bandwidth. Additionally, a mechanism is needed to characterize the level of computational, storage, or bandwidth complexity of ingesting the media format so that the network and the client can jointly or individually determine whether or when the network can employ split rendering to distribute the media to the client.

[0047] If it is determined that the network (or client) should not (or does not need to) assist in the conversion of any media assets (the judgment at S203 is negative), then process 200 proceeds to S206. In S206, media for streaming is prepared. In this case, data ingested for streaming (i.e., media in its original form) is prepared.

[0048] Finally, once the media data is in a streamable format, it will be streamed to the client as prepared in S206 (S207). In some embodiments, (as referenced) Figure 1B If the client has already completed the transformation and / or streaming of the specific media object required or to be presented as part of the work done on the previous scene being presented, the network may skip the transformation and / or streaming of the ingested media (i.e., S204-S207) and assume that the client still has access to or availability of the media previously streamed to the client.

[0049] Figure 2B This is a media conversion decision process with asset reuse logic 210 according to an embodiment. Similar to media conversion decision process 200, the media conversion decision process with asset reuse logic 210 ingests media from the network and employs two decision processes to determine whether the network should convert the media before distributing it to the client.

[0050] The ingested media, represented by format A, is provided to the network by the content provider. Next, steps S221-222 are executed, and S221-222 are similar to S201-S202 (e.g., ...). Figure 2A (As shown). In S221, the network ingests media from the content provider. Then, in S202, if the attributes of the target client are not yet known, the attributes of the target client are obtained. These attributes describe the processing capabilities of the target client.

[0051] If it is determined that the network has previously streamed a specific media object or the current media object (the result of the determination at S223 is yes), the process proceeds to S224. In S224, a proxy is created to replace the previously streamed media object, instructing the client to use a local copy of the previously streamed object.

[0052] If it is determined that the network has not previously streamed a specific media object (the result of the determination at S223 is no), the process proceeds to S225. In S225, it is determined whether the network or the client should perform any format conversion on any media assets contained within the ingested media in S221. For example, the conversion could include converting a specific media object from format A to format B before streaming the media to the client. The processing performed in S225 can be similar to... Figure 2A The process shown is executed in S203.

[0053] If it is determined that the media assets should be converted by the network (the result of the judgment in S225 is yes), the process proceeds to S226. In S226, the media object is converted from format A to format B. Then, preparations are made to stream the converted media to the client (S227).

[0054] If it is determined that the media assets should not be converted by the network (the result of the judgment at S225 is no), the process proceeds to S227. In S227, preparation is then made to stream the media object to the client.

[0055] Finally, once the media data is in a streamable format, the media prepared in S227 will be streamed to the client (S228). Figure 2B The processing performed at S225-S228 can be similar to that at Figure 2A The processes performed at points S203-S207 shown.

[0056] The streaming format of the media can be a timed or untimed heterogeneous immersive media. Figure 3 An example of a timed media representation 300 in a streamable format for heterogeneous immersive media is shown. Timed immersive media can include a set of N scenes. Timed media is media content ordered by time, for example, having a start time and an end time based on a specific clock. Figure 4 An example of an opaque media representation 400 in a streamable format for heterogeneous immersive media is shown. Opaque media is media content organized by spatial, logical, or temporal relationships (e.g., in an interactive experience implemented based on actions taken by one or more users).

[0057] Figure 3 This refers to timing scenarios used for timed media, while Figure 4This refers to untimed scenarios used for untimed media. Timed and untimed scenarios can be represented by various scenario representations or descriptions. Figure 3 and Figure 4 Both employ a single, exemplary included media format that has been adapted from the source media format to match the specific client endpoint. That is, the included media format is a distribution format that can be streamed to the client device. The included media format is structurally robust enough to accommodate a large number of different media attributes, where each attribute can be layered based on the significant amount of information that each layer contributes to media presentation.

[0058] like Figure 3 As shown, the timed media representation 300 includes a list of timed scenes 301. Each timed scene 301 points to a list of components 302, each component 302 individually describing the processing information and media asset types that constitute the timed scene 301. This includes, for example, an asset list and other processing information. The list of components 302 may point to proxy assets 308 corresponding to the asset type (e.g., such as...). Figure 3 The list of visual and audio assets shown is for proxy assets. The list of visual assets points to proxy asset 310 corresponding to the asset type (e.g., such as...). Figure 3 (The proxy visual and audio assets shown). Component 302 points to a list of unique assets that have not been used in other scenes before. For example, Figure 3 The diagram shows a list of unique assets 307 for (timing) scenario 1. Component 302 points to asset 303, which in turn points to a base layer 304 and an attribute enhancement layer 305. The base layer is the nominal representation of the asset, which can be configured to minimize computational resources, the time required to render the asset, and / or the time required to transmit the asset over a network. The enhancement layer can be a set of information that, when applied to the base layer representation of the asset, enhances the base layer to include features or capabilities that may not be supported in the base layer.

[0059] like Figure 4 As shown, the timed media representation 400 includes information about scene 401. Scene 401 is not associated with a start time / duration and an end time / duration (based on clocks, timers, etc.). Scene 401 points to a list of components 402, each individually describing the processing information and media asset types that constitute scene 401. Components 402 point to visual assets, audio assets, haptic assets, and timing assets (collectively referred to as assets 403). Assets 403 also point to the base layer 404 and attribute enhancement layers 405 and 406. Scene 401 can also point to other timed scenes for timed media assets (i.e., in...). Figure 4 This is referred to as the untimed scenarios (2.1-2.4) and / or for timed media scenarios (i.e., in...). Figure 4 Scene 407 (referred to as Timed Scene 3.0) in [the context of the game]. Figure 4 In the example, the timed immersive media contains a set of five scenes (including timed scenes and timed scenes). The list of unique assets 408 identifies unique assets associated with a specific scene that have not been used before in higher-order (e.g., parent) scenes. Figure 4 The list of unique assets shown in 408 includes unique assets for the unpredictable scenario 2.3.

[0060] The media streamed according to the included media format is not limited to traditional visual and audio media. The included media formats can include any type of media information capable of generating signals that interact with machines to stimulate human vision, hearing, taste, touch, and smell. For example... Figures 3-4 As shown, the media streamed, depending on the included media format, can be timed media, non-timed media, or a mixture of both. By using a base layer and enhancement layer architecture to implement a hierarchical representation of media objects, the included media formats are streamable.

[0061] In some embodiments, separate base and enhancement layers are computed by applying multi-resolution or multi-mosaic analysis techniques to media objects in each scene. This computational technique is not limited to raster-based visual formats.

[0062] In some embodiments, the progressive representation of a geometric object can be a multi-resolution representation of the object computed using wavelet analysis techniques.

[0063] In some embodiments of a hierarchical representation media format, enhancement layers can apply different properties to a base layer. For example, one or more enhancement layers can refine the material properties of the surface of a visual object represented by the base layer.

[0064] In some embodiments, in a layered representation media format, attributes can refine the texture of the surface of an object represented by a base layer by, for example, changing the surface from a smooth texture to a porous texture, or from an entangled surface to a smooth surface.

[0065] In some embodiments, in a layered representation media format, the surface of one or more visual objects in a scene can be changed from a Lambertian surface to a ray-traceable surface.

[0066] In some embodiments, in a layered representation media format, the network can distribute the base layer representation to the client, allowing the client to create a nominal representation of the scene while waiting for the transmission of additional enhancement layers to refine the resolution or other features of the base layer.

[0067] In embodiments, the resolution of attribute information or refinement information in the enhancement layer is not explicitly coupled to the resolution of objects in the base layer. Furthermore, the included media formats can support any type of information media that can be rendered or driven by a rendering device or rendering machine, thereby enabling heterogeneous media formats to support heterogeneous client endpoints. In some embodiments, the network distributing the media format will first query the client endpoint to determine the client's capabilities. Based on the query, if the client cannot purposefully ingest the media representation, the network can remove attribute layers that the client does not support. In some embodiments, if the client cannot purposefully ingest the media representation, the network can modify the media from its current format to a format suitable for the client endpoint. For example, the network can adapt the media by using a network-based media processing protocol to convert volumetric visual media assets into a 2D representation of the same visual asset. In some embodiments, the network can adapt the media by employing neural network (NN) processing to reformat the media into an appropriate format, or optionally synthesizing the view required by the client endpoint.

[0068] A complete (or partially complete) immersive experience's scene inventory (replay of live streaming events, games, or on-demand assets) is organized by scenes containing the minimum amount of information needed for rendering and ingestion to create the presentation. The scene inventory includes a list of individual scenes to be rendered for the entire immersive experience requested by the client. Associated with each scene is one or more representations of the geometric objects within the scene, representing a streamable version of the corresponding scene geometry. One embodiment of a scene may refer to a low-resolution version of the scene's geometry. Another embodiment of the same scene may refer to an enhancement layer used for the low-resolution representation of the scene to add additional detail or tessellation to the geometry of the same scene. As described above, each scene may have one or more enhancement layers to progressively increase the detail of the scene's geometry. Each layer of media objects referenced within a scene may be associated with a token (e.g., a uniform resource identifier (URI)) that points to an address on which a resource can be accessed within the network. This resource is similar to a content delivery network (CDN), where content can be retrieved by a client. The token used to represent the geometry may point to a location within the network or to a location on the client. In other words, a client can send a signaling message to the network to indicate that its resources are available for network-based media processing.

[0069] According to embodiments, a scene (timed or untimed) can be represented by a scene graph as a multiplanar image (MPI) or a multispherical image (MSI). Both MPI and MSI techniques are examples of technologies that help create display-agnostic scene representations for natural content (i.e., images of the real world captured simultaneously from one or more cameras). Scene graph techniques, on the other hand, can be used to represent natural and computer-generated images in the form of synthetic representations. However, for cases where content is captured as a natural scene by one or more cameras, computationally intensive methods are particularly needed to create such a representation. Creating a scene graph representation of naturally captured content is both time-consuming and computationally intensive, requiring complex analysis of the natural images using photogrammetry or deep learning techniques, or both, to create a synthetic representation that can then be used to interpolate a sufficient number of views to populate the target immersive client display's visual volume. As a result, such synthetic representations are impractical as candidates for representing natural content because they cannot actually be created in real-time, considering use cases requiring real-time distribution. Therefore, the best representation of a computer-generated image is to use a scene graph with a composite model, since computer-generated images are created using 3D modeling processes and tools, and using a scene graph with a composite model yields the best representation of a computer-generated image.

[0070] Figure 5 An example of a natural media compositing process 500 according to an embodiment is shown. The natural media compositing process 500 converts the ingestion format from a natural scene into a representation of the ingestion format that can be used as a network serving heterogeneous client endpoints. To the left of the dashed line 510 is the content capture portion of the natural media compositing process 500. To the right of the dashed line 510 is the ingestion format compositing of the natural media compositing process 500 (for natural images).

[0071] like Figure 5 As shown, the first camera 501 uses a single camera lens to capture, for example, a person (i.e., Figure 5 The scene depicts the actors shown. The second camera 502 captures the scene with five divergent fields of view by mounting five camera lenses around the circular object. Figure 5 The arrangement of the second camera 502 shown is an exemplary arrangement typically used for capturing omnidirectional content for VR applications. The third camera 503 captures a scene with seven converging fields of view by mounting seven camera lenses on the inner diameter portion of the sphere. The arrangement of the third camera 503 is an exemplary arrangement typically used for capturing light fields or light fields in holographic immersive displays. The embodiments are not limited to this. Figure 5 The configuration shown. The second camera 502 and the third camera 503 may include multiple camera lenses.

[0072] Natural image content 509 is output from the first camera 501, the second camera 502, and the third camera 503 and used as input to the synthesizer 504. The synthesizer 504 can use a set of training images 506 to train an neural network (NN) 505 to produce a capture NN model 508. The training images 506 can be predefined or stored based on previous synthesis processing. The NN model (e.g., the capture NN model 508) is a set of parameters and tensors (e.g., matrices) that define weights (i.e., numerical values) used in well-defined mathematical operations applied to the visual signal to achieve an improved visual output that may include interpolations of new views of visual signals not explicitly provided by the original signal.

[0073] In some embodiments, photogrammetric processing can be implemented instead of NN training 505. If the captured NN model 508 is created during natural media compositing processing 500, then the captured NN model 508 becomes one of the assets in the ingestion format 507 of the natural media content. The ingestion format 507 can be, for example, MPI or MSI. The ingestion format 507 may also include media assets.

[0074] Figure 6 An example of a composite media ingestion creation process 600 according to an embodiment is shown. The composite media ingestion creation process 600 creates ingestion media formats for composite media such as computer-generated images.

[0075] like Figure 6 As shown, camera 601 can capture point cloud 602 of the scene. Camera 601 can be, for example, a LiDAR camera. Computer 603 uses, for example, common gateway interface (CGI) tools, 3D modeling tools, or another animation process to create composite content (i.e., a representation of the composite scene in an ingestion format that can be used as a network serving heterogeneous client endpoints). Computer 603 can create CGI assets 604 over the network. Additionally, sensor 605A can be worn on actor 605 in the scene. Sensor 605A can be, for example, a motion capture suit with attached sensors. Sensor 605A captures digital records of the actor 605's movements to generate animation motion data 606 (or MoCap data). Provided with data from point cloud 602, CGI assets 604, and motion data 606 as input to compositor 607, compositor 607 creates a composite media ingestion format 608. In some embodiments, compositor 607 can use a neural network (NN) and training data to create an NN model to generate the composite media ingestion format 608.

[0076] Both natural and computer-generated (i.e., synthetic) content can be stored in containers. Containers can include serialization formats to store and exchange information representing all natural scenes, all synthetic scenes, or a mixture of synthetic and natural scenes, including scene graphs and all media resources required to render the scenes. Content serialization involves translating data structures or object states into a format that can be stored (e.g., in a file or memory buffer) or transmitted (e.g., via a network link) and subsequently reconstructed in the same or different computing environments. When the resulting sequence of bits is reread according to the serialization format, that sequence of bits can be used to create a semantically identical clone of the original object.

[0077] The dichotomy between the best representations of natural and computer-generated (i.e., synthetic) content suggests that the optimal ingestion format for naturally captured content differs from the optimal ingestion format for computer-generated content or for natural content that is not necessary for real-time distribution applications. Therefore, according to embodiments, the goal of the network is to be robust enough to support multiple ingestion formats for visually immersive media, whether they are naturally created using, for example, physical cameras or computer-generated.

[0078] Technologies such as OTOY's ORBX, Pixar's Universal Scene Description (USD), and the Graphics Language Transmission Format 2.0 (glTF2.0) specification written by the Khronos 3D group represent scene graphics in formats suitable for representing visually immersive media created using computer-generated techniques, or naturally captured content that uses deep learning or photogrammetry techniques to create corresponding synthetic representations of natural scenes (i.e., not required for real-time distribution applications).

[0079] OTOY's ORBX is one of several scene graph technologies capable of supporting any type of visual media, timed or non-timed, including ray-traced, traditional (frame-based), volumetric, and other types of compositing or vector-based visual formats. ORBX differs from other scene graphs because it provides native support for free and / or open-source formats of meshes, point clouds, and textures. ORBX is a carefully designed scene graph intended to facilitate the exchange between multiple vendor technologies operating on scene graphs. Furthermore, ORBX offers a rich material system, support for the Open Shader Language, robust camera systems, and support for Lua scripting. ORBX is also the foundation for immersive technology media formats released under a royalty-free license by the Immersive Digital Experiences Alliance (IDEA). In the context of real-time media distribution, the ability to create and distribute ORBX representations of natural scenes is a function of the availability of computational resources to perform complex analysis of camera-captured data and to composite the same data into a synthetic representation.

[0080] Pixar's USD is a scene graph widely used for visual effects and professional content creation. USD is integrated into Nvidia's Omniverse platform, a toolset for developers to create and render 3D models using Nvidia's graphics processing units (GPUs). A subset of USD released by Apple and Pixar is called USDZ, powered by Apple's ARKit.

[0081] glTF 2.0 is a version of a graphics language transport format specification written by the Khronos 3D group. This format supports simple scene graph formats, typically supporting static (timeless) objects in the scene, including PNG and JPEG image formats. glTF 2.0 supports simple animations, including translation, rotation, and scaling of basic shapes described using glTF primitives (i.e., for geometric objects). glTF 2.0 does not support timed media, therefore it does not support video or audio media input.

[0082] These scene representations for immersive visual media are provided as examples only and do not limit the ability of the disclosed subject matter to adapt input immersive media sources to a format suitable for the specific characteristics of client endpoint devices. Furthermore, any or all of the aforementioned example media representations employ or can employ deep learning techniques to train and create neural network models that enable or facilitate the selection of specific views to populate a particular display's view volume based on a specific size of the truncated volume. The view selected for a particular display's view volume can be interpolated from existing views explicitly provided in the scene representation, such as interpolation based on MSI or MPI techniques. Views can also be rendered directly from the rendering engine based on the specific virtual camera position, filters, or virtual camera descriptions of these rendering engines.

[0083] The methods and apparatus disclosed herein are robust enough to take into account the existence of a relatively small but well-known set of immersive media ingestion formats that are sufficient to meet the requirements for real-time or on-demand (e.g., non-real-time) distribution of media captured naturally (e.g., using one or more cameras) or created using computer-generated techniques.

[0084] The advancements in networking technologies (e.g., 5G for mobile networks) and the deployment of fiber optic cables for fixed networks have further facilitated the interpolation of views from immersive media ingestion formats using neural network models or network-based rendering engines. These advanced networking technologies increase the capacity and capability of commercial networks, as this advanced network infrastructure can support the transmission and delivery of ever-increasing volumes of visual information. Network infrastructure management technologies such as Multi-access Edge Computing (MEC), Software Defined Networking (SDN), and Network Functions Virtualization (NFV) enable commercial network service providers to flexibly configure their network infrastructure to adapt to changing demands for certain network resources, such as responding to dynamic increases or decreases in demand for network throughput, network speed, round-trip latency, and computing resources. Furthermore, this inherent ability to adapt to dynamic network demands also contributes to the network's ability to adapt immersive media ingestion formats to suitable distribution formats, supporting a wide range of immersive media applications with potentially heterogeneous visual media formats for heterogeneous client endpoints.

[0085] Immersive media applications may have varying requirements for network resources, including gaming applications that require significantly lower network latency to respond to real-time updates in the game state, telepresence applications with symmetrical throughput requirements for both the uplink and downlink portions of the network, and passive viewing applications where the type of client endpoint display that consumes data may increase the demand for downlink resources. Generally, any consumer-facing application can be supported by a variety of client endpoints with various onboard client capabilities for storage, computation, and driving, as well as varying requirements for the specific media representation.

[0086] Therefore, embodiments of this disclosure enable a fully equipped network, i.e., a network employing some or all of the features of a modern network, to simultaneously support multiple legacy devices and immersive media-capable devices based on specified features within the device. Thus, the immersive media distribution methods and processes described herein offer flexibility in utilizing media ingestion formats applicable to both real-time and on-demand media distribution use cases, flexibility in supporting naturally generated and computer-generated content for both legacy clients and client endpoints with immersive media capabilities, and support for both timed and untimed media. The methods and processes also dynamically adapt the source media ingestion format to a suitable distribution format based on the characteristics and capabilities of the client endpoints and the requirements of the application. This ensures that the distribution format is streamable over an IP-based network and enables the network to simultaneously serve multiple heterogeneous client endpoints, which may include legacy devices and devices with immersive media capabilities. Furthermore, embodiments provide exemplary media representation frameworks that facilitate the organization of distributed media along scene boundaries.

[0087] According to embodiments of this disclosure, an end-to-end implementation of the improved heterogeneous immersive media distribution described above is provided, which is based on... Figures 3 to 15 The processing and component implementations described in the detailed description will be further elaborated below.

[0088] The techniques described above for representing and streaming heterogeneous immersive media can be implemented in both the source and destination as computer software using computer-readable instructions and physically stored in one or more non-transitory computer-readable media, or implemented by one or more hardware processors with a specific configuration. Figure 7 A computer system 700 suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0089] Computer software can be coded using any suitable machine code or computer language. This machine code or computer language can be assembled, compiled, linked, or similarly used to create code containing instructions that can be executed directly by the computer's central processing unit (CPU), graphics processing unit (GPU), or through interpretation, microcode execution, etc.

[0090] These instructions can be executed on various types of computers or their components, including, for example, personal computers, tablets, servers, smartphones, gaming devices, and Internet of Things (IoT) devices.

[0091] Figure 7 The components of the computer system 700 shown are exemplary in nature and are not intended to suggest any limitation on the scope or functionality of computer software implementing embodiments of this disclosure. The configuration of the components should also not be construed as having any dependencies or requirements relating to any component or combination of components shown in the exemplary embodiments of the computer system 700.

[0092] Computer system 700 may include certain human-computer interface (HCI) input devices. Such HCI input devices may respond to input from one or more HCI users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), and olfactory input. HCI devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographs obtained from still-image cameras), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0093] The input human-machine interface device may include one or more of the following: keyboard 701, touchpad 702, mouse 703, and screen 709 (only one of each is shown), and may be, for example, a touch screen, data glove, joystick 704, microphone 705, camera 706, and scanner 707.

[0094] Computer system 700 may also include certain human-machine interface (HMI) output devices. Such HMI output devices can stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. These HMI output devices may include tactile output devices (e.g., tactile feedback from screen 709, data gloves, or joystick 704, but tactile feedback devices not used as input devices may also exist), audio output devices (e.g., speakers 708, headphones), visual output devices (e.g., screen 709, including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touchscreen input capability, each with or without tactile feedback capability—some of which may be able to output two-dimensional or more than three-dimensional visual output via devices such as stereoscopic output devices; virtual reality glasses, holographic displays, and smoketanks), and printers.

[0095] The computer system 700 may also include human-accessible storage devices and associated media, such as optical media including CD / DVD ROM / RW 711 with CD / DVD or similar media 710, thumb drives 712, removable hard disk drives or solid-state drives 713, conventional magnetic media such as magnetic tapes and floppy disks, devices based on dedicated ROM / ASIC / PLD such as security dongles, etc.

[0096] Those skilled in the art will also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not include transmission media, carrier waves, or other transient signals.

[0097] Computer system 700 may also include an interface 715 connected to one or more communication networks 714. Network 714 may be, for example, wireless, wired, or optical. Network 714 may also be local, wide area, urban, vehicular, and industrial, real-time, latency-tolerant, etc. Examples of network 714 include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., wired or wireless wide area digital TV networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicular and industrial networks including CANBus, etc. Some networks 714 typically require an external network interface adapter (e.g., graphics adapter 725) that connects to some general-purpose data port or peripheral bus 716 (e.g., a USB port of computer system 700; others, as described below, are typically integrated into the core of computer system 700 via connection to a system bus (e.g., an Ethernet interface connected to a PC computer system or a cellular network interface connected to a smartphone computer system)). Using any of these networks 714, computer system 700 can communicate with other entities. This communication can be unidirectional, receive-only (e.g., broadcast television), send-only (e.g., to a CANbus device), or bidirectional, such as to other computer systems using local or wide area digital networks. As mentioned above, certain protocols and protocol stacks can be used on each of these networks and network interfaces.

[0098] The aforementioned human-machine interface device, human-accessible storage device, and network interface can be attached to the core 717 of the computer system 700.

[0099] The core 717 may include one or more central processing units (CPUs) 718, graphics processing units (GPUs) 719, dedicated programmable processing units in the form of field-programmable gate arrays (FPGAs) 720, task-specific hardware accelerators 721, and so on. These devices, along with read-only memory (ROM) 723, random-access memory (RAM) 724, and internal mass storage 722 such as internal non-user-accessible hard disk drives (SDs), can be connected via a system bus 726. In some computer systems, the system bus 726 may be accessed as one or more physical connectors to allow for expansion by adding CPUs, GPUs, etc. Peripheral devices may be directly connected to the core's system bus 726 or connected via a peripheral bus 716. Peripheral bus architectures include PCI, USB, etc.

[0100] The CPU 718, GPU 719, FPGA 720, and accelerator 721 can execute certain instructions, which, when combined, constitute the aforementioned machine code (or computer code). This computer code can be stored in ROM 723 or RAM 724. Transient data can also be stored in RAM 724, while permanent data can be stored, for example, in internal mass storage 722. Fast storage and retrieval of any memory device can be achieved by using a cache memory, which can be closely associated with one or more CPUs 718, GPUs 719, mass storage 722, ROM 723, RAM 724, etc.

[0101] Computer-readable media may contain computer code for performing operations of various computer implementations. For the purposes of this disclosure, the media and computer code may be specially designed and constructed, or they may be of a type well known and available to those skilled in the art of computer software.

[0102] By way of example and not limitation, a computer system having a computer system 700 architecture, particularly a core 717, can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such a computer-readable medium can be a medium associated with a user-accessible mass storage as described above, and some memory of the core 717 having a non-transitory nature, such as a core-internal mass storage 722 or ROM 723. Software implementing various embodiments of this disclosure can be stored in such a device and executed by the core 717. Depending on specific needs, the computer-readable medium may include one or more memory devices or chips. The software can cause the core 717 and specifically cause the processor therein (including a CPU, GPU, FPGA, etc.) to execute a particular process or a particular portion of a particular process described herein, including defining data structures stored in RAM 724 and modifying such data structures according to software-defined processing. Additionally or alternatively, the computer system can provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator 721), which can replace or operate with the software to execute a particular process or a particular portion of a particular process described herein. Where appropriate, references to software may include logic, and vice versa. Where appropriate, references to computer-readable media may include circuitry (e.g., integrated circuits, ICs) storing software for execution, circuitry including logic for execution, or both. This disclosure includes any suitable combination of hardware and software.

[0103] Provided Figure 7This is an example of the number and arrangement of components. In practice, input human-machine interface devices may include... Figure 7 The components shown may be additional components, fewer components, different components, or components with different arrangements compared to the components shown. Additionally or optionally, the component set (e.g., one or more components) of the input human-machine interface device may perform one or more functions described as being performed by another component set of the input human-machine interface device.

[0104] In an embodiment, Figures 1A to 6 and Figures 8 to 15 Any of the operations or processes can be accessed or used Figure 7 Implemented by any of the components shown.

[0105] Figure 8 An exemplary network media distribution system 800 serving multiple heterogeneous client endpoints is illustrated. Specifically, system 800 supports various conventional displays and displays with heterogeneous immersive media capabilities as client endpoints. System 800 may include a content acquisition module 801, a content presentation module 802, and a transmission module 803.

[0106] Content acquisition module 801 uses, for example Figure 6 and / or Figure 5The embodiments described herein are used to capture or create source media. Content rendering module 802 creates an ingestion format, which is then transmitted to a network media distribution system using transmission module 803. Gateway 804 can serve a customer premises device to provide network access to various client endpoints of the network. Set-top box 805 can also be used as a customer premises device to provide network service providers with access to aggregated content. Wireless demodulator 806 can be used as a mobile network access point for mobile devices, such as a mobile handheld display 813. In a particular embodiment of system 800, a conventional 2D television 807 is shown as directly connected to one of gateway 804, set-top box 805, or WiFi (router) 808. Laptop 2D display 809 (i.e., a computer or laptop with a conventional 2D display) is shown as a client endpoint connected to WiFi (router) 808. Head-mounted 2D (raster-based) display 810 is also connected to WiFi (router) 808. Lenticule light field display 811 is shown as one connected to gateway 804. A lenticular light field display 811 may include one or more GPUs 811A, a storage device 811B, and a visual presentation component 811C that creates multiple views using light-based lenticular optics. A holographic display 812 is shown connected to a set-top box 805. The holographic display 812 may include one or more CPUs 812A, GPUs 812B, a storage device 812C, and a visualization component 812D. The visualization component 812D may be a Fresnel-mode, wave-based holographic device / display. An augmented reality (AR) headset 814 is shown connected to a wireless demodulator 806. The AR headset 814 may include a GPU 814A, a storage device 814B, a battery 814C, and a volumetric visual presentation component 814D. A dense light field display 815 is shown connected to a WiFi (router) 808. The dense light field display 815 may include one or more GPUs 815A, CPUs 815B, storage devices 815C, eye-tracking devices 815D, cameras 815E, and dense light field panels 815F.

[0107] Provided Figure 8 The number and arrangement of components shown are for illustrative purposes only. In practice, system 800 may include components related to... Figure 8 The components shown may be additional components, fewer components, different components, or components arranged differently compared to the components shown. Alternatively, the component set of system 800 (e.g., one or more components) may perform one or more functions described as being performed by another component set of the device or a corresponding display.

[0108] Figure 9An exemplary workflow for an immersive media distribution process 900 is shown, which is capable of processing previously... Figure 8 The text describes services provided by conventional displays and displays with heterogeneous immersive media capabilities. Immersive media distribution processing 900 performed by the network can provide adaptation information about specific media represented in a media ingestion format, for example, prior to the network's processing of adapting media for consumption via a specific immersive media client endpoint.

[0109] The immersive media distribution process 900 can be divided into two parts: immersive media production to the left of the dashed line 912 and immersive media network distribution to the right of the dashed line 912. Immersive media production and immersive media network distribution can be performed by network or client devices.

[0110] First, media content 901 is created or retrieved either by a network (or client device) or from a content source. The methods used to create or retrieve data are, for example, embodied in the following ways for natural and synthetic content respectively: Figure 5 and Figure 6 Then, the web ingestion format creation process 902 converts the created content 901 into the ingestion format. The web ingestion creation process 902 also addresses both natural and synthetic content in... Figure 5 and Figure 6 The ingestion format can also be updated to store information about assets that may be reused across multiple scenarios, such as those from the Media Reuse Analyzer 911 (see later). Figure 10 and Figure 14A (Details omitted). The ingested format is transmitted to the network and stored in the ingested media storage 903 (i.e., storage device). In some embodiments, the storage device may be located within the network of the immersive media content creator and remotely accessible to the immersive media network distribution 920. Client-specific and application-specific information, client-specific information 904, may optionally be available in the remote storage device. In some embodiments, client-specific information 904 may reside remotely in an alternative cloud network and may be transmitted to the network.

[0111] Then, the network coordinator 905 is executed. The network coordinator serves as the primary information source and repository for performing the network's main tasks. The network coordinator 905 can be implemented in a uniform format with other network components. The network coordinator 905 can further process client devices using a bidirectional messaging protocol to facilitate all media processing and distribution based on the characteristics of the client devices. Furthermore, the bidirectional protocol can be implemented across different transport channels (e.g., control plane channels and / or data plane channels).

[0112] like Figure 9As shown, network coordinator 905 receives information about the characteristics and attributes of client device 908. Network coordinator 905 collects information about the needs of applications currently running on client device 908. This information can be obtained from client-specific information 904. In some embodiments, this information can be obtained by directly querying client device 908. When directly querying the client device, it is assumed that a bidirectional protocol exists and is operational, allowing client device 908 to communicate directly with network coordinator 905.

[0113] The network coordinator 905 can also initiate the media adaptation and segmentation module 910 (in Figure 10 (As described in the text) and communicate with it. When the media adaptation and segmentation module 910 adapts and segments the ingested media, the media can be transmitted to an intermediate media storage device, such as media 909 prepared for distribution. When media is prepared for distribution and stored in the storage device of media 909 prepared for distribution, the network coordinator 905 enables the client device 908 to receive the distribution media and descriptive information 906 by "push" request, or the client device 908 can initiate a "pull" request for the distribution media and descriptive information 906 from the stored media 909 prepared for distribution. This information can be "pushed" or "pulled" through the network interface 908B of the client device 908. The distribution media and descriptive information 906 "pushed" or "pulled" can be the descriptive information of the corresponding distribution media.

[0114] In some embodiments, the network coordinator 905 uses a bidirectional messaging interface to execute "push" requests or to initiate "pull" requests by the client device 908. The client device 908 may optionally employ a GPU 908C (or a CPU).

[0115] The distribution media format is then stored in a storage device or storage cache 908D included in the client device 908. Finally, the client device 908 visualizes the media via a visualization component 908A.

[0116] Throughout the process of streaming immersive media to client device 908, network coordinator 905 determines the status of client progress via client progress and status feedback channel 907. In some embodiments, the determination of status can be performed through a bidirectional communication message interface.

[0117] Figure 10 An example of media adaptation processing 1000 performed by, for example, media adaptation and segmentation module 910 is shown. By performing media adaptation processing 1000, the ingested source media can be appropriately adapted to match the requirements of the client (e.g., client device 908).

[0118] like Figure 10As shown, the media adaptation process 1000 includes multiple components that help adapt the ingested media to a suitable distribution format for the client device 908. Figure 10 The components shown should be considered exemplary. In practice, the media adaptation process 1000 may include components related to... Figure 10 The components shown may be additional components, fewer components, different components, or components with different arrangements compared to the components shown. Additionally, or optionally, the set of components (e.g., one or more components) of the media adapter processing 1000 may perform one or more functions described as being performed by another set of components.

[0119] exist Figure 10 In this process, the adaptation module 1001 receives input network status 1005 to obtain the current traffic load on the network. As described above, the adaptation module 1001 also receives information from the network coordinator 905. This information may include attribute and characteristic descriptions of the client device 908, application characteristics and descriptions, the current state of the application, and the client NN model (if available) to help map the geometry of the client frustum to the interpolation capability for ingesting immersive media. This information can be obtained through a bidirectional messaging interface. The adaptation module 1001 ensures that the output of the adaptation is stored in a storage device for storing the client adaptation media 1006 when it is created.

[0120] The media reuse analyzer 911 can be an optional process, which can be executed preferentially or as part of a network automation process for distributing media. The media reuse analyzer 911 can store the ingested media format and assets in a storage device (1002). The ingested media format and assets can then be transferred from the storage device (1002) to the adaptation module 1001.

[0121] The adaptation module 1001 can be controlled by a logic controller 1001F. The adaptation module 1001 can also employ a renderer 1001B or a processor 1001C to adapt a specific ingested source media to a format suitable for the client. The processor 1001C can be a neural network-based processor. The processor 1001C uses a neural network model 1001A. Examples of such a processor 1001C include the Deepview Neural Network Model Generator described in MPI and MSI. If the media is in a 2D format, but the client must use a 3D format, the processor 1001C can invoke processing to derive a volumetric representation of the scene depicted in the media using highly correlated images from the 2D video signal.

[0122] A renderer 1001B can be a software-based (or hardware-based) application or process, based on a selective blend of disciplines related to acoustic physics, optical physics, visual perception, audio perception, mathematics, and software development, that, given an input scene graph and asset container, emits visual and / or audio signals suitable for rendering on a target device or conforming to the desired properties specified by the attributes of the rendered target node in the scene graph. For visual-based media assets, the renderer may emit visual signals suitable for target display or storage as intermediate assets (e.g., repackaging into another container and using them in a series of rendering processes in the graphics pipeline). For audio-based media assets, the renderer may emit audio signals for rendering in multi-channel speakers and / or dual-audio headphones, or for repackaging into another (output) container. Renderers include, for example, real-time rendering features of source and cross-platform game engines. Renderers may include scripting languages ​​(i.e., interpreted programming languages) that can be executed by the renderer at runtime to handle dynamic input and variable state changes to scene graph nodes. Dynamic inputs and variable state changes can affect the rendering and evaluation of spatial and temporal object topology (including physics forces, constraints, inverse kinematics, deformation, and collisions) as well as energy propagation and transmission (light, sound). The evaluation of spatial and temporal object topology produces results that transform the output from an abstract result into a concrete result (e.g., similar to the evaluation of a document object model for a webpage).

[0123] Renderer 1001B may be a modified version of, for example, the OTOY Octane renderer, which will be modified to interact directly with adapter module 1001. In some embodiments, renderer 1001B implements computer graphics methods (e.g., path tracing) for rendering 3D scenes, such that the scene's lighting is realistic. In some embodiments, renderer 1001B may employ shaders (i.e., a computer program originally used for shading (producing appropriate levels of light, dark, and color within an image), but now performing various specialized functions in various areas such as computer graphics special effects, video post-processing unrelated to shading, and other functions unrelated to graphics).

[0124] The adaptation module 1001 can perform compression and decompression of media content using a media compressor 1001D and a media decompressor 1001E, respectively, according to the compression and decompression requirements based on the format of the ingested media and the format required by the client device 908. The media compressor 1001D can be a media encoder, and the media decompressor 1001E can be a media decoder. After performing compression and decompression (if necessary), the adaptation module 1001 outputs adapted media 1006, best suited for streaming or distribution to the client, to the client device 908. The client-side adapted media 1006 can be stored in a storage device for storing the adapted media.

[0125] Figure 11 An exemplary distribution format creation process 1100 is shown. For example... Figure 11 As shown, the distribution format creation process 1100 includes an adaptation media grouping module 1103, which packages and stores the media output from the media adaptation process 1000 as client-adapted media 1006. The media grouping module 1103 formats the adaptation media from the client-adapted media 1006 into a robust distribution format 1104. The distribution format can be, for example... Figure 3 or Figure 4 The exemplary format shown is illustrated. Information list 1104A can provide client device 908 with a list 1104B of scene data assets. The list 1104B of scene data assets may also include complexity metadata describing the complexity of all assets in the list 1104B. The list 1104B of scene data assets depicts a list of visual assets, audio assets, and haptic assets, each with its corresponding metadata. Media can be further grouped before streaming. Figure 12 An exemplary packet processing 1200 is illustrated. The packet system 1200 includes a packetizer 1202. The packetizer 1202 can receive a list of scene data assets 1104B as input media 1201 (such as...). Figure 12 (As shown). In some embodiments, client-adapted media 1006 or distribution format 1104 is input to packetizer 1202. Packetizer 1202 separates the input media 1201 into individual packets 1203 suitable for representation and streaming to client device 908 on the network.

[0126] Figure 13 This is a sequence diagram illustrating examples of data and communication flows between components according to an embodiment. Figure 13 The sequence diagram is a network that adapts a specific immersive media in an ingested format to a streamable and suitable distribution format for a specific immersive media client endpoint. The data and communication flows are described below.

[0127] Client device 908 initiates a media request 1308 to network coordinator 905. In some embodiments, the request may be made to the network distribution interface of the client device. Media request 1308 includes information identifying the media requested by client device 908. The media request may be identified by, for example, a uniform resource name (URN) or another standard term. Network coordinator 905 then responds to media request 1308 with a profile request 1309. Profile request 1309 requests the client to provide information about currently available resources (including compute, storage, battery charge percentage, and other information characterizing the client's current operating state). Profile request 1309 also requests the client to provide one or more NN models, which, if available at the client endpoint, the network can use to perform NN inference to extract or interpolate the correct media view to match the characteristics of the client's presentation system.

[0128] Then, client device 908 follows up with a response 1310 from client device 908 to network coordinator 905, which is provided as a client token, an application token, and one or more NN model tokens (if such NN model tokens are available on the client endpoint). Network coordinator 905 then provides a session ID token 1311 to the client device. Network coordinator 905 then requests media ingestion 1312 from ingest media server 1303. Ingest media server 1303 may include, for example, ingest media storage 903 or ingest media format and asset storage device 1002. The request for media ingestion 1312 may also include the URN or other standard name of the media identified in request 1308. Ingest media server 1303 responds to the media ingestion 1312 request with a response 1313 including an ingest media token. Network coordinator 905 then provides the media token from response 1313 to client device 908 in call 1314. Then, the network coordinator 905 initiates adaptation processing for the media requested in request 1315 by providing the adaptation and segmentation module 910 with an ingest media token, a client token, an application token, and an NN model token. The adaptation and segmentation module 910 requests access to the ingested media by providing the ingest media token to the ingest media server 1303 in request 1316 to request access to the ingested media asset.

[0129] In response 1317 to the adaptation and segmentation module 910, the ingesting media server 1303 responds to request 1316 with an ingesting media access token. The adaptation and segmentation module 910 then requests the media adaptation processing 1000 to adapt the ingested media located at the ingesting media access token for the client, application, and NN inference model corresponding to the session ID token created and sent at response 1313. The adaptation and segmentation module 910 issues request 1318 to the media adaptation processing 1000. Request 1318 contains the required token and session ID. The media adaptation processing 1000 provides the adapted media access token and session ID to the network coordinator 905 in an update response 1319. The network coordinator 905 then provides the adapted media access token and session ID to the media packetization module 1103 in an interface call 1320. The media packetization module 1103 provides response 1321 to the network coordinator 905, which contains the packet's media access token and session ID. Then, in response 1322, media packet module 1103 provides packet assets, URN, and packet media access tokens for session ID to packet media server 1307 for storage. Subsequently, client device 908 executes request 1323 to packet media server 1307 to initiate streaming of the media assets corresponding to the packet media access token received in response 1321. Finally, client device 908 executes other requests and provides a status update to network coordinator 905 in message 1324.

[0130] Figure 14A It shows in Figure 9 The workflow of the media reuse analyzer 911 is shown. The media reuse analyzer 911 analyzes metadata related to the uniqueness of objects included in the media data.

[0131] In S1401, media data is obtained from, for example, a content provider or content source. In S1402, initialization is performed. Specifically, the iterator "i" is initialized to zero. The iterator can be, for example, a counter. A set 1420 of unique asset lists for each scene (e.g.) Figure 14B (As shown) is also initialized, and each scene identifier is a unique asset encountered in all scenes, including those being rendered (such as...). Figure 3 and / or Figure 4 (As shown).

[0132] In S1403, a judgment process is performed to determine whether the value of iterator "i" is less than the total number N of scenes to be presented. If the value of iterator "i" is equal to (or greater than) the number N of scenes to be presented (the judgment result at S1403 is negative), the process proceeds to S1404, and the reuse analysis terminates (i.e., the process ends). If the value of iterator "i" is less than the number N of scenes to be presented (the judgment result at S1403 is positive), the process proceeds to S1405. In S1405, the value of iterator "j" is set to zero.

[0133] Subsequently, in S1406, a decision process is performed to determine whether the value of iterator "j" is less than the total number of media assets X (also known as media objects) in the current scene. If the value of iterator "j" is equal to (or greater than) the scene... s The total number of media assets X (the result of the judgment at S1406 is no), then the process proceeds to S1407, where the iterator "i" is incremented by 1 before returning to S1403. If the value of iterator "j" is less than the scene... s The total number of media assets X (the result of the judgment at S1406 is yes), and then the processing proceeds to S1408.

[0134] In S1408, the characteristics of the media assets are compared with those previously obtained from the current scenario (i.e., scenario). s The assets in the previous scenario analysis are compared to determine whether the current media asset has been used previously.

[0135] If the current media asset has already been identified as a unique asset (the result of the judgment at S1408 is no), that is, the current media asset has not been analyzed in the scenario associated with the smaller value of iterator "i" before, then the process proceeds to S1409. In S1409, in the corresponding current scenario (i.e., scenario...) s A unique asset entry is created in set 1420 of the unique asset list. A unique identifier is also assigned to the unique asset entry, and the number of times the asset is used in scenarios 0 to N-1 is set to 1. Then, the process proceeds to S1411.

[0136] If the current media assets have been identified as being in the scene s The assets used in one or more previous scenarios (whose determination result at S1408 is yes) are then processed, proceeding to S1409. In S1410, the assets corresponding to the current scenario (i.e., scenario...) are processed... s In the set 1420 of the unique asset list, the number of times the current media asset has been used in scenarios 0 to N-1 is increased by 1. Then, the process proceeds to S1411.

[0137] In S1411, the value of iterator "j" is incremented by 1. Then, the process returns to S1406.

[0138] In some embodiments, the media reuse analyzer 911 may also signal to a client, such as client device 108, that the client should use a copy of the asset for each instance that uses the asset in the scene set (after the asset is first distributed to the client).

[0139] Note, reference Figure 13 and Figure 14A The steps in the described sequence diagrams and workflows are not intended to limit the configuration of data and communication flows in the embodiments. For example, one or more steps may be performed simultaneously, and data may be stored and / or flow to... Figures 13 to 14A The direction is not explicitly shown in the process.

[0140] Figure 14B This is an example of a set 1420 of unique asset lists initialized in S1402 (and possibly updated in S1409-S1410) for all scenes once the presentation is complete, according to an embodiment. The unique asset lists in set 1420 can be identified or predefined a priori by the network or client device. Set 1420 of unique asset lists shows a sample list of information entries describing assets that are unique relative to the entire presentation, including indicators for the type of media including the asset (e.g., grid, audio, or volume), the asset's unique identifier, and the number of times the asset is used in the scene set including the entire presentation. For example, for scene N-1, its list does not include any assets because all assets required for scene N-1 have been identified as assets also used in scenes 1 and 2.

[0141] Figure 15 This is a block diagram illustrating an example of computer code 1500 for optimizing media distribution according to an embodiment. In the embodiment, the computer code may be, for example, program code or computer program code. According to embodiments of this disclosure, an apparatus / device may be provided including at least one processor and a memory storing computer program code. The computer program code may be configured to perform any number of aspects of this disclosure when executed by the at least one processor.

[0142] like Figure 15 As shown, computer code 1500 may include receiving code 1510, acquiring code 1520, analyzing code 1530, and generating code 1540.

[0143] Receive code 1510 is configured to cause at least one processor to receive immersive media data from the content source for immersive presentation.

[0144] Acquisition code 1520 is configured to cause at least one processor to acquire code, said code being configured to cause at least one processor to acquire asset information corresponding to media assets included in the scene set in the immersive media data.

[0145] Analysis code 1530 is configured to cause at least one processor to analyze the characteristics of the media assets used in the scene set from the asset information to determine whether the corresponding media asset is unique among the media assets.

[0146] The generation code 1540 is configured to cause at least one processor to generate metadata information that uniquely identifies the corresponding media asset based on determining that the corresponding media asset is not unique among the media assets in the scene set included in the immersive media data for immersive presentation, for reuse of the corresponding media asset in the scene set.

[0147] Although Figure 15 Example blocks of code are shown; in some embodiments, the device / apparatus may include... Figure 15 The blocks shown can be different additional blocks, fewer blocks, different blocks, or blocks with different arrangements. Alternatively, two or more blocks of a device / apparatus can be combined. In other words, although... Figure 15 Different code blocks are shown; different code instructions do not have to be different and can be mixed together.

[0148] While several exemplary embodiments have been described in this disclosure, modifications, substitutions, and various equivalent alternatives fall within the scope of this disclosure. Therefore, it should be understood that those skilled in the art will be able to design numerous systems and methods that, while not expressly shown or described herein, embody the principles of this disclosure and thus fall within its spirit and scope.

Claims

1. A method for optimizing media distribution, characterized in that, The method includes: Receive immersive media data from the content source for immersive presentation; Obtain asset information corresponding to media assets included in the scene set of the immersive media data; The characteristics of the media assets used in the scene set are analyzed from the asset information to determine whether the corresponding media asset is unique among the media assets; When the corresponding media asset is identified as a unique media asset, a unique asset entry is created in the unique asset list associated with the scene in the scene set; a unique identifier is assigned to the unique asset entry, and a counter is created to record the number of times the unique media asset is used in the scene set; Based on the determination that the corresponding media asset is not unique among the media assets in the scene set included in the immersive media data used for the immersive presentation, metadata information is generated to uniquely identify the corresponding media asset for reuse in the scene set, and after the corresponding media asset is first used in the scene set, signaling is sent to the client device to use a copy of the corresponding media asset.

2. The method according to claim 1, characterized in that, The signaling instructs the client device to use a copy of the corresponding media asset for each instance after the first instance of using the corresponding media asset in the scene set.

3. The method of claim 2, wherein, The method further includes: Upon receiving the immersive media data for the immersive presentation, initialize a list of unique assets associated with each scene included in the scene set; and The counter is incremented for each instance of a unique media asset identified in the immersive presentation.

4. The method according to claim 1, characterized in that, The asset information includes the basic representation of the corresponding media asset and an asset enhancement layer set, wherein the asset enhancement layer set includes attribute information corresponding to the characteristics of the media asset, and Specifically, when the asset enhancement layer set is applied to the underlying representation of the corresponding media asset, the underlying representation of the corresponding media asset is enhanced to include features not supported in the underlying layer, which contains the underlying representation of the corresponding media asset.

5. The method according to claim 4, characterized in that, The method further includes: identifying two or more scenarios that share the same media asset, and reusing the same media asset in at least one of the two or more scenarios.

6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: determining, based on the complexity of the scene, whether to convert the format of the immersive media data corresponding to the scene in the scene set to a second format before distributing it to the client device.

7. The method according to any one of claims 1 to 5, characterized in that, The method further includes determining whether the corresponding media asset is unique among the media assets previously used in the scene set.

8. A device for optimizing media distribution, characterized in that, The device includes: At least one memory is configured to store computer program code; and At least one processor is configured to read the computer program code and operate according to the instructions of the computer program code to: Receive immersive media data from the content source for immersive presentation; Obtain asset information corresponding to media assets included in the scene set of the immersive media data; The characteristics of the media assets used in the scene set are analyzed from the asset information to determine whether the corresponding media asset is unique among the media assets; When the corresponding media asset is identified as a unique media asset, a unique asset entry is created in the unique asset list associated with the scene in the scene set; a unique identifier is assigned to the unique asset entry, and a counter is created to record the number of times the unique media asset is used in the scene set; Based on the determination that the corresponding media asset is not unique among the media assets in the scene set included in the immersive media data used for the immersive presentation, metadata information is generated to uniquely identify the corresponding media asset for reuse in the scene set, and after the corresponding media asset is first used in the scene set, signaling is sent to the client device to use a copy of the corresponding media asset.

9. The device according to claim 8, characterized in that, The signaling instructs the client device to use a copy of the corresponding media asset for each instance after the first instance of using the corresponding media asset in the scene set.

10. The device according to claim 9, characterized in that, The at least one processor is further configured to read the computer program code and operate in accordance with the instructions of the computer program code to: Upon receiving the immersive media data, initialize a list of unique assets associated with each scene included in the scene set; as well as The counter is incremented for each instance of a unique media asset identified in the immersive presentation.

11. The device according to claim 8, characterized in that, The asset information includes the basic representation of the corresponding media asset and an asset enhancement layer set, wherein the asset enhancement layer set includes attribute information corresponding to the characteristics of the media asset, and Specifically, when the asset enhancement layer set is applied to the underlying representation of the corresponding media asset, the underlying representation of the corresponding media asset is enhanced to include features not supported in the underlying layer, which contains the underlying representation of the corresponding media asset.

12. The device according to claim 8, characterized in that, The at least one processor is further configured to read the computer program code and operate in accordance with the instructions of the computer program code to: Identify two or more scenarios that share the same media asset, and reuse the same media asset in at least one of the two or more scenarios.

13. The device according to any one of claims 8 to 12, characterized in that, The at least one processor is further configured to read the computer program code and operate in accordance with the instructions of the computer program code to: Based on the complexity of the scene, it is determined whether to convert the format of the immersive media data corresponding to the scene in the scene set to the second format before distributing it to the client device.

14. The device according to any one of claims 8 to 12, characterized in that, The at least one processor is also configured to read the computer program code and operate in accordance with the instructions of the computer program code to determine whether the corresponding media asset is unique among media assets previously used in the scene set.

15. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor of a device for optimizing media distribution, cause the at least one processor to perform the method according to any one of claims 1 to 7.