Method and apparatus for object based media
Patent Information
- Application Number
- GB2023019707
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-20
- Publication Date
- 2025-07-09
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUND OF THE INVENTION This invention relates to delivery of media to user devices and in particularly to delivery at large scale, namely to large numbers of devices. Video streams and broadcasts are becoming increasingly complex in order to provide the highest quality user experience. For instance, broadcasters may aim to maximise the viewing quality of streams by adding elements such as graphic overlays, widgets, and alerts to convey additional or supplementary information to a user. From a consumer viewpoint, the ability to view visually appealing and informative images while maintaining a constant frame-rate is widely regarded as a signature of high quality streams. In order to successfully render and process video streams, with increased complexity, natively on a streaming device, for example as result of a number of additional stream elements, a larger amount of compute resources need to be allocated by the streaming device accordingly. Video streaming can occur on a number of user devices such as televisions, set top boxes, smartphones and other devices capable of video stream playback. However, the compute resources of user devices used for video stream playback may vary significantly. For instance, a television will generally have a differing amount of compute resources to a desktop computer. Similarly, a user streaming video to a new smartphone will likely have greater compute resources available than another user with an older smartphone. Streaming devices with differing amounts of compute resources can correspondingly successfully render and process video streams of differing levels of complexity. In order to maintain a baseline Quality of Experience (QoE), the stream complexity may therefore be adjusted such that it is capable of being successfully rendered and processed on devices with small compute resources. However, a broadcaster may also wish to maximise the quality and therefore complexity of a video stream on a device-by-device basis, particularly if a device has more available computational resources that would not be exhausted by rendering the baseline stream quality. Dynamic Adaptive Streaming over HTTP (DASH), also sometimes referred to as MPEG-DASH, is an established adaptive streaming protocol that allows the bitrate of a video stream to be adjusted based on network bandwidth in order to allow continuous viewing of the video stream. As a result, DASH solves a QoE problem, namely providing a minimum level of streaming for users with poor network bandwidth, while also maximising the QoE for users with superior network bandwidth by scaling up the bitrate of a video stream. Instructions detailing the parameters and information for streaming are provided in manifest files to the end user. For DASH these are usually Media Presentation Description (MPD) files. However, as media experiences using video streams become increasingly complex with streams becoming more complex to manage and re-version and larger rendering demands, providing flexible stream quality also becomes an issue of compute resources. In this sense, DASH fails to provide a comprehensive solution, and only allows for description of different resolutions of the same experience and does not sufficiently address the increasing complexity of media experiences such as including or excluding parts of a stream or interactive elements. Furthermore, DASH operates by referring to segments of content that do not exist in real time. A broadcaster may be streaming multimedia content that is intended for consumption in real time, for example an interactive experience around a live event. In particular, a state of temporal coherency is desired between users such that content may be viewed as close to simultaneously as possible to ensure an even QoE between users. As a result of the generation of content segments in DASH (by use of HTTP ranging requests), there can be a delay in viewing of content on the order of 10 seconds, thereby inducing a weak state of temporal coherency between users. Segments may be referred to in a Content Delivery Network (CDN). Existing solutions do not provide sufficient flexibility to properly address object-based media requirements. We have considered existing solutions to the problem of maintaining a flexible QoE in instances where complex multimedia content must be rendered and streamed. Pixel streaming involves real-time rendering of graphical data through a cloud-based server, and the subsequent transmission of this data (often as a static image or video stream) to a user device. In doing so, a user is granted the ability to adjust their multimedia experience without consideration of the native compute resources of the device. This is particularly advantageous in situations where a user device does not have sufficient compute resources to natively render a desired multimedia experience. As a result, pixel streaming has found particular use in gaming where a user can adjust game settings without consideration of the abilities of their device. In this manner, pixel streaming bypasses the issue of varying levels of available compute resources on multimedia playback devices. However, pixel streaming operates by effectively making a certain amount of compute resources in the cloud available to each user on a user-by-user basis. It therefore scales poorly as the number of users making requests increases and computation requirements comparatively grow. It may also be rare that two users would request identical multimedia content configurations such that computation results can be reused. While this can be of lesser concern in the gaming industry where the overall number of active users at a given time is relatively consistent, the same cannot be said for broadcasting (such as television) where the number of viewers can vary dramatically depending on a broadcasting schedule. For instance, a football match may have on the order of 10-million viewers while typical broadcast content may only be viewed by 1 million viewers at a time. There is therefore a desire to scale the required computational resources for determining user experiences sub-linearly. This is particularly the case as the number of permutations of content to be rendered also increases. Consequently, this becomes not just an issue of QoE for the end user, but one of the overall Quality of Service (QoS). Alternative solutions attempt to address the scaling issue, such as geometry streaming. However, in this methodology rendering is passed to the user device in order to reduce server load. This therefore returns to the issue of varying available compute resources as a result of the multiplicity of user devices. There exist methods of optimising rendering processes to reduce overhead for a given runtime in a graphics processing unit (GPU). One such method is the generation of so-called “render command lists”. A part of a scene may be rendered on a first thread while recording on a second thread to create a command list. Said command list may then be played back on the first thread to efficiently recreate a part of a scene. In doing so, complex rendering tasks can be scaled across threads and cores within a GPU. Render command lists are typically generated from higher-level scene descriptions such as scene graphs or Material Exchange Format (MXF) files. SUMMARY OF THE INVENTION We have appreciated that existing arrangements do not adequately allow scaling of delivery of object-based media to heterogenous devices for large numbers of users. In broad terms, the invention provides a method of delivering object based media, comprising: providing computational tasks as compute fragments; providing media as resource fragments; providing a manifest defining how to assert valid combinations of compute fragments and resource fragments; and using the manifest to retrieve the compute and resource fragments to present media. The separation of object-based media into compute fragments and resource fragments and provision of the manifest describing how to validly assert combinations of these, allows client devices to retrieve media and servers and nodes within a network to determine where best to execute computational tasks, and where best to assemble compute and resource fragments in various ways as will be described. We use “assert” in its normal sense to cover a variety of ways in which the client device may consume and use fragments, including executing compute fragments, displaying resource fragments comprising images, providing an output of audio for audio fragments and any other normal way in which devices may use, execute, display or otherwise assert retrieved fragments. BRIEF DESCRIPTION OF THE DRAWINGS The invention will be described in more detail by way of example with reference to the accompanying drawings, in which: Fig. 1 is a graphical representation of the equivalence of resource fragments and compute fragments; Fig. 2 is a graphical representation showing how a compute fragment may be used to generate an output resource fragment; Fig. 3 is a graphical representation showing the interchangeability between compute fragments and resource fragments; Fig. 4 is a graphical representation further showing ways in which compute fragments and resource fragments can provide equivalent functionality; Fig. 5 is a graphical representation showing how fragments may be combined; Fig. 6 is a graphical representation showing how fragments may be combined and considered equivalent; Fig. 7 is a graphical representation showing a first example of a high quality compute intensive approach; Fig. 8 is a graphical representation showing the second example of a lower quality computationally simpler approach; Fig. 9 is a graphical representation showing the equivalence of resource and compute fragments of figures 7 and figure 8; Fig. 10 shows an output fragment; Fig. 11 shows a caching arrangement; Fig. 12 is a logical diagram explaining the inputs there may be provided to a dynamic compute matchmaking service; and Fig. 13 is a diagram of a system embodying the invention. DESCRIPTION OF A PREFERRED EMBODIMENT OF THE INVENTION The invention may be embodied in a method of retrieving object-based media, devices for retrieving object-based media, servers or notes for delivering object-based media and systems involving such delivery of object-based media. An embodiment of the invention will be described in relation to retrieval by user devices such as smartphones, tablets, smart TVs and similar devices. Overview An embodiment of the invention provides the ability to deliver objectbased media across different networks to heterogeneous devices (devices with varying capabilities) in a more flexible and scalable manner. Such devices include smartphones, smart TVs TVs, tablets or other consumer user devices capable of retrieving object-based media and asserting to a user in the sense of providing an output of audio, video or other media. A goal in delivery of media across networks is so-called sub-linear scaling, namely increasing the resources needed in terms of bandwidth and computation of processing at a lower rate than increasing the number of users or size of network. Such sub-linear scaling can provide significant technical advantage in terms of reduced computational or bandwidth requirement. The embodiment described advantageously enables not just the delivery of media content (video, images, sound etc) to be flexibly delivered, but also the computation needed to render or otherwise deliver media output. In order to deliver both media content and computation in a flexible manner, the embodiment divides both of these into “fragments”. The delivery manipulation of media fragments is known in the art starting with the media fragment working group and the media fragment URI 1.0 recommendation in September 2012. As described in the work of that group, a media resource may be defined along a single timeline and can consist of multiple tracks of data that are parallel along this timeline. These tracks can be audio, video, images, text or any other time-aligned data. A media fragment may represent a part of that resource. The present embodiment arranges a media experience using the concept of “fragment” not just for the media itself, but also the computational tasks to deliver that media. For consistency of terminology, we will refer to “compute fragments” as the fragments that describe computations to be performed and “resource fragments” as the fragments that are manipulated to produce the media output. Such “resource fragments” may be used to generate, audio-video, graphics, images and so on, but are not necessarily stored in a form that can be immediately rendered but may require some computation in order to be output. The separation into “compute fragments” and “resource fragments” allows compute to be offloaded to provide high complexity streams to less capable devices with computational tasks being provided upstream of a client device, for example in the cloud. A system delivering media across the network to client devices has a need to describe the compute fragments and resource fragments and, for that purpose, a manifest is provided. An example manifest is provided later. The manifest may be written in XML or other descriptive markup language and provides a description of the computational building blocks that comprise an object-based media experience with the goal of informing how to run instances of the experience efficiently in a distributed compute environment at scale. The manifest describes pre-defined assemblies of compute fragments and resource fragments that are used to generate the media output. A composition system can select different assemblies from the manifest to recompose the software into different arrangements that satisfy certain QoS / QoE (quality of service or quality of experience) expectations and run efficiently on the available infrastructure. The manifest is thus able to describe different permutations of software such that computations can be distributed or offloaded to a network of devices. Different assemblies in the manifest may be selected to: - provide alternative functionality to use in different contexts, such as in response to live content or user preferences; - allow the quality of an experience to be degraded to satisfy service level constraints; - run different computational strategies to support devices with limited capabilities or resources; - facilitate sublinear scaling by sharing or caching computed outputs using session coalescing and continual fragment matchmaking re-evaluation; - solve the media management problem of delivering universal access to media whilst ensuring a guaranteed quality of experience. We will first describe examples of compute and resource fragments that may be used in an embodiment of the invention, and that a manifest by which such compute and resource fragments may be described and finally a system using the computer and resource fragments and manifest. Fragments The embodiment of the invention uses resource and compute fragments described by a manifest as discussed above. Compute fragments provide the code and may be considered small building blocks of code for delivering a presentation. Resource fragments contain the data needed to deliver media such as video, pictures, audio and so on and may be procedural resources in the sense that they represent the underlying media, rather than being the final form in which the media is output. An example of such a procedural resource would be a description of a sound wave, rather than the soundwave itself. The use of procedural resources is known in fields such as in the games industry. An object-based media experience is comprised of resource fragments and compute fragments and such fragments may be combined in a variety of ways to meet the user’s preferences and context. Combining fragments together is a computational process and each fragment requires some sort of computation to be performed on it. These computations are performed by the (typically small) compute fragments which may be considered individual executable computer programs. Within such an object-based media experience, we identify two basic content types: Static content that is pre-recorded or pre-authored e.g., videos, images, buffers of data; and Dynamic content that is procedurally generated by compute fragments. Compute fragments are arranged to perform content generation tasks, with particular types of compute fragment being provided to define specialised algorithms for content creation such as graphics rendering and compositing. A data-driven compute fragment such as this can interpret any list of graphics commands to procedurally generate content. Compute fragments may therefore be provided that can operate with command lists for graphics rendering tasks. We will describe such command list operable compute fragments as the main example of how compute fragments may operate, but it is noted that other compute fragments may be provided. In the example of graphics rendering using command lists, compute fragments can perform lazy evaluation of the [graphics] command-lists they consume. This allows the computation that performs the generation of output content to be deferred and allows a compute fragment to combine command-lists together into larger command-lists to be processed more efficiently in one step in the future. Content may be represented by a command list and compute fragment. As such, a (command-list, compute fragment) pair act as a promise to generate content. This promise can be substituted in place of content fragments and are equivalent to the content they represent. The decision when to execute the compute fragments becomes an independent decision based on the execution context, such as availability of compute capabilities. The fact that the same content may be represented by a command list and a compute fragment provides the advantage that the execution of the command list to produce the media content may be deferred to an appropriate time or appropriate position within a network. For example, a computationally expensive production of a particular piece of media may be undertaken by a more powerful node within a network prior to an end user device. In contrast, a computationally simple production for another piece of media may be deferred until an end user device itself. In essence, instead of delivering objects of already rendered media across a network, the separation of fragments into resource fragments and compute fragments allows both bandwidth and processing capabilities to be taken into account. We have therefore described how a given media presentation may be represented by a combination of compute fragments and resource fragments. The concept of composite fragment is also introduced. A composite fragment also referred to as a fragment assembly defines how fragments fit together and how fragments may call other fragments to produce an output. This arrangement may take into account factors such as availability of network computational resources or storage for caching based on coherence, in particular temporal coherence. Coherence between resources occurs when two or more users request the same or sufficiently similar resources in such a way that delivery in response to these requests may be efficiently managed, for example by computing or storing once and then delivering the same output in response to two a more requests. Temporal coherence is a particular example of such coherence in which two or more users request the same or sufficiently similar resources within a given time window of such a size that any delay or departure from an ideal version does not appear perceivable to the user. The greater such temporal coherence, the greater the opportunity for caching and reusing content and the approach of using compute fragments and resource fragments is particularly beneficial for leveraging temporal coherence. This is because requests for similar compute or resources can be identified within a network and provided once in response to multiple user requests on a fragment by fragment, or composite fragment by composite fragment, basis thereby delivering media experiences in an efficient manner. The approach also allows an assembly to be parameterised so as to allow variants of a given assembly to produce different outputs for different end user devices or to cope with different bandwidth requirements across the network. Figures 1 to 6 provide examples of the way in which a compute fragment containing a command list may represent a resource fragment and so may be used in an interchangeable manner in terms of representation of the media. The example relates to a task such as compositing two images together. In this example, the command-list must therefore reference two pieces of source content which could be one or more of: pre-recorded or pre-authored content; procedurally generated content; another command-list whose [optionally deferred] execution will generate usable content; a compute fragment that amalgamates command-lists together to produce a command-list whose [optionally deferred] execution will generate usable content; a compute fragment that generates content when executed, such as by executing an input commandlist of instructions; a general-purpose compute fragment that outputs usable content directly without using command-lists. For a compute fragment to be able to reference these different sources of content interchangeably, we provide the ability for a system operating with such fragments to treat them in a similar manner. Whether the content is a prerecorded image or a promise to generate an image, they are treated polymorphically as media Fragments. Figure 1 shows how one resource fragment 10 in the form of a command less fragment containing a command list is equivalent to a resource fragment 12 containing the result of executing the command list and can be used interchangeably as an input to a compute fragment. These two fragments are both resource fragments. Figure 2 shows how a compute fragment 14 can be used to generate an output fragment when executed and so may also be considered in place of a resource fragment 12 containing an output image. Unlike the resource fragment 12, though, the compute fragment 14 may require significantly less bandwidth for transmission around the network. The logical treatment of fragments in a similar manner thus allows more flexible use of bandwidth when determining whether to transmit a compute fragment or a resource fragment. An immediate benefit of this approach is the ability to defer instructions; the more that compute can be deferred to an appropriate point the more compute may be combined for deficiency. In short, execution of processes can be deferred within a network to the most appropriate point. Figure 3 further demonstrates the interchangeability between compute fragments and resource fragments. A compute fragment 14 can be arranged to generate a resource fragment in the form of an output image fragment 12 using a resource fragment in the form of a command list fragment 10. The command list and compute fragment together that may be considered to form a composite fragment which can be used entirely in place of the generated image fragment in terms of provision within the network as the result of executing the compute fragment is equivalent to delivering the output image fragment. Fragments can also be parameterised, for example giving parameters such as bandwidth, fidelity, image size and so on allowing the same output fragment to be generated in a variety of different ways. The concept of composite fragment is an important one that can be leveraged in various ways. Compute fragments and resource fragments may be combined into composite fragments which themselves may be treated as compute fragments or resource fragments. Such composite fragments may be treated as compute fragments or resource fragments in the sense that they may be identified in a manifest and identified and retrieved on request by client devices in the same way as any other fragment. Composite fragments may thus provide the functionality of compute fragments, resource fragments or combinations of these. Figure 4 demonstrates further still the ways in which compute fragments and resource fragments can provide equivalent functionality. A compute fragment 14 can generate a command-list fragment 10 as an output. The compute fragment 14 and the command-list fragment 10 together form a composite fragment. The composite fragment is a promise to procedurally generate an image fragment 12 and it can be referenced in place of the image fragment by using deferred evaluation. Similarly, the compute fragment is a promise to procedurally generate a command-list fragment and can be used in place of an actual command-list fragment. It is noted that the same and resource may therefore be represented in a variety of interchangeable ways using resource fragments and compute fragments. Figure 5 demonstrates how fragments may be combined. So far, the examples have demonstrated the interchangeability of simple resource fragments and compute fragments. The arrangement may be expanded further, though, as shown in relation to figure 5. A command-list fragment 10 will typically reference several input fragments. In this case, the command-list fragment 10 and its input fragments of an image resource fragment 11 and a further image resource fragment 13 form a composite fragment which is equivalent to the image fragment 12 that it represents. The composite fragment can be used interchangeably with the image fragment using deferred evaluation. Figure 6 extends the combinations further still. A first resource fragment in the form of the command list fragment 15 and a second command list fragment 17 may be executed by a compute fragment 14 to produce a combined command is fragment 10 which itself may be executed by a compute fragment to produce an output image resource fragment 12. The combined command list fragment 10 may therefore be considered equivalent to the output image resource fragment 12. Specifically, in this example, a command-list that generates a scene's background could be combined with a command-list that generates a scene's foreground. The result would be a new command-list that references the two original command-lists and potentially adds compositing commands to the resulting list describing how the foreground and background results should be composited together. The resulting fragment could be rendered in a single step. There are lots of opportunities for deferred evaluation in this example, because decisions may be made as to where the compute fragments are executed whether at a server, intermediate node or at a client device. In summary, the arrangement shown in figures 1 to 6 demonstrate that the separation of media into a combination of compute fragments and resource fragments allow significant flexibility in delivery. Quality of Service and Quality of Experience Various ways in which the fragments described may be used to flexibly deliver quality of service (QoS) or quality of experience (QoE) across the network to multiple user devices at scale will now be described with further examples in relation to figures 7 to 10. As previously described, since command-lists can be used as promises / proxies for the results they compute, it is useful to consider execution of command-lists as an implicit function in the subsequent figures. Compute fragments can be executed at a time and place that suits the infrastructure and so execution can be considered separately from the representation of content composition, which is the goal of the manifest. This ability to combine and separate fragments facilitates optimisation, composition, abstraction, and deferred evaluation of compute tasks. Some example applications include one or more of: Composition of fragments into assemblies of a suitable granularity to foster worthwhile reuse of server deployments and of computed outputs; Merging computations to ensure they are executed on a single device to leverage locality of data and reduce network traffic overhead (both in terms of fetching input data and in terms of interfragment communication); Keeping a single hungry GPU fed with a larger command-list is more efficient than submitting smaller command-lists; Separating fragments so that intermediate results can be cached to reduce repeated re computation; Fitting the computation efficiently to the network topology used by distributed computing applications and compute offloading scenarios. Several content fragments can be authored that represent the same piece of media but have distinct levels of quality or cost. This permits control over the cost of the computation and the quality of the overall result. For example: Compute Fragments can be parameterised to vary the quality of their output; Simpler source assets can be referenced (e.g. lower resolution, lower quality); Simpler command-lists can be authored that describe simpler compositing and rendering approaches; Less complex fragment compositions can be selected to replace expensive fragment compositions. At the extreme, a single content fragment can be used to replace a large assembly of expensive fragments, trading off quality or available features. Figures 7 to 10 provide one example by which a high quality compute intensive approach may be used (figure 7) or a lower quality, computationally simpler approach (figure 8). The figures may be considered to represent a quality ladder in which each rung is represented by a different composite fragment. The command list fragment 20 and state buffer 22, in figure 7, may together be combined by execution of the command list fragment to produce yet further command list fragment 24. Such command list fragment 24 may be further combined with an already rendered image resource fragment 26 using a command list fragment 28 and so the whole set of fragments may be equivalent to an output image fragment 30. Such equivalency allows for the distribution of both the media and the compute throughout a network, rather than the simple distribution of the end composite image 30. Similarly, the same output image fragment 30 may be represented by a simpler set of fragments, as shown in figure 8, comprising a command list fragment 21 and image resource fragment 23 which may be combined by a command list fragment 28 and so considered equivalent to the output image fragment 30. This combination of representing a given output image fragment as a command list fragment and image resource fragment opens up the possibility of caching as there is a choice to cache the output to be used by devices or cache intermediate outputs. As discussed earlier, the flexibility provided by separation of media into compute fragments and resource fragments allows for the choices of where to cache and where to compute to leverage coherence between requests for compute or resources. In particular, a temporal coherence window may be defined, namely a time period for which a given request may be delayed without a user perception to allow other similar request to be received and delivered against using the same resources. The same output image fragment 30 may be further represented by lower quality representations, as shown in figure 9, in which a first image resource fragment 25 and a second image resource fragment 27 are combined by a command is fragment 28 and so may be considered equivalent to the output image resource fragment 30. The storage and computation may be performed at any point within a network. This allows dynamic choices to be made regarding such storage and computation. Finally, instead of passing any of the compute fragments or resource fragments around the network, the end output image fragment 30 itself is, of course, represented as a result fragment and can itself still be passed around the network in situations that bandwidth is not a concern, but computational complexity is the higher concern. This would be known as pre-canned content as shown as a output fragment 30 in figure 10. The decision of when and where to execute compute fragments is an orthogonal concern. Nevertheless, it can be useful to provide metadata to hint to an execution “engine” which fragments would be good candidates for caching or offloading to a more capable device. For example, fragments can be reused by different computations and the result of a compute fragment can be cached and reused to address scalability of execution. Metadata data associated with fragments can help a matchmaking algorithm setup a cache, cache content and use caching appropriately. Knowledge of the temporal coherency or state coherency of certain content fragments would help inform the caching decisions and policies, should the execution engine choose to honour them. Figure 11 shows an example of caching in which the cached fragment 40 may be combined with another fragment 42 using the command as 44 resulting in a further cached fragment 46. Manifest We have so far described the ways in which compute and resource fragments may be considered as separable representations of a final output resource fragment to be delivered to an end user device. We have also described how such separation allows for computational tasks to be separated and performed at different points within a network as well as the resources being rendered at different points within the network. In order to understand and overall media experience, though, a description is needed of that media experience specifying the fragments and composite fragments. We refer to such a description comprising the pointers to the various fragments as a manifest. The manifest defines how to assert valid combinations of compute fragments and resource fragments. We use “assert” to cover a variety of ways in which the client device may consume and use fragments, including executing compute fragments, displaying resource fragments comprising images, providing an output of audio for audio fragments and any other normal way in which devices may use, execute, display or otherwise assert retrieved fragments. The concept of the manifest is known for protocols such as MPEG-DASH and includes resource locations and descriptions of media. In the present embodiment, the concept of a manifest is the same, but instead of directly reporting media presentation descriptions, the manifest specifies the fragments. In short, a range of composite fragments for a given experience are specified in a manifest. The manifest may be retrieved and by a system that performs composition and compute matchmaking. The goal of such a system is to dynamically select fragments from the manifest according to the user’s preferences and their context of consumption. The use of the manifest also achieves a goal of balancing service level objectives such as cost, quality and reliability with the available network and compute resources (which includes the user’s device capabilities). Such a matchmaking system could dynamically adapt to changing conditions by selecting the most appropriate pre-defined fragment from the manifest. The manifest stores information about the cost (computational and / or bandwidth) of different fragments and defines a quality of experience ladder to permit bounded degradation of experience quality to meet service level objectives. Figure 12 provides a logical diagram explaining the inputs that may be provided to a dynamic compute matchmaking service 50 which may be operable at client devices, and nodes within a network or at a server. The inputs include a manifest 56 for media content, user context and preferences 58, session state 60, service level objectives 52 and compute and context metrics 54. These may be provided to a dynamic algorithm referred to as the dynamic compute and matchmaking service which may be deployed as shown at various points within a network 62. The QoS-aware composition and matchmaking may use a client-based selection algorithm or a server-assisted algorithm with additional knowledge about the current capacity and capabilities of the infrastructure, together with knowledge of other active sessions. A manifest for QoS-aware composition and matchmaking describes computational fragments that can be offloaded and executed remotely. It describes relationships between fragments and how they can be combined, either through compositing fragments together into more complex assemblies or by executing individual fragments and compositing their outputs together. It describes the set of valid / legal permutations of the software that deliver the functionality required. The manifest defines a quality ladder for fragments of a programme or experience defined by producers and content creators. The manifest’s schema has several core concepts: Global configuration; Fragments (resources, buffers, command-lists, programs, etc.); Fragment references (instancing and reuse of content fragments); Feature fragments (modelling of application-level features with content fragments); Variant fragments (parameterisation of fragment instances). There are many different fragment types. The manifest’s schema is extensible enough to accommodate new fragment types in future. Notable fragment types include: Fragments (e.g., compute, command-list, resource, buffer, state data etc.); In-line fragments; Referenced fragment instances; Composite Fragments; Choice fragments (permit selection from a choice of alternative fragments); Feature fragments (describe optional fragment instances); Variant fragments (parameterised fragment instances). Features describe optional fragments that can be enabled or disabled by a feature toggle. Features may need to be disabled if there are insufficient resources to run the computational tasks to generate the feature’s output content. Features may also be disabled according to user preferences or because they are not relevant at certain points on the experience timeline. Variants associate fragments with one or more different parameterisations that control QoE of generated content, among other things. Parameters can define either a quantised set of quality levels or a quality function parameterised with continuous values. It also defines quality ladders for global parameters effecting content generation. An example might be control over rendition such as global illumination, draw distance and choice of shadow technique. System A system embodying the invention shown in figure 13. Various sources 70, 72 a provide fragments via a network 100 comprising nodes 80, 82, 84 to be consumed by user devices 90, 92, 94, 96. As an example, source 70 could be a repository of fragments comprising audio, video and text and source 72 could be a repository of command lists. As fragments are identifiable using a manifest as described, the storage of fragments belonging to a particular audiovideo experience may be distributed across multiple different sources. Typically, the sources 70, 72 will be provided by a broadcaster. The network 100 may be a broadcast network, the Internet and comprise a combination of wired and wireless communications. The nodes may themselves be sources of fragments, processing resources for executing command lists, caches or sources of fragments. As an example, node 80 could be a computing resource, node 82 a source of further video fragments and node 84 a content delivery network comprising multiple audiovideo resources and computational resources. The user devices may comprise any readily available consumer device such as a set-top box, smart TV, computer, smart phone, tablet or any other device capable of requesting and consuming audio, video, text and generally any media content. As an example, device 90 could be a smart phone connected via Wi-Fi, device 92 a set-top box, device 94 a smart TV and device 96 a smart phone connected via 5G. Although only a handful of sources, nodes and devices are shown, it is to be understood that the arrangement is scalable to have many thousands or millions of sources, nodes and end user devices. The operation of the system in broad terms will first be described prior to describing an algorithm operable to select fragments. A user may request content via a user device 90, such as interactive audiovideo content, from a source 70 via one or more nodes 80, 82 across a network 100. In response to the request for content from the user device 90, a manifest is first provided from the source 70 to the user device 90, or to an edge device upstream from the user device 90, defining the compute fragments and resource fragments of the audiovideo content as a fragment assembly. As previously described, the compute fragments define code operable to build the presentation of the audiovideo content and the resource fragments provide the content, either directly as audio, video or images or as procedural resources by which these may be created. Also as previously discussed, compute fragments may generate further compute fragments or resource fragments and the precise combination of computing resource fragments may vary from one presentation to another presentation. The manifest describes how to assemble the compute and resource fragments together to create the audiovideo content. As a result of the manifest describing both the compute and resource fragments, the same audiovideo content may be delivered in different ways to different devices. To give a working example, consider a weather forecast involving audiovideo and subtitles. Fragments in the form of compute fragments and resource fragments may be combined to show the presenter, one or more images and a graphical user interface. In view of the fact that the end devices may have different bandwidth and processing capabilities, the same content may be provided with different fidelity by either changing the experience (reducing bandwidth, providing lower resolution) or changing the experience (removing some features such as graphics). More generally, the separation of a given multimedia experience into fragments of compute fragments and resource fragments allows those fragments to be shared amongst multiple user devices. Multiple devices may request a given audiovideo experience at a similar time. Instead of choosing to deliver all of that experience simultaneously to multiple devices, part of the experience may be simultaneously shared by delivery from the same nose to multiple devices by delivering selected fragments. In this way, different device capabilities may be catered by distributing resource delivery and compute on a fragment by fragment basis. The matching between compute resources and compute need may be algorithmically defined on a fragment by fragment basis thereby making objectbased media scalable and determining features such as which resources to cache, which data to compute at source, at an intermediate node or on a device. Algorithm An example algorithm is included that may use a manifest to make appropriate selections of candidate fragments. As already explained, a manifest for an object-based media experience is a collection of fragment-assemblies, each providing a variant of that experience. This variant can be a composite hierarchy of fragments, each configured with a specific set of parameters. Or the assembly can denote different compositions of fragments based on user preferences. The purpose of the first algorithm is to allow selection of fragments. Later, a second algorithm as discussed which is arranged to compute the “distance” between requests to determine if they are sufficiently similar to allow leverage of temporal or other coherence. The first algorithm may be operated at any point within the network, but is preferably operated by a user device. The user device retrieves the manifest which defines how the fragment assembly has been authored, tested and how it can be delivered. A quality score for a given fragment assembly may be defined as described below allowing the client device to select a particular quality of fragment assembly. This can include a tolerated start delay allowing for reuse of existing resources if a given fragment assembly has already been instantiated as a result of selection by different client device. The algorithm aims to optimise bandwidth, but other features may be optimised such as latency, compute, cache or other resources. A fragment-assembly has the following characteristics that are relevant for its selection: Set F of fragment-ids (simple or composite fragments): {Fo,..., Fn} User preference set, Ur {} (feature-selected, objects-set, story-branch, device orientation, viewpoint, accessibility feature) Perceived quality score, Q (quality perceived by user based on a verified QoE model): [0.. n] Bit rate for the resulting media stream output by the assembly: B Mbps Temporal-coherence-window: {Tstart, Tend} Estimated computation cost: £”=0 cost^Fi) Cold start delay, dF ms A quality estimation may be performed for a fragment assembly. The Perceived quality score, Q for an assembly of media objects is estimated using QoE models for Object Based Media. Given a set of n fragments in the assembly, a composite model to determine the perceived quality for the whole assembly is provided by a function: n r ^Wi.QoECFd QoEp = ------------ n Where the importance of a fragment in the assembly is a function of its size, motion, complexity (Spatial Information) and narrative prominence (author-indicated). i.e. Wi = Size(i) + Motional) + 57 (i) + N arrative(i) And, the individual QoE score of the fragment is determined by a validated QoE model based on the object type. There are validated models for Audio and Video such as P.1203, VMAF, PSNR, etc. The QoE contribution for graphics is a function of its size and placement. A quality selection algorithm may be executed. This first algorithm is arranged to select the most optimal assembly from the list in the manifest, will take as input: User preference set, Ux Estimated bandwidth at the client, bx Buffer occupancy level at client, Ox Frame / sample rate (to calculate time client is behind live edge), rate Minimum quality level threshold, Qm / n Set of Assemblies, manifest Current experience time, tx Function select assembly (Uxr bx ; Ox , tx, rate, Qmin , manifest, deployed assembly set) candidates = [] For each f in manifest If (Ux c Uf) AND (bx <Bf) AND (Qf >Qmin) Then / / if we use run-time info e.g. temporal coherence window If (f. Tstart <= tx <= f.Tend) Add f to candidates End I f End If End for If Candidates[] is NOT Empty Sort candidates[] by increasing Cost, then by increasing start^delay or decreasing quality Return candidates[0] End If The algorithm can be improved by considering run-time factors such as fragments that are already loaded or cached according to the temporalcoherence window of a fragment, and for this purpose a second algorithm is provided. The second algorithm is arranged to compute a “distance” to determine whether there is sufficient similarity (coherence) between selected fragments in order to deliver multiple requests with a single response. Given the large number of variables involved, it is preferable to reduce the amount of data for determining similarity between fragments and for this reason, a representation of fragments may first be determined. Such a representation may be a vector, matrix, hash function or any other function that may be operated to reduce the amount of data or complexity for the purpose of performing the comparison. A hash function will be used as the main example. Each Fragment has a hashing specification to drive a session-coalescence process. Multiple compute fragment instances that share the same hash can be coalesced together and serviced by a single server or cache, thereby avoiding unnecessary duplication of work. The hashing specification defines a subset of compute fragment inputs that form the part of the hash that must match and a set of compute fragment inputs that can optionally match (called “negotiable” inputs). An example of a negotiable input is quality level. This permits greater session-coalescence between experience instances, at the expense of only having an approximate match. Selection of a candidate fragment from a cache requires us to compute the distance between hashes of the fragment input (S1) and the inputs of the cached fragment result (S2) to estimate coherence of computed results: SlInput = { {fragment id}, {mandatory values}, {negotiable values} } S2Input = { {fragment id}, {mandatory values}, {negotiable values} } D = Distance(hash ( SI), hash(S2)) If( D <0 ) / / mandatory values did not match / / Regenerate and cache the compute fragment output Routput = run^fragment (SlInput) cache (hash (SlInput) , ROutPut) Else If ( D == 0 ) / / perfect match / / Use distance func to lookup cached fragment output Routput = lookup cache (hash ( S2Input) ) Else / / mandatory values matched, but variance in negotiable input values Scanciidates = lookup all candidates from cache (hash (S2Input) ) Routput = select_candidate ( Scandldate3) Endif In broad terms, the operation of the first algorithm is to operate a method of selecting an assembly of fragments in a system having computational tasks as compute fragments and media as resource fragments, comprising: determining for each assembly whether the assembly satisfies one or more of bandwidth or compute thresholds; determining for each assembly a sorting of assemblies by one or metrics to produce a sorted list; selecting the top assembly from the sorted list. The method preferably involves selecting the assembly comprises computing a hash of each available assembly defined in a manifest and selecting those assemblies as candidates that satisfy a distance calculation. The method further preferably involves sorting according to one or more metrics selected from computational cost, start delay or quality. Example Manifest As previously described, the manifest for QoS-aware composition and matchmaking describes computational fragments that can be offloaded and executed remotely. It describes relationships between fragments and how they can be combined, either through compositing fragments together into more complex assemblies or by executing individual fragments and compositing their outputs together. It describes the set of valid / legal permutations of the software that deliver the functionality required. The manifest defines a quality ladder for fragments of a programme or experience defined by producers and content creators. The following example is one example of such a manifest. - Computational cost metric - Memory cost metric - Bandwidth cost metric - Latency cost metric - Warmup latency
Claims
1. A method of delivering object based media, comprising:- providing computational tasks as compute fragments;- providing media as resource fragments;- providing a manifest defining how to assert valid combinations of compute fragments and resource fragments; and- using the manifest to retrieve the compute and resource fragments to present media.
2. The method of claim 1, wherein the manifest provides locations of the compute and resource fragments.
3. The method of claim 1 or 2, wherein the manifest provides parameterisation of the resource fragments, preferably one or more of resolution, quality, colour, or preferences in relation to the resource fragments.
4. The method of any preceding claim, wherein the manifest describes how assemblies of compute and resource fragments may represent other resource fragments.
5. The method of any preceding claim, wherein the method is operated to retrieve object-based media across a network comprising multiple nodes, further comprising determining at each node whether to execute retrieved compute fragments or deliver retrieved compute fragments for execution at another node or at a client device.
6. The method according to claim 5, wherein comprising determining at each node whether to execute retrieved compute fragments or deliver retrieved compute fragments for execution at a subsequent node in a network of nodes.
7. The method according to claim 5, wherein comprising determining at each node whether to execute retrieved compute fragments or deliver retrieved compute fragments for execution at a lateral node in a network of nodes.
8. The method of any preceding claim, wherein the manifest defines a set of assemblies of fragments, further comprising operating an algorithm to select one of the assemblies from the set of assemblies.
9. The method of claim 8, wherein the algorithm selects the most optimal assembly taking as an input one or more parameters selected from user preferences, bandwidth at the client, buffer occupancy at the client, frame rate or sample rate, minimum quality threshold and current experience time.
10. The method of claim 8 or 9, wherein the algorithm operates a matchmaking process by determining combinations of multiple compute fragments requested by user devices that may be serviced by a single server or cache, and delivering each such combinations from the single server or cache to multiple client devices.
11. The method according to claim 10, wherein determining the combinations comprises computing a representation of each combination and determining whether a distance between the representation of each combination satisfies a threshold requirement.
12. The method of claim 11, wherein the representation includes only non-negotiable parameters of the combinations of multiple compute fragments, the step of determining combinations of multiple compute fragments further comprising matching by selecting a top candidate of multiple candidates ordered by one or more negotiable parameters.
13. The method of claim 12, wherein the negotiable parameters include one or more of computational cost, start delay or quality.
14. The method of any preceding claim, wherein compute fragments and resource fragments may be combined into composite fragments which themselves may be treated as compute fragments or resource fragments.
15. The method of any preceding claim, wherein providing the manifest comprises retrieving by a client device the manifest from a source across thenetwork and wherein using the manifest comprises requesting an assembly of compute and resource fragments from a set of assemblies defined by the manifest.
16. The method of any preceding claim, wherein the method is operable at a server device or node within a network of nodes.
17. The method of claim 16, comprising receiving a request for a manifest, providing the manifest in response to the request, and using the manifest to retrieve the computer resource fragments comprising receiving a request from a client device.
18. A method of retrieving object based media at a client device, comprising:- requesting a manifest defining how to assert valid combinations of compute fragments and resource fragments;- receiving the requested manifest;- using the retrieved manifest to request the compute and resource fragments to present media; and- receiving at least some requested compute resource fragments and other composite fragments that are the result of executing requested compute fragments.
19. A method operable at a node in a network of nodes for delivering object based media, comprising:- receiving computational tasks as compute fragments;- receiving media as resource fragments;- receiving requests for compute and resource fragments from a client device using a manifest to request the compute and resource fragments; and- determining at the node whether to execute retrieved one or more of the compute fragments or deliver retrieved compute fragments for execution at another node or at a client device.
20. The method of any claim 19, further comprising determining whether to deliver the requested compute and resource fragments or to execute one or moreof the compute fragments in dependence upon coherence of multiple received requests.
21. The method according to claim 20, comprising determining at the node whether to execute retrieved compute fragments or deliver retrieved compute fragments for execution at a subsequent node in a network of nodes.
22. The method according to claim 20, comprising determining at the node whether to execute retrieved compute fragments or deliver retrieved compute fragments for execution at a lateral node in a network of nodes.
23. The method of any of claims 19 to 20, wherein determining at the node whether to execute retrieved one or more of the compute fragments comprises operating a algorithm that provides a matchmaking process by determining combinations of multiple compute fragments requested by user devices that may be serviced by a single server or cache, and delivering each such combinations from the node to multiple client devices.
24. The method according to claim 23, wherein determining the combinations comprises computing a representation of each combination and determining whether a distance between the representation of each combination satisfies a threshold requirement.
25. The method of claim 24, wherein the representation includes only non-negotiable parameters of the combinations of multiple compute fragments, the step of determining combinations of multiple compute fragments further comprising matching by selecting a top candidate of multiple candidates ordered by one or more negotiable parameters.
26. The method of claim 25, wherein the negotiable parameters include one or more of computational cost, start delay or quality.
27. The method of any of claims 19 to 26, wherein compute fragments and resource fragments may be combined into composite fragments which themselves may be treated as compute fragments or resource fragments.
28. A client device configured to retrieve object based media, comprising:- means for requesting a manifest defining howto assert valid combinations of compute fragments and resource fragments;- means for receiving the requested manifest;- means for using the retrieved manifest to request the compute and resource fragments to present media; and- means for receiving at least some requested compute resource fragments and other composite fragments that are the result of executing requested compute fragments.
29. A server or network node for delivering object based media, comprising:- means for providing computational tasks as compute fragments;- means for providing media as resource fragments;- means for providing a manifest defining how to assert valid combinations of compute fragments and resource fragments; and- means for using the manifest to retrieve the compute and resource fragments to present media.
30. A server or network node according to claim 29, comprising means for undertaking the method of any of claims 1 to 17 or 18 to 27.33
Citation Information
Patent Citations
Manifest data for server-side media fragment insertion
US10863211B1