Method and apparatus for processing media content
By defining layout rules and user characteristics, configuring and selecting media objects, the problem of personalizing the rendering of object-based broadcast content on different devices is solved, achieving the effect of personalized media content presentation on different devices.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-14
- Publication Date
- 2026-03-27
AI Technical Summary
When rendering object-based broadcast media content on user devices of different shapes and sizes, it is difficult to simultaneously meet the personalized needs of different users and maintain the advantages of OBB.
By defining layout rules and user characteristics, configuring and selecting media objects to form a suitable presentation, satisfying user-related characteristics and constraints, and using computer-implemented methods to optimize the rendering process of media content.
It enables personalized presentation of media content across different devices, ensuring the prominence of important media objects and the overall quality of user experience, while adapting to the preferences and device constraints of different users.
Smart Images

Figure CN116195260B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates to methods and apparatus for processing media content. In particular, the present invention relates to a computer-implemented method for processing media content that is to be rendered as a presentation for a user at a set of one or more media devices (such as televisions, tablets, smartphones, etc.), the media content comprising media objects, at least some of the media objects comprising video content in a technology known as "object-based broadcasting". BACKGROUND
[0002] Object-Based Broadcasting (OBB) is a term used to describe a mechanism that allows television (TV) programs and other such presentations of media content to be personalized. In this context, an "object" is a different media component that can be brought together to make up a TV program or other such presentation. These media components can include video content that is cut together (e.g. to tell a story, show a sporting event, or present information on a topic), music, speeches and special effects, video replays and slow-motion replays (particularly for sports programs), subtitles, picture-in-picture inset pictures, graphics, commentary, on-screen signers to provide explanations for the deaf, and virtual reality (VR) overlays rendered in a studio.
[0003] In traditional (i.e. non-OBB) television, the presentation and timing of these media "objects" (i.e. whether, when, where and how they appear on screen or are heard) is controlled by those who make the program. By not fixing the arrangement of these objects and giving the viewer some control over what objects can be accessed and how they are presented, the content provider can allow the user to personalize their experience of the program or other such presentation.
[0004] For many years, producers of TV programs, films, etc. have had to make some compromises to scale their content to be presented on different screens. Some common examples are shown in Figures 1(a) and 1(b).
[0005] Figure 1(a) shows a 4:3 image displayed in a 16:9 screen (where X:Y relates to the ratio of horizontal to vertical dimensions or number of pixels). In this case, "pillars" (shown in black) on either side of the 4:3 image fill the rest of the screen. This is known as "pillarboxing".
[0006] The reason for this is more apparent in Figure 1(b), which shows a 4:3 image, composed of 12 blocks (horizontal) by 9 blocks (vertical), displayed in a 16:9 screen. The screen is filled on each side with pillars having a width of two blocks.
[0007] (NB For convenience, the individual blocks within the images in Figure 1(a) are labelled using a "row:column" numbering system, identified by "m:n", where the top left block is numbered "1:1" and the bottom right block is numbered "9:12" - this is purely to allow easy viewing of the number of blocks in each row and column - the numbering system is arbitrary and will be simplified in later Figures 3(a), 3(b) and Figure 4
[0008] Figure 2 Figure 1(b) shows a 21 :9 image, displayed in a 16:9 screen. In this case, the bars above and below the image fill the screen. This is known as "letterboxing".
[0009] While broadcasters have made some efforts to prepare images for screens of different sizes, television manufacturers also provide user options to adjust the image, providing "Fill" or "Zoom" functions, which stretch the image to fill the entire screen, possibly losing the proper aspect ratio within the shot, making faces longer or wider than they are.
[0010] Such options can therefore provide an alternative to pillar and letterbox. Figures 3(a) and 3(b) show the effect of using options which involve stretching the image to fill the screen. The top half of the figure (Figure 3(a)) shows an unstretched 16:9 image on a 4:3 screen, filled at the top and bottom using letterbox, while the bottom half (Figure 3(b)) places the same 16:9 image in the same 4:3 screen, however stretching the image vertically to fill the entire screen, thus avoiding the use of letterbox (while slightly distorting the image).
[0011] Another alternative that has been used involves manually "panning and scanning" (the shape of the target screen) windows over the original image, and then using those cropped images to fill the target screen. This panning and scanning approach helps improve awkward shots where important parts of the image (possibly a "two-shot" (i.e. capturing the faces of two people sitting at a table and talking to each other) are cropped too tightly. While resizing the image for different screen sizes or shapes can result in losing part or all of one of the faces, panning and scanning can allow both to be shown at different times, it can make it difficult to show a line delivered by one character and the reaction from the other character at the same time.
[0012] Screens now appear not only on televisions and in cinemas, but also on smartphones, tablets, phablets and PCs. These screens do not slavishly adhere to a 16:9 aspect ratio (some televisions can even be found with a 21 :9 aspect ratio). Even if they adopt a common aspect ratio, especially telephones are likely to be viewed in "portrait" mode in some circumstances, forcing even more severe top and bottom letterboxing.
[0013] Figure 4 A 16:9 landscape image displayed in a 16:9 screen is shown, with the screen remaining in "portrait" orientation.
[0014] Thus, it is common to view images on "closed format" screens using left and right letterboxing and top and bottom letterboxing. The screen can also offer a function that allows all the pixels on the screen to be lit, but at the cost of seeing all of the image.
[0015] To ensure that important information in the image is visible, there is the concept of a "safe area" - essentially a central area of the defined screen that is assumed to always be (or at least should be) visible, regardless of the screen on which the image is presented. In traditional (non-OBB) content provision, the content creator or provider can ensure that any graphical elements added to the main element (e.g. a leaderboard or scorecard overlaid on a video image of a sporting event) are located in a part of the display that prevents any graphical element from obscuring the central part of the main element, and which generally does not obscure the central part even if the image is stretched, narrowed, cropped or otherwise adjusted for a different size of screen.
[0016] The above approach is characterised in that all image components (video, graphics etc.) of the screen are on a single layer, and all image components are scaled or cropped using a single function.
[0017] Referring to various existing disclosures, from w3schools.com at https: / / www.w3schools.com / html / html_responsive.aspAn online guide to techniques for using Hypertext Markup Language (HTML) and Cascading Style Sheets (CSS) to automatically resize, hide, shrink, or enlarge a website to look good on different types of devices (desktops, tablets, and phones) is provided by a webpage entitled "HTML Responsive Web Design" available at
[0018] In https: / / www.ibc.org / manage / 2-immerse-a-platform-for-production-and- more- / 3316.article A paper entitled "2-IMMERSE: A platform for production, delivery and orchestration of Distributed Media Applications" (dated 27 September 2018) available at describes an overview of the architecture of a multi-screen experience based on MotoGP sports content evaluation developed using object-based broadcast methods.
[0019] In https: / / 2immerse.eu / wp-content / uploads / 2018 / 01 / d2.4-distributed- media-application-platform-description-of-second-release-0.31.final_.pdf A document entitled "2-IMMERSE Deliverable D2.4 (Distributed Media Application Platform - Description of Second Release" (dated 11 January 2018) (in particular section 6.2) describes the 2-IMMERSE Distributed Media Application Platform, the multi-screen experience component and production tools developed for the second service prototype of the project, "Watching MotoGP at Home", and discusses the technical results of the project and details of the current status of the platform, components and key features.
[0020] In https: / / www.youtube.com / watch?v=FZIhrnGzC4I A video entitled "2-IMMERSE MotoGP Service Prototype Video" (dated 17 January 2018) available at introduces the 2-IMMERSE MotoGP service prototype and shows its features in action. In particular, the commentary refers to the ability to adjust and scale the layout of the graphics on the screen.
[0021] In https: / / ir.cwi.nl / pub / 28131 / 28131.pdfA paper titled "Workflow Support for Live Object-Based Broadcasting" by Jack Jansen, Pablo Cesar & Dick Bulterman (DocEng '18, August 28-31, 2018, Halifax, NS, Canada) available at https: / / doi.org / 10.1145 / 3236257.3236270 examines the document aspects of object-based broadcasting. It presents a model and implementation of a dynamic system for supporting object-based broadcasting in the context of sports applications. It defines a multimedia document format that supports dynamic modification during playback, which allows for agent activation at the receiving end of the content by editor decisions.
[0022] Referring now to existing patent literature, US patent US9569501 ("Chedeau et al.") relates to optimization of electronic layout of media content. In one embodiment, a method is described that involves accessing N electronic media content items and a plurality of media content templates, where each media content template includes a predetermined number of surface areas for a predetermined number of media content items. The method includes scoring, for each of one or more media content templates, placement of X electronic media content items in the media content template based on one or more features, where X is equal to the lesser of N and the predetermined number of surface areas of the media content template. The method includes selecting one of the media content templates having the highest score and providing the X electronic media content items in the selected media content template for display to a user.
[0023] While the option of OBB clearly offers potential advantages in terms of user experience and other aspects, presenting media content using OBB technology to users with different requirements and preferences, where each user's presentation of a particular program can include a different set of media objects, and with other possible variable factors, introduces challenges as to how best to provide the media content when the presentation can be rendered and displayed on different possible shapes and sizes of user devices. While some users can be able and / or can prefer to set and / or make their own adjustments to their presentation, this can be done one program at a time by setting general preferences, or in other ways, other users can not be able or can not wish to do so, or can simply prefer to have their presentation provided in a form that does not require setting or adjustment. Providing OBB media content to different users in a way that maintains the benefits provided by OBB while at the same time conforming to the possible requirements / expectations of different users is challenging without knowing the different contexts in which the presentation will be viewed by different users. SUMMARY
[0024] According to a first aspect of the application, there is provided a computer- implemented method for processing media content to be rendered as a presentation for a user at a point in time at an arrangement of one or more media devices, the presentation being based on layout rules defining suitability and configuration of media objects for rendering as part of the presentation, the arrangement and one or more user characteristics and / or attributes constituting a context of the presentation, wherein the presentation is formed from media objects selected from a set of media objects and the context has associated one or more constraints, each constraint defining an attribute of the context affecting rendering of at least a subset of the selected media objects, the method comprising the steps of:
[0025] configuring characteristics of each media object in the set of media objects, the configured characteristics complying with utility conditions based on utility metrics of the media object in the context at the point in time, the utility metrics being evaluated in relation to the constraints of the context; and
[0026] identifying selected media objects in the set of media objects based on the utility metrics associated with each selected media object and the layout rules.
[0027] The set of media objects can include media objects providing one or more of video content, audio content, text content and graphical content. Other types of media objects are also possible.
[0028] Media objects providing video content can provide content such as live video, replay video (sports action replay etc.), computer generated video content (e.g. special effects), on-screen sign language (e.g. for deaf or hard of hearing people), picture-in-picture inset, virtual reality overlays rendered by a radio station etc.
[0029] Media objects providing audio content can provide content such as music, speeches (from a person shown in a video object or otherwise), sound effects, background sounds, commentary (e.g. on a sporting event) etc.
[0030] Media objects providing text content can provide content such as subtitles, information about video, audio or other content, information about a sporting event being broadcast (e.g. scores, scorecards or league tables) etc.
[0031] Media objects providing graphical content can provide content such as graphics, sports team formations or tactical explanations etc.
[0032] According to preferred embodiments, the layout rules defining suitability and configuration of media objects for rendering as part of a presentation can include rules determining whether, when, where and how individual media objects are rendered. These can be based on, for example, requirements / preferences of the provider, producer or director of the overall content, and / or of one or more users / viewers of the content.
[0033] According to preferred embodiments, the characteristics of a media object can include one or more of the size of the object's graphic, screen position, colour scheme, transparency (i.e. whether and how easily an object in front of other objects allows the objects behind to be seen) and layering order (i.e. which visual objects appear in front of or behind other objects).
[0034] According to preferred embodiments, the set of one or more media devices for an arrangement at a particular point in time can include more than one media device for an arrangement. The devices can include large screen objects such as televisions or computer screens and handheld and / or small screen devices such as tablets or smartphones, or other devices such as "dual screen" or multi-screen arrangements. In such embodiments, the characteristics of a media object can include the media device in the set of one or more media devices on which the media object should appear, allowing the user / viewer to ensure that certain objects (e.g. objects bearing statistics, or live chat for example) appear on, for example, a handheld device.
[0035] According to preferred embodiments, the step of identifying selected media objects from the set of media objects can be performed by adding media objects to a list of media objects to be rendered based on the utility values assessed in respect of the media objects until it is determined that the applicable layout rules cannot be adhered to. Such a technique can be used to ensure that objects deemed most important are prioritised (based on a combination of applicable factors which can include any applicable user preferences provided by the user).
[0036] According to preferred embodiments, the step of identifying selected media objects from the set of media objects can be performed by identifying media objects such that the sum of the utility values assessed in respect of the media objects is maximised without breaking the applicable layout rules. Such a technique can be used to ensure that the overall "best compromise" determined (which can be appropriate if certain media objects would be highly desirable if selected) will result in several other objects being missed out or de-emphasised.
[0037] According to preferred embodiments, the steps of configuring and identifying can be performed at least partially before the selected media objects are transmitted to the one or more client media devices. The complete or partially complete presentation can then be transmitted from the provider or an intermediary entity to the media devices of the one or more users.
[0038] According to an alternative embodiment, the steps of configuring and identifying can be performed at least partly after the set of media objects is transmitted to the one or more client media devices. Such an embodiment can be used to allow local expressions or locally available user preferences and / or requirements to be more easily incorporated into the decision making process.
[0039] According to a preferred embodiment, the method can further comprise rendering the selected media object of the set of media objects. Such rendering of the selected media object can be performed after the selected media object is transmitted to the one or more client media devices. In embodiments where the steps of configuring and identifying are performed before the selected media object is transmitted to the one or more client media devices, such rendering can be performed before the selected media object is transmitted to the client media devices.
[0040] According to a preferred embodiment, the method can further comprise providing the selected media object of the set of media objects as a presentation via the one or more client media devices.
[0041] According to a second aspect of the present invention, there is provided an apparatus for processing media content to be rendered as a presentation for a user at a set of one or more media devices arranged for a point in time, the presentation being based on layout rules defining suitability and configuration of media objects for rendering as part of the presentation, the arrangement and one or more user associated characteristics and / or attributes constituting a context of the presentation, wherein the presentation is formed from media objects selected from a set of media objects and the context has associated one or more constraints, each constraint defining a characteristic of the context influencing rendering of at least one subset of the selected media objects, the apparatus comprising a computer system comprising a processor and a memory storing computer program code for performing the steps of the method according to the first aspect.
[0042] According to a third aspect of the present invention, there is provided a computer program element comprising computer program code which, when loaded into a computer system and executed thereon, causes the computer to perform the steps of the method according to the first aspect.
[0043] The various options and preferred embodiments mentioned above in relation to the first aspect can also apply in relation to the second and third aspects. BRIEF DESCRIPTION OF DRAWINGS
[0044] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings, in which:
[0045] Fig. 1 (a), Fig. 1 (b), Figure 2 , Fig. 3 (a), Fig. 3 (b) andFigure 4 Techniques are shown that enable the scale of content to be adapted for presentation on different screens;
[0046] Figure 5 is a block diagram of a computer system suitable for operation of embodiments of the application;
[0047] Figures 6(a) and 6(b) show entities that can be involved in performing a method according to embodiments of the application in accordance with two possible scenarios;
[0048] Figure 7 steps that can be performed in a method according to preferred embodiments of the application are shown; and
[0049] Figure 8 is a chart showing how utility values for presentation values such as the size of rendered on-screen graphics for different constraint values can be calculated based on the quality of a user's sight. DETAILED DESCRIPTION
[0050] With reference to the appended drawings, a method and apparatus according to embodiments will be described.
[0051] First, Figure 5 is a block diagram of a computer system suitable for operation of embodiments of the application. A central processing unit (CPU) 502 is communicably connected to a data storage device 504 and an input / output (I / O) interface 506 via a data bus 508. The data storage device 504 can be any read / write storage device or combination of devices such as random access memory (RAM) or non-volatile storage devices, and can be used to store executable and / or non-executable data. Examples of non-volatile storage devices include disk or tape storage devices. The I / O interface 506 is an interface to a device for inputting or outputting data, or for both inputting and outputting data. Examples of I / O devices that can be connected to the I / O interface 506 include a keyboard, a mouse, a display such as a monitor, and a network connection.
[0052] With reference to Figure 6(a), Figure 6(a) shows entities that can be involved in performing a method according to embodiments of the application in a scenario in which configuration and identification of media objects presented to a particular user is performed at a media device of the particular user.
[0053] In this embodiment, a media device 60 (which can be a television, a smart phone, a tablet, a phablet (a phone / tablet hybrid), a PC, or another such device via which a consumer or user can receive, view, and / or otherwise consume media content) requests and receives media content in the form of media objects 55 from a media content source 50, receiving the media objects via a media content input interface 61. The media content can be for a program such as an object-based broadcast of a sporting event or other such program, a movie, an interactive event, an online computer game, etc. The received media content is passed to a configuration and identification module 62, the function of which will be explained in detail later.
[0054] The configuration and identification module 62 is in communication with a user input interface 64 via which information about user preferences and / or requirements can be received. This information can be actively provided by the user, or can be derived (obtained) from monitoring the user (or users) and / or the environment in which the user and / or media device is located (obtaining information about who the (primary) user / viewer is, how many viewers there are, how large the room is, how far the viewers are from the screen, etc.). The user input interface 64 can also be used to receive other information from the user, including information 58 to be provided from a user output interface 65 of the media device 60 to the media content source 50 or elsewhere. This information 58 can simply include information such as a request for particular media content, but in some embodiments (particularly in embodiments in which some or all of the decisions about proposed layouts are taken at the media content source 50 or at another such device that can provide media content to the user) the information 58 can also include information such as parameters related to the user's display device (e.g., size, shape, resolution, technical capabilities or features) or the user (they are visually impaired, hearing impaired, have a particular interest in a particular person or type of media content that they can request, or in terms of other user preferences and / or requirements, feedback about received media content, and / or other such parameters.
[0055] The configuration and identification module 62 is also in communication with a data store 66 in which data relating to such things as layout rules 66b, constraints 66c, and prioritization factors 66a (discussed below), as well as possibly other types of data 66d, can be stored.
[0056] Based on information received via the user input interface 64 and information retrieved from the data store 66, the configuration and identification module 62 follows a process, which will be described in detail below, in order to form a presentation of media objects selected from those received via the media content input interface 61, the selected media objects being configured based on the information received and / or retrieved by the configuration and identification module 62 in order to meet or optimize the objective criteria indicative of whether the presentation, when rendered and displayed to the user (or users) in question, meets or optimizes the objective user experience criteria.
[0057] The presentation, whose selected media objects are all configured to meet or optimize the objective criteria in question, can then be provided to the media renderer 67 for rendering, and then displayed or otherwise played by the media player 68 as output to the user, the media player 68 itself can be linked to a single display device or a set of display devices (e.g. to allow the presentation to be split between multiple devices such as a television and a tablet).
[0058] The rendering and display / play of the presentation can be performed by modules of the media device 60 itself (as shown in Figure 6(a)), or can be performed by an external media rendering and display / play device using the presentation provided as output from the media player 68 (shown as an alternative selection within the media player 68) to perform one or both of these functions. In another alternative, the presentation determined by the configuration and identification module 62 can be provided to the user as a suggested or default presentation, which the user can simply accept (without having to go through any specific configuration steps), or can further adjust based on their own preferences or their own satisfaction with the automatically provided presentation, which has been personalized to the user as a personal "default" presentation.
[0059] Reference is now made to Figure 6(b), which illustrates the entities that can be involved in performing a method according to an alternative embodiment of the application in a scenario in which the configuration and identification of media objects for a particular user presentation is performed prior to the media content in question being provided to the user's media device.
[0060] In this embodiment, a configuration and identification (C&I) device 600, remote from the user's media device 60 (and possibly co-located with or part of the media content source 50), performs at least some of the functions performed by the media device 60 in the above-described embodiments, in particular, the operations performed by the configuration and identification module 62 in the above-described embodiments with respect to the media content from the media content source 50 prior to providing the media content, including the media content that has been selected and configured (and possibly rendered), to the media device 60. Reference numerals corresponding to those used in Figure 6(a) will be used for entities having corresponding functionality. The functionality of other embodiments and the overall functionality of alternative embodiments will be explained below.
[0061] In an alternative embodiment, shown in Figure 6(b), the C&I device 60 again requests and receives media content in the form of media objects 55 from the media content source 50, the media objects being received via a media content input interface 601. The received media content is passed to a configuration and identification module 602, the functionality of which generally corresponds to that of the configuration and identification module 602 in the above-described embodiments, as will be explained in more detail later.
[0062] The configuration and identification module 602 is in communication with a user information input interface 604 via which information 58 can be received from the user's media device. This can include information such as requests for particular media content (which can be passed to the media content source 50) and can also include information such as parameters relating to the user's display device (e.g. size, shape, resolution, technical capabilities or features) or parameters relating to the user (whether they are visually impaired, hearing impaired, have particular interests in terms of the personalities or types of media content they can request, feedback with respect to received media content and / or other user preferences and / or requirements). As previously mentioned, this information can be actively provided by the user or can be derived from monitoring the user (or users) and / or the environment in which the user and / or media device is located.
[0063] The configuration and identification module 602 is also in communication with a data store 606 in which data relating to matters such as layout rules, constraints and prioritisation factors (to be discussed below) and possibly other types of data can be stored.
[0064] Based on information received via the user information input interface 604 and information retrieved from the data store 606, the configuration and identification module 602 follows a process to be described in detail below, in order to form a presentation of media objects selected from those received via the media content input interface 601, the selected media objects being configured based on the information received and / or retrieved by the configuration and identification module 602, in order to meet or optimize the objective criteria indicative of whether the objective user experience criteria will be met or optimized when the presentation is rendered and displayed to the user in question.
[0065] The presentation can then be provided to the media renderer 607 for rendering, and then to the user's media device 60 via the media output 608 for display or otherwise playing for the user, with the selected media objects each being configured to meet or optimize the objective criteria in question. Alternatively, the presentation, which has not yet been rendered, can then be provided to the user's media device 60 via the media output 608, for rendering by the renderer 67 in the user's media device 60 before display or otherwise playing for the user, with each selected media object being configured to meet or optimize the objective criteria in question. In any case, it will be appreciated that the user's media device 60 receives from the C&I device 600 the (possibly rendered) media content that has been selected and configured.
[0066] The above-described embodiments thus relate to performing a process that can manipulate the scale and / or layout (and possibly other features such as color palette, transparency and / or layering order) of object-based graphics in a presentation on a set of one or more screens (i.e. TV screens or other devices), for example so that these graphics are always visible and legible (as far as this is possible given the applicable constraints associated with the context in question), regardless of the screen size and shape of the device in question and the position of the viewer relative to the device. In the case of complex arrangements of multiple graphical objects, it enables the presentation to be prioritized so that the graphics that are essential for interaction, or that are most important for the viewer in question (e.g. a user with impaired vision, a person with poor hearing, a viewer with a particular interest in a character or aspect of the content in question), are given more prominence in terms of scale and / or layout and / or other features.
[0067] The key advantage of the preferred embodiments is that they can continuously determine the objective value of each media component throughout the TV program or other media content experience, by taking into account the relationship between the properties controlling its rendering and the constraints of the available devices and users. This method creates at any moment a "utility value" for each component, which can then be used to select a set of media objects (i.e. as components of the rendering) and their properties, and where possible, both:
[0068] a) meet predefined thresholds of individual objective quality of experience (QoE); and
[0069] b) meet predefined requirements on the overall rendering, e.g. according to broadcaster / author style guides.
[0070] Reference is also made to Figure 7 A method according to the preferred embodiments is described. This builds on the layout service defined and implemented in the EU funded collaborative project 2-IMMERSE (http: / / www.2-immerse.eu / ), www.2immerse.eu and subsequently released as open source code within the 2-IMMERSE GitHub organization (under Apache 2 license) (https: / / github.com / 2-Immerse / layout-service). https: / / github.com / 2-IMMERSE / layout- service
[0071] As defined in D2.2 (Platform - Component Interface Specification) deliverable of 2-IMMERSE and https: / / 2immerse.eu / wiki / layout /
[0072] The layout service is responsible for managing and optimizing the rendering of a set of DMApp components [media objects] across a set of participating devices (i.e. a context).
[0073] The resources exposed by the layout service through its API are:
[0074] • Context - one or more connected devices cooperating together to render a media experience
[0075] • DMApp (Distributed Media Application) - a set of software components that can be flexibly distributed across multiple participating multi-screen devices. DMApps run in a context.
[0076] • Component - a DMApp software component
[0077] For a running DMApp (including a set of media objects / DMApp components that changes over time); given its authored layout requirements, user preferences and the set of participating devices (and their capabilities) in context, the layout service will determine the optimal layout of components for that configuration. The layout can not be able to accommodate the presentation of all available components simultaneously.
[0078] The service instance maintains a model of the participating devices (context) and their capabilities, e.g. video: screen size, resolution, color depth, audio, number of channels, interaction: touch, etc.
[0079] The layout requirements will specify for each media object / DMApp component: layout constraints, such as minimum / maximum size, audio capabilities, interaction support, and whether the user can override these constraints. Some of these constraints can be expressed relative to other components (priority, position, etc.).
[0080] The layout model that the layout service will employ will be determined, but there is a range of options from very simple (a single component is displayed full screen with a simple selector) to a non-overlapping grid-based arrangement, overlapping models (such as picture-in-picture), to full 3D composition of arbitrary shaped components.
[0081] In particular, the present embodiment can be seen as changing the concept of applying fixed constraints to media objects as part of the authoring process by providing a technology that systematically expresses and evaluates complex and dynamic relationships between a set of constraints and "presentation variables" that define how a media object is presented on a device.
[0082] With respect to this embodiment, the following terms are defined:
[0083] "Presentation" is the rendering of a personalized object-based experience for one or more users on one or more devices that provide audio and / or video playback capabilities as well as the possibility of user interaction (e.g. mouse, keyboard, touch, voice). The presentation is the result of a process that continuously determines the media content of the object-based experience as well as how the media content should be presented; it is the finished product that is seen / consumed by one or more users.
[0084] A "context" is the name given to describe the set of devices available at any moment to one or more users in question in order to render a personalized object-based presentation for the one or more users in question (typically viewers, but one or more users can be only listeners). Thus, a "context" is made of the arrangement of one or more media devices on which a presentation can be rendered (and indeed displayed / played) in association with one or more user-related features and / or attributes, examples of which are given below. As will be appreciated, a context can thus change constantly or continuously as devices become available or are removed at any time, and also if one or more user-related features and / or attributes change. As will be explained, examples of user-related features and / or attributes include the identity or number of users / viewers, their position (or distance from the display device) relative to the display device, and stored topics of interest for a particular user / viewer.
[0085] The "arrangement" of one or more users in question and one or more such user-related features and / or attributes together constitute the "context" of the presentation in question.
[0086] A "media object" is a component rendered on a device within a "context" as part of a "presentation". Media objects include audio streams, video streams, textual content and on-screen graphics.
[0087] A "presentation variable" or "attribute" PV is an attribute of a media object that defines one aspect of the rendering of the media object on a device. Presentation variables can include the physical dimensions on which a graphic will be rendered on a screen; whether the presentation should include audio description, subtitles or sign language actors; how a scaled image should be centered on the screen; the color chosen for a graphic; the volume level, etc.
[0088] A "presentation constraint" c is a property of the "context" in which a "presentation" is being delivered and can be a continuous or categorical variable. Constraints can include:
[0089] - a categorical variable that indicates the quality of vision of the user, for example {unimpaired, low vision, blind}.
[0090] - a continuous variable that defines parameters such as the size, aspect ratio and pixel resolution of the device.
[0091] - a categorical variable that indicates the functional capabilities of the device, such as the type of interaction supported on the device, for example {none, single-point touch, multi-point touch, remote control buttons}, etc.
[0092] - a categorical variable that indicates the functional capabilities of the device, such as the type of interaction supported on the device, for example {none, single-point touch, multi-point touch, remote control buttons}, etc.
[0093] The concept of "utility" is defined as an objective measure of the degree to which a user can understand and interact with the experience. The "utility value" u is a value that indicates the contribution of a media object to the overall "utility" of the "context" when faced with a specific value (or multiple values) of a "presentation variable" (or multiple variables).
[0094] The "priority factor" w is a numerical scaling factor that can be used to assign weights to a specific "constraint" when used as part of the "utility value" calculation.
[0095] Used to combine media objects i and constraints c j and presentation variable PV k utility value u i,j,k It can be represented as:
[0096] u i,j,k =f(c j ,PV k )×w j
[0097] f(c j ,PV k The function can be a continuous function or a combination of discrete functions. For example, in Figure 8 The diagram illustrates how utility values for rendering values (such as the size of graphics rendered on the screen) can be calculated based on different constraint values according to the user's visual quality. For users with unimpaired vision, the step function C0 indicates zero utility for attribute values up to a certain threshold, but maximum utility for attribute values above that threshold. For users with low vision, the tilt function C1 indicates zero utility for attribute values up to a certain threshold, then increases to its maximum as the attribute value increases. For blind users, the function C2 indicates zero utility regardless of the value of the attribute in question (e.g., regardless of the size of the graphic object, which would not benefit the user in question).
[0098] These features can be determined per object or for a group of objects, and enable the incorporation of production decisions based on the context in question, such as the minimum (and / or maximum) acceptable size of graphics on the screen.
[0099] The above allows the utility u of media object i to be calculated using the following function. i :
[0100]
[0101] Therefore, the "overall utility" U of a set of media objects in a specific context is:
[0102]
[0103] It is important to note that this utility calculation depends on the time at which it is performed, and that the overall utility U can change when media objects are added to or removed from a presentation, and that the overall utility U can also change when devices are added to or removed from a context.
[0104] The "layout model" (constructed according to the above definition) is a set of rules that are evaluated within the "context" for all media objects to be rendered. The layout model constrains how a set of media objects and their selected presentation variables can be combined to create a presentation according to the broadcaster / author's style guide. For example, rules can be used in order to:
[0105] - define areas in which certain media object types can be displayed
[0106] - define minimum space between graphical media objects on the screen (thus preventing occlusion)
[0107] - define a maximum number of objects to be displayed simultaneously (to avoid complexity)
[0108] On this basis, Figure 7 The illustrated process can be used to create and maintain an optimal object-based presentation, while accommodating a set of dynamic constraints.
[0109] As indicated in the description of the embodiments provided earlier with reference to Figures 6(a) and 6(b), the processing of media content to be rendered as a presentation for a user according to the preferred embodiments can be performed by (or in association with) the modules 62 of the user's media device 60, or by a device such as the C&I device 600 illustrated in Figure 6(b) (which can be co-located with the media content source 50 or located elsewhere), the selected and configured media objects then being rendered before or after being provided to the user's media device 60. Thus, Figure 7 The process illustrated in Figure 6(c) indicates the basic steps that can be involved in processing media content to be rendered as a presentation for a user according to the preferred embodiments, whether these steps are performed by (or in association with) the modules 62 of the user's media device 60, or by a device such as the C&I device 600 illustrated in Figure 6(b), and excluding additional steps that can occur before and / or after the steps illustrated, in part to avoid over-complicating the flowchart, and in part because the nature of these additional steps is generally dependent on the overall type of embodiment. Such additional steps (e.g. preliminary steps such as initial request and initial provision of media content in the form of media objects; and subsequent steps such as, once a set of media objects has been selected and configured for a user, displaying the content) have been discussed earlier in connection with the figures illustrating particular exemplary embodiments. Thus, Figure 7 The steps illustrated can occur in different ways before and / or after the steps illustrated, in part to avoid over-complicating the flowchart, and in part because the nature of these additional steps is generally dependent on the overall type of embodiment. Such additional steps (e.g. preliminary steps such as initial request and initial provision of media content in the form of media objects; and subsequent steps such as, once a set of media objects has been selected and configured for a user, displaying the content) have been discussed earlier in connection with the figures illustrating particular exemplary embodiments. Thus,Figure 7 Starting at the point where an entity configured to perform a method according to an embodiment has received media content in the form of media objects, in the absence of processing according to the illustrated procedure, that media content would typically be simply rendered for display in a default layout or presentation suggested by the content provider or producer for all users, or rendered for display in a layout or presentation that each user can first need to configure by scratch.
[0110] The illustrated procedure starts with receiving (at step s70) a layout change trigger. This can be a request from a user to incorporate one or more additional media objects into the user's presentation (or an indication from a user that one or more media objects are to be removed from the user's existing presentation), or an indication that the user has started or is about to start using one or more additional screens (e.g. a tablet device supplementing content being displayed on a television) and wishes to transfer one or more media objects (e.g. a leaderboard, a camera-feed following a particular player in a sporting event) to the additional screen (or wishes to stop using one or more additional screens, etc.). Alternatively or additionally, the layout change trigger to incorporate one or more additional media objects into the user's presentation can come from the broadcaster. This can be because the broadcaster wishes to add a new bottom third graphic to indicate a target score, which supersedes an existing user-requested graphic, for example, due to layout rules or priorities. Another option is that the layout change trigger can be based on an indication of context, such as an indication that at least one user has poor hearing (so media objects displaying subtitles are needed) or impaired vision (so media objects displaying text need to be presented larger or using more legible colour schemes), or can be a determination that a user has moved closer to or further from a display and so can benefit from a presentation in which a primary media object or text-based media object occupies more or less of the entire screen area. Other types of layout change trigger are also possible.
[0111] At step s71, the entity performing the procedure checks the constraint values c j and updates the constraint values as appropriate with respect to the current context for the presentation, where the context for the presentation includes the current arrangement of one or more media devices associated with the current user (or users) viewing the presentation. Thus, the constraints can relate to, for example, the number of displays used, and their size, shape and capabilities, or characteristics of the user. Each constraint c j defines a characteristic of the context that can affect the manner in which at least one subset of the selected media objects should be rendered to best meet the objective quality of the experience criteria.
[0112] At step s72, the entity performing the process can check one or more prioritization factors and update them if appropriate, based on user preferences or other factors that can influence the weight w j to be used with respect to a particular constraint. Thus, the prioritization factors can relate to factors that the user has indicated are important (such as the presentation of a leaderboard that is to be displayed at a particular location in the home screen and that covers less than a sixteenth of the screen area, or the presentation of a leaderboard that is to be displayed on a tablet (i.e. a "second screen"), or the presentation of a video input that is to be displayed in a particular off-center portion of the home screen and that is focused on a particular player). The "prioritization factor" w is a numerical scaling factor that can be used to assign a weight to a particular "constraint" when used as part of the calculation of the "utility value" described below.
[0113] At step s73, the entity performing the process determines a presentation value PV k for each media object (i.e. those that have been or can be included in the presentation) that maximizes the utility value u i of the media object in question. As explained previously, "utility" is an objective measure of the extent to which a user can understand and interact with an experience, and the "utility value" u of a media object is a value that indicates the contribution that the media object makes to the overall "utility" of the "context" when faced with one or more presentation variables at a particular value (or values). The result is a list of utility values for the various media objects in the current context, allowing the media objects to be ordered with respect to the current context. Data such as presentation variables, utility values, lists of media objects, etc. can be temporary and can be (temporarily) saved in the configuration and identification module 62.
[0114] It should be noted that in some embodiments, the utility value u can be evaluated using a utility function that depends on only one presentation variable and / or only one constraint. However, in other embodiments, the utility value can be evaluated using a utility function that depends on multiple presentation variables and / or multiple constraints.
[0115] At step s74, the entity performing the process determines which media object (in the current context) has the highest utility value, and then adds that media object to a list of media objects that can be rendered. (The list can initially include one or more default media objects with a default presentation value, or can start with no media objects and build from there).
[0116] At step s75, reference is made to the stored layout rules to determine whether, given the current situation, the list of selected media objects can still satisfy the applicable layout rules. As explained previously, layout rules can limit how a set of media objects and their selected presentation variables can be combined to create a presentation according to the broadcaster / author's style guide or according to declared user preferences. They can define areas in which certain media object types can or cannot be displayed, or define minimum space between graphical media objects on the screen (thus preventing occlusion), or define a maximum number of objects to be displayed simultaneously (to avoid complexity). Other types of layout rules are also possible.
[0117] If the applicable layout rules can still be satisfied with the additional object added to the list, the process proceeds to step s76, where it is determined whether further media objects are available for possible selection. If so, at step s77, the media object with the next-highest utility value (in the current context) is identified and added to the list of media objects that can be rendered, and the process returns to step s75, where reference is again made to the stored layout rules to determine whether the applicable layout rules can still be satisfied.
[0118] If it is found at step s75 that adding another media object to the list of media objects that can be rendered makes it impossible to satisfy the applicable layout rules, the process proceeds to step s78, where the last media object to be added to the list is removed. The process then proceeds to step s79, where the presentation is complete, ready to be rendered. The presentation can then be rendered to make it ready for display, or can be provided to an entity that is to render the presentation and make it ready for display.
[0119] Accordingly, if it is found at step s76 that no further media objects are available for possible selection, the process proceeds to step s79, where the presentation is complete, ready to be rendered. As set out in the preceding paragraph, the presentation can then be rendered to make it ready for display, or can be provided to an entity that is to render the presentation and make it ready for display.
[0120] According to the above procedure, a presentation is thus prepared which includes as many media objects as possible (each configured to have its maximum utility in the current context) without making it impossible to comply with the applicable layout rules, optionally also taking into account specific user preferences if any have been provided. An alternative to this is to provide the determined presentation as an initial suggested presentation to the user (based on the current context) and then allow the user to request changes to it by specifying user preferences (which can then be considered as layout change triggers or as direct commands), by directly requesting changes to it or otherwise.
[0121] As mentioned previously, additional types of data are included within the data store (66 in Fig. 6(a), 606 in Fig. 6(b)). In some embodiments, a history of media object lists, presentation variables and utility values can be stored, allowing such "history" data to be used as part of the decision making process in order to "smooth" changes in the user experience and avoid possible instability (i.e. several changes at once, or a presentation oscillating between several states). Alternatively, such issues can be handled by appropriate design of the utility functions and layout rules.
[0122] According to other embodiments, other procedures can be used in order to prepare a suggested presentation for the user in the current context. Instead of the procedure shown (where additional objects are added one by one to the list of media objects that can be rendered until it becomes impossible to comply with the applicable layout rules, an alternative optimization-based approach can be used to select the set of objects that has the highest sum of utility values overall (without violating the applicable layout rules). This can result in selecting and configuring a presentation that does not include the media object with the highest individual utility value, if so, allowing to include several other media objects that cannot be included next to the media object with the highest individual utility value. In some cases, this can be Figure 7 the preferred procedure shown.
[0123] In preferred embodiments, one or more utility functions are typically selected in order to ensure consistency between the utility functions and the layout model rules, ensuring that the above procedure (or a similar procedure) does not result in undesirable situations, such as situations where no media objects are selected for presentation. Other techniques can be used to ensure that there is always at least one media object selected. A default media object can initially be included in the set of selected media objects, with rules ensuring that if the default media object would be removed from the list as a result of the execution of the procedure, this can only happen if the total number of selected media objects is still at least one.
[0124] Where the described implementations of the application can be implemented, at least in part, using a software-controlled programmable processing device such as a microprocessor, digital signal processor or other processing device, data processing apparatus or system, it will be appreciated that a computer program for configuring a programmable device, apparatus or system to implement the foregoing described methods is envisaged as an aspect of the present application. The computer program can be implemented in, for example, a high level procedural or object-oriented programming language, and / or in assembly or machine language. Furthermore, the software can be
[0125] Suitably, the computer program is stored on a carrier medium in machine or device readable form, for example in solid-state memory, magnetic or optical storage media such as diskette or magnetic disc; optical storage media such as compact disc read-only memory (CD-ROM); or in a ROM, for use by, or in connection with, the processing equipment. The processing equipment can be a computer, terminal, mobile terminal or apparatus. Alternatively, the processing equipment can be configured to be capable of operating in accordance with the computer program. The computer program can be supplied from a remote source embodied in a communications medium such as an electronic signal, radio frequency carrier wave or optical carrier wave. Such carrier media are also envisaged as aspects of the present application.
[0126] Those skilled in the art will appreciate that, although the application has been described in relation to the above exemplary embodiments, the application is not limited thereto and that there are many possible variations and modifications which fall within the scope of the application.
[0127] The scope of the application can include other novel features or combinations of features disclosed herein. Accordingly, Applicant notes that new claims can be formulated to any such features or combinations of features during the prosecution of this application or any such further applications derived therefrom. In particular, with reference to the appended claims, features from dependent claims can be combined with those of the independent claims and features of respective independent claims can be combined in any appropriate manner in, without being limited to only the specific combinations enumerated in the claims.
Claims
1. A computer-implemented method for processing media content, said media content being rendered for a user presentation at a point in time across a group of one or more media devices, said presentation being based on layout rules defining the suitability and configuration of media objects rendered as part of said presentation, wherein, The layout rules include rules for determining whether, when, where, and how to render the corresponding media objects. The layout and one or more user-associated features and / or attributes constitute a context for the rendering, wherein the rendering is formed by media objects selected from a set of media objects, and the context has one or more associated constraints, each constraint defining the characteristics of the context that influence the rendering of at least one subset of the selected media objects. The method includes the following steps: Configure features for each media object in the set of media objects, wherein the configured features conform to a utility condition based on a utility metric of the media object in the context at the specified time point, the utility metric being estimated as a function of the constraints of the context and attributes defining the aspects of the media object's presentation on the media device; and The selected media objects in the set of media objects are identified based on the utility metric associated with each selected media object and the layout rules, wherein the media objects are added to a list of media objects to be rendered based on a utility value evaluated with respect to the media object, until it is determined that they cannot comply with the applicable layout rules. Adding the media object to the list of media objects to be rendered includes: Determine which media object has the highest performance value and then add the media object with the highest performance value to the list of media objects to be rendered; Refer to the layout rules to determine whether the list of selected media objects still conforms to the applicable layout rules; and When it is determined that the list of selected media objects still conforms to the applicable layout rules, the media object with the second efficiency value is determined and then the media object with the second efficiency value is added to the list of media objects to be rendered.
2. The method according to claim 1, wherein, The set of media objects includes media objects that provide one or more of video content, audio content, text content, and graphic content.
3. The method according to claim 1, wherein, The characteristics of a media object include one or more of the object's graphic size, screen position, color scheme, transparency, and layering order.
4. The method according to claim 1, wherein, The group of one or more media devices arranged for a specific point in time includes more than one media device arranged in one manner.
5. The method according to claim 4, wherein, The media object is characterized by a specific media device on which the media object should appear in the arrangement of one or more media devices in the group.
6. The method according to claim 1, wherein, The step of identifying the selected media objects in the set of media objects is performed by identifying media objects such that the sum of the utility values evaluated for the media objects is maximized without violating applicable layout rules.
7. The method according to claim 1, wherein, The configuration and identification steps are performed before the selected media object is transmitted to one or more client media devices.
8. The method according to claim 1, wherein, The configuration and identification steps are performed after the set of media objects has been transferred to one or more client media devices.
9. The method according to claim 1, wherein, The method further includes rendering the selected media object from the set of media objects.
10. The method according to claim 9, wherein, After the selected media object is transmitted to one or more client media devices, the rendering of the selected media object is performed.
11. The method according to claim 1, wherein, The method further includes: providing the selected media object from the set of media objects for presentation via one or more client media devices.
12. An apparatus for processing media content, the media content to be rendered for a user at a point in time across a group of one or more media devices arranged as part of the presentation, the presentation being based on layout rules defining the suitability and configuration of media objects rendered as part of the presentation, the arrangement and one or more user-associated features and / or attributes constituting the context of the presentation, wherein... The rendering is formed by media objects selected from a set of media objects, and the context has one or more associated constraints, each constraint defining the characteristics of the context that influence the rendering of at least a subset of the selected media objects. The apparatus includes a computer system comprising a processor and memory storing computer program code for executing computer program code, such that when the computer program code is executed by the processor, the apparatus is at least configured to: For each media object in the set of media objects, the features of the media object are configured to conform to a utility condition based on a utility metric of the media object in the context at the point in time, the utility metric being estimated as a function of the constraints of the context and the attributes of the media object that define its presentation on the media device. as well as The selected media objects in the set of media objects are identified based on the utility metric associated with each selected media object and the layout rules, wherein the media objects are added to a list of media objects to be rendered based on a utility value evaluated with respect to the media object, until it is determined that they cannot comply with the applicable layout rules. In order to add the media object to the list of media objects to be rendered, the device is further configured to: Determine which media object has the highest performance value and then add the media object with the highest performance value to the list of media objects to be rendered; Refer to the layout rules to determine whether the list of selected media objects still conforms to the applicable layout rules; and When it is determined that the list of selected media objects still conforms to the applicable layout rules, the media object with the second efficiency value is determined and then the media object with the second efficiency value is added to the list of media objects to be rendered.
13. A non-transitory computer-readable storage medium storing computer program code that, when loaded into and executed on a computer system, causes the computer to perform the steps of the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Optimizing electronic layouts for media content
US9569501B2
Method for the combined broadcasting of a television programme and an additional multimedia content
CN111095943A
Optimizing Electronic Layouts for Media Content
US20150019545A1
Audio Authoring and Compositing
US20190180785A1