Method and device for delivering and displaying video streams
The method addresses the challenges of adapting audiovisual content to diverse devices and networks by detecting regions of interest and transmitting data through separate channels, ensuring high-quality and synchronized display on OTT platforms.
Patent Information
- Application Number
- PCT/EP2025/070378
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-18
- Filing Date
- 2025-07-16
- Publication Date
- 2026-01-22
AI Technical Summary
The distribution of audiovisual content on OTT platforms faces challenges in adapting to varying screen sizes, managing network resources efficiently, and ensuring interoperability between different devices and formats, leading to video quality degradation and user experience issues.
A method for broadcasting a video stream that detects regions of interest within the stream, adapts the content to target devices, and transmits data through separate channels for audiovisual and graphic elements, allowing for synchronized display without post-production overlays.
The method enables high-quality video adaptation to diverse devices, optimizes bandwidth usage, and ensures smooth streaming with synchronized display of audiovisual and graphic elements, enhancing user experience.
Smart Images

Figure EP2025070378_22012026_PF_FP_ABST
Abstract
Description
Method and device for broadcasting and displaying video streams technical field
[0001] The present invention relates to the distribution and display of video streams and more particularly to an OTT service enabling the display of video streams on different types of devices. Technological background
[0002] The technical field of the present invention relates to OTT services (Over-The-Top service, or services bypassing the Internet service provider's offer). OTT services enable the distribution of digital media, such as audiovisual content, via the Internet, independently of traditional cable or satellite television operators.
[0003] The current landscape of video streaming on OTT platforms is characterized by a growing diversity of devices used by users to view audiovisual content, such as televisions, smartphones, tablets, and computers. Each device has its own unique characteristics in terms of screen size, image resolution, processing power, and network access. This heterogeneity presents significant challenges for content providers, who must guarantee a high-quality user experience regardless of the viewing platform used or the network access conditions.
[0004] One of the technical problems to solve is adapting audiovisual content to screens of varying sizes without compromising the user experience. High-definition televisions, computer monitors, tablets, and smartphones have different screen resolutions and aspect ratios (the ratio between the screen's width and height), requiring sophisticated transcoding and resizing techniques.
[0005] Current transcoding and resizing solutions can lead to video quality degradation, image distortion, or prolonged loading times, negatively impacting the user experience.
[0006] Network resource management presents another challenge. Streaming audiovisual content requires significant bandwidth, and variations in network conditions can affect the quality of the received video stream. Stream interruptions, resolution drops, and high latency are common problems when audiovisual content is streamed over unstable or congested networks. Existing technologies need to be improved to optimize bandwidth usage and ensure smooth streaming, even under fluctuating network conditions.
[0007] Another critical aspect is interoperability between different platforms and video formats. Content providers must ensure that their audiovisual content is compatible with a multitude of operating systems, browsers, and applications, each with its own specifications and constraints.
[0008] In summary, the distribution of audiovisual content on OTT platforms faces significant technical challenges related to adapting content to a variety of screens, efficiently managing network resources, and personalizing the user experience. Summary of the present invention
[0009] One object of the present invention is to solve at least one of the problems of the technological background described above.
[0010] According to a first aspect, the present invention relates to a method for broadcasting a video stream without post-production graphic overlays and originating from an audiovisual production. The video stream has a source audiovisual composition defining at least one audiovisual content, a spatial definition of each audiovisual content within a frame of the video stream, and a display format for each audiovisual content. The method comprises the following steps: obtaining at least one audiovisual composition model based on the source audiovisual composition, each audiovisual composition model defining a target audiovisual composition, each audiovisual composition model graphically defining a display area of the target audiovisual composition and spatially defining at least a part of the display area dedicated to displaying audiovisual content from the source audiovisual composition in a display format different from that defined for said audiovisual content by the source audiovisual composition; - detection of at least one area of interest in an image of the video stream based on the source audiovisual composition; - determination of initial data representative of a spatial definition of each area of interest detected in the video stream image; - determination of second data for each audiovisual composition model defining a target audiovisual composition, the second data, determined for an audiovisual composition model, graphically defining the display area of the target audiovisual composition; a part of the display area for each detected area of interest and an audiovisual content display format for each part of the display area; - obtaining third data points representative of the video stream; - dissemination of the first data, the second data determined for each audiovisual composition model, and the third data representative of the video stream through a first communication channel, - dissemination of fourth data points representing at least one post-production graphic element associated with the video stream via a second communication channel distinct from the first communication channel; and - broadcasting synchronization data from a clock signal through the first communication channel in relation to the first, second and third data and the same synchronization data through the second communication channel in relation to the fourth data.
[0011] The method obtains at least one audiovisual composition model based on the source audiovisual composition of the video stream. Furthermore, the method detects at least one region of interest within a frame of the video stream based on the source audiovisual composition. Furthermore, the process determines first data representing a spatial definition of each area of interest detected in the image of the video stream and, for each audiovisual composition model, second data defining a target audiovisual composition, said second data graphically defining the display area of the target audiovisual composition; a part of the display area for each detected area of interest and an audiovisual content display format for each part of the display area.In addition, the process disseminates first data, second data determined for each audiovisual composition model and third data representative of the video stream through a first communication channel, disseminates fourth data representative of at least one post-production graphics element associated with the video stream through a second communication channel distinct from the first communication channel; and disseminates synchronization data from a clock signal through the first communication channel in relation to the first, second and third data and the same synchronization data through the second communication channel in relation to the fourth data.
[0012] The process allows for the transmission of a video stream without post-production graphic overlays, which can be displayed (viewed) as audiovisual compositions formed from at least a portion of the transmitted video stream. Each portion of the video stream is derived from a region of interest that is automatically detected within a frame of the video stream and positioned within the display area according to the spatial location of this region of interest within the display area.
[0013] The process is advantageous because it allows for the definition of a target audiovisual composition that can be adapted to the characteristics of a target device, such as the screen dimensions of the device on which the target audiovisual composition is displayed, or the orientation of that display screen. The adaptation can relate to the resolution of the video stream segments that will be displayed in the target audiovisual composition, their number, or their arrangement within the display area defined by each audiovisual composition template.
[0014] The process is a solution to the problem of heterogeneity of devices intended to display a target audiovisual composition formed from the video stream.
[0015] The process is advantageous because it does not significantly increase the bandwidth required to broadcast the video stream, since the first and second additional data represent only a small percentage of the third data broadcast, which represents the video stream.
[0016] The process is advantageous because each audiovisual composition template can then define a target audiovisual composition combining both parts of the video stream and one or more post-production graphic elements such as logos or banner ads. The resolution and dimensions of these post-production graphic elements are adapted to each audiovisual composition template; that is, the post-production graphic elements specific to each audiovisual composition template are defined in resolution and dimensions from the moment they are generated.
[0017] The process is advantageous because it allows the fourth data point to be transmitted without disrupting the transmission of the first, second, and third data points. Thus, some target audiovisual compositions may be formed solely from the first, second, and third data points, while others may also include fourth data points. Whether or not the fourth data point is used to form a target audiovisual composition may depend on the communication capabilities of a target device or its permissions to access the content represented by this fourth data point.
[0018] The process is advantageous because it allows the display of audiovisual content from different parts of the target audiovisual composition to be synchronized.
[0019] According to a particular and non-limiting embodiment of the present invention, the method may further include a step of detecting scene changes in the video stream, and wherein the steps of obtaining at least one audiovisual composition model, detecting at least one area of interest, determining first, second, third and fourth data and broadcasting are executed following the detection of each scene change.
[0020] This embodiment is advantageous because it allows the audiovisual composition model and regions of interest to be adapted to a new scene in the video stream whose source audiovisual composition and / or regions of interest may differ from those of the previous scene. According to a second aspect, the present invention relates to a method for displaying a video stream on a display screen. The method receives first data representing a spatial definition of at least one region of interest in an image of the video stream and, for at least one audiovisual composition model, receives second data defining a target audiovisual composition. The second data received for each audiovisual composition model graphically defines a display area of the target audiovisual composition, and for each region of interest,a portion of the display area and a display format for audiovisual content from said portion of the display area. Furthermore, the method receives, via the first communication channel, third data representing the video stream without post-production graphic overlays; selects an audiovisual composition model from among at least one audiovisual composition model defined by the second data received; obtains from the display area defined by the second data received the selected audiovisual composition model; for each region of interest spatially defined by the first data received, the method obtains a portion of audiovisual content from the video stream delimited by the region of interest; and adds the portion of audiovisual content obtained to a portion of the display area defined for the region of interest by the second data received for the selected audiovisual composition model.depending on the display format defined for said part of the display area by the second data received for the selected audiovisual composition model; receives, through a second communication channel distinct from the first communication channel, fourth data representing at least one post-production graphic element associated with the video stream; adds said at least one post-production graphic element to at least a part of the resulting display area; receives synchronization data from a clock signal through the first communication channel in relation to the first, second, and third data and the same data from, Synchronization via the second communication channel in relation to the fourth data point; controls a synchronized display, according to the synchronization data, on the display screen, of the audiovisual content and / or the post-production graphic overlay element of each part of the display area obtained according to a display format defined for said part of the display area by the second data point received for the selected audiovisual composition model
[0021] The process offers a solution to a user to personalize their display experience by giving them the ability to select a target audiovisual composition from among several target audiovisual compositions.
[0022] The process makes it possible to maintain synchronization of the audiovisual content of the video stream without post-production graphic elements when this video stream and its post-production graphic elements are broadcast by two separate communication channels.
[0023] According to a particular and non-limiting embodiment of the present invention, the audiovisual composition model is selected from said at least one audiovisual composition model from a graphical interface or according to a display screen orientation or a preferred display format.
[0024] This embodiment allows either manual selection of a target audiovisual composition or automatic selection based on the orientation of the target device. This embodiment is particularly advantageous for mobile communication devices such as smartphones or tablets because it allows for automatic switching of the target audiovisual composition depending on the device's orientation. For example, when the device is horizontal, a first target audiovisual composition can be viewed. This first target audiovisual composition can consist of a large horizontal portion of the display area in which the wide-angle video stream is viewed.When the device is vertical, a second target audiovisual composition can be viewed. This second target audiovisual composition can be formed from one or more more tightly enclosed parts of the display area in which or which. is / are viewed one or more parts of the video stream corresponding to one or more regions of interest.
[0025] This implementation is advantageous because it allows a user to save a preferred video stream display format when their device offers a choice of several formats. For example, a display area aspect ratio can be saved in a smartphone's user preferences, and the selected target audiovisual composition will correspond to a display area with an aspect ratio equal to (or as close as possible to) the saved one.
[0026] According to a particular and non-limiting embodiment of the present invention, the audiovisual composition model is selected from said at least one audiovisual composition model based on video stream display capacity and / or based on data reception capabilities.
[0027] This example of implementation is advantageous because it allows for latency-free and stutter-free display of the target audiovisual composition, optimizing the user experience.
[0028] According to a particular and non-limiting embodiment of the present invention, the display format defined for a part of the display area by the second received data for the selected audiovisual composition model defines a cropping of the audiovisual content of said part of the display area according to the dimensions of said part of the display area and the dimensions of said audiovisual content.
[0029] This example of implementation is advantageous because it allows the dimensions of a part of the video stream corresponding to a region of interest defined in the video stream to be adapted to the dimensions of a part of the display area of a target audiovisual composition.
[0030] According to a third aspect, the present invention relates to a device for broadcasting a video stream, the device comprising a memory associated with a processor configured for the implementation of the steps of the process according to the first aspect of the present invention.
[0031] According to a fourth aspect, the present invention relates to a device for displaying a video stream, the device comprising a memory associated with a processor configured for the implementation of the steps of the process according to the second aspect of the present invention.
[0032] According to a fifth aspect, the present invention relates to a system for broadcasting and displaying a video stream comprising a device according to the third aspect of the present invention and at least one device according to the fourth aspect of the present invention.
[0033] According to a sixth aspect, the present invention relates to a computer program which includes instructions adapted for carrying out the steps of the process according to the first and / or the second aspect of the present invention, in particular when the computer program is executed by at least one processor.
[0034] Such a computer program can use any programming language, and be in the form of source code, object code, or an intermediate form between source code and object code, such as in a partially compiled form, or in any other desirable form.
[0035] According to a seventh aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the process according to the first and / or second aspect of the present invention.
[0036] On the one hand, the recording medium can be any entity or device capable of storing the program. For example, the medium can include a storage means, such as ROM, RAM, CD-ROM or microelectronic circuit-type ROM, or a magnetic recording means or a hard drive.
[0037] On the other hand, this recording medium can also be a transmissible medium such as an electrical or optical signal, such a signal being able to be transmitted via an electrical or optical cable, by conventional or Hertzian radio, by self-directing laser beam, or by other means. The computer program according to the present invention can, in particular, be downloaded onto an Internet-type network.
[0038] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to execute or to be used in the execution of the process in question. Brief description of the figures
[0039] Other features and advantages of the present invention will become apparent from the description of the specific and non-limiting embodiments of the present invention below, with reference to the attached Figures 1 to 9, in which:
[0040] [Fig. 1] schematically illustrates an example of a video streaming system according to a particular and non-limiting embodiment of the present invention.
[0041] [Fig. 2] schematically illustrates a device configured for the implementation of an OTT service for broadcasting a video stream according to a particular and non-limiting embodiment of the present invention.
[0042] [FIG. 3] illustrates a flowchart of the different stages of the video stream broadcasting process according to a particular embodiment of the present invention.
[0043] [FIG. 4A] illustrates an example of an audiovisual composition model corresponding to a source audiovisual composition, according to a particular and non-limiting embodiment of the present invention.
[0044] [FIG. 4B] illustrates an example of an audiovisual composition model corresponding to a source audiovisual composition, according to a particular and non-limiting embodiment of the present invention.
[0045] [FIG. 4G] illustrates an example of an audiovisual composition model corresponding to a source audiovisual composition, according to a particular and non-limiting embodiment of the present invention.
[0046] [FIG. 4D] illustrates an example of an audiovisual composition model corresponding to a source audiovisual composition, according to a particular and non-limiting embodiment of the present invention.
[0047] [FIG. 4E] illustrates an example of an audiovisual composition model corresponding to a source audiovisual composition, according to a particular and non-limiting embodiment of the present invention.
[0048] [FIG. 4F] illustrates an example of an audiovisual composition model corresponding to a source audiovisual composition, according to a particular and non-limiting embodiment of the present invention.
[0049] [FIG. 4G] illustrates an example of an audiovisual composition model corresponding to a source audiovisual composition, according to a particular and non-limiting embodiment of the present invention.
[0050] [FIG. 4H] illustrates an example of an audiovisual composition model corresponding to a source audiovisual composition, according to a particular and non-limiting embodiment of the present invention.
[0051] [FIG. 5] illustrates an example of encapsulation of the first, second and third data according to the MPEG2-TS protocol, according to a particular and non-limiting embodiment of the present invention.
[0052] [Fig. 6] schematically illustrates a device configured to display a video stream broadcast by an OTT service according to a particular and non-limiting embodiment of the present invention.
[0053] [FIG. 7] schematically illustrates a diagram of the steps of a method for displaying a video stream according to a particular and non-limiting embodiment of the present invention.
[0054] [FIG. 8A] illustrates an example of target audiovisual composition of a requested video stream, according to a particular and non-limiting embodiment of the present invention.
[0055] [FIG. 8B] illustrates an example of target audiovisual composition of a requested video stream, according to a particular and non-limiting embodiment of the present invention.
[0056] [FIG. 8C] illustrates an example of target audiovisual composition of a requested video stream, according to a particular and non-limiting embodiment of the present invention.
[0057] [FIG. 8D] illustrates an example of target audiovisual composition of a requested video stream, according to a particular and non-limiting embodiment of the present invention.
[0058] [FIG. 9] schematically illustrates an information system implementing the present invention, according to a particular and non-limiting example of the present invention. Description of examples of achievements
[0059] A method and device for broadcasting and displaying video streams will now be described in what follows with joint reference to Figures 1 to 9. The same elements are identified with the same reference symbols throughout the description that follows.
[0060] In the following description, numerous specific details are set forth to provide a thorough understanding of the present invention. However, it will be evident to those skilled in the art that the present invention, including structures, devices, systems, and methods, can be implemented without all of these specific details. The description and drawings below are means commonly used by persons experienced or qualified in the technical field relevant to the present invention to convey the substance of the present invention as effectively as possible to other persons experienced or qualified in the prior art. In other cases, well-known methods, components, and circuits have not been described in detail so as not to unnecessarily obscure certain aspects of the present invention.
[0061] In the description, references to "an embodiment," "an example of an embodiment," etc., indicate that the described embodiment or example may include a particular feature, structure, or element, but that not every embodiment or example necessarily includes that particular feature, structure, or element. Furthermore, these expressions do not necessarily refer to the same embodiment. Moreover, when a feature, structure, or characteristic is described in relation to an embodiment or example, it is assumed that a person skilled in the art knows how to apply that feature, structure, or characteristic to other embodiments or examples, whether or not they are explicitly described.
[0062] The terms "first(s)", "second(s)" (or "first(s)", "second(s)"), etc. are used in the by arbitrary convention to allow identification and distinction of different elements (such as operations, means, frequency pairs, etc.) implemented in the modes or embodiments described below.
[0063] It is also worth noting that the various modes or embodiments can be described as a process or procedure represented as a flowchart, data flow diagram, structure diagram, or block diagram. Although a flowchart may depict operations as a sequential process or procedure, many operations may be performed in parallel or simultaneously. Furthermore, the order of operations can be changed. A process or procedure ends when its operations are completed, but it may include additional steps not shown in a diagram. A process or procedure can correspond to a method, function, procedure, subroutine, subprogram, etc. When a process or procedure corresponds to a function, its end may be a return of the function to the calling function or the main function.
[0064] The term "computer-readable media" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or transporting instructions and / or data. A computer-readable medium may include a non-transient medium in which data can be stored and which does not include carrier waves and / or transient electronic signals propagating wirelessly or over wired connections. Examples of non-transient media include, but are not limited to, magnetic disks or tapes, optical storage media such as compact discs (CDs) or digital multipurpose discs (DVDs), flash memory, memory, or memory devices. A computer-readable medium may contain machine-executable code and / or instructions that may represent a procedure, function, subroutine, program, routine, subprogram, module, software package, class, or any combination of instructions, data structures, or program statements.A code segment can be coupled to another code segment or to a hardware circuit by transmitting and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., can be transmitted by any appropriate means, including memory sharing, message transmission, token transmission, network transmission, etc.
[0065] Furthermore, the embodiments or examples can be implemented by hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented in software, firmware, middleware, or microcode, the program code or code segments that perform the necessary tasks (e.g., a computer program product) can be stored on computer-readable or machine-readable media. One or more processors can perform the necessary tasks.
[0066] Figure 1 schematically illustrates an example of a video stream distribution system 1 according to a particular and non-limiting embodiment of the present invention.
[0067] The broadcast system 1 includes a video stream capture device 110, a captured video stream processing and broadcast device 120, and a broadcast video stream display device 130.
[0068] Devices 120 and 130 can be connected via a cloud-based information infrastructure (cloud) that enables communication, data exchange, and management of both devices over the internet. This infrastructure can utilize remote servers to provide data processing, storage, and management services, thereby facilitating connection and interaction between devices 120 and 130.
[0069] The 110 device can be a high-definition camera or a high-definition camera system that allows for continuous capture of a high-quality, high-resolution video stream. For example, the captured video stream is displayed in Full HD (1920 x 1080 pixels, 16:9), 4K UHD (3840 x 2160 pixels, 16:9), or 9K UHD (7680 x 4320 pixels, 16:9). The video stream can be captured at a refresh rate of 60 Hz or higher.
[0070] Device 120 is configured to implement one or more processing steps on the captured video stream and to broadcast the processed video stream according to the present invention, as will be seen in detail later.
[0071] As an example, devices 110 and 120 can be implemented on separate pieces of equipment. In this case, devices 110 and 120 are connected by a communication channel within a network. This communication channel can be dedicated to the video stream between device 110 and device 120.
[0072] This example corresponds, for instance, to situations where one or more cameras are distributed throughout a space and a control room coordinates the production of the video feed from other video feeds captured by these cameras. Typically, these situations occur for the live broadcast of events, such as sporting events.
[0073] According to another example, devices 110 and 120 can be implemented in the same equipment such as a television, computer or device mobile communication such as a smartphone, tablet, laptop, etc.
[0074] According to a particular and non-limiting embodiment of the present invention, the broadcasting system 1 can implement an OTT service for broadcasting audiovisual content.
[0075] The OTT service relies on a series of operational modules enabling the implementation of a video stream distribution process 111 of figure 3.
[0076] Figure 2 schematically illustrates a device 120 configured for the implementation of an OTT service for broadcasting a video stream 1 11 according to a particular and non-limiting embodiment of the present invention.
[0077] Device 110 is configured to implement the capture of video stream 111 and device 120 is configured to implement the video stream broadcasting method of Figure 3
[0078] For this purpose, the device 120 includes a module 121 for processing the captured video stream 111 and a module 122 for data broadcasting.
[0079] The 111 video feed is clean of any post-production graphics; that is, the 111 video feed is a video signal without any superimposed post-production graphics. A post-production graphics element is a graphic that can be superimposed on audiovisual content within the clean 111 video feed. A post-production graphics element can be an image of a barcode, an advertising banner, or a logo, such as those seen on video feeds displayed by television channels. A post-production graphics element is not a graphic that is integrated into audiovisual content (called a baked-in or burned-in graphical element).
[0080] The video stream 111 originates from an audiovisual production. The images in the video stream 111 are clean images as captured by one or more cameras, such as the camera of device 110, before the addition of post-production graphics.
[0081] A clean video stream (111) can be used as a working template to allow for the possible addition of post-production graphics. This clean video stream (111) is very often recorded to facilitate its future reuse without constraints related to specific or outdated post-production graphics. Module 121 receives the video stream (111) as input and outputs first data (1211), second data (1212), and third data (1213).
[0082] Module 121 also receives, as input, a clock signal 112 and data 113 and also provides, as output, fourth data 1214 which are broadcast in particular to device 130 via infrastructure 140.
[0083] Module 122 receives the first data 1211, the second data 1212 and the third data 1213 as input and broadcasts data 122i, 1222 and 1223 to, among others, device 130 via infrastructure 140.
[0084] Figure 3 illustrates a flowchart of the different stages of the video stream broadcasting process according to a particular embodiment of the present invention.
[0085] The video stream 1 11 has a source audiovisual composition defining at least one audiovisual content, a spatial definition of each audiovisual content in an image of the video stream and a display format of each audiovisual content.
[0086] The display format of audiovisual content defines display parameters such as the dimensions of a display area in which the visual part of the audiovisual content is displayed, a spatio-temporal resolution, a spatial resolution of the images of the audiovisual content and / or a frame rate and / or parameters related to the number of components of the images displayed or a number of bits per pixel of said images.
[0087] According to an example implementation, the source audiovisual composition of video stream 111 can be selected from a predefined set of source audiovisual compositions.
[0088] According to an example implementation, the source audiovisual composition of video stream 111 can be selected from a content type of video stream 111.
[0089] According to a specific, but not limiting, implementation example, the audiovisual content type of video stream 111 can be selected from a predefined set of audiovisual content types. For example, this predefined set might include a 'TV news' type, a 'TV debate' type, and a 'match' type.
[0090] According to a particular and non-limiting embodiment of the present invention, the type of audiovisual content of the video stream 111 can be obtained by manually selecting one of the types from the predefined set of types.
[0091] For example, a user can use a human-machine interface of device 120 to select the source audiovisual composition and / or the type of audiovisual content of the video stream 1 11.
[0092] According to a particular and non-limiting embodiment of the present invention, the source audiovisual composition of the video stream 111 can be obtained from a memory in which it has been previously stored.
[0093] Each source audiovisual composition in the predefined set of source audiovisual compositions is associated with at least one audiovisual composition model.
[0094] Each audiovisual composition model defines a target audiovisual composition.
[0095] An audiovisual composition template graphically defines a display area for the target audiovisual composition and spatially defines at least a portion of that display area dedicated to displaying audiovisual content from the source audiovisual composition in a display format different from that defined for said audiovisual content by the source audiovisual composition, or dedicated to displaying a post-production graphic element. Audiovisual content from the source audiovisual composition and a post-production graphic element can be displayed in the same portion of the display area by overlaying the post-production graphic element onto the audiovisual content.
[0096] According to a particular and non-limiting embodiment of the present invention, the display area is defined by its dimensions.
[0097] As an example, the display area can be defined by its height and width when the display area is a rectangle.
[0098] According to another example, the display area can be defined by spatial positions in the video stream image of two opposite vertices of a rectangle when the display area is a rectangle.
[0099] The present invention is not limited to a particular shape of the display area or to its representation by the spatial positions of two of its opposite vertices. The display area can have any shape and be, for example, represented by the spatial positions of pixels in an image of the video stream 1 11.
[0100] When the video stream is encoded by a video codec, the display area can also be defined by references to syntax elements defined by that video codec, such as coding units (Coding Units or Coding-Tree Units) as defined in video codecs implementing an MPEG coding standard, such as AVC (ISO / IEC 14496-10 Advanced Video Coding for generic audio-visual services, ITU-T Recommendation H.264, https: / / www.itu.int / rec / T-REC-H.264-202108-P / en), EVC (ISO / IEC 23094-1 Essential video coding), HEVC (ISO / IEC 23008-2 High Efficiency Video Coding, ITU-T Recommendation H.265, https: / / www.itu.int / rec / T-REC-H.265-202108-P / en), or VVC. (ISO / IEC 23090-3 Versatile Video Coding, ITU-T Recommendation H.266, https: / / www.itu.int / rec / T-REC-H.266-202008-l / en).
[0101] Each audiovisual composition model can be represented by a set of at least one parameter.
[0102] According to a particular and non-limiting embodiment of the present invention, an audiovisual composition model can be represented by a set of parameters that defines a display area of a target audiovisual composition and a spatial definition of at least a portion of the display area dedicated to displaying audiovisual content of the composition. audiovisual source according to a display format different from that defined for said audiovisual content by the source audiovisual composition.
[0103] An audiovisual composition model can be defined for specific characteristics of a device's display screen, such as its dimensions (for example, for a particular aspect ratio of the display screen). Typically, possible aspect ratios of a smartphone's display area are 16:9, 1:1, 4:5, 2:3, and 9:16.
[0104] An audiovisual composition model can be defined for particular characteristics of the displayed audiovisual content such as spatio-temporal resolution, spatial resolution of the images of the audiovisual content, frame rate and / or parameters related to the number of components of the displayed images or the number of bits per pixel of said images.
[0105] Audiovisual composition templates can also be defined to display the video stream 111 either in portrait mode or in landscape mode on the display screen of the device 130.
[0106] Figures 4A to 4H illustrate examples of audiovisual composition models according to different source audiovisual compositions, according to particular and non-limiting examples of embodiment of the present invention.
[0107] Figure 4A illustrates an example of an audiovisual composition model corresponding to a source audiovisual composition associated with a 'TV news' type which corresponds to a landscape mode display on the display screen of device 130. This example of an audiovisual composition model defines a first part 31 of a display area 30 located in the upper part and in the center of this display area 30 and a second display area 32 located in the lower part of the display area 30.
[0108] Figure 4B illustrates an example of an audiovisual composition model corresponding to a source audiovisual composition associated with a 'TV news' type, which corresponds to a portrait mode display on the display screen of device 130. This example of an audiovisual composition model defines a first part 35 of a display area 33 located partly in the center of the display area 33 and a second part 34 of the display area 33 located in the lower part of the display area 33.
[0109] Figure 4C illustrates an example of an audiovisual composition model corresponding to a source audiovisual composition associated with a 'TV news' type which corresponds to a landscape mode display on the display screen of device 130. This example of an audiovisual composition model defines the first part 31 of the display area 30 located in the upper part and in the center of the display area 30, the second part 32 of the display area 30 located in the lower part of the display area 30 and a third part 36 located in the upper left of the display area 30.
[0110] Figure 4D illustrates an example of an audiovisual composition model corresponding to a source audiovisual composition associated with a 'TV news' type which corresponds to a portrait mode display on the display screen of device 130. This example of an audiovisual composition model defines a first part 37 of the display area 33 located in the center of the display area 33, the second part 34 of the display area 33 located in the lower part of the display area 33 and a third part 38 located at the top and left of the display area 33.
[0111] Figure 4E illustrates an example of an audiovisual composition model corresponding to a source audiovisual composition associated with a 'televised debate' type which corresponds to a landscape mode display on the display screen of device 130. This example of an audiovisual composition model defines a first part 39 of the display area 30 located in the upper part and in the centre of the display area 30 and a set of several parts 40i to 404 (here 4) of the display area 30 located in the lower part of the display area 30.
[0112] Figure 4F illustrates an example of an audiovisual composition model corresponding to a source audiovisual composition associated with a 'televised debate' type, which corresponds to a portrait mode display on the display screen of device 130. This example of an audiovisual composition model defines a first part 41 of the display area 33 located in the upper part and in the centre of the display area 33 and a set of several parts 42i to 424 (here 4) of display area 33 located in the lower part of display area 33.
[0113] Figure 4G illustrates an example of an audiovisual composition model corresponding to a source audiovisual composition associated with a 'match' type which corresponds to a landscape mode display on the display screen of device 130. This example of an audiovisual composition model defines a first part 43 of the display area 30 located in the upper left part of the display area 30 and a second part 44 of the display area 30 located in the center of the display area 30.
[0114] Figure 4H illustrates an example of an audiovisual composition model corresponding to a source audiovisual composition associated with a 'match' type which corresponds to a portrait mode display on the display screen of device 130. This example of an audiovisual composition model defines a first part 45 of the display area 33 located in the upper left part of the display area 33 and a second part 46 of the display area 33 located in the center of the display area 33.
[0115] The present invention is not limited to particular types of audiovisual content or to particular examples of audiovisual composition patterns such as those illustrated above, but extends to any type of audiovisual content and any type of audiovisual composition pattern that defines at least a part of a display area intended for the display of audiovisual content.
[0116] In step 310 of Figure 3, module 121 obtains at least one audiovisual composition model based on the source audiovisual composition.
[0117] Based on the examples in Figures 4A to 4H, if the video stream is of the 'TV news' type, then four audiovisual composition models (illustrated by Figures 4A to 4D) can be obtained. If the video stream is of the 'televised debate' type, then two audiovisual composition models (illustrated by Figures 4E and 4F) can be obtained. If the video stream is of the 'match' type, then two audiovisual composition models (illustrated by Figures 4G and 4H) can be obtained.
[0118] In step 320, module 121 detects at least one area of interest in an image of the video stream 111 based on the source audiovisual composition.
[0119] According to a particular and non-limiting embodiment of the present invention, a region of interest is spatially defined in the image of the video stream 111 as a function of two opposite vertices of a rectangle.
[0120] According to a particular and non-limiting embodiment of the present invention, a region of interest can be detected in at least a part of the display area of at least one obtained audiovisual composition model (step 320).
[0121] This example of implementation is advantageous because it allows methods for defining regions of interest to be dedicated to specific parts of the display area of an audiovisual composition model.
[0122] For example, in the case of the audiovisual composition model in Figures 4A to 4F, parts 31, 35, 37, 39, 40i to 404, 41, and 42i to 424 can each be dedicated to face display. In this case, a face detection method can be applied to an image of video stream 111 by limiting face detection within the image of video stream 111 to these parts of the audiovisual composition model. One example is the face detection method by Pau Viola and Michael Jones ("Robust Real-time Face Detection," IJVC, 2004, pp. 137-154).
[0123] In the case of part 44 or 46 of display area 30 or 33, an object detection and tracking method can be used to detect and track an object in the video stream. For example, the method by Sayed Mohammed Majidi Dorcheh et al. (SmartCrop: AI-based cropping of soccer videos, 2023, IEEE International Symposium on Multimedia (ISM)) describes a method that enables the detection and tracking of a soccer player and the ball. This method can be used, for example, to detect and track a region of interest around the ball and dedicate part 44 or 46 of display area 30 or 33 to displaying the content of the video stream delimited by this region of interest.
[0124] In a step 330, module 121 determines the first data 1211 which are representative of a spatial definition of each area of interest detected in the image of the video stream 111.
[0125] In step 340, module 121 determines the second data 1212 for each audiovisual composition model obtained (step 320).
[0126] The second set of 1212 data points determined for an audiovisual composition model defines a target audiovisual composition. The second set of 1212 data points determined for an audiovisual composition model graphically defines a display area for the target audiovisual composition, a portion of the display area for each detected area of interest, and an audiovisual content display format for each portion of the display area.
[0127] Defining a portion of the display area for a detected region of interest allows for the partial generation of the target audiovisual composition by establishing a relationship between the region of interest and a portion of the target audiovisual composition's display area. The audiovisual content of the video stream delimited by the region of interest can then be displayed in this portion of the display area according to the display format for that specific portion of the display area.
[0128] In step 350, module 121 obtains the third data 1213 representative of video stream 111.
[0129] The first data 1211, the second data 1212 determined for each audiovisual composition model and the third data 1213 are presented as input to module 122 (figure 2).
[0130] In a 360 step, module 122 broadcasts the 122i data, the 1222 data determined for each audiovisual composition model and the 1223 data which are obtained, respectively from the first 1211 data, the second 1222 data determined for each audiovisual composition model and the third 1223 data, through a first communication channel.
[0131] According to an example implementation, the 122i data, the 1222 data determined for each audiovisual composition model and the 1223 data are broadcast cyclically.
[0132] In a step 370, module 121 obtains data 113 representative of at least one post-production graphics element associated with the video stream 111. Module 121 then broadcasts the fourth data 1214 representative of said at least one post-production graphics element associated with the video stream through a second communication channel distinct from the first communication channel.
[0133] A post-production graphic element can represent any type of visual information related to the video stream 1 11. It can, for example, represent a brand logo, be presented as a scrolling banner ad or as an area in which textual information or thumbnails representing still images are displayed or in which videos are viewed.
[0134] The dimensions of a post-production graphics element are defined according to each audiovisual composition model, that is to say that the dimensions of a post-production graphics element are adapted to those of the part of the display area dedicated to it.
[0135] According to an example of implementation, each post-production graphic element can be represented in an HTML5 format.
[0136] As an example, the first and second communication channels can be established via the Internet network.
[0137] In a step 380, module 122 receives the clock signal 112 and broadcasts synchronization data obtained from the clock signal 112 through the first communication channel in relation to data 122i, 1222 and 1223 and the same synchronization data through the second communication channel in relation to data 1224.
[0138] As an example, synchronization data can represent time instants obtained from the clock signal 112.
[0139] According to a particular and non-limiting embodiment of the present invention, in a step 300, the module 121 detects whether a scene change appears in the video stream 11. If a scene change is detected then steps 310-380 are executed.
[0140] Module 122 is configured to perform processing on the third data points representing the video stream 111. For example, one such processing step is adaptive bitrate (ABR), which allows the video stream 111 to be adapted to different formats and resolutions required for viewing on various devices 130. Adaptive bitrate also allows the quality of the video stream 111 to be automatically adjusted based on the available bandwidth of the user. For example, a 4K UHD video stream can be transcoded into several resolutions, such as 1080p, 720p, and 480p. Thus, if the user's connection is fast and stable, they will receive the video stream 111 in high definition. Conversely, if the connection is slow or unstable, the quality of the video stream 111 will be dynamically reduced to prevent playback interruptions.This ensures a smooth and uninterrupted user experience, regardless of bandwidth variations.
[0141] The different versions of the 111 video stream output from the transcoding process can be represented in an MBR (Multiple BitRates) format. In this case, several versions of the 111 video stream at different bitrates are available. These different versions of the 111 video stream can be encoded using a video codec such as HEVC or VVC, for example. The third data can then represent image data encoded according to this video codec.
[0142] The third data 1213 may represent the different versions of the video stream 1 11 (possibly encoded).
[0143] Module 122 is also configured to package the first 1211, second 1212, third 1213 data, and synchronization data into formats suitable for streaming. Packaged data involves segmenting these first, second, and third data elements and the synchronization data into chunks and encapsulating them in formats Specific formats are used for streaming. The most commonly used formats are HLS / TS (HTTP Live Streaming with MPEG-TS type packet) and HLS / CMAF (HTTP Live Streaming associated with Common Media Application Format). The HLS / TS format segments the first, second, third, and synchronization data into chunks encapsulated in the MPEG-TS format, which is widely compatible with many devices but relatively inefficient in terms of file size. In contrast, the HLS / CMAF format uses the CMAF format, which reduces latency and improves streaming efficiency by enabling faster startup times and better network resource management. This latter format is increasingly adopted due to its performance and flexibility advantages.
[0144] According to a particular and non-limiting embodiment of the present invention, the first 1211, second 1212, and third 1213 data and synchronization data can be encapsulated using the MPEG-TS (MPEG Transport Stream) protocol defined by the MPEG-2 Part 1 standard (System, ISO / IEC 13818-1) of the Moving Picture Experts Group. The MPEG-2 standard defines the transport aspects across networks for digital television but is widely extended to the transport of other audiovisual content. Its primary purpose is to enable the multiplexing of video and audio in order to synchronize them. An MPEG-TS stream (named after the protocol used) can include several audio / video programs as well as program description and service data.
[0145] Figure 5 illustrates an example of encapsulation of first, second, third data and synchronization data according to the MPEG2-TS protocol, according to a particular and non-limiting embodiment of the present invention.
[0146] According to this example, three MPEG2 streams 541, 542, and 543 can be multiplexed. Stream 541 can include several programs 531 to 534, here four. A program, for example program 531, is made up of at least one Packet Elementary Stream (PES), here a pack of three elementary streams. (English Elementary Stream ES). One of these elementary streams, for example 511, can encapsulate the visual part of the third data 1213 representative of the video stream 111, another, for example 512, can encapsulate the audio part of the third data 1213 representative of the video stream 111 and another, for example 513, can encapsulate the first 1211, the second 1212 data and the synchronization data.
[0147] According to a particular and non-limiting embodiment of the present invention, the first 1211, second 1212 data and the synchronization data are of type ID3 or KLV.
[0148] The encapsulation of the first, second, third data and synchronization data by the MPEG2-TS protocol ensures the persistence of these first, second data and synchronization data during various processing (transcoding, packetization, etc.) carried out during the broadcast of the first, second, third data and synchronization data by module 122.
[0149] The first, second, and third data streams, along with the synchronization data, are packaged and stored. This is typically done in data centers using content management systems (CMS). These systems allow for the classification, security, and accessibility of on-demand video streams.
[0150] Once stored, the first, second, third, and packaged synchronization data are distributed. Content Delivery Networks (CDNs) can be used for this purpose. CDNs have servers located worldwide, which brings available video streams closer to end users and reduces latency. Thus, when a user requests a video stream from the UI of device 130 via the OTT service, it is delivered from the server closest to their location.
[0151] Figure 6 schematically illustrates a device configured to display a video stream broadcast by an OTT service according to a particular and non-limiting embodiment of the present invention.
[0152] The OTT video streaming service can, for example, implement the process shown in Figure 3.
[0153] The 130 video stream display device 111 includes a display screen 131.
[0154] In addition, the device 130 includes a human-machine interface (HMI) 133 configured so that a user can interact with a graphical interface 134 displayed on the display screen 131.
[0155] The 133 HMI is designed to be intuitive so that a user can easily navigate among audiovisual content, create lists of favorites, access display options and benefit from personalized recommendations based on their viewing habits.
[0156] For example, this graphical interface 134 may include graphical representations of videos available for viewing on demand. A user can then navigate through these graphical representations and select one to display the corresponding video stream using a video player (not shown).
[0157] For example, the 134 graphical interface may be similar to those of video-on-demand systems currently found on various online platforms.
[0158] According to a particular and non-limiting embodiment of the present invention, the display screen 131 may include a touch functionality which is used by a user as a means of interaction with the graphical interface 134. A user can then navigate among the graphical representations and select a video stream to be displayed by direct action on the graphical interface 134.
[0159] The 131 touch display screen corresponds for example to an LCD type screen (from the English "Liquid Crystal Display" or in French "Display à cristals liquide"), for example of type TFT (from the English "Thin-Film Transistor" or in French "Transistor en film mince"), or OLED (from the English "Organic Light-Emitting Diode" or in French "Diode électroluminescente organique").
[0160] Furthermore, device 130 includes a communication interface 135 configured to send requests for video to be displayed to a multimedia data server such as, for example, a CDN server and to receive data relating to the requested video via a communication network.
[0161] According to examples, the communication network can be a wired, wireless or hybrid network of the Internet type.
[0162] Communication interface 135 is configured to receive a request to display a video represented by a graphical representation of the graphical interface 134 selected by a user via a remote control 135 (or directly from the touchscreen). Communication interface 135 is also configured to send a request to display the selected video to the nearest CDN server to device 130. The CDN server then offers several versions of the requested video stream to device 130. Communication interface 135 automatically selects the highest bandwidth that the connection to this server can support without interruption and selects one of the offered versions.For example, if a fast connection is established between this CDN server and device 130, then a 1080p version at 5 Mbps (for example) of the requested video stream can be selected, while if a slower connection is established, then the 480p version at 1 Mbps (for example) is selected. Once the requested video stream version is selected, the CDN server broadcasts 122i, 1222, 123s data and synchronization data related to the selected version of the requested video stream.
[0163] According to one variant, 1224 data and synchronization data are also broadcast to device 130.
[0164] The communication interface 135 receives the data 122i, 1222, 123s and 1224 and obtains from the first data 135i representative of the spatial definition of at least one area of interest detected in the image of the requested video stream, from the second data 1352 defining, for each audiovisual composition model, graphically the display area of a target audiovisual composition, a part of the display area for each detected area of interest and a display format of audiovisual content for each part of the display area, third data 1353 representing the selected version of the selected video stream and fourth data 1354 representing at least one post-production graphic element relating to the selected version of the requested video stream, depending respectively on data 122i, 1222, 123s and 1224.
[0165] The communication interface 135 receives synchronization data related to data 122i, 1222 and 1223 through a first communication channel and the same synchronization data through a second communication channel related to data 1224.
[0166] According to a particular and non-limiting embodiment of the present invention, the data 122i, 1222, 123s and 1224 and the synchronization data are received in a binary file and the communication interface 135 obtains the first 1351, second 1352, third 135s, fourth 1354 data and the synchronization data by parsing the received binary file.
[0167] For example, this binary file is in the HLS / TS or HLS / CMAF format discussed previously.
[0168] For example, the MPEG-TS protocol can be used to obtain the first 135i, second 1222, third 1223, fourth 1224 data and synchronization data from a binary stream carrying this data.
[0169] In addition, device 130 includes a module 136 for creating a target audiovisual composition providing data 132i when the first data 135i, the second data 1352, the third data 135s and the fourth data 1354 are presented as input to module 136.
[0170] Module 136 is configured to implement the process shown in Figure 7.
[0171] Device 130 further includes a module 132 for controlling the display of a selected target audiovisual composition of a requested video stream on the display screen 131 according to the data 132i provided by module 136 and the synchronization data.
[0172] According to examples, device 130 can be a computer, possibly a laptop, a tablet, a smartphone, a television, or any other mobile communication device configured to display a video stream.
[0173] Figure 7 schematically illustrates a diagram of the steps of a method for displaying a video stream according to a particular and non-limiting embodiment of the present invention.
[0174] In a step 71, module 136 receives, through the first communication channel, the first 135i data representative of a spatial definition of at least one area of interest of an image of the video stream.
[0175] In step 72, module 136 receives the second data 1352 for at least one audiovisual composition model through the first communication channel. Each audiovisual composition model defines a target audiovisual composition; the second data 1352 received for each audiovisual composition model graphically defines a display area of the target audiovisual composition, and for each area of interest, a portion of the display area and a display format for the audiovisual content of said portion of the display area.
[0176] The display area spatially delimits a zone on the display screen 131 in which a target audiovisual composition is displayed. The dimensions of the display area may, for example, correspond to those of the display screen 131 or be smaller than them.
[0177] In step 73, module 136 receives the third data 135s representative of the video stream without post-production graphic elements.
[0178] In step 74, module 36 selects an audiovisual composition model from said at least one audiovisual composition model defined by the second data 1352.
[0179] In step 75, module 136 obtains the display area defined by the second data received 1352 for the selected audiovisual composition model.
[0180] In a step 76, for each region of interest spatially defined by the first 135i data received, the module 136 obtains a portion of audiovisual content from the video stream delimited by the region of interest and, in a step 77, adds the portion of audiovisual content obtained to a portion of the display area defined for the region of interest by the second 1352 data received for the selected audiovisual composition model, according to the display format defined for said portion of the display area by the second 1352 data received for the selected audiovisual composition model.
[0181] Module 136 then provides the 132i data as input to module 132.
[0182] In step 78, module 136 receives the fourth data 1354 representing at least one post-production graphic element associated with the video stream via a second communication channel distinct from the first communication channel, and in step 79, module 136 adds said at least one post-production graphic element to at least a portion of the resulting display area. Each post-production graphic element is added to a portion of the resulting display area, for example, as an overlay on audiovisual content displayed in said portion of the display area.
[0183] In one step (80), the module 136 receives synchronization data from a clock signal through the first communication channel in relation to the first 135i, second 1352 and third 135s data and the same synchronization data through the second communication channel in relation to the fourth 1354 data.
[0184] In step 81, module 132 controls a synchronized display, according to the synchronization data, on display screen 131, of the audiovisual content and / or the post-production graphic overlay element of each part of the display area obtained according to a display format defined for said part of the area display by the second data 1352 received for the selected audiovisual composition model.
[0185] Display control for audiovisual content within the display area of a target audiovisual composition includes rendering that audiovisual content. Such rendering corresponds to a set of operations performed by one or more processors on the pixels of one or more images of the audiovisual content to be displayed on the display screen. For example, rendering involves associating pixel data (e.g., color data expressed in an RGB (Red, Green, Blue) color space) with each graphic object within a set of pixels in an image. Rendering may include data decoding to obtain pixel data according to, for example, MPEG-type video codecs such as HEVC or VVC.
[0186] The control of the display of audiovisual content thus includes the transmission by module 132 of control signals to the display screen 131 to modify the values associated with the pixels of the display screen 131 at the location intended to display the audiovisual content.
[0187] According to a particular and non-limiting embodiment of the present invention, the audiovisual composition model is selected from said at least one audiovisual composition model defined by the second data 1352 from a human-machine interface, for example the HMI 133.
[0188] According to a particular and non-limiting embodiment of the present invention, the audiovisual composition model is selected from said at least one audiovisual composition model defined by the second data 1352 according to an orientation of the display screen 131.
[0189] For example, when module 136 is configured to detect an orientation (landscape or portrait mode) of display screen 131 and selects an audiovisual composition model from said at least one audiovisual composition model defined by the second data 1352 depending on whether display screen 131 is in landscape or portrait mode.
[0190] According to a particular and non-limiting embodiment of the present invention, the audiovisual composition model is selected from said at least one audiovisual composition model defined by the second data 1352 according to a preferred display format.
[0191] For example, this preferred display format is chosen beforehand by a user.
[0192] For example, a preferred display format may define display parameters such as the dimensions of a display area in which the visual parts of audiovisual content are displayed, spatio-temporal resolutions, spatial resolutions of the images of the audiovisual content and / or a frame rate and / or parameters related to the number of components of the images displayed or a number of bits per pixel of said images.
[0193] According to a particular and non-limiting embodiment of the present invention, the audiovisual composition model is selected from said at least one audiovisual composition model based on video stream display capacity and / or based on data reception capabilities.
[0194] This example of implementation is advantageous because it allows limiting the number of audiovisual composition models that a user can choose to display a video stream to those that allow a smooth display experience (without stuttering or unexpected blocking).
[0195] According to a particular and non-limiting embodiment of the present invention, the display format defined for a part of the display area by the second 1352 received data for the selected audiovisual composition model defines a cropping of the audiovisual content of said part of the display area according to the dimensions of said part of the display area and the dimensions of said audiovisual content.
[0196] The process in Figure 7 allows for obtaining a target audiovisual composition from the selected video stream and possibly one or more post-production graphic elements.
[0197] The audiovisual content that makes up this target audiovisual composition is displayed in a synchronized manner by the synchronization data.
[0198] When the display of the target audiovisual composition is implemented by an application installed on device 130 running an operating system such as Android or iOS, a software component integrated into the application performs the operations of obtaining the requested portions of the video stream (cropping) and inserts post-production graphic elements (overlay), frame by frame. The video frames thus generated are then displayed on the display screen 131. It should be noted that audio tracks can also be rendered synchronously with the display of the generated video frames.
[0199] When the display of the target audiovisual composition is implemented by a web browser's display (web player), a calculation module such as a JavaScript library that can be integrated into a portal web page performs the operations of obtaining the requested parts of the video stream (cropping) and inserting the post-production graphic elements (overlay), then obtains a new video stream in a standard format such as .ts or .mp4, which is then sent to the native decoder of a web browser for the display of the selected target audiovisual composition of the requested video stream.
[0200] Figure 8A illustrates an example of a target audiovisual composition of a requested video stream, according to a particular and non-limiting embodiment of the present invention.
[0201] According to this example, the target audiovisual composition is generated by the process in Figure 7 according to the audiovisual composition model in Figure 4A. The target audiovisual composition is therefore adapted for display on the display screen 131 in a landscape format. The target audiovisual composition includes the audiovisual content of a region of interest 91 extracted from the video stream 11 and displayed in part 31 of the display area 30 defined by the audiovisual composition model in Figure 4A, and a post-production graphic element 92, for example, an announcement banner, displayed in part 32 of the display area 30.
[0202] Figure 8B illustrates an example of a target audiovisual composition of a requested video stream, according to a particular and non-limiting embodiment of the present invention.
[0203] According to this example, the target audiovisual composition is generated by the process in Figure 7 according to the audiovisual composition model in Figure 4B. The target audiovisual composition is therefore adapted for display on the display screen 131 in a portrait format. The target audiovisual composition includes the audiovisual content of a region of interest 93 extracted from the requested video stream 11 and displayed in part 35 of the display area 33 defined by the audiovisual composition model in Figure 4B, and a post-production graphic element 94, for example, an announcement banner, displayed in part 34 of the display area 33.
[0204] The target audiovisual composition in Figure 8A is intended to be displayed in landscape format, while the target audiovisual composition in Figure 8B is intended to be displayed in portrait format. Video stream segments 91 and 93 and elements 92 and 94 are adapted to the display area dimensions depending on whether the display format is portrait or landscape.
[0205] Figure 8C illustrates an example of a target audiovisual composition of a requested video stream, according to a particular and non-limiting embodiment of the present invention.
[0206] According to this example, the target audiovisual composition is generated by the process shown in Figure 7, based on the audiovisual composition model in Figure 4E. The target audiovisual composition is therefore adapted for display on the display screen 131 in a landscape format. The target audiovisual composition comprises the content of a region of interest 96 extracted from the requested video stream and displayed in part 39 of the display area 30 defined by the audiovisual composition model in Figure 4E, and the content of four regions of interest 97 to 100 extracted from the requested video stream and displayed in parts 40i to 404 of the display area 30.
[0207] Figure 8D illustrates an example of a target audiovisual composition of a requested video stream, according to a particular and non-limiting embodiment of the present invention.
[0208] According to this example, the target audiovisual composition is generated by the process in Figure 7 according to the audiovisual composition model in Figure 4E. The composition is therefore adapted for display on the display screen 131 in a portrait format. The target audiovisual composition includes the content of a region of interest 102 extracted from the requested video stream and displayed in part 41 of the display area 33 defined by the audiovisual composition model in Figure 4F, and the content of four regions of interest 103 to 106 extracted from the requested video stream and displayed in parts 42i to 424 of the display area 33.
[0209] The target audiovisual composition in Figure 8C is intended to be displayed in landscape format, while the target audiovisual composition in Figure 8D is intended to be displayed in portrait format. Video stream segments 96-100 and 102-106 are adapted to the display area dimensions depending on whether the display format is portrait or landscape.
[0210] Of course, the present invention is not limited to the embodiments described above but extends to a method for broadcasting and displaying video streams that would include secondary steps without falling outside the scope of the present invention. The same would apply to a device configured for implementing such a method.
[0211] The embodiments or examples of the present invention can be implemented using hardware comprising one or more analog and / or digital circuits, and / or software, by executing instructions by one or more general-purpose or special-purpose processors, or as a combination of hardware and software. Therefore, the embodiments or examples of the present invention can be implemented in the environment of a computer system or other processing system. An example of such a computer system is shown in Figure 9. The blocks described in the figures above, such as the blocks in Figures 3 and 7, can run on one or more computer systems. Furthermore, Each of the steps in the flowcharts described above can be implemented on one or more computer systems. When more than one computer system is used to implement embodiments of the present invention, the computer systems can be interconnected by one or more networks to form a cluster of computer systems that can act as a single pool of homogeneous resources. The interconnected computer systems can form a "cloud" of computers.
[0212] The computer system 900 includes one or more processors, such as the processor 904. The processor 904 may be, for example, a dedicated-purpose processor, a general-purpose processor, a microprocessor, or a digital signal processor. The processor 904 may be connected to a communication infrastructure 902 (for example, a bus or a network). The computer system 900 may also include main memory 906, such as random access memory (RAM), and may also include secondary memory 908. The secondary memory 908 may include, for example, a hard disk drive 910 and / or a removable storage unit 912, representing a magnetic tape drive, an optical disk drive, or other. The removable storage unit 912 can read from and / or write to a removable storage unit 916 in a well-known manner.Removable storage unit 916 represents a magnetic tape, optical disc or other, which is read and written by removable storage unit 912. As will be understood by persons competent in the relevant fields, removable storage unit 916 comprises a storage medium usable by a computer on which software and / or data are stored.
[0213] In other embodiments, the secondary memory 908 may include other similar means for loading computer programs or other instructions into the computer system 900. These means may include, for example, a removable storage unit 918 and an interface 914. This could be, for example, a program cartridge and cartridge interface (such as those found in video game devices), a removable memory chip (such as an EPROM or PROM) and associated connector, a USB flash drive and USB port, and other storage units. removable 918 and 914 interfaces that allow software and data to be transferred from the removable storage unit 918 to the computer system 900.
[0214] The computer system 900 may also include a communication interface 920. The communication interface 920 allows the transfer of software and data between the computer system 900 and external devices such as a display screen. The communication interface 920 may include, for example, a modem, a network interface (such as an Ethernet card), a communication port, etc. The software and data transferred via the communication interface 920 are in the form of signals that can be electronic, electromagnetic, optical, or other signals capable of being received by the communication interface 920. These signals are transmitted to the communication interface 920 via a communication path 922.The 3122 communication path carries the signals and can be implemented using wire or cable, optical fiber, telephone line, cellular phone link, RF link and other communication channels.
[0215] The 900 computer system can also include one or more 924 sensors. The 924 sensor(s) can measure or detect one or more physical quantities and convert the measured or detected physical quantities into an electrical signal in digital and / or analog form. For example, the 924 sensor(s) may include a camera to capture a video stream.
[0216] In this document, the terms "computer program media" and "computer-readable media" refer to tangible storage media, such as removable storage units 916 and 918 or a hard disk drive installed in the hard disk drive 910. These computer program products are means of delivering software to the computer system 900. Computer programs (also called computer control logic) can be stored in main memory 906 and / or secondary memory 908. Computer programs can also be received via the communication interface 920. These computer programs, when executed, enable the computer system 900 to implement this disclosure as described herein. In particular, the Computer programs, when executed, allow the processor 904 to implement the processes of this disclosure, such as the methods described herein. Consequently, these computer programs represent controllers of the computer system 900.
[0217] In another embodiment, the features of the present invention can be implemented in hardware using, for example, hardware components such as application-specific integrated circuits (ASICs) and gate arrays. The implementation of a hardware state machine to perform the functions described above will also be obvious to those competent in the relevant field.
Claims
DEMANDS 1. A method for broadcasting a video stream without post-production graphic overlay elements and originating from an audiovisual production, the video stream having a source audiovisual composition defining at least one audiovisual content, a spatial definition of each audiovisual content within an image of the video stream, and a display format for each audiovisual content, the method comprising the following steps: obtaining (310) at least one audiovisual composition model based on the source audiovisual composition, each audiovisual composition model defining a target audiovisual composition,each audiovisual composition model graphically defining a display area of the target audiovisual composition and spatially defining at least a portion of the display area dedicated to displaying audiovisual content from the source audiovisual composition in a display format different from that defined for said audiovisual content by the source audiovisual composition; - detection (320) of at least one area of interest in an image of the video stream based on the source audiovisual composition; - determination (330) of initial data representative of a spatial definition of each area of interest detected in the image of the video stream; - determination (340) of second data for each audiovisual composition model defining a target audiovisual composition, the second data, determined for an audiovisual composition model, graphically defining the display area of the target audiovisual composition; a part of the display area for each detected area of interest and an audiovisual content display format for each part of the display area; - obtaining (350) third data points representative of the video stream; - dissemination (360) of the first data, the second data determined for each audiovisual composition model, and the third representative data video stream through a primary communication channel, - dissemination (370) of fourth data representing at least one post-production graphic element associated with the video stream through a second communication channel distinct from the first communication channel; and - broadcasting (380) of synchronization data from a clock signal through the first communication channel in relation to the first, second and third data and the same synchronization data through the second communication channel in relation to the fourth data.
2. A method according to claim 1, which further comprises a step (300) of scene change detection in the video stream, and for which the steps of obtaining at least one audiovisual composition model, detecting at least one area of interest, determining first, second, third and fourth data and broadcasting are executed following the detection of each scene change.
3. A method for displaying a video stream on a display screen, the method comprising the following steps: - reception (71), through a first communication channel, of first data representative of a spatial definition of at least one area of interest of an image of the video stream; - reception (72) of second data for at least one audiovisual composition model through the first communication channel, each audiovisual composition model defining a target audiovisual composition, the second data received for each audiovisual composition model graphically defining a display area of the target audiovisual composition, and for each area of interest, a part of the display area and an audiovisual content display format of said part of the display area; - reception (73), through the first communication channel, of third data representative of the video stream without post-production graphic elements; - selection (74) of an audiovisual composition model from said at least one audiovisual composition model defined by the second data received; - obtaining (75) the display area defined by the second data received for the selected audiovisual composition model; - for each region of interest spatially defined by the first data received: obtaining (76) a part of an audiovisual content from the video stream delimited by the region of interest; and adding (77) the part of the audiovisual content obtained into a part of the display area defined for the region of interest by the second data received for the selected audiovisual composition model, according to the display format defined for said part of the display area by the second data received for the selected audiovisual composition model; - reception (78), through a second communication channel distinct from the first communication channel, of fourth data representing at least one post-production graphic element associated with the video stream; - addition (79) of said at least one post-production graphic element in at least part of the resulting display area; - reception (80) of synchronization data from a clock signal through the first communication channel in relation to the first, second and third data points and the same synchronization data through the second communication channel in relation to the fourth data point; and - control (81) of a synchronized display according to the synchronization data, on the display screen, of the audiovisual content and / or the post-production graphic element of each part of the display area obtained according to a display format defined for said part of the display area by the second data received for the selected audiovisual composition model.
4. A method according to claim 3, wherein the audiovisual composition model is selected from said at least one audiovisual composition model of a preferred display format.
5. A method according to any one of claims 3 to 4, wherein the audiovisual composition model is selected from said at least one audiovisual composition model based on data reception capabilities.
6. A method according to any one of claims 3 to 5, wherein the display format defined for a portion of the display area by the second received data for the The selected audiovisual composition template defines a cropping of the audiovisual content of said part of the display area based on the dimensions of said part of the display area and the dimensions of said audiovisual content.
7. Device for broadcasting a video stream implementing one of the methods according to any one of claims 1 to 2.
8. A device for displaying a video stream comprising a display screen implementing one of the methods according to any one of claims 3 to 6.
9. Video streaming and display system comprising a video streaming device according to claim 7 and at least one video streaming display device according to claim 8.
10. Computer program which includes instructions adapted for carrying out the steps of the process according to any one of claims 1 to 6, when the computer program is executed by at least one processor.
11. Computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the process according to any one of claims 1 to 6.
Citation Information
Patent Citations
Multiple camera video system which displays selected images
US20020049979A1
Information reading apparatus
US20070229706A1
User terminal apparatus, display apparatus, system and control method thereof
US20160050449A1
Method and apparatus for transreceiving broadcast signal for panorama service
US20160337706A1
Multi-source video navigation
US20180167685A1