System and method for generating a video stream - Patent application

JP2024538087A5Pending Publication Date: 2025-10-17LIVEARENA TECH AB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024522190
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-10-15
Filing Date
2022-10-14
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Digital video conferencing systems face challenges in dynamically determining what information to display due to varying delays, frame rates, resolutions, aspect ratios, and encodings among different video streams, leading to unsynchronized feeds and computational intensity, which affect user experience.

Method used

A method and system for generating shared digital video streams by collecting, time-synchronizing, and pattern-detecting multiple input streams using a central server with AI components to align and combine them into a synchronized output stream, considering events and patterns for optimal display.

Benefits of technology

This approach ensures low-latency, synchronized, and high-quality video streams that adapt to dynamic meeting scenarios, improving user experience by intelligently managing display content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method for providing a shared digital video stream is disclosed, the method comprising the steps of: collecting a plurality of primary digital video streams (210) from at least two digital video sources (120), respectively, in a collection step; individually analyzing the plurality of primary digital video streams (210) to detect at least one event (211) selected from a first set of events in an event detection step; time synchronizing the plurality of primary digital video streams (210) to a common time reference (260), in a synchronization step; analyzing the plurality of time-synchronized primary digital video streams (210) in a pattern detection step. detecting at least one pattern (212) selected from a first pattern set, wherein the detection of the pattern is based on the at least one detected event (211); generating, in a generating step, the shared digital video stream as an output digital video stream (230) based on consecutively considered frames (213) of the time-synchronized multiple primary digital video streams (210) and the detected pattern (212); and, in a publishing step, continuously providing the output digital video stream (230) to consumers of the shared digital video stream. The present invention also relates to a system and a computer software product.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a system, a computer software product and a method for generating a digital video stream, in particular for generating a digital video stream based on two or more different digital input video streams. In a preferred embodiment, the digital video stream is generated in the context of a digital video conference, in particular a digital video conference or meeting system, in which a number of different simultaneously connected users participate. The generated digital video stream may be published externally or within the digital video conference or digital video conference system.

[0002] In other embodiments, the invention is applied to contexts that are not digital video conferencing, but where multiple digital video input streams are simultaneously processed and combined into a digital video stream to be generated. For example, such a context may be educational or instructional. [Background technology]

[0003] Many digital video conferencing systems are known, such as Microsoft® Teams®, Zoom®, and Google® Meet®, that allow two or more participants to meet virtually using digital video and audio that is recorded locally and broadcast to all participants, emulating a physical meeting.

[0004] There is a general need to improve such digital video conferencing solutions, particularly with regard to the generation (production) of viewing content: what content to show, at what time, to whom, and through what distribution channel.

[0005] For example, some systems automatically detect who is currently speaking and display the corresponding video feed of that participant to the other participants. Many systems allow for the sharing of graphics such as the currently displayed screen, a viewing window, or a digital presentation. But as virtual meetings become more complex, it will soon become difficult for services to know what, of all the information currently available, should be shown to each participant at any given time.

[0006] In another example, a presenting participant moves around on stage while talking about slides in a digital presentation, in which case the system needs to decide whether to show the presentation, the presenter, or both, or switch between the two.

[0007] It may be desirable to generate, by an automated generation process, one or more output digital video streams based on multiple input digital video streams, and provide such generated digital video stream or streams to one or more consumers. DISCLOSURE OF THEINVENTION [Problem to be solved by the invention]

[0008] However, in many cases, due to the many technical challenges faced by such digital video conferencing systems, it is difficult for dynamic conference screen layout managers and other automated generators to select what information to display.

[0009] First, low latency is important because digital video conferencing is real-time sensitive. This becomes problematic when different incoming digital video streams are associated with different latency, different frame rates, different aspect ratios, or different resolutions, such as when different participants join using different hardware. Often, such incoming digital video streams require processing for a well-formed user experience.

[0010] Second, there is the issue of time synchronization: the various input digital video streams, such as external digital video streams and digital video streams provided by the participants, are typically fed into a central server or the like, and there is no absolute time to synchronize each of such digital video feeds with. Similar to too much delay, unsynchronized digital video feeds lead to a poor user experience.

[0011] Third, digital video conferences between multiple participants may involve different digital video streams with different encodings or formats, which require decoding and re-encoding, creating problems in terms of latency and synchronization, and such encoding is computationally intensive and expensive in terms of hardware requirements.

[0012] Fourth, the fact that different digital video sources may be associated with different frame rates, different aspect ratios, and different resolutions can result in unpredictable changes in memory allocation needs that require continual balancing, potentially resulting in additional latency and synchronization issues, which in turn necessitates the need for large buffers.

[0013] Fifth, participants may experience a variety of difficulties in terms of connectivity fluctuations, drop-off / reconnection, etc., posing additional challenges to automatically generating a shaped user experience.

[0014] These issues are amplified in more complex meeting situations, including those with large numbers of participants; participants connecting using different hardware and / or software; using externally provided digital video streams; screen sharing; and multiple hosts.

[0015] Corresponding problems arise in other contexts, when an output digital video stream is to be generated based on multiple input digital video streams, such as in digital video generation systems for education and instruction.

[0016] The present invention is directed to solving one or more of the problems set forth above. [Means for solving the problem]

[0017] Therefore, the present invention relates to a method for providing a shared digital video stream, the method comprising the steps of: in a collecting step, collecting a plurality of primary digital video streams from at least two digital video sources respectively; in an event detecting step, analyzing the plurality of primary digital video streams individually to detect at least one event selected from a first event set; in a synchronizing step, time synchronizing the plurality of primary digital video streams to a common time reference; in a pattern detecting step, analyzing the time-synchronized plurality of primary digital video streams to detect at least one pattern selected from a first pattern set, where the detection of the pattern is based on the at least one detected event; in a generating step, generating the shared digital video stream as an output digital video stream based on a successively considered plurality of frames of the time-synchronized plurality of primary digital video streams and the detected pattern; and in a publishing step, continuously providing the output digital video stream to consumers of the shared digital video stream.

[0018] The present invention also relates to a computer software product for providing a shared digital video stream, which when executed performs the following steps: in a collection step, a plurality of primary digital video streams are collected from at least two digital video sources respectively; in an event detection step, the plurality of primary digital video streams (210) are individually analyzed to detect at least one event (211) selected from a first set of events; in a synchronization step, the plurality of primary digital video streams (210) are time-synchronized with respect to a common time reference (260); in a pattern detection step, the time-synchronized plurality of primary digital video streams (210) are analyzed to detect at least one pattern (212) selected from a first set of patterns, where the detection of the pattern is based on the at least one detected event (211); in a generation step, the shared digital video stream is generated as an output digital video stream based on successively considered frames of the time-synchronized plurality of primary digital video streams and the detected pattern; and in a publishing step, the output digital video stream is continuously provided to consumers of the shared digital video stream.

[0019] The present invention further relates to a system for providing a shared digital video stream, comprising a central server having the following functions: a collection function configured to collect, in a collection step, a plurality of primary digital video streams from at least two digital video sources respectively; an event detection function configured to analyze the plurality of primary digital video streams individually to detect at least one event selected from a first set of events; a synchronization function configured to time-synchronize the plurality of primary digital video streams to a common time reference; a pattern detection function configured to analyze the time-synchronized plurality of primary digital video streams to detect at least one pattern selected from a first set of patterns, where the detection of the pattern is based on the at least one detected event; a generation function configured to generate the shared digital video stream as an output digital video stream based on successively considered frames of the time-synchronized plurality of primary digital video streams and the detected pattern; and a publishing function configured to continuously provide the output digital video stream to consumers of the shared digital video stream.

[0020] The invention will now be described in detail with reference to exemplary embodiments thereof and the enclosed drawings, in which: [Brief description of the drawings]

[0021] [Figure 1] FIG. 1 is a diagram showing a first system according to the present invention. [Diagram 2] FIG. 2 shows a second system according to the present invention. [Diagram 3] FIG. 3 shows a third system according to the present invention. [Figure 4] FIG. 4 is a diagram illustrating a central server according to the present invention. [Diagram 5] FIG. 5 shows a central server for use in the system according to the present invention. [Figure 6a]FIG. 6a illustrates subsequent states associated with different method steps in the method illustrated in FIG. [Figure 6b] FIG. 6b illustrates subsequent states associated with different method steps in the method illustrated in FIG. [Figure 6c] FIG. 6c illustrates subsequent states associated with different method steps in the method illustrated in FIG. [Figure 6d] FIG. 6d illustrates subsequent states associated with different method steps in the method illustrated in FIG. [Figure 6e] FIG. 6e illustrates subsequent states associated with different method steps in the method illustrated in FIG. [Figure 6f] FIG. 6f illustrates subsequent states associated with different method steps in the method shown in FIG. [Figure 7] FIG. 7 is a diagram conceptually showing a common protocol used in the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0022] All figures share the same or corresponding part reference numbers.

[0023] FIG. 1 shows a system 100 according to the invention, adapted to carry out a method according to the invention for providing a shared digital video stream.

[0024] The system 100 may include a video communication service 110, which may be external to the system 100 in some embodiments.

[0025] The system 100 may include one or more participant clients 121, although one, some, or all of the participant clients 121 may be external to the system 100 in some embodiments.

[0026] The system 100 includes a central server 130 .

[0027] As used herein, the term "central server" refers to a computer-implemented function configured to be accessible in a logically centralized manner, such as through a well-defined API (Application Programming Interface). Such central server functionality may be implemented purely in computer software, or in a combination of software and virtual and / or physical hardware, in a standalone physical or virtual server computer, or distributed across multiple interconnected physical and / or virtual server computers.

[0028] The physical or virtual hardware on which central server 130 runs, in other words the computer software that defines the functionality of central server 130, may be comprised of a per se conventional CPU, a per se conventional GPU, per se conventional RAM / ROM memory, per se conventional computer buses, and per se conventional external communication capabilities, such as an Internet connection.

[0029] The video communication service 110 , in so far as that is used, is also a central server in the above sense, which may be a different central server from the central server 130 or may be part of the central server 130 .

[0030] Correspondingly, each of the participant clients 121 may be a central server in the above sense, in a corresponding interpretation, in which the physical or virtual hardware on which each participant client 121 runs, in other words the computer software defining the functionality of the participant client 121, comprises a CPU / GPU per se, conventional RAM / ROM memory per se, a computer bus per se, and external communication capabilities per se conventional, such as an Internet connection.

[0031] Each participant client 121 also typically comprises, or is in communication with, a computer screen arranged to display video content that is provided to the participant client 121 as part of an ongoing video communication, speakers arranged to emit sound content that is provided to the participant client 121 as part of the video communication, a video camera, and a microphone arranged to record sound local to a human participant 122 to the video communication, who uses that participant client 121 to participate in the video communication.

[0032] In other words, the human-machine interface of each participant client 121 enables each participant 122 to interact with other participants in a video communication and / or with audio / video streams provided from various sources at that client 121.

[0033] Typically, each participant client 121 comprises a respective input means 123, which may consist of said video camera, said microphone, a keyboard, a computer mouse or trackpad, and / or an API for receiving digital video streams, digital audio streams and / or other digital data. The input means 123 are in particular configured to receive video and / or audio streams from the video communication service 110 and / or a central server, such as the central server 130, such video and / or audio streams being provided as part of the video communication and preferably generated based on corresponding digital data input streams provided to said central server from at least two sources of such digital data input streams, e.g. the participant client 121 and / or an external source (described below).

[0034] More generally, each participant client 121 includes a respective output means 124 that may consist of the computer screen, the speakers, and an API that emits digital video and / or audio streams that are representative of the locally captured video and / or audio to a participant 122 using that participant client 121.

[0035] In practice, each participant client 121 may be a mobile device, such as a mobile phone, equipped with a screen, speakers, microphone, and Internet connection, running computer software locally or accessing remotely executed computer software to perform the functions of that participant client 121. Correspondingly, a participant client 121 may be a thick or thin laptop or stationary computer, running locally installed applications and also using functions accessed remotely via a web browser.

[0036] There may be one or more participant clients 121, for example at least three, used in one and the same video communication in this embodiment.

[0037] The video communications may be provided at least in part by the video communications service 110 and at least in part by the central server 130, as described and illustrated herein.

[0038] As the term is used herein, a "video communication" is a two-way digital communication session that includes at least two, and preferably at least three, video streams, preferably also coinciding with an audio stream that is used to generate one or more mixed or collaborative digital video / audio streams. Such video communication may be real-time, with or without a certain latency or delay. At least one, and preferably at least two, participants 122 who participate in such a video communication engage in the video communication in a two-way manner, providing and consuming video / audio information.

[0039] At least one of the participant clients 121, or all of the participant clients 121, includes a local synchronization software function 125, which is described in more detail below.

[0040] The video communication services 110 may be provided with or have access to a common time reference, as described in more detail below.

[0041] The central server 130 may include an API 137 for digitally communicating with entities external to the central server 130. Such communications may include both inputs and outputs.

[0042] The system 100, such as the central server 130, may be configured to digitally communicate with an external information source 300, such as an externally provided video stream, and in particular to receive digital information, such as audio and / or video stream data, from the external information source 300. By "external," the information source 300, it is meant that it is not provided by or as part of the central server 130. Preferably, the digital data provided by the external information source 300 is independent of the central server 130, and the central server 130 cannot influence its information content. For example, the external information source 300 may be live captured video and / or audio, such as a public sporting event or an ongoing news event or report. Also, the external information source 300 may be captured by a webcam or the like, rather than by any of the participant clients 121. Thus, such captured video may depict the same locality as any one of the participant clients 121, but is not captured as part of the participant client 121's activities. One possible difference between the externally provided information source 300 and the internally provided information source 120 is that the internally provided information source may be provided in that capacity as a participant in a video communication of the type defined above, whereas the externally provided information source 300 is not, but instead is provided as part of a context that is external to the video conference.

[0043] There may also be multiple external information sources 300 providing digital information of that type, such as audio and / or video streams, in parallel to the central server 130 .

[0044] As shown in FIG. 1, each participant client 121 constitutes a source of information (video and / or audio) streams 120 that are provided by that participant client 121 to the video communication service 110, as described.

[0045] The system 100, such as the central server 130, may be further configured to digitally communicate with the external consumers 150, and in particular to emit digital information to the external consumers 150. For example, digital video and / or audio streams generated by the central server 130 may be provided continuously, in real time or near real time, to one or more external consumers 150 via the API 137 described above. Again, the consumer 150 being "external" means that the consumer 150 is not provided as part of the central server 130 and / or is not a party to the video communication in question.

[0046] Unless otherwise noted, all functions and communications herein are provided digitally and electronically, implemented by computer software running on appropriate computer hardware and communicated over a digital communications network or channel, such as the Internet.

[0047] 1, multiple participant clients 121 participate in a digital video communication provided by the video communication service 110. Each participant client 121 therefore has an ongoing login, session, or the like to the video communication service 110 and can participate in one and the same ongoing video communication provided by the video communication service 110. In other words, the video communication is "shared" among the participant clients 121 and therefore also by the corresponding human participants 122.

[0048] 1 , the central server 130 comprises an auto-join client 140, which is an auto-client that corresponds to the participant client 121, but is not associated with a human participant 122. Instead, the auto-join client 140 is added as a participant client to the video communication service 110 to participate in the same shared video communication as the participant client 121. As such a participant client, the auto-join client 140 is given access to continuously generated digital video and / or audio stream(s) provided as part of an ongoing video communication by the video communication service 110, and such streams can be consumed by the central server 130 via the auto-join client 140. Preferably, the automatic participant client 140 receives from the video communication service 110 a common video and / or audio stream that is distributed or can be distributed to each participant client 121; respective video and / or audio streams that are provided from each of one or more participant clients 121 to the video communication service 110 and relayed by the video communication service 110 to all participant clients 121 or to requesting participant clients 121 in raw or modified form; and / or a common time reference.

[0049] The central server 130 includes a collection function 131 configured to receive a plurality of video and / or audio streams of the above types from the auto-participant clients 140, and possibly also from the above-mentioned external information source(s) 300, for processing as described below, and then provide a shared video stream via the API 137. For example, this shared video stream may be consumed by external consumers 150 and / or by the video communication service 110, which may then distribute it to all or any requesting ones of the participant clients 121.

[0050] FIG. 2 is similar to FIG. 1, but instead of using an auto-join client 140, the central server 130 receives video and / or audio stream data from an ongoing video communication via the API 112 of the video communication service 110.

[0051] 3 is similar to FIG. 1, but the video communication service 110 is not shown. In this case, the participant clients 121 communicate directly with the API 137 of the central server 130, for example, to provide video and / or audio stream data to the central server 130 and / or to receive video and / or audio stream data from the central server 130. The generated shared streams may then be provided to external consumers 150 and / or to one or more of the client participants 121.

[0052] 4 shows the central server 130 in more detail. As shown, the collection function 131 may be composed of one or, preferably, multiple, format-specific collection functions 131a. Each of the format-specific collection functions 131a is configured to receive video and / or audio streams having a predefined format, such as a predefined binary encoding format and / or a predefined stream data container, and in particular to parse and classify the binary video and / or audio data of said format into individual video frames, sequences of video frames and / or time slots.

[0053] The central server 130 further comprises an event detection function 132 configured to receive video and / or audio stream data, such as binary stream data, from the collection function 131 and perform respective event detection on each individual one of the received data streams. The event detection function 132 may comprise an AI (artificial intelligence) component 132a for performing event detection. The event detection may be performed without first time synchronizing the collected individual streams.

[0054] The central server 130 further comprises a synchronization function 133 configured to time-synchronize the multiple data streams provided by the collection function 131 and processed by the event detection function 132. The synchronization function 133 may comprise an AI component 133a for performing the time synchronization.

[0055] The central server 130 further comprises a pattern detection function 134 configured to perform pattern detection based on a combination of at least one, but often at least two, e.g. at least three, e.g. all, of the received data streams. The pattern detection may further be based on one, possibly at least two or more events detected by the event detection function 132 for each one of said data streams. Such detected events considered by the pattern detection function 134 may be distributed over time with respect to the individual collected streams. The pattern detection function 134 may comprise an AI component 134a for performing pattern detection.

[0056] The central server 130 further comprises a generating function 135 configured to generate a shared digital video stream based on the multiple data streams provided by the collecting function 131 and further based on any detected events and / or patterns. The shared video stream includes at least a generated video stream comprising one or more of the raw data, reformatted or converted video streams provided by the collecting function 131, and may include corresponding audio stream data.

[0057] The central server 130 further comprises a publishing function 136 configured to publish the generated shared digital video stream, such as via the API 137 described above.

[0058] It should be noted that while Figures 1, 2 and 3 show three different examples of how the central server 130 can be used to implement the principles described herein, and in particular to provide methods in accordance with the present invention, other configurations are possible, either with or without one or more video communication services 110.

[0059] Thus, Figure 5 illustrates a method according to the invention for providing a shared digital video stream. Figures 6A to 6F show the states of the different digital video / audio data streams resulting from the method steps shown in Figure 5.

[0060] In a first step, the method begins.

[0061] In a subsequent collection step, a respective plurality of primary digital video streams 210, 301 are collected, for example by a collection function 131, from at least two of said digital video sources 120, 300. Each such plurality of primary data streams 210, 301 may comprise an audio portion 214 and / or a video portion 215. It is understood that "video" in this context denotes the moving image and / or still image content of such data streams. Each primary data stream 210, 301 may be encoded according to any video / audio encoding standard (using the respective codec used by the entity providing said primary stream 210, 301), and the encoding format may differ between different ones of said plurality of primary streams 210, 301 used simultaneously in one and the same video communication. At least one, for example all, of the plurality of primary data streams 210, 301 are preferably provided as a stream of binary data, possibly in a data container data structure that is itself conventional. Preferably, at least one, such as at least two, or even all, of the multiple primary data streams 210, 301 are provided as respective live video recordings.

[0062] It should be noted that the multiple primary data streams 210, 301 may not be synchronized in time when they are received by the collection function 131. This may mean that they are associated with different latencies or delays relative to each other. For example, if two primary video streams 210, 301 are live recordings, this may mean that they are associated with different latencies relative to the recording time when they are received by the collection function 131.

[0063] It should also be noted that the multiple primary data streams 210, 301 may themselves be respective live camera feeds from webcams; a screen or presentation currently being shared; a film clip being viewed; or any combination of these arranged in various ways within one and the same screen.

[0064] The collection steps are illustrated in Figures 6a and 6b. Figure 6b also illustrates how the collection function 131 can store each primary video stream 210, 301 as bundled audio / video information or as audio stream data separated from the associated video stream data. Figure 6b illustrates how the data of the primary video streams 210, 301 is stored as individual frames 213 or collections / clusters of frames, where a "frame" refers here to a time-limited portion of image data and / or any associated audio data, e.g., each frame being an individual still image or a continuous series of images (e.g., a series of images that constitutes up to one second of moving images) that together form the moving image video content.

[0065] In a subsequent event detection step performed by the event detection functionality 132, the multiple primary digital video streams 210, 301 are analyzed by the event detection functionality 132, particularly the AI ​​component 132a etc., to detect at least one event 211 selected from the first set of events. This is illustrated in Figure 6c.

[0066] This event detection step is preferably performed for at least one, e.g. at least two, e.g. all, of the primary video streams 210, 301, and individually for each of said primary video streams 210, 301. In other words, the event detection step is preferably performed for each of said individual primary video streams 210, 301, taking into account only information contained as part of that particular primary video stream 210, 301, and in particular without taking into account information contained as part of other primary video streams. Furthermore, event detection is preferably performed without taking into account any common time reference 260 associated with multiple primary video streams 210, 301.

[0067] However, preferably, event detection takes into account information contained as part of the individually analyzed primary video stream over a time interval, for example over a historical time interval of the primary video stream that is greater than 0 seconds, for example at least 0.1 seconds, for example at least 1 second.

[0068] Event detection may take into account information contained in the audio and / or video data included as part of the primary video stream 210,301.

[0069] The first set of events may include any number of types of events, such as a change in a slide in a slide presentation that constitutes or is part of the primary video stream 210, 301, a change in connection quality of the source 120, 300 providing the primary video stream 210, 301 that results in a change in image quality, loss of image data, or reacquisition of image data, and physical events of movement detected in the primary video stream 210, 301, such as movement of a person or object in the video, a change in lighting in the video, a sudden sharp noise in the audio, or a change in audio quality. It should be understood that this is not intended to be an exhaustive list, and that these examples are provided to understand the applicability of the presently described principles.

[0070] In a subsequent synchronization step performed by the synchronization function 133, the primary digital video streams 210, 310 are time-synchronized to a common time reference 260. As shown in Fig. 6d, this time synchronization involves aligning the primary video streams 210, 301 to each other using the common time reference 260 so that they can be combined to form a time-synchronized context. The common time reference 260 may be a stream of data, a heartbeat signal or other pulse data, or a time anchor that is applicable to each of the individual primary video streams 210, 301. What is important is that by making the common time reference applicable to each of the individual primary video streams 210, 301, the information content of the primary video streams 210, 301 can be uniquely related to the common time reference with respect to a common time axis. In other words, the common time reference aligns the primary video streams 210, 301 to be time-synchronized in a present sense via time shifting.

[0071] As shown in FIG. 6d, time synchronization may include determining one or more timestamps 261 relative to a common time reference 260 for each of the multiple primary video streams 210, 301.

[0072] In a subsequent pattern detection step performed by pattern detection function 134, the multiple time synchronized primary digital video streams 210, 301 are analyzed to detect at least one pattern 212 selected from the first pattern set. This is shown in Figure 6e.

[0073] In contrast to the event detection step, the pattern detection step is preferably performed based on video and / or audio information included as part of at least two of the multiple time-synchronized primary video streams 210,301.

[0074] The first set of patterns may include any number of types of patterns, such as multiple participants speaking in turn or simultaneously, or a change in a presentation slide occurring simultaneously as another event, such as another participant speaking, etc. This list is not exhaustive but is exemplary.

[0075] In alternative embodiments, the detected pattern 212 may relate to information contained in only one of the multiple primary video streams 210, 301, rather than information contained in more than one of the multiple primary video streams 210, 301. In such cases, such pattern 212 is preferably detected based on video and / or audio information contained in that single primary video stream 210, 301 spanning at least two detected events 211, e.g., two or more consecutive detected presentation slide changes or connection quality changes. As an example, multiple consecutive slide changes that rapidly follow one another over time may be detected as one single slide change pattern, as opposed to one distinct slide change pattern for each detected slide change event.

[0076] It is understood that the first set of events and the first set of patterns may comprise a predefined type of event / pattern defined using a respective set of parameters and parameter intervals. As described below, the set of events / patterns may also be defined and detected using various AI tools.

[0077] In a subsequent generation step performed by the generation function 135, a shared digital video stream is generated as an output digital video stream 230 based on a plurality of consecutively considered frames 213 of a plurality of time-synchronized primary digital video streams 210, 301 and the detected pattern 212.

[0078] As explained and detailed below, the present invention allows for fully automated generation of the output digital video stream 230.

[0079] For example, such generation may include selection of what video and / or audio information from which primary video streams 210, 301 to use in the output video stream 230, and to what extent; the video screen layout of the output video stream 230; the switching pattern between different such uses or layouts over time; etc.

[0080] This is also illustrated in Figure 6f, which shows one or more additional portions of time-related (relative to the common time reference 260) digital video information 220, such as additional digital video information streams that may be time-synchronized with the common time reference 260 and used in concert with the time-synchronized multiple primary video streams 210, 301 in generating the output video stream 230. For example, the additional streams 220 may include information regarding any video and / or audio special effects to use, such as dynamically based on detected patterns; a planned time schedule for the video communication; etc.

[0081] In a subsequent publishing step performed by the publishing function 136, the generated output digital video stream 230 is continuously provided to the consumers 110, 150 of the shared digital video stream, as described above.

[0082] In the subsequent steps, the method ends. However, initially, the method may be repeated any number of times to generate the output video stream 230 as a continuously provided stream, as shown in FIG. 5. Preferably, the output video stream 230 is generated to be consumed in real-time or near real-time (taking into account the sum of the latencies added by all steps along the way) and continuously (published as soon as more information becomes available, but not counting the intentionally added latencies described below). In this way, the output video stream 230 may be consumed in an interactive manner, whereby the output video stream 230 is fed back to the video communication service 110 or to other contexts that form the basis for the generation of the primary video stream 210 that is fed back to the collection function 131 to form a closed feedback loop; or the output video stream 230 is consumed in a different context (outside the system 100, or at least outside the central server 130), where it may form the basis for real-time two-way video communication.

[0083] As mentioned above, in some embodiments, at least two of the multiple primary digital video streams 210, 301 are provided as part of a shared digital video communication such as provided by a video communication service 110, which video communication includes respective remotely connected participant clients 121 providing said primary digital video streams 210. In such cases, the collecting step may consist of collecting at least one of said primary digital video streams 210 from the shared digital video communication service 110 itself, via an auto-participant client 140 that is in turn granted access to video and / or audio stream data from within said video communication service 110, and / or via the API 112 of the video communication service 110.

[0084] Additionally, in this and other cases, the collecting step may comprise collecting at least one of said plurality of primary digital video streams 210, 301 as a respective external digital video stream 301 collected from an information source 300 that is external to the shared digital video communication service 110. It should be noted that one or more of such external video sources 300 may be external to the central server 130.

[0085] In some embodiments, the multiple primary video streams 210, 301 are not formatted in the same manner. Such different formats could be the formats in which they are provided to the collection function 131 in different types of data containers (such as AVI or MPEG), but in preferred embodiments, at least one of the multiple primary video streams 210, 301 is formatted according to a deviating format (with respect to at least one other of the primary video streams 210, 301) in that the deviating primary digital video streams 210, 301 have deviating video encodings; deviating fixed or variable frame rates; deviating aspect ratios; deviating video resolutions; and / or deviating audio sample rates.

[0086] The collection function 131 is preferably pre-configured to read and interpret all encoding formats, container standards, etc. occurring in all collected primary video streams 210, 301. This allows processing as described herein to be performed without requiring decoding until a relatively later stage in these processes (such as until the primary streams in question are in their respective buffers; or until after the event detection step; etc.). However, in the rare case where one or more of the primary video feeds 210, 301 are encoded using a codec that the collection function 131 cannot interpret without decoding, the collection function 131 may be configured to perform decoding and analysis of such primary video streams 210, 301, followed by conversion to a format that can be processed, for example, by the event detection function. Note that even in this case, it is preferable not to perform re-encoding at this stage.

[0087] For example, a primary video stream 220 fetched from a multi-party video event, such as that provided by the video communication service 110, typically has a requirement for low latency and is therefore typically associated with variable frame rates and variable pixel resolutions to enable participants 122 to communicate effectively. In other words, the overall video and audio quality is degraded as necessary for low latency.

[0088] On the other hand, the external video feed 301 typically has a more stable frame rate and higher image quality, but may therefore have a higher delay.

[0089] Thus, the video communication service 110 may at each point in time use a different encoding and / or container than the external video source 300. Thus, the analysis and video generation process described herein must combine these multiple streams 210, 301 of different formats into a new single stream for a combined experience.

[0090] As mentioned above, the collection functionality 131 may comprise a set of format-specific collection functionality 131a, each configured to process a particular type of format of the primary video streams 210, 301. For example, each one of these format-specific collection functionality 131a may be configured to process multiple primary video streams 210, 301 encoded using different respective video encoding methods / codecs, such as Windows® Media® or DivX®.

[0091] However, in a preferred embodiment, the collecting step involves converting at least two, eg, all, of the multiple primary digital video streams 210 , 301 to a common protocol 240 .

[0092] As used in this context, the term "protocol" refers to an information structuring standard or data structure that specifies how to store the information contained in the digital video / audio stream. However, the common protocol preferably does not prescribe how to store the digital video and / or audio information, for example at a binary level (i.e., encoded / compressed data that indicates the sounds and images themselves), but instead forms a structure of a predefined format for storing such data. In other words, the common protocol prescribes storing digital video data in a raw binary format without performing any digital video decoding or encoding in connection with such storage, and possibly without modifying the existing binary format in any way apart from concatenating and / or splitting the binary format byte strings. Instead, the raw (encoded / compressed) binary data content of said primary video stream 210, 301 is preserved while repacking said raw binary data in a data structure defined by the protocol. In some embodiments, the common protocol defines a video file container format.

[0093] FIG. 7 shows, by way of example, multiple primary video streams 210, 301 as shown in FIG. 6a reconstructed by respective format-specific acquisition functions 131a and using the common protocol 240 described above.

[0094] Thus, the common protocol 240 provides for storing digital video and / or audio data in data sets 241 that are preferably divided into discrete, contiguous sets of data along a time axis relative to the primary video stream in question 210, 301. Each such data set may contain one or several frames of video and associated audio data.

[0095] The common protocol 240 may also provide for storing, in association with the stored digital video and / or audio data set 241, metadata 242 associated with a specified point in time.

[0096] The metadata 242 may include information about the binary format of the raw data of the primary digital video stream 210, such as about the digital video encoding method or codec used to generate the binary data of the raw data, the resolution of the video data, the video frame rate, the frame rate variation flag, the video resolution, the video aspect ratio, the audio compression algorithm, or the audio sampling rate. The metadata 242 may also include information about the timestamps of the stored data relative to the time base of the primary video stream 210, 301.

[0097] The use of format-specific collection functions 131a in combination with the common protocol 240 allows for rapid collection of the information content of the primary video streams 210, 301 without the added latency (delay) of decoding / re-encoding the received video / audio data.

[0098] The collecting step may thus comprise collecting a plurality of primary digital video streams 210, 301 encoded using different binary video and / or audio encoding formats using different ones of the plurality of format specific collecting functions 131a in order to parse said primary video streams 210, 301 and store the parsed raw binary data, together with any associated metadata, in a data structure using a common protocol. Obviously, the decision as to which format specific collecting function 131a to use for which primary video stream 210, 301 may be performed by the collecting function 131 based on predefined and / or dynamically detected characteristics of each of said primary video streams 210, 301.

[0099] Each primary video stream 210, 301 collected in this manner may be stored in its own separate memory buffer, such as a RAM memory buffer within the central server 130.

[0100] The conversion of the primary video streams 210, 301 performed by each format-specific collection function 131a may therefore comprise splitting the raw binary data of each primary digital video stream 210, 301 thus converted into an ordered set of smaller data sets 241.

[0101] Furthermore, the conversion may also comprise associating each of the smaller sets 241 (or a subset, for example a subset regularly distributed along a time axis of each of the primary streams 210, 301 in question) with a respective time of the common time reference 260. This association may be performed by analysis of the binary video and / or audio data of the raw data, in any of the principle methods described below or in other ways, and may be performed in order to be able to perform a subsequent time synchronization of the primary video streams 210, 301. Depending on the type of common time reference 260 used, at least a part of this association of each data set 241 may also be performed by or instead of the synchronization function 133. In the latter case, the collecting step may instead comprise associating each of the smaller sets 241, or a subset thereof, with a respective time of a time axis specific to the primary stream 210, 301 in question.

[0102] In some embodiments, the collecting step also includes converting the raw binary video and / or audio data collected from the multiple primary video streams 210, 301 to a uniform quality and / or updating frequency. This may include downsampling or upsampling the raw binary digital video and / or audio data of the multiple primary digital video streams 210, 301 to a common video frame rate; a common video resolution; or a common audio sampling rate, as appropriate. It should be noted that such resampling can be performed without performing a full decoding / re-encoding, or even without performing any decoding at all, since the format-specific collecting function 131a can directly process the raw binary data according to the correct binary encoding target format.

[0103] Preferably, each of the multiple primary digital video streams 210, 301 is stored in an individual data storage buffer 250 as an individual frame 213 or a sequence of frames 213, as described above, and each is associated with a corresponding timestamp that is in turn associated with a common time reference 260.

[0104] In a specific example provided for illustrative purposes, the video communication service 110 is Microsoft® Teams® and is conducting a video conference involving multiple simultaneous participants 122. The auto-join client 140 is registered as a conference participant in the Teams® conference.

[0105] The primary video input signals 210 are then provided to the collection function 130 via the auto-join client 140 and are acquired by the collection function 130. These are raw data signals in H264 format and include timestamp information for each video frame.

[0106] The associated format-specific collection function 131a picks up the raw data over IP (cloud LAN network) on a configurable predefined TCP port. Every Teams® meeting participant and associated audio data is associated with a separate port. The collection function 131 then uses the timestamp from the audio signal (50Hz) and downsamples the video data to a fixed output signal of 25Hz before storing the video streams 220 in their respective individual buffers 250.

[0107] As mentioned above, the common protocol 240 stores data in a raw binary format. It can be designed to process, at a very low level, the raw bits and bytes of video / audio data. In a preferred embodiment, the data is stored in the common protocol 240 as a simple byte array or corresponding data structure (such as a slice). This means that the data does not need to be put into a traditional video container at all (the common protocol 240 does not constitute such a traditional container in this context). Also, video encoding and decoding is computationally heavy, thus inducing delays and requiring expensive hardware. Moreover, this problem scales with the number of participants.

[0108] The common protocol 240 allows for reserving memory in the collection function 131 for the primary video stream 210 associated with each Teams® conference participant 122 and any external video sources 300, and changing the amount of allocated memory on the fly during the process. In this way, it is possible to change the number of input streams, so that each buffer can remain valid. For example, information such as resolution, frame rate, etc., is variable, but is stored as metadata in the common protocol 240, so this information can be used to quickly change the size of each buffer as needed.

[0109] The following is an example of a specification for this type of common protocol 240:

[0110] [Table 1]

[0111] In the above table, the "Detected event in, if any" data is included as part of the specification of common protocol 260. However, in some embodiments, this information (regarding detected events) may instead be placed in a separate memory buffer.

[0112] In some embodiments, the at least one additional portion of the digital video information 220, which may be an overlay or effect, is also stored in each individual buffer 250 as individual frames or sequences of frames each associated with a corresponding timestamp that is in turn associated with the common time reference 260.

[0113] As illustrated above, the event detection step may include using a common protocol 240 to store metadata 242 describing the detected event 211 in association with the primary digital video stream 210, 301 in which the event 211 was detected.

[0114] Event detection can be performed in different ways. In some embodiments performed by the AI ​​component 132a, the event detection step includes a first trained neural network or other machine learning component individually analyzing at least one, e.g., some or all, of the multiple primary digital video streams 210, 301 to automatically detect any of said events 211. This may include the AI ​​component 132a classifying the data of the primary video streams 210, 301 into a set of predefined events in a supervised classification and / or into a dynamically determined set of events in an unsupervised classification.

[0115] In some embodiments, the detected event 211 is a change in a presentation slide of a presentation that is, or is contained in, the primary video stream 210, 301.

[0116] For example, if a presenter of a presentation decides to change the slide of the presentation that he or she is currently making to the audience, this means that what is interesting for a given viewer may change. The newly displayed slide may just be a general-level image that is best viewed for a short time in a so-called "butterfly" mode (e.g., displaying the slide side-by-side with the video of the presenter in the output video stream 230). Or the slide may contain a lot of detail, text with a small font size, etc. In the latter case, the slide will be displayed full screen and will be displayed for a somewhat longer time than would normally be the case. The slide in this case may be more interesting to the viewer of the presentation than the face of the presenter, so butterfly mode may not be as appropriate.

[0117] In practice, the event detection step consists of at least one of the following:

[0118] Firstly, the event 211 may be detected based on image analysis of the difference between a first image of a detected slide and a subsequent second image of the detected slide. The nature of the primary video stream 220, 301 as being indicative of a slide may be determined automatically using digital image processing which is per se conventional, such as using motion detection combined with OCR (Optical Character Recognition).

[0119] This may involve using automatic computer image processing techniques to check whether the detected slide has changed sufficiently to be classified as a real slide change. This can be done by checking the delta between the current slide and the previous slide in terms of RGB color values. For example, one can evaluate how much the RGB values ​​have changed globally in the screen area covered by the slide in question, and at the same time evaluate whether it is possible to find groups of adjacent pixels that change in concert with this. In this way, relevant slide changes can be detected, while filtering out irrelevant changes, such as, for example, computer mouse movements across the screen. This approach allows for full composability. For example, it may be desirable to be able to capture computer mouse movements, for example if a presenter wants to present something in detail while pointing to different things with the computer mouse.

[0120] Second, the event 211 may be detected based on image analysis of the information complexity of the second image itself to determine the type of event with greater specificity.

[0121] This might involve, for example, assessing the total amount of textual information on the slide in question and the associated font sizes. This can be done using traditional OCR methods, including deep learning-based character recognition techniques.

[0122] Note that because the raw binary format of the evaluated video streams 210, 301 is known, this may be performed directly in the binary domain without first decoding or re-encoding the video data. For example, the event detection function 132 may invoke an associated format-specific collection function for an image interpretation service, or the event detection function 132 itself may include functionality for evaluating image information, such as to the individual pixel level, for a number of different supported raw binary video data formats.

[0123] In another example, the detected event 211 is a loss of a communication connection of a participant client 121 to the digital video communication service 110. In this case, the detecting step may include detecting that the participant client 121 has lost the communication connection based on image analysis of a series of subsequent video frames 213 of the primary digital video stream 210 corresponding to that participant client 121.

[0124] Because participant clients 121 are associated with different physical locations and different Internet connections, it may occur that someone loses connection to the video communication service 110 or the central server 130. In such a situation, it is desirable to avoid a black or blank screen appearing in the generated output video stream 230.

[0125] Alternatively, such loss of connection can be detected as an event by the event detection function 132, for example by applying a two-class classification algorithm where the two classes used are connected / not connected (no data). In this case, "no data" is understood to be different from the presenter intentionally sending a black screen. Since a short duration black screen, such as just one or two frames, may not be noticeable in the final generated stream 230, the two-class classification algorithm can be applied over time to create a time series. A threshold specifying the minimum length of a connection interruption can then be used to determine whether a connection has been lost.

[0126] As described below, detected events of the types illustrated above may be used by pattern detection function 134 to take various responses, as appropriate and desired.

[0127] As noted above, the individual primary video streams 210, 301 are each associated with a common time reference 260 and can be time-synchronized relative to one another by the synchronization function 133.

[0128] In some embodiments, the common time reference 260 is based on or comprises a common audio signal 111 (see Figures 1 to 3), which, as described above, is common to a shared digital video communication service 110 participating in at least two remotely connected participant clients 121, each providing a respective one of the primary digital video streams 210.

[0129] In the Microsoft® Teams® example discussed above, a common audio signal may be generated and captured by the central server 130 via the auto-join client 140 and / or via the API 112. In this and other examples, such a common audio signal may be used as a heartbeat signal to time-synchronize the individual primary video streams 220 by combining the individual primary video streams 220 at specific times based on the heartbeat signal. Such a common audio signal may be provided as a separate (with respect to each of the other primary video streams 210) signal, such that each of the other primary video streams 210 may be individually time-correlated to the common audio signal based on audio contained in the other primary video streams 210 or based on image information contained therein (such as using automatic image processing-based lip-sync techniques).

[0130] In other words, to handle the variable and / or different latencies associated with the individual primary video streams 210 and to achieve time synchronization of the combined video output stream 230, such a common audio signal is used as a heartbeat for all primary video streams 210 within the central server 130 (but perhaps not the external primary video stream 301). In other words, all other signals are mapped to this common audio time heartbeat to make sure they are all time synchronized.

[0131] In another example, time synchronization is achieved using a time synchronization element 231 that is introduced into the output digital video stream 230 and detected by a respective local time synchronization software function 125 provided as part of one or more individual ones of the participant clients 121. The local software function 125 is configured to detect the time of arrival of the time synchronization element 231 in the output video stream 230. As will be appreciated, in such an embodiment, the output video stream 230 is fed back to the video communication service 110 or otherwise made available to each participant client 121 and its local software function 125.

[0132] For example, the time synchronization elements 231 may be visual markers, such as pixels that change color in a predetermined order or manner, that are placed or updated in the output video 230 at regular time intervals; a visual clock that is updated and displayed in the output video 230; an audio signal (which may be designed to be inaudible to the participants 122, for example, by having a sufficiently low amplitude and / or a sufficiently high frequency) that is added to the audio that forms part of the output video stream 230. The local software functionality 125 is configured to automatically detect the arrival time of each of the time synchronization elements (of each) 231 using appropriate image and / or audio processing.

[0133] The common time reference 260 may then be determined based, at least in part, on the detected arrival times. For example, each of the local software functions 125 may communicate respective information indicative of the detected arrival times to the central server 130.

[0134] Such communication may occur via a direct communication link between the participant client 121 and the central server 130. However, communication may also occur via a primary video stream 210 associated with the participant client 121. For example, the participant client 121 may introduce a visual or audible code, such as the type described above, into the primary video stream 210 generated by the participant client 121 for automatic detection by the central server 130 and use to determine the common time reference 260.

[0135] In yet further embodiments, each participant client 121 may perform image detection in a common video stream viewable by all participant clients 121 for the video communication service 110 and relay the results of such image detection to the central server 130 in a manner corresponding to that described above, where they are used to determine the respective offsets of each participant client 121 relative to one another over time. In this manner, a common time reference 260 may be determined as a set of individual relative offsets. For example, a selected reference pixel of the commonly available video stream may be monitored by some or all of the participating clients 121, such as by local software functions 125, and the current color of that pixel may be communicated to the central server 130. The central server 130 may generate an estimated set of relative time offsets across the different participating clients 121 by calculating respective time series based on such color values ​​received successively from each of many (or all) of the participating clients 121 and performing cross-correlation.

[0136] In practice, the output video stream 230 provided to the video communication service 110 may be included as part of the shared screen of all participant clients of that video communication and may therefore be used to evaluate such time offsets associated with the participant clients 121. In particular, the output video stream 230 provided to the video communication service 110 may be made available again to the central server via the auto-join client 140 and / or API 112.

[0137] In some embodiments, the common time reference 260 may be determined at least in part based on a detected discrepancy between an audio portion 214 of a first one of the multiple primary digital video streams 210, 301 and an image portion 215 of said first one of the multiple primary digital video streams 210, 301. Such a discrepancy may for example be based on a digital lip-sync video image analysis of a speaking participant 122 viewed in said first primary digital video stream 210, 301. Such a lip-sync analysis may be conventional per se and may for example use a trained neural network. The analysis may be performed by the synchronization function 133 for each primary video stream 210, 301 in relation to the available common audio information, and a relative offset across the individual primary video streams 210, 301 may be determined based on this information.

[0138] In some embodiments, the synchronization step includes intentionally introducing a latency of up to 30 seconds, e.g. up to 5 seconds, e.g. up to 1 second, e.g. up to 0.5 seconds, but more than 0 seconds, to provide at least that latency in the output digital video stream 230. Whatever the length, the intentionally introduced latency is at least a number of video frames, such as at least 3, or at least 5, or even 10, such as this number of frames (or individual images) stored after any resampling in the collection step. As used herein, the term "intentionally" means that the latency is introduced independently of the need to introduce such latency based on synchronization issues or the like. In other words, the intentionally introduced latency is in addition to the latency introduced as part of the synchronization of the primary video streams 210, 301, in order to time-synchronize the multiple primary video streams 210, 301 with each other. The intentionally introduced latency may be predetermined, fixed, or variable with respect to the common time reference 260. Latency may be measured relative to the least latent one of the multiple primary video streams 210, 301, and as a result of the above time synchronization, the more latent ones of these streams 210, 301 may be associated with relatively small intentionally added latency.

[0139] In some embodiments, a relatively small latency, such as 0.5 seconds or less, is introduced that is barely noticeable to participants in the video communication service 110 using the output video stream 230. In other embodiments, a larger latency may be introduced, such as when the output video stream 230 is not used in an interactive context, but instead is exposed in a one-way communication to an external consumer 150.

[0140] This intentionally introduced latency achieves sufficient time for the synchronization function 133 to map the collected video frames of the individual primary streams 210, 301 to the correct common time reference 260 timestamps 261. It also provides sufficient time to perform the event detection described above to detect lost primary stream 210, 301 signals, slide changes, resolution changes, etc. Additionally, the intentionally introduced latency allows for improvements to the pattern detection function 134, as described below.

[0141] It will be appreciated that introducing latency involves buffering 250 each of the collected and time-synchronized multiple primary video streams 210, 301 before publishing the output video stream 230 using that buffered frame 213. In other words, the video and / or audio data of at least one, some or all of the multiple primary video streams 210, 301 will be present in the central server 130 in a buffered manner for the reasons discussed above, in particular for use by the pattern detection function 134, rather than being used like a cache, but with the intent of being able to handle varying bandwidth situations (as in a traditional cache buffer).

[0142] Thus, in some embodiments, the pattern detection step involves considering certain information of at least one, e.g. some or all, of the multiple primary digital video streams 210, 301, which is present in a frame 213 that is later than a frame of the time-synchronized primary digital video stream 210 that has not yet been used in generating the output digital video stream 230. Thus, the newly added frame 213 resides in said buffer 250 for a certain latency period before forming part of (or the basis of) the output video stream 230. During this period, the information of said frame 213 constitutes "future" information in relation to the frame currently being used to generate the current frame of the output video stream 230. When the timeline of the output video stream 230 reaches said frame 213, said frame is used in generating the corresponding frame of the output video stream 230 and may thereafter be discarded.

[0143] In other words, the pattern detection function 134 has at its disposal a set of video / audio frames 213 that have not yet been used to generate the output video stream 230, and uses this data to detect said patterns.

[0144] Pattern detection can be performed in different ways: In some embodiments performed by the AI ​​component 134a, the pattern detection step includes a second trained neural network or other machine learning component analyzing in concert at least two, e.g., at least three, or even all, of the multiple primary digital video streams 120, 301 to automatically detect the pattern 212.

[0145] In some embodiments, the detected pattern 212 comprises a speech pattern including at least two different active participants 122, each associated with a respective participant client 121 for the shared video communication service 110, each of which is visually viewed in a respective one of the multiple primary digital video streams 210, 301.

[0146] Preferably, the generating step includes determining, tracking, and updating a current generating state of the output video stream 230. For example, such state may dictate which participants 122 (if any) are visible in the output video stream 230 and where on screen they are visible; which external video streams 300 are visible in the output video stream 230 and where on screen they are visible; which slides or shared screens are displayed in full screen mode or in combination with any live video streams; etc. Thus, the generating function 135 can be viewed as a state machine for the generated output video stream 230.

[0147] In order to generate the output video stream 230 as a combined video experience to be viewed, for example, by the end consumer 150, it is advantageous for the central server 130 to be able to understand what is happening at a deeper level than simply detecting individual events associated with the individual primary video streams 210, 301.

[0148] In a first example, the presenting participant client 121 changes the currently displayed slide. This slide change is detected by the event detection function 132 as described above, and metadata 242 is added to the frame indicating that a slide change has occurred. This happens many times as the presenting participant client 121 is found to be skipping forward a number of slides in rapid succession, resulting in a series of "slide change" events that are also detected by the detection function 132 and stored with the corresponding metadata 242 in a separate buffer 250 of the primary video stream 210. In practice, each such rapidly skipped forward slide may only be displayed for a few seconds.

[0149] The pattern detection function 134 looks at the information in the buffer 250 across these detected slide changes and detects a pattern that corresponds to one single slide change rather than multiple or rapidly executed slide changes (i.e., a single slide change to the last slide in the forward skip, where the last slide remains visible once the fast skip ends). In other words, the pattern detection function 134 notes that there were, for example, ten slide changes in a very short time, why treat them as a detected pattern that means one single slide change. As a result, the generation function 135 has access to the pattern detected by the pattern detection function 134 and can choose to display this last slide in full screen mode for a few seconds in the output video stream 230 because it determines that this last slide is potentially important in the state machine. It can also choose not to display the intermediately viewed slides in the output stream 230 at all.

[0150] Detection of patterns having multiple rapid slide changes may be detected by a simple rule-based algorithm, but alternatively may be detected using a neural network designed and trained to detect such patterns in video images by classification.

[0151] In another example, it may be desirable to quickly switch visual attention between current speakers, while still providing a relevant viewing experience for the consumer 150 by generating and presenting a calm and smooth output video stream 230, which may be useful, for example, when the video communication is a talk show, panel discussion, or the like. In this case, the event detection function 132 may continuously analyze each primary video stream 210, 301 to determine at any time whether the person being viewed in that particular primary video stream 210, 301 is currently speaking or not. This may be performed as described above, for example, using image processing tools conventional per se. The pattern detection function 134 may then be operable to detect certain overall patterns involving multiple primary video streams 210, 301, which patterns are useful for generating a smooth output video stream 230. For example, the pattern detection function 134 may detect a pattern of very frequent switches between current speakers and / or a pattern involving multiple simultaneous speakers.

[0152] The generation functionality 135 can then take such detected patterns into account when making automated decisions related to the generation state, such as, for example, not automatically switching visual focus to a speaker who speaks for only half a second before going silent again, or switching to a state in which multiple speakers are displayed side-by-side during a period in which they are alternating or speaking simultaneously. This state determination process can itself be performed using time series pattern recognition techniques or using trained neural networks, but can also be based at least in part on a predetermined set of rules.

[0153] In some embodiments, there may be multiple patterns that are detected in parallel and form input to the state machine of the generation function 135. Such multiple patterns may be used by the generation function 135 in different AI components, computer vision detection algorithms, etc. As an example, a permanent slide change may be detected while simultaneously detecting unstable connections of some participant clients 121, while other patterns detect the current main speaking participant 122. Using all such available pattern data, a classifier neural network may be trained and / or a set of rules may be developed to analyze the time series of such pattern data. Such classification may be supervised, at least in part, e.g., completely, to result in determined desired state changes used in the generation. For example, different such predefined classifiers may be generated that are specifically configured to automatically generate the output video stream 230 according to various different generation styles and desires. Training may be based on known generation state change sequences as the desired output, and known pattern time series data as training data. In some embodiments, a Bayesian model may be used to generate such classifiers. In a concrete example, a priori information can be obtained from experienced producers, who can provide input such as "In talk shows, we never switch directly from speaker A to speaker B, but always give an overview first before focusing on other speakers, unless the other speaker is very dominant and speaks loudly". This generation logic is expressed as a Bayesian model of the general form "If X is true | given the fact that Y is true | do Z". The actual detection (e.g. whether someone is speaking loudly) can be done using classifiers or threshold-based rules.

[0154] Given a large dataset (of pattern time series data), deep learning techniques can be used to develop correct and compelling generative formats for use in the automatic generation of video streams.

[0155] In summary, by using a combination of event detection based on multiple individual primary video streams 210, 301; intentionally introduced latency; pattern detection based on multiple time-synchronized primary video streams 210, 301 and detected events; and a generation process based on detected patterns, it is possible to achieve automatic generation of the output digital video stream 230 according to a wide range of possible choices of taste and style. This result is valid across a wide range of possible neural network and / or rule-based analysis techniques used by the event detection function 132, pattern detection function 134, and generation function 135.

[0156] As exemplified above, the generating step may include generating the output digital video stream 230 based on a set of predetermined and / or dynamically variable parameters relating to the visibility of each one of the multiple primary digital video streams 210, 301 in the output digital video stream 230, the arrangement of the visual and / or auditory video content, the visual or auditory effects used, and / or the output mode of the output digital video stream 230. Such parameters may be automatically determined by a state machine in the generating functionality 135, and / or set by an operator controlling the generation (semi-automated), and / or predetermined based on some a priori configurational desires (such as a minimum time between layout changes of the output video stream 230 or state changes of the types exemplified above).

[0157] In a practical example, the state machine may support a set of predefined standard layouts that may be applied to the output video stream 230, such as a full-screen presenter view (showing the currently speaking participant 122 in full screen); a slide view (showing the currently shared presentation slide in full screen); a "butterfly view" (showing both the currently speaking participant 122 and the currently shared presentation slide in a side-by-side view); a multi-speaker view (showing all or a selected subset of the participants 122 side-by-side or in a matrix layout). Various available production formats may be defined by a set of state machine state change rules (as exemplified above) along with a set of available states (such as the set of standard layouts above). For example, one such production format may be "panel discussion", another may be "presentation", etc. By selecting a particular production format via a GUI or other interface to the central server 130, an operator of the system 100 may quickly select one of a set of such predefined production formats, and then enable the central server 130 to generate the output video stream 230 in accordance with that production format in a fully automatic manner based on available information as described above.

[0158] Furthermore, during production, for each conference participant client 121 or external video source 300, a respective in-memory buffer is created and maintained, as described above. These buffers can be easily deleted, added, and modified on the fly. The central server 130 may then be configured to receive information regarding added / dropped-off participant clients 121 and participants 122 scheduled to speak, scheduled or unexpected pauses / resumes of the presentation, desired changes to the currently used production format, etc., during the production of the output video stream 230. Such information may be provided to the central server 130, for example, via an operator GUI or interface, as described above.

[0159] As illustrated above, in some embodiments, at least one of the multiple primary digital video streams 210, 301 may be provided to a digital video communication service 110, and the publishing step may then include providing the output digital video stream 230 to that same communication service 110. For example, the output video stream 230 may be provided to a participant client 121 of the video communication service 110, or may be provided as an external video stream to the video communication service 110 via the API 112. In this manner, the output video stream 230 may be made available to multiple or all of the participants of the video communication event currently being facilitated by the video communication service 110.

[0160] As also mentioned above, the output video stream 230 may additionally or alternatively be provided to one or more external consumers 150 .

[0161] Generally, the generating step is performed by a central server 130, and the output digital video stream 230 can be provided as a live video stream via an API 137 to one or more concurrent consumers.

[0162] The present invention also relates to computer software functionality for providing a shared digital video stream in accordance with the above disclosure, such computer software functionality being configured, when executed, to perform the collection, event detection, synchronization, pattern detection, generation and publishing steps described above, the computer software functionality being configured to run on the physical or virtual hardware of the central server 130 as described above.

[0163] The present invention also relates to such a system 100 for providing shared digital video streams, which in turn comprises a central server 130. The central server 103 is in turn configured to perform the above-mentioned steps of collection, event detection, synchronization, pattern detection, generation and publishing, for example by the central server 130 executing computer software functions for performing the above-mentioned steps.

[0164] Although preferred embodiments have been described above, it will be apparent to those skilled in the art that many modifications can be made to the disclosed embodiments without departing from the essential concepts of the invention.

[0165] For example, many additional features not described herein may be provided as part of the system 100 described herein. In general, the solutions described herein provide a framework within which detailed functionality and features can be built to accommodate a wide variety of specific applications in which streams of video data are used for communication.

[0166] One example is a demonstration situation, where the primary video stream includes the presenter's view, a shared digital slide-based presentation, and a live video of the product being demonstrated.

[0167] Another example is an educational situation, where the primary video stream includes a teacher's view, live video of a physical entity that is the topic of instruction, and live video of each of multiple students asking questions and interacting with the teacher.

[0168] In either of these two examples, a video communication service (which may or may not be part of the system) may provide one or more of the primary video streams and / or some of the primary video streams may be provided as external video sources of the type disclosed herein.

[0169] In general, anything disclosed with respect to the present method is applicable to the present system and computer software product, and vice versa.

[0170] Therefore, the invention is not limited to the described embodiments, but can be modified within the scope of the appended claims.

Claims

1. 1. A method for providing a shared digital video stream, said method comprising the steps of: In a collecting step, collecting a plurality of primary digital video streams (210) from at least two digital video sources (120), respectively; In an event detection step, the plurality of primary digital video streams (210) are individually analyzed to detect at least one event (211) selected from a first set of events; In a synchronizing step, the plurality of primary digital video streams (210) are time-synchronized with respect to a common time reference (260); a pattern detection step of analyzing the plurality of time-synchronized primary digital video streams (210) to detect at least one pattern (212) selected from a first pattern set, wherein the detection of the pattern is based on the at least one detected event (211); generating, in a generating step, the shared digital video stream as an output digital video stream (230) based on consecutively considered frames (213) of the time-synchronized plurality of primary digital video streams (210) and the detected pattern (212); and In a publishing step, the output digital video stream (230) is continuously provided to consumers of the shared digital video stream, wherein: The detection of the pattern is based, first, on video and / or audio information included as part of at least two of the time-synchronized primary video streams (210) considered together, or, second, on video and / or audio information included in a single primary video stream spanning at least two detected events.

2. 10. The method of claim 1, 1. A method according to claim 1, wherein at least two of the plurality of primary digital video streams (210) are provided as part of a shared digital video communication service (110) involving respective remotely connected participant clients (121) providing the primary digital video streams (210).

3. 3. The method of claim 2, The method, wherein the collecting step includes collecting at least one of the plurality of primary digital video streams (210) from the shared digital video communication service (110).

4. 4. The method of claim 2 or 3, The method, wherein the collecting step includes collecting at least one of the plurality of primary digital video streams (210) as an external digital video stream (301) collected from a source (300) external to the shared digital video communication service (110).

5. 10. The method of claim 1, The method of claim 1, wherein at least one of the plurality of primary digital video streams (210) has a displaced video encoding, frame rate, aspect ratio, and / or resolution.

6. 10. The method of claim 1, The method, wherein the collecting step includes converting at least two of the plurality of primary digital video streams (210) into a common protocol (240), the common protocol (240) providing for storing digital video data in a raw binary format without performing digital video decoding or digital video encoding, and the common protocol (240) also providing for storing metadata (242) associated with specified points in time related to the stored digital video data.

7. 10. The method of claim 1, the common time reference (260) comprises a common audio signal (111); The method, wherein the common audio signal (111) is common to a shared digital video communication service (110) involving at least two remotely connected participant clients (121), each providing a respective one of the plurality of primary digital video streams (210).

8. 10. The method of claim 1, at least two of the plurality of primary digital video streams (210) are provided to the shared digital video communication service (110) by respective participant clients (121), each such participant client (121) having respective local synchronization software (125) configured to detect the time of arrival of a time synchronization element (231) provided as part of the output digital video stream (230) provided to that participant client (121); and The method of claim 1, wherein the common time reference (260) is determined based at least in part on the detected arrival times.

9. 10. The method of claim 1, The method of claim 1, wherein the common time reference (260) is determined based at least in part on a detected mismatch between an audio portion (214) of a first one of the plurality of primary digital video streams (210) and an image portion (215) of the first primary digital video stream (210), the mismatch being based on digital lip-sync video analysis of a speaking participant (122) viewed in the first primary digital video stream (210).

10. 10. The method of claim 1, The method includes the step of synchronizing intentionally introducing a latency of up to 30 seconds, for example up to 5 seconds, for example up to 1 second, for example up to 0.5 seconds, thereby providing at least said latency in the output digital video stream (230).

11. 11. The method of claim 10, The method, wherein the pattern detection step includes considering information from the plurality of primary digital video streams (210), the information being present in frames (213) that are later than frames of the time-synchronized primary digital video streams (210) that have not yet been used to generate the output digital video stream (230).

12. 10. The method of claim 1, The method, wherein the generating step further includes generating the output digital video stream (230) based on a set of predetermined and / or dynamically variable parameters relating to: the visibility of each of the plurality of primary digital video streams (210) in the output digital video stream (230); the placement of visual and / or audio video content; the visual or audio effects used; and / or the output mode of the output digital video stream (230).

13. 10. The method of claim 1, At least one of the plurality of primary digital video streams (210) is provided to a digital video communication service (110); and The method, wherein the publishing step includes providing the output digital video stream (230) to the communication service (110), such as a participant client (121) of the communication service (110) or an external consumer (150).

14. 1. A computer program for providing a shared digital video stream, the computer program being configured to perform the following steps when executed by a computer: In a collecting step, a plurality of primary digital video streams (210) are collected from at least two digital video sources (120), respectively; In an event detection step, the plurality of primary digital video streams (210) are individually analyzed to detect at least one event (211) selected from a first set of events; In a synchronizing step, the plurality of primary digital video streams (210) are time-synchronized with respect to a common time reference (260); In a pattern detection step, the time-synchronized plurality of primary digital video streams (210) are analyzed to detect at least one pattern (212) selected from a first pattern set, wherein the detection of the pattern is based on the detected at least one event (211); In a generating step, the shared digital video stream is generated as an output digital video stream (230) based on consecutively considered frames (213) of the time-synchronized plurality of primary digital video streams (210) and the detected pattern (212); and In a publishing step, the output digital video stream (230) is continuously provided to consumers of the shared digital video stream, wherein: The detection of the pattern is based, first, on video and / or audio information contained as part of at least two of the time-synchronized primary video streams (210) considered together, or, second, on video and / or audio information contained in a single primary video stream spanning at least two detected events.

15. A system (100) for providing a shared digital video stream, said system (100) comprising a central server (130), said central server having the following functions: a collection function (131) configured to collect a plurality of primary digital video streams (210) from at least two digital video sources (120), respectively; an event detection function (132) configured to individually analyze the plurality of primary digital video streams (210) to detect at least one event (211) selected from a first set of events; a synchronization function (133) configured to time-synchronize the plurality of primary digital video streams (210) to a common time reference (260); a pattern detection function (134) configured to analyze the plurality of time-synchronized primary digital video streams (210) to detect at least one pattern (212) selected from a first pattern set, wherein the detection of the pattern is based on the at least one detected event (211); a generating function (135) configured to generate the shared digital video stream as an output digital video stream (230) based on consecutively considered frames (213) of the plurality of time-synchronized primary digital video streams (210) and the detected pattern (212); and a publishing function (136) configured to continuously provide the output digital video stream (230) to consumers of the shared digital video stream, wherein: The detection of the pattern is based, first, on video and / or audio information contained as part of at least two of the time-synchronized primary video streams (210) considered together, or, second, on video and / or audio information contained in a single primary video stream spanning at least two detected events.