System and method for producing video stream
Patent Information
- Application Number
- JP2025080868
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-11-30
- Filing Date
- 2025-05-14
- Publication Date
- 2025-11-10
AI Technical Summary
Digital video conferencing systems face challenges in managing dynamic conference screen layouts due to varying latency, frame rates, aspect ratios, resolutions, and connectivity issues among multiple input streams, leading to synchronization problems and poor user experiences, especially in complex meetings with multiple participants and diverse hardware.
A method and system that generates two video streams with different delays, allowing one stream to be continuously provided to a first participant and temporarily shared with a second participant, while the second stream is continuously provided to the second participant after an interaction step, using a central server to synchronize and process multiple input streams.
This approach enhances user experience by addressing synchronization and latency issues, enabling effective management of multiple input streams with varying characteristics, resulting in a well-formed and synchronized output video stream for all participants.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a system, computer software product and method for generating a digital video stream, in particular for generating a digital video stream based on two or more different digital input video streams. In a preferred embodiment, the digital video stream is generated in the context of a digital video conference or a digital video conference or meeting system, in particular in which multiple different concurrent users participate. The generated digital video stream may be published externally or within the digital video conference or digital video conference system.
[0002] In other embodiments, the invention is applied to contexts that are not digital video conferencing, but where multiple digital video input streams are simultaneously processed and combined into a digital video stream to be generated. For example, such a context may be educational or instructional. [Background technology]
[0003] Many digital video conferencing systems are known, such as Microsoft® Teams®, Zoom®, and Google® Meet®, that allow two or more participants to meet virtually using locally recorded digital video and audio that is broadcast to all participants, emulating a physical meeting.
[0004] There is a general need to improve these digital video conferencing solutions, particularly with regard to the generation (production) of viewing content: what content to show, at what time, to whom, and through what distribution channel.
[0005] For example, some systems automatically detect the currently speaking participant and display the corresponding video feed of that participant to other participants. Many systems also allow for the sharing of graphics such as the currently displayed screen, viewing window, or digital presentation. However, as virtual meetings become more complex, it will quickly become difficult for services to know which of all the information currently available should be displayed to each participant at any given time.
[0006] In another example, a presenting participant moves around the stage while talking about the slides in a digital presentation, in which case the system must decide whether to show the presentation, the presenter, or both, or switch between the two.
[0007] It may be desirable to generate one or more output digital video streams based on multiple input digital video streams through an automated generation process and provide such generated digital video stream or streams to one or more consumers. DISCLOSURE OF THE INVENTION [Problem to be solved by the invention]
[0008] However, due to the many technical challenges faced by such digital video conferencing systems, it is often difficult for dynamic conference screen layout managers and other automated generators to select what information to display.
[0009] First, low latency is important because digital video conferencing is real-time sensitive. This becomes problematic when different incoming digital video streams are associated with different latency, different frame rates, different aspect ratios, or different resolutions, such as when different participants join using different hardware. Often, these incoming digital video streams require processing for a well-formed user experience.
[0010] Second, there is the issue of time synchronization. Because various input digital video streams, such as external digital video streams and digital video streams provided by participants, are typically fed into a central server or the like, there is no absolute time to synchronize each of these digital video feeds. Similar to excessive delay, unsynchronized digital video feeds lead to a poor user experience.
[0011] Third, digital video conferences between multiple participants may involve different digital video streams with different encodings or formats, which require decoding and re-encoding, creating problems in terms of latency and synchronization, and such encoding is computationally intensive and expensive in terms of hardware requirements.
[0012] Fourth, the fact that different digital video sources may be associated with different frame rates, different aspect ratios, and different resolutions can result in unpredictable and variable memory allocation needs that require continuous balancing, potentially resulting in additional latency and synchronization issues, resulting in the need for large buffers.
[0013] Fifth, participants may experience varying difficulties in terms of connectivity fluctuations, dropouts / reconnections, etc., which poses additional challenges to automatically generating a well-formed user experience.
[0014] These issues are amplified in more complex meeting situations, such as those with large numbers of participants; participants connecting using different hardware and / or software; externally provided digital video streams; screen sharing; and multiple hosts.
[0015] Corresponding problems arise in other contexts when an output digital video stream is to be generated based on multiple input digital video streams, such as in digital video generation systems for education and instruction.
[0016] Swedish patent application SE2151267-8 (unpublished at the effective priority date of the present application) discloses various solutions to the above mentioned problems.
[0017] Digital video environments involving multiple participants present additional latency challenges. In particular, latency requirements may vary among different participants. In such environments, providing a good time-synchronized experience for all participants so that time delays do not adversely affect communication can prove challenging. This is especially true in complex video environments, such as those involving intermediately generated video streams involving multiple participants and / or those involving multiple types of participants.
[0018] The present invention addresses one or more of the problems set forth above. [Means for solving the problem]
[0019] Therefore, the present invention relates to a method for providing a second generated video stream, said method comprising the following steps: in a first generating step (135'), generating a first generated digital video stream based on a first primary digital video stream and a second primary digital video stream, wherein said first generated digital video stream is generated consecutively for publication with a first delay; in a second generating step (135''), generating a second generated video stream based on said first and second primary digital video streams, wherein said second generated digital video stream is , continuously generated to be published with a second delay, the second delay being greater than the first delay; and in a publishing step (136'), the first generated video stream is continuously provided to a first participant client (121) and the second generated video stream is continuously provided to a second participant client (121); in an interacting step, the first generated video stream is temporarily provided to the first and second participant clients (121); and after the interacting step, the second digitally generated video stream is again provided to the second participant client (121).
[0020] The present invention also relates to a computer software product for providing a second generated video stream, the computer program function of which, when executed, performs the following steps: in a first generating step (135'), generating a first generated digital video stream based on a first primary digital video stream and a second primary digital video stream, wherein said first generated digital video stream is generated consecutively for release with a first delay; in a second generating step (135''), generating said second generated video stream based on said first and second primary digital video streams, wherein said second generated digital video stream is generated consecutively for release with a first delay; The video streams are continuously generated for publication with a second delay, the second delay being greater than the first delay; and in a publishing step (136'), the first generated video stream is continuously provided to a first participant client (121) and the second generated video stream is continuously provided to a second participant client (121); in an interacting step, the first generated video stream is temporarily provided to the first and second participant clients (121); and in a providing step, after the interacting step, the second digitally generated video stream is again provided to the second participant client (121).
[0021] The present invention also relates to a system (100) for providing a second generated video stream, the system (100) comprising a central server (130) having the following functions: a first generating function (135') configured to generate a first generated digital video stream based on a first primary digital video stream and a second primary digital video stream, wherein the first generated digital video stream is generated consecutively for publication with a first delay; a second generating function (135'') configured to generate the second generated video stream based on the first and second primary digital video streams, wherein the first generated digital video stream is generated consecutively for publication with a first delay; The second generated digital video stream is continuously generated to be published with a second delay, the second delay being greater than the first delay; and a publishing function (136') configured to continuously provide the first generated video stream to a first participant client (121) and continuously provide the second generated video stream to a second participant client (121); an interactive function to temporarily provide the first generated video stream to the first and second participant clients (121); and a providing function to again provide the second digitally generated video stream to the second participant client (121) after the interactive step.
[0022]
[0023]
[0024]
[0025]
[0026]
[0027]
[0028] Furthermore, the present invention relates to a system.
[0029] The present invention will now be described in detail with reference to exemplary embodiments thereof and the accompanying drawings. [Brief explanation of the drawings]
[0030] [Figure 1] FIG. 1 illustrates a first exemplary system. [Figure 2] FIG. 2 illustrates a second exemplary system. [Figure 3] FIG. 3 illustrates a third exemplary system. [Figure 4] FIG. 4 is a diagram illustrating the central server. [Figure 5] FIG. 5 is a diagram illustrating the first method. [Figure 6a] FIG. 6a shows subsequent states associated with different method steps in the method shown in FIG. [Figure 6b] FIG. 6b shows subsequent states associated with different method steps in the method shown in FIG. [Figure 6c] FIG. 6c shows subsequent states associated with different method steps in the method shown in FIG. [Figure 6d] FIG. 6d shows subsequent states associated with different method steps in the method shown in FIG. [Figure 6e] FIG. 6e shows subsequent states associated with different method steps in the method shown in FIG. [Figure 6f] FIG. 6f illustrates subsequent states associated with different method steps in the method shown in FIG. [Figure 7] FIG. 7 is a diagram conceptually illustrating a common protocol. [Figure 8a] FIG. 8a illustrates the second method. [Figure 8b] FIG. 8b illustrates a third method. [Figure 8c] FIG. 8c illustrates a fourth method. [Figure 8d] FIG. 8d illustrates a fifth method. [Figure 9] FIG. 9 illustrates a fourth exemplary system. DETAILED DESCRIPTION OF THE INVENTION
[0031] All figures share the same or corresponding part reference numbers.
[0032] FIG. 1 shows a system 100 according to the invention, configured to perform a method according to the invention for providing a digital video stream, for example a shared digital video stream.
[0033] System 100 may include video communication service 110, which in some embodiments may be external to system 100. As described below, multiple video communication services 110 may be included.
[0034] The system 100 may include one or more participant clients 121, although one, some, or all of the participant clients 121 may be external to the system 100 in some embodiments.
[0035] The system 100 includes a central server 130 .
[0036] As used herein, the term "central server" refers to a computer-implemented function configured to be accessible in a logically centralized manner, such as through a well-defined API (application programming interface). Such central server functionality may be implemented purely in computer software, or in a combination of software and virtual and / or physical hardware. It may also be implemented in a standalone physical or virtual server computer, or distributed across multiple interconnected physical and / or virtual server computers.
[0037] The physical or virtual hardware on which the central server 130 runs, in other words, the computer software that defines the functionality of the central server 130, may be comprised of a per se conventional CPU, a per se conventional GPU, per se conventional RAM / ROM memory, per se conventional computer buses, and per se conventional external communication capabilities such as an Internet connection.
[0038] The video communication service 110 is also a central server in the above sense, insofar as that is used, which may be a different central server from the central server 130 or may be part of the central server 130 .
[0039] Correspondingly, each of the participant clients 121 may be a central server in the above sense, in a corresponding sense, in which the physical or virtual hardware on which each participant client 121 runs, in other words the computer software that defines the functionality of the participant client 121, comprises a per se conventional CPU / GPU, a per se conventional RAM / ROM memory, a per se conventional computer bus, and per se conventional external communication capabilities such as an Internet connection.
[0040] Each participant client 121 also typically includes or is in communication with a computer screen arranged to display video content that is provided to the participant client 121 as part of an ongoing video communication, speakers arranged to emit sound content that is provided to the participant client 121 as part of the video communication, a video camera, and a microphone arranged to record sound locally to a human participant 122 to the video communication, who uses that participant client 121 to participate in the video communication.
[0041] In other words, each participant client 121's respective human-machine interface allows each participant 122 to interact with other participants in a video communication and / or with audio / video streams provided from various sources at that client 121.
[0042] Generally, each participant client 121 comprises a respective input means 123, which may consist of the video camera, the microphone, a keyboard, a computer mouse or trackpad, and / or an API for receiving digital video streams, digital audio streams, and / or other digital data. The input means 123 are configured, inter alia, to receive video and / or audio streams from the video communication service 110 and / or a central server, such as central server 130, where such video and / or audio streams are provided as part of the video communication and are preferably generated based on corresponding digital data input streams provided to the central server from at least two sources of such digital data input streams, e.g., the participant client 121 and / or an external source (described below).
[0043] More generally, each participant client 121 includes a respective output means 124, which may consist of the computer screen, the speakers, and an API that emits digital video and / or audio streams that are representative of the locally captured video and / or audio for a participant 122 using that participant client 121.
[0044] In practice, each participant client 121 may be a mobile device, such as a mobile phone, equipped with a screen, speakers, microphone, and Internet connection, running computer software locally or accessing remotely executed computer software to perform the functions of that participant client 121. Correspondingly, a participant client 121 may be a thick or thin laptop or a desktop computer, running locally installed applications or using functions remotely accessed via a web browser.
[0045] There may be one or more participant clients 121, for example at least three or at least four, used in one and the same video communication in this embodiment.
[0046] There may be at least two different groups of participating clients. Each of the participating clients may be assigned to a respective such group. The groups may reflect different roles of the participating clients, different virtual or physical locations of the participating clients, and / or different interaction rights of the participating clients.
[0047] A variety of such roles are available, such as "leader" or "conference organizer", "speaker", "panelist", "interactive audience", and "remote listener".
[0048] Such physical locations may vary widely, for example, "on stage," "in a panel," "physically present audience," or "physically distant audience."
[0049] Virtual locations may be defined in terms of physical locations, but may also include virtual groupings that may overlap with physical locations, for example, physically present audience participants may be divided into a first virtual group and a second virtual group, with some physically present audience participants grouped together with some physically remote audience participants in the same virtual group.
[0050] Such interaction rights can be varied and may be, for example, "full interaction" (no restrictions), "can speak, but only after requesting the microphone" (e.g., raising a virtual hand in a video conferencing service), "cannot speak, but can write in a common chat", or "only watch / listen".
[0051] In some implementations, each defined role and / or physical / virtual location may be defined with respect to certain predetermined interaction permissions. In other examples, all participants with the same interaction permissions form a group. Thus, the defined roles, locations, and / or interaction permissions can reflect various group assignments, and different groups may differ from or overlap with each other as desired.
[0052] This is exemplarily explained below.
[0053] The video communications may be provided at least in part by the video communications service 110 and at least in part by the central server 130 as described and illustrated herein.
[0054] As the term is used herein, a "video communication" is a two-way digital communication session involving at least two, preferably at least three, or at least four, video streams, preferably used to generate one or more mixed or collaborative digital video / audio streams, and a matching audio stream, for consumption by one or more consumers (e.g., participant clients of the types described above) who may or may not contribute to the video communication via video and / or audio. Such video communication may be real-time, with or without a certain latency or delay. At least one, preferably at least two, or at least four participants 122 participating in such a video communication engage in the video communication in a two-way manner, providing and consuming video / audio information.
[0055] At least one of the participant clients 121, or all of the participant clients 121, includes local synchronization software functionality 125, which is described in more detail below.
[0056] The video communication services 110 may have or access a common time reference, as described in more detail below.
[0057] Each of the at least one central server 130 may include an API 137 for digitally communicating with entities external to that central server 130. Such communication may include both input and output.
[0058] The system 100, such as the central server 130, may be configured to digitally communicate with external information sources 300, such as externally provided video streams, and in particular to receive digital information, such as audio and / or video stream data, from the external information sources 300. The information sources 300 are "external" in the sense that they are not provided by or as part of the central server 130. Preferably, the digital data provided by the external information sources 300 is independent of the central server 130, and the central server 130 cannot influence its information content. For example, the external information sources 300 may be live-captured video and / or audio, such as a public sporting event or ongoing news event or coverage. The external information sources 300 may also be captured by a webcam or the like, rather than by one of the participant clients 121. Thus, such captured video may depict the same locale as any one of the participant clients 121, but is not captured as part of the participant client 121's activities. One possible difference between an externally provided information source 300 and an internally provided information source 120 is that an internally provided information source may be provided in that capacity as a participant in a video communication of the type defined above, whereas an externally provided information source 300 is not, but instead is provided as part of a context that is external to the video conference.
[0059] There may also be multiple external information sources 300 providing digital information of that type, such as audio and / or video streams, to the central server 130 in parallel.
[0060] As shown in FIG. 1, each participant client 121 constitutes the source of a respective information (video and / or audio) stream 120 provided by that participant client 121 to the video communication service 110 as described.
[0061] The system 100, such as the central server 130, may be further configured to digitally communicate with, and in particular emit digital information to, the external consumers 150. For example, digital video and / or audio streams generated by the central server 130 may be provided continuously, in real time or near real time, to one or more external consumers 150 via the API 137. Again, the consumer 150 being "external" means that the consumer 150 is not provided as part of the central server 130 and / or is not a party to the video communication.
[0062] Unless otherwise noted, all functions and communications herein are provided digitally and electronically, implemented by computer software running on appropriate computer hardware and communicated over digital communications networks or channels such as the Internet.
[0063] 1 , multiple participant clients 121 participate in a digital video communication provided by video communication service 110. Each participant client 121 therefore has an ongoing login, session, or the like with video communication service 110 and can participate in one and the same ongoing video communication provided by video communication service 110. In other words, the video communication is "shared" among the participant clients 121 and, therefore, by the corresponding human participants 122.
[0064] 1 , the central server 130 comprises an auto-join client 140, which is an automatic client corresponding to the participant client 121, but which is not associated with a human participant 122. Instead, the auto-join client 140 is added as a participant client to the video communication service 110 to participate in the same shared video communication as the participant client 121. As such a participant client, the auto-join client 140 is given access to continuously generated digital video and / or audio stream(s) provided by the video communication service 110 as part of the ongoing video communication, and such streams can be consumed by the central server 130 via the auto-join client 140. Preferably, the automatic participant client 140 receives from the video communication service 110 a common video and / or audio stream that is or can be distributed to each participant client 121; respective video and / or audio streams that are provided from each of one or more participant clients 121 to the video communication service 110 and relayed by the video communication service 110 to all participant clients 121 or to requesting participant clients 121 in raw or modified form; and / or a common time reference.
[0065] The central server 130 may include a collection function 131 configured to receive multiple video and / or audio streams of the above types from the auto-attendee clients 140, and possibly also from the above-mentioned external information source(s) 300, for processing as described below, and then provide a generated video stream, e.g., a shared video stream, via the API 137. For example, this generated video stream may be consumed by external consumers 150 and / or by the video communication service 110, which may distribute it to all or any requesting one of the participant clients 121.
[0066] FIG. 2 is similar to FIG. 1, but instead of using an auto-join client 140, the central server 130 receives video and / or audio stream data from an ongoing video communication via the API 112 of the video communication service 110.
[0067] 3 is similar to FIG. 1, but does not show the video communication service 110. In this case, the participant clients 121 communicate directly with the API 137 of the central server 130, for example, to provide video and / or audio stream data to and / or receive video and / or audio stream data from the central server 130. The generated shared streams may then be provided to external consumers 150 and / or to one or more of the client participants 121.
[0068] 4 shows the central server 130 in more detail. As shown, the collection function 131 may be composed of one or, preferably, multiple, format-specific collection functions 131 a. Each of the format-specific collection functions 131 a is configured to receive video and / or audio streams having a predetermined format, such as a predetermined binary encoding format and / or a predetermined stream data container, and specifically to parse and classify the binary video and / or audio data in said format into individual video frames, sequences of video frames, and / or time slots.
[0069] The central server 130 further includes an event detection function 132 configured to receive video and / or audio stream data, such as binary stream data, from the collection function 131 and perform respective event detection on each of the received data streams. The event detection function 132 may include an AI (artificial intelligence) component 132a for performing event detection. The event detection may be performed without first time-synchronizing the collected individual streams.
[0070] The central server 130 further comprises a synchronization function 133 configured to time-synchronize multiple data streams that may be provided by the collection function 131 and processed by the event detection function 132. The synchronization function 133 may comprise an AI component 133a for performing the time synchronization.
[0071] The central server 130 may further include a pattern detection function 134 configured to perform pattern detection based on a combination of at least one, but often at least two, e.g., at least three, or at least four, e.g., all, of the received data streams. Pattern detection may further be based on one, and possibly at least two or more, events detected by the event detection function 132 for each of the data streams. Such detected events considered by the pattern detection function 134 may be distributed over time for each collected stream. The pattern detection function 134 may include an AI component 134a for performing pattern detection. Pattern detection may further be based on the groupings described above, and may be configured to detect a specific pattern occurring only for one group, for some groups but not all groups, or for all groups.
[0072] The central server 130 further comprises a generating function 135 configured to generate a generated digital video stream, e.g., a shared digital video stream, based on the multiple data streams provided from the collecting function 131, and possibly based on any detected events and / or patterns. The generated video stream includes at least a generated video stream comprising one or more of the raw, reformatted, or converted video streams provided by the collecting function 131, and may also include corresponding audio stream data. As exemplified below, there may be multiple generated video streams, and one such generated video stream may be generated in the manner described above, but may also be generated based on another already-generated video stream.
[0073] All generated video streams are preferably generated continuously, and preferably in near real time (after subtracting latencies and delays of the type described later herein).
[0074] The central server 130 may further include a publishing function 136 configured to publish the generated shared digital video stream, such as via the API 137 described above.
[0075] It should be noted that while Figures 1, 2 and 3 show three different examples of how the central server 130 can be used to implement the principles described herein, and in particular to provide methods in accordance with the present invention, other configurations are possible, with or without one or more video communication services 110.
[0076] Thus, Figure 5 illustrates a method for providing the generated digital video stream. Figures 6a to 6f show the states of the different digital video / audio data streams resulting from the method steps shown in Figure 5.
[0077] In the first step, the method begins.
[0078] In a subsequent collection step, a respective plurality of primary digital video streams 210, 301 are collected from at least two of the digital video sources 120, 300, e.g., by a collection function 131. Each of these plurality of primary data streams 210, 301 may comprise an audio portion 214 and / or a video portion 215. "Video" in this context is understood to refer to the moving image and / or still image content of such data streams. Each primary data stream 210, 301 may be encoded according to any video / audio coding standard (using a respective codec used by the entity providing that primary stream 210, 301), and the encoding format may differ between different ones of the plurality of primary streams 210, 301 used simultaneously in one and the same video communication. At least one, e.g., all, of the plurality of primary data streams 210, 301 are preferably provided as a stream of binary data, possibly itself provided in a conventional data container data structure. Preferably, at least one, for example at least two, or even all of the plurality of primary data streams 210, 301 are provided as respective live video recordings.
[0079] It should be noted that the multiple primary data streams 210, 301 may not be synchronized in time when received by the collection function 131. This may mean that they are associated with different latencies or delays relative to each other. For example, if two primary video streams 210, 301 are live recordings, this may mean that they are associated with different latencies relative to the recording time when received by the collection function 131.
[0080] It should also be noted that the multiple primary data streams 210, 301 may themselves be respective live camera feeds from webcams; a currently shared screen or presentation; a film clip being viewed; or any combination of these arranged in various ways within one and the same screen.
[0081] The collection steps are illustrated in Figures 6a and 6b. Figure 6b also shows how the collection function 131 can store each primary video stream 210, 301 as bundled audio / video information or as audio stream data separated from the associated video stream data. Figure 6b illustrates how the data for the primary video streams 210, 301 is stored as individual frames 213 or collections / clusters of frames, where "frame" here refers to a time-limited portion of image data and / or any associated audio data, for example, each frame being an individual still image or a consecutive series of images (e.g., a series of images that constitute up to one second of moving images) that together form moving image video content.
[0082] In a subsequent event detection step performed by the event detection function 132, the multiple primary digital video streams 210, 301 are analyzed by the event detection function 132, particularly the AI component 132a, etc., to detect at least one event 211 selected from the first set of events. This is shown in Figure 6c.
[0083] This event detection step is preferably performed for at least one, e.g., at least two, e.g., all, of the primary video streams 210, 301, and individually for each of said primary video streams 210, 301. In other words, the event detection step is preferably performed for each of said primary video streams 210, 301, taking into account only the information contained as part of that particular primary video stream 210, 301, and in particular without taking into account information contained as part of other primary video streams. Furthermore, event detection is preferably performed without taking into account any common time reference 260 associated with multiple primary video streams 210, 301.
[0084] Preferably, however, event detection takes into account information contained as part of the individually analyzed primary video stream over a time interval, for example over a historical time interval of the primary video stream that is greater than 0 seconds, for example at least 0.1 seconds, for example at least 1 second.
[0085] Event detection may take into account information contained in the audio and / or video data included as part of the primary video stream 210, 301.
[0086] The first set of events may include any number of types of events, such as a change of slide in a slide presentation that constitutes or is part of the primary video stream 210, 301; a change in connection quality of the source 120, 300 providing the primary video stream 210, 301 that results in a change in image quality, loss of image data, or reacquisition of image data; and physical events of movement detected in the primary video stream 210, 301, such as movement of a person or object in the video, a change in lighting in the video, a sudden sharp noise in the audio, or a change in audio quality. It should be understood that this is not intended to be an exhaustive list, and that these examples are provided to help understand the applicability of the presently described principles.
[0087] In a subsequent synchronization step performed by the synchronization function 133, the multiple primary digital video streams 210 are time-synchronized. This time synchronization may be performed with respect to a common time reference 260. As shown in FIG. 6d, this time synchronization may include aligning the multiple primary video streams 210, 301 with respect to one another using the common time reference 260, for example, so that they can be combined to form a time-synchronized context. The common time reference 260 may be a stream of data, a heartbeat signal or other pulse data, or a time anchor applicable to each of the multiple individual primary video streams 210, 301. By making the common time reference applicable to each of the multiple individual primary video streams 210, 301, the information content of the primary video streams 210, 301 can be uniquely associated with the common time reference with respect to a common time axis. In other words, the common time reference aligns the multiple primary video streams 210, 301 so that they are time-synchronized in a present sense via time shifting. In other embodiments, time synchronization may be based on known information about the time difference between the primary video streams 210, 301, such as measurements.
[0088] As shown in FIG. 6d, time synchronization may include determining one or more timestamps 261 for each of multiple primary video streams 210, 301, for example, relative to a common time reference 260, or for each video stream 210, 301 relative to the other video stream 210, 301 or relative to other multiple video streams 210, 301.
[0089] In a subsequent pattern detection step performed by pattern detection function 134, the time-synchronized multiple primary digital video streams 210, 301 are analyzed to detect at least one pattern 212 selected from the first pattern set. This is shown in Figure 6e.
[0090] In contrast to the event detection step, the pattern detection step is preferably performed based on video and / or audio information included as part of at least two of the multiple time-synchronized primary video streams 210, 301.
[0091] The first set of patterns may include any number of types of patterns, such as multiple participants speaking alternately or simultaneously, or a change in presentation slides occurring simultaneously as another event, such as another participant speaking, etc. This list is not exhaustive but is exemplary.
[0092] In alternative embodiments, the detected pattern 212 may relate to information contained in only one of the multiple primary video streams 210, 301, rather than information contained in more than one of the multiple primary video streams 210, 301. In such cases, such pattern 212 is preferably detected based on video and / or audio information contained in that single primary video stream 210, 301 that spans at least two detected events 211, e.g., two or more consecutively detected presentation slide changes or connection quality changes. As an example, multiple consecutive slide changes that follow one another rapidly over time may be detected as one single slide change pattern, as opposed to one distinct slide change pattern for each detected slide change event.
[0093] It is understood that the first set of events and the first set of patterns may comprise predetermined types of events / patterns defined using respective sets of parameters and parameter intervals. As described below, the sets of events / patterns may also be defined and detected using various AI tools.
[0094] In a subsequent generation step performed by the generation function 135, a shared digital video stream is generated as an output digital video stream 230 based on the consecutively considered multiple frames 213 of the multiple time-synchronized primary digital video streams 210, 301 and the detected pattern 212.
[0095] As will be explained and detailed below, the present invention allows for fully automated generation of video streams, such as output digital video stream 230.
[0096] For example, such generation may include selection of what video and / or audio information from which primary video streams 210, 301 to use in the output video stream 230, and to what extent; the video screen layout of the output video stream 230; the switching pattern between different such uses or layouts over time; etc.
[0097] This is also illustrated in Figure 6f, which shows one or more additional portions of time-related (which may be relative to the common time reference 260) digital video information 220, such as additional digital video information streams that may be time-synchronized (e.g., relative to the common time reference 260) and used in conjunction with the time-synchronized multiple primary video streams 210, 301 in generating the output video stream 230. For example, the additional streams 220 may include information regarding any video and / or audio special effects to use, such as dynamically based on detected patterns; a planned time schedule for the video communication; etc.
[0098] In a subsequent publishing step performed by the publishing function 136, the generated output digital video stream 230 is continuously provided to the consumers 110, 150 of the shared digital video stream, as described above. The generated digital video stream may be provided to one or more participant clients 121, for example, via the video communication service 110.
[0099] At the subsequent step, the method ends. However, initially, the method may be iterated any number of times to generate output video stream 230 as a continuously provided stream, as shown in FIG. 5. Preferably, output video stream 230 is generated to be consumed in real time or near real time (taking into account the total latency added by all steps along the way) and continuously (published as more information becomes available, but not counting the intentionally added latency described below). In this manner, output video stream 230 may be consumed in an interactive manner, whereby output video stream 230 is fed back to video communication service 110 or to another context that forms the basis for generating primary video stream 210, which is fed back to collection function 131 to form a closed feedback loop; or output video stream 230 is consumed in a different context (external to system 100, or at least external to central server 130), where it may form the basis for real-time two-way video communication.
[0100] As mentioned above, in some embodiments, at least two, e.g., at least three, e.g., at least four or at least five of the multiple primary digital video streams 210, 301 are provided as part of a shared digital video communication such as provided by a video communication service 110, which video communication includes respective remotely connected participant clients 121 providing such primary digital video streams 210. In such cases, the collecting step may consist of collecting at least one of such primary digital video streams 210 from the shared digital video communication service 110 itself, via an auto-join client 140 that has in turn been granted access to the video and / or audio stream data from within such video communication service 110, and / or via the API 112 of the video communication service 110.
[0101] Additionally, in this and other cases, the collecting step may comprise collecting at least one of the plurality of primary digital video streams 210, 301 as a respective external digital video stream 301 collected from an information source 300 that is external to the shared digital video communication service 110. Note that one or more of such external video sources 300 may be external to the central server 130.
[0102] In some embodiments, the multiple primary video streams 210, 301 are not formatted in the same manner. While such different formats can be in the form in which they are supplied to the collection function 131 in different types of data containers (such as AVI or MPEG), in preferred embodiments, at least one of the multiple primary video streams 210, 301 is formatted according to a deviating format (relative to at least one other of the primary video streams 210, 301), in that the deviating primary digital video streams 210, 301 have deviating video encodings; deviating fixed or variable frame rates; deviating aspect ratios; deviating video resolutions; and / or deviating audio sample rates.
[0103] The collection function 131 is preferably pre-configured to read and interpret all encoding formats, container standards, etc. occurring in all collected primary video streams 210, 301. This allows processing as described herein to be performed without requiring decoding until relatively late in these processes (e.g., until the primary streams in question are placed in their respective buffers; or after the event detection step; or until after the event detection step). However, in the rare case where one or more of the primary video feeds 210, 301 are encoded using a codec that the collection function 131 cannot interpret without decoding, the collection function 131 may be configured to perform decoding and analysis of such primary video streams 210, 301, followed by conversion to a format that can be processed, for example, by the event detection function. Note that even in this case, re-encoding is preferably not performed at this stage.
[0104] For example, a primary video stream 220 fetched from a multi-party video event, such as that provided by video communication service 110, typically has low latency requirements and is therefore typically associated with variable frame rates and variable pixel resolutions to enable participants 122 to communicate effectively. In other words, the overall video and audio quality is degraded as necessary for low latency.
[0105] On the other hand, the external video feed 301 typically has a more stable frame rate and higher image quality, but may therefore have a higher latency.
[0106] Thus, the video communication service 110 may at each point in time use a different encoding and / or container than the external video source 300. Therefore, the analysis and video generation process described herein must then combine these multiple streams 210, 301 of different formats into a new single stream for a combined experience.
[0107] As mentioned above, the collection functionality 131 may comprise a set of format-specific collection functionality 131a, each configured to process a particular type of format of the primary video stream 210, 301. For example, each one of these format-specific collection functionality 131a may be configured to process multiple primary video streams 210, 301 encoded using different respective video encoding methods / codecs, such as Windows® Media® or DivX®.
[0108] However, in a preferred embodiment, the collecting step includes converting at least two, eg, all, of the plurality of primary digital video streams 210 , 301 to a common protocol 240 .
[0109] As used in this context, the term "protocol" refers to an information structuring standard or data structure that specifies how information contained in a digital video / audio stream is stored. However, the common protocol preferably does not prescribe how digital video and / or audio information is stored, e.g., at a binary level (i.e., encoded / compressed data indicating the sounds and images themselves), but instead forms a structure of a predetermined format for storing such data. In other words, the common protocol prescribes storing digital video data in a raw binary format without performing any digital video decoding or encoding in connection with such storage, and without modifying the existing binary format in any way apart from concatenating and / or splitting binary byte sequences, as the case may be. Instead, the raw (encoded / compressed) binary data content of the primary video stream 210, 301 is preserved by repacking the raw binary data in a data structure defined by the protocol. In some embodiments, the common protocol defines a video file container format.
[0110] FIG. 7 shows, by way of example, multiple primary video streams 210, 301 shown in FIG. 6a reconstructed by respective format-specific collection functions 131a and using the common protocol 240 described above.
[0111] Thus, the common protocol 240 provides for storing digital video and / or audio data in data sets 241 that are preferably divided into discrete, contiguous sets of data along a time axis relative to the primary video stream 210, 301. Each such data set may contain one or several frames of video and associated audio data.
[0112] The common protocol 240 may also provide for storing, in association with the stored digital video and / or audio data set 241, metadata 242 associated with a specified point in time.
[0113] The metadata 242 may include information about the binary format of the raw data of the primary digital video stream 210, such as about the digital video encoding method or codec used to generate the raw binary data, the resolution of the video data, the video frame rate, a frame rate variation flag, the video resolution, the video aspect ratio, the audio compression algorithm, or the audio sampling rate. The metadata 242 may also include information about timestamps of the stored data, for example, related to the time base of the primary video stream 210, 301 itself or related to a different video stream as mentioned above.
[0114] The use of format-specific collection functions 131a in combination with the common protocol 240 allows for the rapid collection of information content of the primary video streams 210, 301 without the added latency (delay) of decoding / re-encoding the received video / audio data.
[0115] Thus, the collecting step may comprise collecting a plurality of primary digital video streams 210, 301 encoded using different binary video and / or audio encoding formats using different ones of a plurality of format-specific collection functions 131a to parse the primary video streams 210, 301 and store the parsed raw binary data, along with any associated metadata, in a data structure using a common protocol. It will be appreciated that the decision as to which format-specific collection function 131a to use for which primary video stream 210, 301 may be made by the collection function 131a based on predetermined and / or dynamically detected characteristics of each of the primary video streams 210, 301.
[0116] Each primary video stream 210, 301 collected in this manner may be stored in its own separate memory buffer, such as a RAM memory buffer within the central server 130.
[0117] The conversion of the primary video streams 210, 301 performed by each format-specific collection function 131a may therefore comprise splitting the raw binary data of each thus converted primary digital video stream 210, 301 into an ordered set of smaller data sets 241.
[0118] Furthermore, the conversion may also comprise associating each of the smaller sets 241 (or subsets, e.g., subsets regularly distributed along the time axis of each of the primary streams 210, 301 in question) with a respective time along a shared time axis, e.g., with respect to a common time reference 260. This association may be performed by analysis of the raw binary video and / or audio data, in one of the principle ways described below, or in other ways, and may be performed to enable subsequent time synchronization of the primary video streams 210, 301 to be performed. Depending on the type of common time reference 260 used, at least part of this association of each data set 241 may also be performed by or instead of the synchronization function 133. In the latter case, the collecting step may instead comprise associating each of the smaller sets 241, or subsets thereof, with a respective time on a time axis specific to the primary stream 210, 301 in question.
[0119] In some embodiments, the collecting step also includes converting the raw binary video and / or audio data collected from the multiple primary video streams 210, 301 to uniform quality and / or frequency updating. This may include downsampling or upsampling the raw binary digital video and / or audio data of the multiple primary digital video streams 210, 301 to a common video frame rate; a common video resolution; or a common audio sampling rate, as appropriate. Note that such resampling can be performed without performing full decoding / re-encoding, or even without performing any decoding at all, since the format-specific collecting function 131a can directly process the raw binary data according to the correct binary encoding target format.
[0120] Preferably, each of the multiple primary digital video streams 210, 301 is stored in an individual data storage buffer 250 as an individual frame 213 or sequence of frames 213, as described above, and each is associated with a corresponding timestamp that is in turn associated with a common time base 260.
[0121] In the specific example provided for illustrative purposes, the video communication service 110 is Microsoft® Teams® and is conducting a video conference involving multiple simultaneous participants 122. The auto-join client 140 is registered as a conference participant in the Teams® conference.
[0122] Next, the primary video input signals 210 are provided to the collection function 130 via the auto-join client 140 and are acquired by the collection function 130. These are raw data signals in H264 format, and include timestamp information for each video frame.
[0123] The associated format-specific collection function 131a picks up the raw data over IP (cloud LAN network) on a configurable predefined TCP port. Every Teams® meeting participant and associated audio data is associated with a separate port. The collection function 131 then uses timestamps from the audio signal (50 Hz) and downsamples the video data to a fixed 25 Hz output signal before storing the video streams 220 in their respective individual buffers 250.
[0124] As mentioned above, common protocol 240 stores data in a raw binary format. It can be designed to process raw bits and bytes of video / audio data at a very low level. In a preferred embodiment, data is stored in common protocol 240 as a simple byte array or corresponding data structure (such as a slice). This means that the data does not need to be placed in a traditional video container at all (common protocol 240 does not constitute such a traditional container in this context). Also, video encoding and decoding is computationally intensive, causing delays and requiring expensive hardware. Furthermore, this problem scales with the number of participants.
[0125] The common protocol 240 allows memory to be reserved within the collection function 131 for the primary video stream 210 associated with each Teams® conference participant 122 and any external video sources 300, and the amount of allocated memory can be changed on the fly during the process. In this way, the number of input streams can be changed, thereby keeping each buffer valid. For example, information such as resolution and frame rate is variable, but is stored as metadata in the common protocol 240, so this information can be used to quickly change the size of each buffer as needed.
[0126] The following is an example of the specification of this type of common protocol 240:
[0127] [Table 1]
[0128] In the above table, the "Detected event in, if any" data is included as part of the specification of common protocol 260. However, in some embodiments, this information (regarding detected events) may instead be placed in a separate memory buffer.
[0129] In some embodiments, the at least one additional portion of the digital video information 220, which may be an overlay or effect, is also stored in each individual buffer 250 as an individual frame or sequence of frames, each associated with a corresponding timestamp that is in turn associated with the common time reference 260.
[0130] As illustrated above, the event detection step may include using a common protocol 240 to store metadata 242 describing the detected event 211 in association with the primary digital video stream 210, 301 in which the event 211 was detected.
[0131] Event detection can be performed in different ways. In some embodiments performed by the AI component 132a, the event detection step involves a first trained neural network or other machine learning component individually analyzing at least one, e.g., some, or all, of the multiple primary digital video streams 210, 301 to automatically detect any of the events 211. This may involve the AI component 132a classifying the data of the primary video streams 210, 301 into a predefined set of events in a supervised classification and / or into a dynamically determined set of events in an unsupervised classification.
[0132] In some embodiments, the detected event 211 is a change in a presentation slide of a presentation that is or is included in the primary video stream 210, 301.
[0133] For example, if a presenter of a presentation decides to change the slide in the presentation that they are currently making to the audience, this means that what is interesting to a given viewer may change. The newly displayed slide may simply be a general-level image that is best viewed for a short time in so-called "butterfly" mode (e.g., displaying the slide side-by-side with the presenter's video in output video stream 230). Or the slide may contain a lot of detail, text in a small font size, etc. In the latter case, the slide will be displayed full screen and for a somewhat longer time than would normally be the case. In this case, the slide may be more interesting to the viewer than the presenter's face, so butterfly mode may not be as appropriate.
[0134] In practice, the event detection step consists of at least one of the following:
[0135] First, the event 211 may be detected based on an image analysis of the difference between a first image of a detected slide and a subsequent second image of the detected slide. The nature of the primary video stream 220, 301 as being indicative of a slide may be determined automatically using conventional digital image processing, such as using motion detection in combination with OCR (Optical Character Recognition).
[0136] This may involve using automated computer image processing techniques to check whether the detected slide has changed significantly enough to be classified as a true slide change. This can be done by checking the delta between the current and previous slides in terms of RGB color values. For example, one can evaluate how much the RGB values have changed globally in the screen area covered by the slide in question, and simultaneously evaluate whether it is possible to find groups of adjacent pixels that change in concert with this. This allows relevant slide changes to be detected while filtering out irrelevant changes, such as computer mouse movements across the screen. This approach allows for complete configurability. For example, it may be desirable to be able to capture computer mouse movements, for example, if a presenter wants to present something in detail while pointing at different things with the computer mouse.
[0137] Second, the event 211 may be detected based on image analysis of the information complexity of the second image itself to determine the type of event with greater specificity.
[0138] This might involve, for example, assessing the amount of textual information on the slide in question and the associated font size, which can be done using traditional OCR methods, including deep learning-based character recognition techniques.
[0139] Note that because the raw binary format of the evaluated video streams 210, 301 is known, this may be performed directly in the binary domain without first decoding or re-encoding the video data. For example, the event detection function 132 may invoke an associated format-specific collection function for an image interpretation service, or the event detection function 132 itself may include functionality for evaluating image information, such as to the individual pixel level, for a number of different supported raw binary video data formats.
[0140] In another example, the detected event 211 is a loss of a communication connection of a participant client 121 to the digital video communication service 110. In this case, the detecting step may include detecting that the participant client 121 has lost the communication connection based on image analysis of a series of subsequent video frames 213 of the primary digital video stream 210 corresponding to that participant client 121.
[0141] Because participant clients 121 are associated with different physical locations and different Internet connections, it may occur that someone loses connection to the video communication service 110 or the central server 130. In such a situation, it is desirable to avoid a black or blank screen appearing in the generated output video stream 230.
[0142] Alternatively, such loss of connection can be detected as an event by the event detection function 132, for example, by applying a two-class classification algorithm where the two classes used are connected / not connected (no data). In this case, "no data" is understood to be different from the presenter intentionally sending a black screen. Because a short black screen, such as just one or two frames, may not be noticeable in the final generated stream 230, the two-class classification algorithm can be applied over time to create a time series. A threshold specifying the minimum length of a connection interruption can then be used to determine whether a connection has been lost.
[0143] As will be explained below, detected events of the types exemplified above may be used by pattern detection function 134 to take various responses, as appropriate and desired.
[0144] As noted above, the individual primary video streams 210, 301 are each associated with a common time reference 260 and can be time-synchronized relative to one another by the synchronization function 133.
[0145] In some embodiments, the common time reference 260 is based on or comprises a common audio signal 111 (see Figures 1 to 3), which, as described above, is common to a shared digital video communication service 110 participating in at least two remotely connected participant clients 121, each providing a respective one of the primary digital video streams 210.
[0146] In the Microsoft® Teams® example mentioned above, a common audio signal may be generated and captured by the central server 130 via the auto-join client 140 and / or via the API 112. In this and other examples, such common audio signal may be used as a heartbeat signal to time-synchronize the individual primary video streams 220 by combining the individual primary video streams 220 at specific points in time based on the heartbeat signal. Such a common audio signal may be provided as a separate signal (relative to each of the other primary video streams 210), such that each of the other primary video streams 210 may be individually time-correlated to the common audio signal based on audio contained in the other primary video streams 210 or based on image information contained therein (e.g., using automatic image processing-based lip-sync techniques).
[0147] In other words, to handle the variable and / or different latencies associated with the individual primary video streams 210 and to achieve time synchronization of the combined video output stream 230, such a common audio signal is used as a heartbeat for all primary video streams 210 within the central server 130 (but possibly not the external primary video stream 301). In other words, all other signals are mapped to this common audio time heartbeat to ensure they are all time synchronized.
[0148] In another example, time synchronization is achieved using a time synchronization element 231 that is introduced into the output digital video stream 230 and detected by a respective local time synchronization software function 125 provided as part of one or more individual ones of the participant clients 121. The local software function 125 is configured to detect the time of arrival of the time synchronization element 231 in the output video stream 230. As will be appreciated, in such an embodiment, the output video stream 230 is fed back to the video communication service 110 or otherwise made available to each participant client 121 and its local software function 125.
[0149] For example, the time synchronization elements 231 may be visual markers, such as pixels that change color in a predetermined order or manner, that are placed or updated in the output video 230 at regular time intervals; a visual clock that is updated and displayed in the output video 230; or an audio signal (which may be designed to be inaudible to the participants 122, for example, by having a sufficiently low amplitude and / or a sufficiently high frequency) that is added to the audio that forms part of the output video stream 230. The local software function 125 is configured to automatically detect the arrival time of each of the time synchronization elements 231 using appropriate image and / or audio processing.
[0150] The common time reference 260 may then be determined, at least in part, based on the detected arrival times. For example, each of the local software functions 125 may communicate respective information indicative of the detected arrival times to the central server 130.
[0151] Such communication may occur via a direct communication link between the participant client 121 and the central server 130. However, communication may also occur via a primary video stream 210 associated with the participant client 121. For example, the participant client 121 may introduce a visual or audible code, such as the type described above, into the primary video stream 210 generated by the participant client 121 for automatic detection by the central server 130 and use to determine the common time reference 260.
[0152] In yet additional embodiments, each participant client 121 may perform image detection on a common video stream viewable by all participant clients 121 for the video communication service 110 and relay the results of such image detection to the central server 130 in a manner corresponding to that described above, where they may be used to determine the respective offsets of each participant client 121 relative to one another over time. In this manner, a common time reference 260 may be determined as a set of individual relative offsets. For example, a selected reference pixel of the commonly available video stream may be monitored by some or all participant clients 121, such as by local software functions 125, and the current color of that pixel may be communicated to the central server 130. The central server 130 may generate an estimated set of relative time offsets across the different participant clients 121 by calculating respective time series based on such color values received consecutively from each of many (or all) of the participant clients 121 and performing cross-correlation.
[0153] In practice, the output video stream 230 provided to the video communication service 110 may be included as part of the shared screen of all participant clients of that video communication and may therefore be used to evaluate such time offsets associated with the participant clients 121. In particular, the output video stream 230 provided to the video communication service 110 may be made available back to the central server via the auto-join client 140 and / or API 112.
[0154] In some embodiments, the common time reference 260 may be determined based at least in part on a detected discrepancy between the audio portion 214 of a first one of the multiple primary digital video streams 210, 301 and the image portion 215 of said first one of the multiple primary digital video streams 210, 301. Such discrepancy may be based, for example, on a digital lip-sync video image analysis of the speaking participant 122 viewed in said first primary digital video stream 210, 301. Such lip-sync analysis may be conventional per se and may, for example, use a trained neural network. The analysis may be performed by the synchronization function 133 for each primary video stream 210, 301 with respect to available common audio information, and the relative offset across the individual primary video streams 210, 301 may be determined based on this information.
[0155] In some embodiments, the synchronization step includes intentionally introducing a delay (in this context, "delay" and "latency" are intended to mean the same thing) of up to 30 seconds, e.g., up to 5 seconds, e.g., up to 1 second, e.g., up to 0.5 seconds, but more than 0 seconds, to provide at least that delay in the output digital video stream 230. Whatever the length, the intentionally introduced delay is at least a plurality of video frames, such as at least three, or at least five, or even ten, such as this number of frames (or individual images) stored after any resampling in the acquisition step. As used herein, the term "intentionally" means that the delay is introduced regardless of the need to introduce such a delay based on synchronization issues or the like. In other words, the intentionally introduced delay is in addition to the delay introduced as part of the synchronization of the multiple primary video streams 210, 301, in order to time-synchronize the multiple primary video streams 210, 301 with each other. The intentionally introduced delay may be predetermined, fixed, or variable relative to the common time reference 260. The delay time may be measured relative to the least latent one of the multiple primary video streams 210, 301, and as a result of the time synchronization, the more latent ones of these streams 210, 301 may be associated with a relatively small intentionally added delay.
[0156] In some embodiments, a relatively small delay, such as 0.5 seconds or less, is introduced, which delay is barely noticeable to participants in the video communication service 110 using the output video stream 230. In other embodiments, a larger delay may be introduced, such as when the output video stream 230 is not used in an interactive context but instead is published in a one-way communication to an external consumer 150.
[0157] This intentionally introduced delay may be sufficient to allow the synchronization function 133 sufficient time to map the collected video frames of the individual primary streams 210, 301 to the correct timestamps 261 of the common time reference 260. It may also be sufficient to provide sufficient time to perform the event detection described above to detect missing primary stream 210, 301 signals, slide changes, resolution changes, etc. Furthermore, the intentionally introduced delay may be sufficient to improve the pattern detection function 134, as described below.
[0158] It will be understood that introducing a delay involves buffering 250 each of the collected and time-synchronized multiple primary video streams 210, 301 before publishing the output video stream 230 using that buffered frame 213. In other words, the video and / or audio data of at least one, some, or all of the multiple primary video streams 210, 301 may be present in the central server 130 in a buffered manner, for the reasons discussed above, particularly for use by the pattern detection function 134, rather than being used like a cache but intended to be able to handle varying bandwidth situations (as in a traditional cache buffer).
[0159] Thus, in some embodiments, the pattern detection step involves considering specific information of at least one, e.g., some, e.g., at least four, or all, of the multiple primary digital video streams 210, 301, where this specific information resides in a frame 213 that follows a frame 213 of the time-synchronized primary digital video stream 210 that has not yet been used in generating the output digital video stream 230. Thus, a newly added frame 213 resides in the buffer 250 for a specific latency period before forming part of (or the basis for) the output video stream 230. During this period, the information of that frame 213 constitutes "future" information with respect to the frame currently being used to generate the current frame of the output video stream 230. When the timeline of the output video stream 230 reaches that frame 213, that frame is used to generate the corresponding frame of the output video stream 230 and may thereafter be discarded.
[0160] In other words, the pattern detection function 134 has at its disposal a set of video / audio frames 213 that have not yet been used to generate the output video stream 230, and uses this data to detect the above patterns.
[0161] Pattern detection can be performed in different ways: In some embodiments performed by the AI component 134a, the pattern detection step includes a second trained neural network or other machine learning component analyzing in concert at least two, e.g., at least three, e.g., at least four, or all of the multiple primary digital video streams 120, 301 to automatically detect the pattern 212.
[0162] In some embodiments, the detected pattern 212 comprises a speech pattern including at least two, e.g., at least three, e.g., at least four different active participants 122, each associated with a respective participant client 121, for the shared video communication service 110, and each of these active participants 122 is visually seen in a respective one of the multiple primary digital video streams 210, 301.
[0163] Preferably, the generating step includes determining, tracking, and updating the current generation state of the output video stream 230. For example, such state can dictate which participants 122 (if any) are visible in the output video stream 230 and where on-screen they are visible; which external video streams 300 are visible in the output video stream 230 and where on-screen they are visible; whether any slides or shared screens are displayed in full-screen mode or in combination with any live video streams; etc. Thus, the generating function 135 can be viewed as a state machine for the generated output video stream 230.
[0164] In order to generate the output video stream 230 as a combined video experience to be viewed, for example, by the end consumer 150, it is advantageous for the central server 130 to be able to understand what is happening at a deeper level than simply detecting individual events associated with the individual primary video streams 210, 301.
[0165] In a first example, the presenting participant client 121 changes the currently displayed slide. This slide change is detected by the event detection function 132 as described above, and metadata 242 is added to the frame indicating that a slide change has occurred. This happens multiple times as the presenting participant client 121 is found to be skipping forward a number of slides in rapid succession, resulting in a series of "slide change" events that are also detected by the detection function 132 and stored along with the corresponding metadata 242 in a separate buffer 250 of the primary video stream 210. In practice, each such rapidly skipped forward slide may only be displayed for a few seconds.
[0166] The pattern detection function 134 looks at the information in the buffer 250 across these detected slide changes and detects a pattern that corresponds to one single slide change rather than multiple or rapidly executed slide changes (i.e., a single slide change to the last slide in a forward skip, with that last slide remaining visible once the fast skip ends). In other words, the pattern detection function 134 notes, for example, that there were 10 slide changes in a very short period of time, and so treats them as a detected pattern representing one single slide change. As a result, the generation function 135 has access to the pattern detected by the pattern detection function 134 and, because it determines that this last slide is potentially important in the state machine, it can choose to display that last slide in full-screen mode for a few seconds in the output video stream 230. It can also choose not to display intermediately viewed slides at all in the output stream 230.
[0167] Detection of patterns with multiple rapid slide changes may be detected by simple rule-based algorithms, but alternatively may be detected using neural networks designed and trained to detect such patterns in video images by classification.
[0168] In another example, it may be desirable to quickly switch visual attention between current speakers while still providing a relevant viewing experience for the consumer 150 by generating and presenting a calm and smooth output video stream 230, as may be useful, for example, when the video communication is a talk show, panel discussion, or the like. In this case, the event detection function 132 may continuously analyze each primary video stream 210, 301 to determine at any given time whether the person being viewed in that particular primary video stream 210, 301 is currently speaking. This may be performed, for example, as described above, using conventional image processing tools. The pattern detection function 134 may then be operable to detect certain overall patterns involving multiple primary video streams 210, 301, which patterns are useful for generating a smooth output video stream 230. For example, the pattern detection function 134 may detect patterns of very frequent switches between current speakers and / or patterns involving multiple simultaneous speakers.
[0169] The generation functionality 135 can then take such detected patterns into account when making automated decisions related to the generation state, such as, for example, not automatically switching visual focus to a speaker who speaks for only half a second and then goes silent again, or switching to a state where multiple speakers are displayed side by side during a period where they are alternating or speaking simultaneously. This state determination process can itself be performed using time-series pattern recognition techniques or using trained neural networks, but can also be based at least in part on a predetermined set of rules.
[0170] In some embodiments, there may be multiple patterns detected in parallel and form input to the state machine of the generation function 135. Such multiple patterns may be used by the generation function 135 in different AI components, computer vision detection algorithms, etc. As an example, a permanent slide change may be detected while simultaneously detecting unstable connections for some participant clients 121, and other patterns may detect the current main speaking participant 122. Using all such available pattern data, a classifier neural network may be trained and / or a set of rules may be developed to analyze the time series of such pattern data. Such classification may be supervised, at least in part, e.g., completely, to yield determined desired state changes used in the generation. For example, different such predetermined classifiers may be generated that are specifically configured to automatically generate the output video stream 230 according to a variety of different generation styles and desires. Training may be based on known generation state change sequences as the desired output and known pattern time series data as training data. In some embodiments, a Bayesian model may be used to generate such classifiers. In a concrete example, a priori information can be obtained from experienced producers, who can provide input such as "In talk shows, we never switch directly from speaker A to speaker B, but always give an overview first before focusing on other speakers unless they are very dominant and loud." This generation logic is expressed as a Bayesian model of the general form "If X is true, then | given the fact that Y is true, | do Z." The actual detection (e.g., whether someone is speaking loudly) can be done using classifiers or threshold-based rules.
[0171] Given large datasets (of pattern time series data), deep learning techniques can be used to develop correct and attractive generative formats for use in the automatic generation of video streams.
[0172] In summary, by using a combination of event detection based on multiple individual primary video streams 210, 301; intentionally introduced delays; pattern detection based on multiple time-synchronized primary video streams 210, 301 and detected events; and a generation process based on the detected patterns, it is possible to achieve automatic generation of the output digital video stream 230 according to a wide range of possible tastes and styles. This result is valid across a wide range of possible neural network and / or rule-based analysis techniques used by the event detection function 132, pattern detection function 134, and generation function 135. This is particularly valid in the embodiment described below, which features a first generated video stream being used to automatically generate a second generated video stream; and the use of different intentionally added delays for different groups of participant clients.
[0173] As illustrated above, the generating step may include generating the output digital video stream 230 based on a set of predetermined and / or dynamically variable parameters relating to the visibility of each of the plurality of primary digital video streams 210, 301 in the output digital video stream 230, the arrangement of visual and / or audio video content, the visual or audio effects used, and / or the output mode of the output digital video stream 230. Such parameters may be determined automatically by a state machine in the generating functionality 135, and / or set by an operator controlling the generation (semi-automated), and / or predetermined based on some a priori configurational desires (such as a minimum time between layout changes of the output video stream 230 or state changes of the types illustrated above).
[0174] In a practical example, the state machine may support a set of predefined standard layouts that can be applied to the output video stream 230, such as a full-screen presenter view (showing the currently speaking participant 122 in full screen); a slide view (showing the currently shared presentation slide in full screen); a "butterfly view" (showing both the currently speaking participant 122 and the currently shared presentation slide in a side-by-side view); or a multi-speaker view (showing all or a selected subset of the participants 122 in a side-by-side or matrix layout). Various available production formats can be defined by a set of available states (such as the set of standard layouts described above) along with a set of state machine state change rules (such as those illustrated above). For example, one such production format could be "panel discussion," another "presentation," etc. By selecting a particular production format via a GUI or other interface to the central server 130, an operator of the system 100 can quickly select one of a set of such predefined production formats and then enable the central server 130 to fully automatically generate the output video stream 230 in accordance with that production format based on available information such as those described above.
[0175] Additionally, during production, a respective in-memory buffer is created and maintained for each conference participant client 121 or external video source 300, as described above. These buffers can be easily deleted, added, and modified on the fly. The central server 130 may then be configured to receive information regarding added / dropped-off participant clients 121 and participants 122 scheduled to speak during production of the output video stream 230; planned or unexpected pauses / resumes of the presentation; desired changes to the currently used production format; and the like. Such information may be provided to the central server 130, for example, via an operator GUI or interface, as described above.
[0176] As illustrated above, in some embodiments, at least one of the multiple primary digital video streams 210, 301 may be provided to a digital video communication service 110, and the publishing step may then include providing the output digital video stream 230 to that same communication service 110. For example, the output video stream 230 may be provided to a participant client 121 of the video communication service 110, or may be provided as an external video stream to the video communication service 110 via the API 112. In this manner, the output video stream 230 may be made available to several or all of the participants in the video communication event currently being facilitated by the video communication service 110.
[0177] As also mentioned above, the output video stream 230 may additionally or alternatively be provided to one or more external consumers 150 .
[0178] Generally, the generating step is performed by a central server 130, and the output digital video stream 230 can be provided as a live video stream via an API 137 to one or more concurrent consumers.
[0179] Figure 8a illustrates a method according to a first aspect of the present invention, which will now be described with reference to the disclosures above: in the method shown in Figure 8a for providing a digital video stream (hereinafter referred to as the "second" digital video stream), all of the mechanisms and principles discussed above regarding collection, event detection, synchronization, pattern detection, generation and publishing of digital video streams can be applied.
[0180] The same is generally true for Figure 8b, which shows a method according to the second aspect of the invention, Figure 8c, which shows a method according to the third aspect of the invention, and Figure 8d, which shows a method according to the fourth aspect of the invention.
[0181] The first, second, third and fourth aspects can be freely combined, and in particular the method according to the fourth aspect can be used in combination with any one of the methods according to the first, second and third aspects.
[0182] Furthermore, FIG. 9 is a simplified diagram of a system 100 having an arrangement for carrying out the method illustrated in FIGS. 8a to 8d.
[0183] The central server 130 includes the collection function 131 as described above.
[0184] The central server 130 also includes a first generating function 135′, a second generating function 135″, and a third generating function 135′′. Each such generating function 135′, 135″, 135′′ corresponds to a generating function 135, and what has been discussed above with respect to generating function 135 applies equally to generating functions 135′, 135″, 135′′. Depending on the detailed configuration of the central server 130, generating functions 135′, 135″, 135′′ may be distinct, multiple functions may be shared within a single logical function, and there may be more than three generating functions. Generating functions 135′, 135″, 135′′ may in some cases be different functional aspects of the same generating function 135. Various communications between generating functions 135′, 135″, 135′′ and other entities may occur via appropriate APIs.
[0185] It will further be appreciated that there may be a separate collection function 131 for each of the production functions 135′, 135″, 135′″ or groups of such production functions, and depending on the detailed configuration, there may be multiple logically separated central servers 130, each with its own collection function 131.
[0186] Additionally, the central server 130 includes a first publishing function 136′, a second publishing function 136″, and a third publishing function 136′′. Each such publishing function 136′, 136″, 136′′ corresponds to a publishing function 136, and what has been described above with respect to publishing function 136 applies equally to publishing functions 136′, 136″, 136′′. Depending on the detailed configuration of the central server 130, publishing functions 136′, 136″, 136′′ may be separate functions, may be co-located in a single logical function with multiple functions, or there may be more than three publishing functions. Publishing functions 136′, 136″, 136′′ may in some cases be different functional aspects of the same publishing function 136.
[0187] In FIG. 9, for purposes of illustrating the principles disclosed herein, three sets or groups of participant clients are shown, each corresponding to a participant client 121 described above. Thus, there is a first group 121′ of such participant clients 121, a second group 121″ of such participant clients, and a third group 121′″ of such participant clients. Each of these groups may consist of one, or preferably at least two, participant clients. Depending on the detailed configuration, there may be only two such groups, or there may be more than two. The assignment between groups 121′, 121″, 121′″ may be exclusive, in the sense that each participant client 121 is assigned to at most one group 121′, 121″, 121′″. In an alternative configuration, at least one participant client 121 may be assigned to more than one such group 121′, 121″, 121′″ simultaneously.
[0188] FIG. 9 also shows an external consumer 150, and as noted above, it is understood that there may be multiple such external consumers 150.
[0189] While FIG. 9 does not show the video communication service 110 for simplicity's sake, it will be understood that a video communication service of the general type described above may be used in conjunction with the central server 130, for example, to provide a shared video communication service to each participant client 121 using the central server 130 in the manner described above.
[0190] Returning to FIG. 8a, the method begins with a first step.
[0191] In a subsequent collection step, each of the multiple primary video streams, in this exemplary case at least a first primary digital video stream, a second primary digital video stream, and a third primary digital video stream, are collected from a respective participant client 121. Thus, the first primary digital video stream is collected from the first participant client, the second primary digital video stream is collected from the second participant client, and the third primary digital video stream is collected from the third participant client.
[0192] In a subsequent publishing step, at least one video stream is provided to at least one of the first participant client and the second participant client, i.e., the video stream is at least one of the first primary digital video stream, the second primary digital video stream, and a first generated video stream generated based on at least one of the first and second primary video streams. The generation of such a primary video stream may be performed by a first generation function 135′, as described below, and may include, for example, introducing a delay in the first generated digital video stream as a result of the generation.
[0193] Such provision or publication may be continuous and may be in real time.
[0194] For example, a first and second participant may participate in the same video communication service, such as a video conference, as described elsewhere herein. For example, a first participant client 121 may be provided with a second digital video stream for viewing on the screen 124 of the first participant client 121, or vice versa, thereby allowing users 122 of the first and second participant clients 121 to see and interact with each other. Additionally or alternatively, each or one of the first and second participant clients may be provided with a first generated digital video stream for viewing on the respective screen 124 of that participant client 121. When both the first generated digital video stream and one of the first and second primary video streams are provided in conjunction, the primary video stream may be delayed, as described below, to time-synchronize the video streams displayed at that participant client 121.
[0195] In a subsequent second generation step performed by second generation function 135'', a second generated video stream is generated as a digital video stream that is based on the first primary digital video stream and the second primary digital video stream, and also based on a third primary digital video stream. Note that the third primary digital video stream is preferably not provided to the first or second participant clients, either as is or as part of the generated digital video stream. As described elsewhere herein, the first and second participant clients may be assigned to different groups of participant clients compared to the third participant client.
[0196] The second generating step includes introducing a time delay such that the second generated video stream is asynchronous in time with any of the video streams that may be provided to the first or second participant clients in the publishing step. This time delay may be added intentionally and / or be a direct result of generating the second generated digital video stream in any of the ways described herein. Preferably, the second generated digital video stream is available for publication with a delay relative to any published video streams at the first and / or second participant clients. One way to think about this is that any consumer client of the second generated digital video stream will consume this second generated digital video stream in a "time zone" that is slightly later (in time) than the video stream consumption "time zones" of the first and second participant clients.
[0197] For example, when one or more primary digital video streams are provided to a first and / or second participant client, such provision may be direct (without the use of an intentionally introduced time delay) and / or may involve only relatively computationally lightweight processing prior to provision to the participant client, while generation of a second generated digital video stream may involve an intentionally introduced time delay and / or relatively heavyweight processing, resulting in the second generated digital video stream being generated with a delay for earliest release relative to the earliest delay for release of the first and / or second primary digital video stream. When a first generated video stream is provided to a first and / or second participant client, the first generated digital video stream is generated with a relatively short intentionally applied time delay and / or relatively lightweight processing, while the second generated digital video stream is generated with a relatively long intentionally applied time delay and / or relatively heavyweight processing, resulting in the second generated digital video stream being generated with a delay for earliest release relative to the earliest delay of the first generated digital video stream.
[0198] Typically, the second generated digital video stream is not provided for publication at the first or second participant client, but is provided at a participating client that is assigned to a different group (such as the first group 121′) from the group to which the first and second clients belong, such as, for example, a third participant client (assigned to a different group, such as the second group 121″) and / or an external consumer client 151.
[0199] Therefore, as shown in FIG. 8a, the publishing step further includes continuously providing the second generated video stream to at least one consumer client 121, 150 that is not the first or second participant client.
[0200] As also shown in Figure 8a, the method may be repeated to continuously generate and provide / publish the digital video stream.
[0201] The method ends in the following steps.
[0202] FIG. 8b shows a method according to a second embodiment.
[0203] In the first step the method starts.
[0204] In a subsequent collection step, multiple primary video streams are collected, in this illustrative example, at least a first primary digital video stream and a second primary digital video stream collected from each participant client 121′ selected from a first group of participant clients.
[0205] As with the method shown in Figure 8a, collection may be performed as described above, with collection function 131 processing the raw data, for example, without performing re-encoding. There may also be event detection, synchronization, and pattern detection steps of the general type described above that are applied to the primary digital video streams collected from the first group of participant clients 121' for the purpose of generating the first resulting digital video stream.
[0206] That is, in a subsequent first generation step, the first generation function 135′ receives the first and second primary video streams as respective digital video streams from the collection function 131 and generates a first generated digital video stream based on the first and second primary digital video streams. Preferably, the first digital video stream is not generated based on other participant clients 121 connected to the same video communication service 110 other than the participant clients assigned to the first group 121′, and is not generated in a manner that allows such other participant clients 121 to interact with members of the first group 121′ over the video communication service 110. On the other hand, the first generated video stream may be generated based on other information, such as an external video feed, static data, or graphics. For clarity, these and other aspects described in connection with FIG. 8b may also apply to the methods shown in FIGS. 8a, 8c, and 8d.
[0207] The result of this first generation is thus a generated digital video stream of the type described above, which may visually include, for example, as a subpart, one or more of said primary video streams in processed or unprocessed form. This first generated video stream may include live-captured video streams, slides, externally provided video or images, etc., as generally described above in connection with the video output streams generated by the central server 130. The first generated video stream may also be generated based on detected events and / or patterns in the intentionally delayed or real-time first and / or second primary video streams provided by the participant clients of the first group 121′, in the manner generally described above.
[0208] In a subsequent second generation step, a second generated digital video stream is generated as a digital video stream based on both the first generated video stream and the first and second primary digital video streams collected from the first group of participant clients 121′. The first and second primary digital video streams may be provided from the collection function 131 to the second generation function 135″, while the first generated video stream may be provided from the first generation function 135′ to the second generation function 135″. If the first and second generation functions 135′, 135″ are (and may be) one and the same logical unit, generation is simply performed in two successive steps in that generation function.
[0209] The method ends in the following steps.
[0210] It will be appreciated that the first and / or second primary video streams provided to the second generation step 135'' may be pre-formatted in a variety of ways prior to the second generation step 135'', and may also be intentionally delayed to detect events and / or patterns, as outlined above.
[0211] The second generation step may be similar to any of the generation steps described above, and everything described above with respect to the functionality of generation functions 135, 135' is also correspondingly applicable to second generation function 135". For example, second generation function 135" may generate the second generated video stream by formatting the primary video stream in various ways as part of the generation process.
[0212] As described above, the first and second primary video streams may be time-synchronized with each other before being supplied to the first generation function 135', for example, using a common time reference in any of the ways described above.
[0213] However, in the second generation step, the first and second primary digital video streams may be intentionally time-delayed (e.g., in addition to any already applied time delays implemented to time-synchronize the multiple primary video streams with one another and / or to enable event and / or pattern detection for use by the first generation function 135′). The purpose and result of the intentionally introduced time delay here is to time-synchronize the first generated video stream with the first generated video stream before being used in the second generated video stream. Thus, the additional delay introduced with respect to the first and second primary video streams is then determined to be equal to, substantially equal to, or at least as a function of, the latency associated with performing the first generation step. The exact latency to be added may be determined, for example, based on a detected common time reference of the general type described above.
[0214] That is, the first generation step, in which the first generated video stream is generated, is typically associated with some latency (due to data processing in the first generation step 135' itself), which may depend, for example, on the available computing power and the complexity of the first generation step 135'. This latency is typically not present in the first and second video streams themselves (or any latency is small in any case), which are simply captured, optionally processed in the manner described above, and then provided by the collection function 131 to and used by the second generation function 135''.
[0215] Taking into account the latency of the first generated video stream resulting from the first generation step, intentionally introducing this (additional) delay into the first and second primary video streams to time-synchronize these three video streams makes it possible to generate the second video stream without synchronization issues, even if the second generated video stream is generated not only based on the first and second primary video streams but also based on the first generated video stream which is generated based on the same first and second primary video streams, i.e. the second generated video stream is generated based on time-delayed first and second primary digital video streams.
[0216] The first generated video stream may thus be fed to a second generation step 135'', which is thus generated using two (or more) generation steps 135', 135'', where the same primary video stream is used in at least two such generation steps associated with different latencies relative to a common time reference of the primary video stream provided by the collection function 131.
[0217] In an exemplary embodiment, the participant clients of a first group 121′ are part of a discussion panel and communicate using the video communication service 110 with relatively low latency, and each of these participant clients is continuously supplied with a first generated video stream (or each other's respective primary video stream, as described above in connection with FIG. 8a). The audience of the discussion panel is made up of the participant clients of a second group 121″, which are continuously supplied with a second generated video stream, which in turn is associated with slightly higher latency. The second generated video stream may be automatically generated in the general manner described above to automatically shift between a view of an individual speaker of the discussion panel (participant clients assigned to the first group 121′; such view is provided directly from the collection function 131) and a generated view showing all the speakers of the discussion panel (this view being the first generated video stream). Using the present invention according to the first and / or second aspects, the panel speakers can interact with each other with minimal latency, while the audience can enjoy a well-produced experience.
[0218] The delay intentionally added to the first and second primary video streams in relation to the second generation step may be at least 0.1 seconds, for example at least 0.2 seconds, for example at least 0.5 seconds, and may be up to 5 seconds, for example up to 2 seconds, for example up to 1 second.
[0219] It will be appreciated that the first and second primary video streams, and the first generated video stream, may all additionally be intentionally delayed to improve pattern detection for use in the second generation function 135'' in the general manner described above.
[0220] FIG. 9 illustrates a number of alternative or concurrent ways of publishing the various generated video streams generated by the central server 130.
[0221] Generally, in a subsequent publishing step performed by a first publishing function 136′ configured to receive the first generated video stream from the first generating function 135′, the first generated video stream may be continuously provided to at least one of a first participant client 121 and a second participant client 121. For example, this first participant client may be a participant client from the group 121′ that provides the first primary digital video stream, and / or the second participant client may be a participant client from the group 121′ that provides the second primary digital video stream.
[0222] In other words, the first resulting video stream may be continuously provided to at least one of the first participant client and the second participant client.
[0223] In some embodiments, one or more of the participant clients of the group 121′ may receive the second generated video stream by a second publishing function 136″ configured to receive the second generated video stream from the second generating function 135″.
[0224] Thus, each participant client assigned to the first group 121′ that provides a primary video stream, if not directly provided with the primary digital video stream, may be provided with a first generated video stream that includes a certain delay or latency due to synchronization between multiple primary video streams and may further include a delay or latency that may be intentionally added to allow sufficient time for event and / or pattern detection, as described above.
[0225] Correspondingly, each of the participant clients assigned to the second group 121″ may be provided with a second generated video stream that also includes an intentionally added delay associated with the second generation step, added for the purpose of time-synchronizing the first generated video stream with the first and second primary video streams. This excess delay may or may not cause communication failures between the participant clients in the second group 121″, for example, due to the participant clients in the second group 121″ interacting with the video communication service 110 differently than the participant clients in the first group 121′ (see below). In other embodiments (e.g., where the participant clients in the first group 121′ are provided with the primary digital video stream directly), each of the participant clients assigned to the second group 121″ may be provided with the first generated video stream directly.
[0226] Thus, the participant clients in the first group 121′ form a subgroup of all participant clients 121 currently participating in the video communication service 110 and are present at and using the service in a “time zone” slightly ahead (e.g., 1-3 seconds ahead) of any participant clients being continuously provided with a generated video stream, such as the first generated video stream or the second generated video stream. Nonetheless, other participant clients (not assigned to the first group 121′ but instead assigned to the second group 121″) are continuously provided with a second generated video stream, which is generated based on the first and second primary video streams (which may include either or both at each point in time), but in a slightly later “time zone.” Because the first generated video stream is generated directly based on the first and second primary video streams, no delay or latency is added to time-synchronize it with a video stream already generated based on the primary video streams themselves, providing these participant clients 121 with a more direct, low-latency experience of the video communication service 110.
[0227] Again, this may mean that participant clients 121 assigned to the first group 121' are not provided with access to the second generated video stream.
[0228] That is, the first and second primary digital video streams may be provided as part of a shared digital video communication service 1100 of the general type described above, and the first participant client and the second participant client (belonging to the same first group 121′) may both be respective remotely connected participant clients to the shared digital video communication service 1100. The participant clients in the second group 121″ (and also the participant clients in the third group 121″) may be participant clients remotely connected to the shared digital video communication service 1100.
[0229] In this context, it is understood that "remote connection" does not mean that such participant client 121 or corresponding user 122 is necessarily located in a different room, premises, or geographic location, but rather that the user 122 uses the participant client 121 to interact audibly / visually with the video communication service 110.
[0230] The collecting step may include collecting the first and / or second primary digital video streams from the shared digital video communication service 110, for example, in any of the ways described above.
[0231] Figure 8c shows a method according to a third embodiment. As mentioned above, the method shown in Figure 8c is similar to the method shown in Figure 8d, but is also similar to the methods shown in Figures 8a and 8b, and these four embodiments of the invention can be freely combined. Everything said in relation to one of these embodiments can easily be applied to the other embodiments in a corresponding manner, so long as they are compatible.
[0232] In the first step the method starts.
[0233] In a subsequent collection step, a first primary digital video stream is collected from a first participant client assigned to the first group 121′, and a second primary digital video stream is collected from a second participant client also assigned to the first group 121′. Additionally, a third digital video stream is collected from a third participant client that may not be assigned to the first group 121′. For example, the third participant client may be assigned to the second group 121″. This collection step may be similar to the collection step described in connection with FIG. 8b.
[0234] In a subsequent first generation step, which may be similar to the first generation step described in connection with Figure 8b, a first generated video stream may be generated as a digital video stream based on the collected first and second primary digital video streams, where it should be noted that the first generated video stream does not have to be generated based on the third primary video stream.
[0235] The first originating digital video stream is continuously generated with a first latency for presentation to a number of consumer clients. In other words, according to this third aspect, the first originating digital video stream is generated such that presentation of each newly generated frame of the first originating digital video stream occurs with a first latency if the frame is presented immediately upon its generation.
[0236] A subsequent second generation step may be similar to the second generation step described in relation to Figure 8b, but the second generated video stream is generated as a digital video stream based on all three primary digital streams, in other words, all of the first, second and third primary digital video streams.
[0237] A second generated digital video stream is continuously generated for publication at a second latency, in a manner corresponding to the first generated video stream and the first latency, the second latency being greater than the first latency, meaning that if the first and second generated video streams both include, for example, frames from the first primary video stream, such frames will be displayed earlier in the instantaneous publication of the first generated video stream compared to the instantaneous publication of the second generated video stream.
[0238] In a subsequent publishing step, which may be similar to the publishing step described in connection with Figure 8b, at least one of the first primary digital video stream, the second primary digital video stream, and the first generated video stream (e.g., any set of one or more of these multiple streams) is provided sequentially to at least one of the first participant client and the second participant client, similar to the method described in connection with Figure 8a above.
[0239] The second resulting video stream is also continuously provided to at least one other participant client.
[0240] The method ends in the following steps.
[0241] The same example used to illustrate the practical application of the solution of the second aspect can also be used to illustrate how this third aspect can be implemented: Since the third primary video stream is collected from participant clients of the second group 121′ that have lower latency requirements than the participant clients of the first group 121′ that provide the first and second primary video streams, the second generated video stream can be provided in a higher latency manner to achieve the desired automatic generation, while the panel discussion speakers of the first group 121′ can interact with lower latency.
[0242] Of course, in addition to the third primary video stream, there may be more primary video streams provided by the second group 121'', which may be used correspondingly.
[0243] 8b and 8c, the second generated video stream may be continuously provided to at least one consumer client that is not the first or second participant client. More generally, the second generated video stream may be continuously provided to participant clients 121 that are not assigned to the first group 121′ and / or external consumers 150.
[0244] As noted above, the collecting step 131 may include collecting at least one of the primary digital video streams, e.g., an additional primary video stream in addition to the first and second primary video streams, as an external digital video stream 301 of the type described above, collected from an information source 300 that is external to the shared digital video communication service 110. Also as noted above, such external video stream 301 may be time-synchronized to the first and second primary video streams by a synchronization function 133 logically located (in terms of data flow) between the collecting function 131 and the first generating function 135′. Correspondingly, the same applies to the third, fourth, and fifth primary video streams described herein. In that case, the first and / or second generated video streams may be generated based on the external digital video stream 301.
[0245] Also, as generally described above, the first generating step 135′ and / or the second generating step 135″ may further include generating the respective generated (first and / or second) video stream based on a set of predetermined and / or dynamically variable parameters relating to: the visibility of individual ones of the first and / or second primary digital video streams 210 in the generated digital video stream; the placement of visual and / or audio video content; the visual or audio effects used; and / or the output mode of the generated digital video stream.
[0246] Also as mentioned above, the first generation step 135′ and / or the second generation step 135″ may be performed by the central server 130, which provides the second generated video stream as a live video stream to one or more concurrent (external and / or participating) consumer clients via an API 137 of the general type described above.
[0247] Therefore, different groups 121′, 121″, 121′″ of participant clients 121 may have different requirements in terms of delay tolerance. This is especially true when participating in one and the same live video communication service 110 as participant clients 121. This is further illustrated below.
[0248] Figure 8d shows a method according to a fourth aspect of the invention.
[0249] In the first step the method starts.
[0250] In general, as shown in Figure 8d, the method for generating a second generated digital video stream may include a subsequent allocation step, which may be an initial step but may also be performed at any point during the method, for example as a reallocation step.
[0251] This allocation step may allocate a plurality of participant clients 121 across at least two groups 121′, 121″, 121′″ of such participant clients 121. In this embodiment, the allocation is made to at least a first group 121′ and a third group 121″, although participant clients 121 may of course also be allocated to the third group 121″.
[0252] More specifically, the first and second primary digital video streams may be collected, for example, by the collection function 131 and in a subsequent collection step, from each participant client 121 assigned to the first group of participant clients 121′. However, the fourth and fifth primary digital video streams may also be collected by the collection function 131 and in the above collection step from each participant client 121 assigned to the third group of participant clients 121′″.
[0253] 9, the participant clients 121 assigned to the third group 121'' may have less stringent latency requirements than the participant clients 121 assigned to the first group 121'. For example, the participant clients 121 in the first group 121' may be members of the discussion panel described above (which interact with each other in real time and therefore require low latency), while the participant clients 121 in the third group 121'' may constitute an expert panel or similar panel that interacts with the panel but in a more structured manner (such as using clear questions / answers) and therefore can tolerate greater latency than the first group 121'.
[0254] The first resulting video stream may be generated by the first generation function 135' based on the first and second primary video streams (as well as any additional input content as described), as described above. The third resulting video stream is generated in a corresponding manner, but by the third generation function 135''' and based on (at least) the fourth and fifth primary video streams.
[0255] Both the first generated video stream and the third generated video stream may optionally be provided to a second generation function 135'' to form the basis for generating a second generated video stream.
[0256] However, according to this fourth aspect, in a second generating step performed by second generating function 135″, the second generated video stream is generated based on at least one of the first and second primary video streams and further based on at least one of the fourth and fifth primary video streams in a manner as described above. Second generating function 135″ may be provided with the fourth and fifth primary video streams from collection function 131 in a manner corresponding to the first and second primary video streams, including any cross-stream time synchronization, event detection, etc. It is particularly noted that the second generated video stream may be based directly or indirectly on the first and / or second primary video streams, for example, by basing the second generated video stream on the first generated video stream which is based on the first and second primary video streams; correspondingly, the same is true for the fourth and fifth primary video streams and the third generated video stream.
[0257] The third generated video stream is generated in a subsequent third generation step.
[0258] According to a fourth aspect, the third generating step includes intentionally introducing a time delay into the fourth and fifth primary video streams such that they are time-synchronized with each other but time-asynchronous (not time-synchronized) with respect to the first generated video stream. It will be appreciated that this time delay can be introduced in the third generating function 135'" itself or in a corresponding synchronization function 133 upstream of said third generating function 135'".
[0259] Thus, the first generating step 135' may involve introducing an intentional delay or latency of the type described above that is introduced in addition to the delay introduced as part of the synchronization of the first and second primary video streams, e.g., to achieve sufficient time to perform efficient event and / or pattern detection. The introduction of such an intentional delay or latency may be performed as part of the synchronization performed by the synchronization function 133 described above (not shown in FIG. 9 for reasons of simplicity). Similarly, the third generating step 135''' may introduce an intentional delay or latency that is different from the delay or latency intentionally introduced for the first generating step 135'.
[0260] In particular, the intentionally introduced delay or latency results in a temporal desynchronization between the first and third generated video streams, meaning that the first and third generated video streams would not follow a common timeline if they were exposed immediately and consecutively upon generation of each individual frame.
[0261] As mentioned above, the second generated video stream may be associated with a larger delay than the first generated video stream, and in some cases may also be associated with a larger delay than the third generated video stream. Accordingly, the second generation functionality 135'' may be configured to synchronize the first, second, fourth, and fifth primary video streams by adding respective additional delays to the first, second, fourth, and fifth primary video streams before combining them into the second generated video stream.
[0262] In the publishing step, the third generated video stream may be continuously provided to at least one participant client assigned to the third group 121''', where it may be continuously published to that user 122. Similarly, the first generated video stream may be continuously provided to at least one participant client assigned to the first group 121', where it may be published to that user 122; and / or the second generated video stream may be provided and published as described above.
[0263] The method ends in the following steps.
[0264] Thus, in this fourth aspect, three independently generated video streams may be simultaneously generated and consumed / published in different “time zones.” Although they are based at least in part on the same primary video material, the generated video streams are published with different latencies. The first group 121′ requires the lowest latency and can interact using the first generated video stream, offering very low latency. The third group 121′″ is willing to accept slightly higher latency and can interact using the second generated video stream, offering more latency, but on the other hand, offering greater flexibility in terms of intentionally added delay to achieve better automatic generation, as disclosed elsewhere herein. Meanwhile, the second group 121″, which is less sensitive to delay, can incorporate material from the first group 121′ and the third group 121′″ and can enjoy interacting using the second generated video stream, which is automatically generated in a very flexible manner. It is particularly noted that all these groups of participant users 121', 121'', 121''' interact with each other using the video communication service 110, albeit using the various delays described above and therefore operating in different "time zones". However, due to the synchronization of the individual input video streams at each production function, the participant users 121 are unaware of the different latencies from their respective perspectives.
[0265] The first generating step 135' may include time delaying the first and second primary video streams so as to time synchronize them with one another, as described above.
[0266] Correspondingly, the third generating step 135'' (or corresponding synchronizing step 133) above may include time delaying the fourth and fifth primary video streams so that they are time synchronized with one another. However, here a maximum time delay is used that is greater than the maximum time delay used to time delay the first and second primary video streams in the first generating step 135' (or corresponding synchronizing step 133), so that the first generated video stream is not time synchronized with the third generated video stream in the manner described above.
[0267] As described above, each participant client 121 assigned to each of the above groups 121′, 121″, 121′″ can participate in one and the same video communication service 110 in which the second generated video stream is continuously published.
[0268] Different ones of the groups 121', 121'', 121''' may be associated with different participant interaction privileges in the video communication service 110, and different ones of the groups 121', 121'', 121''' may be associated with different maximum time delays (latencies) used to generate the respective generated video streams that are made available to the participant clients 121 assigned to the groups 121', 121'', 121'''.
[0269] For example, a first group of participant clients 121' in a panel discussion may be associated with full interaction privileges, and may speak whenever they wish. A third group of participant clients 121''' may be associated with slightly more restricted interaction privileges, such as requiring the video communication service 110 to request the floor before they can speak by unmuting their microphones. A second group of audience participant users 121'' may be associated with even more restricted interaction privileges, such as only being able to pose questions via text postings to a common chat room, but not being able to speak.
[0270] Thus, different groups of participant users may be associated with different interaction privileges and different latency times for the respective generated video streams exposed to them, such that latency time is an increasing function of decreasing interaction privileges. The more freely a given participant user 121 is permitted by the video communication service 110 to interact with other users, the lower the tolerable latency time. The lower the tolerable latency time, the less likely the corresponding auto-generated feature will take detected events, patterns, etc. into account.
[0271] The group with the greatest latency may be a viewer-only group with no right to interact other than passively participating in the video communication service.
[0272] In particular, the respective maximum time delay (latency) for each of the groups 121′, 121″, 121′″ may be determined as the difference in maximum latency across all primary video streams and any resulting video streams that are continuously exposed to participant clients in that group. This sum may include additional time delays intentionally added for the purpose of detecting events and / or patterns, as described above.
[0273] As used herein, the terms “generation” and “generated digital video stream” may refer to different types of generation. In one example, a single, well-defined digital video stream is generated by a central entity, such as a central server 130, to form the generated digital video stream for provision and publication to each of a specific set of participant clients 121 that will consume the generated digital video stream. In other aspects, different individual such participant clients 121 may view slightly different versions of the generated digital video stream. For example, a generated digital video stream may include multiple individual or combined digital video streams that a local software function 125 of a participant client 121 may allow the user 122 to switch between, place on a screen 124, or otherwise configure or process. Often, what matters is in what “time zone” (i.e., what latency) a generated digital video stream containing time-synchronized subcomponents is provided. Thus, the case described above in connection with FIG. 8a in which the first and second participant clients are provided with each other's primary video streams can be considered as a first generated digital video stream being provided to the first and second participant clients (in the sense that a time-synchronized set of raw or processed first and second primary digital video streams is made available to both the first and second participant clients).
[0274] To further clarify and illustrate the use of the above-described groups of participant clients 121', 121'', 121''', the following example is described in the form of a video communication services meeting involving three different simultaneous "time zones".
[0275] A first group of participant clients 121′ experience interactions in real time, or at least near real time (depending on unavoidable hardware and software delays). These participant clients are provided with video, including audio, from one another to facilitate such interactions and communication between the users 122. The first group 121′ can serve the core users 122 of the meeting, whose interactions may be of interest to other participant clients (other than the first group 121′) to participate in.
[0276] A second group of such other participant clients 121″ participates in the same meeting but is in a different “time zone,” further away from real time than the first group of participant clients 121′. For example, the second group 121″ may be an audience with interactive privileges, such as being able to pose questions to the first group 121′. The “time zone” of the second group 121″ may have a delay relative to the “time zone” of the first group 121′, such that posed questions and answers are associated with a noticeable but short delay. On the other hand, this slightly larger delay allows the participant clients of this second group 121″ to experience the automatically generated digital video stream in a more complex way, providing a more pleasant user experience.
[0277] A third group of such other participant clients 121''' also participates in the same meeting, but only as viewers. This third group 121''' consumes a generated digital video stream, which may be generated automatically in a more sophisticated and complex manner, that is consumed in a third "time zone" having an even greater delay than the second "time zone." However, because the third group 121''' cannot provide input to the communication service in a way that affects the first group 121' and the second group 121'', the third group 121''' experiences the conference as occurring in "real time" and in a convincingly staged manner.
[0278] Of course, there may be more than two such groups of participant clients, each associated with a meeting "time zone" of increasing time delay and complexity to generate, using the principles described herein.
[0279] The present invention also relates to computer software functionality for providing a second digital video stream in accordance with the above, and such computer software functionality may be configured, when executed, to perform at least some of the collection, event detection, synchronization, pattern detection, generation, and publishing steps described above, particularly with respect to the first, second, third, and / or fourth aspects. The computer software functionality may be configured to run on physical or virtual hardware of the central server 130, as described above.
[0280] The present invention also relates to a system for providing a second digital video stream, such a system 100, in turn comprising a central server 130. The central server 103, in turn, may be configured to perform at least some of the collection, event detection, synchronization, pattern detection, generation and publishing steps described above, particularly with respect to the first, second, third and / or fourth aspects. For example, these steps may be performed by the central server 130 executing said computer software functions for performing said steps as described above.
[0281] It will be appreciated that the principles of auto-generation based on a set of available input video streams described above, including time synchronization of such input video streams, event and / or pattern detection, etc., may be applied simultaneously at different levels, and thus one such auto-generated video stream may form an available input video stream for a downstream auto-generation function that generates a video stream.
[0282] The central server 130 may be configured to control the assignment of groups 121′, 121″, 121′″ to individual participant clients 121. For example, dynamically changing the group assignment for a particular such participant client during the course of a live video communication service session may be part of the automatic generation of that video communication service by the central server 130. Such reallocation may be triggered dynamically based on a predetermined timetable or in response to requests (provided via the individual participant client's user 122) of that client 121, for example, as a function of parameter data that may change dynamically over time.
[0283] Correspondingly, the central server 130 may be configured to dynamically change the group composition during the course of the video communication service, such as using a particular group only for a predetermined time period (e.g., during a scheduled panel discussion).
[0284] One possible practical solution for group assignment is to use the concept of "breakout rooms" available in some videoconferencing systems. Participant clients 121 assigned to a particular group 121', 121'', 121''' are assigned to such a breakout room, from which the central server 130 can then retrieve video stream data, such as individual primary or generated video streams, for use in downstream generation steps at the central server 130. The extraction of such video streams may itself be performed as described above.
[0285] In all the above aspects, the present invention may further comprise an interaction step in which at least one participant client of a first group (the first group associated with a first latency) interacts bi-directionally (two-way) with at least one participant client of a second group (the second group associated with a second latency), where the second latency is different from the first latency, it being understood that all these participant clients may be participants in one and the same communication service of the type described above.
[0286] In such cases, it is preferable for participant clients associated with different latencies (or "time zones" as discussed above) to be temporarily associated with the same "time zone," or in other words, the same latency. For example, this may be done by temporarily providing one of the participant clients temporarily associated with a higher latency with one or more primary / generated digital video streams that have been generated using a lower latency than the latency associated with that participant client. In other words, if a participant client that is normally continuously provided with one or more video streams having a higher latency wishes to interact with a participant client that is continuously provided with one or more video streams having a lower latency, the former participant client is instead temporarily provided with one or more of the latter video streams. Thus, the high-latency participant client temporarily switches to the low-latency "time zone" associated with the low-latency participant client. After the interaction ends, the high-latency participant client returns to the high-latency communication environment it was using before the interaction.
[0287] For example, an audience member in the panel discussion described above may wish to pose a question. In this case, that audience member is given the opportunity to speak and switched into the panel discussion "time zone." This means that that audience member will be watching the panel with a lower latency, but less elaborate presentation. More specifically, the audience member may watch the same video stream or streams provided to the panel members during the interaction. The rest of the audience will not notice the difference because they remain in the higher latency audience "time zone." After the interaction between the speaking audience member and the panel, the speaking audience member will again be provided with the higher latency video stream or streams as they were before the interaction.
[0288] Switching between different "time zones" may be performed automatically by the central server 130.
[0289] Although preferred embodiments have been described above, it will be apparent to those skilled in the art that many modifications can be made to the disclosed embodiments without departing from the essential concepts of the invention.
[0290] For example, many additional features not described herein may be provided as part of the system 100 described herein. In general, the solutions disclosed herein provide a framework upon which detailed functionality and features can be built to accommodate a wide variety of specific applications in which streams of video data are used for communication.
[0291] An example is a demonstration situation, where the primary video stream includes the presenter's view, a shared digital slide-based presentation, and live video of the product being demonstrated.
[0292] Another example is an educational situation, where the primary video stream includes footage of the teacher, live footage of the physical entity being taught, and live footage of multiple students asking questions and interacting with the teacher.
[0293] In either of these two examples, a video communication service (which may or may not be part of the system) may provide one or more of the primary video streams, and / or multiple primary video streams may be provided as external video sources of the type disclosed herein.
[0294] The various groups have been exemplified as a discussion panel, an expert panel, and an audience. However, it is possible to divide the participant users in a digital video communication service into two or more groups, reflecting the current focus and structure of the communication being conducted. For example, one or more groups may include participant users who remotely access the video communication service from different geographic locations, while one or more other groups may include participant users who access the video communication service from a common central location, such as a lecture hall. In all such cases, the same principles as described above apply.
[0295] In general, anything disclosed with respect to the present method is applicable to the present system and computer software product, and vice versa.
[0296] Therefore, the invention is not limited to the described embodiments, but can be modified within the scope of the appended claims.
Claims
1. 1. A method for providing a second resulting video stream, the method comprising the steps of: In a first generating step (135'), a first primary digital video stream and generating a first generated video stream based on the first primary digital video stream and the second primary digital video stream, wherein the first generated video stream is generated consecutively for presentation with a first delay; In a second generating step (135''), generating the second generated video stream based on the first and second primary digital video streams, wherein the second generated video stream is generated consecutively for presentation with a second delay, the second delay being greater than the first delay; and In a publishing step (136'), the first generated video stream is published to a first participant. continuously providing the first resulting video stream to a first participant client (121) and continuously providing the second resulting video stream to a second participant client (121); In an interactive step, temporarily providing the first generated video stream to the first and second participant clients (121); and After the interaction step, the second generated video stream is provided again to the second participant client (121).
2. 10. The method of claim 1, The first participant client (121) belongs to a first group and is associated with the first delay, and / or the second participant client (121) belongs to a second group and is associated with the second delay.
3. 10. The method of claim 1, further comprising: The second resulting video stream is continuously provided to at least one other participant client (121), for example belonging to the second group.
4. 4. The method of any one of claims 1 to 3, further comprising: In said publishing step (136''), said second generated video stream is continuously provided to at least one consumer client (121; 150) that is not said first or second participating client.
5. 4. The method of any one of claims 1 to 3, further comprising: The first and second primary digital video streams are provided as part of a shared digital video communication service (110), and the first participant client (121) and the second participant client (121) are both participant clients remotely connected to the shared digital video communication service (110).
6. 6. The method of claim 5, The collecting step (131) includes collecting the first and / or second primary digital video streams from the shared digital video communication service (110).
7. 6. The method of claim 5, The collecting step (131) includes collecting at least one primary digital video stream as an external digital video stream (301) collected from a source (300) external to the shared digital video communication service (110), wherein: The first and / or second generated video streams are generated based on the external digital video stream (301).
8. 4. The method according to any one of claims 1 to 3, The first (135') and / or second (135'') generation step is generating each of the resulting video streams based on a set of predetermined and / or dynamically variable parameters relating to: the visibility of each of the first and / or second primary digital video streams in the visual and / or audio video content arrangement of the video streams; the visual or audio effects used; and / or the output mode of the resulting video streams.
9. 4. The method according to any one of claims 1 to 3, The first (135') and / or second (135'') generation step may be performed by a central server. The second generated video stream (230) is executed by a server (130) to provide the second generated video stream (230) as a live video stream to one or more concurrent consumer clients via an API (137).
10. 4. The method of any one of claims 1 to 3, further comprising: In an allocation step, a plurality of participant clients (121) are allocated across at least two groups (121', 121'', 121''') of such participant clients (121), wherein in the collecting step (131), the first and second primary digital video streams are allocated to a first group (121') of participant clients (121). ), and assigning the fourth and fifth primary digital video streams to a third group (121''') of participating clients (121). Collected from the participating clients (121); generating, in the second generating step (135''), the second generated video stream based on at least one of the first and second primary digital video streams and further based on at least one of the fourth and fifth primary digital video streams; in a third generating step (135'''), generating a third generated video stream based on the fourth and fifth primary digital video streams, wherein the third generating step (135''') time-delays the fourth and fifth primary digital video streams such that the third generated video stream is asynchronous in time with respect to the first generated video stream; and In the publishing step (136'''), the small number of people assigned to the third group and continuously providing the third resulting video stream to at least one participant client.
11. 11. The method of claim 10, The participant clients (121) assigned to each of the groups (121′, 121″, 121′″) participate in a video communication service (110) to which the second generated video stream is published, wherein the method further includes: Associating different ones of the groups (121', 121'', 121''') with different participant interaction privileges in the video communication service (110); and Different ones of the groups (121', 121'', 121''') are associated with different maximum time delays used to generate the respective generated video streams that are exposed to the participant clients (121) assigned to that group (121', 121'', 121''').
12. 12. The method of claim 11, The maximum time delay for each of the groups (121', 121'', 121''') is across all primary digital video streams and any resulting video streams that are consecutively exposed to participating clients (121) in that group (121', 121'', 121'''). It is determined as the maximum differential delay.
13. 4. The method of any one of claims 1 to 3, further comprising: In a collecting step (131), a first primary digital video stream is collected from the first participant client (121) and a second primary digital video stream is collected from the second participant client (121).
14. 1. A computer program for providing a second resulting video stream, the computer program performing the following steps when executed by a computer: In a first generating step (135'), a first primary digital video stream and generating a first generated video stream based on the first primary digital video stream and the second primary digital video stream, wherein the first generated video stream is generated consecutively for presentation with a first delay; In a second generating step (135''), generating the second generated video stream based on the first and second primary digital video streams, wherein the second generated video stream is generated consecutively for presentation with a second delay, the second delay being greater than the first delay; and In a publishing step (136'), the first generated video stream is published to a first participant. continuously providing the first resulting video stream to a first participant client (121) and continuously providing the second resulting video stream to a second participant client (121); In an interactive step, temporarily providing the first generated video stream to the first and second participant clients (121); and In a providing step, after the interacting step, the second generated video stream is provided again to the second participant client (121).
15. A system (100) for providing a second generated video stream, the system (100) comprising a central server (130), the central server (130) comprising the following functions: a first generating function (135′) configured to generate a first generated video stream based on the first primary digital video stream and the second primary digital video stream, wherein the first generated video stream is generated consecutively for release with a first delay; a second generating function (135'') configured to generate the second generated video stream based on the first and second primary digital video streams, wherein the second generated video stream is generated consecutively for release at a second delay, the second delay being greater than the first delay; and a publishing function (136') configured to continuously provide the first generated video stream to a first participant client (121) and continuously provide the second generated video stream to a second participant client (121); an interactive function for temporarily providing the first generated video stream to the first and second participant clients (121); and A presentation function that presents the second generated video stream back to the second participant client (121) after the interaction step.