System and method for generating a video stream
The system addresses synchronization and processing challenges in digital video conferencing by dynamically adjusting video stream generation based on detected triggers, resulting in a seamless and synchronized output stream for improved user experience.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- LIVEARENA TECH AB
- Filing Date
- 2026-02-04
- Publication Date
- 2026-05-19
AI Technical Summary
Existing digital video conferencing systems face challenges in managing multiple input streams with varying latency, frame rates, aspect ratios, resolutions, encodings, and connectivity issues, leading to unsynchronized and computationally intensive processing that affects user experience, especially in complex meeting scenarios.
A system and method for generating a shared digital video stream that automatically detects triggers such as motion, lighting changes, or sound patterns to dynamically adjust cropping, zooming, panning, and focus, synchronizing and combining multiple input streams into a synchronized output stream for real-time consumption.
The solution provides a seamless and synchronized digital video stream that enhances user experience by addressing latency, synchronization, and hardware requirements, improving the automated generation of complex meeting scenarios.
Smart Images

Figure 2026082966000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a system, computer software product, and method for generating a digital video stream, particularly for generating a digital video stream based on two or more different digital input video streams. In a preferred embodiment, the digital video stream is generated particularly in the context of a digital video conference in which a plurality of different simultaneous connecting users participate, or a digital video conference or meeting system. The generated digital video stream may be published externally or may be published within a digital video conference or digital video conference system.
[0002] In other embodiments, the present invention is applicable to contexts where a plurality of digital video input streams are processed simultaneously and combined into a digital video stream to be generated, although not in a digital video conference. For example, such a context may be educational or instructional.
Background Art
[0003] Many digital video conference systems such as Microsoft® Teams®, Zoom®, Google® Meet® are known, and two or more participants can hold a virtual meeting using locally recorded digital video and audio and broadcast it to all participants to emulate a physical meeting.
[0004] There is a general need to improve such digital video conference solutions, particularly with regard to the generation (production) of viewing content, such as what content to show, at what time, to whom, and through what distribution channels.
[0005] For example, some systems automatically detect the participant currently speaking and display their corresponding video feed to other participants. Many systems allow sharing graphics such as the currently displayed screen, viewing windows, and digital presentations. However, as virtual meetings become more complex, it will quickly become difficult for service providers to determine which information from all currently available information should be displayed to each participant at each point in time.
[0006] In another example, a presenter moves around the stage while talking about slides from their digital presentation. In this case, the system needs to decide whether to display the presentation, the presenter, both, or switch between the two.
[0007] It may be desirable to generate one or more output digital video streams based on a large number of input digital video streams through an automated generation process, and to provide such generated digital video streams or sets of streams to one or more consuming entities. [Disclosure of the Invention] [Problems that the invention aims to solve]
[0008] However, in many cases, due to the numerous technical difficulties faced by such digital video conferencing systems, it is difficult for dynamic meeting screen layout managers and other automated generation features to choose what information should be displayed.
[0009] Firstly, because digital video conferencing prioritizes real-time performance, low latency is crucial. This becomes problematic when different incoming digital video streams are associated with different latency, frame rates, aspect ratios, or resolutions, such as when different participants join using different hardware. Often, such incoming digital video streams require processing for a well-formed user experience.
[0010] Secondly, there is the issue of time synchronization. Since various input digital video streams, such as external digital video streams and digital video streams provided by participants, are typically supplied to a central server, there is no absolute time to synchronize each of these digital video feeds. As with excessive latency, unsynchronized digital video feeds lead to a degraded user experience.
[0011] Thirdly, digital video conferencing with multiple participants may involve different digital video streams with different encodings or formats, which require decoding and re-encoding, leading to problems in terms of latency and synchronization. Furthermore, such encodings are computationally intensive and expensive in terms of hardware requirements.
[0012] Fourth, the fact that different digital video sources may be associated with different frame rates, aspect ratios, and resolutions means that memory allocation needs can change unpredictably, requiring continuous balancing. As a result, additional latency and synchronization issues may arise. Consequently, large buffers may be required.
[0013] Fifth, participants may experience various difficulties in terms of connectivity fluctuations, disconnections / reconnections, etc., which further complicates the automatic generation of a structured user experience.
[0014] These issues are amplified in more complex meeting scenarios, such as those involving a large number of participants; participants connecting using different hardware and / or software; using externally provided digital video streams; screen sharing; or multiple hosts.
[0015] When a digital video generation system for education or instruction needs to generate an output digital video stream based on multiple input digital video streams, corresponding problems arise in other contexts.
[0016] Swedish patent application SE2151267-8 (not published as of the priority date of this application) discloses various solutions to the above-mentioned problems.
[0017] Swedish patent application SE2151461-7, which is not yet published as of the effective date of this application, discloses a variety of solutions specifically for handling wait times in multi-participant digital video environments, such as when different participant groups are associated with different general wait times.
[0018] There are still challenges in achieving the automated generation of the type of meeting described above. In particular, when multiple cameras are available in such automated generation, it has proven difficult to use the image output from each of those cameras in a way that is natural and intuitive for meeting participants.
[0019] The present invention solves one or more of the above-mentioned problems. [Means for solving the problem]
[0020] Accordingly, the present invention relates to a method for providing a shared digital video stream, the method comprising the following steps: an acquisition step, The system automatically detects triggers which are patterns, wherein a) the event is an event of physical motion detected on a person or object in the first and / or second digital video stream, or b) the event is a change in lighting detected in the first and / or second digital video stream, or c) the pattern includes a predetermined image pattern having a relative change in general motion in the first and / or second digital video stream, and / or d) the pattern includes a predetermined sound pattern characterized by its sound position in the first and / or second digital video stream; wherein the trigger instructs the system to change the generation mode of the shared digital video stream according to a predetermined generation rule;In a second generation step initiated in response to the detection of the trigger, the shared digital video stream is generated as an output digital video stream based on a successively considered plurality of frames of the second digital video stream, and / or the shared video stream is generated as an output digital video stream based on a successively considered plurality of frames of the first digital video source, but with at least one of different cropping, different zooming, different panning, or different focus plane selection of the first digital video stream compared to the first generation step; and in a publishing step, the output digital video stream is continuously provided to consumers of the shared digital video stream.
[0021] The present invention also relates to a computer program that causes a computer to perform a process to provide a shared digital video stream, wherein, when the process is performed, the computer program causes the computer to perform the following steps: in a collection step, collect a first digital video stream from a first digital video source and collect a second digital stream from a second digital video source; in a first generation step, generate the shared digital video stream as an output digital video stream based on a number of sequentially considered frames of the first digital video stream such that image information from the first digital video source is visible in the shared digital video stream and image information from the second digital video source is not visible in the shared digital video stream; in a trigger detection step, the first and / or second digital The system automatically detects triggers, which are events or patterns, by digitally analyzing the digital video stream, wherein a) the event is an event of physical motion detected on a person or object in the first and / or second digital video stream, or b) the event is a change in lighting detected in the first and / or second digital video stream, or c) the pattern includes a predetermined image pattern having a relative change in general motion in the first and / or second digital video stream, and / or d) the pattern includes a predetermined sound pattern characterized by its sound position in the first and / or second digital video stream; wherein the trigger instructs to change the generation mode of the shared digital video stream according to a predetermined generation rule;In a second generation step initiated in response to the detection of the trigger, the shared digital video stream is generated as an output digital video stream based on a successively considered plurality of frames of the second digital video stream, and / or the shared video stream is generated as an output digital video stream based on a successively considered plurality of frames of the first digital video source, but with at least one of different cropping, different zooming, different panning, or different focus plane selection of the first digital video stream compared to the first generation step; and in a publishing step, the output digital video stream is continuously provided to consumers of the shared digital video stream.
[0022] The present invention also relates to a system for providing a shared digital video stream, the system comprising a central server, the central server comprising the following functions: a collection function configured to collect a first digital video stream from a first digital video source and a second digital stream from a second digital video source; a first generation function configured to generate the shared digital video stream as an output digital video stream based on a successively considered plurality of frames of the first digital video stream such that image information from the first digital video source is visible in the shared digital video stream and image information from the second digital video source is not visible in the shared digital video stream; and a digital analysis of the first and / or second digital video streams to determine events or patterns. A trigger detection function that automatically detects a trigger, wherein a) the event is an event of physical motion detected on a person or object in the first and / or second digital video stream, or b) the event is a change in lighting detected in the first and / or second digital video stream, or c) the pattern includes a predetermined image pattern having a relative change in general motion in the first and / or second digital video stream, and / or d) the pattern includes a predetermined sound pattern characterized by its sound position in the first and / or second digital video stream; wherein the trigger instructs to change the generation mode of the shared digital video stream according to a predetermined generation rule;configured to be started in response to detection of the trigger and to generate the shared digital video stream as an output digital video stream based on a plurality of consecutively considered frames of the second digital video stream such that image information from the second digital video source is visible in the shared digital video stream and / or to generate the shared video stream as an output digital video stream based on a plurality of consecutively considered frames of the first digital video source, but configured to perform at least one of different cropping, different zooming, different panning or different focus plane selection of the first digital video stream as compared to the first generation function; and a publishing function configured to continuously provide the output digital video stream to a consumer of the shared digital video stream.;
[0023] Hereinafter, the present invention will be described in detail with reference to exemplary embodiments of the present invention and the accompanying drawings.
Brief Description of the Drawings
[0024] [Figure 1] FIG. 1 is a diagram showing a first exemplary system. [Figure 2] FIG. 2 is a diagram showing a second exemplary system. [Figure 3] FIG. 3 is a diagram showing a third exemplary system. [Figure 4] FIG. 4 is a diagram showing a central server. [Figure 5] FIG. 5 is a diagram showing a first method. [Figure 6a] FIG. 6a is a diagram showing subsequent states related to different method steps in the method shown in FIG. 5. [Figure 6b] FIG. 6b is a diagram showing subsequent states related to different method steps in the method shown in FIG. 5. [Figure 6c] FIG. 6c is a diagram showing subsequent states related to different method steps in the method shown in FIG. 5. [Figure 6d] Figure 6d shows the subsequent states associated with different method steps in the method shown in Figure 5. [Figure 6e] Figure 6e shows the subsequent states associated with different method steps in the method shown in Figure 5. [Figure 6f] Figure 6f shows the subsequent states associated with different method steps in the method shown in Figure 5. [Figure 7] Figure 7 is a conceptual diagram illustrating the common protocol. [Figure 8] Figure 8 shows a fourth exemplary system. [Figure 9] Figure 9 shows the second method. [Figure 10a] Figure 10a shows the cameras and participants in the above and various configurations. [Figure 10b] Figure 10b shows the cameras and participants in the above and various configurations. [Figure 10c] Figure 10c shows the cameras and participants in the above and various configurations. [Figure 11a] Figure 11a shows the cameras and participants in the above and various configurations. [Figure 11b] Figure 11b shows the cameras and participants in the above and various configurations. [Modes for carrying out the invention]
[0025] All figures share the same or corresponding reference numerals.
[0026] Figure 1 shows a system 100 according to the present invention configured to perform a method according to the present invention for providing a digital video stream, for example, a shared digital video stream.
[0027] The system 100 may include a video communication service 110, but in some embodiments, the video communication service 110 may be located outside the system 100. Multiple video communication services 110 may be provided, as will be described below.
[0028] The system 100 may include one or more participant clients 121, but one, some, or all of the participant clients 121 may be outside the system 100 in some embodiments.
[0029] System 100 includes a central server 130.
[0030] As used herein, the term “central server” refers to a computer-implemented function configured to be accessible in a logically centralized manner, such as through a clearly defined API (Application Programming Interface). Such a central server function may be implemented purely in computer software, or in a combination of software, virtual hardware, and / or physical hardware. It may also be implemented in a standalone physical or virtual server computer, or distributed across multiple interconnected physical and / or virtual server computers.
[0031] The physical or virtual hardware that the central server 130 runs on, in other words, the computer software that defines the functions of the central server 130, may consist of a conventional CPU, a conventional GPU, conventional RAM / ROM memory, a conventional computer bus, and conventional external communication functions such as an internet connection.
[0032] The video communication service 110 is also a central server in the sense described above, insofar as it is used, and it may be a different central server from the central server 130, or it may be part of the central server 130.
[0033] In response to this, each participant client 121 may, in the corresponding interpretation, be a central server in the sense described above, where the physical or virtual hardware that each participant client 121 runs on, in other words, the computer software that defines the functions of the participant client 121, itself comprises a conventional CPU / GPU, conventional RAM / ROM memory, a conventional computer bus, and conventional external communication functions such as an internet connection.
[0034] Each participant client 121 also typically includes, or communicates with, a computer screen positioned to display video content provided to the participant client 121 as part of an ongoing video communication, a speaker positioned to emit sound content provided to the participant client 121 as part of the video communication, a video camera, and a microphone positioned to record sound locally to a human participant 122 for the video communication, the participant 122 using the participant client 121 to participate in the video communication.
[0035] In other words, each participant 122 can interact with audio / video streams provided by other participants and / or various sources through the human-machine interface of each participant client 121 during video communication.
[0036] Generally, each participant client 121 is equipped with its own input means 123, which may consist of the video camera; the microphone; a keyboard; a computer mouse or trackpad; and / or an API for receiving digital video streams, digital audio streams, and / or other digital data. The input means 123 is configured in particular to receive video streams and / or audio streams from a central server, such as a video communication service 110 and / or a central server 130, and such video streams and / or audio streams are provided as part of the video communication and are preferably generated based on corresponding digital data input streams provided to the central server from at least two sources of such digital data input streams, for example, participant client 121 and / or an external source (described later).
[0037] More generally, each participant client 121 comprises output means 124 which may consist of the computer screen; the speaker; and an API that emits a digital video and / or audio stream. Such streams represent locally captured video and / or audio to participant 122 using the participant client 121.
[0038] In practice, each participant client 121 may be a mobile device such as a cell phone equipped with a screen, speaker, microphone, and internet connectivity, and the mobile device performs the functions of the participant client 121 by running computer software locally or accessing computer software that is run remotely. Correspondingly, the participant client 121 may be a thick or thin laptop or a desktop computer, and may run locally installed applications or use functions accessed remotely via a web browser.
[0039] There may be one or more participant clients 121 used in one of the same video communications in this embodiment, for example, at least three or at least four.
[0040] There may be at least two different groups of participating clients. Each participating client may be assigned to one of these groups. The groups may reflect different roles of the participating clients, different virtual or physical locations of the participating clients, and / or different interaction rights of the participating clients.
[0041] Such roles can be diverse and may include, for example, "leader" or "conference organizer," "speaker," "panelist," "interactive audience," or "remote listener."
[0042] Such physical locations can be diverse and may include, for example, "on stage," "within a panel," "a physically present audience," or "a physically distant audience."
[0043] A virtual location may be defined in terms of a physical location, but it may also include virtual groupings that partially overlap with physical locations. For example, physically present audience members may be divided into a first virtual group and a second virtual group, and some physically present audience members may be grouped together with some physically separated audience members into the same virtual group.
[0044] A variety of dialogue permissions are available, such as "full dialogue" (no restrictions), "can speak but only after requesting a microphone" (like raising a virtual hand in a video conferencing service), "cannot speak but can write in a shared chat," or "watch / listen only."
[0045] In one embodiment, each defined role and / or physical / virtual location may be defined with respect to a certain predetermined interaction permission. In another example, all participants with the same interaction permission form a group. Thus, defined roles, locations, and / or interaction permissions can reflect diverse group assignments, and different groups may be distinct from or overlap with each other as needed.
[0046] This is illustrated with an example below.
[0047] Video communication may be provided, at least in part, by a video communication service 110 and at least in part by a central server 130, as described and illustrated herein.
[0048] Where used herein, “video communication” means a two-way digital communication session comprising at least two, preferably at least three or at least four video streams, and preferably also coinciding with an audio stream used to generate one or more mixed or joint digital video / audio streams, which may or may not contribute to the video communication via video and / or audio, and are consumed by one or more consumers (e.g., participant clients of the type described above). Such video communication is real-time and may or may not have a certain latency or delay. At least one, preferably at least two or at least four participants 122 participating in such video communication engage in the video communication in a two-way manner, providing and consuming video / audio information.
[0049] At least one, or all, of the participant clients 121 are equipped with local synchronization software functionality 125, the details of which will be described later.
[0050] The video communication service 110 may have or may have access to a common time standard, as will be described in more detail below.
[0051] Each of at least one central server 130 may include an API 137 for digital communication with entities outside of the central server 130. Such communication may include both input and output.
[0052] System 100, such as a central server 130, may be configured to digitally communicate with an external information source 300, such as a video stream, provided from an external source, and in particular to receive digital information such as audio and / or video stream data from the external information source 300. The fact that the information source 300 is “external” means that it is not provided by or as part of the central server 130. Preferably, the digital data provided by the external information source 300 is independent of the central server 130, and the central server 130 cannot influence its information content. For example, the external information source 300 may be live-captured video and / or audio, such as a public sporting event or an ongoing news event or broadcast. Furthermore, the external information source 300 may be captured by a webcam or the like, rather than by any of the participant clients 121. Thus, such captured video may depict the same local area as any one of the participant clients 121, but it is not captured as part of the participant client 121's activities. One possible difference between an externally provided information source 300 and an internally provided information source 120 is that an internally provided information source may be provided in its capacity as a participant in the video communication of the type defined above, whereas an externally provided information source 300 is not, but rather provided as part of a context that is outside the video conference.
[0053] Furthermore, there may be multiple external information sources 300 that provide the central server 130 in parallel with other digital information of that type, such as audio and / or video streams.
[0054] As shown in Figure 1, each participant client 121 constitutes the source of its respective information (video and / or audio) stream 120 that it provides to the video communication service 110, as described above.
[0055] The system 100, such as the central server 130, may be further configured to communicate digitally with external consumers 150 and, in particular, to transmit digital information to the external consumers 150. For example, digital video and / or audio streams generated by the central server 130 may be continuously provided to one or more external consumers 150 in real time or near real time via the API 137. Here again, "external" means that the consumer 150 is not provided as part of the central server 130 and / or is not a party to the video communication.
[0056] Unless otherwise specified, all functions and communications in this specification are implemented by computer software running on appropriate computer hardware and are provided digitally and electronically, communicated via digital communication networks or channels such as the Internet.
[0057] Therefore, in the configuration of system 100 shown in Figure 1, a number of participant clients 121 participate in digital video communication provided by the video communication service 110. Each participant client 121, therefore, has an ongoing login, session, or similar to the video communication service 110 and can participate in one of the same ongoing video communications provided by the video communication service 110. In other words, the video communication is “shared” among the participant clients 121 and therefore also “shared” by the corresponding human participants 122.
[0058] In Figure 1, the central server 130 includes an automated participant client 140, which is an automated client corresponding to participant client 121 but is not associated with a human participant 122. Instead, the automated participant client 140 is added to the video communication service 110 as a participant client to participate in the same shared video communication as participant client 121. As such a participant client, the automated participant client 140 is granted access to continuously generated digital video and / or audio streams(s) provided by the video communication service 110 as part of the ongoing video communication, and these streams can be consumed by the central server 130 via the automated participant client 140. Preferably, the automated participant client 140 receives from the video communication service 110 a common video and / or audio stream that is distributed or may be distributed to each participant client 121; each video and / or audio stream that is provided to the video communication service 110 from each of one or more participant clients 121 and relayed by the video communication service 110 to all participant clients 121 or requesting participant clients 121 in raw data or modified form; and / or a common time reference.
[0059] The central server 130 may have a collection function 131 configured to receive multiple video streams and / or audio streams of the above type from the automated participant clients 140 and possibly from the external information sources 300, in order to process them as described later, and then provide a generated video stream, such as a shared video stream, via the API 137. For example, this generated video stream is consumed by an external consumer 150 and / or a video communication service 110, and is distributed by the video communication service 110 to all or any one of the participant clients 121 that requests it.
[0060] Figure 2 is similar to Figure 1, but instead of using the automated participant client 140, the central server 130 receives video and / or audio stream data from the ongoing video communication via the API 112 of the video communication service 110.
[0061] Figure 3 is similar to Figure 1, but the video communication service 110 is not shown. In this case, participant clients 121 communicate directly with the API 137 of the central server 130, for example, by providing video and / or audio stream data to the central server 130 and / or receiving video and / or audio stream data from the central server 130. The generated shared stream can then be provided to external consumers 150 and / or to one or more of the client participants 121.
[0062] Figure 4 shows the central server 130 in more detail. As shown, the collection function 131 may consist of one or preferably more format-specific collection functions 131a. Each of the format-specific collection functions 131a is configured to receive video and / or audio streams having a predetermined format such as a predetermined binary encoding format and / or a predetermined stream data container, and is specifically configured to parse the binary video and / or audio data of the above format and classify it into individual video frames, sequences of video frames and / or time slots.
[0063] The central server 130 further comprises an event detection function 132 configured to receive video and / or audio stream data, such as binary stream data, from the collection function 131, and to perform event detection on each of the multiple data streams received. The event detection function 132 may include an AI (artificial intelligence) component 132a for performing event detection. Event detection may be performed without first time-synchronizing the multiple individual streams that have been collected.
[0064] The central server 130 further includes a synchronization function 133 configured to synchronize multiple data streams, which may be provided by the collection function 131 and processed by the event detection function 132. The synchronization function 133 may include an AI component 133a for performing time synchronization.
[0065] The central server 130 may further include a pattern detection function 134 configured to perform pattern detection on a combination of at least one, but often at least two, e.g., at least three or at least four, e.g., all, of the multiple data streams received. Pattern detection may further be based on one, and possibly at least two or more, events detected for each individual of the multiple data streams by the event detection function 132. Such detected events considered by the pattern detection function 134 may be distributed over time with respect to the individual collected streams. The pattern detection function 134 may further include an AI component 134a for performing pattern detection. Pattern detection may further be based on the groupings described above, and in particular may be configured to detect a particular pattern occurring with respect to only one group, or a particular pattern occurring with respect to some groups but not all groups, or a particular pattern occurring with respect to all groups.
[0066] The central server 130 further includes a generation function 135 configured to generate generated digital video streams, such as a shared digital video stream, based on multiple data streams provided by the collection function 131, and possibly on any detected events and / or patterns. The generated video streams include at least one or more raw, reformatted, or transformed video streams provided by the collection function 131, and may include corresponding audio stream data. There may be multiple generated video streams, as illustrated below, one of which may be generated in the manner described above, but may also be generated based on another already generated video stream.
[0067] All generated video streams are preferably generated sequentially and preferably in near real-time (after subtracting latency and delays of the types described herein).
[0068] The central server 130 may also include a publishing function 136 configured to publish the generated shared digital video stream, for example, via the API 137 described above.
[0069] Figures 1, 2, and 3 show three different examples of how the central server 130 can be used to implement the principles described herein, and in particular to provide the method according to the present invention. However, it should be noted that other configurations are also possible, with or without one or more video communication services 110.
[0070] Therefore, Figure 5 illustrates a method for providing the generated digital video stream. Figures 6a to 6f show different states of digital video / audio data streams resulting from the method steps shown in Figure 5.
[0071] The first step is where this method begins.
[0072] In the subsequent collection step, each of the multiple primary digital video streams 210, 301 is collected from at least two of the digital video sources 120, 300 by, for example, a collection function 131. Each of these multiple primary data streams 210, 301 may comprise an audio portion 214 and / or a video portion 215. In this context, “video” is understood to refer to the video and / or still image content of such data streams. Each primary data stream 210, 301 may be encoded according to any video / audio encoding standard (using the respective codecs used by the entity providing the primary stream 210, 301), and the encoding format may differ among different primary streams 210, 301 used simultaneously in the same video communication. At least one, for example all, of the multiple primary data streams 210, 301 are preferably provided as a stream of binary data, and in some cases are provided themselves in a conventional data container data structure. Preferably, at least one, for example, at least two, or all of the multiple primary data streams 210, 301 are provided as their respective live video recordings.
[0073] It should be noted that the multiple primary data streams 210 and 301 may not be temporally synchronized when received by the collection function 131. This may mean that they are associated with different latency or delays relative to each other. For example, if the two primary video streams 210 and 301 are live recordings, this may mean that when received by the collection function 131, they are associated with different latency with respect to the recording time.
[0074] Furthermore, it should be noted that the multiple primary data streams 210, 301 themselves may be the respective live camera feeds from the webcams; the currently shared screen or presentation; the film clip being viewed; or any combination of these arranged in various ways within a single screen.
[0075] The collection steps are shown in Figures 6a and 6b. Figure 6b also shows how the collection function 131 can store each primary video stream 210, 301 as bundled audio / video information or as audio stream data separated from the associated video stream data. Figure 6b illustrates how the data of the primary video streams 210, 301 is stored as individual frames 213 or as agglomerates / clusters of frames, where “frame” here refers to a time-limited portion of image data and / or any associated audio data, for example, each frame may be an individual still image or a sequence of images (for example, a sequence of images that make up a video up to 1 second long) that together form a video content.
[0076] In a subsequent event detection step performed by the event detection function 132, multiple primary digital video streams 210, 301 are analyzed by the event detection function 132, particularly by the AI component 132a, to detect at least one event 211 selected from the first set of events. This is shown in Figure 6c.
[0077] This event detection step is performed on at least one, for example, at least two, or for example, all, primary video streams 210, 301, preferably individually for each of the primary video streams 210, 301. In other words, the event detection step is preferably performed on each of the primary video streams 210, 301, taking into account only the information contained as part of that particular primary video stream 210, 301, and in particular without considering the information contained as part of other primary video streams. Furthermore, event detection is preferably performed without considering any common time reference 260 associated with the multiple primary video streams 210, 301.
[0078] On the other hand, preferably, event detection considers information included as part of the individually analyzed primary video stream over a certain time interval, for example, over a historical time interval of the primary video stream, such as longer than 0 seconds, for example, at least 0.1 seconds, for example, at least 1 second.
[0079] Event detection may take into account information contained in audio and / or video data included as part of the primary video streams 210, 301.
[0080] The first set of events described above may include any number of types of events, such as changes in slides in a slide presentation that constitutes or is part of the primary video streams 210, 301; changes in connection quality of sources 120, 300 providing the primary video streams 210, 301, resulting in changes in image quality, loss of image data, or reacquisition of image data; and physical events of motion detected within the primary video streams 210, 301, such as movement of people or objects in the video, changes in lighting in the video, sudden sharp noises in the audio, or changes in audio quality. It should be understood that this is not intended to be an exhaustive list, and these examples are provided to illustrate the applicability of the principles described above.
[0081] In a subsequent synchronization step performed by the synchronization function 133, the multiple primary digital video streams 210 are time-synchronized. This time synchronization may be performed against a common time reference 260. As shown in Figure 6d, this time synchronization may include aligning the multiple primary video streams 210, 301 relative to one another, for example, using a common time reference 260, so that they can be combined to form a time-synchronized context. The common time reference 260 may be a stream of data, a heartbeat signal or other pulse data, or a time anchor applicable to each of the individual multiple primary video streams 210, 301. By making the common time reference applicable to each of the individual multiple primary video streams 210, 301, the information content of the primary video streams 210, 301 can be uniquely associated with the common time reference with respect to a common time axis. In other words, the common time reference aligns the multiple primary video streams 210, 301 so that they are time-synchronized in the present sense via time shift. In other embodiments, time synchronization may be based on known information regarding the time difference between the primary video streams 210, 301, such as measured values.
[0082] As shown in Figure 6d, time synchronization may include determining one or more timestamps 261 for each of the multiple primary video streams 210, 301, for example, in relation to a common time reference 260, or for each of the video streams 210, 301, in relation to the other video stream 210, 301 or other multiple video streams 210, 301.
[0083] In a subsequent pattern detection step performed by the pattern detection function 134, multiple time-synchronized primary digital video streams 210, 301 are analyzed to detect at least one pattern 212 selected from a first pattern set. This is shown in Figure 6e.
[0084] In contrast to the event detection step, the pattern detection step is preferably performed based on video and / or audio information included as part of at least two of a plurality of time-synchronized primary video streams 210, 301.
[0085] The first set of patterns described above may include any number of patterns, such as multiple participants speaking in turn or simultaneously, or a change in a presentation slide occurring simultaneously as another event, such as another participant speaking. This list is not exhaustive and is illustrative.
[0086] In an alternative embodiment, the detected pattern 212 may relate to information contained in only one of the multiple primary video streams 210, 301, rather than information contained in multiple of the multiple primary video streams 210, 301. In such a case, it is preferable that such pattern 212 is detected based on video and / or audio information contained in that single primary video stream 210, 301 that spans at least two detected events 211, for example, two or more consecutively detected presentation slide changes or connection quality changes. As an example, multiple consecutive slide changes that rapidly follow each other over time may be detected as a single slide change pattern, in contrast to one individual slide change pattern for each detected slide change event.
[0087] It is understood that the first event set and the first pattern set may comprise a given type of events / patterns defined using their respective sets of parameters and parameter intervals. As described below, the events / patterns in the above sets can also be defined and detected using various AI tools.
[0088] In a subsequent generation step performed by the generation function 135, the shared digital video stream is generated as an output digital video stream 230 based on a plurality of frames 213 that are sequentially considered from a plurality of time-synchronized primary digital video streams 210, 301 and a detected pattern 212.
[0089] As described and detailed below, the present invention makes it possible to completely automatically generate video streams, such as the output digital video stream 230.
[0090] For example, such generation may include selecting what video and / or audio information from which primary video streams 210, 301 to use and to what extent in the output video stream 230; the video screen layout of the output video stream 230; and patterns of switching between different such uses or layouts over time.
[0091] This is also shown in Figure 6f, which shows one or more additional portions of time-related (may be relative to the common time reference 260) digital video information 220, such as an additional digital video information stream, which can be used in conjunction with a plurality of time-synchronous primary video streams 210, 301 (for example, relative to the common time reference 260) in the generation of the output video stream 230. For example, the additional stream 220 may include information about any video and / or audio special effects to be used, such as dynamically based on detected patterns; a planned time schedule for video communication, etc.
[0092] In subsequent publishing steps performed by the publishing function 136, the generated output digital video stream 230 is successively provided to consumers 110, 150 of the shared digital video stream, as described above. The generated digital video stream may be provided to one or more participant clients 121, for example, via the video communication service 110.
[0093] In a subsequent step, the method terminates. However, initially, the method may be repeated any number of times to generate the output video stream 230 as a continuously provided stream, as shown in Figure 5. Preferably, the output video stream 230 is generated to be consumed in real time or near real time (taking into account the total latency added by all intermediate steps) and continuously (published as soon as more information becomes available, but not counting intentionally added latency as described below). In this way, the output video stream 230 may be consumed in a bidirectional (interactive) manner, thereby feeding the output video stream 230 back to the video communication service 110 or to other contexts that form the basis for generating the primary video stream 210 which is then supplied back to the collection function 131 to form a closed feedback loop; or the output video stream 230 may be consumed in a different context (outside system 100, or at least outside the central server 130) where it can form the basis for real-time bidirectional video communication.
[0094] As described above, in some embodiments, at least two, for example, at least three, for example, at least four or at least five of a plurality of primary digital video streams 210, 301 are provided as part of a shared digital video communication, such as provided by a video communication service 110, which includes each remotely connected participant client 121 providing the primary digital video stream 210. In such cases, the collection step may consist of collecting at least one of the primary digital video streams 210 from the shared digital video communication service 110 itself, via an automated participant client 140 which is sequentially granted access to video and / or audio stream data from within the video communication service 110, and / or via the API 112 of the video communication service 110.
[0095] Furthermore, in this case and other cases, the collection step may include collecting at least one of the plurality of primary digital video streams 210, 301 as each external digital video stream 301 collected from an information source 300 which is outside the shared digital video communication service 110. Note that one or more of these external video sources 300 may be outside the central server 130.
[0096] In some embodiments, the multiple primary video streams 210, 301 are not formatted in the same way. Such different formats may be the formats in which they are supplied to the acquisition function 131 in different types of data containers (such as AVI or MPEG), but in a preferred embodiment, at least one of the multiple primary video streams 210, 301 is formatted according to a deviating format (relative to at least one other of the primary video streams 210, 301), in that the deviating primary digital video stream 210, 301 has a deviating video encoding; a deviating fixed or variable frame rate; a deviating aspect ratio; a deviating video resolution; and / or a deviating audio sample rate.
[0097] The collection function 131 is preferably pre-configured to read and interpret all encoding formats, container standards, etc., occurring in all collected primary video streams 210, 301. This enables the execution of processing as described herein, without requiring decoding until relatively later stages of these processes (until the primary streams are placed in their respective buffers; after the event detection step; or after the event detection step, etc.). However, in the rare case where one or more of the primary video feeds 210, 301 are encoded using a codec that the collection function 131 cannot interpret without decoding, the collection function 131 may be configured to decode and analyze such primary video streams 210, 301 and then convert them to a format that can be processed by, for example, the event detection function. Note that even in this case, it is preferable not to perform re-encoding at this stage.
[0098] For example, a primary video stream 220 fetched from a multi-party video event, such as one provided by a video communication service 110, typically has a requirement for low latency and is therefore typically associated with a variable frame rate and variable pixel resolution to enable participants 122 to communicate effectively. In other words, the overall video and audio quality is reduced as necessary for low latency.
[0099] On the other hand, the external video feed 301 typically has a more stable frame rate and higher image quality, but therefore may have greater latency.
[0100] Therefore, the video communication service 110 may use a different encoding and / or container than the external video source 300 at each point in time. Thus, the analysis and video generation processes described herein need to combine these multiple streams 210, 301 of different formats into a new single stream for a combined experience.
[0101] As described above, the collection function 131 may comprise a set of format-specific collection functions 131a, each configured to process primary video streams 210, 301 of a particular type of format. For example, each of these format-specific collection functions 131a may be configured to process multiple primary video streams 210, 301 encoded using different video encoding methods / codecs, such as Windows® Media® or DivX®.
[0102] However, in a preferred embodiment, the acquisition step includes converting at least two, for example all, of a plurality of primary digital video streams 210, 301 to a common protocol 240.
[0103] As used in this context, the term “protocol” refers to an information structuring standard or data structure that specifies how the information contained in a digital video / audio stream is stored. However, a common protocol preferably does not specify how the digital video and / or audio information is stored, for example, at the binary level (i.e., encoded / compressed data that indicates the sound and image itself), but rather forms a structure of a predetermined format for storing such data. In other words, a common protocol specifies that digital video data is stored in the raw binary format without performing any digital video decoding or digital video encoding in connection with such storage, and without modifying the existing binary format at all, apart from possibly concatenating and / or splitting byte sequences in binary format. Instead, the (encoded / compressed) binary data content of the raw data of the primary video streams 210, 301 is preserved while repacking this raw binary data in a data structure defined by the protocol. In some embodiments, the common protocol defines a video file container format.
[0104] Figure 7 shows, as an example, multiple primary video streams 210, 301 shown in Figure 6a, which are reconstructed by their respective format-specific acquisition functions 131a and use the common protocol 240 described above.
[0105] Therefore, the common protocol 240 specifies that digital video and / or audio data be stored in datasets 241, which are preferably divided into discrete and continuous sets of data along the time axis associated with the primary video streams 210, 301. Each such dataset may contain one or more video frames and associated audio data.
[0106] The common protocol 240 may also specify that metadata 242 associated with a specified point in time be stored in relation to the stored digital video and / or audio dataset 241.
[0107] The metadata 242 may include information about the binary format of the raw data of the primary digital video stream 210, such as the digital video encoding method or codec used to generate the binary data of the raw data; the resolution of the video data; the video frame rate; a frame rate variation flag; the video aspect ratio; an audio compression algorithm; or an audio sampling rate. The metadata 242 may also include information about the timestamp of the stored data, for example, relating to the time reference of the primary video stream 210, 301 itself, or relating to the different video streams mentioned above.
[0108] By using the format-specific collection function 131a in combination with the common protocol 240, it becomes possible to quickly collect the information content of the primary video streams 210 and 301 without adding waiting time (delay) due to decoding / re-encoding the received video / audio data.
[0109] Therefore, the collection step may include collecting multiple primary digital video streams 210, 301 encoded using different binary video and / or audio encoding formats, using different collection functions 131a from a plurality of format-specific collection functions 131a, in order to analyze the primary video streams 210, 301 and store the analyzed raw binary data, along with any associated metadata, in a data structure using a common protocol. Obviously, the decision of which format-specific collection function 131a to use for which primary video streams 210, 301 may be made by the collection function 131 based on predetermined and / or dynamically detected characteristics of each primary video stream 210, 301.
[0110] Each of the primary video streams 210, 301 collected in this manner may be stored in its own separate memory buffer, such as a RAM memory buffer in the central server 130.
[0111] The conversion of primary video streams 210, 301 performed by each format-specific acquisition function 131a may therefore involve dividing the binary data of the raw data of each thus converted primary digital video stream 210, 301 into smaller datasets 241 of ordered sets.
[0112] Furthermore, the transformation may also include associating each of the smaller sets 241 (or subsets, e.g., subsets regularly distributed along the respective time axes of the primary streams 210, 301) with their respective times along a shared time axis, for example, in relation to a common time reference 260. This association may be performed by analysis of the raw binary video and / or audio data, either by one of the principle methods described below or otherwise, in order to enable subsequent time synchronization of the primary video streams 210, 301. Depending on the type of common time reference 260 used, at least a portion of this association for each dataset 241 may also be performed by the synchronization function 133, or instead. In the latter case, the collection step may instead include associating each of the smaller sets 241 or a subset thereof with their respective times on the time axes specific to the primary streams 210, 301.
[0113] In some embodiments, the acquisition step also includes converting the raw binary video and / or audio data collected from multiple primary video streams 210, 301 to a uniform quality and / or updating the frequencies. This may include, if necessary, downsampling or upsampling the raw binary digital video and / or audio data of multiple primary digital video streams 210, 301 to a common video frame rate; a common video resolution; or a common audio sampling rate. It should be noted that such resampling can be performed without performing full decoding / re-encoding, or even without performing decoding at all, because the format-specific acquisition function 131a can directly process the raw binary data according to the correct binary encoding target format.
[0114] Preferably, each of the multiple primary digital video streams 210, 301 is stored in an individual data storage buffer 250 as an individual frame 213 or a sequence of frames 213, as described above, and each is associated with a corresponding timestamp sequentially associated with a common time reference 260.
[0115] In the example provided for illustrative purposes, the video communication service 110 is Microsoft Teams® and is running a video conference involving multiple participants 122 with simultaneous connections. The automated participant client 140 is registered as a meeting participant in the Teams® meeting.
[0116] Next, the primary video input signal 210 is provided to the acquisition function 130 via the automatic participant client 140 and acquired by the acquisition function 130. These are raw data signals in H264 format and include timestamp information for each video frame.
[0117] The associated format-specific collection function 131a picks up raw data via IP (the cloud's LAN network) on a configurable, predefined TCP port. All Teams® meeting participants and associated audio data are associated with separate ports. The collection function 131 then uses a timestamp from the audio signal (50Hz) to downsample the video data to a fixed output signal of 25Hz, and then stores the video stream 220 into its respective separate buffer 250.
[0118] As mentioned above, Common Protocol 240 stores data in raw binary format. This is at a very low level and can be designed to handle the bits and bytes of raw video / audio data. In a preferred embodiment, data is stored in Common Protocol 240 as a simple byte array or corresponding data structure (such as a slice). This means that the data does not need to be placed in a conventional video container at all (Common Protocol 240 does not constitute such a conventional container in this context). Furthermore, encoding and decoding video is computationally intensive, causing delays and requiring expensive hardware. Moreover, this problem scales with the number of participants.
[0119] The Common Protocol 240 allows for the allocation of memory within the acquisition function 131 for the primary video stream 210 associated with each Teams® meeting participant 122 and for any external video source 300, and enables on-the-fly modification of the allocated memory amount during the process. In this way, the number of input streams can be changed, and as a result, each buffer can be kept active. For example, information such as resolution and frame rate is variable, but since it is stored as metadata in the Common Protocol 240, this information can be used to quickly change the size of each buffer as needed.
[0120] The following is an example of the specifications for this type of common protocol 240.
[0121] [Table 1]
[0122] In the table above, the "Detected event in, if any" data is included as part of the Common Protocol 260 specification. However, in some embodiments, this information (regarding detected events) may instead be stored in a separate memory buffer.
[0123] In some embodiments, the above-mentioned at least one additional portion of the digital video information 220, which may be an overlay or effect, is also stored in its respective individual buffer 250 as individual frames or sequences of frames, each associated with a corresponding timestamp sequentially associated with a common time reference 260.
[0124] As illustrated above, the event detection step may include using a common protocol 240 to store metadata 242 describing the detected event 211, associated with the primary digital video streams 210, 301 in which the event 211 was detected.
[0125] Event detection can be performed in different ways. In some embodiments performed by the AI component 132a, the event detection step includes a first trained neural network or other machine learning component individually analyzing at least one, for example, some or all of a plurality of primary digital video streams 210, 301 to automatically detect any of the events 211. This may include the AI component 132a classifying the data of the primary video streams 210, 301 into a predefined set of events in a managed classification, and / or into a dynamically determined set of events in an unmanaged classification.
[0126] In some embodiments, the detected event 211 is the primary video stream 210, 301, or a change in the presentation slides of the presentation contained therein.
[0127] For example, if a presenter decides to change the slides they are currently showing to the audience, this means that what is interesting to a given audience may change. The newly displayed slide might be just a full-level image that is best viewed briefly in so-called "butterfly" mode (for example, displaying the slide alongside the presenter's video in output video stream 230). Or, the slide might contain a lot of detail and small font size text. In the latter case, the slide would be displayed in full screen and would be shown for a slightly longer time than usual. Butterfly mode might not be so appropriate in this case, as the slide might be of more interest to the viewers than the presenter's face.
[0128] In practice, the event detection step consists of at least one of the following:
[0129] Firstly, event 211 can be detected based on image analysis of the difference between the first image of the detected slide and the second image of the subsequently detected slide. The nature of the primary video streams 220, 301 to represent slides can itself be automatically determined using conventional digital image processing, such as using motion detection combined with OCR (optical character recognition).
[0130] This may involve using automated computer image processing techniques to check whether the detected slide has changed significantly enough to be classified as a slide change. This can be done by checking the delta between the current slide and the previous slide with respect to RGB color values. For example, it is possible to assess how globally the RGB values have changed in the screen area covered by the slide in question, and at the same time, whether it is possible to find groups of adjacent pixels that have changed in coordination with this. In this way, relevant slide changes can be detected, while irrelevant changes, such as computer mouse movements across the entire screen, can be filtered out. Full configurability is achieved with this approach. For example, it may be desirable to be able to capture computer mouse movements, for instance, if a presenter wants to explain something in detail while pointing to different things with the computer mouse.
[0131] Secondly, event 211 may be detected based on image analysis of the information complexity of the second image itself in order to determine the type of event with higher specificity.
[0132] This might involve, for example, evaluating the total amount of text information on the slide in question and the associated font size. This can be done using conventional OCR methods, such as deep learning-based character recognition technology.
[0133] It should be noted that, since the binary format of the raw data of the evaluated video streams 210, 301 is known, this may be performed directly in the binary domain without first decoding or re-encoding the video data. For example, the event detection function 132 may invoke relevant format-specific collection functions for the image interpretation service, or the event detection function 132 itself may include functions for evaluating image information down to the individual pixel level, etc., for a number of different supported raw data binary video data formats.
[0134] In another example, the detected event 211 is the loss of communication connectivity of a participant client 121 to the digital video communication service 110. In this case, the detection step may include detecting that the participant client 121 has lost communication connectivity based on image analysis of a series of subsequent video frames 213 of the primary digital video stream 210 corresponding to the participant client 121.
[0135] Since participant clients 121 are associated with different physical locations and different internet connections, it is possible that someone may lose their connection to the video communication service 110 or the central server 130. In such situations, it is desirable that the generated output video stream 230 does not display a black or blank screen.
[0136] Alternatively, such loss of connection can be detected as an event by the event detection function 132, for example, by applying a two-class classification algorithm where the two classes used are connected / disconnected (no data). In this case, "no data" is understood to be different from the presenter intentionally sending a black screen. Since a short black screen, such as just one or two frames, may not be noticeable in the final generated stream 230, the two-class classification algorithm can be applied over time to create a time series. Then, a threshold specifying the minimum length of connection interruption can be used to determine whether the connection has been lost.
[0137] As described below, the types of detected events exemplified above can be used by the pattern detection function 134 to take various appropriate and desired actions.
[0138] As described above, each primary video stream 210, 301 is related to a common time reference 260, and the synchronization function 133 can synchronize their time relative to each other.
[0139] In some embodiments, the common time reference 260 is based on or comprises a common audio signal 111 (see Figures 1 to 3), the common audio signal 111 being common to a shared digital video communication service 110 in which at least two remotely connected participant clients 121 participate, each providing one of the primary digital video streams 210.
[0140] In the Microsoft® Teams® example described above, a common audio signal is generated and can be captured by the central server 130 via the automated participant client 140 and / or via API 112. In this example and others, such a common audio signal can be used as a heartbeat signal to time-synchronize individual primary video streams 220 by combining them at specific points in time based on this heartbeat signal. Such a common audio signal may also be provided as a separate signal (in relation to each of the other primary video streams 210), so that each of the other primary video streams 210 may be individually time-correlated to the common audio signal based on the audio contained in the other primary video stream 210 or based on the image information contained therein (such as using an automated image processing-based lip-sync technique).
[0141] In other words, in order to handle the variable and / or different latency associated with each individual primary video stream 210 and achieve time synchronization of the combined video output stream 230, such a common audio signal is used as a heartbeat for all primary video streams 210 in the central server 130 (but possibly not for the external primary video stream 301). In other words, all other signals are mapped to this common audio time heartbeat to ensure that everything is time-synchronized.
[0142] In another example, time synchronization is achieved using a time synchronization element 231 introduced into the output digital video stream 230 and detected by each local time synchronization software function 125 provided as part of one or more individual participant clients 121. The local software function 125 is configured to detect the time of arrival of the time synchronization element 231 in the output video stream 230. As understood, in such an embodiment, the output video stream 230 is either fed back to the video communication service 110 or, otherwise, made available to each participant client 121 and the local software function 125.
[0143] For example, the time synchronization element 231 may be a visual marker, such as a pixel whose color changes in a predetermined order or manner, which is placed or updated on the output video 230 at regular time intervals; a visual clock which is updated and displayed on the output video 230; or an audio signal (which may be designed not to be heard by the participant 122, for example, by having a sufficiently low amplitude and / or a sufficiently high frequency) which is added to the audio that forms part of the output video stream 230. The local software function 125 is configured to automatically detect the arrival time of each of the time synchronization elements 231 using appropriate image processing and / or audio processing.
[0144] Next, the common time standard 260 may be determined, at least in part, based on the detected arrival time. For example, each of the local software functions 125 may communicate its respective information indicating the detected arrival time to the central server 130.
[0145] Such communication may take place via a direct communication link between the participant client 121 and the central server 130. However, the communication may also take place via a primary video stream 210 associated with the participant client 121. For example, the participant client 121 may introduce visual or audible code of the type described above into the primary video stream 210 generated by the participant client 121 for automatic detection by the central server 130, and use this to determine the common time reference 260.
[0146] In a further additional embodiment, each participant client 121 may perform image detection on a common video stream viewable by all participant clients 121 to the video communication service 110, relaying the results of such image detection to the central server 130 in a manner corresponding to that described above, where they may be used over time to determine the respective offsets of each participant client 121 relative to one another. In this way, the common time reference 260 may be determined as a set of individual relative offsets. For example, a selected reference pixel of the commonly available video stream may be monitored by some or all participant clients 121 by a local software function 125, etc., and the current color of that pixel may be transmitted to the central server 130. The central server 130 may generate an estimated set of relative time offsets across different participant clients 121 by calculating each time series based on such color values received sequentially from each of a number of (or all) participant clients 121 and performing cross-correlation.
[0147] In practice, the output video stream 230 supplied to the video communication service 110 may be included as part of the shared screen of all participant clients of the video communication and may therefore be used to evaluate such time offsets associated with participant client 121. In particular, the output video stream 230 supplied to the video communication service 110 may be made available again to the central server via the automated participant client 140 and / or API 112.
[0148] In some embodiments, the common time criterion 260 may be determined at least in part on the detected discrepancy between the audio portion 214 of a first of a plurality of primary digital video streams 210, 301 and the image portion 215 of the first of the plurality of primary digital video streams 210, 301. Such discrepancies may be based, for example, on digital lip-sync video image analysis of participant 122 during speech viewed in the first primary digital video stream 210, 301. Such lip-sync analysis is conventional in itself and may, for example, use a trained neural network. The analysis may be performed by a synchronization function 133 for each primary video stream 210, 301 in relation to available common audio information, and the relative offset across the individual primary video streams 210, 301 may be determined based on this information.
[0149] In some embodiments, the synchronization step includes intentionally introducing a delay of up to 30 seconds, e.g., up to 5 seconds, e.g., up to 1 second, e.g., up to 0.5 seconds, but longer than 0 seconds (in this context, “delay” and “wait time” are intended to mean the same thing), thereby providing the output digital video stream 230 with at least that delay. Whatever the length, the intentionally introduced delay is at least several video frames, such as at least three, or at least five, or even ten, of which this number of frames (or individual images) are stored after any resampling in the acquisition step. As used herein, the term “intentionally” means that the delay is introduced regardless of the need to introduce such a delay based on synchronization issues or the like. In other words, the intentionally introduced delay is introduced in addition to the delay introduced as part of the synchronization of the primary video streams 210, 301 in order to synchronize the time between the multiple primary video streams 210, 301. The intentionally introduced delay may be predetermined, fixed, or variable in relation to the common time reference 260. The latency may be measured in relation to the least latent of the multiple primary video streams 210, 301, and as a result of the above time synchronization, more latent streams 210, 301 may be associated with a relatively small intentionally added latency.
[0150] In some embodiments, a relatively small delay of 0.5 seconds or less is introduced. This delay is barely noticeable to participants in the video communication service 110 using the output video stream 230. In other embodiments, a larger delay may be introduced, for example, when the output video stream 230 is not used in an interactive context and is instead exposed in one-way communication to an external consumer 150.
[0151] This intentionally introduced delay may be sufficient to allow the synchronization function 133 enough time to map the collected individual primary streams 210, 301 video frames to the correct common time reference 260 timestamps 261. It may also be sufficient to allow enough time to perform the event detection described above to detect lost primary streams 210, 301 signals, slide changes, resolution changes, etc. Furthermore, the intentional introduction of the delay may be sufficient to improve the pattern detection function 134, as described below.
[0152] The introduction of delay is understood to involve buffering each of the multiple collected and time-synchronized primary video streams 210, 301 before publishing the output video stream 230 using the buffered frame 213. In other words, at least one, some, or all of the video and / or audio data from the multiple primary video streams 210, 301 may reside in the central server 130 in a buffered manner, particularly as used by the pattern detection function 134, for the reasons stated above, rather than being used like a cache, but not intended to be able to handle situations where bandwidth changes (like a conventional cache buffer).
[0153] Therefore, in some embodiments, the pattern detection step includes considering specific information of at least one, for example, several, for example, at least four, or all of a plurality of primary digital video streams 210, 301, where this specific information resides in a frame 213 later than a frame of the time-synchronized primary digital video stream 210 that has not yet been used in the generation of the output digital video stream 230. Thus, the newly added frame 213 resides in the buffer 250 for a certain waiting period before forming part (or the basis) of the output video stream 230. During this period, the information of the frame 213 constitutes "future" information in relation to the frame currently being used to generate the current frame of the output video stream 230. When the timeline of the output video stream 230 reaches the frame 213, the frame is used to generate the corresponding frame of the output video stream 230 and may thereafter be discarded.
[0154] In other words, the pattern detection function 134 has free access to a set of video / audio frames 213 that have not yet been used to generate the output video stream 230, and uses this data to detect the pattern described above.
[0155] Pattern detection can be performed in different ways. In some embodiments performed by the AI component 134a, the pattern detection step includes a second trained neural network or other machine learning component in concert analyzing at least two, e.g., at least three, e.g., at least four, or all of a plurality of primary digital video streams 120, 301 to automatically detect the pattern 212.
[0156] In some embodiments, the detected pattern 212 comprises an utterance pattern to the shared video communication service 110, each associated with its respective participant client 121, including at least two, e.g., at least three, e.g., at least four different utterance participants 122, each of which is visually viewed in one of the multiple primary digital video streams 210, 301.
[0157] Preferably, the generation step includes determining, tracking, and updating the current generation state of the output video stream 230. For example, such a state can determine which participants 122 (if any) are visible in the output video stream 230 and where they are visible on the screen; which external video streams 300 are visible in the output video stream 230 and where they are visible on the screen; whether any slides or shared screens are displayed in full-screen mode or in combination with any live video streams, and so on. Thus, the generation function 135 can be viewed as a state machine with respect to the generated output video stream 230.
[0158] To generate the output video stream 230 as a combined video experience viewed, for example, by the end consumer 150, it is advantageous for the central server 130 to be able to understand what is happening at a deeper level than simply detecting the individual events associated with the individual primary video streams 210, 301.
[0159] In the first example, participant client 121, who is giving a presentation, changes the currently displayed slide. This slide change is detected by the event detection function 132 as described above, and metadata 242 indicating that a slide change has occurred is added to the frame. This can happen multiple times, as it is found that participant client 121 is rapidly skipping forward through a number of slides, and as a result, a series of “slide change” events are also detected by the detection function 132 and stored along with metadata 242 corresponding to the individual buffers 250 of the primary video stream 210. In practice, each slide that is rapidly skipped forward may only be displayed for a few seconds.
[0160] The pattern detection function 134 refers to the information in the buffer 250 that spans multiple of these detected slide changes and detects a pattern that corresponds to a single slide change rather than a large number of slide changes that are performed rapidly (i.e., a single slide change to the last slide in a forward skip, which remains visible once the rapid skip ends). In other words, the pattern detection function 134 notices, for example, that there were 10 slide changes in a very short time and why they are treated as a detected pattern that means a single slide change. As a result, the generation function 135 has access to the pattern detected by the pattern detection function 134 and determines that this last slide is potentially important in the state machine, so it can choose to display that last slide in fullscreen mode for a few seconds in the output video stream 230. Alternatively, it can choose not to display any intermediately viewed slides at all in the output stream 230.
[0161] The detection of patterns with multiple rapid slide changes may be performed by a simple rule-based algorithm, or alternatively, by using a neural network designed and trained to detect such patterns in video by classification.
[0162] In another example, it may be useful when the video communication is a talk show, panel discussion, or similar, where it may be desirable to quickly switch visual attention between current speakers, while on the one hand providing a relevant viewing experience to consumers 150 by generating and publishing a smooth output video stream 230. In this case, the event detection function 132 can continuously analyze each primary video stream 210, 301 to always determine whether the person being viewed in that particular primary video stream 210, 301 is currently speaking. This can be done, for example, using conventional image processing tools themselves, as described above. Next, the pattern detection function 134 may be capable of detecting certain overall patterns, including multiple primary video streams 210, 301, which are useful for generating a smooth output video stream 230. For example, the pattern detection function 134 can detect patterns of very frequent switching between current speakers and / or patterns including multiple simultaneous speakers.
[0163] Next, the generation function 135 can take such detected patterns into consideration when making automated decisions related to the generated state, such as not automatically switching the visual focus to a speaker who speaks for only half a second and then becomes silent again, or switching to a state where multiple speakers are displayed side by side during a certain period of time when they are speaking alternately or simultaneously. This state determination process can itself be performed using time-series pattern recognition techniques or a trained neural network, but it can also be based at least in part on a predetermined set of rules.
[0164] In some embodiments, there may be multiple patterns detected in parallel and forming input to the state machine of the generation function 135. Such multiple patterns can be used by the generation function 135 for different AI components, computer vision detection algorithms, etc. For example, it may be possible to detect permanent slide changes while simultaneously detecting unstable connections of some participant clients 121, and other patterns to detect participant 122 in the current main utterance. Using all such available pattern data, a classifier neural network can be trained and / or a set of rules can be developed to analyze the time series of such pattern data. Such classifications can be supervised, at least partially, for example, entirely, to result in determined desired state changes used in the generation described above. For example, different such given classifiers can be generated, specifically configured to automatically generate output video streams 230 according to various different generation styles and requirements. Training can be performed based on known generation state change sequences as desired outputs and known pattern time series data as training data. In some embodiments, a Bayesian model can be used to generate such classifiers. In a concrete example, we can obtain a priori information from an experienced producer and provide input such as, "In a talk show, we don't switch directly from speaker A to speaker B, but we always give an overview before focusing on other speakers, unless another speaker is very dominant and speaking loudly." This generative logic can be expressed as a general form of Bayesian model: "If X is true, then | given the fact that Y is true, then | do Z." Actual detection (such as whether someone is speaking loudly) can be done using classifiers or threshold-based rules.
[0165] With a large dataset (of patterned time-series data), deep learning techniques can be used to develop a correct and attractive generation format for use in the automatic generation of video streams.
[0166] In summary, by using a combination of event detection based on multiple individual primary video streams 210, 301; intentionally introduced delays; pattern detection based on multiple time-synchronized primary video streams 210, 301 and detected events; and a generation process based on detected patterns, it becomes possible to achieve automatic generation of output digital video streams 230 according to a wide range of possible tastes and styles. This result is valid across a wide range of possible neural network and / or rule-based analysis techniques used by the event detection function 132, pattern detection function 134, and generation function 135. This is particularly valid in the embodiments described below, which feature that a first generated video stream is used for the automatic generation of a second generated video stream, and that different intentionally added delays are used for different groups of participant clients. This is also particularly valid in later embodiments, where detected triggers result in a switch of which video stream is used in the generated output video stream, or result in automatic cropping or zooming of the video stream used in the output video stream.
[0167] As illustrated above, the generation step may include generating the output digital video stream 230 based on a predetermined and / or dynamically variable set of parameters relating to the visibility of individual primary digital video streams 210, 301 in the output digital video stream 230; the arrangement of visual and / or auditory video content; the visual or auditory effects used; and / or the output mode of the output digital video stream 230. Such parameters may be automatically determined by the state machine of the generation function 135 and / or set (semi-automatically) by an operator controlling the generation, and / or predetermined based on some prior configuration requirement (such as the shortest time between layout changes of the output video stream 230 or state changes of the type illustrated above).
[0168] In practice, the state machine may support a set of predetermined standard layouts that can be applied to the output video stream 230, such as a full-screen presenter view (showing the currently speaking participant 122 in full screen); a slide view (showing the currently shared presentation slide in full screen); a "butterfly view" (showing both the currently speaking participant 122 and the currently shared presentation slide side-by-side); and a multi-speaker view (showing all or a selected subset of participants 122 side-by-side or in a matrix layout). The various available generation formats can be defined by a set of state machine state change rules (such as those illustrated above), along with a set of available states (such as the set of standard layouts above). For example, one such generation format might be a "panel discussion," another a "presentation," and so on. By selecting a specific generation format via a GUI or other interface to the central server 130, the operator of system 100 can quickly select one of a predefined set of such generation formats, and then the central server 130 can automatically generate the output video stream 230 according to that generation format based on the available information as described above.
[0169] Furthermore, during generation, as described above, a memory buffer is created and maintained for each conference participant client 121 or external video source 300. These buffers can be easily deleted, added, and modified on the fly. The central server 130 may then be configured to receive information during the generation of the output video stream 230, such as information about participant clients 121 that have been added / dropped off and participants 122 scheduled to give speeches; scheduled or unexpected pauses / resumes of presentations; and desired changes to the currently used generation format. Such information may be supplied to the central server 130, for example, via an operator GUI or interface, as described above.
[0170] As illustrated above, in some embodiments, at least one of a plurality of primary digital video streams 210, 301 is provided to a digital video communication service 110, and the publishing step may then include providing an output digital video stream 230 to the same communication service 110. For example, the output video stream 230 may be provided to participant clients 121 of the video communication service 110, or it may be provided as an external video stream to the video communication service 110 via an API 112. In this way, the output video stream 230 can be made available to several or all participants of a video communication event currently being implemented by the video communication service 110.
[0171] As mentioned above, the output video stream 230 may be provided to one or more external consumers 150, either additionally or alternatively.
[0172] Generally, the generation step is performed by the central server 130, and the output digital video stream 230 can be provided as a live video stream to one or more simultaneous consumers via the API 137.
[0173] As described above, participant clients 121 can be organized into groups of two or more participant clients 121. Figure 8 is a simplified diagram of system 100 in a configuration that performs automatic generation of output video streams when such groups exist.
[0174] In Figure 8, the central server 130 is equipped with the collection function 131 described above.
[0175] The central server 130 also includes a first generation function 135', a second generation function 135'', and a third generation function 135'''. Each of these generation functions 135', 135'', and 135''' corresponds to generation function 135, and what has been described above in relation to generation function 135 also applies to generation functions 135', 135'', and 135'''. Depending on the detailed configuration of the central server 130, generation functions 135', 135'', and 135''' may be separate, multiple functions may be shared in a single logical function, and there may be more than three generation functions. Generation functions 135', 135'', and 135''' may, in some cases, be different functional aspects of the same generation function 135. Various communications between generation functions 135', 135'', and 135''' and other entities may be conducted via appropriate APIs.
[0176] Furthermore, a separate collection function 131 may exist for each of the generation functions 135', 135'', 135'''' or groups of such generation functions, and depending on the detailed configuration, there may be multiple logically separated central servers 130, each having its own collection function 131.
[0177] Furthermore, the central server 130 includes a first public function 136', a second public function 136'', and a third public function 136'''. Each of these public functions 136', 136'', and 136''' corresponds to public function 136, and what has been described above in relation to public function 136 also applies to public functions 136', 136'', and 136'''. Depending on the detailed configuration of the central server 130, public functions 136', 136'', and 136''' may be separate functions, may be shared and located in a single logical function with multiple functions, and there may be more than three public functions. Public functions 136', 136'', and 136''' may, in some cases, be different functional aspects of the same public function 136.
[0178] In Figure 8, for illustrative purposes, three sets or groups of participant clients are shown, each corresponding to the participant client 121 described above. Thus, there exists a first group 121' of such participant clients 121, a second group 121'' of such participant clients, and a third group 121'''' of such participant clients. Each of these groups may consist of one or preferably at least two participant clients. Such groups may consist of only two or three or more, depending on the detailed configuration. The assignments between groups 121', 121'', and 121'''' may be exclusive, in the sense that each participant client 121 is assigned to at most one group 121', 121'', or 121''''. In an alternative configuration, at least one participant client 121 may be assigned to multiple such groups 121', 121'', or 121'''' simultaneously.
[0179] Figure 8 also shows external consumers 150, and as mentioned above, it is understood that there may be multiple such external consumers 150.
[0180] Although Figure 8 does not show the video communication service 110 for the purpose of simplification, it is understood that the general type of video communication service described above may be used with the central server 130, and for example, the central server 130 may be used to provide a shared video communication service to each participant client 121 in the manner described above.
[0181] Each primary video stream may be collected by the collection function 131 from each participant client 121, such as the participant clients of the groups 121', 121'', and 121''' described above. Based on the provided primary video streams, the generation functions 135', 135'', and 135''' may generate their respective digital video output streams.
[0182] As shown in Figure 8, one or more such generated output streams are supplied as respective input digital video streams from one or more respective generating functions 135', 135''' to another generating function 135'', which may, in turn, generate a secondary digital output video stream for publication by the publishing function 136''. This secondary digital output video stream is thus generated based not only on one or more input primary digital video streams, but also on one or more pre-generated digital input digital video streams.
[0183] Two or more different generation steps 135', 135'', 135'''' may include the introduction of a time delay. In some embodiments, one or more of the generated output digital video streams from these generation steps 135', 135'', 135'''' may not be time-synchronized with any other video streams that may be supplied to other participant clients in the publishing step, due to the introduction of the time delay. Such a time delay may be intentionally added in any of the methods described herein, and / or may be a direct result of the generation of the generated digital video stream. As a result, any participant client consuming the time-asynchronous generated output digital video stream will consume it in a “time zone” that is slightly (time-) offset with the video stream consumption “time zone” of the other participant clients.
[0184] For example, one of the participant client groups 121', 121'', and 121'''' may consume its generated video stream in a first such "time zone," while another participant client 121 of the same group 121', 121'', and 121'''' may consume its generated video stream in a second such "time zone." Since both of these generated video streams may be generated based on at least partially the same primary video stream, all such participant clients 121 are active in the same video communication but are in different "time zones" relative to one another. In other words, each timeline for consuming the generated video stream may be temporally offset between different groups 121', 121'', and 121''''.
[0185] For example, some generation steps (135', 135''', etc.) may be direct (without using intentionally introduced time delays) and / or may involve only computationally relatively lightweight processing before supply for publication, while other generation steps (135'', etc.) may involve intentionally introduced time delays and / or relatively heavy processing, which would result in the generated digital video stream being generated for the earliest publication with a delay related to the earliest publication delay of each digital video stream in the former generation steps 135', 135'''.
[0186] Therefore, each participant client 121 in one or more of the above groups 121', 121'', and 121'''' may be able to interact with each other with the same perceived time delay. At the same time, groups associated with larger time delays can use the generated video streams from groups with smaller time delays as input video streams when the groups with larger time delays generate output video streams that will be viewed in a later "time zone" for that group.
[0187] The result of this first greater time delay generation (in step 135'') is a generated digital video stream of the type described above, which may visually include, for example, one or more of the primary video streams as subparts, in processed or unprocessed form. The generated video stream may include live-captured video streams, slides, externally supplied video or images, etc., as generally described above in relation to the video output stream generated by the central server 130. The generated video stream may also be generated based on detected events and / or patterns of an intentionally delayed or real-time input primary video stream supplied by the participant client 121, in the general manner described above.
[0188] In an exemplary embodiment, participant clients of the first group 121' are part of the discussion panel and communicate using the video communication service 110 with relatively low latency, and each of these participant clients is continuously supplied with the generated video stream (or each other's primary video stream) from the publication step 136'. The audience of the discussion panel consists of participant clients of the second group 121'' and is continuously supplied with the generated video stream from the generation step 135'', here associated with slightly higher latency. The generated video stream from the generation step 135'' can be automatically generated in the general manner described above, so as to automatically shift between views of individual discussion panel speakers (participant clients assigned to the first group 121', such views are supplied directly from the collection function 131) and a generated view showing all discussion panel speakers (this view is the first generated video stream). Thus, the audience can enjoy a well-staged experience while the panel speakers can interact with each other with minimal latency.
[0189] The intentional delay added to each primary video stream used in generation step 136'' may be at least 0.1 seconds, e.g., at least 0.2 seconds, e.g., at least 0.5 seconds, and at most 5 seconds, e.g., at most 2 seconds, e.g., at most 1 second. The intentionally added delay may also depend on an inheritance latency associated with each primary video stream to achieve perfect time synchronization between each primary video stream used and the generated video stream input from generation step 135' to generation step 135''.
[0190] It is understood that, as with the generated video streams from generation step 135', all such primary video streams may be additionally and intentionally delayed to improve pattern detection for use in the second generation function 135'' in the general manner described above.
[0191] Figure 8 further illustrates numerous alternative or simultaneous methods for publishing various generated video streams produced by the central server 130.
[0192] In general, in a publishing step performed by a first publishing function 136' configured to receive a first generated video stream from a first generating function 135', the first generated video stream may be continuously supplied to at least one of a first participant client 121 and a second participant client 121. For example, this first participant client may be a participant client from a group 121' that supplies its respective primary digital video stream to the first generating function 135'.
[0193] In some embodiments, one or more participant clients of group 121' may receive the second generated video stream by a second publishing function 136'' configured to receive the second generated video stream from the second generating function 135''.
[0194] Therefore, each of the primary video stream supply participant clients assigned to the first group 121' may be supplied with a first generated video stream if the primary digital video streams of each other are not supplied directly, which will be accompanied by a certain delay or latency due to the synchronization between the above-mentioned multiple primary video streams, and further, as described above, an additional delay or latency may be intentionally added to ensure sufficient time for event and / or pattern detection.
[0195] In response to this, each participating client assigned to the second group 121'' may be supplied with a second generated video stream, which includes the intentionally added delay in relation to the second generation step, added for the purpose of time-synchronizing the first generated video stream with the first and second primary video streams. This additional delay may, for example, make communication between the participating clients of the second group 121'' difficult, or may not be difficult, by causing the participating clients of the second group 121'' to interact with the video communication service 110 in a different way than the participating clients of the first group 121''.
[0196] Therefore, the participant clients of the first group 121' form a subgroup of all participant clients 121 currently participating in the video communication service 110, and are in and using the service in a “time zone” slightly ahead (e.g., 1-3 seconds ahead) of any participant clients to whom generated video streams, such as the first or second generated video stream, are instead continuously supplied. Nevertheless, other participant clients (not assigned to the first group 121', but instead assigned to the second group 121'') will be continuously supplied with the second generated video stream. This second generated video stream is generated based on at least one of the primary video streams on which the second generated video stream is generated (and at each point in time may include one or both of multiple primary video streams), but is in a slightly later “time zone”. Since the first generated video stream is generated directly based on at least one of these primary video streams, there is no additional delay or latency to synchronize it with the video stream already generated based on the primary video stream itself, providing these participant clients 121 with a more direct and low-latency video communication service 110 experience.
[0197] In this case as well, it may mean that participant client 121 assigned to the first group 121' will not be provided with access to the second generated video stream.
[0198] As shown in Figure 8, the second generated video stream may also be generated as a digital video stream based in addition to the generated output video (third generated output digital video stream) of the third generation step 135'''. Thus, the generated video streams from each of the generation steps 135', 135'', and 135''' are generated using different intentionally added latency ("time zones") to the above primary video stream, although they are based on the primary input digital video stream, which is at least partially overlapping, and are supplied to the different groups 121', 121'', and 121''' of the participant clients 121.
[0199] Participant clients 121 assigned to the third group 121''' may have less stringent latency requirements than participant clients 121 assigned to the first group 121'. For example, participant clients 121 in the first group 121' may be members of the aforementioned discussion panel (requiring low latency because they interact with each other in real time), while participant clients 121 in the third group 121'' may constitute an expert panel or similar panel that interacts with a panel but in a more structured way (e.g., using clear questions / answers), and therefore may tolerate higher latency than the first group 121'.
[0200] Both the first and third generated video streams may, in some cases, be supplied to a second generation function 135'' to form the basis for the generation of a second generated video stream.
[0201] Therefore, the first generation step 135' may include introducing an intentional delay or latency of the type described above, in addition to the delay introduced as part of the synchronization of the first and second primary video streams, thereby achieving, for example, sufficient time to perform efficient event and / or pattern detection. The introduction of such an intentional delay or latency may be done as part of the synchronization performed by the synchronization function 133 (not shown in Figure 8 for reasons of simplification). The same applies to the third generation step 135'', but it may introduce an intentional delay or latency different from the delay or latency introduced for the first generation step 135'.
[0202] In particular, intentionally introduced delays or latency result in a time asynchronous relationship between the first and third generated video streams. This means that the first and third generated video streams do not follow a common timeline when they are published immediately and sequentially as each individual frame is generated.
[0203] Therefore, three separately generated video streams may be generated and consumed / published simultaneously, but in different "time zones." Although they are at least partially based on the same primary video material, the multiple generated video streams will be published with different latency. The first group 121', which requires the lowest latency, can interact using the first generated video stream, which provides very low latency. Meanwhile, the third group 121''', which is willing to accept slightly higher latency, can interact using the second generated video stream, which provides higher latency but on the other hand offers greater flexibility in terms of intentionally added delay, thereby achieving better auto-generation as described elsewhere in this specification. Meanwhile, the second group 121'', which is not so sensitive to latency, can incorporate material from both the first group 121' and the third group 121''' and can also enjoy interacting using the second generated video stream, which is automatically generated in a very flexible manner. It should be noted that all these groups of participant users 121', 121'', and 121''' interact with each other using the video communication service 110, despite using the various latency periods described above and therefore operating in different "time zones". However, due to the synchronization of individual input video streams in each generation function, participant users 121 will not notice the different latency periods from their respective perspectives.
[0204] As described above, each participant client 121 assigned to each of the groups 121', 121'', and 121'''' can participate in the same video communication service 110 from which the second generated video stream is continuously published.
[0205] Furthermore, the differences among the groups 121', 121'', and 121''' may be associated with different participant interaction permissions in the video communication service 110. In these embodiments and other embodiments, the differences among the groups 121', 121'', and 121''' may be associated with different maximum time delays (latency) used to generate the respective generated video streams that are exposed to the participant clients 121 assigned to the groups 121', 121'', and 121'''.
[0206] For example, the first group 121' of panel discussion participant clients may be associated with full dialogue privileges and can speak at any time. The third group 121'' of participant clients may be associated with slightly restricted dialogue privileges, for example, requiring them to request the floor before they can speak by the video communication service 110 unmuting their microphones. The second group 121'' of audience participant users may be associated with even more restricted dialogue privileges, for example, only being able to ask questions in text in a common chat room and not being able to speak.
[0207] Therefore, different groups of participant users may be associated with different dialogue privileges and different latency for their respective generated video streams, such that latency is an increasing function of decreasing dialogue privileges. The more freely a participant user 121 is allowed to interact with other users through the video communication service 110, the lower the acceptable latency (waiting time). The lower the acceptable latency (waiting time), the less likely the corresponding auto-generating function is to consider detected events, patterns, etc.
[0208] The group with the longest waiting time may be a group of viewers only, who have no right to participate in the video communication service other than passively attending.
[0209] In particular, the maximum time delay (latency) for each of the above groups 121', 121'', and 121'''' can be determined as the maximum latency difference across all primary video streams and any generated video streams that are continuously exposed to the participant clients of that group. This total may include time delays that have been intentionally added for the purpose of detecting events and / or patterns, as described above.
[0210] As used herein, the terms “generated” and “generated digital video stream” may refer to different types of generation. In one example, a single, clearly defined digital video stream is generated by a central entity, such as a central server 130, and forms the generated digital video stream for supply and exposure to each set of specific participant clients 121 that will consume the generated digital video stream. In another example, different individual participant clients 121 may view slightly different versions of the generated digital video stream. For example, the generated digital video stream may include multiple separate or combined digital video streams that a local software function 125 of a participant client 121 can cause the user 122 to switch, place on a screen 124, or otherwise configure or process. Often, the important question is in which “time zone” (i.e., at what latency) the generated digital video stream, including time-synchronized subcomponents, is supplied. Therefore, in relation to Figure 8, the case described above where different participant clients 120 of the first group 121' are supplied with each other's primary video streams can be considered as the first generated digital video stream being supplied to these participant clients (meaning that a time-synchronized set of raw or processed first and second primary digital video streams becomes available to both the first and second participant clients).
[0211] To further clarify and illustrate the use of the participant client groups 121', 121'', and 121'''' described above, the following example is provided in the form of a video communication service conference involving three different simultaneous "time zones":
[0212] The first group of participant clients 121' are experiencing the interaction in real time, or at least near real time (depending on unavoidable hardware and software latency). These participant clients are supplied with video, including audio, from each other to enable such interaction and communication between the users 122. The first group 121' may service the core user 122 of the meeting with interactions that other participant clients (other than the first group 121') may be interested in joining.
[0213] Such a second group 121'' of other participant clients participates in the same meeting but is in a different "time zone" that is further removed from real-time than the first group of participant clients 121''. For example, the second group 121'' may be an audience with dialogue privileges, such as being able to ask questions to the first group 121''. The "time zone" of the second group 121'' may have a delay in relation to the "time zone" of the first group 121'' such that the questions and answers raised are noticed but with a short delay. On the other hand, this slightly larger delay allows the participant clients of this second group 121'' to experience a generated digital video stream that is automatically generated in a more complex way, providing a more comfortable user experience.
[0214] A third group 121''' of other participant clients also participates in the same conference, but only as viewers. This third group 121''' consumes a generated digital video stream that can be automatically generated in a more elaborate and complex way, consumed in a third "time zone" with even greater latency than the second "time zone". However, since the third group 121''' cannot supply input to the communication service in a way that affects the first group 121' and the second group 121'', the third group 121''' experiences the conference as being conducted "in real time" and conducted in a convincing manner.
[0215] Naturally, there may be four or more such groups of participant clients, each associated with a different meeting “time zone” where the time delays become progressively larger and the complexity of generation increases, using the principles described herein.
[0216] Figure 9 illustrates a method according to the present invention, which is for supplying a shared digital video stream.
[0217] In the first step S1, the method is initiated.
[0218] In the subsequent collection step S2, a first digital video stream is collected from a first digital video source, and a second digital stream is collected from a second digital video source, which are generally done in the manner described above. Thus, the first and / or second digital video streams may be collected from their respective participant clients 121, or from an external source 300, or may be performed by the collection function 131 of the central server 130.
[0219] In the subsequent first generation step S4, the shared digital video stream is generated as an output digital video stream. This generation can generally be performed as described above by generation steps 135, 135', 135'', 135''', etc.
[0220] In the first generation step S4, the shared digital video stream is generated based on a series of sequentially considered frames of the first digital video stream such that image information from the first digital video source is visible in the shared digital video stream, but image information from the second digital video source is not visible in the shared digital video stream. In other words, the shared video stream includes, at least to some extent, visual material derived from the first digital video stream, but does not include visual material derived from the second digital video stream.
[0221] In the subsequent trigger detection step S5, the first and / or second digital video streams are digitally analyzed to detect at least one trigger.
[0222] This analysis and detection may be performed by the same generation steps 135, 135', 135'', 135''' that generate the shared video stream, and is based on the automatic detection of a predetermined type of image and / or sound pattern.
[0223] The triggers may be events or patterns of the types described and illustrated above (performed by the event detection function 132 and the pattern detection function 134, respectively), and their detection is typically performed using automated digital processing of audio and / or image / video data contained in the digital video stream(s). For example, an automated image processing algorithm, such as using a trained neural network or other machine learning tool as illustrated above, may be employed to automatically detect the presence of a particular trigger based on images contained in the first and / or second video feeds. Correspondingly, a corresponding type of conventional automated audio processing algorithm can be used to detect the presence of a particular trigger based on audio contained in the first and / or second video feeds.
[0224] For an image and / or audio pattern to be of a "predetermined type" means that the pattern in question is characterized in terms of a set of one or more absolute or relative parameter values defined prior to the detection. This is illustrated below.
[0225] Generally, the presence of the above audio or image pattern constitutes the corresponding trigger. Furthermore, the trigger is specifically predetermined to instruct the automated generation steps 135, 135', 135'', 135''' that generate the shared video stream to change the generation mode (rules) of the shared digital video stream when the trigger is detected, according to predetermined generation rules.
[0226] Therefore, generation steps 135, 135', 135'', 135'''' may include or have access to a database that defines one or more triggers, either immediately or over time, with respect to corresponding parameter values that characterize the corresponding image and / or sound pattern.
[0227] In a subsequent second generation step S6 or S7 initiated in response to the detection of the above trigger, the shared digital video stream is then generated again by the same (or different) generation steps 135, 135', 135'', 135'''', but not in the same manner as in the first generation step S4.
[0228] In the first alternative second generation step S6, the shared digital video stream is generated as an output digital video stream based on a series of consecutively considered frames of the second digital video stream, such that image information from the second digital video source is visible in the shared digital video stream. Note that in this case, the output digital video stream may be generated based on or without a series of consecutively considered frames of the first digital video stream, such that image information from the first digital video source is visible in the shared digital video stream. In other words, when switching from the first generation step S4 to the second generation step S6, the shared video stream may change from a state where it displays content from the first video stream but not from the second video stream, to a state where it displays content from the second video stream but not from the first video stream, or to a state where it displays content from both the first and second video streams.
[0229] In a second alternative generation step S7, the shared digital video stream is generated as an output digital video stream based on a sequentially considered multiple frames of the first digital video source, but with respect to the first generation step S4, at least one of the following is performed on the first digital video stream: different cropping, different zooming, different panning, and different focus plane selection. In other words, the content of the video stream displayed in the shared video stream is cropped, decropped, zoomed in, zoomed out, panned vertically and / or horizontally, and / or the focus plane of the video stream is shifted relative to the current crop / zoom / pan / focus plane state of the video stream as used in the first generation step S4.
[0230] It is understood that such cropping / panning / zooming / focusing planes may be performed by the generation steps 135, 135', 135'', 135''' based on an existing video stream (which may itself contain multiple possible focus planes having different image information at the pixel level) and / or by the generation steps 135, 135', 135'', 135''' communicating commands to a video source (such as a digital video camera) capturing the video stream to modify the corresponding capture parameters accordingly. For example, this may then involve the corresponding camera capturing the video stream in question zooming, panning, and / or shifting its focus plane in accordance with the instructions provided by the generation steps 135, 135', 135'', 135'''.
[0231] In the subsequent publishing step S8, the output digital video stream is continuously supplied to consumers of the shared digital video stream, such as participant clients 121 and / or external consumers 150, in the general manner described above.
[0232] Subsequently, as shown in Figure 9, this method can be repeated by returning to step S2.
[0233] The method concludes in the following step S9.
[0234] A first digital video stream may be continuously captured by a first digital camera, and a second digital video stream may be continuously captured by a second different digital camera (therefore, these constitute the primary video stream using the terminology already used herein). Alternatively, the first and / or second digital video streams may constitute their respective previously generated digital video streams, such as when using multiple different groups 121', 121'', 121'''' of participant clients 121 associated with the different latency ("time zones") described above, each of which is generated with a different latency ("time zone") compared to the currently generated shared video stream (see above for further details regarding such "time zones").
[0235] If the first video stream is an already generated video stream, it is preferable (though not required) for the crop / zoom / pan settings to be performed based on the already existing first video stream, rather than instructing the upstream camera to change the crop / zoom / pan settings.
[0236] This method allows for the automated generation of shared video streams, which can provide a more intuitive and natural experience for consumers of the generated shared videos, because the actual audio / video content of each video stream is used to detect triggers that would cause the automated generation to shift from one automated generation format to a different one.
[0237] Triggers can be predefined with appropriate parameters to address various needs. For example, the actions of individual people depicted in the first and / or second video streams may be automatically evaluated in relation to such triggers, and the generation (performance) format may be modified depending on how such actions are performed. In another example, certain predefined triggers may be used as manual cues given by people depicted in the first and / or second video streams to modify the generation format on the fly while the generation is in progress.
[0238] Below are some examples of such triggers and the corresponding changes to the generated format.
[0239] In the first example shown in Figures 10a and 10b, a given pattern includes a first (human or, for example, machine) participant 430 depicted in the first digital video stream, as illustrated in Figure 10a, as captured by a first digital video camera 410. In the example of Figure 10a, the first participant 430 gazes toward an object 440, which is a second (human or machine) participant, and the second participant 440 is subsequently depicted in a second digital video stream. In this example, the second digital video stream is captured by a second digital video camera 420.
[0240] It is understood that the second object 440 may be anything else, such as a group of human participants or any physical object of general interest for ongoing communication. For example, a shared video stream may be material for a medical procedure, thereby object 440 may be a part of the patient. In other examples, object 440 may be an object in an educational session or a sales presentation. Also, object 440 may be, for example, a whiteboard or a slide presentation screen.
[0241] The first and second video streams may be of any of the types described herein.
[0242] In the cases shown in Figures 10a and 10b, a predetermined image and / or audio pattern is detected based on information regarding the relative orientation of the first camera 410, the participant 430, and the object 440. This information may be present within the system 100 (particularly within the central server 130), such as being supplied in advance during setup / configuration and / or being automatically detected while generation is in progress.
[0243] For example, the positions of the first camera 410 and the second camera 420 may be supplied from each camera 410, 420 to the central server 130 by cameras equipped with measuring means such as a MEMS circuit with an accelerometer and a gyroscope, or conventional position measuring means such as a stepping motor arranged to continuously or intermittently measure the current position of the cameras 410, 420 based on some appropriate geometric criteria. In other examples, the orientation of the cameras 410, 420 may be detected by a third camera (not shown) using an appropriate automated digital image processing algorithm, the third camera capturing at least one of the cameras 410, 420 in an image, and using digital image processing to determine the relative orientation based on this captured image information.
[0244] In this context, it should be noted that "orientation" can encompass both the location and direction components.
[0245] The positions of the first participant 430 and object 440 relative to some appropriate reference frame (such as the first camera 410 and / or the second camera 420) can be determined using digital image processing based on the video stream(s) captured by the first camera 410 and / or the second camera 420.
[0246] As shown in Figure 10a, the first participant 430 is looking downwards at the figure, rather than at the second participant 430.
[0247] In this example, a predetermined image and / or audio pattern is further detected based on at least one digital image-based determination of the body orientation, head orientation and gaze orientation of a first participant 430, which is based on a first digital video stream.
[0248] As shown in Figure 10b, the first participant 430 turns to face the second participant 440 and is looking at (gazing at) the second participant 440.
[0249] The orientation of the body and head of the first participant 430 may be determined based on digital image processing of the first video stream captured by the first camera 410. Such an algorithm is conventional in itself and can use, for example, prior knowledge of the expected shape of the first participant 430 in the video stream captured by the first camera 410 when it turns in various directions. This can be achieved using a trained neural network or other machine learning components. Along with the relative orientation of the first camera 410, the relative position of the first participant 430 and the object 440, and the determined body or head orientation of the first participant 430, the central server 130 can determine whether the first participant 430 is facing (head or body) toward the object 440 in question.
[0250] The gaze direction of the first participant 430 can be achieved in a similar manner, such as based on images captured by the first video camera 410. Such gaze tracking techniques are known in themselves and may be based, for example, on identifying the position of the pupil and light reflection visible in the eye of the first participant 430.
[0251] The trigger can be defined as the detection of a transition pattern, such as a transition by the first participant 430 from a state where it is not directed towards or looking at object 440 to a state where the first participant 430 is actually directed towards and / or looking at object 440. Therefore, the central server 130 may continuously monitor for such transitions based on appropriately configured corresponding absolute or relative parameter values, and a trigger may be detected when such a transition occurs.
[0252] In this embodiment and other embodiments, multiple different predetermined image and / or sound patterns may be monitored simultaneously, and such detected predetermined image and / or sound patterns then constitute a detected corresponding trigger, which in turn leads to automatic generation switching to each different corresponding mode according to a predetermined parameterized generation logic.
[0253] Figure 10c shows a second example, similar to that shown in Figure 10b, but the predetermined image and / or sound pattern corresponding to the trigger in question is not the first participant 430 turning to or gazing at object 440. Instead, the predetermined image and / or sound pattern includes a participant (such as the first participant 430) depicted in a first digital video stream (such as from a first camera 410) performing a predetermined gesture.
[0254] The gesture may be any gesture of a predetermined parameterized type, such as a gesture geometrically related to the object 440 depicted in the second digital video stream. Specifically, the gesture may be the first participant 430 pointing towards the object 440 (as shown in Figure 10c, the arm 431 of the first participant 430 is pointing towards the object 440). However, the gesture may also be based solely on the hand or fingers of the first participant 430, for example.
[0255] A predetermined pattern of this gesture type (and, in particular, its orientation in space, if any) may be detected in a manner corresponding to the situation described in relation to Figure 10b, and thus detected based on information regarding the relative orientation of the first camera 410, the first participant 430, and the object 440, and further detected based on a digital image-based determination of the gesture orientation of the first participant 430 based on the first digital video stream.
[0256] In both the case shown in Figure 10b and the case shown in Figure 10c, the orientation of the second camera 420 is also detected and can be used to determine whether a trigger is detected. For example, it can be used to determine whether object 440 is visible to the second camera 420, which may constitute a condition for the trigger to be detected if the detected trigger is related to switching to the second camera 420. In some embodiments, the second generation step S6 may include determining one second camera 420 (of a plurality of possible second cameras) that is currently displaying object 440 and selecting the second camera 420 to supply the second video stream in the second generation step S6.
[0257] In the example shown in Figure 10d, there is only one camera, namely the first camera 410 (of course, it is understood that in various embodiments there may be more cameras and other video stream generation components). The first camera 410 captures both the first participant 430 and the object 440. When the first participant 430 turns their body, head, or gaze toward the object 440, this detected image pattern constitutes a detected trigger. In this case, the second generation step may include panning and / or zooming and / or cropping (trimming) the video stream captured by the first camera 410 in the generated output video stream so as to focus the viewer's attention on the object 440.
[0258] In a real-world example, the detected attention of the first participant 430 (embodied in the body, head, or gaze direction of the first participant 430, as illustrated above) triggers automatic generation (in the second generation step above) to somehow emphasize or shift focus based on an existing video stream or by instructing a unit that generates the primary video stream, thereby increasing the visual focus on the object 440 in question.
[0259] In another example shown in Figures 11a and 11b, a given image and / or audio pattern includes a first participant 430, depicted in a second digital video stream, gazing toward a second camera 420 that is continuously capturing the second digital video stream. Specifically, in Figure 11a, the first participant 430 is not gazing toward the second camera 420, whereas in Figure 11b, the first participant 430 is gazing toward the direction of the second camera 420. Thus, this detected switching of gaze direction may constitute the detection of the corresponding trigger in question.
[0260] In other examples, a given image and / or audio pattern includes relative changes in motion in a first digital video stream and / or a second digital video stream. For example, the amount of general motion shown in the first video stream may be parameterized, and the zoom of the first video stream and / or the second video stream may increase as a function of the decrease in general motion, or vice versa. Alternatively, the second generation step may switch to the second video stream in the case of an increase in detected general motion, the second video stream being captured by a second camera 420 showing a wider-angle or further-view of the scene depicted by the first camera 410. Correspondingly, the performance of such zoom-in / zoom-out / camera switching can be determined using recorded participant 430 speech audio. For example, if participant 430 is recorded speaking louder, there may be a zoom-out of the first video stream in the output shared video stream, and vice versa.
[0261] In yet another example, a given image and / or sound pattern sequentially includes a given sound pattern characterized by its frequency and / or amplitude and / or amplitude time derivative and / or absolute amplitude change and / or sound position, which is determined, for example, by the relative microphone volume of a particular sound-capturing microphone. Such a microphone may be, for example, part of a first camera 410 or a separate microphone. Such a microphone is positioned to record sound occurring in or directly related to a scene displayed in a first video stream.
[0262] For example, a predetermined voice pattern may consist of a predetermined phrase containing at least one orally spoken word. The voice may be provided as part of a video stream that is recorded and supplied to a first generation step, and the predetermined pattern may be detected by a central server 130, after which, if a corresponding trigger is detected, generation may switch to a second generation step. The voice analysis can use appropriate digital voice processing algorithms, such as a rule-based decision engine that uses various voice information (pitch, amplitude, pattern matching, etc.) or a trained neural network, to determine whether a predetermined sound pattern has been detected.
[0263] As described in detail above, the method may also include a delay step (see Figure 9), where a delay is intentionally introduced for at least the first and second digital video streams, and this delay exists in the shared digital video stream. The trigger detection step may then be performed based on the first and / or second digital video streams before introducing the aforementioned delay.
[0264] The waiting time is a maximum of 30 seconds, for example, a maximum of 5 seconds, for example, a maximum of 1 second, for example, a maximum of 0.5 seconds.
[0265] Using such intentionally added delays, automatic generation can plan an automatic switch from the first camera 410 to the second camera 420 based on a detected trigger (e.g., participant 430 looking into the second camera 420), and this plan is made a certain amount of time (e.g., 0.5 seconds) before this event joins the generated shared video stream. In this way, such a switch can be made precisely at the time the trigger actually occurs, or at a time that best fits other parameterized generation parameters, such as timing with the speech rhythm of participant 430 in question. If there are multiple groups of the type described above 121', 121'', 121'''', such plans may be made on different time horizons with respect to the generated output video streams generated for each participant client 121 of the different groups 121', 121'', 121'''' (for consumption by them).
[0266] A predetermined image and / or sound pattern may constitute (or be determined to constitute) an "event" and / or "pattern" of the type described above in relation to the event detection function 132 and the pattern detection function 132.
[0267] The following are some examples illustrating how the present invention can be put into practice:
[0268] In multicam presentations, talk shows, and panel discussions, different camera angles of the presenter can be displayed using different cameras. The auto-generation feature can be configured to automatically select different camera angles depending on which camera the presenter is looking at, and / or triggered by gestures or audio cues.
[0269] In video podcasts and talkinghead videos, the auto-generation feature can be configured to automatically switch between multiple different cameras depending on which camera the current speaker is facing.
[0270] In town hall meetings, the auto-generating feature can be configured to switch to a camera facing the audience to add input and questions from participants. This can be triggered by monitoring the audio feed associated with that camera and when it reaches a certain level to go live, or by a voice command such as "question from the audience."
[0271] In product presentations and reviews, the auto-generation feature can be configured to automatically switch to a camera pointed at the product when motion is detected at the source, or based on another trigger.
[0272] In robotic surgery, the video stream captured by the robotic camera recording the surgery can be replaced with a normal informational presentation when a predetermined gesture, audio cue is detected, or when it is recognized that the surgeon is not using the surgical console or has raised their head from the surgical console.
[0273] In an educational context, the camera can be set to point towards a regular whiteboard or blackboard, and the auto-generating function can be configured to switch to that camera when the teacher makes gestures or gives instructions via voice commands towards the board.
[0274] In cultural events such as concerts, an automated function can be configured to switch between multiple cameras focused on the singer, band members, and orchestra. This can be triggered by gestures or by which camera the talent is looking at.
[0275] In theatrical performances, the auto-generation function can be configured to cut between different camera angles depending on who is speaking, based on face tracking, voice cues, gestures, or according to a predetermined schedule (rundown).
[0276] Therefore, in addition to the types of trigger detection described herein, automatic generation can also switch from one format (generation rule) to another based on a predetermined schedule (rundown), and of course, can be manually overridden in some cases.
[0277] The present invention also relates to a computer software function for providing a shared digital video stream in accordance with the above. Such a computer software function may be configured to perform at least some of the above-described steps of collection, delay, first generation, trigger detection, second generation, and publishing during runtime. The computer software function may be configured to run on the physical or virtual hardware of the central server 130, as described above.
[0278] The present invention also relates to a system 100 that provides a shared digital video stream, comprising, in order, a central server 130. The central server 130 may be configured to perform, in order, at least some of the above-described steps of collection, delay, first generation, trigger detection, second generation, and publishing. For example, these steps may be performed by the central server 130 performing the above-described computer software functions for performing the above-described steps. Collection may be performed by a collection function 131. Detection of predetermined image and / or sound patterns and triggers as described above may be performed by event pattern detection functions 132, 134, or generation function 135 of the central server 130. Any intentional delay may be performed by the collection function 131, etc., in the manner generally described above. Publishing may be performed by a publishing function 136.
[0279] The principle of automatic generation based on the available set of the input video stream described above, including time synchronization of such an input video stream, event and / or pattern detection, and trigger detection, etc., may be applied simultaneously and in parallel at different levels. Therefore, one such automatically generated video stream may form the available input video stream for a downstream automatic generation function that generates video streams in sequence.
[0280] The central server 130 may be configured to control the assignment of groups 121', 121'', 121''' to individual participant clients 121. For example, during the process of a live video communication service session, dynamically changing the group assignment for a particular such participant client may be part of the automatic generation of the video communication service by the central server 130. Such re-assignment may be dynamically triggered based on a predetermined time table or, for example, in response to the requests of individual participant client users 122 (provided via the client 121), e.g., as a function of parameter data that may change dynamically over time.
[0281] Correspondingly, the central server 130 may be configured to dynamically change the group configuration during the process of the video communication service, such as using a particular group only for a predetermined time frame (e.g., during a scheduled panel discussion).
[0282] Changing the group assignment in a predetermined manner may be an automatic result of the detection of a particular trigger, in a manner corresponding to what was described above in relation to FIGS. 9 to 11b.
[0283] In all of the above-described aspects, the present invention may further include an interaction step, in which at least one participant client of the first group (the first group is associated with a first waiting time) interacts bidirectionally with at least one participant client of the same first group or the second group (the second group is associated with a second waiting time), and the second waiting time is different from the first waiting time. It is understood that all of these participating clients may be participants of one and the same communication service of the type described above.
[0284] As described above, the preferred embodiments have been described. However, it will be apparent to those skilled in the art that many changes can be made to the disclosed embodiments without departing from the basic idea of the present invention.
[0285] For example, as part of the system 100 described herein, many additional features not described herein can be provided. Generally, the solution currently being described provides a framework that can build detailed functionality and features to accommodate a wide variety of specific applications in which video data streams are used for communication.
[0286] Generally, all that has been said with respect to the method is applicable to the system and computer software product, and vice versa.
[0287] Therefore, the present invention is not limited to the described embodiments, and various changes are possible within the scope of the appended claims.
Claims
1. A method for providing a shared digital video stream, the method comprising the following steps: In the collection step, a first digital video stream is collected from a first digital video source, and a second digital stream is collected from a second digital video source; In the first generation step, the shared digital video stream is generated as an output digital video stream based on a series of sequentially considered frames of the first digital video stream such that image information from the first digital video source is visible in the shared digital video stream, and image information from the second digital video source is not visible in the shared digital video stream; In the trigger detection step, the trigger is automatically detected by digitally analyzing the first and / or second digital video streams to automatically detect an event or pattern, where, a) The event is a physical motion event detected in the first and / or second digital video stream for a person or object, or b) The event is a change in lighting detected in the first and / or second digital video stream, or c) The pattern includes a predetermined image pattern having relative changes in general motion in the first and / or second digital video stream, and / or d) The pattern includes a predetermined audio pattern characterized by its audio position in the first and / or second digital video stream; Here, the trigger instructs to change the generation mode of the shared digital video stream according to a predetermined generation rule; In a second generation step initiated in response to the detection of the trigger, the shared digital video stream is generated as an output digital video stream based on a successively considered plurality of frames of the second digital video stream such that image information from the second digital video source is visible in the shared digital video stream, and / or the shared video stream is generated as an output digital video stream based on a successively considered plurality of frames of the first digital video source, but with respect to the first generation step, at least one of different cropping, different zooming, different panning or different focus plane selection is performed on the first digital video stream; and, In the publishing step, the output digital video stream is continuously provided to consumers of the shared digital video stream.
2. The method according to claim 1, The pattern is a method based on at least two events detected for the first and / or second digital video stream.
3. The method according to claim 1, The first digital video stream described above is continuously captured by the first digital camera. A method in which a second digital video stream is continuously captured by a second digital camera.
4. A method according to claim 1, further comprising: The positions of the participant and the object are determined relative to the first camera and / or the second camera using digital image processing based on the first and / or second digital video streams.
5. The method according to claim 1, The method comprising the second generation step of determining a specific second camera from among a plurality of possible second cameras that are currently showing the object, and selecting the second camera to provide the second video stream.
6. The method according to claim 1, A method wherein the generation in the second generation step is configured to enhance or shift the focus of the output digital video stream to increase the visual focus on the object.
7. A method according to claim 1, further comprising: A delay step in which a delay is intentionally introduced for at least the first and second digital video streams, wherein the delay is present in the shared digital video stream, where The trigger detection step is performed based on the first and / or second digital video stream before introducing the latency.
8. The method according to claim 7, The aforementioned waiting time is a maximum of 30 seconds.
9. A computer program that causes a computer to perform a process of providing a shared digital video stream, wherein the computer program causes the computer to perform the following steps when the process is being performed: In the collection step, a first digital video stream is collected from a first digital video source, and a second digital stream is collected from a second digital video source; In the first generation step, the shared digital video stream is generated as an output digital video stream based on a series of sequentially considered frames of the first digital video stream such that image information from the first digital video source is visible in the shared digital video stream, and image information from the second digital video source is not visible in the shared digital video stream; In the trigger detection step, the trigger is automatically detected by digitally analyzing the first and / or second digital video streams to automatically detect an event or pattern, where, a) The event is a physical motion event detected in the first and / or second digital video stream for a person or object, or b) The event is a change in lighting detected in the first and / or second digital video stream, or c) The pattern includes a predetermined image pattern having relative changes in general motion in the first and / or second digital video stream, and / or d) The pattern includes a predetermined audio pattern characterized by its audio position in the first and / or second digital video stream; Here, the trigger instructs to change the generation mode of the shared digital video stream according to a predetermined generation rule; In a second generation step initiated in response to the detection of the trigger, the shared digital video stream is generated as an output digital video stream based on a successively considered plurality of frames of the second digital video stream such that image information from the second digital video source is visible in the shared digital video stream, and / or the shared video stream is generated as an output digital video stream based on a successively considered plurality of frames of the first digital video source, but with respect to the first generation step, at least one of different cropping, different zooming, different panning or different focus plane selection is performed on the first digital video stream; and, In the publishing step, the output digital video stream is continuously provided to consumers of the shared digital video stream.
10. A system that provides a shared digital video stream, the system comprising a central server, the central server having the following functions: A collection function configured to collect a first digital video stream from a first digital video source and a second digital stream from a second digital video source; A first generation function configured to generate the shared digital video stream as an output digital video stream based on a series of sequentially considered frames of the first digital video stream, such that image information from the first digital video source is visible in the shared digital video stream, and image information from the second digital video source is not visible in the shared digital video stream; A trigger detection function that automatically detects triggers, which are events or patterns, by digitally analyzing the first and / or second digital video streams, wherein a) The event is a physical motion event detected in the first and / or second digital video stream for a person or object, or b) The event is a change in lighting detected in the first and / or second digital video stream, or c) The pattern includes a predetermined image pattern having relative changes in general motion in the first and / or second digital video stream, and / or d) The pattern includes a predetermined audio pattern characterized by its audio position in the first and / or second digital video stream; Here, the trigger instructs to change the generation mode of the shared digital video stream according to a predetermined generation rule; A second generation function configured to be initiated in response to the detection of the trigger, and which generates the shared digital video stream as an output digital video stream based on a successively considered plurality of frames of the second digital video stream such that image information from the second digital video source is visible in the shared digital video stream, and / or which generates the shared video stream as an output digital video stream based on a successively considered plurality of frames of the first digital video source, but which is configured to perform at least one of different cropping, different zooming, different panning or different focus plane selection of the first digital video stream compared to the first generation function; and, A publishing function configured to continuously provide the output digital video stream to consumers of the shared digital video stream.