System and method for generating a video stream - Patent application

JP2025506378A5Pending Publication Date: 2026-02-12LIVEARENA TECH AB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024545896
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-02-04
Filing Date
2023-02-03
Publication Date
2026-02-12

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The method for providing a shared digital video stream includes: in a collection step (S2), collecting a first and a second digital video stream; in a first generation step (S4), generating the shared stream based on the first stream such that the first source is visible in the shared stream and the second source is invisible in the shared stream; in a trigger detection step (S5), analyzing the first and / or second stream to detect a trigger based on detection of a predetermined image and / or audio pattern, the trigger being a detected gaze or gesture; in a second generation step (S6, S7), generating the shared stream based on the second stream and / or the first source, but with at least one of different cropping, different zooming, different panning, or different focus plane selection; in a publishing step (S8), providing the output stream to a consumer. The present invention also relates to a computer software product and system.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a system, a computer software product and a method for generating a digital video stream, in particular for generating a digital video stream based on two or more different digital input video streams. In a preferred embodiment, the digital video stream is generated in the context of a digital video conference, in particular a digital video conference or meeting system, in which a number of different simultaneously connected users participate. The generated digital video stream may be published externally or within the digital video conference or digital video conference system.

[0002] In other embodiments, the invention is applied to contexts that are not digital video conferencing, but where multiple digital video input streams are simultaneously processed and combined into a digital video stream to be generated. For example, such a context may be educational or instructional. [Background technology]

[0003] Many digital video conferencing systems are known, such as Microsoft® Teams®, Zoom®, and Google® Meet®, that allow two or more participants to meet virtually using digital video and audio that is recorded locally and broadcast to all participants, emulating a physical meeting.

[0004] There is a general need to improve such digital video conferencing solutions, particularly with regard to the generation (production) of viewing content: what content to show, at what time, to whom, and through what distribution channel.

[0005] For example, some systems automatically detect who is currently speaking and display the corresponding video feed of that participant to the other participants. Many systems allow for the sharing of graphics such as the currently displayed screen, a viewing window, or a digital presentation. But as virtual meetings become more complex, it will soon become difficult for services to know what, of all the information currently available, should be shown to each participant at any given time.

[0006] In another example, a presenting participant moves around on stage while talking about slides in a digital presentation, in which case the system needs to decide whether to show the presentation, the presenter, or both, or switch between the two.

[0007] It may be desirable to generate, by an automated generation process, one or more output digital video streams based on multiple input digital video streams, and provide such generated digital video stream or streams to one or more consuming entities. DISCLOSURE OF THEINVENTION [Problem to be solved by the invention]

[0008] However, in many cases, due to the many technical challenges faced by such digital video conferencing systems, it is difficult for dynamic conference screen layout managers and other automated generators to select what information to display.

[0009] First, low latency is important because digital video conferencing is real-time sensitive. This becomes problematic when different incoming digital video streams are associated with different latency, different frame rates, different aspect ratios, or different resolutions, such as when different participants join using different hardware. Often, such incoming digital video streams require processing for a well-formed user experience.

[0010] Second, there is the issue of time synchronization: the various input digital video streams, such as external digital video streams and digital video streams provided by the participants, are typically fed into a central server or the like, and there is no absolute time to synchronize each of such digital video feeds with. Similar to too much delay, unsynchronized digital video feeds lead to a poor user experience.

[0011] Third, digital video conferences between multiple participants may involve different digital video streams with different encodings or formats, which require decoding and re-encoding, creating problems in terms of latency and synchronization, and such encoding is computationally intensive and expensive in terms of hardware requirements.

[0012] Fourth, the fact that different digital video sources may be associated with different frame rates, different aspect ratios, and different resolutions can result in unpredictable changes in memory allocation needs that require continual balancing, potentially resulting in additional latency and synchronization issues, which in turn necessitates the need for large buffers.

[0013] Fifth, participants may experience a variety of difficulties in terms of connectivity fluctuations, drop-off / reconnection, etc., posing additional challenges to automatically generating a shaped user experience.

[0014] These issues are amplified in more complex meeting situations, including those with large numbers of participants; participants connecting using different hardware and / or software; using externally provided digital video streams; screen sharing; and multiple hosts.

[0015] Corresponding problems arise in other contexts, when an output digital video stream is to be generated based on multiple input digital video streams, such as in digital video generation systems for education and instruction.

[0016] Swedish patent application SE2151267-8 (unpublished at the effective priority date of the present application) discloses various solutions to the above mentioned problems.

[0017] Swedish patent application SE2151461-7, which is unpublished as of the effective date of this application, discloses various solutions specific to dealing with latency in multi-participant digital video environments, where different participant groups are associated with different general latency times.

[0018] Achieving automated generation of the above-mentioned type of conferences remains problematic: in particular, when multiple cameras are available in such automated generation, it has proven difficult to use the image output of each of such cameras in a way that is natural and intuitive for the conference participants.

[0019] The present invention is directed to solving one or more of the problems set forth above. [Means for solving the problem]

[0020] Accordingly, the present invention relates to a method for providing a shared digital video stream, the method comprising the steps of: in a collecting step, collecting a first digital video stream from a first digital video source and collecting a second digital stream from a second digital video source; in a first generating step, generating the shared digital video stream as an output digital video stream based on a plurality of consecutively considered frames of the first digital video stream such that image information from the first digital video source is visible in the shared digital video stream and image information from the second digital video source is not visible in the shared digital video stream; in a trigger detecting step, automatically detecting a trigger by performing at least one of the following: a) digitally analyzing the first digital video stream to detect an image depicted in the first digital video stream; automatically detecting a trigger in the form of a participant gazing towards or making a predefined gesture in relation to an object depicted in the second digital video stream, wherein the detection is based on information regarding the relative orientations of a first camera, the participant and the object, the detection being further based on a digital image-based determination of at least one of the participant's body direction, head direction, gaze direction and gesture direction; and b) digitally analysing the first and / or second digital video streams to automatically detect a trigger in the form of a gaze direction of a participant depicted in the second digital video stream transitioning from not gazing at a second camera to gazing at the second camera, wherein the trigger instructs to change a generation mode of the shared digital video stream according to predefined generation rules;in a second generating step, initiated in response to detection of the trigger, generating the shared digital video stream as an output digital video stream based on a plurality of consecutively considered frames of the second digital video stream such that image information from the second digital video source is visible in the shared digital video stream and / or generating the shared video stream as an output digital video stream based on a plurality of consecutively considered frames of the first digital video source, but with at least one of different cropping, different zooming, different panning or different focus plane selection of the first digital video stream compared to the first generating step; and in a publishing step, continuously providing the output digital video stream to consumers of the shared digital video stream;

[0021] The present invention also relates to a computer software product for providing a shared digital video stream, the computer software functions, when executed, performing the following steps: in a collecting step, collecting a first digital video stream from a first digital video source and collecting a second digital stream from a second digital video source; in a first generating step, generating the shared digital video stream as an output digital video stream based on a plurality of consecutively considered frames of the first digital video stream such that image information from the first digital video source is visible in the shared digital video stream and image information from the second digital video source is not visible in the shared digital video stream; in a trigger detection step, automatically detecting a trigger by performing at least one of the following: a) digitally analyzing the first digital video stream to generate a first digital video stream from the first digital video source and a second digital video stream from the second digital video source; a) automatically detecting a trigger in the form of a participant depicted in a first digital video stream gazing towards or making a predefined gesture in relation to an object depicted in the second digital video stream, wherein the detection is based on information regarding relative orientations of a first camera, the participant and the object, the detection being further based on a digital image based determination of at least one of the participant's body direction, head direction, gaze direction and gesture direction; and b) digitally analyzing the first and / or second digital video streams to automatically detect a trigger in the form of a gaze direction of a participant depicted in the second digital video stream transitioning from not gazing at a second camera to gazing at the second camera, wherein the trigger instructs to change a generation mode of the shared digital video stream according to predefined generation rules;in a second generating step, initiated in response to detection of the trigger, generating the shared digital video stream as an output digital video stream based on a plurality of consecutively considered frames of the second digital video stream such that image information from the second digital video source is visible in the shared digital video stream and / or generating the shared video stream as an output digital video stream based on a plurality of consecutively considered frames of the first digital video source, but with at least one of different cropping, different zooming, different panning or different focus plane selection of the first digital video stream compared to the first generating step; and in a publishing step, continuously providing the output digital video stream to consumers of the shared digital video stream;

[0022] The present invention also relates to a system for providing a shared digital video stream, the system comprising a central server, the central server comprising: a collection function configured to collect a first digital video stream from a first digital video source and to collect a second digital stream from a second digital video source; a first generation function configured to generate the shared digital video stream as an output digital video stream based on a plurality of consecutively considered frames of the first digital video stream such that image information from the first digital video source is visible in the shared digital video stream and image information from the second digital video source is not visible in the shared digital video stream; a trigger detection function configured to automatically detect a trigger by performing at least one of the following: a) digitally analyzing the first digital video stream to generate a first digital video stream from the first digital video source and a second digital video stream from the second digital video source; a) automatically detecting a trigger in the form of a participant depicted in the first digital video stream gazing towards or making a predefined gesture in relation to an object depicted in the second digital video stream, wherein the detection is based on information regarding relative orientations of a first camera, the participant and the object, the detection being further based on digital image based determination of at least one of the participant's body direction, head direction, gaze direction and gesture direction; and b) digitally analyzing the first and / or second digital video streams to automatically detect a trigger in the form of a gaze direction of a participant depicted in the second digital video stream transitioning from not gazing at a second camera to gazing at the second camera, wherein the trigger instructs to change a generation mode of the shared digital video stream according to predefined generation rules;a second generating function configured to be initiated in response to detection of the trigger and configured to generate the shared digital video stream as an output digital video stream based on a plurality of consecutively considered frames of the second digital video stream such that image information from the second digital video source is visible in the shared digital video stream and / or to generate the shared video stream as an output digital video stream based on a plurality of consecutively considered frames of the first digital video source, but with at least one of a different cropping, a different zooming, a different panning or a different focus plane selection of the first digital video stream compared to the first generating function; and a publishing function configured to continuously provide the output digital video stream to consumers of the shared digital video stream.

[0023] The present invention will now be described in detail with reference to exemplary embodiments thereof and the accompanying drawings. [Brief description of the drawings]

[0024] [Figure 1] FIG. 1 illustrates a first exemplary system. [Diagram 2] FIG. 2 illustrates a second exemplary system. [Diagram 3] FIG. 3 illustrates a third exemplary system. [Figure 4] FIG. 4 is a diagram illustrating the central server. [Diagram 5] FIG. 5 is a diagram illustrating the first method. [Figure 6a] FIG. 6a illustrates subsequent states associated with different method steps in the method illustrated in FIG. [Figure 6b] FIG. 6b illustrates subsequent states associated with different method steps in the method illustrated in FIG. [Figure 6c] FIG. 6c illustrates subsequent states associated with different method steps in the method illustrated in FIG. [Figure 6d] FIG. 6d illustrates subsequent states associated with different method steps in the method illustrated in FIG. [Figure 6e] FIG. 6e illustrates subsequent states associated with different method steps in the method illustrated in FIG. [Figure 6f] FIG. 6f illustrates subsequent states associated with different method steps in the method shown in FIG. [Figure 7] FIG. 7 is a diagram conceptually illustrating a common protocol. [Figure 8] FIG. 8 illustrates a fourth exemplary system. [Figure 9] FIG. 9 is a diagram illustrating the second method. [Figure 10a] FIG. 10a shows the cameras and participants in the above and various other configurations. [Figure 10b] FIG. 10b shows the cameras and participants in the above and various other configurations. [Figure 10c] FIG. 10c shows the cameras and participants in the above and various other configurations. [Figure 11a] FIG. 11a shows the cameras and participants in the above and various other configurations. [Figure 11b] FIG. 11b shows the cameras and participants in the above and various other configurations. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0025] All figures share the same or corresponding part reference numbers.

[0026] FIG. 1 shows a system 100 according to the invention, arranged to carry out a method according to the invention for providing a digital video stream, for example a shared digital video stream.

[0027] The system 100 may include a video communication service 110, which in some embodiments may be external to the system 100. As described below, multiple video communication services 110 may be included.

[0028] The system 100 may include one or more participant clients 121, although one, some, or all of the participant clients 121 may be external to the system 100 in some embodiments.

[0029] The system 100 includes a central server 130 .

[0030] As used herein, the term "central server" refers to a computer-implemented function configured to be accessible in a logically centralized manner, such as through a well-defined API (Application Programming Interface). Such central server functionality may be implemented purely in computer software, or in a combination of software and virtual and / or physical hardware, in a standalone physical or virtual server computer, or distributed across multiple interconnected physical and / or virtual server computers.

[0031] The physical or virtual hardware on which central server 130 runs, in other words the computer software that defines the functionality of central server 130, may be comprised of a per se conventional CPU, a per se conventional GPU, per se conventional RAM / ROM memory, per se conventional computer buses, and per se conventional external communication capabilities, such as an Internet connection.

[0032] The video communication service 110 , in so far as that is used, is also a central server in the above sense, which may be a different central server from the central server 130 or may be part of the central server 130 .

[0033] Correspondingly, each of the participant clients 121 may be a central server in the above sense, in a corresponding interpretation, in which the physical or virtual hardware on which each participant client 121 runs, in other words the computer software defining the functionality of the participant client 121, comprises a CPU / GPU per se, conventional RAM / ROM memory per se, a computer bus per se, and external communication capabilities per se conventional, such as an Internet connection.

[0034] Each participant client 121 also typically comprises, or is in communication with, a computer screen arranged to display video content that is provided to the participant client 121 as part of an ongoing video communication, speakers arranged to emit sound content that is provided to the participant client 121 as part of the video communication, a video camera, and a microphone arranged to record sound local to a human participant 122 to the video communication, who uses that participant client 121 to participate in the video communication.

[0035] In other words, the human-machine interface of each participant client 121 enables each participant 122 to interact with other participants in a video communication and / or with audio / video streams provided from various sources at that client 121.

[0036] Typically, each participant client 121 comprises a respective input means 123, which may consist of said video camera, said microphone, a keyboard, a computer mouse or trackpad, and / or an API for receiving digital video streams, digital audio streams and / or other digital data. The input means 123 are in particular configured to receive video and / or audio streams from the video communication service 110 and / or a central server, such as the central server 130, such video and / or audio streams being provided as part of the video communication and preferably generated based on corresponding digital data input streams provided to said central server from at least two sources of such digital data input streams, e.g. the participant client 121 and / or an external source (described below).

[0037] More generally, each participant client 121 includes a respective output means 124 that may consist of the computer screen, the speakers, and an API that emits digital video and / or audio streams that are representative of the locally captured video and / or audio to a participant 122 using that participant client 121.

[0038] In practice, each participant client 121 may be a mobile device, such as a mobile phone, equipped with a screen, speakers, microphone, and Internet connection, running computer software locally or accessing remotely executed computer software to perform the functions of that participant client 121. Correspondingly, a participant client 121 may be a thick or thin laptop or stationary computer, running locally installed applications and also using functions accessed remotely via a web browser.

[0039] There may be one or more participant clients 121, for example at least three or at least four, used in one and the same video communication in this embodiment.

[0040] There may be at least two different groups of participating clients. Each of the participating clients may be assigned to a respective such group. The groups may reflect different roles of the participating clients, different virtual or physical locations of the participating clients, and / or different interaction rights of the participating clients.

[0041] A variety of such roles are available and may be, for example, "leader" or "conference host", "speaker", "panelist", "interactive audience", "remote listener".

[0042] Such physical locations may vary widely and may be, for example, "on stage," "in a panel," "physically present audience," or "physically distant audience."

[0043] A virtual location may be defined in terms of a physical location, but may also include virtual groupings that may overlap with the physical locations, for example, physically present audience participants may be divided into a first virtual group and a second virtual group, with some physically present audience participants grouped together with some physically remote audience participants in the same virtual group.

[0044] Such interaction permissions can be varied and may be, for example, "full interaction" (no restrictions), "can speak, but only after requesting the microphone" (e.g., raising a virtual hand in a video conferencing service), "cannot speak, but can write in a common chat", or "only watch / listen".

[0045] In some implementations, each defined role and / or physical / virtual location may be defined with respect to certain predefined interaction permissions. In other examples, all participants with the same interaction permissions form a group. Thus, the defined roles, locations, and / or interaction permissions may reflect various group assignments, and different groups may differ from or overlap with each other as desired.

[0046] This is exemplarily explained below.

[0047] The video communications may be provided at least in part by the video communications service 110 and at least in part by the central server 130, as described and illustrated herein.

[0048] As the term is used herein, a "video communication" is a two-way digital communication session that includes at least two, and preferably at least three or at least four, video streams, preferably used to generate one or more mixed or collaborative digital video / audio streams, and also coincident audio streams, that are consumed by one or more consumers (e.g., participant clients of the types described above) who may or may not contribute to the video communication via video and / or audio. Such video communication may be in real time, with or without a certain latency or delay. At least one, and preferably at least two, or at least four participants 122 who participate in such a video communication are involved in the video communication in a two-way manner, providing and consuming video / audio information.

[0049] At least one of the participant clients 121, or all of the participant clients 121, includes a local synchronization software function 125, which is described in more detail below.

[0050] The video communication services 110 may be provided with or have access to a common time reference, as described in more detail below.

[0051] Each of the at least one central server 130 may include an API 137 for digitally communicating with entities external to that central server 130. Such communications may include both inputs and outputs.

[0052] The system 100, such as the central server 130, may be configured to digitally communicate with an external information source 300, such as an externally provided video stream, and in particular to receive digital information, such as audio and / or video stream data, from the external information source 300. By "external," the information source 300, it is meant that it is not provided by or as part of the central server 130. Preferably, the digital data provided by the external information source 300 is independent of the central server 130, and the central server 130 cannot influence its information content. For example, the external information source 300 may be live captured video and / or audio, such as a public sporting event or an ongoing news event or report. Also, the external information source 300 may be captured by a webcam or the like, rather than by any of the participant clients 121. Thus, such captured video may depict the same locality as any one of the participant clients 121, but is not captured as part of the participant client 121's activities. One possible difference between the externally provided information source 300 and the internally provided information source 120 is that the internally provided information source may be provided in that capacity as a participant in a video communication of the type defined above, whereas the externally provided information source 300 is not, but instead is provided as part of a context that is external to the video conference.

[0053] There may also be multiple external information sources 300 providing digital information of that type, such as audio and / or video streams, in parallel to the central server 130 .

[0054] As shown in FIG. 1, each participant client 121 constitutes the source of a respective information (video and / or audio) stream 120 that is provided by that participant client 121 to the video communication service 110, as described.

[0055] The system 100, such as the central server 130, may be further configured to digitally communicate with the external consumers 150, and in particular to emit digital information to the external consumers 150. For example, digital video and / or audio streams generated by the central server 130 may be provided continuously, in real time or near real time, to one or more external consumers 150 via the API 137 described above. Again, the consumer 150 being "external" means that the consumer 150 is not provided as part of the central server 130 and / or is not a party to the video communication in question.

[0056] Unless otherwise noted, all functions and communications herein are provided digitally and electronically, implemented by computer software running on appropriate computer hardware and communicated over a digital communications network or channel, such as the Internet.

[0057] 1, multiple participant clients 121 participate in a digital video communication provided by the video communication service 110. Each participant client 121 therefore has an ongoing login, session, or the like to the video communication service 110 and can participate in one and the same ongoing video communication provided by the video communication service 110. In other words, the video communication is "shared" among the participant clients 121 and therefore also by the corresponding human participants 122.

[0058] 1 , the central server 130 comprises an auto-join client 140, which is an auto-client that corresponds to the participant client 121, but is not associated with a human participant 122. Instead, the auto-join client 140 is added as a participant client to the video communication service 110 to participate in the same shared video communication as the participant client 121. As such a participant client, the auto-join client 140 is given access to continuously generated digital video and / or audio stream(s) provided as part of an ongoing video communication by the video communication service 110, and such streams can be consumed by the central server 130 via the auto-join client 140. Preferably, the automatic participant client 140 receives from the video communication service 110 a common video and / or audio stream that is distributed or can be distributed to each participant client 121; respective video and / or audio streams that are provided from each of one or more participant clients 121 to the video communication service 110 and relayed by the video communication service 110 to all participant clients 121 or to requesting participant clients 121 in raw or modified form; and / or a common time reference.

[0059] The central server 130 may include a collection function 131 configured to receive a plurality of video and / or audio streams of the above types from the auto-participant clients 140, and possibly also from the above-mentioned external information source(s) 300, for processing as described below, and then provide a generated video stream, e.g., a shared video stream, via the API 137. For example, this generated video stream may be consumed by external consumers 150 and / or by the video communication service 110, which may distribute it to all or any requesting ones of the participant clients 121.

[0060] FIG. 2 is similar to FIG. 1, but instead of using an auto-join client 140, the central server 130 receives video and / or audio stream data from an ongoing video communication via the API 112 of the video communication service 110.

[0061] 3 is similar to FIG. 1, but the video communication service 110 is not shown. In this case, the participant clients 121 communicate directly with the API 137 of the central server 130, for example, to provide video and / or audio stream data to the central server 130 and / or to receive video and / or audio stream data from the central server 130. The generated shared streams may then be provided to external consumers 150 and / or to one or more of the client participants 121.

[0062] 4 shows the central server 130 in more detail. As shown, the collection function 131 may be composed of one or, preferably, multiple, format-specific collection functions 131a. Each of the format-specific collection functions 131a is configured to receive video and / or audio streams having a predefined format, such as a predefined binary encoding format and / or a predefined stream data container, and in particular to parse and classify the binary video and / or audio data of said format into individual video frames, sequences of video frames and / or time slots.

[0063] The central server 130 further comprises an event detection function 132 configured to receive video and / or audio stream data, such as binary stream data, from the collection function 131 and perform respective event detection on each individual one of the received data streams. The event detection function 132 may comprise an AI (artificial intelligence) component 132a for performing event detection. The event detection may be performed without first time synchronizing the collected individual streams.

[0064] The central server 130 further comprises a synchronization function 133 configured to time-synchronize multiple data streams provided by the collection function 131 and that may be processed by the event detection function 132. The synchronization function 133 may comprise an AI component 133a for performing the time synchronization.

[0065] The central server 130 may further comprise a pattern detection function 134 configured to perform pattern detection based on a combination of at least one, but often at least two, for example at least three or at least four, for example all, of the received multiple data streams. The pattern detection may further be based on one, possibly at least two or more events detected by the event detection function 132 for each individual one of said multiple data streams. Such detected events considered by the pattern detection function 134 may be distributed over time for the individual collected streams. The pattern detection function 134 may comprise an AI component 134a for performing pattern detection. The pattern detection may further be based on the groupings mentioned above and may in particular be configured to detect a specific pattern occurring only for one group, to detect a specific pattern occurring only for some groups but not all groups, or to detect a specific pattern occurring for all groups.

[0066] The central server 130 further comprises a generating function 135 configured to generate a generated digital video stream, e.g., a shared digital video stream, based on the multiple data streams provided from the collecting function 131 and possibly based on any detected events and / or patterns. The generated video stream includes at least a generated video stream comprising one or more of the raw data, reformatted or converted video streams provided by the collecting function 131, and may include corresponding audio stream data. As exemplified below, there may be multiple generated video streams, one of which may be generated in the manner described above, but may also be generated based on another already generated video stream.

[0067] All generated video streams are preferably generated continuously, and preferably in near real-time (after subtracting latency and delays of the type described later herein).

[0068] The central server 130 may further include a publishing function 136 configured to publish the generated shared digital video stream, such as via the API 137 described above.

[0069] It should be noted that while Figures 1, 2 and 3 show three different examples of how the central server 130 can be used to implement the principles described herein, and in particular to provide methods in accordance with the present invention, other configurations are possible, either with or without one or more video communication services 110.

[0070] Thus, Figure 5 illustrates a method for providing the generated digital video stream. Figures 6a to 6f show the states of the different digital video / audio data streams resulting from the method steps shown in Figure 5.

[0071] In a first step, the method begins.

[0072] In a subsequent collection step, a respective plurality of primary digital video streams 210, 301 are collected, for example by a collection function 131, from at least two of said digital video sources 120, 300. Each such plurality of primary data streams 210, 301 may comprise an audio portion 214 and / or a video portion 215. It is understood that "video" in this context denotes the moving image and / or still image content of such data streams. Each primary data stream 210, 301 may be encoded according to any video / audio encoding standard (using the respective codec used by the entity providing said primary stream 210, 301), and the encoding format may differ between different ones of said plurality of primary streams 210, 301 used simultaneously in one and the same video communication. At least one, for example all, of the plurality of primary data streams 210, 301 are preferably provided as a stream of binary data, possibly in a data container data structure that is itself conventional. Preferably, at least one, such as at least two, or even all, of the multiple primary data streams 210, 301 are provided as respective live video recordings.

[0073] It should be noted that the multiple primary data streams 210, 301 may not be synchronized in time when they are received by the collection function 131. This may mean that they are associated with different latencies or delays relative to each other. For example, if two primary video streams 210, 301 are live recordings, this may mean that they are associated with different latencies relative to the recording time when they are received by the collection function 131.

[0074] It should also be noted that the multiple primary data streams 210, 301 may themselves be respective live camera feeds from webcams; a screen or presentation currently being shared; a film clip being viewed; or any combination of these arranged in various ways within one and the same screen.

[0075] The collection steps are illustrated in Figures 6a and 6b. Figure 6b also illustrates how the collection function 131 can store each primary video stream 210, 301 as bundled audio / video information or as audio stream data separated from the associated video stream data. Figure 6b illustrates how the data of the primary video streams 210, 301 is stored as individual frames 213 or collections / clusters of frames, where a "frame" refers here to a time-limited portion of image data and / or any associated audio data, e.g., each frame being an individual still image or a continuous series of images (e.g., a series of images that constitutes up to one second of moving images) that together form the moving image video content.

[0076] In a subsequent event detection step performed by the event detection functionality 132, the multiple primary digital video streams 210, 301 are analyzed by the event detection functionality 132, particularly the AI ​​component 132a etc., to detect at least one event 211 selected from the first set of events. This is illustrated in Figure 6c.

[0077] This event detection step is preferably performed for at least one, e.g. at least two, e.g. all, of the primary video streams 210, 301, and individually for each of said primary video streams 210, 301. In other words, the event detection step is preferably performed for each of said individual primary video streams 210, 301, taking into account only information contained as part of that particular primary video stream 210, 301, and in particular without taking into account information contained as part of other primary video streams. Furthermore, event detection is preferably performed without taking into account any common time reference 260 associated with multiple primary video streams 210, 301.

[0078] However, preferably, event detection takes into account information contained as part of the individually analyzed primary video stream over a time interval, for example over a historical time interval of the primary video stream that is greater than 0 seconds, for example at least 0.1 seconds, for example at least 1 second.

[0079] Event detection may take into account information contained in the audio and / or video data included as part of the primary video stream 210,301.

[0080] The first set of events may include any number of types of events, such as a change in a slide in a slide presentation that constitutes or is part of the primary video stream 210, 301, a change in connection quality of the source 120, 300 providing the primary video stream 210, 301 that results in a change in image quality, loss of image data, or reacquisition of image data, and physical events of movement detected in the primary video stream 210, 301, such as movement of a person or object in the video, a change in lighting in the video, a sudden sharp noise in the audio, or a change in audio quality. It should be understood that this is not intended to be an exhaustive list, and these examples are provided to understand the applicability of the presently described principles.

[0081] In a subsequent synchronization step performed by the synchronization function 133, the primary digital video streams 210 are time-synchronized. This time synchronization may be performed with respect to a common time reference 260. As shown in Fig. 6d, this time synchronization may include aligning the primary video streams 210, 301 with respect to each other using, for example, the common time reference 260, so that they can be combined to form a time-synchronized context. The common time reference 260 may be a stream of data, a heartbeat signal or other pulse data, or a time anchor that is applicable to each of the individual primary video streams 210, 301. By making the common time reference applicable to each of the individual primary video streams 210, 301, the information content of the primary video streams 210, 301 can be uniquely related to the common time reference with respect to a common time axis. In other words, the common time reference aligns the primary video streams 210, 301 to be time-synchronized in a present sense via time shifting. In other embodiments, time synchronization may be based on known information about the time difference between the primary video streams 210, 301, such as measurements.

[0082] As shown in FIG. 6d, time synchronization may include determining one or more timestamps 261 for each of multiple primary video streams 210, 301, e.g., relative to a common time reference 260, or for each video stream 210, 301 relative to the other video streams 210, 301 or relative to other multiple video streams 210, 301.

[0083] In a subsequent pattern detection step performed by pattern detection function 134, the multiple time synchronized primary digital video streams 210, 301 are analyzed to detect at least one pattern 212 selected from the first pattern set. This is shown in Figure 6e.

[0084] In contrast to the event detection step, the pattern detection step is preferably performed based on video and / or audio information included as part of at least two of the multiple time-synchronized primary video streams 210,301.

[0085] The first set of patterns may include any number of types of patterns, such as multiple participants speaking in turn or simultaneously, or a change in a presentation slide occurring simultaneously as another event, such as another participant speaking, etc. This list is not exhaustive but is exemplary.

[0086] In alternative embodiments, the detected pattern 212 may relate to information contained in only one of the multiple primary video streams 210, 301, rather than information contained in more than one of the multiple primary video streams 210, 301. In such cases, such pattern 212 is preferably detected based on video and / or audio information contained in that single primary video stream 210, 301 spanning at least two detected events 211, e.g., two or more consecutive detected presentation slide changes or connection quality changes. As an example, multiple consecutive slide changes that rapidly follow one another over time may be detected as one single slide change pattern, as opposed to one distinct slide change pattern for each detected slide change event.

[0087] It is understood that the first set of events and the first set of patterns may comprise a predefined type of event / pattern defined using a respective set of parameters and parameter intervals. As described below, the set of events / patterns may also be defined and detected using various AI tools.

[0088] In a subsequent generation step performed by the generation function 135, a shared digital video stream is generated as an output digital video stream 230 based on a plurality of consecutively considered frames 213 of a plurality of time-synchronized primary digital video streams 210, 301 and the detected pattern 212.

[0089] As explained and detailed below, the present invention allows for fully automated generation of video streams, such as output digital video stream 230.

[0090] For example, such generation may include selection of what video and / or audio information from which primary video streams 210, 301 to use in the output video stream 230, and to what extent; the video screen layout of the output video stream 230; the switching pattern between different such uses or layouts over time; etc.

[0091] This is also illustrated in Figure 6f, which shows one or more additional portions of time-related (which may be relative to the common time reference 260) digital video information 220, such as additional digital video information streams that may be time-synchronized (e.g., relative to a common time reference 260) and used in concert with the time-synchronized multiple primary video streams 210, 301 in generating the output video stream 230. For example, the additional streams 220 may include information regarding any video and / or audio special effects to use, such as dynamically based on detected patterns; a planned time schedule for the video communication; etc.

[0092] In a subsequent publishing step performed by the publishing function 136, the generated output digital video stream 230 is continuously provided to the consumers 110, 150 of the shared digital video stream, as described above. The generated digital video stream may be provided to one or more participant clients 121, for example via the video communication service 110.

[0093] In the subsequent steps, the method ends. However, initially, the method may be repeated any number of times to generate the output video stream 230 as a continuously provided stream, as shown in FIG. 5. Preferably, the output video stream 230 is generated to be consumed in real-time or near real-time (taking into account the sum of the latencies added by all steps along the way) and continuously (published as soon as more information becomes available, but not counting the intentionally added latencies described below). In this way, the output video stream 230 may be consumed in an interactive manner, whereby the output video stream 230 is fed back to the video communication service 110 or to other contexts that form the basis for the generation of the primary video stream 210 that is fed back to the collection function 131 to form a closed feedback loop; or the output video stream 230 is consumed in a different context (outside the system 100, or at least outside the central server 130), where it may form the basis for real-time two-way video communication.

[0094] As mentioned above, in some embodiments, at least two, e.g., at least three, e.g., at least four or at least five of the multiple primary digital video streams 210, 301 are provided as part of a shared digital video communication as provided by a video communication service 110, which video communication includes respective remotely connected participant clients 121 providing said primary digital video streams 210. In such a case, the collecting step may consist of collecting at least one of said primary digital video streams 210 from the shared digital video communication service 110 itself, via an auto-participant client 140 that is in turn granted access to the video and / or audio stream data from within said video communication service 110, and / or via the API 112 of the video communication service 110.

[0095] Additionally, in this and other cases, the collecting step may comprise collecting at least one of said plurality of primary digital video streams 210, 301 as a respective external digital video stream 301 collected from an information source 300 that is external to the shared digital video communication service 110. It should be noted that one or more of such external video sources 300 may be external to the central server 130.

[0096] In some embodiments, the multiple primary video streams 210, 301 are not formatted in the same manner. Such different formats could be the formats in which they are provided to the collection function 131 in different types of data containers (such as AVI or MPEG), but in preferred embodiments, at least one of the multiple primary video streams 210, 301 is formatted according to a deviating format (with respect to at least one other of the primary video streams 210, 301) in that the deviating primary digital video streams 210, 301 have deviating video encodings; deviating fixed or variable frame rates; deviating aspect ratios; deviating video resolutions; and / or deviating audio sample rates.

[0097] The collection function 131 is preferably pre-configured to read and interpret all encoding formats, container standards, etc. occurring in all collected primary video streams 210, 301. This allows processing as described herein to be performed without requiring decoding until a relatively later stage in these processes (such as until the primary streams in question are in their respective buffers; or until after the event detection step; etc.). However, in the rare case where one or more of the primary video feeds 210, 301 are encoded using a codec that the collection function 131 cannot interpret without decoding, the collection function 131 may be configured to perform decoding and analysis of such primary video streams 210, 301, followed by conversion to a format that can be processed, for example, by the event detection function. Note that even in this case, it is preferable not to perform re-encoding at this stage.

[0098] For example, a primary video stream 220 fetched from a multi-party video event, such as that provided by the video communication service 110, typically has a requirement for low latency and is therefore typically associated with variable frame rates and variable pixel resolutions to enable participants 122 to communicate effectively. In other words, the overall video and audio quality is degraded as necessary for low latency.

[0099] On the other hand, the external video feed 301 typically has a more stable frame rate and higher image quality, but may therefore have a higher delay.

[0100] Thus, the video communication service 110 may at each point in time use a different encoding and / or container than the external video source 300. Thus, the analysis and video generation process described herein must then combine these multiple streams 210, 301 of different formats into a new single stream for a combined experience.

[0101] As mentioned above, the collection functionality 131 may comprise a set of format-specific collection functionality 131a, each configured to process a particular type of format of the primary video streams 210, 301. For example, each one of these format-specific collection functionality 131a may be configured to process multiple primary video streams 210, 301 encoded using different respective video encoding methods / codecs, such as Windows® Media® or DivX®.

[0102] However, in a preferred embodiment, the collecting step involves converting at least two, eg, all, of the multiple primary digital video streams 210 , 301 to a common protocol 240 .

[0103] As used in this context, the term "protocol" refers to an information structuring standard or data structure that specifies how to store the information contained in the digital video / audio stream. However, the common protocol preferably does not prescribe how to store the digital video and / or audio information, for example at a binary level (i.e., encoded / compressed data that indicates the sounds and images themselves), but instead forms a structure of a predefined format for storing such data. In other words, the common protocol prescribes storing digital video data in a raw binary format without performing any digital video decoding or encoding in connection with such storage, and possibly without modifying the existing binary format in any way apart from concatenating and / or splitting the binary format byte strings. Instead, the raw (encoded / compressed) binary data content of said primary video stream 210, 301 is preserved while repacking said raw binary data in a data structure defined by the protocol. In some embodiments, the common protocol defines a video file container format.

[0104] FIG. 7 shows, by way of example, multiple primary video streams 210, 301 as shown in FIG. 6a reconstructed by respective format-specific acquisition functions 131a and using the common protocol 240 described above.

[0105] Thus, the common protocol 240 provides for storing digital video and / or audio data in data sets 241 that are preferably divided into discrete, contiguous sets of data along a time axis relative to the primary video stream in question 210, 301. Each such data set may contain one or several frames of video and associated audio data.

[0106] The common protocol 240 may also provide for storing, in association with the stored digital video and / or audio data set 241, metadata 242 associated with a specified point in time.

[0107] The metadata 242 may include information about the binary format of the raw data of the primary digital video stream 210, such as about the digital video encoding method or codec used to generate the binary data of the raw data, the resolution of the video data, the video frame rate, a frame rate variation flag, the video resolution, the video aspect ratio, the audio compression algorithm, or the audio sampling rate. The metadata 242 may also include information about timestamps of the stored data, e.g. related to the time base of the primary video stream 210, 301 itself or related to the different video streams mentioned above.

[0108] The use of format-specific collection functions 131a in combination with the common protocol 240 allows for rapid collection of the information content of the primary video streams 210, 301 without the added latency (delay) of decoding / re-encoding the received video / audio data.

[0109] The collecting step may thus comprise collecting a plurality of primary digital video streams 210, 301 encoded using different binary video and / or audio encoding formats using different ones of the plurality of format specific collecting functions 131a in order to parse said primary video streams 210, 301 and store the parsed raw binary data, together with any associated metadata, in a data structure using a common protocol. Obviously, the decision as to which format specific collecting function 131a to use for which primary video stream 210, 301 may be performed by the collecting function 131 based on predefined and / or dynamically detected characteristics of each of said primary video streams 210, 301.

[0110] Each primary video stream 210, 301 collected in this manner may be stored in its own separate memory buffer, such as a RAM memory buffer within the central server 130.

[0111] The conversion of the primary video streams 210, 301 performed by each format-specific collection function 131a may therefore comprise splitting the raw binary data of each primary digital video stream 210, 301 thus converted into an ordered set of smaller data sets 241.

[0112] Furthermore, the conversion may also comprise associating each of the smaller sets 241 (or a subset, e.g. a subset regularly distributed along the time axis of each of the primary streams 210, 301 in question) with a respective time along a shared time axis, e.g. with respect to a common time reference 260. This association may be performed by analysis of the binary video and / or audio data of the raw data, in any of the principle methods described below or otherwise, and may be performed in order to enable subsequent time synchronization of the primary video streams 210, 301 to be performed. Depending on the type of common time reference 260 used, at least a part of this association of each data set 241 may also be performed by or instead of the synchronization function 133. In the latter case, the collecting step may instead comprise associating each of the smaller sets 241, or a subset thereof, with a respective time of a time axis specific to the primary stream 210, 301 in question.

[0113] In some embodiments, the collecting step also includes converting the raw binary video and / or audio data collected from the multiple primary video streams 210, 301 to a uniform quality and / or updating frequency. This may include downsampling or upsampling the raw binary digital video and / or audio data of the multiple primary digital video streams 210, 301 to a common video frame rate; a common video resolution; or a common audio sampling rate, as appropriate. It should be noted that such resampling can be performed without performing a full decoding / re-encoding, or even without performing any decoding at all, since the format-specific collecting function 131a can directly process the raw binary data according to the correct binary encoding target format.

[0114] Preferably, each of the multiple primary digital video streams 210, 301 is stored in an individual data storage buffer 250 as an individual frame 213 or a sequence of frames 213, as described above, and each is associated with a corresponding timestamp that is in turn associated with a common time reference 260.

[0115] In a specific example provided for illustrative purposes, the video communication service 110 is Microsoft® Teams® and is conducting a video conference involving multiple simultaneous participants 122. The auto-join client 140 is registered as a conference participant in the Teams® conference.

[0116] The primary video input signals 210 are then provided to the collection function 130 via the auto-join client 140 and are acquired by the collection function 130. These are raw data signals in H264 format and include timestamp information for each video frame.

[0117] The associated format-specific collection function 131a picks up the raw data over IP (cloud LAN network) on a configurable predefined TCP port. Every Teams® meeting participant and associated audio data is associated with a separate port. The collection function 131 then uses the timestamp from the audio signal (50Hz) and downsamples the video data to a fixed output signal of 25Hz before storing the video streams 220 in their respective individual buffers 250.

[0118] As mentioned above, the common protocol 240 stores data in a raw binary format. It can be designed to process, at a very low level, the raw bits and bytes of video / audio data. In a preferred embodiment, the data is stored in the common protocol 240 as a simple byte array or corresponding data structure (such as a slice). This means that the data does not need to be put into a traditional video container at all (the common protocol 240 does not constitute such a traditional container in this context). Also, video encoding and decoding is computationally heavy, thus inducing delays and requiring expensive hardware. Moreover, this problem scales with the number of participants.

[0119] The common protocol 240 allows for reserving memory in the collection function 131 for the primary video stream 210 associated with each Teams® conference participant 122 and any external video sources 300, and changing the amount of allocated memory on the fly during the process. In this way, the number of input streams can be changed, so that each buffer can remain valid. For example, information such as resolution, frame rate, etc., is variable, but is stored as metadata in the common protocol 240, so that this information can be used to quickly change the size of each buffer as needed.

[0120] The following is an example of a specification for this type of common protocol 240:

[0121] [Table 1]

[0122] In the above table, the "Detected event in, if any" data is included as part of the specification of common protocol 260. However, in some embodiments, this information (regarding detected events) may instead be placed in a separate memory buffer.

[0123] In some embodiments, the at least one additional portion of the digital video information 220, which may be an overlay or effect, is also stored in each individual buffer 250 as individual frames or sequences of frames each associated with a corresponding timestamp that is in turn associated with the common time reference 260.

[0124] As illustrated above, the event detection step may include using a common protocol 240 to store metadata 242 describing the detected event 211 in association with the primary digital video stream 210, 301 in which the event 211 was detected.

[0125] Event detection can be performed in different ways. In some embodiments performed by the AI ​​component 132a, the event detection step includes a first trained neural network or other machine learning component individually analyzing at least one, e.g., some or all, of the multiple primary digital video streams 210, 301 to automatically detect any of said events 211. This may include the AI ​​component 132a classifying the data of the primary video streams 210, 301 into a set of predefined events in a supervised classification and / or into a dynamically determined set of events in an unsupervised classification.

[0126] In some embodiments, the detected event 211 is a change in a presentation slide of a presentation that is, or is contained in, the primary video stream 210, 301.

[0127] For example, if a presenter of a presentation decides to change the slide of the presentation that he or she is currently making to the audience, this means that what is interesting for a given viewer may change. The newly displayed slide may just be a general-level image that is best viewed for a short time in a so-called "butterfly" mode (e.g., displaying the slide side-by-side with the video of the presenter in the output video stream 230). Or the slide may contain a lot of detail, text with a small font size, etc. In the latter case, the slide will be displayed full screen and will be displayed for a somewhat longer time than would normally be the case. The slide in this case may be more interesting to the viewer of the presentation than the face of the presenter, so butterfly mode may not be as appropriate.

[0128] In practice, the event detection step consists of at least one of the following:

[0129] Firstly, the event 211 may be detected based on image analysis of the difference between a first image of a detected slide and a subsequent second image of the detected slide. The nature of the primary video stream 220, 301 as being indicative of a slide may be determined automatically using digital image processing which is per se conventional, such as using motion detection combined with OCR (Optical Character Recognition).

[0130] This may involve using automatic computer image processing techniques to check whether the detected slide has changed sufficiently to be classified as a real slide change. This can be done by checking the delta between the current slide and the previous slide in terms of RGB color values. For example, one can evaluate how much the RGB values ​​have changed globally in the screen area covered by the slide in question, and at the same time evaluate whether it is possible to find groups of adjacent pixels that change in concert with this. In this way, relevant slide changes can be detected, while filtering out irrelevant changes, such as, for example, computer mouse movements across the screen. This approach allows for full composability. For example, it may be desirable to be able to capture computer mouse movements, for example if a presenter wants to present something in detail while pointing to different things with the computer mouse.

[0131] Second, the event 211 may be detected based on image analysis of the information complexity of the second image itself to determine the type of event with greater specificity.

[0132] This might involve, for example, assessing the amount of textual information on the slide in question and the associated font sizes. This can be done using traditional OCR methods, including deep learning-based character recognition techniques.

[0133] Note that because the raw binary format of the evaluated video streams 210, 301 is known, this may be performed directly in the binary domain without first decoding or re-encoding the video data. For example, the event detection function 132 may invoke an associated format-specific collection function for an image interpretation service, or the event detection function 132 itself may include functionality for evaluating image information, such as to the individual pixel level, for a number of different supported raw binary video data formats.

[0134] In another example, the detected event 211 is a loss of a communication connection of a participant client 121 to the digital video communication service 110. In this case, the detecting step may include detecting that the participant client 121 has lost the communication connection based on image analysis of a series of subsequent video frames 213 of the primary digital video stream 210 corresponding to that participant client 121.

[0135] Because participant clients 121 are associated with different physical locations and different Internet connections, it may occur that someone loses connection to the video communication service 110 or the central server 130. In such a situation, it is desirable to avoid a black or blank screen appearing in the generated output video stream 230.

[0136] Alternatively, such loss of connection can be detected as an event by the event detection function 132, for example by applying a two-class classification algorithm where the two classes used are connected / not connected (no data). In this case, "no data" is understood to be different from the presenter intentionally sending a black screen. Since a short duration black screen, such as just one or two frames, may not be noticeable in the final generated stream 230, the two-class classification algorithm can be applied over time to create a time series. A threshold specifying the minimum length of a connection interruption can then be used to determine whether a connection has been lost.

[0137] As described below, detected events of the types illustrated above may be used by pattern detection function 134 to take various responses, as appropriate and desired.

[0138] As noted above, the individual primary video streams 210, 301 are each associated with a common time reference 260 and can be time-synchronized relative to one another by the synchronization function 133.

[0139] In some embodiments, the common time reference 260 is based on or comprises a common audio signal 111 (see Figures 1 to 3), which, as described above, is common to a shared digital video communication service 110 participating in at least two remotely connected participant clients 121, each providing a respective one of the primary digital video streams 210.

[0140] In the Microsoft® Teams® example discussed above, a common audio signal may be generated and captured by the central server 130 via the auto-join client 140 and / or via the API 112. In this and other examples, such a common audio signal may be used as a heartbeat signal to time-synchronize the individual primary video streams 220 by combining the individual primary video streams 220 at specific times based on the heartbeat signal. Such a common audio signal may be provided as a separate (with respect to each of the other primary video streams 210) signal, such that each of the other primary video streams 210 may be individually time-correlated to the common audio signal based on audio contained in the other primary video streams 210 or based on image information contained therein (such as using automatic image processing-based lip-sync techniques).

[0141] In other words, to handle the variable and / or different latencies associated with the individual primary video streams 210 and to achieve time synchronization of the combined video output stream 230, such a common audio signal is used as a heartbeat for all primary video streams 210 within the central server 130 (but perhaps not the external primary video stream 301). In other words, all other signals are mapped to this common audio time heartbeat to make sure they are all time synchronized.

[0142] In another example, time synchronization is achieved using a time synchronization element 231 that is introduced into the output digital video stream 230 and detected by a respective local time synchronization software function 125 provided as part of one or more individual ones of the participant clients 121. The local software function 125 is configured to detect the time of arrival of the time synchronization element 231 in the output video stream 230. As will be appreciated, in such an embodiment, the output video stream 230 is fed back to the video communication service 110 or otherwise made available to each participant client 121 and its local software function 125.

[0143] For example, the time synchronization elements 231 may be visual markers, such as pixels that change color in a predetermined order or manner, that are placed or updated in the output video 230 at regular time intervals; a visual clock that is updated and displayed in the output video 230; an audio signal (which may be designed to be inaudible to the participants 122, for example, by having a sufficiently low amplitude and / or a sufficiently high frequency) that is added to the audio that forms part of the output video stream 230. The local software functionality 125 is configured to automatically detect the arrival time of each of the time synchronization elements (of each) 231 using appropriate image and / or audio processing.

[0144] The common time reference 260 may then be determined based, at least in part, on the detected arrival times. For example, each of the local software functions 125 may communicate respective information indicative of the detected arrival times to the central server 130.

[0145] Such communication may occur via a direct communication link between the participant client 121 and the central server 130. However, communication may also occur via a primary video stream 210 associated with the participant client 121. For example, the participant client 121 may introduce a visual or audible code, such as the type described above, into the primary video stream 210 generated by the participant client 121 for automatic detection by the central server 130 and use to determine the common time reference 260.

[0146] In yet further embodiments, each participant client 121 may perform image detection in a common video stream viewable by all participant clients 121 for the video communication service 110 and relay the results of such image detection to the central server 130 in a manner corresponding to that described above, where they are used to determine the respective offsets of each participant client 121 relative to one another over time. In this manner, a common time reference 260 may be determined as a set of individual relative offsets. For example, a selected reference pixel of the commonly available video stream may be monitored by some or all of the participating clients 121, such as by local software functions 125, and the current color of that pixel may be communicated to the central server 130. The central server 130 may generate an estimated set of relative time offsets across the different participating clients 121 by calculating respective time series based on such color values ​​received successively from each of many (or all) of the participating clients 121 and performing cross-correlation.

[0147] In practice, the output video stream 230 provided to the video communication service 110 may be included as part of the shared screen of all participant clients of that video communication and may therefore be used to evaluate such time offsets associated with the participant clients 121. In particular, the output video stream 230 provided to the video communication service 110 may be made available again to the central server via the auto-join client 140 and / or API 112.

[0148] In some embodiments, the common time reference 260 may be determined at least in part based on a detected discrepancy between an audio portion 214 of a first one of the multiple primary digital video streams 210, 301 and an image portion 215 of said first one of the multiple primary digital video streams 210, 301. Such a discrepancy may for example be based on a digital lip-sync video image analysis of a speaking participant 122 viewed in said first primary digital video stream 210, 301. Such a lip-sync analysis may be conventional per se and may for example use a trained neural network. The analysis may be performed by the synchronization function 133 for each primary video stream 210, 301 in relation to the available common audio information, and a relative offset across the individual primary video streams 210, 301 may be determined based on this information.

[0149] In some embodiments, the synchronization step includes intentionally introducing a delay (in this context "delay" and "latency" are intended to mean the same thing) of up to 30 seconds, e.g. up to 5 seconds, e.g. up to 1 second, e.g. up to 0.5 seconds, but more than 0 seconds, to provide at least that delay in the output digital video stream 230. Whatever the length, the intentionally introduced delay is at least a number of video frames, such as at least 3, or at least 5, or even 10, such as this number of frames (or individual images) stored after any resampling in the acquisition step. As used herein, the term "intentionally" means that the delay is introduced independently of the need to introduce such a delay based on synchronization issues or the like. In other words, the intentionally introduced delay is in addition to the delay introduced as part of the synchronization of the primary video streams 210, 301, in order to time-synchronize the multiple primary video streams 210, 301 with each other. The intentionally introduced delay may be predetermined, fixed, or variable with respect to the common time reference 260. The delay time may be measured relative to the least latent one of the multiple primary video streams 210, 301, and as a result of the above time synchronization, the more latent ones of these streams 210, 301 may be associated with a relatively small intentionally added delay.

[0150] In some embodiments, a relatively small delay, such as 0.5 seconds or less, is introduced that is barely noticeable to participants in the video communication service 110 using the output video stream 230. In other embodiments, a larger delay may be introduced, such as when the output video stream 230 is not used in an interactive context, but instead is exposed in a one-way communication to an external consumer 150.

[0151] This intentionally introduced delay may be sufficient to allow the synchronization function 133 sufficient time to map the collected video frames of the individual primary streams 210, 301 to the correct common time reference 260 timestamps 261. It may also be sufficient to provide sufficient time to perform the event detection described above to detect lost primary stream 210, 301 signals, slide changes, resolution changes, etc. Additionally, the intentionally introduced delay may be sufficient to improve the pattern detection function 134, as described below.

[0152] It will be appreciated that introducing a delay involves buffering 250 each of the collected and time-synchronized multiple primary video streams 210, 301 before publishing the output video stream 230 using that buffered frame 213. In other words, the video and / or audio data of at least one, some or all of the multiple primary video streams 210, 301 may end up being present in the central server 130 in a buffered manner, for the reasons discussed above, particularly for use by the pattern detection function 134, rather than being used like a cache, but with the intent of being able to handle varying bandwidth situations (as in a traditional cache buffer).

[0153] Thus, in some embodiments, the pattern detection step involves considering certain information of at least one, e.g. several, e.g. at least four, or all, of the multiple primary digital video streams 210, 301, which is present in a frame 213 that is later than a frame of the time-synchronized primary digital video stream 210 that has not yet been used in generating the output digital video stream 230. Thus, the newly added frame 213 resides in said buffer 250 for a certain latency period before forming part of (or the basis of) the output video stream 230. During this period, the information of said frame 213 constitutes "future" information in relation to the frame currently being used to generate the current frame of the output video stream 230. When the timeline of the output video stream 230 reaches said frame 213, said frame is used in generating the corresponding frame of the output video stream 230 and may thereafter be discarded.

[0154] In other words, the pattern detection function 134 has at its disposal a set of video / audio frames 213 that have not yet been used to generate the output video stream 230, and uses this data to detect said patterns.

[0155] Pattern detection can be performed in different ways: In some embodiments performed by the AI ​​component 134a, the pattern detection step includes a second trained neural network or other machine learning component analyzing in concert at least two, e.g., at least three, e.g., at least four, or even all, of the multiple primary digital video streams 120, 301 to automatically detect said pattern 212.

[0156] In some embodiments, the detected pattern 212 comprises a speech pattern including at least two, e.g., at least three, e.g., at least four different active participants 122, each associated with a respective participant client 121 for the shared video communication service 110, and each of these active participants 122 is visually viewed in a respective one of the multiple primary digital video streams 210, 301.

[0157] Preferably, the generating step includes determining, tracking, and updating a current generating state of the output video stream 230. For example, such state may dictate which participants 122 (if any) are visible in the output video stream 230 and where on screen they are visible; which external video streams 300 are visible in the output video stream 230 and where on screen they are visible; which slides or shared screens are displayed in full screen mode or in combination with any live video streams; etc. Thus, the generating function 135 can be viewed as a state machine for the generated output video stream 230.

[0158] In order to generate the output video stream 230 as a combined video experience to be viewed, for example, by the end consumer 150, it is advantageous for the central server 130 to be able to understand what is happening at a deeper level than simply detecting individual events associated with the individual primary video streams 210, 301.

[0159] In a first example, the presenting participant client 121 changes the currently displayed slide. This slide change is detected by the event detection function 132 as described above, and metadata 242 is added to the frame indicating that a slide change has occurred. This happens many times as the presenting participant client 121 is found to be skipping forward a number of slides in rapid succession, resulting in a series of "slide change" events that are also detected by the detection function 132 and stored with the corresponding metadata 242 in a separate buffer 250 of the primary video stream 210. In practice, each such rapidly skipped forward slide may only be displayed for a few seconds.

[0160] The pattern detection function 134 looks at the information in the buffer 250 across these detected slide changes and detects a pattern that corresponds to one single slide change rather than multiple or rapidly executed slide changes (i.e., a single slide change to the last slide in the forward skip, where the last slide remains visible once the fast skip ends). In other words, the pattern detection function 134 notes that there were, for example, ten slide changes in a very short time, why treat them as a detected pattern that means one single slide change. As a result, the generation function 135 has access to the pattern detected by the pattern detection function 134 and can choose to display this last slide in full screen mode for a few seconds in the output video stream 230 because it determines that this last slide is potentially important in the state machine. It can also choose not to display the intermediately viewed slides in the output stream 230 at all.

[0161] Detection of patterns having multiple rapid slide changes may be detected by a simple rule-based algorithm, but alternatively may be detected using a neural network designed and trained to detect such patterns in video images by classification.

[0162] In another example, it may be desirable to quickly switch visual attention between current speakers, while still providing a relevant viewing experience for the consumer 150 by generating and presenting a calm and smooth output video stream 230, which may be useful, for example, when the video communication is a talk show, panel discussion, or the like. In this case, the event detection function 132 may continuously analyze each primary video stream 210, 301 to determine at any time whether the person being viewed in that particular primary video stream 210, 301 is currently speaking or not. This may be performed as described above, for example, using image processing tools conventional per se. The pattern detection function 134 may then be operable to detect certain overall patterns involving multiple primary video streams 210, 301, which patterns are useful for generating a smooth output video stream 230. For example, the pattern detection function 134 may detect a pattern of very frequent switches between current speakers and / or a pattern involving multiple simultaneous speakers.

[0163] The generation functionality 135 can then take such detected patterns into account when making automated decisions related to the generation state, such as, for example, not automatically switching visual focus to a speaker who speaks for only half a second before going silent again, or switching to a state in which multiple speakers are displayed side-by-side during a period in which they are alternating or speaking simultaneously. This state determination process can itself be performed using time series pattern recognition techniques or using trained neural networks, but can also be based at least in part on a predetermined set of rules.

[0164] In some embodiments, there may be multiple patterns that are detected in parallel and form input to the state machine of the generation function 135. Such multiple patterns may be used by the generation function 135 in different AI components, computer vision detection algorithms, etc. As an example, a permanent slide change may be detected while simultaneously detecting unstable connections of some participant clients 121, while other patterns detect the current main speaking participant 122. Using all such available pattern data, a classifier neural network may be trained and / or a set of rules may be developed to analyze the time series of such pattern data. Such classification may be supervised, at least in part, e.g., completely, to result in determined desired state changes used in the generation. For example, different such predefined classifiers may be generated that are specifically configured to automatically generate the output video stream 230 according to various different generation styles and desires. Training may be based on known generation state change sequences as the desired output, and known pattern time series data as training data. In some embodiments, a Bayesian model may be used to generate such classifiers. In a concrete example, a priori information can be obtained from experienced producers, who can provide input such as "In talk shows, we never switch directly from speaker A to speaker B, but always give an overview first before focusing on other speakers, unless the other speaker is very dominant and speaks loudly". This generation logic is expressed as a Bayesian model of the general form "If X is true | given the fact that Y is true | do Z". The actual detection (e.g. whether someone is speaking loudly) can be done using classifiers or threshold-based rules.

[0165] Given a large dataset (of pattern time series data), deep learning techniques can be used to develop correct and compelling generative formats for use in the automatic generation of video streams.

[0166] In summary, by using a combination of event detection based on multiple individual primary video streams 210, 301; intentionally introduced delays; pattern detection based on multiple time-synchronized primary video streams 210, 301 and detected events; and a generation process based on detected patterns, it is possible to realize an automatic generation of the output digital video stream 230 according to a wide possible selection of tastes and styles. This result is valid across a wide range of possible neural network and / or rule-based analysis techniques used by the event detection function 132, the pattern detection function 134, and the generation function 135. This is especially valid in the embodiment described below, which is characterized by a first generated video stream being used for the automatic generation of a second generated video stream; and by using different intentionally added delays for different groups of participant clients. It is also especially valid in the embodiment described below, in which a detected trigger results in a switch of which video stream is used in the generated output video stream, or in an automatic crop or zoom of the video stream used in the output video stream.

[0167] As exemplified above, the generating step may include generating the output digital video stream 230 based on a set of predetermined and / or dynamically variable parameters relating to the visibility of each one of the multiple primary digital video streams 210, 301 in the output digital video stream 230, the arrangement of the visual and / or auditory video content, the visual or auditory effects used, and / or the output mode of the output digital video stream 230. Such parameters may be automatically determined by a state machine in the generating functionality 135, and / or set by an operator controlling the generation (semi-automated), and / or predetermined based on some a priori configurational desires (such as a minimum time between layout changes of the output video stream 230 or state changes of the types exemplified above).

[0168] In a practical example, the state machine may support a set of predefined standard layouts that may be applied to the output video stream 230, such as a full-screen presenter view (showing the currently speaking participant 122 in full screen); a slide view (showing the currently shared presentation slide in full screen); a "butterfly view" (showing both the currently speaking participant 122 and the currently shared presentation slide in a side-by-side view); a multi-speaker view (showing all or a selected subset of the participants 122 side-by-side or in a matrix layout). Various available production formats may be defined by a set of state machine state change rules (as exemplified above) along with a set of available states (such as the set of standard layouts above). For example, one such production format may be "panel discussion", another may be "presentation", etc. By selecting a particular production format via a GUI or other interface to the central server 130, an operator of the system 100 may quickly select one of a set of such predefined production formats, and then enable the central server 130 to generate the output video stream 230 in accordance with that production format in a fully automatic manner based on available information as described above.

[0169] Furthermore, during production, for each conference participant client 121 or external video source 300, a respective in-memory buffer is created and maintained, as described above. These buffers can be easily deleted, added, and modified on the fly. The central server 130 may then be configured to receive information regarding added / dropped-off participant clients 121 and participants 122 scheduled to speak, scheduled or unexpected pauses / resumes of the presentation, desired changes to the currently used production format, etc., during the production of the output video stream 230. Such information may be provided to the central server 130, for example, via an operator GUI or interface, as described above.

[0170] As illustrated above, in some embodiments, at least one of the multiple primary digital video streams 210, 301 may be provided to a digital video communication service 110, and the publishing step may then include providing the output digital video stream 230 to that same communication service 110. For example, the output video stream 230 may be provided to a participant client 121 of the video communication service 110, or may be provided as an external video stream to the video communication service 110 via the API 112. In this manner, the output video stream 230 may be made available to multiple or all of the participants of the video communication event currently being facilitated by the video communication service 110.

[0171] As also mentioned above, the output video stream 230 may additionally or alternatively be provided to one or more external consumers 150 .

[0172] Generally, the generating step is performed by a central server 130, and the output digital video stream 230 can be provided as a live video stream via an API 137 to one or more concurrent consumers.

[0173] As mentioned above, participant clients 121 may be organized into groups of two or more participant clients 121. Figure 8 is a simplified diagram of the system 100 in a configuration that performs automatic generation of output video streams when such groups exist.

[0174] In this FIG. 8, a central server 130 has a collection function 131 as described above.

[0175] The central server 130 also comprises a first generating function 135', a second generating function 135'', and a third generating function 135'''. Each such generating function 135', 135'', 135''' corresponds to a generating function 135, and what has been described above in relation to generating function 135 applies equally to generating functions 135', 135'', 135'''. The generating functions 135', 135'', 135''' may be separate or may share multiple functions in one single logical function, and there may be more than three generating functions, depending on the detailed configuration of the central server 130. The generating functions 135', 135'', 135''' may in some cases be different functional aspects of one and the same generating function 135. Various communications between the generating functions 135', 135'', 135''' and other entities may be performed via appropriate APIs.

[0176] It will further be appreciated that there may be a separate collection function 131 for each production function 135', 135'', 135''' or group of such production functions, and depending on the detailed configuration, there may be multiple logically separated central servers 130, each having its own collection function 131.

[0177] Furthermore, the central server 130 comprises a first publishing function 136', a second publishing function 136'', and a third publishing function 136'''. Each such publishing function 136', 136'', 136''' corresponds to a publishing function 136, and what has been described above in relation to the publishing function 136 applies equally to the publishing functions 136', 136'', 136'''. Depending on the detailed configuration of the central server 130, the publishing functions 136', 136'', 136''' may be separate functions, may be co-located in one single logical function with multiple functions, and there may be more than three publishing functions. The publishing functions 136', 136'', 136''' may in some cases be different functional aspects of one and the same publishing function 136.

[0178] In FIG. 8, for illustrative purposes, three sets or groups of participant clients are shown, each corresponding to a participant client 121 described above. Thus, there is a first group 121′ of such participant clients 121, a second group 121″ of such participant clients, and a third group 121′″ of such participant clients. Each of these groups may consist of one or preferably at least two participant clients. There may be only two such groups or there may be more than two such groups, depending on the detailed configuration. The assignment between the groups 121′, 121″, 121′″ may be exclusive in the sense that each participant client 121 is assigned to at most one group 121′, 121″, 121′″. In an alternative configuration, at least one participant client 121 may be assigned to more than one such group 121′, 121″, 121′″ simultaneously.

[0179] FIG. 8 also shows an external consumer 150, and as mentioned above, it is understood that there may be multiple such external consumers 150.

[0180] While FIG. 8 does not show the video communication service 110 for purposes of simplicity, it will be appreciated that a video communication service of the general type described above may be used in conjunction with the central server 130, for example to provide a shared video communication service to each participant client 121 using the central server 130 in the manner described above.

[0181] Respective primary video streams may be collected by a collection function 131 from respective participant clients 121, such as participant clients of the above groups 121′, 121″, 121′″. Based on the provided primary video streams, a generation function 135′, 135″, 135′″ may generate respective digital video output streams.

[0182] 8, one or more such generated output streams may be provided as respective input digital video streams from one or more respective generating functions 135', 135''' to another generating function 135'', which in turn generates a secondary digital output video stream for presentation by a publishing function 136''. This secondary digital output video stream is thus generated based on the one or more input primary digital video streams as well as one or more pre-generated digital input digital video streams.

[0183] The two or more different generating steps 135', 135'', 135''' may include the introduction of a respective time delay. In some embodiments, one or more of the respective generated output digital video streams from these generating steps 135', 135'', 135''' may not be synchronized in time with any other of the video streams that may be provided to other participant clients in the publishing step due to the introduction of the time delay. Such a time delay may be added intentionally in any of the ways described herein and / or may be a direct result of the generation of the generated digital video stream. As a result, any participant client consuming the time-unsynchronized generated output digital video stream will do so in a "time zone" that is slightly offset (in time) relative to the video stream consumption "time zones" of the other participant clients.

[0184] For example, one of a group 121', 121'', 121''' of participant clients 121 may consume a respective produced video stream in a first such "time zone", while another participant client 121 of that group 121', 121'', 121''' may consume a respective produced video stream in a second such "time zone". Because both of these respective produced video streams may be generated at least in part based on the same primary video stream, all such participant clients 121 are active in the same video communication but in different "time zones" relative to one another. In other words, the respective timelines for produced video stream consumption may be offset in time between different groups 121', 121'', 121'''.

[0185] For example, some generation steps (such as 135', 135''') may be direct (without the use of intentionally introduced time delays) and / or may involve only relatively computationally lightweight processing prior to provision for publication, while other generation steps (such as 135'') may involve intentionally introduced time delays and / or relatively heavy-duty processing, resulting in the resulting digital video stream being generated for earliest publication at a delay relative to the earliest delay for publication of the respective digital video stream of the former generation steps 135', 135'''.

[0186] Thus, each participant client 121 in one or more of the groups 121', 121'', 121''' may be able to interact with one another with the same perceived time delay. At the same time, the group associated with the respective larger time delay may use the generated video stream from the group with the smaller time delay as an input video stream to be used when generating an output video stream for the group with the larger time delay to view in a "time zone" after the group.

[0187] The result of this first, greater time delay generation (in step 135'') is thus a generated digital video stream of the type described above, which may visually include, for example, one or more of said primary video streams as subparts, in processed or unprocessed form. This said generated video stream may include live captured video streams, slides, externally supplied video or images, etc., as generally described above in connection with the video output streams generated by the central server 130. The said generated video stream may also be generated based on detected events and / or patterns in an intentionally delayed or real-time input primary video stream supplied by the participant client 121, in the general manner described above.

[0188] In an exemplary embodiment, the participant clients of the first group 121' are part of a discussion panel and communicate using the video communication service 110 with relatively low latency, each of which is continuously fed with the generated video stream from the publishing step 136' (or each other's respective primary video stream). The audience of the discussion panel is constituted by the participant clients of the second group 121'' and is continuously fed with the generated video stream from the generating step 135'', now associated with a slightly higher latency. The generated video stream from the generating step 135'' can be automatically generated in the general manner described above to automatically shift between the view of the individual discussion panel speakers (participant clients assigned to the first group 121', such view being fed directly from the collection function 131) and a generated view showing all the discussion panel speakers (this view being the first generated video stream). Thus, the audience can enjoy a well-staged experience while the panel speakers can interact with each other with minimal latency.

[0189] The intentionally added delay to each primary video stream used in generating step 136'' may be at least 0.1 seconds, such as at least 0.2 seconds, such as at least 0.5 seconds, and may be at most 5 seconds, such as at most 2 seconds, such as at most 1 second. The intentionally added delay may also depend on the inherit latency associated with each primary video stream used, so as to achieve perfect time synchronization between each of the primary video streams used and the generated video streams input from generating step 135' to generating step 135''.

[0190] It will be appreciated that all such primary video streams, as well as the generated video streams from the generation step 135', may additionally be intentionally delayed to improve pattern detection for use in the second generation function 135'' in the general manner described above.

[0191] FIG. 8 further illustrates a number of alternative or simultaneous ways of publishing the various generated video streams generated by the central server 130.

[0192] In general, in a publishing step performed by a first publishing function 136′ configured to receive the first generated video stream from the first generation function 135′, said first generated video stream may be continuously provided to at least one of a first participant client 121 and a second participant client 121. For example, this first participant client may be a participant client from a group 121′ providing a respective primary digital video stream to the first generation function 135′.

[0193] In some embodiments, one or more of the participant clients of the group 121′ may receive the second generated video stream by a second publishing function 136″ configured to receive the second generated video stream from the second generating function 135″.

[0194] Thus, each of the primary video stream supplying participant clients assigned to the first group 121' may be supplied with the first generated video stream if they are not directly supplied with each other's primary digital video streams, with a certain delay or latency due to synchronization between the multiple primary video streams, and further with a delay or latency that may be intentionally added to allow sufficient time for event and / or pattern detection, as described above.

[0195] Correspondingly, each of the participant clients assigned to the second group 121'' may be provided with a second generated video stream that also includes the intentionally added delay associated with the second generation step, added for the purpose of time-synchronizing the first generated video stream with the first and second primary video streams. This additional delay may or may not make communication between the participant clients of the second group 121'' difficult, for example, by the participant clients of the second group 121'' interacting with the video communication service 110 in a different way than the participant clients of the first group 121'.

[0196] Thus, the participant clients of the first group 121' form a subgroup of all participant clients 121 currently participating in the video communication service 110 and are present at and using the service in a "time zone" slightly ahead (e.g., 1-3 seconds ahead) of any participant clients to which a generated video stream, such as the first generated video stream or the second generated video stream, is instead continuously provided. The other participant clients (not assigned to the first group 121', but instead assigned to the second group 121'') will nevertheless be continuously provided with a second generated video stream, which is generated based on at least one of the primary video streams from which the second generated video stream is generated (and which may at each point in time include either or both of multiple primary video streams), but in a slightly later "time zone". Because the first generated video stream is generated directly based on at least one of these primary video streams, there is no added delay or latency to time-synchronize the video stream already generated based on the primary video stream itself, providing a more direct, low-latency video communication service 110 experience to these participant clients 121.

[0197] Again, this may mean that participant clients 121 assigned to the first group 121' are not provided with access to the second produced video stream.

[0198] 8, the second resultant video stream may also be generated as a digital video stream additionally based on the resultant output video (third resultant output digital video stream) of the third generation step 135'''. Thus, the resultant video stream from each of the generation steps 135', 135'', 135''' is generated based on an at least partially overlapping primary input digital video stream but with a different respective intentionally added latency ("time zone") relative to said primary video stream and provided to different respective groups 121', 121'', 121''' of participant clients 121.

[0199] The participant clients 121 assigned to the third group 121''' may have less strict latency requirements than the participant clients 121 assigned to the first group 121'. For example, the participant clients 121 in the first group 121' may be members of the discussion panel described above (which interact with each other in real time and therefore require low latency), while the participant clients 121 in the third group 121''' may constitute an expert panel or similar panel that interacts with the panel but in a more structured manner (such as using clear questions / answers) and therefore may tolerate a higher latency than the first group 121'.

[0200] Both the first generated video stream and the third generated video stream may optionally be provided to a second generation function 135'' for use as a basis for generating a second generated video stream.

[0201] Thus, the first generating step 135' may include introducing an intentional delay or latency of the type described above that is introduced in addition to the delay introduced as part of the synchronization of the first and second primary video streams, so that, for example, sufficient time may be achieved to perform efficient event and / or pattern detection. The introduction of such an intentional delay or latency may be done as part of the synchronization performed by the synchronization function 133 described above (not shown in FIG. 8 for reasons of simplicity). Similarly for the third generating step 135''', but may introduce an intentional delay or latency that is different from the delay or latency intentionally introduced for the first generating step 135'.

[0202] In particular, the intentionally introduced delay or latency results in a time desynchronization between the first and third generated video streams, meaning that the first and third generated video streams would not follow a common timeline if they were exposed immediately and consecutively upon generation of each individual frame.

[0203] Thus, three separately generated video streams may be generated and consumed / published simultaneously, but in different "time zones". Despite the fact that they are based at least in part on the same primary video material, the multiple generated video streams are published with different latencies. The first group 121', which requires the lowest latency, can interact using the first generated video stream, which offers very low latency. Meanwhile, the third group 121''', which is willing to accept slightly higher latency, can interact using the second generated video stream, which offers higher latency but on the other hand offers more flexibility in terms of intentionally added delays, thereby achieving better automatic generation as described elsewhere herein. Meanwhile, the second group 121'', which is less sensitive to latency, can enjoy interacting using the second generated video stream, which can incorporate material from both the first group 121' and the third group 121''', and which is also automatically generated in a very flexible manner. It is noted that all these groups of participant users 121', 121'', 121''' interact with each other using the video communication service 110, albeit using various latencies and therefore operating in different "time zones". However, due to the synchronization of the individual input video streams at each production facility, the participant users 121 are unaware of the different latencies from their respective perspectives.

[0204] As described above, each participant client 121 assigned to each of the groups 121', 121'', 121''' can participate in one and the same video communication service 110 in which the second generated video stream is continuously published.

[0205] And different ones of the groups 121', 121'', 121''' may be associated with different participant interaction privileges in the video communication service 110. In these and other embodiments, different ones of the groups 121', 121'', 121''' may be associated with different maximum time delays (latencies) used to generate the respective generated video streams that are exposed to the participant clients 121 assigned to that group 121', 121'', 121''''.

[0206] For example, a first group of panel discussion participant clients 121' may be associated with full interaction privileges and can speak at any time. A third group of participant clients 121''' may be associated with slightly more restricted interaction privileges, such as having to request the floor by the video communication service 110 before they can speak by unmuting their microphones. A second group of audience participant users 121'' may be associated with even more restricted interaction privileges, such as only being able to pose written questions in a common chat room, but not being able to speak.

[0207] Thus, various groups of participant users may be associated with different interaction privileges and different latencies for the respective generated video streams exposed to them, such that the latency is an increasing function of decreasing interaction privileges. The more freely a given participant user 121 is permitted by the video communication service 110 to interact with other users, the lower the tolerable latency. The lower the tolerable latency, the less likely the corresponding auto-generated functionality will take detected events, patterns, etc. into account.

[0208] The group with the greatest latency may be a viewer-only group with no interactive privileges other than passive participation in the video communication service.

[0209] In particular, a respective maximum time delay (latency) for each of the groups 121', 121'', 121''' may be determined as the maximum latency difference between all primary video streams and any resulting video streams that are continuously exposed to participant clients of that group. To this sum may be added time delays intentionally added for the purpose of detecting events and / or patterns, as described above.

[0210] As used herein, the terms "generation" and "generated digital video stream" may refer to different types of generation. In one example, a single, well-defined digital video stream is generated by a central entity, such as a central server 130, which forms the generated digital video stream for delivery and publication to a set of specific participant clients 121 that will consume the generated digital video stream. In another example, different individual such participant clients 121 may view slightly different versions of the generated digital video stream. For example, the generated digital video stream may include multiple separate or combined digital video streams that a local software function 125 of a participant client 121 may allow the user 122 to switch between, place on a screen 124, or otherwise configure or process. Often, what matters is in what "time zone" (i.e., with what latency) the generated digital video stream, including its time-synchronized subcomponents, is delivered. Thus, the case described above in connection with FIG. 8 in which different participant clients 120 of a first group 121′ are supplied with each other's primary video streams can be considered as a first generated digital video stream being supplied to these participant clients (in the sense that a time-synchronized set of raw data or processed first and second primary digital video streams is made available to both the first and second participant clients).

[0211] To further clarify and illustrate the use of the participant client groups 121', 121'', 121''' described above, the following example is provided in the form of a video communication services conference involving three different concurrent "time zones":

[0212] The first group of participant clients 121' experience interactions in real time, or at least near real time (depending on unavoidable hardware and software latencies). These participant clients are fed video with audio from each other to facilitate such interactions and communication between the users 122. The first group 121' may serve interactions to the core users 122 of the conference that other participant clients (other than the first group 121') may be interested in participating in.

[0213] A second group 121'' of such other participant clients participates in the same conference, but is in a different "time zone", further away from real time than the first group of participant clients 121'. For example, the second group 121'' may be an audience with interactive rights, such as being able to pose questions to the first group 121'. The "time zone" of the second group 121'' may have a delay in relation to the "time zone" of the first group 121' such that posed questions and answers are delivered with a noticeable but short delay. On the other hand, this slightly larger delay allows the participant clients of this second group 121'' to experience a generated digital video stream, which is automatically generated in a more complex way, providing a more pleasant user experience.

[0214] A third group 121''' of such other participant clients also participates in the same conference, but only as viewers. This third group 121''' consumes a generated digital video stream, which may be generated automatically in a more elaborate and complex manner, that is consumed in a third "time zone" that has an even greater delay than the second "time zone". However, since the third group 121''' cannot provide input to the communication service in a manner that affects the first group 121' and the second group 121'', the third group 121''' experiences the conference as taking place in "real time" and in a convincingly staged manner.

[0215] Of course, there may be more than three such groups of participant clients, each associated with a respective conference "time zone" of increasingly larger time delays and increasing complexity to generate, using the principles described herein.

[0216] FIG. 9 illustrates a method according to the present invention for providing a shared digital video stream.

[0217] In a first step S1 the method starts.

[0218] In a subsequent collection step S2, a first digital video stream is collected from a first digital video source and a second digital stream is collected from a second digital video source, generally in the manner described above. Thus, the first and / or second digital video stream may each be collected from a respective participant client 121 or from an external source 300, and may be performed by a collection function 131 of the central server 130.

[0219] In a subsequent first generation step S4, said shared digital video stream is generated as an output digital video stream, which may be performed generally as described above by generation steps 135, 135', 135'', 135''' etc.

[0220] In a first generating step S4, a shared digital video stream is generated based on consecutively considered frames of the first digital video stream such that image information from the first digital video source is visible in the shared digital video stream but image information from the second digital video source is not visible in the shared digital video stream, in other words the shared video stream contains, at least to some extent, visual material originating from the first digital video stream but not visual material originating from the second digital video stream.

[0221] In a subsequent trigger detection step S5, the first and / or second digital video streams are digitally analysed to detect at least one trigger.

[0222] This analysis and detection may be performed by the same generation step 135, 135', 135'', 135''' that generates the shared video stream and is based on automatic detection of predefined types of image and / or audio patterns.

[0223] A trigger may be an event or pattern of the type described and illustrated above (performed by event detection function 132 and pattern detection function 134, respectively), the detection of which is typically performed using automatic digital processing of audio and / or image / video data contained in the digital video stream(s). For example, as illustrated above, automatic image processing algorithms, such as those using trained neural networks or other machine learning tools, may be employed to automatically detect the presence of a particular trigger based on images contained in the first and / or second video feeds. Correspondingly, a corresponding type of automatic audio processing algorithm, conventional per se, may be used to detect the presence of a particular trigger based on audio contained in the first and / or second video feeds.

[0224] That an image and / or audio pattern is of a "predetermined type" means that the pattern in question is characterized in terms of a set of one or more absolute or relative parameter values ​​that are defined prior to said detection, as exemplified below.

[0225] Typically, the presence of said audio or visual pattern constitutes said corresponding trigger, which is further specifically predefined in order to instruct the automatic generation step 135, 135', 135'', 135''' for generating the shared video stream to change the generation mode (rules) of the shared digital video stream upon detection of said trigger according to predefined generation rules.

[0226] Thus, the generation step 135, 135', 135'', 135''' may comprise or have access to a database that defines one or more triggers, either in real time or over time, with respect to corresponding parameter values ​​that characterize corresponding image and / or audio patterns.

[0227] In a subsequent second generation step S6 or S7, initiated in response to detection of the above trigger, the shared digital video stream is then generated again by the same (or different) generation step 135, 135', 135'', 135''', but not in the same way as the first generation step S4.

[0228] In the first alternative second generating step S6, the shared digital video stream is generated as an output digital video stream based on a number of frames of the second digital video stream considered consecutively such that image information from the second digital video source is visible in the shared digital video stream. It should be noted that in this case, the output digital video stream may be further generated based on a number of frames of the first digital video stream considered consecutively such that image information from the first digital video source is visible in the shared digital video stream, or may not be generated based on the number of frames of the first digital video stream considered consecutively such that image information from the first digital video source is visible in the shared digital video stream. In other words, when switching from the first generating step S4 to the second generating step S6, the shared video stream may go from a state of displaying content from the first video stream but not from the second video stream to a state of displaying content from the second video stream but not from the first video stream, or a state of displaying content from both the first video stream and the second video stream.

[0229] In a second alternative second generating step S7, a shared digital video stream is generated as an output digital video stream based on consecutively considered frames of said first digital video source, but with at least one of different cropping, different zooming, different panning and different focus plane selection of said first digital video stream compared to the first generating step S4. In other words, the content of the video stream displayed in the shared video stream is cropped, uncropped, zoomed in, zoomed out, panned vertically and / or horizontally and / or the focus plane of said video stream is shifted with respect to the current crop / zoom / pan / focus plane state of the video stream as used in the first generating step S4.

[0230] It will be appreciated that such crop / pan / zoom / focus planes may be performed by said generating step 135, 135', 135'', 135''' based on an existing video stream (which itself may contain multiple possible focus planes with different image information at pixel level) and / or by said generating step 135, 135', 135'', 135''' communicating commands to the video source (such as a digital video camera) capturing said video stream to modify corresponding capture parameters accordingly. For example, this may then entail that the corresponding camera capturing the video stream in question zooms, pans, and / or shifts its focus plane according to instructions provided by the generating step 135, 135', 135'', 135'''.

[0231] In a subsequent publishing step S8, the output digital video stream is continuously provided to consumers of the shared digital video stream, such as participant clients 121 and / or external consumers 150, in the general manner described above.

[0232] The method may then be repeated, returning to step S2, as shown in FIG.

[0233] In the following step S9, the method ends.

[0234] The first digital video stream may be captured continuously by a first digital camera, and the second digital video stream may be captured continuously by a second, different digital camera (thus constituting the primary video stream using the terminology already used herein). Alternatively, the first and / or second digital video streams may constitute respective previously generated digital video streams, such as in the case of using multiple different groups 121', 121'', 121''' of participant clients 121 associated with different respective latencies ("time zones") as described above, which previously generated digital video streams are, in turn, generated at different latencies ("time zones") compared to the currently generated shared video stream (see above for further details regarding such "time zones").

[0235] If the first video stream is an already generated video stream, it is preferable (but not required) that the crop / zoom / pan settings are performed based on the already existing first video stream, rather than instructing an upstream camera to change the crop / zoom / pan settings.

[0236] Using this method, automatic generation of shared video streams can be achieved that can provide a more intuitive and natural experience for consumers of the generated shared video, since the actual audio / video content of each video stream is used to detect triggers that will result in the automatic generation transitioning from one automatically generated format to a different such format.

[0237] The triggers can be predefined with appropriate parameters to accommodate various needs. For example, the actions of individual people depicted in the first and / or second video streams can be automatically evaluated with respect to such triggers, allowing the production format to be changed depending on how such actions are performed. In other examples, certain predefined triggers can be used as manual cues given by people depicted in the first and / or second video streams to change the production format on the fly while the production is ongoing.

[0238] Below we describe some examples of such triggers and the corresponding changes to the generated format.

[0239] In a first example shown in Figures 10a and 10b, the predefined pattern includes a first (human or e.g. machine) participant 430 depicted in said first digital video stream, illustrated in Figure 10a as captured by a first digital video camera 410. The first participant 430 gazes towards an object 440, which in the example of Figure 10a is a second (human or machine) participant, who in turn is depicted in a second digital video stream, which in this example is captured by a second digital video camera 420.

[0240] It is understood that the second object 440 may be something else, such as a group of human participants or any physical object of general interest for ongoing communication. For example, the shared video stream may be a medical procedure document, whereby the object 440 may be part of a patient. In another example, the object 440 may be an object of an educational session or a sales presentation. Also, the object 440 may be, for example, a whiteboard or a screen for a slide presentation.

[0241] The first and second video streams may be of any of the types described herein above.

[0242] 10a and 10b, a predefined image and / or audio pattern is detected based on information about the relative orientations of the first camera 410, the participant 430, and the object 440. This information may be present in the system 100 (particularly in the central server 130), such as pre-supplied during setup / configuration and / or automatically detected while production is in progress.

[0243] For example, the respective positions of the first camera 410 and the second camera 420 may be provided to the central server 130 by the respective cameras 410, 420 equipped with measuring means such as MEMS circuits with accelerometers and gyros, or conventional position measuring means per se such as stepper motors arranged to continuously or intermittently measure the current position of the respective cameras 410, 420 based on some suitable geometrical reference. In another example, the orientation of the respective cameras 410, 420 may be detected by a third camera (not shown) using suitable automatic digital image processing algorithms, which captures at least one of the cameras 410, 420 in an image and uses digital image processing to determine the relative orientation based on this captured image information.

[0244] Note that in this context, "orientation" can encompass both position and orientation components.

[0245] The positions of the first participant 430 and the object 440 relative to any suitable reference system (such as relative to the first camera 410 and / or the second camera 420) can be determined using digital image processing based on the video stream(s) captured by the first camera 410 and / or the second camera 420.

[0246] As shown in FIG. 10a, the first participant 430 is gazing downward in the diagram, rather than towards the second participant 430.

[0247] In this example, the predetermined image and / or audio pattern is further detected based on a digital image-based determination of at least one of a body orientation, head orientation and gaze direction of the first participant 430 based on the first digital video stream.

[0248] As shown in FIG. 10b, the first participant 430 has turned towards the second participant 440 and is looking (gazing) at the second participant 440.

[0249] The body and head orientation of the first participant 430 may be determined based on digital image processing of the first video stream captured by the first camera 410. Such algorithms may be conventional in themselves and may, for example, use a priori knowledge of the expected shape of the first participant 430 in the video stream captured by the first camera 410 when turning in various directions. This may be achieved using a trained neural network or other machine learning components. Together with the relative orientation of the first camera 410, the relative position of the first participant 430 and the object 440, and the determined body or head orientation of the first participant 430, the central server 130 may determine whether the first participant 430 has turned (head or body) towards the object 440 in question.

[0250] The gaze direction of the first participant 430 can be achieved in a similar manner, such as based on an image captured by the first video camera 410. Such eye-tracking techniques are known per se and may for example be based on identifying the position of the pupil and light reflection visible in the eye of the first participant 430.

[0251] The trigger may be defined as the detection of a transition pattern, such as a transition by the first participant 430 from a state in which the first participant 430 is not pointing at or gazing at the object 440 to a state in which the first participant 430 is actually pointing at and / or gazing at the object 440. The central server 130 may thus continuously monitor for such transitions based on corresponding appropriately set absolute or relative parameter values, and a trigger may be detected when such a transition occurs.

[0252] It will be appreciated that in this and other embodiments, multiple different predefined image and / or audio patterns may be monitored simultaneously, and such detected predefined image and / or audio patterns then constitute detected corresponding triggers which in turn lead to switching of auto-generation to different respective corresponding modes in accordance with predefined parameterized generation logic.

[0253] Figure 10c shows a second example, similar to that shown in Figure 10b, except that the predetermined image and / or audio pattern corresponding to the trigger in question is not the first participant 430 turning to turn or gazing at the object 440. Instead, the predetermined image and / or audio pattern includes a participant depicted in the first digital video stream (e.g., by the first camera 410) (e.g., the first participant 430) making a predetermined gesture.

[0254] The gesture may be any gesture of a predefined parameterized type, such as a gesture geometrically related to an object 440 depicted in said second digital video stream. In particular, the gesture may be a pointing of the first participant 430 towards said object 440 (as depicted in FIG. 10c, the arm 431 of the first participant 430 points towards the object 440). However, the gesture may also be based solely on the hand or fingers of the first participant 430, for example.

[0255] The predetermined pattern of this gesture type (and in particular its direction in space, if any) may be detected in a manner corresponding to the situation described in relation to FIG. 10b, and is thus detected based on information about the relative orientations of the first camera 410, the first participant 430 and the object 440, and further based on a digital image-based determination of the gesture direction of the first participant 430 based on the first digital video stream.

[0256] For both the cases shown in Fig. 10b and 10c, the orientation of the second camera 420 may also be detected and used to determine whether a trigger is detected. For example, it may be used to determine that the object 440 is visible in the second camera 420, which may constitute a condition for a trigger to be detected if the detected trigger is related to switching to the second camera 420. In some embodiments, the second generating step S6 may include determining one second camera 420 (out of multiple possible second cameras) that currently displays the object 440 and selecting that second camera 420 as the one to supply the second video stream in the second generating step S6.

[0257] In the example shown in Fig. 10d, there is only one camera, the first camera 410 (it is understood, of course, that in various embodiments there may be more cameras and other video stream generation components). The first camera 410 captures both the first participant 430 and the object 440. When the first participant 430 turns his body, head or gaze towards the object 440, this detected image pattern constitutes a detected trigger. In this case, the second generation step may include panning and / or zooming and / or cropping the video stream captured by the first camera 410 in the generated output video stream so as to focus the attention of a viewer of said generated output video stream on the object 440.

[0258] In a practical example, the detected attention of the first participant 430 (embodied in the body, head, or gaze direction of the first participant 430, as exemplified above) triggers automatic generation (in the second generation step above), based on an existing video stream or by instructing a unit generating the primary video stream to shift emphasis or focus in some way, thereby increasing the visual focus on the object in question 440.

[0259] In another example shown in Figures 11a and 11b, the predetermined image and / or audio pattern includes a first participant 430 depicted in the second digital video stream gazing towards a second camera 420 that is continuously capturing the second digital video stream. Specifically, in Figure 11a, the first participant 430 is not gazing at the second camera 420, whereas in Figure 11b, the first participant 430 is gazing in the direction of the second camera 420. This detected switch in gaze direction may thus constitute detection of a corresponding trigger of the problem.

[0260] In another example, the predetermined image and / or audio pattern includes a relative change in motion in the first digital video stream and / or the second digital video stream. For example, the amount of general motion shown in the first video stream may be parameterized and the zoom of the first video stream and / or the second video stream may be increased as a function of a decrease in general motion or vice versa. Alternatively, the second generating step may switch to a second video stream in case of a detected increase in general motion, the second video stream being captured by a second camera 420 showing a wider or more distant view of the scene depicted by the first camera 410. Correspondingly, the recorded speech audio of the participant 430 may be used to determine the performance of such zoom in / zoom out / camera switching. For example, if the participant 430 is recorded speaking louder, there may be a zoom out of the first video stream in the output shared video stream and vice versa.

[0261] In yet another example, the predefined image and / or sound pattern in turn includes a predefined sound pattern characterized by its frequency and / or amplitude and / or amplitude time derivative and / or absolute amplitude change and / or sound location, for example as determined by the relative microphone volume of a particular sound capturing microphone. Such a microphone may be, for example, part of the first camera 410 or may be a separate microphone. Such a microphone is positioned to record sound occurring in or directly associated with a scene displayed in the first video stream.

[0262] For example, the predefined audio pattern may consist of a predefined phrase including at least one verbally uttered word. The audio may be recorded and provided as part of a video stream that is fed to a first generation step, and the predefined pattern may be detected by the central server 130, after which generation may switch to a second generation step when a corresponding trigger is detected. The audio analysis may use any suitable digital audio processing algorithm, such as a rule-based decision engine, a trained neural network, etc., that uses various audio information (pitch, amplitude, pattern matching, etc.) to determine whether a predefined sound pattern has been detected.

[0263] As explained in detail above, the method may also include a delay step (see FIG. 9 ), where a latency is intentionally introduced for at least the first and second digital video streams, which latency is present in the shared digital video stream, and the trigger detection step may be performed based on the first and / or second digital video streams before introducing said latency.

[0264] The waiting time is maximum 30 seconds, such as maximum 5 seconds, such as maximum 1 second, such as maximum 0.5 seconds.

[0265] Using such an intentionally added delay, the automatic generation can, for example, plan an automatic switch from the first camera 410 to the second camera 420 based on a detected trigger (such as a participant 430 looking into the second camera 420), with this planning occurring a certain time (such as 0.5 seconds) before this event joins the generated shared video stream. In this way, such a switch can be performed exactly at the time the trigger actually occurs, or at a time that best fits with other parameterized generation parameters, such as timing with the speech rhythm of the participant 430 in question. In the case of multiple groups 121', 121'', 121''' of the type described above, such planning can be done at different time horizons with respect to the generated output video streams generated for (for consumption by) the participant clients 121 of the different ones of such groups 121', 121'', 121'''.

[0266] The predetermined image and / or audio patterns may constitute (or be determined to be) "events" and / or "patterns" of the type described above in connection with event detection function 132 and pattern detection function 132.

[0267] Below are some examples of how the present invention may be embodied in practice:

[0268] In multi-cam presentations, talk shows and panel discussions, different cameras can present different camera angles of the presenter. The auto-generation feature can be configured to automatically select different camera angles depending on which camera the presenter is looking at and / or triggered by gestures or audio cues.

[0269] For video podcasts or talking head videos, the auto-generation feature can be configured to automatically switch between different cameras depending on which camera the current speaker is facing.

[0270] In town hall meetings, the auto-generation feature can be configured to switch to an audience-facing camera for additional input or questions from participants by monitoring the audio feed associated with that camera and triggering it when a certain level is reached that the camera goes live, or by a voice command such as "question from audience."

[0271] For product presentations or reviews, the auto-generation feature can be configured to automatically switch to a camera pointed at the product when it senses motion at that source, or based on another trigger.

[0272] In robotic surgery, the video stream captured by the robotic camera filming the surgery can be replaced with normal information presentation when it detects certain gestures, audio cues, or recognizes that the surgeon is not using or has looked up from the surgeon console.

[0273] In an educational context, the camera can be set to point at a regular whiteboard or chalkboard, and the auto-generation feature can be set to switch to that camera when the teacher gestures at the board or gives instructions via voice commands.

[0274] At cultural events such as concerts, the auto-generation feature can be configured to switch between multiple cameras aimed at the singer, band members, and orchestra, which can be triggered by gestures or which camera the talent is looking at.

[0275] In a theatrical performance, the auto-generation feature can be configured to cut between different camera angles depending on who is speaking, by face tracking, audio cues, gestures, or according to a pre-determined schedule (rundown).

[0276] Thus, in addition to trigger detection of the type described herein, automatic generation can also switch from one format (generation rules) to another based on a predefined schedule (rundown), and of course can be manually overridden in some cases.

[0277] The present invention also relates to computer software functionality for providing a shared digital video stream in accordance with the above, and such computer software functionality may be configured, when executed, to perform at least some of the collecting, delaying, first generating, trigger detection, second generating, and publishing steps described above. The computer software functionality may be configured to run on physical or virtual hardware of the central server 130, as described above.

[0278] The present invention also relates to such a system 100 for providing a shared digital video stream, which in turn comprises a central server 130. The central server 130 may in turn be configured to perform at least some of the collection, delay, first generation, trigger detection, second generation, and publishing steps described above. For example, these steps may be performed by the central server 130 executing the computer software functions described above for performing the steps described above. Collection may be performed by a collection function 131. Detection of the predefined image and / or audio patterns and triggers described above may be performed by an event pattern detection function 132, 134, or a generation function 135 of the central server 130. Any intentional delays may be performed by the collection function 131 or the like in the manner generally described above. Publishing may be performed by a publishing function 136.

[0279] It will be appreciated that the principles of auto-generation based on an available set of input video streams described above, including time synchronization of such input video streams, event and / or pattern detection, trigger detection, etc., may be applied simultaneously at different levels, such that one such auto-generated video stream may form an available input video stream for a downstream auto-generation function that in turn generates a video stream.

[0280] The central server 130 may be configured to control the assignment of groups 121′, 121″, 121′″ to individual participant clients 121. For example, dynamically changing the group assignments for a particular such participant client during the course of a live video communication service session may be part of the automated generation of that video communication service by the central server 130. Such reallocation may be triggered dynamically based on a predefined timetable or as a function of parameter data that may dynamically change over time, for example, in response to requests of individual participant client users 122 (provided via those clients 121).

[0281] Correspondingly, the central server 130 may be configured to dynamically change the group composition over the course of the video communication service, such as using particular groups only for predefined time periods (e.g., during a scheduled panel discussion).

[0282] Changing the group assignments in a predetermined manner may be an automatic result of detection of a particular trigger, in a manner corresponding to that described above in relation to Figures 9 to 11b.

[0283] In all the above-mentioned aspects, the invention may further comprise an interaction step, in which at least one participant client of a first group, the first group being associated with a first latency, interacts bidirectionally with at least one participant client of the same first group or of a second group, the second group being associated with a second latency, the second latency being different from the first latency. It is understood that all these participant clients may be participants of one and the same communication service of the above-mentioned type.

[0284] Although preferred embodiments have been described above, it will be apparent to those skilled in the art that many modifications can be made to the disclosed embodiments without departing from the essential concepts of the invention.

[0285] For example, many additional features not described herein may be provided as part of the system 100 described herein. In general, the presently described solution provides a framework within which detailed functionality and features can be built to accommodate a wide variety of specific applications in which streams of video data are used for communication.

[0286] In general, anything said about the present method is applicable to the present system and computer software product, and vice versa.

[0287] Therefore, the invention is not limited to the described embodiments, but can be modified within the scope of the appended claims.

Claims

1. 1. A method for providing a shared digital video stream, the method comprising the steps of: In a collecting step (S2), a first digital video stream is collected from a first digital video source, and a second digital stream is collected from a second digital video source; In a first generating step (S4), generating the shared digital video stream as an output digital video stream based on consecutively considered frames of the first digital video stream such that image information from the first digital video source is visible in the shared digital video stream and image information from the second digital video source is not visible in the shared digital video stream; In a trigger detection step (S5), a trigger is automatically detected by performing at least one of the following: a) digitally analyzing the first digital video stream to automatically detect the trigger in the form of a participant (430) depicted in the first digital video stream gazing towards or making a predetermined gesture in relation to an object (440) depicted in the second digital video stream, wherein the detection is based on information regarding the relative orientations of the first camera (410), the participant (430), and the object (440), and the detection is further based on digital image-based determination of at least one of the participant's (430) body direction, head direction, gaze direction, and gesture direction; and b) digitally analyzing the first and / or second digital video streams to automatically detect a trigger in the form of a transition in the gaze direction of a participant (430) depicted in the second digital video stream from not gazing at a second camera (420) to gazing at the second camera (420); wherein the trigger instructs changing a generation mode of the shared digital video stream according to a predetermined generation rule; In a second generating step (S6, S7) initiated in response to detecting the trigger, the shared digital video stream is generated as an output digital video stream based on a plurality of consecutively considered frames of the second digital video stream, and / or the shared video stream is generated as an output digital video stream based on a plurality of consecutively considered frames of the first digital video source, such that image information from the second digital video source is visible in the shared digital video stream, but with at least one of different cropping, different zooming, different panning, or different focus plane selection of the first digital video stream compared to the first generating step (S4); and In a publishing step (S8), the output digital video stream is continuously made available to consumers of the shared digital video stream.

2. 10. The method of claim 1, the first digital video stream is continuously captured by a first digital camera (410); The method wherein the second digital video stream is continuously captured by a second digital camera (420).

3. 3. The method of claim 1 or 2, The method of claim 1, wherein the second object (440) is a person, a group of people, a machine, a whiteboard, or a presentation screen.

4. 10. The method of claim 1, The method, wherein the gesture is to point to the object (440).

5. 10. The method of claim 1, The method, wherein said detecting said trigger comprises detecting a predetermined image pattern comprising a change in relative motion in said first and / or second digital video streams.

6. 10. The method of claim 1, further comprising: Digital image processing based on the first and / or second digital video streams is used to determine the positions of the participants (430) and the object (440) relative to the first camera (410) and / or the second camera (420).

7. 10. The method of claim 1, The method, wherein the second generating step (S6, S7) includes determining a particular second camera (420) from among a plurality of possible second cameras currently showing the object (440) and selecting the second camera (420) to provide the second video stream.

8. 10. The method of claim 1, The method of claim 1, wherein the generating in the second generating step (S6, S7) is configured to enhance or shift focus of the output digital video stream so as to increase visual focus on the object (440).

9. 10. The method of claim 1, further comprising: a delay step (A3) in which latency is intentionally introduced for at least said first and second digital video streams, said latency being present in said shared digital video stream, wherein: The trigger detection step is performed based on the first and / or second digital video stream before introducing the latency period.

10. 10. The method of claim 9, A method wherein the waiting time is at most 30 seconds, such as at most 5 seconds, such as at most 1 second, such as at most 0.5 seconds.

11. 1. A computer program product causing a computer to perform a process for providing a shared digital video stream, the process causing the computer to perform the following steps when executed: In a collecting step (S2), a first digital video stream is collected from a first digital video source, and a second digital stream is collected from a second digital video source; In a first generating step (S4), generating the shared digital video stream as an output digital video stream based on consecutively considered frames of the first digital video stream such that image information from the first digital video source is visible in the shared digital video stream and image information from the second digital video source is not visible in the shared digital video stream; In a trigger detection step (S5), a trigger is automatically detected by performing at least one of the following: a) digitally analyzing the first digital video stream to automatically detect a trigger in the form of a participant (430) depicted in the first digital video stream gazing towards or making a predetermined gesture in relation to an object (440) depicted in the second digital video stream, wherein the detection is based on information regarding the relative orientations of the first camera (410), the participant (430), and the object (440), and the detection is further based on digital image-based determination of at least one of the participant's (430) body direction, head direction, gaze direction, and gesture direction; and b) digitally analyzing the first and / or second digital video streams to automatically detect a trigger in the form of a transition in the gaze direction of a participant (430) depicted in the second digital video stream from not gazing at a second camera (420) to gazing at the second camera (420); wherein the trigger instructs changing a generation mode of the shared digital video stream according to a predetermined generation rule; In a second generating step (S6, S7) initiated in response to detecting the trigger, the shared digital video stream is generated as an output digital video stream based on a plurality of consecutively considered frames of the second digital video stream, and / or the shared video stream is generated as an output digital video stream based on a plurality of consecutively considered frames of the first digital video source, such that image information from the second digital video source is visible in the shared digital video stream, but with at least one of different cropping, different zooming, different panning, or different focus plane selection of the first digital video stream compared to the first generating step (S4); and In a publishing step (S8), the output digital video stream is continuously made available to consumers of the shared digital video stream.

12. A system (100) for providing a shared digital video stream, the system (100) comprising a central server (130), the central server (130) having the following functions: a collection function configured to collect a first digital video stream from a first digital video source and a second digital stream from a second digital video source; a first generating function (135, 135', 135'', 135''') configured to generate the shared digital video stream as an output digital video stream based on consecutively considered frames of the first digital video stream such that image information from the first digital video source is visible in the shared digital video stream and image information from the second digital video source is not visible in the shared digital video stream; A trigger detection function configured to automatically detect a trigger by performing at least one of the following: a) digitally analyzing the first digital video stream to automatically detect a trigger in the form of a participant (430) depicted in the first digital video stream gazing towards or making a predetermined gesture in relation to an object (440) depicted in the second digital video stream, wherein the detection is based on information regarding the relative orientations of the first camera (410), the participant (430), and the object (440), and the detection is further based on a digital image-based determination of at least one of a body direction, a head direction, a gaze direction, and a gesture direction of the participant (430); and b) digitally analyzing the first and / or second digital video streams to automatically detect a trigger in the form of a transition in the gaze direction of a participant (430) depicted in the second digital video stream from not gazing at a second camera (420) to gazing at the second camera (420); wherein the trigger instructs changing a generation mode of the shared digital video stream according to a predetermined generation rule; a second generation function (135, 135', 135'', 135''') configured to be initiated in response to detection of the trigger and to generate the shared digital video stream as an output digital video stream based on a plurality of consecutively considered frames of the second digital video stream such that image information from the second digital video source is visible in the shared digital video stream, and / or to generate the shared video stream as an output digital video stream based on a plurality of consecutively considered frames of the first digital video source, but configured to perform at least one of different cropping, different zooming, different panning, or different focus plane selection of the first digital video stream compared to the first generation function (135, 135', 135'', 135'''); and A publishing function configured to continuously provide the output digital video stream to consumers of the shared digital video stream.