System and method for generating a video stream
By implementing the collection, synchronization and generation steps in the central server, the time delay technology is used to solve the delay and synchronization problems of video streams in the digital video conferencing system, achieving low-latency and efficient video stream generation, adapting to different encoding formats and resolutions, and improving user experience.
Patent Information
- Application Number
- CN202510635095.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2021-11-30
- Filing Date
- 2022-11-07
- Publication Date
- 2025-08-08
AI Technical Summary
Existing digital video conferencing systems face problems such as delay, synchronization, inconsistent encoding formats, frame rate and resolution differences when generating multiple input video streams, resulting in poor user experience, especially in multiple participants, complex configurations, and different hardware environments, which are difficult to achieve efficient and low-latency video stream generation.
By implementing the collection, synchronization, generation and publishing steps in the central server, the time delay technology is used to synchronize different video streams with the generation stream in time, and generate multiple output video streams with different delays as needed to adapt to the needs of different participants.
It realizes low-latency and efficient video stream generation in a multi-participant environment, improves user experience, adapts to video stream processing in different encoding formats and resolutions, and reduces hardware cost and computing complexity.
Smart Images

Figure CN120455723A_ABST
Abstract
Description
[0001] This disclosure is a divisional application of the invention patent application with application number 202280079471.9 filed on May 30, 2024, and invention name “System and method for generating video streams”. Technical Field
[0002] The present invention relates to a system, computer software product, and method for generating a digital video stream, and more particularly, for generating a digital video stream based on two or more different digital input video streams. In a preferred embodiment, the digital video stream is generated in the context of a digital video conference or a digital video meeting or conferencing system, particularly involving multiple different concurrent users. The generated digital video stream can be distributed externally or within the digital video conference or digital video conferencing system. Background Art
[0003] In other embodiments, the present invention is applied to scenarios other than digital video conferencing, but where several digital video input streams are processed concurrently and combined into a generated digital video stream. For example, such scenarios may be educational or instructional.
[0004] Many digital video conferencing systems are available, such as Microsoft ® Teams ® , Zoom ® and Google ® Meet ® These digital video conferencing systems enable two or more people to meet virtually with digital video and audio that is recorded locally and broadcast to all participants, simulating a physical meeting.
[0005] These digital video conferencing solutions generally need improvement, especially in the generation of viewing content, such as what is shown to whom at what time and via what distribution channels. For example, some systems automatically detect the participant who is currently speaking and display the corresponding video feed of the speaking participant to other participants. In many systems, graphics such as the currently displayed screen, window, or digital presentation can be shared. However, as virtual meetings become more and more complex, it quickly becomes more difficult for services to know which of all the currently available information to show to each participant at each point in time.
[0006] In other examples, participants in a presentation move around the stage while talking about slides in a digital presentation. The system then needs to decide whether to show the presentation, the presenter, or both, or switch between the two.
[0007] It may be desirable to generate one or several output digital video streams based on a plurality of input digital video streams by an automatic generation process and to provide the one or several digital video streams thus generated to one or several consumers.
[0008] However, in many cases, due to the many technical difficulties faced by such digital video conferencing systems, it is difficult for a dynamic conference screen layout manager or other automated generation function to select what information to show.
[0009] First, low latency is crucial due to the real-time nature of digital video conferencing. This creates a problem when different incoming digital video streams (such as from different participants joining using different hardware) are associated with different latencies, frame rates, aspect ratios, or resolutions. Often, these incoming digital video streams need to be processed to achieve a good user experience.
[0010] Secondly, there is the issue of time synchronization. Because various input digital video streams (such as external digital video streams or digital video streams provided by participants) are typically fed into a central server, there is no absolute time synchronization between each of these digital video feeds. Like excessive latency, unsynchronized digital video feeds will result in a poor user experience. Third, multi-party digital video conferencing may involve different digital video streams with different encodings or formats, which need to be decoded and re-encoded, thereby generating delay and synchronization issues. Such encoding is also computationally heavy and therefore costly in terms of hardware requirements.
[0011] Fourth, the fact that different digital video sources may be associated with different frame rates, aspect ratios, and resolutions may also result in memory allocation requirements that may vary unpredictably, requiring constant balancing. This encoding involves heavy computational work, resulting in high hardware costs.
[0012] Fifth, participants may encounter various challenges in terms of variable connectivity, leaving / reconnecting, etc., thus posing further challenges in automatically generating a good user experience.
[0013] These issues are magnified in more complex meeting situations, such as those involving many participants, participants connecting using different hardware and / or software, externally provided digital video streams, screen sharing, or multiple hosts. Corresponding problems arise in said other scenarios, where an output digital video stream is to be generated based on several input digital video streams, such as in digital video generation systems for education and teaching.
[0014] Swedish application SE 2151267-8 (which has not yet been published at the date of effectiveness of the present application) discloses various solutions to the problems discussed above.
[0015] Additional latency-related issues exist in multi-participant digital video environments. In particular, latency requirements may vary between participants. In such environments, it has proven difficult to provide a well-synchronized experience for all participants, without negatively impacting communication due to time delays. This is particularly true in video environments with complex configurations, such as those using intermediately generated multi-participant video streams and / or involving several types of participants. Summary of the Invention
[0016] The present invention solves one or more of the above problems.
[0017] Therefore, the present invention relates to a method for providing a second digital video stream, the method comprising: in a collecting step, collecting a first main digital video stream from a first participant client, collecting a second main digital video stream from a second participant client, and collecting a third main digital video stream from a third participant client; in a publishing step, providing the first main digital video stream, the second main digital video and at least one of the first generated video streams generated based on at least one of the first main video stream and the second main video stream to at least one of the first participant client and the second participant client; in a second generating step, generating the second generated video stream as a digital video stream based on the first main digital video stream, the second main digital video stream and the third main digital video stream, the second generating step introducing a time delay so that the second generated video stream is not synchronized with any video stream provided to the first participant client or the second participant client in the publishing step, and the publishing step also includes continuously providing the second generated video stream to at least one consuming client that is not the first participant client or the second participant client. The present invention also relates to a method for providing a second digital video stream, the method comprising: in a collecting step, collecting a first main digital video stream and a second main digital video stream from at least two different digital video sources; in a first generating step, generating a first generated video stream as a digital video stream based on the first main digital video stream and the second main digital video stream; in a second generating step, generating a second generated video stream as a digital video stream based on the first generated video stream and also based on the first main digital video stream and the second main digital video stream; and in the second generating step, taking into account the delay of the first generated video stream caused by the first generating step, time-delaying the first main digital video stream and the second main digital video stream so that they are time-synchronized with the first generated video stream, and the second generated video stream is generated based on the time-delayed first main digital video stream and the second main digital video stream.
[0018] The present invention also relates to a method for providing a second digital video stream, the method comprising: in a collecting step, collecting a first main digital video stream from a first participant client, collecting a second main digital video stream from a second participant client, and collecting a third main digital video stream from a third participant client; in a first generating step, generating a first generated video stream as a digital video stream based on the first main digital video stream and the second main digital video stream, the first generated digital video stream being continuously generated for publishing with a first delay; in a second generating step, generating a second generated video stream as a digital video stream based on the first main digital video stream, the second main digital video stream and the third main digital video stream, the second generated digital video stream being continuously generated for publishing with a second delay, the second delay being greater than the first delay; and in a publishing step, continuously providing at least one of the first main digital video stream, the second main digital video stream and the first generated video stream to at least one of the first participant client and the second participant client, and continuously providing the second generated video stream to at least one other participant client.
[0019] The invention further relates to a computer software product for providing a second digital video stream, the computer software being arranged to, when run, perform: A collecting step, wherein a first main digital video stream is collected from a first participant client, a second main digital video stream is collected from a second participant client, and a third main digital video stream is collected from a third participant client; a publishing step, wherein at least one of the first main digital video stream, the second main digital video, and a first generated video stream that has been generated based on at least one of the first main video stream and the second main video stream is provided to at least one of the first participant client and the second participant client; a second generating step, wherein a second generated video stream is generated as a digital video stream based on the first main digital video stream, the second main digital video stream, and the third main digital video stream, the second generating step introducing a time delay so that the second generated video stream is not synchronized with any video stream provided to the first participant client or the second participant client in the publishing step, wherein the publishing step further comprises continuously providing the second generated video stream to at least one consuming client that is not the first participant client or the second participant client.
[0020] The present invention also relates to a computer software product for providing a shared digital video stream, the computer software being arranged to, when run, perform: A collecting step, wherein a first main digital video stream and a second main digital video stream are collected from at least two different digital video sources; a first generating step, wherein a first generated video stream is generated as a digital video stream based on the first main digital video stream and the second main digital video stream; a second generating step, wherein a second generated video stream is generated as a digital video stream based on the first generated video stream and also based on the first main digital video stream and the second main digital video stream; and wherein in the second generating step, the first main digital video stream and the second main digital video stream are time-delayed so as to synchronize them with the first generated video stream in time, taking into account a delay of the first generated video stream caused by the first generating step, and the second generated video stream is generated based on the time-delayed first main digital video stream and the second main digital video stream.
[0021] The present invention also relates to a computer software product for providing a shared digital video stream, the computer software being arranged to, when run, perform: A collecting step, wherein a first main digital video stream is collected from a first participant client, a second main digital video stream is collected from a second participant client, and a third main digital video stream is collected from a third participant client; a first generating step, wherein a first generated video stream is generated as a digital video stream based on the first main digital video stream and the second main digital video stream, the first generated digital video stream being continuously generated for publishing with a first delay; a second generating step, wherein a second generated video stream is generated as a digital video stream based on the first main digital video stream, the second main digital video stream and the third main digital video stream, the second generated digital video stream being continuously generated for publishing with a second delay, the second delay being greater than the first delay; and a publishing step, wherein at least one of the first main digital video stream, the second main digital video stream and the first generated video stream is continuously provided to at least one of the first participant client and the second participant client, and the second generated video stream is continuously provided to at least one other participant client.
[0022] The present invention also relates to a system for providing a second digital video stream, which includes a central server, which in turn includes: a collection function, in which a first main digital video stream is collected from a first participant client, a second main digital video stream is collected from a second participant client, and a third main digital video stream is collected from a third participant client; a publishing function, in which the first main digital video stream, the second main digital video and at least one of the first generated video streams generated based on at least one of the first main video stream and the second main video stream are provided to at least one of the first participant client and the second participant client; a second generation function, in which a second generated video stream is generated as a digital video stream based on the first main digital video stream, the second main digital video stream and the third main digital video stream, the second generation step introducing a time delay so that the second generated video stream is not synchronized with any video stream provided to the first participant client or the second participant client in the publishing step, wherein the publishing function includes continuously providing the second generated video stream to at least one consuming client that is not the first participant client or the second participant client.
[0023] Moreover, the present invention relates to a system for providing a shared digital video stream, which includes a central server, which in turn includes: a collection function, wherein a first main digital video stream and a second main digital video stream are collected from at least two different digital video sources; a first generation function, wherein a first generated video stream is generated as a digital video stream based on the first main digital video stream and the second main digital video stream; a second generation function, wherein a second generated video stream is generated as a digital video stream based on the first generated video stream and also based on the first main digital video stream and the second main digital video stream; and wherein in the second generation function, the first main digital video stream and the second main digital video stream are time-delayed so as to synchronize them with the first generated video stream, taking into account the delay of the first generated video stream caused by the first generation function, and the second generated video stream is generated based on the time-delayed first main digital video stream and the second main digital video stream.
[0024] The present invention also relates to a system for providing shared digital video streams, the system comprising a central server, which further comprises: a collection function, wherein a first main digital video stream is collected from a first participant client, a second main digital video stream is collected from a second participant client, and a third main digital video stream is collected from a third participant client; a first generation function, wherein a first generated video stream is generated as a digital video stream based on the first main digital video stream and the second main digital video stream, and the first generated digital video stream is continuously generated for publishing with a first delay; a second generation function, wherein a second generated video stream is generated as a digital video stream based on the first main digital video stream, the second main digital video stream and the third main digital video stream, and the second generated digital video stream is continuously generated for publishing with a second delay, the second delay being greater than the first delay; and a publishing function, wherein at least one of the first main digital video stream, the second main digital video stream and the first generated video stream is continuously provided to at least one of the first participant client and the second participant client, and the second generated video stream is continuously provided to at least one other participant client.
[0025] Furthermore, the invention relates to a system. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Hereinafter, the present invention will be described in detail with reference to exemplary embodiments thereof and the accompanying drawings, in which: Figure 1 A first exemplary system is shown; Figure 2 A second exemplary system is shown; Figure 3 A third exemplary system is shown; Figure 4 A central server is shown; Figure 5 A first method is shown; Figures 6a to 6f Shown with Figure 5 Subsequent states associated with different method steps in the method shown in; Figure 7 The general protocol is shown conceptually; Figures 8a to 8d A second method, a third method, a fourth method, and a fifth method are shown; and Figure 9 A fourth exemplary system is shown.
[0027] All figures share reference numerals for identical or corresponding parts. DETAILED DESCRIPTION
[0028] Figure 1A system 100 is shown which is arranged according to the invention to perform the method according to the invention for providing a digital video stream, such as a shared digital video stream.
[0029] System 100 may include a video communication service 110, but in some embodiments, video communication service 110 may also be external to system 100. As will be discussed, more than one video communication service 110 may be present.
[0030] The system 100 may include one or several participant clients 121 , but in some embodiments one, some, or all of the participant clients 121 may also be external to the system 100 .
[0031] The system 100 may include a central server 130 .
[0032] As used herein, the term "central server" refers to computer-implemented functionality that is arranged to be accessed in a logically centralized manner, such as via a well-defined API (Application Programming Interface). This central server functionality can be implemented entirely in computer software, or in a combination of software and virtual and / or physical hardware. It can be implemented on a standalone physical or virtual server computer, or distributed across several interconnected physical and / or virtual server computers.
[0033] The physical or virtual hardware on which the central server 130 runs (in other words, the computer software that defines the functionality of the central server 130) may include a conventional CPU, a conventional GPU, conventional RAM / ROM memory, a conventional computer bus, and conventional external communication capabilities, such as an Internet connection.
[0034] Each video communication service 110 is also a central server in the sense that it is used, which may be a central server different from the central server 130 or a part of the central server 130 .
[0035] Accordingly, each of the participant clients 121 may be a central server with a corresponding interpretation in the described sense, and the physical or virtual hardware on which each participant client 121 runs (in other words, the computer software that defines the functionality of the participant client 121) may also include a conventional CPU / GPU per se, conventional RAM / ROM memory per se, a conventional computer bus per se, and conventional external communication functionality per se, such as an Internet connection.
[0036] Each participant client 121 typically also includes or is in communication with: a computer screen arranged to display video content provided to the participant client 121 as part of an ongoing video communication; speakers arranged to emit sound content provided to the participant client 121 as part of said video communication; a camera; and a microphone arranged to record sound locally to a human participant 122 of said video communication who is participating in said video communication using the participant client 121 in question. In other words, the respective human-machine interface of each participant client 121 allows the respective participant 122 to interact with the client 121 in question in video communications with other participants and / or audio / video streams provided by various sources. Typically, each of the participant clients 121 includes a respective input device 123, which may include the camera; the microphone; a keyboard; a computer mouse or trackpad; and / or an API for receiving a digital video stream, a digital audio stream, and / or other digital data. The input device 123 is particularly arranged to receive a video stream and / or an audio stream from a central server (such as the video communication service 110 and / or the central server 130), such video stream and / or audio stream being provided as part of the video communication and preferably generated based on corresponding digital data input streams provided to the central server from at least two sources of such digital data input streams (e.g., the participant client 121 and / or an external source (see below)).
[0037] More generally, each of the participant clients 121 includes a corresponding output device 124, which may include the computer screen; the speakers; and an API for emitting a digital video and / or audio stream representing the video and / or audio captured locally to the participant 122 using the participant client 121 in question. In practice, each participant client 121 may be a mobile device, such as a mobile phone, arranged with a screen, speakers, microphone, and internet connection, which runs computer software locally or accesses remotely run computer software to perform the functionality of the participant client 121 in question. Accordingly, a participant client 121 may also be a thick or thin laptop or stationary computer running a locally installed application, with remote access functionality via a web browser, etc. (as the case may be).
[0038] There may be more than one (such as at least three or even at least four) participant clients 121 used in the same video communication of the present type. There may be at least two different groups of participant clients. Each of the participant clients may be assigned to such a corresponding group. These groups may reflect the different roles of the participant clients, the different virtual or physical locations of the participant clients, and / or the different interaction permissions of the participant clients.
[0039] The various such roles available may be, for example, "leader" or "participant", "speaker", "panel participant", "interactive audience" or "remote listener".
[0040] Various available such physical locations may be, for example, "on stage," "in a panel," "in a physically present audience," or "in a physically remote audience."
[0041] The virtual location can be defined based on the physical location, but can also involve virtual groupings that may partially overlap with the physical location. For example, physically present audience members can be divided into a first virtual group and a second virtual group, and some physically present audience participants can be grouped in the same virtual group as some physically distant audience participants.
[0042] The various such interaction permissions available may be, for example, "full interaction" (no restrictions), "can speak but only after requesting a microphone" (such as raising a virtual hand in a video conferencing service), "cannot speak but write in normal chat" or "watch / listen only".
[0043] In some cases, each defined role and / or physical / virtual location can be defined according to certain predetermined interaction permissions. In other cases, all participants with the same interaction permissions form a group. Thus, any defined role, location, and / or interaction permissions can reflect various group assignments, and different groups can be disjoint or overlapping, as appropriate.
[0044] This will be exemplified below. Video communications may be provided at least in part by video communications service 110 and at least in part by central server 130 , as will be described and illustrated herein.
[0045] As the term is used herein, a "video communication" is an interactive digital communication session involving at least two, preferably at least three, or even at least four video streams, and preferably also a matching audio stream for generating one or more mixed or combined digital video / audio streams, which are in turn consumed by one or more consumers (such as participant clients of the type in question), who may or may not contribute to the video communication via video and / or audio. Such video communication is real-time, with or without some delay or latency. At least one, preferably at least two, or even at least four, participants 122 of such a video communication interactively participate in the video communication, providing and consuming video / audio information.
[0046] At least one of the participant clients 121 or all of the participant clients 121 may include local synchronization software functionality 125, which will be described in more detail below.
[0047] The video communication service 110 may include or have access to a common time reference, as will also be described in more detail below.
[0048] Each of the at least one central server 130 may include a respective API 137 for digitally communicating with entities external to the central server 130 in question. Such communication may involve both input and output.
[0049] The system 100 (such as the central server 130) may also be arranged to digitally communicate with an external information source 300 (such as an externally provided video stream), and in particular, receive digital information (such as audio and / or video stream data) from this external information source. The "external" in the external information source 300 refers to the information source not provided by or as part of the central server 130. Preferably, the digital data provided by the external information source 300 is independent of the central server 130, and the central server 130 cannot influence its information content. For example, the external information source 130 may be live video and / or audio, such as video and / or audio from a public sporting event or an ongoing news event or report. The external information source 130 may also be captured by a webcam, etc., but not by any of the participant clients 121. Thus, such captured video may depict the same location as any of the participant clients 121, but not be captured as part of the participant client 121's own activities. One possible difference between the externally provided information source 300 and the internally provided information source 120 is that the internally provided information source can be and is provided in its capabilities as a participant in a video communication of the type defined above, whereas the externally provided information source 300 is not such, but is provided as part of a scene external to the video conference.
[0050] There may also be several external information sources 300 providing digital information of the type described, such as audio and / or video streams, to the central server 130 in parallel.
[0051] like Figure 1 As shown, each of the participant clients 121 may constitute a source of a respective information (video and / or audio) stream 120 that is provided to the video communication service 110 by the participant client 121 in question as described.
[0052] The system 100 (such as the central server 130) can also be arranged to digitally communicate with external consumers 150, and in particular to send digital information to the external consumers. For example, a digital video and / or audio stream generated by the central server 130 can be continuously provided to one or more external consumers 150 in real time or near real time via the API 137. Likewise, the "external" in the external consumer 150 means that the consumer 150 is not provided as part of the central server 130 and / or is not a party to the video communication.
[0053] Unless otherwise stated, all functions and communications herein are provided digitally and electronically, implemented by computer software running on appropriate computer hardware, and communicated over a digital communications network or channel, such as the Internet.
[0054] Therefore, in Figure 1 In the configuration of system 100 shown in FIG, multiple participant clients 121 participate in a digital video communication provided by a video communication service 110. Each participant client 121 may therefore have an ongoing login, session, etc. with the video communication service 110 and may participate in the same ongoing video communication provided by the video communication service 110. In other words, the video communication is "shared" between the participant clients 121 and, therefore, also by the corresponding human participants 122.
[0055] exist Figure 1In the example embodiment, central server 130 includes automated participant client 140, which is an automated client corresponding to participant client 121 but not associated with human participant 122. Instead, automated participant client 140 is added to video communication service 110 as a participant client to participate in the same shared video communication as participant client 121. As such a participant client, automated participant client 140 is granted access to (multiple) continuously generated digital video and / or audio streams provided by video communication service 110 as part of the ongoing video communication and consumed by central server 130 via automated participant client 140. Preferably, automated participant client 140 receives from video communication service 110: a common video and / or audio stream that is or can be distributed to each participant client 121; corresponding video and / or audio streams provided to video communication service 110 from one or more of participant clients 121 and relayed by video communication service 110 in their original or modified form to all or the requesting participant clients 121; and / or a common time reference.
[0056] The central server 130 may include a collection function 131 arranged to receive video and / or audio streams of the type described from the automated participant clients 140 and possibly also from the external information source(s) 300, for processing as described below, and then provide a generated (such as shared) video stream via the API 137. For example, this generated video stream may be consumed by the external consumer 150 and / or by the video communication service 110 for in turn distribution by the video communication service 110 to all or any requested one of the participant clients 121.
[0057] Figure 2 Similar to Figure 1 , but instead of using automated client participants 140 , the central server 130 receives video and / or audio stream data from the ongoing video communication via the API 112 of the video communication service 110 .
[0058] Figure 3 Also similar to Figure 1 , but the video communication service 110 is not shown. In this case, the participant clients 121 communicate directly with the API 137 of the central server 130, for example, to provide video and / or audio streaming data to the central server 130 and / or receive video and / or audio streaming data from the central server 130. The generated shared stream can then be provided to the external consumer 150 and / or to one or more of the client participants 121.
[0059] Figure 4The central server 130 is shown in more detail. As shown, the collection function 131 may include one or preferably several format-specific collection functions 131a. Each of the format-specific collection functions 131a may be arranged to receive a video and / or audio stream having a predetermined format, such as a predetermined binary encoding format and / or a predetermined stream data container, and may be particularly arranged to parse the binary video and / or audio data in the format into individual video frames, video frame sequences and / or time slots. The central server 130 may also include an event detection function 132, which is configured to receive video and / or audio stream data (such as binary stream data) from the collection function 131 and perform corresponding event detection on each individual data stream in the received data stream. The event detection function 132 may include an AI (Artificial Intelligence) component 132a for performing the event detection. This event detection can be performed without first time-synchronizing the individual collected streams. The central server 130 further comprises a synchronization function 133 arranged to time synchronize the data streams provided by the collection function 131 and possibly processed by the event detection function 132. The synchronization function 133 may comprise an AI component 133a for performing said time synchronization.
[0060] The central server 130 may also include a pattern detection function 134 that is arranged to perform pattern detection based on a combination of at least one, but in many cases at least two, such as at least three or even at least four, such as all, of the received data streams. Pattern detection may also be based on one, or in some cases at least two or more events detected by the event detection function 132 for each individual one of the data streams. Such detected events considered by the pattern detection function 134 may be distributed over time relative to each individual collected stream. The pattern detection function 134 may include an AI component 134a for performing the pattern detection. Pattern detection may also be based on the grouping discussed above, and may be specifically arranged to detect specific patterns that occur only with respect to one group, only with respect to certain groups but not all groups, or with respect to all groups.
[0061] The central server 130 also includes a generation function 135 that is configured to generate a generated digital video stream, such as a shared digital video stream, based on the data stream provided from the collection function 131 and possibly also based on any detected events and / or patterns. Such a generated video stream may include at least a video stream generated to include one or more video streams (original, reformatted, or transformed) provided by the collection function 131, and may also include corresponding audio stream data. As will be illustrated below, there may be several generated video streams, one of which may be generated in the manner discussed above, but further based on another already generated video stream. The entire generated video stream is preferably generated continuously, and preferably in near real time (after deducting any delays and latencies of the type discussed below).
[0062] The central server 130 may further comprise a publishing function 136 arranged to publish the generated digital video stream, for example via an API 137 as described above.
[0063] Notice, Figure 1 、 Figure 2 and Figure 3 It shows how the principles described herein may be implemented using a central server 130 and in particular provides three different examples of the method according to the invention, although other configurations are possible with or without the use of one or more video communication services 110 .
[0064] therefore, Figure 5 A method for providing the generated digital video stream is shown. Figures 6a to 6f Shown by Figure 5 The method steps shown generate different digital video / audio data stream states.
[0065] In a first step, the method starts.
[0066] In a subsequent collection step, respective primary digital video streams 210, 301 are collected from at least two of the digital video sources 120, 300, such as by the collection function 131. Each such primary data stream 210, 301 may include an audio portion 214 and / or a video portion 215. It should be understood that, in this context, "video" refers to the moving and / or still image content of such data streams. Each primary data stream 210, 301 may be encoded according to any video / audio coding specification (using the respective codec used by the entity providing the primary data stream 210, 301 in question), and the encoding format may differ between different streams in the primary data stream 210, 301 used concurrently in the same video communication. Preferably, at least one (such as all) of the primary data streams 210, 301 is provided as a binary data stream, possibly in a conventional data container data structure. Preferably, at least one (such as at least two) or even all of the primary data streams 210, 301 are provided as respective live video recordings.
[0067] Note that the main streams 210, 301 may not be synchronized in time when they are received by the collection function 131. This may mean that they are associated with different delays or latencies relative to each other. For example, if the two main video streams 210, 301 are live recordings, this may mean that they are associated with different delays relative to the time of recording when they are received by the collection function 131.
[0068] Note also that the main streams 210, 301 themselves can be respective live camera feeds from a webcam, a currently shared screen or presentation, a watched movie clip, etc., or any combination of these arranged in the same screen in various ways.
[0069] The collection step is Figure 6a and Figure 6b In Figure 6b , it is also shown how the collection function 131 can store each primary video stream 210, 301 as bundled audio / video information or as separate audio stream data from the associated video stream data. Figure 6b It shows how the main video stream 210, 301 data is stored as individual frames 213 or collections / clusters of frames, where "frame" refers here to a time-bound portion of image data and / or any associated audio data, such as each frame being an individual still image or a continuous series of images that together form moving image video content (such as a series of moving images constituting up to 1 second).
[0070] In a subsequent event detection step performed by the event detection function 132, the primary digital video stream 210, 301 is analyzed, such as by the event detection function 132 and in particular the AI component 132a, to detect at least one event 211 selected from the first set of events. Figure 6c Shown in.
[0071] Preferably, this event detection step can be performed for at least one, such as at least two, such as all primary video streams 210, 301, and can be performed separately for each such primary video stream 210, 301. In other words, the event detection step is preferably performed for said individual primary video stream 210, 301 taking into account only information contained as part of this particular primary video stream 210, 301 in question, and in particular without taking into account information contained as part of other primary video streams. Furthermore, event detection is preferably performed without taking into account any common time reference 260 associated with the several primary video streams 210, 301.
[0072] On the other hand, event detection preferably considers information contained as part of the separately analyzed main video stream in question within a certain time interval, such as a historical time interval of the main video stream longer than 0 seconds, such as at least 0.1 seconds, such as at least 1 second. Event detection may take into account information contained in the audio and / or video data contained as part of the primary video stream 210 , 301 .
[0073] The first set of events may include any number of event types, such as a change of slides in a slide presentation that constitutes or is part of the primary video stream 210, 301 in question; a change in the quality of connectivity of the source 120, 300 providing the primary video stream 210, 301 in question that results in a change in image quality, a loss of image data, or restoration of image data; and a detected moving physical event in the primary video stream 210, 301 in question, such as movement of a person or object in the video, a change in lighting in the video, a sudden sharp noise in the audio, or a change in audio quality. It should be understood that these examples are not intended to be exhaustive, but rather to facilitate an understanding of the applicability of the presently described principles.
[0074] In a subsequent synchronization step performed by the synchronization function 133, the main digital video stream 210 is time synchronized. This time synchronization can be relative to a common time reference 260. Figure 6dAs shown, time synchronization may involve, for example, aligning the primary video streams 210, 301 with each other using the common time reference 260 so that they can be combined to form a time-synchronized scene. The common time reference 260 can be a data stream, a heartbeat signal or other pulse data, or a time anchor applicable to each of the individual primary video streams 210, 301. The common time reference can be applied to each of the individual primary video streams 210, 301 in such a way that the information content of the primary video stream 210, 301 in question can be clearly related to the common time reference relative to a common time axis. In other words, the common time reference can allow the primary video streams 210, 301 to be aligned by time shifting so as to be time-synchronized in the current sense. In other embodiments, time synchronization can be based on known information about the time difference between the primary video streams 210, 301 in question, such as based on measurements.
[0075] like Figure 6d As shown, time synchronization may include determining one or several timestamps 261 for each main video stream 210 , 301 , such as relative to a common time reference 260 ; or determining one or several timestamps for each video stream 210 , 301 relative to another video stream 210 , 301 or relative to other video streams 210 , 301 .
[0076] In a subsequent pattern detection step performed by the pattern detection function 134, the thus time-synchronized primary digital video stream 210, 301 is analyzed to detect at least one pattern 212 selected from the first set of patterns. Figure 6e Shown in.
[0077] In contrast to the event detection step, the pattern detection step may preferably be performed based on video and / or audio information contained as part of at least two of the jointly considered time-synchronized primary video streams 210, 301.
[0078] The first set of patterns may include any number of types of patterns, such as several participants talking interchangeably or simultaneously, or presentation slide changes occurring concurrently as different events (such as different participants speaking). This list is not exhaustive, but illustrative.
[0079] In an alternative embodiment, the detected pattern 212 may not relate to information contained in several of the primary video streams 210, 301, but rather to information contained in only one of the primary video streams 210, 301. In this case, it is preferred to detect such a pattern 212 based on video and / or audio information contained in a single primary video stream 210, 301 spanning at least two detected events 211 (e.g., two or more consecutively detected presentation slide changes or connection quality changes). As an example, several consecutive slide changes that follow each other quickly over time may be detected as one single slide change pattern, rather than a separate slide change pattern for each detected slide change event.
[0080] It is recognized that the first event set and the first pattern set may include events / patterns of predetermined types defined using corresponding parameter groups and parameter intervals. As will be explained below, the events / patterns in the set may also or additionally be defined and detected using various AI tools.
[0081] In a subsequent generation step performed by the generation function 135 , a shared digital video stream is generated as an output digital video stream 230 based on the successively considered frames 213 of the time-synchronized primary digital video streams 210 , 301 and said detected patterns 212 .
[0082] As will be explained and detailed below, the present invention allows for fully automatic generation of video streams, such as the output digital video stream 230 .
[0083] For example, such generation may involve selecting what video and / or audio information from what primary video stream 210, 301 to use and to what extent in such output video stream 230, the video screen layout of the output video stream 230, switching patterns between different such uses or layouts over time, etc.
[0084] This is Figure 6f , which also shows one or more additional time-related (which may be related to the common time reference 260) digital video information 220, such as additional digital video information streams, that may be time-synchronized (such as to a common time reference 260) and used with the time-synchronized primary video stream 210, 301 in generating the output video stream 230. For example, the additional stream 220 may include information about any video and / or audio effects to be used, such as dynamically based on detected patterns, a planned schedule for the video communication, etc.
[0085] The generated output digital video stream 230 is continuously provided to the generated digital video stream consumers 110, 150 in a subsequent publishing step performed by the publishing function 136, as described above. The generated digital video stream may be provided to one or several participant clients 121, such as via the video communication service 110.
[0086] In the subsequent step, the method ends. However, first, the method can be iterated any number of times, such as Figure 5 As shown, the output video stream 230 is generated as a continuously provided stream. Preferably, the output video stream 230 is generated for consumption in real time or near real time (taking into account the total delay added by all the steps along the way) and continuously (publishing occurs immediately when more information becomes available, but not counting the intentionally added delays or latencies described below). In this way, the output video stream 230 can be consumed in an interactive manner, so that the output video stream 230 can be fed back into the video communication service 110 or into any other scenario that forms the basis for the generation of the main video stream 210 that is fed back to the collection function 131 to form a closed loop feedback; or so that the output video stream 230 can be consumed in a different scenario (external to the system 100 or at least external to the central server 130), but there forming the basis for real-time, interactive video communication.
[0087] As mentioned above, in some embodiments, at least two (such as at least three, such as at least four, or even at least five) of the primary digital video streams 210, 301 are provided as part of a shared digital video communication (such as provided by the video communication service 110), the video communication involving the respective remotely connected participant clients 121 providing the primary digital video streams 210 in question. In this case, the collecting step may include collecting at least one of the primary digital video streams 210 from the shared digital video communication service 110 itself, such as via an automated participant client 140, which in turn is authorized to access video and / or audio stream data from within the video communication service 110 in question; and / or via the API 112 of the video communication service 110.
[0088] Furthermore, in this and other cases, the collecting step may include collecting at least one of the primary digital video streams 210, 301 as a corresponding external digital video stream 301 collected from an information source 300 external to the shared digital video communication service 110. Note that one or several of such external video sources 300 used may also be external to the central server 130.
[0089] In some embodiments, the primary video streams 210, 301 are not formatted in the same manner. This different formatting can be in the form of them being delivered to the collection function 131 in different types of data containers (such as AVI or MPEG), but in a preferred embodiment, at least one of the primary video streams 210, 301 is formatted according to a deviated format (compared to at least one other of the primary video streams 210, 301) insofar as the deviated primary digital video streams 210, 301 have a deviated video encoding, a deviated fixed or variable frame rate, a deviated aspect ratio, a deviated video resolution, and / or a deviated audio sampling rate.
[0090] Preferably, the collection function 131 is pre-configured to read and interpret all encoding formats, container standards, etc. present in all collected primary video streams 210, 301. This makes it possible to perform the processing described herein without requiring any decoding until relatively late in the process (such as without requiring any decoding until after the primary stream in question has been placed into the corresponding buffer, without requiring any decoding until after the event detection step, or even without requiring any decoding until after the event detection step). However, in the rare case where one or more of the primary video feeds 210, 301 are encoded using a codec that the collection function 131 cannot interpret without decoding, the collection function 131 can be arranged to perform decoding and analysis of such primary video streams 210, 301, followed by conversion into a format that can be processed by, for example, the event detection function. Note that even in this case, it is preferred that no re-encoding is performed at this stage.
[0091] For example, the primary video stream 220 obtained from a multi-party video event, such as that provided by the video communication service 110, typically has requirements regarding low latency and, therefore, is typically associated with a variable frame rate and variable pixel resolution to enable effective communication for the participants 122. In other words, the overall video and audio quality will be degraded as necessary for low latency.
[0092] On the other hand, the external video feed 301 will typically have a more stable frame rate, higher quality, but therefore may have higher latency.
[0093] Therefore, the video communication service 110 may use a different encoding and / or container at each moment than the external video source 300. Therefore, in this case, the analysis and video generation process described herein needs to combine these differently formatted streams 210, 301 into a new stream to obtain a combined experience.
[0094] As mentioned above, the collection function 131 may comprise a set of format-specific collection functions 131 a, each of which is arranged to process a specific type of format of the primary video stream 210, 301. For example, each of these format-specific collection functions 131 a may be arranged to process a video stream that has been encoded using a different video encoding method / codec (such as Windows ® Media ® or DivX) encoded main video stream 210, 301. However, in a preferred embodiment, the collecting step comprises converting at least two (eg all) of the primary digital video streams 210 , 301 to a common protocol 240 .
[0095] As used in this context, the term "protocol" refers to an information structuring standard or data structure that specifies how the information contained in a digital video / audio stream is stored. However, the common protocol preferably does not specify how digital video and / or audio information (i.e., the encoded / compressed data indicative of the sounds and images themselves) is stored at the binary level, but instead forms a structure in a predetermined format for storing such data. In other words, the common protocol provides for the storage of digital video data in its raw binary form without performing any digital video decoding or digital video encoding associated with such storage, possibly by not modifying the existing binary form at all, other than possibly concatenating and / or splitting byte sequences in the binary form. Instead, the original (encoded / compressed) binary data content of the primary video stream 210, 301 in question is preserved, while repackaging this raw binary data into a data structure defined by the protocol. In some embodiments, the common protocol defines a video file container format.
[0096] As an example, Figure 7 shows the reconstruction and use of the common protocol 240 by the corresponding format-specific collection function 131a. Figure 6a The main video streams 210, 301 are shown in FIG.
[0097] Thus, the common protocol 240 provides for the storage of digital video and / or audio data in data sets 241, which are preferably divided into discrete, continuous data sets along a timeline associated with the primary video stream 210, 301 in question. Each such data set may include one or several video frames, and associated audio data.
[0098] The common protocol 240 may also provide for storing metadata 242 associated with a specified point in time associated with the stored digital video and / or audio data set 241 .
[0099] The metadata 242 may include information about the original binary format of the primary digital video stream 210 in question, such as information about the digital video encoding method or codec used to generate the raw binary data, the resolution of the video data, the video frame rate, the frame rate variability flag, the video resolution, the video aspect ratio, the audio compression algorithm, or the audio sampling rate. The metadata 242 may also include information about the timestamps of the stored data (such as information related to the time base of the primary video stream 210, 301 in question or related to different video streams as discussed above).
[0100] The use of the format-specific collection function 131a in conjunction with the common protocol 240 allows for rapid collection of the information content of the primary video stream 210, 301 by decoding / re-encoding the received video / audio data without increasing delay.
[0101] Thus, the collecting step may comprise collecting primary digital video streams 210, 301 that are encoded using different binary video and / or audio encoding formats using different ones of the format-specific collecting functions 131a to parse the primary video stream 210, 301 in question and store the parsed raw binary data and any associated metadata in a data structure using a common protocol. It goes without saying that the determination of which format-specific collecting function 131a to use for which primary video stream 210, 301 may be performed by the collecting function 131 based on predetermined and / or dynamically detected properties of each primary video stream 210, 301 in question.
[0102] Each primary video stream 210 , 301 thus collected may be stored in its own separate storage buffer (such as a RAM storage buffer) in the central server 130 .
[0103] The conversion of the primary video stream 210 , 301 performed by each format-specific collection function 131 a may thus comprise splitting the raw binary data of each such converted primary digital video stream 210 , 301 into an ordered set of said smaller data sets 241 .
[0104] Furthermore, the conversion may also include associating each of the smaller sets 241 (or a subset, such as a regularly distributed subset along the respective timeline of the main stream 210, 301 in question) with a respective time along a shared timeline (such as with respect to the common time reference 260). This association may be performed by analyzing the raw binary video and / or audio data in any of the principle manners described below, or otherwise, and may be performed to enable subsequent time synchronization of the main video streams 210, 301. Depending on the type of common time reference used, at least a portion of this association of each of the data sets 241 may also or alternatively be performed by the synchronization function 133. In the latter case, the collection step may alternatively include associating each or a subset of the smaller sets 241 with a respective time along the particular timeline of the main stream 210, 301 in question.
[0105] In some embodiments, the collection step further includes converting the raw binary video and / or audio data collected from the primary video streams 210, 301 to a uniform quality and / or update frequency. This may involve downsampling or upsampling the raw binary digital video and / or audio data of the primary digital video streams 210, 301 to a common video frame rate, a common video resolution, or a common audio sampling rate, as needed. Note that such resampling can be performed without performing a full decoding / re-encoding, or even without performing any decoding at all, since the format-specific collection function 131a in question can directly process the raw binary data according to the correct binary encoding target format.
[0106] Each of the primary digital video streams 210 , 301 may be stored in a separate data storage buffer 250 as a separate frame 213 or sequence of frames 213 as described above and each associated with a corresponding timestamp, which in turn is associated with the common time reference 260 .
[0107] In the specific example provided to illustrate these principles, the video communication service 110 is a Microsoft Windows XP Professional running a video conference involving concurrent participants 122. ® Teams ® The automated participant client 140 is registered as Teems ® Meeting participants in a meeting.
[0108] The main video input signals 210 are then available to and obtained by the collection function 130 via the automated participant client 140. These are raw signals in H264 format and contain time stamp information for each video frame.
[0109] The associated format-specific collection function 131a collects raw data over IP (LAN network in the cloud) on a configurable predefined TCP port. ® The meeting participants and associated audio data are associated with individual ports. The collection function 131 then uses the timestamps from the audio signal (which is at 50 Hz) and downsamples the video data to a fixed output signal of 25 Hz before storing the video stream 220 in its respective individual buffer 250 .
[0110] As mentioned, the public protocol 240 can store data in raw binary form. It can be designed to be very low-level and used to process the raw bits and bytes of video / audio data. In a preferred embodiment, the data is stored in the public protocol 240 as a simple byte array or corresponding data structure (such as a slice). This means that there is no need to put the data into a conventional video container (the public protocol 240 does not constitute a conventional container in this scenario). Moreover, encoding and decoding video is computationally intensive, which means it causes delays and requires expensive hardware. Moreover, this problem scales with the number of participants.
[0111] Using the common protocol 240, it becomes possible to provide a common protocol for each Teams in the collection function 131. ® Memory is reserved for the primary video stream 210 associated with the meeting participant 122, as well as for any external video sources 300, and the amount of allocated memory is then dynamically changed during the process. In this way, the number of input streams can be changed and each buffer can be kept valid as a result. For example, since information such as resolution, frame rate, etc. may be variable but stored as metadata in the common protocol 240, this information can be used to quickly adjust the size of each buffer as may be needed.
[0112] The following is an example of a specification of a public protocol 240 of this type: Byte Example Description 1 byte 1 0 = video; 1 = audio 4 bytes 1234567 buffer length (int) 8 bytes 424234234 timestamp from the incoming audio / video buffer Measured in ticks, 1 tick = 100 ns. (long int) 1 byte 0 VideoColorFormat { NV12 = 0, Rgb24 = 1, Yuy2 = 2, H264 = 3} 4 bytes 720 video frame pixel height (int) 4 bytes 640 video frame pixel width (int) 4 bytes 25.0 Video frame rate Frames per second (floating point) 1 byte 0 Is the audio muted? 1 = true; 0 = false 1 byte 0 AudioFormat { 0 = Pcm16K 1 = Pcm44KStereo } 1 byte 0 Detected event (if any) 0 = No event 1, 2, 3, etc. = Events of the specified type detected 30 bytes reserved for future use 8 bytes 1000000 Length of binary data in bytes (long int) Variable 0x87A879… Raw binary video / audio data for this frame(s) 4 bytes 1234567 Primary speaker port 4 bytes 1234567 active speaker In the above, "detected events (if any)" data is included as part of the specification of the common protocol 260. However, in some embodiments, this information (regarding detected events) may instead be placed in a separate memory buffer.
[0113] In some embodiments, the at least one piece of additional digital video information 220 , which may be an overlay or effect, is also stored in respective separate buffers 250 as separate frames or sequences of frames each associated with a corresponding timestamp, which in turn is associated with the common time reference 260 .
[0114] As exemplified above, the event detection step may comprise storing, using said common protocol 240, metadata 242 describing the detected event 211 in association with the primary digital video stream 210, 301 in which the event 211 in question was detected.
[0115] Event detection can be performed in different ways. In some embodiments, the event detection step performed by the AI component 132a includes a first trained neural network or other machine learning component individually analyzing at least one (such as several or even all) of the primary digital video streams 210, 301 to automatically detect any of the events 211. This can involve the AI component 132a classifying the primary video stream 210, 301 data into a set of predefined events in a managed classification and / or into a set of dynamically determined events in an unmanaged classification. In some embodiments, the detected event 211 is a change of the primary video stream 210, 301 in question or a presentation slide in a presentation included in the primary video stream in question.
[0116] For example, if the presenter of a presentation decides to change the slides in the presentation that he / she is currently giving to the audience, this means that the content of interest to a given audience may change. It is possible that the newly shown slides are just high-level pictures that can be best viewed concisely in so-called "butterfly" mode (e.g., displaying the slides alongside the presenter's video in the output video stream 230). Alternatively, the slides may contain a lot of details, text with a smaller font size, etc. In this latter case, the slides should instead be presented in full screen and for a longer period of time than usual. Butterfly mode may not be appropriate because, in this case, the viewer of the presentation may be more interested in the slides than in the presenter's face.
[0117] In practice, the event detection step may include at least one of the following: First, the event 211 can be detected based on an image analysis of the difference between a first detected image of the slide and a subsequent second detected image of the slide. Conventional digital image processing per se (such as using motion detection in conjunction with OCR (Optical Character Recognition)) can be used to automatically determine the nature of the main video stream 220, 301 as showing a slide.
[0118] This may involve using automatic computer image processing techniques to check whether the detected slide has changed significantly enough to actually classify it as a slide change. This can be done by checking the difference between the current slide and the previous slide with respect to RGB color values. For example, one can evaluate the degree to which the RGB values have changed globally in the area of the screen covered by the slide in question, and whether it is possible to find groups of pixels that belong together and change consistently. In this way, relevant slide changes can be detected while, for example, filtering out irrelevant changes, such as computer mouse movements shown on the screen. This approach also allows for full configurability - for example, it is sometimes desirable to be able to capture computer mouse movements, such as when a presenter wants to demonstrate something in detail using a computer mouse pointing at different things.
[0119] Secondly, the event 211 may be detected based on image analysis of the information complexity of the second image itself, thereby determining the type of event with greater specificity.
[0120] For example, this might involve assessing the amount of textual information on the slide in question and the associated font size. This can be accomplished using conventional OCR methods, such as deep learning-based character recognition techniques.
[0121] Note that since the raw binary format of the video stream 210, 301 being evaluated is known, this can be performed directly in the binary domain without first decoding or re-encoding the video data. For example, the event detection function 132 can call relevant format-specific collection functions for image interpretation services, or the event detection function 132 itself can include functionality for evaluating image information (such as at the individual pixel level) for a plurality of different supported raw binary video data formats.
[0122] In another example, the detected event 211 is a loss of a communication connection of a participant client 121 to the digital video communication service 110. The detecting step may then comprise detecting that the participant client 121 has lost the communication connection based on an image analysis of a series of subsequent video frames 213 of the primary digital video stream 210 corresponding to the participant client 121 in question.
[0123] Because participant clients 121 are associated with different physical locations and different internet connections, it may happen that someone will lose connection with the video communication service 110 or with the central server 130. In such a situation, it is desirable to avoid displaying a black or blank screen in the generated output video stream 230.
[0124] Instead, such a connection loss can be detected as an event by event detection functionality 132, such as by applying a two-class classification algorithm where the two classes used are connected / not connected (no data). In this case, it is understood that "no data" is different from a presenter intentionally displaying a black screen. Because a brief black screen (such as a black screen of only one or two frames) may not be noticeable in the final generated stream 230, one can apply the two-class classification algorithm over time to create a time series. A threshold specifying a minimum length for a connection interruption can then be used to determine whether a lost connection has occurred. As will be explained below, these illustrative types of detected events may be used by pattern detection functionality 134 to take various actions as appropriate and desired.
[0125] As mentioned above, the separate primary video streams 210, 301 may each be related to a common time reference 260 or to each other in the time domain, which enables the synchronization function 133 to time synchronize them with each other.
[0126] In some embodiments, the common time reference 260 is based on or includes the common audio signal 111 (see Figures 1 to 3 ), the common audio signal 111 is common to a shared digital video communication service 110 involving at least two remotely connected participant clients 121 each providing a respective one of said primary digital video streams 210 as described above.
[0127] In the Microsoft discussed above ® Teams ® In the example of FIG, a common audio signal is generated via the automated participant client 140 and / or via the API 112 and can be captured by the central server 130. In this and other examples, this common audio signal can be used as a heartbeat signal to time-synchronize the various primary video streams by tying the individual primary video streams 220 to a specific point in time based on this heartbeat signal. This common audio signal can be provided as a separate signal (relative to each of the other primary video streams 210), whereby the other primary video streams 210 can each be individually time-correlated with the common audio signal based on the audio contained in the other primary video stream 210 in question, or even based on the image information contained therein (such as using lip sync techniques based on automatic image processing).
[0128] In other words, this common audio signal is used as a heartbeat for all primary video streams 210 (but possibly not the external primary video stream 301) in the central server 130 in order to handle any variable and / or different delays associated with the individual primary video streams 210, and to achieve time synchronization for the combined video output stream 230. In other words, all other signals are mapped to this common audio time heartbeat to ensure that everything is in time synchronization. In a different example, time synchronization is achieved using a time synchronization element 231 that is introduced into the outgoing digital video stream 230 and detected by a respective local time synchronization software function 125 provided as part of one or several individual participant clients 121, the local software function 125 being arranged to detect the time of arrival of the time synchronization element 231 in the outgoing video stream 230. As will be appreciated, in such an embodiment, the outgoing video stream 230 is fed back into the video communication service 110 or otherwise made available to each participant client 121 and the local software function 125 in question.
[0129] For example, the time synchronization elements 231 may be visual markers, such as pixels placed or updated at regular time intervals in the output video 230 that change color in a predetermined sequence or manner; a visual clock that is updated and displayed in the output video 230; or an acoustic signal (which may be designed to be inaudible to the participant 122 by, for example, having a sufficiently low amplitude and / or a sufficiently high frequency) added to the audio forming part of the output video stream 230. The local software function 125 is arranged to automatically detect the respective arrival time of each of the time synchronization element(s) 231 (or detect each of the time synchronization element(s)) using appropriate image and / or audio processing.
[0130] A common time reference 260 can then be determined based at least in part on the detected arrival times. For example, each of the local software functions 125 can communicate to the central server 130 respective information representing the detected arrival times. Such communication may occur via a direct communication link between the participant client 121 in question and the central server 130. However, communication may also occur via a primary video stream 210 associated with the participant client 121 in question. For example, the participant client 121 may introduce a visual or audible code, such as the type discussed above, into the primary video stream 210 generated by the participant client 121 in question for automatic detection by the central server 130 and for use in determining the common time reference 260.
[0131] In yet another additional example, each participant client 121 can perform image detection in a common video stream available for viewing by all participant clients 121 of the video communication service 110, and relay the results of this image detection to the central server 130 in a manner corresponding to the manner discussed above, for use therein in determining the respective offsets of each participant client 121 relative to one another over time. In this way, the common time reference 260 can be determined as a set of individual relative offsets. For example, a selected reference pixel of a commonly available video stream can be monitored by several or all participant clients 121, such as by the local software function 125, and the current color of this pixel can be communicated to the central server 130. The central server 130 can calculate a corresponding time series based on such color values continuously received from each of the multiple (or all) participant clients 121 and perform cross-correlations, thereby resulting in a set of estimates of the relative time offsets between different participant clients 121.
[0132] In practice, the output video stream 230 fed into the video communication service 110 may be included as part of the shared screen of each participant client of the video communication in question and may therefore be used to assess such a time offset associated with the participant client 121. In particular, the output video stream 230 fed into the video communication service 110 may be made available again to the central server via the automated participant client 140 and / or the API 112.
[0133] In some embodiments, the common time reference 260 can be determined at least in part based on a detected difference between the audio portion 214 of a first primary digital video stream in the primary digital video streams 210, 301 and the image portion 215 of the first primary digital video stream 210, 301. For example, such a difference can be based on an analysis of a digital lip-synced video image of the speaking participant 122 viewed in the first primary digital video stream 210, 301 in question. Such lip-sync analysis is also conventional and can, for example, use a trained neural network. The analysis can be performed by the synchronization function 133 for each primary video stream 210, 301 relative to the available common audio information, and based on this information, the relative offset between the respective primary video streams 210, 301 can be determined.
[0134] In some embodiments, the synchronization step includes intentionally introducing a delay of up to 30 seconds (such as up to 5 seconds, such as up to 1 second, such as up to 0.5 seconds, but longer than 0 seconds) (in this context, the terms "delay" and "latency" are intended to mean the same thing) so that the output digital video stream 230 is provided with at least this delay. In any case, the intentionally introduced delay is at least a few video frames, such as at least three, or even at least five, or even 10 video frames, such as this number of frames (or individual images) stored after any resampling in the collection step. As used herein, the term "intentionally" refers to the introduction of the delay regardless of whether it is necessary due to synchronization issues, etc. In other words, the intentionally introduced delay is in addition to any delay introduced as part of synchronization of the primary video streams 210, 301 in order to synchronize the primary video streams relative to each other. The intentionally introduced delay can be predetermined, fixed, or variable relative to the common time reference 260. The delay time may be measured relative to the least delayed of the primary video streams 210, 301, such that the more delayed of these streams 210, 301 is associated with a relatively smaller intentionally increased delay as a result of said time synchronization.
[0135] In some embodiments, a relatively small delay is introduced, such as a delay of 0.5 seconds or less. This delay will be barely noticeable to participants of the video communication service 110 using the output video stream 230. In other embodiments, such as when the output video stream 230 is not to be used in an interactive scenario but is instead published to external consumers 150 in a one-way communication, a larger delay may be introduced.
[0136] This intentionally introduced delay may be sufficient to allow sufficient time for the synchronization function 133 to map the collected various main stream 210, 301 video frames onto the correct common time base 260 timestamps 261. It may also be sufficient to allow sufficient time to perform the event detection described above, such as to detect a lost main stream 210, 301 signal, slide changes, resolution changes, etc. Additionally, the intentionally introduced delay may be sufficient to allow for an improved pattern detection function 134, as will be described below.
[0137] It is recognized that the introduction of said delay may involve buffering 250 each of the collected and time-synchronized primary video streams 210, 301 before publishing the output video stream 230 using the buffered frames 213 in question. In other words, the video and / or audio data of at least one, several or even all of the primary video streams 210, 301 may then be present in a buffered manner in the central server 130, much like a cache but not (like a conventional cache buffer) for the purpose of being able to handle varying bandwidth conditions, but used (in particular to be used by the pattern detection function 134) for the reasons described above.
[0138] Thus, in some embodiments, the pattern detection step includes taking into account specific information of at least one (such as several, such as at least four, or even all) of the primary digital video streams 210, 301, which is present in a frame 213 subsequent to the frame of the time-synchronized primary digital video stream 210 that will also be used to generate the output digital video stream 230. Thus, during a specific delay before forming part of (or serving as a basis for) the output video stream 230, the newly added frame 213 will be present in the buffer 250 in question. During this period of time, the information in the frame 213 in question will constitute "future" information relative to the frame currently being used to generate the current frame of the output video stream 230. Once the output video stream 230 timeline reaches the frame 213 in question, it will be used to generate the corresponding frame of the output video stream 230 and can thereafter be discarded.
[0139] In other words, the pattern detection function 134 can have a set of video / audio frames 213 that have not yet been used to generate the output video stream 230 and can use this data to detect the pattern.
[0140] Pattern detection can be performed in different ways. In some embodiments, the pattern detection step performed by the AI component 134a includes a second trained neural network or other machine learning component analyzing at least two (such as at least three, such as at least four, or even all) of the primary digital video streams 120, 301 together to automatically detect the pattern 212.
[0141] In some embodiments, the detected pattern 212 includes speaking patterns involving at least two (such as at least three, such as at least four) different speaking participants 122 (each associated with a respective participant client 121) of the shared video communication service 110, each of the speaking participants 122 being visually viewed in a respective one of the primary digital video streams 210, 301.
[0142] The generation step preferably includes determining, tracking, and updating the current generated state of the output video stream 230. For example, such a state may indicate which (if any) participants 122 are visible in the output video stream 230 and at what location on the screen; whether any external video streams 300 are visible in the output video stream 230 and at what location on the screen; whether any slides or shared screens are shown in full-screen mode or in conjunction with any real-time video streams, etc. Thus, the generation function 135 can be viewed as a state machine with respect to the generated output video stream 230.
[0143] In order to generate the output video stream 230 as a combined video experience to be viewed by, for example, the end consumer 150 , it is advantageous for the central server 130 to be able to understand what is happening at a deeper level than simply detecting individual events associated with the individual primary video streams 210 , 301 .
[0144] In the first example, presentation participant client 121 is changing the currently viewed slide. As described above, this slide change is detected by event detection functionality 132, and metadata 242 is added to the frame in question, indicating that a slide change has occurred. This occurs multiple times because presentation participant client 121 ends up skipping forward multiple slides in rapid succession, resulting in a series of "slide change" events detected by event detection functionality 132 and stored in a separate buffer 250 of the primary video stream 210 in question, along with the corresponding metadata 242. In reality, each of these rapidly skipped slides may only be visible for a fraction of a second. The pattern detection function 134, reviewing the information in the buffer 250 in question, spanning several of these detected slide changes, will detect a pattern corresponding to a single slide change (i.e., corresponding to the last slide in a forward skip, which remains visible once the fast skip is complete), rather than multiple or rapidly executed slide changes. In other words, the pattern detection function 134 will notice, for example, that ten slide changes occurred within a very short period of time, and why they are being processed as a detected pattern representing a single slide change. As a result, the generation function 135, having access to the pattern detected by the pattern detection function 134, may choose to display the last slide in full-screen mode for a few seconds in the output video stream 230, because it determines that this slide is potentially important in the state machine. It may also choose not to display the intermediate slides at all in the output stream 230.
[0145] The detection of a pattern with several rapidly changing slides can be detected by a simple rule-based algorithm, but can alternatively be detected using a trained neural network designed and trained to detect such patterns in moving images by classification.
[0146] In a different example, this may be useful, for example, in situations where the video communication is a talk show, panel debate, or the like, where it may be desirable to quickly switch visual attention between current speakers while still providing the consumer 150 with a relevant viewing experience by generating and publishing a calm and smooth output video stream 230. In this case, the event detection function 132 may continuously analyze each primary video stream 210, 301 to determine at all times whether the person being viewed in this particular primary video stream 210, 301 is currently speaking. This may be performed, for example, using conventional image processing tools as described above. The pattern detection function 134 may then be operable to detect specific overall patterns involving several of the primary video streams 210, 301 that are useful for generating a smooth output video stream 230. For example, the pattern detection function 134 may detect patterns of very frequent switching between current speakers and / or patterns involving several concurrent speakers. The generation function 135 can then take this detected pattern into account when making automatic decisions related to the generated state, for example by not automatically switching the visual focus to a speaker who speaks for only half a second before going silent again, or switching to a state in which several speakers are displayed side by side during a period when two speakers are speaking alternately or concurrently. This state determination process itself can be performed using time series pattern recognition techniques or using a trained neural network, but can also be based at least in part on a predetermined set of rules.
[0147] In some embodiments, multiple patterns may be detected in parallel and form inputs to the state machine of the generation function 135. These multiple patterns can be used by the generation function 135 via various AI components, computer vision detection algorithms, and the like. For example, a permanent slide change can be detected while concurrently detecting unstable connections for some participant clients 121, while another pattern detects the current primary speaking participant 122. Using all of this available pattern data, a classifier neural network can be trained, and / or a set of rules can be developed for analyzing the time series of this pattern data. This classification can be at least partially (e.g., fully) supervised to result in a determined desired state change for use in the generation. For example, different such predefined classifiers can be generated, specifically configured to automatically generate output video streams 230 according to various production styles and expectations. Training can be based on known production state change sequences as desired outputs and known pattern time series data as training data. In some embodiments, a Bayesian model can be used to generate such classifiers. In a specific example, information can be gathered a priori from experienced producers, providing input such as, "In a talk show, I never switch directly from speaker A to speaker B. Instead, I always show an overview before focusing on another speaker, unless the other speaker is very dominant and speaking loudly." This generative logic can then be expressed as a Bayesian model of the general form "If X is true | Assume fact Y is true | Do Z." Actual detection (e.g., whether someone is speaking loudly) can be performed using a classifier or threshold-based rules.
[0148] With larger datasets (larger datasets of pattern time series data), one can use deep learning methods to develop correct and attractive generative formats for use in automated generation of video streams.
[0149] In summary, the combination of event detection based on individual primary video streams 210, 301, intentionally introduced delays, pattern detection based on several time-synchronized primary video streams 210, 301 and detected events, and a generative process based on the detected patterns enables the automatic generation of an output digital video stream 230 according to a variety of possible taste and style choices. This result is valid across a variety of possible neural network and / or rule-based analysis techniques used by the event detection function 132, the pattern detection function 134, and the generation function 135. In particular, it is valid in the embodiments described below, which are characterized by the use of a first generated video stream for the automatic generation of a second generated video stream, and the use of different intentionally added delays for different groups of participant clients.
[0150] As exemplified above, the generating step may include generating the output digital video stream 230 based on a set of predetermined and / or dynamically variable parameters regarding the visibility of the individual video streams in the primary digital video streams 210, 301 in the output digital video stream 230, the visual and / or auditory video content arrangement, the visual or auditory effects used, and / or the output mode of the output digital video stream 230. These parameters may be automatically determined by the generation function 135 state machine and / or set by an operator controlling the generation (making it semi-automatic) and / or predetermined based on certain a priori configuration expectations (such as the minimum time between output video stream 230 layout changes or state changes of the types exemplified above).
[0151] In a practical example, the state machine may support a set of predefined standard layouts that can be applied to the output video stream 230, such as a full-screen presenter view (showing the currently speaking participant 122 full screen), a slide view (showing the currently shared presentation slide full screen), a "butterfly view" (showing the currently speaking participant 122 and the currently shared presentation slide in a side-by-side view), a multi-speaker view (showing all participants 122 or a selected subset of participants side-by-side or in a matrix layout), and the like. The various available generation formats may be defined by a set of state machine state change rules (as described above) and a set of available states (such as the set of standard layouts described above). For example, one generation format may be "group discussion," another generation format may be "presentation," and so on. By selecting a particular generation format via a GUI or other interface to the central server 130, an operator of the system 100 can quickly select a generation format from this predefined set of such generation formats and then allow the central server 130 to fully automatically generate the output video stream 230 according to the generation format in question based on the available information as described above.
[0152] In addition, during the generation, as described above, a corresponding in-memory buffer is created and maintained for each meeting participant client 121 or external video source 300. These buffers can be easily removed, added, and changed on-site. The central server 130 can then be arranged to receive information about added / removed participant clients 121 and participants 122 scheduled for voice delivery, planned or unexpected presentation pauses / resumptions, desired changes to the currently used production format, etc. during the generation of the output video stream 230. As described above, such information can be fed to the central server 130, for example, via an operator GUI or interface.
[0153] As illustrated above, in some embodiments, at least one of the primary digital video streams 210, 301 is provided to the digital video communication service 110, and then the publishing step may include providing the output digital video stream 230 to the same communication service 110. For example, the output video stream 230 may be provided to a participant client 121 of the video communication service 110, or provided to the video communication service 110 as an external video stream via the API 112. In this way, the output video stream 230 may be made available to several or all of the participants of the video communication event currently being facilitated by the video communication service 110. As also discussed above, the output video stream 230 may additionally or alternatively be provided to one or more external consumers 150 .
[0154] In general, the generating step may be performed by the central server 130 to provide the output digital video stream 230 as a live video stream to one or several concurrent consumers via the API 137 .
[0155] Figure 8a The method according to the first aspect of the present invention is shown, which will be described below with reference to the content described above. Figure 8a In the method for providing a digital video stream (which is denoted as the "second" digital video stream in the following) shown in , all the above mechanisms and principles regarding digital video stream collection, event detection, synchronization, pattern detection, generation and publication can be applied.
[0156] Regarding the method according to the second aspect of the present invention Figure 8b , shows the method according to the third aspect of the present invention Figure 8c , and a method according to the fourth aspect of the present invention Figure 8d Overall, they correspond.
[0157] The first aspect, the second aspect, the third aspect and the fourth aspect can be combined freely. In particular, the method according to the fourth aspect can be used in combination with the method according to any one of the first aspect, the second aspect and the third aspect.
[0158] and, Figure 9 is in the process of executing Figures 8a to 8d A simplified view of a system 100 in a configuration of the method is shown. The central server 130 includes a collection function 131, which may be as described above.
[0159] The central server 130 also includes a first generation function 135', a second generation function 135", and a third generation function 135'". Each such generation function 135', 135", 135'' corresponds to the generation function 135, and the above description of the generation function 135 also applies to the generation functions 135', 135", and 135''. Depending on the detailed configuration of the central server 130, the generation functions 135', 135", 135'' can be different, or can be arranged together with several functions in a single logical function, and there can also be more than three generation functions. The generation functions 135', 135", 135'' can in some cases be different functional aspects of the same generation function 135, as the case may be. Various communications between the generation functions 135', 135", 135'' and other entities can be carried out via appropriate APIs.
[0160] It is also recognized that, depending on the detailed configuration, for each of the generation functions 135', 135", 135''' or a group of such generation functions, there may be a separate collection function 131 and there may be several logically separate central servers 130, each central server having a corresponding collection function 131.
[0161] Moreover, the central server 130 includes a first publishing function 136', a second publishing function 136", and a third publishing function 136'". Each such publishing function 136', 136", 136'' corresponds to the publishing function 136, and the above description of the publishing function 136 also applies to the publishing functions 136', 136''' and 136'''. Depending on the detailed configuration of the central server 130, the publishing functions 136', 136", 136'' can be different, or can be arranged together with several functions in a single logical function, and there can also be more than three publishing functions. The publishing functions 136', 136", 136'' can in some cases be different functional aspects of the same publishing function 136, depending on the circumstances.
[0162] exist Figure 9, three sets or groups of participant clients are shown to illustrate the principles described herein, each set or group corresponding to a participant client 121 as described above. Thus, there is a first group of such participant clients 121', a second group of such participant clients 121", and a third group of such participant clients 121'''. Each of these groups may comprise one or preferably at least two participant clients. Depending on the detailed configuration, there may be only two such groups, or more than three such groups. The allocation between the groups 121', 121", 121''' may be exclusive in the sense that each participant client 121 is allocated to at most one such group 121', 121", 121'''. In an alternative configuration, at least one participant client 121 may be allocated to more than one such group 121', 121", 121''' at the same time.
[0163] Figure 9 An external consumer 150 is also shown, and it is appreciated that there may be more than one external consumer 150 as described above. To keep it simple, Figure 9 Video communication service 110 is not shown, but it will be appreciated that video communication services of the general type discussed above may be used with central server 130, such as to provide a shared video communication service to each participant client 121 using central server 130 in the manner discussed above.
[0164] Back to Figure 8a ,In a first step, the method starts.
[0165] In a subsequent collecting step, a plurality of the primary video streams are collected, in this exemplary case at least a first primary digital video stream, a second primary digital video stream, and a third primary digital video stream, each collected from a corresponding participant client 121. Thus, the first primary digital video stream is collected from the first participant client, the second primary digital video stream is collected from the second participant client, and the third primary digital video stream is collected from the third participant client.
[0166] In a subsequent publishing step, at least one video stream is provided to at least one of the first participant client and the second participant client. That is, the video stream is at least one of the first primary digital video stream, the second primary digital video, and a first generated video stream that has been generated based on at least one of the first primary video stream and the second primary video stream. This generation of the primary video stream may be performed by a first generation function 135' as described below, and may, for example, include introducing a delay to the first generated digital video stream as a result of the generation in question.
[0167] The provisioning and publishing in question may be continuous and may be in real time.
[0168] For example, a first participant and a second participant may participate in the same video communication service (such as a video conference), as also described elsewhere herein. Then, for example, the second digital video stream may be provided to the first participant client 121 for viewing on the screen 124 of the first participant client 121, and vice versa, so that the first and second participant client 121 users 122 can see and interact with each other. Additionally or alternatively, the first generated digital video stream may be provided to each or one of the first and second participant clients for viewing on the respective screen 124 of the participant client 121 in question. Where both the first generated digital video stream and either the first primary video stream or the second primary video stream are provided in tandem, the primary video stream in question may be delayed (as described below) to time-synchronize the video streams displayed on the participant client 121 in question.
[0169] In a subsequent second generating step performed by the second generating function 135″, a second generated video stream is generated as a digital video stream based on the first primary digital video stream, the second primary digital video stream and also based on the third primary digital video stream. Note that the third primary digital video stream is preferably not provided to the first participant client or the second participant client (neither as is nor as part of the generated digital video stream). As described elsewhere herein, the first participant client and the second participant client may be assigned to a different participant client group than the third participant client.
[0170] The second generation step includes introducing a time delay so that the second generated video stream is not time-synchronized with any of the video streams that may be provided to the first participant client or the second participant client in the publishing step. This time delay can be intentionally added and / or as a direct result of the generation of the second generated digital video stream in any of the ways described herein. Preferably, the second generated digital video stream can be used to publish with a certain delay relative to any video stream published at the first participant client and / or the second participant client. A way of thinking about this is that any consuming client of the second generated digital video stream consumes this second generated digital video stream in a "time zone" slightly after the video stream consumption "time zone" of the first participant client and the second participant client.
[0171] For example, where one or more primary digital video streams are provided to a first participant client and / or a second participant client, such provision may be direct (without any intentionally introduced temporal delay) and / or involve only relatively computationally lightweight processing prior to provision to the participant client in question; whereas generation of a second generated digital video stream may involve an intentionally introduced temporal delay and / or relatively heavyweight processing, resulting in the second generated digital video stream being generated for earliest delivery with a delay relative to the earliest delay for delivery of the first primary digital video stream and / or the second primary digital video stream. Where a first generated video stream is provided to a first participant client and / or a second participant client, the first generated digital video stream is generated using a relatively short intentionally introduced temporal delay and / or relatively lightweight processing, whereas the second generated digital video stream is generated using a relatively long intentionally introduced temporal delay and / or relatively heavyweight processing, resulting in the second generated digital video stream being generated for earliest delivery with a delay correspondingly relative to the earliest delay of the first generated digital video stream.
[0172] Typically, the second generated digital video stream is not provided for distribution at the first participant client or the second participant client, but instead is distributed at a participant client (such as a third participant client (assigned to a different group, such as the second group 121') and / or the external consuming client 151) that is assigned to a group different from the group to which the first client and the second client belong (such as the first group 121').
[0173] Therefore, if Figure 8a As shown, the publishing step further includes continuously providing the second generated video stream to at least one consuming client 121, 150 that is not the first participant client or the second participant client.
[0174] Likewise Figure 8a As shown, the method can iteratively and continuously generate and provide / publish the digital video stream in question.
[0175] In the following steps, the method ends.
[0176] Figure 8b A method according to the second aspect is shown.
[0177] In a first step, the method starts.
[0178] In a subsequent collecting step, a plurality of said primary video streams are collected, in this exemplary case at least a first primary digital video stream and a second primary digital video stream collected from respective participant clients 121 ′ selected from said first group of participant clients.
[0179] like Figure 8aAs in the case of the method shown in FIG. 1 , the collection may be as described above, with the collection function 131 processing the raw data, for example, without performing any re-encoding. There may also be event detection steps, synchronization steps, and pattern detection steps of the general type described above applied to the primary digital video stream collected from the first group of participant clients 121′ for the purpose of generating a first generated digital video stream.
[0180] That is, in a subsequent first generation step, the first generation function 135' receives the first and second primary video streams as respective digital video streams from the collection function 131 and generates the first generated digital video stream based on the first and second primary digital video streams. Preferably, the first digital video stream is not generated based on any other participant client 121 other than the participant client assigned to the first group 121' (the any other participant client being connected to the same video communication service 110 in a manner allowing such other participant client 121 to interact with the members of the first group 121' in the video communication service 110). On the other hand, the first generated video stream may be generated based on other information, such as an external video feed, static data or graphics. For the sake of clarity, with respect to Figure 8b These and other things described can also be applied to Figure 8a 、 Figure 8c and Figure 8d The method shown in .
[0181] Thus, the result of this first generation is a generated digital video stream of the type described above, which may, for example, visually include one or more of the primary video streams discussed as sub-portions in processed or unprocessed form. This first generated video stream may include live captured video streams, slides, externally provided videos or images, etc., as generally described above with respect to the video output streams generated by the central server 130. The first generated video stream may also be generated in the general manner described above based on detected events and / or patterns of the first and / or second primary video streams, whether intentionally delayed or in real time, provided by the participant clients of the first group 121′.
[0182] In a subsequent second generation step, a second generated digital video stream is generated as a digital video stream based on the first generated video stream and also based on both the first and second primary digital video streams collected from the first participant client group 121'. The first and second primary digital video streams can be provided from the collection function 131 to the second generation function 135", while the first generated video stream can be provided from the first generation function 135' to the second generation function 135". In the case where the first generation function 135' and the second generation function 135" are the same logical unit (which may be the case), generation simply occurs in two consecutive steps within this generation function.
[0183] In the following steps, the method ends.
[0184] It is appreciated that the first and / or second primary video streams fed to the second generating step 135" may be pre-formatted in various ways prior to the second generating step 135". They may also be intentionally delayed in order to detect events and / or patterns as described above.
[0185] The second generation step can be similar to any of the above-mentioned generation steps, and all the contents described above regarding the operation of the generation functions 135, 135' also apply accordingly to the second generation function 135". For example, as part of the generation process, the second generation function 135" can generate a second generated video stream by formatting the main video stream in various ways.
[0186] As mentioned, the first and second primary video streams may be time synchronized with each other before being fed to the first generation function 135 ′, such as using a common time reference in any of the manners discussed above.
[0187] However, in the second generation step, the first and second primary digital video streams can be intentionally time-delayed (e.g., in addition to any already-applied time delays implemented to time-synchronize the primary video streams with each other and / or to enable detection of events and / or patterns for use in the first generation function 135'). The purpose and result of this now intentionally introduced time delay is to time-synchronize the primary video streams with the first generated video stream before use in the second generated video stream. Thus, the additional delay introduced relative to the first and second primary video streams is equal to, substantially equal to, or at least determined as a function of the delay associated with performing the first generation step. For example, the exact delay to be added can be determined based on a detected common time reference of the general type described above.
[0188] That is, the first generation step, in which the first generated video is generated, is usually associated with a certain delay (due to the data processing of the first generation step 135' itself), which may depend, for example, on the available computer power and the complexity of the first generation step 135'. For the first and second video streams themselves, such delay is usually absent (or any delay is in any case much smaller), which are simply captured, optionally processed in the manner mentioned, and then provided by the collection function 131 to the second generation function 135" and used by it.
[0189] By intentionally introducing this (additional) delay into the first and second primary video streams while taking into account the delay of the first generated video stream caused by the first generation step in order to time-synchronize the three video streams, even in a case where the second generated video stream is generated not only based on the first and second primary video streams but also based on the first generated video stream (which is in turn generated based on the same first and second primary video streams), the second video stream can be generated without any synchronization problems. That is, the second generated video stream is then generated based on the time-delayed first and second primary digital video streams. Thus, the first generated video stream can be fed to a second generation step 135", which is thus generated using two (or more) generation steps 135', 135", wherein the same main video stream is used in at least two such generation steps associated with different delays relative to a common time base of the main video stream provided by the collection function 131.
[0190] In the illustrative example, a first group 121′ of participant clients are part of a debate team that communicates with relatively low latency using the video communication service 110, each of which is continuously fed a first generated video stream (or each other's respective primary video streams, as described above in conjunction with Figure 8a As described). The audience for the debate team consists of a second group 121' of participant clients that are continuously fed a second generated video stream, which in turn is associated with a slightly higher latency. The second generated video stream can be automatically generated in the general manner discussed above to automatically switch between a view of an individual debate team speaker (assigned to the first group 121' of participant clients, such view being provided directly from the collection function 131) and a generated view showing all debate team speakers (this view being the first generated video stream). Using the invention according to the first and / or second aspects, the audience can receive a good experience while the group speakers can interact with each other with minimal latency.
[0191] The delay intentionally added to the first and second primary video streams in conjunction with the second generation step may be at least 0.1 s, such as at least 0.2 s, such as at least 0.5 s; and may be at most 5 s, such as at most 2 s, such as at most 1 s. It may also depend on the inherited delay associated with each primary video stream in order to achieve complete time synchronization between the first and second primary video streams and also the first generated video stream.
[0192] It should be understood that the first and second primary video streams and the first generated video stream may all be additionally intentionally delayed to improve pattern detection for use in the second generation function 135" in the general manner described above.
[0193] Figure 9 A number of alternative or concurrent ways of publishing the various generated video streams generated by the central server 130 are shown. In general, in a subsequent publishing step performed by a first publishing function 136′ arranged to receive the first generated video stream from the first generating function 135′, the first generated video stream may be continuously provided to at least one of a first participant client 121 and a second participant client 121. For example, this first participant client may be a participant client from a group 121′ providing the first primary digital video stream, and / or the second participant client may be a participant client from a group 121′ providing the second primary digital video stream.
[0194] In other words, the first generated video stream may be continuously provided to at least one of the first participant client and the second participant client.
[0195] In some embodiments, one or several of the participant clients of the group 121 ′ may also receive the second generated video stream via a second publishing function 136 ″ which in turn is arranged to receive the second generated video stream from a second generating function 135 ″. Thus, the main video stream assigned to the first group 121' provides that each of the participant clients can be provided with the first generated video stream if not directly provided with the main digital video stream, which involves a certain delay or latency due to the synchronization between the main video streams and also the possibility of intentionally adding a delay or latency in order to allow sufficient time for event and / or pattern detection, as described above. Accordingly, each of the participant clients assigned to the second group 121" may be provided with the second generated video stream, also including said intentionally added delay associated with the second generation step, added for the purpose of time-synchronizing the first generated video stream with said first and second main video streams. This additional delay may or may not cause communication difficulties between the participant clients of the second group 121", for example because they interact with the video communication service 110 in a different manner than the participants of the first group 121' (see below). In other embodiments (such as when the participant clients of the first group 121' are provided with the main digital video stream directly), each of the participant clients assigned to the second group 121" may be provided with the first generated video stream directly.
[0196] Thus, the first group 121' of participant clients forms a subgroup of all participant clients 121 currently participating in the video communication service 110 in question, existing and using the service in a "time zone" that is slightly ahead (such as 1 to 3 seconds ahead) of any other participant clients, rather than being continuously provided with a generated video stream (such as the first generated video stream or the second generated video stream). However, other participant clients (not assigned to the first group 121' but assigned to the second group 121") will be continuously provided with a second generated video stream that is based on the first and second primary video streams but generated in a slightly later "time zone" (and may contain either or both of the first and second primary video streams at each point in time). Because the first generated video stream is generated directly based on the first and second primary video streams without any latency or delay added to time-synchronize them with the already generated video streams based on the primary video streams themselves, these participant clients 121 can obtain a more direct, low-latency video communication service 110 experience. Likewise, this may also mean that the participant clients 121 assigned to the first group 121 ′ are not provided with access to the second generated video stream.
[0197] That is, the first and second primary digital video streams may be provided as part of a shared digital video communication service 110 of the general type discussed above, and both the first and second participant clients (belonging to the same first group 121′) may be participant clients of respective remote connections to the shared digital video communication service 110. The second group 121″ participant clients (and also the third group 121′″ participant clients) may also be participant clients of remote connections to the shared digital video communication service 1110.
[0198] It is recognized that in this scenario, “remote connection” does not necessarily mean that such participant client 121 or corresponding user 122 is located in a different room, place or geographical location, but rather that the user 122 uses the participant client 121 in question to conduct audio / visual interaction with the video communication service 110. The collecting step may include collecting the first primary digital video stream and / or the second primary digital video stream from the shared digital video communication service 110, such as in any of the manners discussed above.
[0199] Figure 8c The method according to the third aspect is shown. As mentioned, Figure 8c The method shown in ( Figure 8d This is also the case with the method shown in Figure 8a and Figure 8b The method shown in FIG, and these four aspects of the invention can be freely combined. According to compatibility, all contents related to one of these aspects can be easily applied to other aspects in a corresponding manner.
[0200] In a first step, the method starts. In a subsequent collecting step, a first primary digital video stream is collected from the first participant client assigned to the first group 121', and a second primary digital video stream is collected from a second participant client also assigned to the same first group 121'. Furthermore, a third digital video stream is collected from a third participant client that may not be assigned to the first group 121'. For example, the third participant client may be assigned to the second group 121". This collecting step may be similar to the one in conjunction with Figure 8b Describe the collection steps.
[0201] In the subsequent first generation step (which can be similar to the Figure 8b In the first generation step described in the embodiment of the present invention, the first generated video stream can be generated as a digital video stream based on the collected first main digital video stream and the second main digital video stream. It should be noted that the first generated video stream may not be generated based on the third main video stream. The first generated digital video stream is continuously generated with a first delay for distribution to a consuming client. In other words, according to this third aspect, the first generated digital video stream is generated such that if each newly generated frame of the first generated digital video stream is distributed immediately after the generation of the frame in question, the distribution of this frame occurs with the first delay.
[0202] In the subsequent second generation step (which can be similar to Figure 8bIn the second generating step described in the foregoing, a second generated video stream is generated as a digital video stream based on all three main digital streams (ie, all of the first, second and third main digital video streams). A second generated digital video stream is continuously generated for distribution with a second delay in a manner corresponding to the first generated video stream and the first delay. The second delay is greater than the first delay. This means that if both the first generated video stream and the second generated video stream contain frames from, for example, the first primary video stream, such frames will be shown earlier in the immediate distribution of the first generated video stream than in the immediate distribution of the second generated video stream.
[0203] In a subsequent release step (which may be similar to Figure 8b In the publishing step described above, at least one of the first primary digital video stream, the second primary digital video stream, and the first generated video stream (such as any set of one or more of these streams) is continuously provided to at least one of the first participant client and the second participant client. This is similar to the above description of Figure 8a Described method.
[0204] Furthermore, the second generated video stream is continuously provided to at least one other participant client.
[0205] In the following steps, the method ends.
[0206] The same example used to illustrate the practical application of the second aspect of the solution can also be used to illustrate how this third aspect can be put into practice. Because the third primary video stream is collected from the second group 121″ participant client, which has a lower latency requirement than the first group 121′ participant client that provides the first and second primary video streams, the second generated video stream is provided with more latency, which enables the desired automatic generation, while the first group 121′ panelists can interact with each other with even lower latency.
[0207] Naturally, in addition to the third main video stream, there may be more main video streams provided by the second group 121 ″ which will then be used accordingly.
[0208] In About Figure 8b and Figure 8c In the described publishing step, the second generated video stream can be continuously provided to at least one consuming client that is not the first participant client or the second participant client. More generally, it can be continuously provided to participant clients 121 that are not assigned to the first group 121 ′ and / or the external consumer 150.
[0209] As mentioned above, the collecting step 131 may include collecting at least one of the primary digital video streams (such as an additional primary video stream in addition to the first and second primary video streams) as an external digital video stream 301 of the type discussed above, collected from an information source 300 external to the shared digital video communication service 110. Also as described above, such an external video stream 301 may be time-synchronized with the first and second primary video streams via a synchronization function 133 logically located (in terms of data streams) between the collecting function 131 and the first generating function 135'. The same applies to the third, fourth, and fifth primary video streams described herein. A first generated video stream and / or a second generated video stream may then be generated based on the external digital video stream 301.
[0210] As also generally discussed above, the first generating step 135′ and / or the second generating step 135″ may also include generating the respective generated (first and / or second) video stream in question based on a set of predetermined and / or dynamically variable parameters regarding the visibility of individual primary digital video streams in the first and / or second primary digital video streams 210 in the generated digital video streams in question, the visual and / or auditory video content arrangement, the visual or auditory effects used, and / or the output mode of the generated digital video stream in question.
[0211] Also as discussed, the first generating step 135′ and / or the second generating step 135″ may be performed by the central server 130, thereby providing the second generated video stream as a live video stream to one or more concurrent (external and / or participating) consumer clients via an API 137 of the general type discussed above.
[0212] Therefore, different groups 121 ′, 121 ″, 121 ′″ of participant clients 121 may have different requirements regarding delay tolerance. This may be particularly true if they participate in the same live video communication service 110 as participant clients 121 of such service 110. This will be further exemplified below.
[0213] Figure 8d A method according to the fourth aspect of the invention is shown.
[0214] In a first step, the method starts.
[0215] Overall, and as Figure 8d As shown, the method for generating said second generated digital video stream may comprise a subsequent allocation step, which may be an initial step but may also be performed at any time during the method, for example as a reallocation step. In this allocation step, a plurality of participant clients 121 may be allocated between at least two groups 121 ′, 121 ″, 121 ′″ of such participant clients 121. In this example, the participant clients 121 are allocated to at least a first group 121 ′ and a third group 121 ′″, but the participant clients 121 may of course also be allocated to the third group 121 ′″.
[0216] More specifically, the first and second primary digital video streams may be collected, such as by the collection function 131 and in a subsequent collection step, from the respective participant clients 121 assigned to the first participant client group 121′. However, the fourth and fifth primary digital video streams may also be collected, such as by the collection function 131 and in the collection step, but from the respective participant clients 121 assigned to the third participant client group 121′″.
[0217] exist Figure 9 In the example shown in , the participant clients 121 assigned to the third group 121" may have less stringent latency requirements than the participant clients 121 assigned to the first group 121'. For example, the first group 121' participant clients 121 may be members of the debate team discussed above (interacting with each other in real time and therefore requiring low latency), while the third group 121" participant clients 121 may be a panel of experts who constitute but interact with the team in a more structured manner (such as using clear questions / answers) and therefore be able to tolerate higher latency than the first group 121'. The first generated video stream may be generated as described above by the first generation function 135' and based on the first and second primary video streams (and any additional input content, as discussed). A third generated video stream is also generated in a corresponding manner, but by the third generation function 135''' and based on (at least) the fourth and fifth primary video streams.
[0218] Both the first generated video stream and the third generated video stream may be fed to a second generation function 135 ″ as a basis for the generation of a second generated video stream, as the case may be.
[0219] Then, however, according to this fourth aspect, in a second generation step performed by the second generation function 135", a second generated video stream is generated based on at least one of the first and second main video streams and further based on at least one of the fourth and fifth main video streams, such as in the manner described above. The fourth and fifth main video streams from the collection function 131 can be provided to the second generation function 135" in a manner corresponding to the first and second main video streams (including any cross-stream time synchronization, event detection, etc.). It is particularly noted that the second generated video stream can be based directly or indirectly on the first and / or second main video streams, for example, the second generated video stream is based on the first generated video stream, which in turn is based on the first and second main video streams, and correspondingly for the fourth and fifth main video streams and the third generated video stream.
[0220] The third generated video stream is generated in a subsequent third generation step. According to the fourth aspect, the third generation step includes intentionally introducing a time delay relative to the fourth and fifth primary video streams so that they are time-synchronized with each other but not time-synchronized with the first generated video stream (not time-synchronized). It should be understood that this time delay can be introduced in the third generation function 135''' itself, or in a corresponding synchronization function 133 upstream of the third generation function 135''' in question.
[0221] Thus, the first generation step 135' may involve introducing an intentional delay or latency of the type discussed above, which is introduced in addition to any delay introduced as part of the synchronization of the first and second primary video streams, and is introduced in order to obtain sufficient time, for example, to perform effective event and / or pattern detection. This introduction of an intentional delay or latency may occur as part of the synchronization performed by the synchronization function 133 (for simplicity, in Figure 9 The same may be true for the third generating step 135'', but with the introduction of an intentionally introduced delay or delay that is different from the intentionally introduced delay or delay for the first generating step 135'. In particular, the intentionally introduced delay or latency results in a time desynchronization between the first generated video stream and the third generated video stream. This means that while the first generated video stream and the third generated video stream are both published immediately and continuously as each individual frame is generated, they do not follow a common timeline.
[0222] As discussed above, the second generated video stream may be associated with a higher delay than the first generated video stream, and may also be associated with a higher delay than the third generated video stream. Therefore, the second generation function 135" may be arranged to synchronize the first primary video stream, the second primary video stream, the fourth primary video stream, and the fifth primary video stream by adding additional corresponding delays to the first primary video stream, the second primary video stream, the fourth primary video stream, and the fifth primary video stream before merging them into the second generated video stream. In the publishing step, the third generated video stream is continuously provided to at least one participant client assigned to the third group 121''', and at the at least one participant client, the third generated video stream can be continuously published to the user 122 in question. Similarly, the first generated video stream can be continuously provided to at least one participant client assigned to the first group 121', and at the at least one participant client, the first generated video stream can be published to the user 122 in question; and / or the second generated video stream can be provided and published as described above.
[0223] In the following steps, the method ends.
[0224] Thus, in this fourth aspect, three separate generated video streams can be generated and consumed / published simultaneously, but in different "time zones." Even though they are based at least in part on the same primary video material, the generated video streams are published with different delays. A first group 121', requiring the lowest latency, can interact using the first generated video stream, which offers very low latency. A third group 121''', willing to accept slightly greater latency, can interact using the second generated video stream, which offers greater latency but also provides greater flexibility in intentionally adding delays to enable better automatic generation, as described elsewhere herein. Meanwhile, a second group 121'', less sensitive to latency, can enjoy interacting using the second generated video stream, which can combine material from both the first group 121' and the third group 121''', and can also be automatically generated in a very flexible manner. It is important to note that despite using these varying delays and thus operating in different "time zones," all of these participant user groups 121', 121", and 121''' interact with each other using the video communication service 110. However, due to the synchronization of the respective input video streams in each generating function, the participant users 121 will not notice the different delays from their respective perspectives.
[0225] Said first generating step 135' may comprise time-delaying the first primary video stream and the second primary video stream in order to time-synchronize them with each other, as described above.
[0226] Accordingly, the third generation step 135''' (or the corresponding synchronization step 133) may include time delaying the fourth main video stream and the fifth main video stream so that they are time synchronized with each other, but using a maximum time delay that is greater than the maximum time delay used to time delay the first main video stream and the second main video stream in the first generation step 135' (or the corresponding synchronization step 133), thereby causing the first generated video stream to not be time synchronized with the third generated video stream in the described manner.
[0227] As described above, the respective participant clients 121 assigned to each of the groups 121 ′, 121 ″, 121 ′″ may participate in the same video communication service 110 in which the second generated video stream is continuously published.
[0228] Different ones of the groups 121', 121", 121''' may then be associated with different participant interaction rights in the video communication service 110, and different ones of the groups 121', 121", 121''' may be associated with different maximum time delays (latencies) for generating the respective generated video streams published to the participant clients 121 assigned to the group 121', 121", 121''' in question.
[0229] For example, a first group 121′ of panel debate participant clients may be associated with full interaction rights and may speak whenever they wish. A third group 121′″ of participant clients may be associated with slightly more limited interaction rights, such as requiring them to unmute their microphones to speak via the video communication service 110 before requesting the floor. A second group 121″ of audience participant users may be associated with even more limited interaction rights, such as being able to ask questions in writing in a public chat room but not being able to speak.
[0230] Thus, different groups of participating users can be associated with different interaction privileges and different delays for the corresponding generated video streams published to them, in such a way that the delay is an increasing function of decreasing interaction privileges. The more freely the video communication service 110 allows the participating user 121 in question to interact with other users, the lower the acceptable delay. The shorter the acceptable delay, the less likely the corresponding automatic generation function will take into account events or patterns, such as those detected.
[0231] The group with the greatest delay may be a viewers-only group with no interaction rights other than passive participation in the video communication service.
[0232] In particular, the respective maximum time delay (latency) for each of said groups 121', 121", 121''' can be determined as the maximum delay difference between the entire primary video stream and any generated video streams that are continuously published to the participant clients in the group in question. To this can be added any additional time delay that is intentionally added in order to detect events and / or patterns as described above.
[0233] As used herein, the terms "generate" and "generate digital video stream" may refer to different types of generation. In one example, a single well-defined digital video stream is generated by a central entity (such as a central server 130) to form the generated digital video stream in question for provision to and publication at each of a particular set 121 of participant clients that are to consume the generated digital video stream in question. In other cases, different individual such participant clients 121 may view slightly different versions of the generated digital video stream in question. For example, the generated digital video stream may include several separate or combined digital video streams, and the local software functionality 125 of the participant client 121 may allow the user 122 in question to switch between these digital video streams; arrange these digital video streams on a screen 124; or configure or process these digital video streams in any other manner. Often times, it is important to provide the "time zone" (i.e., at what delay) the generated digital video stream (including any time-synchronized subcomponents). Thus, the above in conjunction with Figure 8a The described situation of providing each other's main video stream to the first participant client and the second participant client can be regarded as providing the first generated digital video stream to the first participant client and the second participant client (in the sense that a set of time-synchronized original or processed first main digital video stream and second main digital video stream are available to both the first participant client and the second participant client).
[0234] To further clarify and illustrate the use of the above-described participant client groups 121 ′, 121 ″, 121 ′″, the following example is provided in the form of a video communication service meeting involving three different concurrent “time zones”: The first group of participant clients 121′ experience real-time, or at least near-real-time (depending on unavoidable hardware and software delays), interactions with one another. These participant clients are provided with video (including audio) from one another to enable this interaction and communication between the user 122 in question. The first group 121′ can serve the user 122 at the heart of the conference, where other participant clients (not in the first group 121′) may be interested in participating.
[0235] Such a second group of other participant clients 121 ' participates in the same meeting, but in a different "time zone" that is further away from real time than the first group of participant clients 121 '. The second group 121 ' may for example be an audience with interactive privileges such as the possibility to ask questions to the first group 121 '. The "time zone" of the second group 121 ' may have a delay relative to the "time zone" of the first group 121 ', such that asked questions and answers are associated with a noticeable but short delay. On the other hand, this slightly larger delay allows this second group 121 ' participant clients to experience a generated digital video stream that is automatically generated in a more sophisticated way, thereby providing a more satisfying user experience.
[0236] Such a third group of other participant clients 121 also participates in the same meeting, but only as viewers. This third group 121''' consumes a generated digital video stream, which may be automatically generated in a more sophisticated and complex manner and consumed in a third "time zone" with greater latency than the second "time zone." However, because the third group 121''' cannot provide input to the communication service in a manner that affects the first group 121' and the second group 121", the third group 121''' will experience the meeting as if it were performed in "real time", with a satisfactory outcome.
[0237] Of course, using the principles described herein, there may be more than three such participant client groups associated with corresponding meeting "time zones" of increasing time delays and increasing production complexity.
[0238] The present invention also relates to a computer software function for providing a second digital video stream according to the above-described content. Such a computer software function can then be arranged to perform, when running, at least some of the above-described collecting steps, event detection steps, synchronization steps, pattern detection steps, generation steps, and publishing steps, in particular with respect to the first, second, third, and / or fourth aspects. The computer software function can be arranged to run on the physical or virtual hardware of the central server 130, as described above.
[0239] The present invention also relates to a system 100 for providing a second digital video stream and further comprising a central server 130. The central server 103 can be configured to perform at least some of the collecting steps, event detection steps, synchronization steps, pattern detection steps, generation steps, and publishing steps described, in particular, with respect to the first, second, third, and / or fourth aspects. For example, these steps can be performed by the central server 130 running the computer software functionality described above to perform the steps described above.
[0240] It should be understood that the principles described above for automatic generation of an available set of input video streams (such as time synchronization, event and / or pattern detection involving such input video streams, etc.) can be applied concurrently at different levels. Thus, one such automatically generated video stream can form an available input video stream to a downstream automatic generation function that in turn generates a video stream.
[0241] The central server 130 may be arranged to control the assignment of groups 121′, 121″, 121′″ to the various participant clients 121. For example, dynamically changing the group assignment to a particular such participant client during the course of a live video communication service session may be part of automatically generating said video communication service by the central server 130. Such reallocation may be based on a predetermined schedule or triggered dynamically, for example as a function of parameter data that may vary dynamically over time, for example upon request by an individual participant client user 122 (provided via the client 121 in question).
[0242] Accordingly, the central server 130 may be arranged to dynamically change the group structure during the course of the video communication service, such as using a certain group only during predetermined time slots (such as during a planned group debate).
[0243] One practical solution for group assignment is to use the concept of "breakout rooms" available on some video conferencing systems. Participant clients 121 assigned to a particular group 121', 121", or 121'" can then be assigned to such a breakout room, and central server 130 can then obtain video stream data (such as a separate primary video stream or a generated video stream) from this breakout room for use in downstream generation steps within central server 130. This video stream extraction itself can occur as already described above. In all of the above aspects, the present invention may further include an interaction step in which at least one participant client of a first group interacts in a bidirectional (two-way) manner with at least one participant client of a second group, the first group being associated with a first delay and the second group being associated with a second delay, the second delay being different from the first delay. It should be understood that these participant clients may all be participants in the same communication service of the type described above.
[0244] In this case, it is preferable that the participant clients associated with different delays (or "time zones" as discussed above) are temporarily placed in the same "time zone," in other words, associated with the same delay. For example, this can occur by temporarily providing one of the participant clients temporarily associated with a higher delay with one or more primary / generated digital video streams that have been generated using a lower delay than the delay associated with the participant client in question. In other words, if a participant client that is normally provided with one or more video streams with a higher delay wishes to interact with a participant client that is normally provided with one or more video streams with a lower delay, the former participant client is instead temporarily provided with one or more video streams with a lower delay. Thus, the higher-latency participant client is temporarily switched to the lower-latency "time zone" associated with the lower-latency participant client. After the interaction, the higher-latency participant client is then switched back to the higher-latency communication environment used prior to the interaction.
[0245] For example, a member of the audience of the panel debate discussed above might want to ask a question. In this case, the audience gets a command and switches to the panel debate "time zone." This means the audience will see the panel with lower latency, but in a less refined generation. More specifically, the audience member may see one or more of the same video streams presented to the panelists during the interaction. The rest of the audience will remain in the higher-latency audience "time zone" and therefore will not notice any difference. After the interaction between the speaking audience member and the panel, the speaking audience member will again be presented with the same high-latency video stream or streams as before the interaction.
[0246] Switching between different “time zones” can be automatically implemented by the central server 130 .
[0247] In the above, the preferred embodiment has been described. However, it is obvious to those skilled in the art that many modifications can be made to the disclosed embodiment without departing from the basic idea of the invention. For example, many additional features may be provided as part of the system 100 described herein and are not described herein. In summary, the presently described solution provides a framework upon which detailed functionality and features may be built to meet a variety of specific applications in which video data streams are used for communication.
[0248] An example is a presentation scenario where the main video stream includes the presenter's view, a shared digital slide-based presentation, and live video of the product being demonstrated.
[0249] Another example is a teaching situation, where the main video stream includes a view of the teacher, a live video of the physical entity that is the subject of the teaching, and respective videos of several students who may ask questions and engage in dialogue with the teacher. In either of these two examples, a video communication service (which may or may not be part of the system) may provide one or several of the primary video streams, and / or several of the primary video streams may be provided as external video sources of the type discussed herein.
[0250] The various groups are exemplified as a debate team, a panel of experts, and an audience. However, it should be appreciated that it is possible to divide the participant users in a digital video communication service into two or more groups, reflecting the current goals and structure of the communication to be performed. For example, one or more groups may include participant users accessing the video communication service remotely from different geographical locations, while one or more other groups may include participant users accessing the video communication service from a common central location (such as a lecture hall). The same principles as described above apply to all such scenarios.
[0251] In general, everything described in relation to the method applies to the system and computer software product, and vice versa.
[0252] The invention is thus not limited to the described embodiments but may be varied within the scope of the appended claims.
Claims
1. A method for providing a second generated video stream, the method comprising: In a first generating step (135'), a first generated digital video stream is generated based on the first main digital video stream and the second main digital video stream, wherein the first generated digital video stream is continuously generated for distribution with a first delay; In a second generating step (135"), a second generated video stream is generated based on the first main digital video stream and the second main digital video stream, the second generated digital video stream being continuously generated for distribution with a second delay, the second delay being greater than the first delay; as well as In a publishing step (136'), the first generated video stream is continuously provided to a first participant client (121), and the second generated video stream is provided to a second participant client (121); In the interaction step, temporarily providing the first generated video stream to the first participant client (121) and the second participant client (121); as well as After the interaction step, the second digitally generated video stream is again provided to the second participant client (121).
2. The method according to claim 1, wherein The first participant client (121) belongs to a first group associated with the first delay, and / or the second participant client (121) belongs to a second group associated with the second delay.
3. The method according to claim 1 or 2, wherein: The method comprises continuously providing the second generated video stream to at least one other participant client (121), for example belonging to a second group.
4. The method according to any one of claims 1 to 3, further comprising: In the publishing step (136"), the second generated video stream is continuously provided to at least one consuming client (121; 150) that is not the first participant client or the second participant client.
5. The method according to any one of claims 1 to 4, further comprising: The first primary digital video stream and the second primary digital video stream are provided as part of a shared digital video communication service (110), and the first participant client (121) and the second participant client (121) are both participant clients with respective remote connections to the shared digital video communication service (110).
6. The method according to claim 5, wherein: The collecting step (131) comprises collecting the first primary digital video stream and / or the second primary digital video stream from the shared digital video communication service (110).
7. The method according to claim 5 or 6, wherein: The collecting step (131) comprises collecting at least one primary digital video stream as an external digital video stream (301) collected from an information source (300) external to the shared digital video communication service (110), and wherein, The first generated video stream and / or the second generated video stream are generated based on the external digital video stream (301).
8. A method according to any one of the preceding claims, wherein The first generating step (135') and / or the second generating step (135") comprise generating the respective generated video stream based on a set of predetermined and / or dynamically variable parameters regarding visibility of individual ones of the first and / or second main digital video streams in said generated digital video streams, the visual and / or auditory video content arrangement, the used visual or auditory effects, and / or the output mode of said generated digital video streams.
9. A method according to any one of the preceding claims, wherein The first generating step (135') and / or the second generating step (135") are performed by the central server (130) so as to provide the second generated video stream (230) as a live video stream to one or several parallel consumer clients via an application programming interface API (137).
10. The method according to any one of the preceding claims, further comprising: In the allocating step, a plurality of participant clients (121) are allocated between at least two groups (121', 121", 121''') of such participant clients (121), wherein in the collecting step (131) the first primary video stream and the second primary video stream are collected from the participant clients (121) of the first group (121') allocated to the participant clients (121), and a fourth primary video stream and a fifth primary video stream are collected from the participant clients (121) of the third group (121''') allocated to the participant clients (121); In the second generating step (135"), the second generated video stream is generated based on at least one of the first primary video stream and the second primary video stream and further based on at least one of the fourth primary video stream and the fifth primary video stream; as well as In a third generating step (135'''), a third generated video stream is generated based on the fourth main video stream and the fifth main video stream, and the third generating step (135''') includes time-delaying the fourth main video stream and the fifth main video stream so that the third generated video stream is not synchronized with the first generated video stream; as well as In a publishing step (136''), the third generated video stream is continuously provided to at least one participant client assigned to the third group.
11. The method according to claim 10, wherein: The participant client (121) assigned to each of the groups (121', 121", 121''') participates in the video communication service (110) in which the second generated video stream is published; the method further comprises: Associating different groups in the groups (121', 121", 121''') with different participant interaction rights in the video communication service (110), and Different ones of the groups (121', 121", 121''') are associated with different maximum delays for generating respective generated video streams published to participant clients (121) assigned to the group (121', 121", 121''') in question.
12. The method according to claim 11, wherein The respective maximum delay for each of the groups (121', 121", 121''') is determined as the maximum delay difference between the overall primary video stream and any generated video stream that is consecutively published to a participant client (121) in the group (121', 121", 121''') in question.
13. A method according to any one of the preceding claims, wherein The method comprises: In a collecting step (131), a first primary digital video stream is collected from the first participant client (121), and a second primary digital video stream is collected from the second participant client (121).
14. A computer software product for providing a second generated video stream, said computer software functionality being arranged to, when run, perform: A first generation step (135'), in which generating a first generated digital video stream based on the first primary digital video stream and the second primary digital video stream, wherein the first generated digital video stream is continuously generated for distribution with a first delay; A second generating step (135"), wherein a second generated video stream is generated based on the first main digital video stream and the second main digital video stream, the second generated digital video stream being continuously generated for distribution with a second delay, the second delay being greater than the first delay; and a publishing step (136'), wherein the first generated video stream is continuously provided to a first participant client (121), and the second generated video stream is continuously provided to a second participant client (121); an interaction step, wherein the first generated video stream is temporarily provided to the first participant client (121) and the second participant client (121); and A providing step follows the interacting step, wherein the second digitally generated video stream is again provided to the second participant client (121).
15. A system (100) for providing a second generated video stream, the system (100) comprising a central server (130), the central server (130) further comprising: a first generating function (135'), wherein a first generated digital video stream is generated based on the first primary digital video stream and the second primary digital video stream, the first generated digital video stream being continuously generated for distribution with a first delay; A second generation function (135"), wherein a second generated video stream is generated based on the first main digital video stream and the second main digital video stream, the second generated digital video stream being continuously generated for distribution with a second delay, the second delay being greater than the first delay; and a publishing function (136'), wherein the first generated video stream is continuously provided to a first participant client (121), and the second generated video stream is provided to a second participant client (121); an interactive function, wherein the first generated video stream is temporarily provided to the first participant client (121) and the second participant client (121); and A functionality is provided, wherein, after the interaction step, the second digitally generated video stream is again provided to the second participant client (121).