System and method for generating a video stream
The method addresses synchronization and content selection issues in digital video conferencing by analyzing real-time streams for events and patterns, generating synchronized output streams with minimal latency, enhancing user experience in complex meetings.
Patent Information
- Application Number
- JP2025504708
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-02
- Filing Date
- 2023-08-01
- Publication Date
- 2025-08-19
AI Technical Summary
Digital video conferencing systems face challenges in dynamically selecting what information to display due to varying latency, frame rates, aspect ratios, and resolutions among different incoming video streams, leading to unsynchronized feeds and poor user experience, especially in complex meetings with multiple participants and diverse hardware.
A method and system that continuously collect real-time video streams, perform digital image analysis to identify events or patterns, establish generation control parameters based on these detections with minimal delay, and generate synchronized output streams without additional latency, using AI components for event and pattern detection.
Enables efficient, synchronized generation of digital video streams in real-time, improving user experience by dynamically adjusting content display based on detected events and patterns, even with diverse input streams from different participants and hardware.
Smart Images

Figure 2025527075000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a system, computer software product and method for generating a digital video stream, in particular for generating a digital video stream based on a digital input video stream. In a preferred embodiment, the digital video stream is generated in the context of a digital video conference, in particular a digital video conference or meeting system, in which multiple different concurrent users participate. The generated digital video stream may be published externally or within the digital video conference or digital video conference system.
[0002] In other embodiments, the invention is applied to contexts that are not digital video conferencing, but where multiple digital video input streams are simultaneously processed and combined into a digital video stream to be generated. For example, such a context may be educational or instructional. [Background technology]
[0003] Many digital video conferencing systems are known, such as Microsoft® Teams®, Zoom®, and Google® Meet®, that allow two or more participants to meet virtually using locally recorded digital video and audio that is broadcast to all participants, emulating a physical meeting.
[0004] There is a general need to improve these digital video conferencing solutions, particularly with regard to the generation (production) of viewing content: what content to show, at what time, to whom, and through what distribution channel.
[0005] For example, some systems automatically detect the currently speaking participant and display the corresponding video feed of that participant to other participants. Many systems also allow for the sharing of graphics such as the currently displayed screen, viewing window, or digital presentation. However, as virtual meetings become more complex, it will quickly become difficult for services to know which of all the information currently available should be displayed to each participant at any given time.
[0006] In another example, a presenting participant moves around while talking through slides in a digital presentation, in which case the system must decide whether to show the presentation, the presenter, or both, or switch between the two.
[0007] It may be desirable to generate one or more output digital video streams based on multiple input digital video streams through an automated generation process and provide such generated digital video stream or streams to one or more consuming entities. DISCLOSURE OF THE INVENTION [Problem to be solved by the invention]
[0008] However, in many cases, due to the many technical challenges faced by such digital video conferencing systems, it is difficult for dynamic conference screen layout managers and other automated generators to select what information to display.
[0009] First, low latency is important because digital video conferencing is real-time sensitive. This becomes problematic when different incoming digital video streams are associated with different latency, different frame rates, different aspect ratios, or different resolutions, such as when different participants join using different hardware. Often, these incoming digital video streams require processing for a well-formed user experience.
[0010] Second, production, in the sense of processing, selecting, and formatting video images, introduces delays that may be undesirable in real-time video communication between participants.
[0011] Third, there's the issue of time synchronization: like too much latency, unsynchronized digital video feeds lead to a poor user experience.
[0012] These issues are amplified in more complex meeting situations, such as those with large numbers of participants; participants connecting using different hardware and / or software; externally provided digital video streams; screen sharing; and multiple hosts.
[0013] These issues arise particularly in situations where a large number of participants participate in a video conference, where all participants are locally present in the same room or venue, or where some participants are locally present and some are remote.
[0014] Corresponding problems arise in other contexts when an output digital video stream is to be generated based on multiple input digital video streams, such as in digital video generation systems for education and instruction.
[0015] Swedish patent application SE2151267-8 (unpublished at the effective priority date of the present application) discloses various solutions to the above mentioned problems.
[0016] Swedish patent application SE2151461-7 (unpublished at the effective priority date of the present application) discloses various solutions specifically directed to handling latency in multi-participant digital video environments, including where different participant groups are associated with different general latency times.
[0017] Swedish patent application SE2250113-4 (unpublished at the effective priority date of the present application) discloses various solutions that specialize in using one or more cameras to track one or more people.
[0018] The present invention addresses one or more of the problems set forth above. [Means for solving the problem]
[0019] Accordingly, the present invention relates to a method for providing an output digital video stream, the method comprising: continuously collecting a real-time first primary digital video stream; performing a first digital image analysis of the first primary digital video stream to identify at least one first event or pattern within the first primary digital video stream, wherein the first digital image analysis results in establishing first generation control parameters based on detection of the first event or pattern, and the first digital image analysis establishes the first generation control parameters after a first time delay related to a time of occurrence of the first event or pattern within the first primary digital video stream. applying the first generation control parameters to the real-time first primary digital video stream, wherein application of the first generation control parameters results in the first primary digital video stream being modified based on the first generation control parameters to generate a first generated digital video stream without being delayed by the first time delay; and continuously providing the output digital video stream to at least one participant client, wherein the output digital video stream is provided in the form of or based on the first generated digital video stream.
[0020] Furthermore, the invention relates to a computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out the above method for providing an output digital video stream.
[0021] The present invention also relates to a system for providing an output digital video stream, the system comprising: a collection function configured to continuously collect a real-time first primary digital video stream; a generation function configured to perform a first digital image analysis of the first primary digital video stream to identify at least one first event or pattern within the first primary digital video stream, wherein the first digital image analysis results in first generation control parameters being established based on detection of the first event or pattern, and the first digital image analysis establishes the generation control parameters with a first time delay related to a time of occurrence of the first event or pattern within the first primary digital video stream. a generating function configured to generate a first generated digital video stream in the form of, or based on, the first generated digital video stream; and a publishing function configured to continuously provide the output digital video stream to at least one participant client, the output digital video stream being provided in the form of, or based on, the first generated digital video stream.
[0022] The present invention will now be described in detail with reference to exemplary embodiments thereof and the accompanying drawings. [Brief explanation of the drawings]
[0023] [Figure 1] FIG. 1 illustrates a first exemplary system. [Figure 2] FIG. 2 illustrates a second exemplary system. [Figure 3] FIG. 3 illustrates a third exemplary system. [Figure 4]FIG. 4 is a diagram illustrating the central server. [Figure 5] FIG. 5 is a diagram illustrating the first method. [Figure 6a] FIG. 6a shows subsequent states associated with different method steps in the method shown in FIG. [Figure 6b] FIG. 6b shows subsequent states associated with different method steps in the method shown in FIG. [Figure 6c] FIG. 6c shows subsequent states associated with different method steps in the method shown in FIG. [Figure 6d] FIG. 6d shows subsequent states associated with different method steps in the method shown in FIG. [Figure 6e] FIG. 6e shows subsequent states associated with different method steps in the method shown in FIG. [Figure 6f] FIG. 6f illustrates subsequent states associated with different method steps in the method shown in FIG. [Figure 7] FIG. 7 is a diagram conceptually illustrating a common protocol. [Figure 8] FIG. 8 is a diagram illustrating the second method. [Figure 9] FIG. 9 illustrates a fourth exemplary system. [Figure 10] FIG. 10 illustrates a fifth exemplary system. DETAILED DESCRIPTION OF THE INVENTION
[0024] All figures share the same or corresponding part reference numbers.
[0025] FIG. 1 shows a system 100 according to the present invention configured to perform a method according to the present invention for providing an output digital video stream, for example a shared digital video stream.
[0026] System 100 may include video communication service 110, which in some embodiments may be external to system 100. As described below, multiple video communication services 110 may be included.
[0027] The system 100 may include one or more participant clients 121, although one, some, or all of the participant clients 121 may be external to the system 100 in some embodiments.
[0028] The system 100 includes a central server 130 .
[0029] As used herein, the term "central server" refers to a computer-implemented function configured to be accessible in a logically centralized manner, such as through a well-defined API (application programming interface). Such central server functionality may be implemented purely in computer software, or in a combination of software and virtual and / or physical hardware. It may also be implemented in a standalone physical or virtual server computer, or distributed across multiple interconnected physical and / or virtual server computers.
[0030] As exemplified below, in some embodiments, the central server comprises or entirely consists of a piece of hardware that is locally located relative to one or more of the participant clients 121. As used herein, two entities being "locally located" relative to one another means that they are located on the same premises, such as in the same building, e.g., in the same room, and are preferably interconnected for local communication using a dedicated cable or local area network connection rather than over the open internet.
[0031] The physical or virtual hardware on which the central server 130 runs, in other words, on which the computer software defining the functionality of the central server 130 executes, may consist of a per se conventional CPU, a per se conventional GPU, per se conventional RAM / ROM memory, per se conventional computer buses, and per se conventional external communication capabilities such as an Internet connection.
[0032] Each video communication service 110, insofar as that is used, is also a central server in the above sense, which may be a different central server from central server 130 or may be part of central server 130. In particular, the or each video communication service 110 may be located locally relative to one, more or all of the participant clients 121.
[0033] Correspondingly, each of the participant clients 121 may, in a corresponding interpretation, be a central server in the above sense, with a per se conventional CPU / GPU, per se conventional RAM / ROM memory, per se conventional computer bus, and per se conventional external communication capabilities, such as an Internet connection, on which the physical or virtual hardware on which each participant client 121 executes, in other words the computer software defining the functionality of the participant client 121, is executed.
[0034] Each participant client 121 also typically includes or is in communication with a computer screen arranged to display video content that is provided to the participant client 121 as part of an ongoing video communication, speakers arranged to emit sound content that is provided to the participant client 121 as part of the video communication, a video camera, and a microphone arranged to record sound locally to a human participant 122 to the video communication, who uses that participant client 121 to participate in the video communication.
[0035] In other words, each participant client 121's respective human-machine interface allows each participant 122 to interact with other participants in a video communication and / or with audio / video streams provided from various sources at that client 121.
[0036] Generally, each participant client 121 comprises a respective input means 123, which may consist of the video camera, the microphone, a keyboard, a computer mouse or trackpad, and / or an API for receiving digital video streams, digital audio streams, and / or other digital data. The input means 123 are configured, inter alia, to receive video and / or audio streams from the video communication service 110 and / or a central server, such as central server 130, where such video and / or audio streams are provided as part of the video communication and are preferably generated based on corresponding digital data input streams provided to the central server from at least two sources of such digital data input streams, e.g., the participant client 121 and / or an external source (described below).
[0037] More generally, each participant client 121 includes a respective output means 124, which may consist of the computer screen, the speakers, and an API that emits digital video and / or audio streams that are representative of the locally captured video and / or audio for a participant 122 using that participant client 121.
[0038] In practice, each participant client 121 may be a mobile device, such as a mobile phone, equipped with a screen, speakers, microphone, and Internet connection, running computer software locally or accessing remotely executed computer software to perform the functions of that participant client 121. Correspondingly, a participant client 121 may be a thick or thin laptop or a desktop computer, running locally installed applications or using functions remotely accessed via a web browser.
[0039] There may be one or more participant clients 121, for example at least three or at least four, used in one and the same video communication in this embodiment.
[0040] There may be at least two different groups of participating clients. Each of the participant clients may be assigned to a respective such group. The groups may reflect different roles of the participant clients, different virtual or physical locations of the participant clients, and / or different interaction rights of the participant clients.
[0041] A variety of such roles are available, such as "leader" or "conference organizer", "speaker", "panelist", "interactive audience", and "remote listener".
[0042] Such physical locations may vary widely, for example, "physically in the room," "listening remotely," "on stage," "at a panel," "in a physically present audience," or "in a physically distant audience."
[0043] Virtual locations may be defined in terms of physical locations, but may also include virtual groupings that may overlap with physical locations, for example, physically present audience participants may be divided into a first virtual group and a second virtual group, with some physically present audience participants grouped together with some physically remote audience participants in the same virtual group.
[0044] Such interaction rights can be varied and may be, for example, "full interaction" (no restrictions), "can speak, but only after requesting the microphone" (e.g., raising a virtual hand in a video conferencing service), "cannot speak, but can write in a common chat," or "only watch / listen."
[0045] In some implementations, each defined role and / or physical / virtual location may be defined with respect to certain predetermined interaction permissions. In other examples, all participants with the same interaction permissions form a group. Thus, the defined roles, locations, and / or interaction permissions can reflect various group assignments, and different groups may differ from or overlap with each other as desired.
[0046] The video communications may be provided at least in part by the video communications service 110 and at least in part by the central server 130 as described and illustrated herein.
[0047] As the term is used herein, a "video communication" is a two-way digital communication session involving at least two, preferably at least three, or at least four, video streams, preferably used to generate one or more mixed or collaborative digital video / audio streams, and a matching audio stream, for consumption by one or more consumers (e.g., participant clients of the types described above) who may or may not contribute to the video communication via video and / or audio. Such video communication may be real-time, with or without a certain latency or delay. At least one, preferably at least two, or at least four participants 122 participating in such a video communication engage in the video communication in a two-way manner, providing and consuming video / audio information.
[0048] At least one of the participant clients 121, or all of the participant clients 121, includes local synchronization software functionality 125, which is described in more detail below.
[0049] The video communication services 110 may have or access a common time reference, as described in more detail below.
[0050] Each of the at least one central server 130 may include an API 137 for digitally communicating with entities external to that central server 130. Such communication may include both input and output.
[0051] The system 100, such as the central server 130, may be configured to digitally communicate with external information sources 300, such as externally provided video streams, and in particular to receive digital information, such as audio and / or video stream data, from the external information sources 300. The information sources 300 are "external" in the sense that they are not provided by or as part of the central server 130. Preferably, the digital data provided by the external information sources 300 is independent of the central server 130, and the central server 130 cannot influence its information content. For example, the external information sources 300 may be live-captured video and / or audio, such as a public sporting event or an ongoing news event or report. The external information sources 300 may also be captured by a webcam or the like, rather than by one of the participant clients 121. Thus, such captured video may show the same locale as any one of the participant clients 121, but is not captured as part of the participant client 121's activities. One possible difference between the externally provided information source 300 and the internally provided information source 120 is that the internally provided information source may be provided in that capacity as a participant in the above-defined type of video communication, while the externally provided information source 300 is not, but instead is provided as part of a context that is external to the video conference. In other embodiments, one or more externally provided information sources 300 are in the form of respective digital cameras or microphones configured to capture respective digital image / video and / or audio streams in a manner controlled by the central server 130 in the same locality as one or more of the participant clients 121 and / or corresponding users 122. Thus, the central server 130 can control the on / off state of such digital image / video / audio capture devices 300 and / or other capture states, such as currently applied physical or virtual panning or zooming.
[0052] There may also be multiple external information sources 300 providing digital information of that type, such as audio and / or video streams, to the central server 130 in parallel.
[0053] As shown in FIG. 1, each participant client 121 constitutes the source of a respective information (video and / or audio) stream 120 provided by that participant client 121 to the video communication service 110 as described.
[0054] The system 100, such as the central server 130, may be further configured to digitally communicate with, and in particular emit digital information to, the external consumers 150. For example, digital video and / or audio streams generated by the central server 130 may be provided continuously, in real time or near real time, to one or more external consumers 150 via the API 137. Again, the consumer 150 being "external" means that the consumer 150 is not provided as part of the central server 130 and / or is not a party to the video communication.
[0055] Unless otherwise noted, all functions and communications herein are provided digitally and electronically, implemented by computer software running on appropriate computer hardware and communicated over local or global digital communications networks or channels such as the Internet.
[0056] 1 , multiple participant clients 121 participate in a digital video communication provided by video communication service 110. Each participant client 121 therefore has an ongoing login, session, or the like with video communication service 110 and can participate in one and the same ongoing video communication provided by video communication service 110. In other words, the video communication is "shared" among the participant clients 121 and, therefore, by the corresponding human participants 122.
[0057] 1 , the central server 130 comprises an auto-join client 140, which is an automatic client corresponding to the participant client 121, but which is not associated with a human participant 122. Instead, the auto-join client 140 is added as a participant client to the video communication service 110 to participate in the same shared video communication as the participant client 121. As such a participant client, the auto-join client 140 is given access to continuously generated digital video and / or audio stream(s) provided by the video communication service 110 as part of the ongoing video communication, and such streams can be consumed by the central server 130 via the auto-join client 140. Preferably, the automatic participant client 140 receives from the video communication service 110 a common video and / or audio stream that is or can be distributed to each participant client 121; respective video and / or audio streams that are provided from each of one or more participant clients 121 to the video communication service 110 and relayed by the video communication service 110 to all participant clients 121 or to requesting participant clients 121 in raw or modified form; and / or a common time reference.
[0058] The central server 130 may include a collection function 131 configured to receive multiple video and / or audio streams of the above types from the auto-attendee clients 140, and possibly also from the above-mentioned external information source(s) 300, for processing as described below, and then provide a generated video stream, e.g., a shared video stream, via the API 137. For example, this generated video stream may be consumed by external consumers 150 and / or by the video communication service 110, which may distribute it to all or any requesting one of the participant clients 121.
[0059] FIG. 2 is similar to FIG. 1, but instead of using an auto-join client 140, the central server 130 receives video and / or audio stream data from an ongoing video communication via the API 112 of the video communication service 110.
[0060] 3 is similar to FIG. 1, but does not show the video communication service 110. In this case, the participant clients 121 communicate directly with the API 137 of the central server 130, for example, to provide video and / or audio stream data to and / or receive video and / or audio stream data from the central server 130. The generated shared streams may then be provided to external consumers 150 and / or to one or more of the client participants 121.
[0061] 4 shows the central server 130 in more detail. As shown, the collection function 131 may be composed of one or, preferably, multiple, format-specific collection functions 131 a. Each of the format-specific collection functions 131 a is configured to receive video and / or audio streams having a predetermined format, such as a predetermined binary encoding format and / or a predetermined stream data container, and specifically to parse and classify the binary video and / or audio data in said format into individual video frames, sequences of video frames, and / or time slots.
[0062] The central server 130 further includes an event detection function 132 configured to receive video and / or audio stream data, such as binary stream data, from the collection function 131 and perform respective event detection on each of the received data streams. The event detection function 132 may include an AI (artificial intelligence) component 132a for performing event detection. The event detection may be performed without first time-synchronizing the collected individual streams.
[0063] The central server 130 further comprises a synchronization function 133 configured to time-synchronize multiple data streams that may be provided by the collection function 131 and processed by the event detection function 132. The synchronization function 133 may comprise an AI component 133a for performing the time synchronization.
[0064] The central server 130 may further include a pattern detection function 134 configured to perform pattern detection based on a combination of at least one, but often at least two, e.g., at least three, or at least four, e.g., all, of the received data streams. Pattern detection may further be based on one, and possibly at least two or more, events detected by the event detection function 132 for each of the data streams. Such detected events considered by the pattern detection function 134 may be distributed over time for each collected stream. The pattern detection function 134 may include an AI component 134a for performing pattern detection. Pattern detection may further be based on the groupings described above, and may be configured to detect a specific pattern occurring only for one group, for some groups but not all groups, or for all groups.
[0065] The central server 130 further comprises a generating function 135 configured to generate a generated digital video stream, such as a shared digital video stream, based on one or more data streams provided from the collecting function 131, and possibly based on any detected events and / or patterns. The generated video stream includes at least a generated video stream comprising one or more of the raw, reformatted, or converted video streams provided by the collecting function 131, and may also include corresponding audio stream data. As exemplified below, there may be multiple generated video streams, and one such generated video stream may be generated in the manner described above, but may also be generated based on another already-generated video stream.
[0066] All generated video streams are preferably generated continuously, and preferably in near real time (after subtracting latencies and delays of the type described later herein).
[0067] The central server 130 may further include a publishing function 136 configured to publish the generated shared digital video stream, such as via the API 137 described above.
[0068] It should be noted that while Figures 1, 2 and 3 show three different examples of how the central server 130 can be used to implement the principles described herein, and in particular to provide methods in accordance with the present invention, other configurations are possible, with or without one or more video communication services 110.
[0069] Figure 5 illustrates a method for providing a generated digital video stream, and Figures 6a to 6f show the states of the different digital video / audio data streams resulting from the method steps shown in Figure 5.
[0070] In a first step S500, the method begins.
[0071] In a subsequent collection step S501, a respective plurality of primary digital video streams 210, 301 are collected from at least two of the digital video sources 120, 300, for example by a collection function 131. Each of these plurality of primary data streams 210, 301 may comprise an audio portion 214 and / or a video portion 215. "Video" in this context will be understood to refer to the moving image and / or still image content of such data streams. Each primary data stream 210, 301 may be encoded according to any video / audio coding standard (using a respective codec used by the entity providing that primary stream 210, 301), and the coding format may differ between different ones of the plurality of primary streams 210, 301 used simultaneously in one and the same video communication. At least one, for example all, of the plurality of primary data streams 210, 301 is preferably provided as a stream of binary data, possibly itself provided in a conventional data container data structure. Preferably, at least one, for example at least two, or even all of the plurality of primary data streams 210, 301 are provided as respective live video recordings.
[0072] It should be noted that the multiple primary data streams 210, 301 may not be synchronized in time when received by the collection function 131. This may mean that they are associated with different latencies or delays relative to each other. For example, if two primary video streams 210, 301 are live recordings, this may mean that they are associated with different latencies relative to the recording time when received by the collection function 131.
[0073] It should also be noted that the multiple primary data streams 210, 301 may themselves be respective live camera feeds from webcams; a currently shared screen or presentation; a film clip being viewed; or any combination of these arranged in various ways within one and the same screen.
[0074] The collection step S501 is illustrated in Figures 6a and 6b. Figure 6b also shows how the collection function 131 can store each primary video stream 210, 301 as bundled audio / video information or as audio stream data separated from the associated video stream data. Figure 6b illustrates how the data for the primary video streams 210, 301 is stored as individual frames 213 or collections / clusters of frames, where "frame" refers here to a time-limited portion of image data and / or any associated audio data, for example, each frame being an individual still image or a consecutive series of images (e.g., a series of images that constitute up to one second of moving images) that together form moving image video content.
[0075] In a subsequent event detection step S502 performed by the event detection function 132, the plurality of primary digital video streams 210, 301 are analyzed by the event detection function 132, such as the AI component 132a, to detect at least one event 211 selected from the first set of events. This is shown in Figure 6c.
[0076] This event detection step S502 is preferably performed for at least one, e.g., at least two, e.g., all, primary video streams 210, 301, and is preferably performed separately for each of said primary video streams 210, 301. In other words, the event detection step S502 is preferably performed for each primary video stream 210, 301, taking into account only the information contained as part of that particular primary video stream 210, 301, and in particular without taking into account information contained as part of other primary video streams. Furthermore, event detection is preferably performed without taking into account any common time reference 260 associated with multiple primary video streams 210, 301.
[0077] Preferably, however, event detection takes into account information contained as part of the individually analyzed primary video stream over a time interval, for example over a historical time interval of the primary video stream that is greater than 0 seconds, for example at least 0.1 seconds, for example at least 1 second.
[0078] Event detection may take into account information contained in the audio and / or video data included as part of the primary video stream 210, 301.
[0079] The first set of events may include any number of types of events, such as a change of slide in a slide presentation that constitutes or is part of the primary video stream 210, 301; a change in connection quality of the source 120, 300 providing the primary video stream 210, 301 that results in a change in image quality, loss of image data, or reacquisition of image data; and physical events of movement detected in the primary video stream 210, 301, such as movement of a person or object in the video, a change in lighting in the video, a sudden sharp noise in the audio, or a change in audio quality. It should be understood that this is not intended to be an exhaustive list, and that these examples are provided to help understand the applicability of the presently described principles.
[0080] In a subsequent synchronization step S503 performed by the synchronization function 133, the multiple primary digital video streams 210 may be time-synchronized. This time synchronization may be performed with respect to a common time reference 260. As shown in FIG. 6d, this time synchronization may include aligning the multiple primary video streams 210, 301 with respect to one another using the common time reference 260, for example, so that they can be combined to form a time-synchronized context. The common time reference 260 may be a stream of data, a heartbeat signal or other pulse data, or a time anchor applicable to each of the multiple individual primary video streams 210, 301. By making the common time reference applicable to each of the multiple individual primary video streams 210, 301, the information content of the primary video streams 210, 301 can be uniquely associated with the common time reference with respect to a common time axis. In other words, the common time reference aligns the multiple primary video streams 210, 301 so that they are time-synchronized in a present sense via time shifting. In other embodiments, time synchronization may be based on known information about the time difference between the primary video streams 210, 301, such as measurements.
[0081] As shown in FIG. 6d, time synchronization may include determining one or more timestamps 261 for each of multiple primary video streams 210, 301, for example, relative to a common time reference 260, or for each video stream 210, 301 relative to the other video stream 210, 301 or relative to other multiple video streams 210, 301.
[0082] In a subsequent pattern detection step S504 performed by pattern detection function 134, the time-synchronized multiple primary digital video streams 210, 301 are analyzed to detect at least one pattern 212 selected from the first pattern set. This is shown in Figure 6e.
[0083] In contrast to the event detection step S502, the pattern detection step S504 may be performed based on video and / or audio information included as part of at least two of the multiple time-synchronized primary video streams 210, 301.
[0084] The first set of patterns may include any number of types of patterns, such as multiple participants speaking alternately or simultaneously, or a change in presentation slides occurring simultaneously as another event, such as another participant speaking, etc. This list is not exhaustive but is exemplary.
[0085] In some embodiments, the detected pattern 212 may relate to information contained in only one of the multiple primary video streams 210, 301, rather than information contained in more than one of the multiple primary video streams 210, 301. In such cases, such pattern 212 is preferably detected based on video and / or audio information contained in that single primary video stream 210, 301 across at least two detected events 211, such as two or more consecutively detected presentation slide changes or connection quality changes. As one example, multiple consecutive slide changes that follow each other rapidly over time may be detected as one single slide change pattern, as opposed to one distinct slide change pattern for each detected slide change event. Other examples include the movement of displayed entities or people, recognition of audio phrases spoken by participating users, etc.
[0086] It is understood that the first set of events and the first set of patterns may comprise predetermined types of events / patterns defined using respective sets of parameters and parameter intervals. As described below, the sets of events / patterns may also be defined and detected using various AI tools.
[0087] In a subsequent generation step S505 performed by the generation function 135, a shared digital video stream is generated as an output digital video stream 230 based on a plurality of consecutively considered frames 213 of a plurality of primary digital video streams 210, 301 which may be time-synchronized, and further based on the detected events 211 and / or the detected patterns 212.
[0088] As will be explained and detailed below, the present invention allows for fully automated generation of video streams, such as output digital video stream 230.
[0089] For example, such generation may include selection of what video and / or audio information from which primary video streams 210, 301 to use in the output video stream 230, and to what extent; the video screen layout of the output video stream 230; the switching pattern between different such uses or layouts over time; etc.
[0090] This is also illustrated in Figure 6f, which shows one or more additional portions of time-related (which may be relative to the common time reference 260) digital video information 220, such as additional digital video information streams that may be time-synchronized (e.g., relative to the common time reference 260) and used in conjunction with the time-synchronized multiple primary video streams 210, 301 in generating the output video stream 230. For example, the additional streams 220 may include information regarding any video and / or audio special effects to use, such as dynamically based on detected patterns; a planned time schedule for the video communication; etc.
[0091] In a subsequent publishing step S506 performed by the publishing function 136, the generated output digital video stream 230 is continuously provided to the consumers 110, 150 of the shared digital video stream, as described above. The generated digital video stream may be provided to one or more participant clients 121, for example, via the video communication service 110.
[0092] At the following step S507, the method ends. However, initially, the method may be iterated any number of times to generate the output video stream 230 as a continuously provided stream, as shown in FIG. 5. Preferably, the output video stream 230 is generated to be consumed in real time or near real time (taking into account the total latency added by all steps along the way) and continuously (published as soon as more information becomes available, but not counting any intentionally added latency, see below). In this way, the output video stream 230 may be consumed in an interactive manner, whereby the output video stream 230 is fed back to the video communication service 110 or to another context that forms the basis for generating the primary video stream 210, which is fed back to the collection function 131 to form a closed feedback loop; or the output video stream 230 is consumed in a different context (external to the system 100, or at least external to the central server 130), where it may form the basis for real-time two-way video communication.
[0093] As mentioned above, in some embodiments, at least two, e.g., at least three, e.g., at least four or at least five of the multiple primary digital video streams 210, 301 are provided as part of a shared digital video communication such as provided by a video communication service 110, the video communication including respective remotely connected participant clients 121 providing said primary digital video streams 210. In such cases, the collecting step S501 may comprise collecting at least one of said primary digital video streams 210 from the shared digital video communication service 110 itself, via an auto-participant client 140 that has in turn been granted access to the video and / or audio stream data from within said video communication service 110, and / or via the API 112 of the video communication service 110.
[0094] Additionally, in this and other cases, the collecting step S501 may comprise collecting at least one of the plurality of primary digital video streams 210, 301 as a respective external digital video stream 301 collected from an information source 300 that is external to the shared digital video communication service 110. Note that one or more of such external video sources 300 may be external to the central server 130.
[0095] In some embodiments, the multiple primary video streams 210, 301 are not formatted in the same manner. While such different formats can be in the form in which they are supplied to the collection function 131 in different types of data containers (such as AVI or MPEG), in preferred embodiments, at least one of the multiple primary video streams 210, 301 is formatted according to a deviating format (relative to at least one other of the primary video streams 210, 301), in that the deviating primary digital video streams 210, 301 have deviating video encodings; deviating fixed or variable frame rates; deviating aspect ratios; deviating video resolutions; and / or deviating audio sample rates.
[0096] The collection function 131 is preferably pre-configured to read and interpret all encoding formats, container standards, etc. occurring in all collected primary video streams 210, 301. This allows processing as described herein to be performed without requiring decoding until relatively late in these processes (e.g., until the primary streams in question are placed in their respective buffers; or until after the event detection step S502; or until after the event detection step S502). However, in the rare case where one or more of the primary video feeds 210, 301 are encoded using a codec that the collection function 131 cannot interpret without decoding, the collection function 131 may be configured to perform decoding and analysis of such primary video streams 210, 301, followed by conversion to a format that can be processed, for example, by the event detection function. Note that even in this case, re-encoding is preferably not performed at this stage.
[0097] For example, a primary video stream 220 fetched from a multi-party video event, such as that provided by video communication service 110, typically has low latency requirements and is therefore typically associated with variable frame rates and variable pixel resolutions to enable participants 122 to communicate effectively. In other words, the overall video and audio quality is degraded as necessary for low latency.
[0098] On the other hand, the external video feed 301 typically has a more stable frame rate and higher image quality, but may therefore have a higher latency.
[0099] Thus, the video communication service 110 may at each point in time use a different encoding and / or container than the external video source 300. Therefore, the analysis and video generation process described herein must then combine these multiple streams 210, 301 of different formats into a new single stream for a combined experience.
[0100] As mentioned above, the collection functionality 131 may comprise a set of format-specific collection functionality 131a, each configured to process a particular type of format of the primary video stream 210, 301. For example, each one of these format-specific collection functionality 131a may be configured to process multiple primary video streams 210, 301 encoded using different respective video encoding methods / codecs, such as Windows® Media® or DivX®.
[0101] However, in some embodiments, the collecting step S501 includes converting at least two, eg, all, of the plurality of primary digital video streams 210, 301 to a common protocol 240.
[0102] As used in this context, the term "protocol" refers to an information structuring standard or data structure that specifies how information contained in a digital video / audio stream is stored. However, the common protocol preferably does not prescribe how digital video and / or audio information is stored, e.g., at a binary level (i.e., encoded / compressed data indicating the sounds and images themselves), but instead forms a structure of a predetermined format for storing such data. In other words, the common protocol prescribes storing digital video data in a raw binary format without performing any digital video decoding or encoding in connection with such storage, and without modifying the existing binary format in any way apart from concatenating and / or splitting binary byte sequences, as the case may be. Instead, the raw (encoded / compressed) binary data content of the primary video stream 210, 301 is preserved by repacking the raw binary data in a data structure defined by the protocol. In some embodiments, the common protocol defines a video file container format.
[0103] FIG. 7 shows, by way of example, multiple primary video streams 210, 301 shown in FIG. 6a reconstructed by respective format-specific collection functions 131a and using the common protocol 240 described above.
[0104] Thus, the common protocol 240 provides for storing digital video and / or audio data in data sets 241 that are preferably divided into discrete, contiguous sets of data along a time axis relative to the primary video stream 210, 301. Each such data set may contain one or several frames of video and associated audio data.
[0105] The common protocol 240 may also provide for storing, in association with the stored digital video and / or audio data set 241, metadata 242 associated with a specified point in time.
[0106] The metadata 242 may include information about the binary format of the raw data of the primary digital video stream 210, such as about the digital video encoding method or codec used to generate the raw binary data, the resolution of the video data, the video frame rate, a frame rate variation flag, the video resolution, the video aspect ratio, the audio compression algorithm, or the audio sampling rate. The metadata 242 may also include information about timestamps of the stored data, for example, related to the time base of the primary video stream 210, 301 itself or related to a different video stream as mentioned above.
[0107] The use of format-specific collection functions 131a in combination with the common protocol 240 allows for the rapid collection of information content of the primary video streams 210, 301 without the added latency (delay) of decoding / re-encoding the received video / audio data.
[0108] Thus, the collecting step S501 may comprise collecting a plurality of primary digital video streams 210, 301 encoded using different binary video and / or audio encoding formats using different ones of a plurality of format-specific collection functions 131a in order to parse the primary video streams 210, 301 and store the parsed raw binary data, along with any associated metadata, in a data structure using a common protocol. It will be appreciated that the decision as to which format-specific collection function 131a to use for which primary video stream 210, 301 may be made by the collection functions 131a based on predetermined and / or dynamically detected characteristics of each of the primary video streams 210, 301.
[0109] Each primary video stream 210, 301 collected in this manner may be stored in its own separate memory buffer, such as a RAM memory buffer within the central server 130.
[0110] The conversion of the primary video streams 210, 301 performed by each format-specific collection function 131a may therefore comprise splitting the raw binary data of each thus converted primary digital video stream 210, 301 into an ordered set of smaller data sets 241.
[0111] Furthermore, the conversion may also comprise associating each of the smaller sets 241 (or subsets, e.g., subsets regularly distributed along the time axis of each of the primary streams 210, 301 in question) with a respective time along a shared time axis, e.g., with respect to a common time reference 260. This association may be performed by analysis of the raw binary video and / or audio data, in one of the principle ways described below, or in other ways, and may be performed to enable subsequent time synchronization of the primary video streams 210, 301 to be performed. Depending on the type of common time reference 260 used, at least part of this association of each data set 241 may also be performed by or instead of the synchronization function 133. In the latter case, the collecting step S501 may instead comprise associating each of the smaller sets 241, or subsets thereof, with a respective time on a time axis specific to the primary stream 210, 301 in question.
[0112] In some embodiments, the collecting step S501 also includes converting the raw binary video and / or audio data collected from the multiple primary video streams 210, 301 to uniform quality and / or frequency updating. This may include downsampling or upsampling the raw binary digital video and / or audio data of the multiple primary digital video streams 210, 301 to a common video frame rate; a common video resolution; or a common audio sampling rate, as appropriate. Note that such resampling can be performed without performing full decoding / re-encoding, or even without performing any decoding at all, because the format-specific collecting function 131a can directly process the raw binary data according to the correct binary encoding target format.
[0113] Preferably, each of the multiple primary digital video streams 210, 301 is stored in an individual data storage buffer 250 as an individual frame 213 or sequence of frames 213, as described above, and each is associated with a corresponding timestamp that is in turn associated with a common time base 260.
[0114] In the specific example provided for illustrative purposes, the video communication service 110 is Microsoft® Teams® and is conducting a video conference involving multiple simultaneous participants 122. The auto-join client 140 is registered as a conference participant in the Teams® conference.
[0115] Next, the primary video input signals 210 are provided to the collection function 130 via the auto-join client 140 and are acquired by the collection function 130. These are raw data signals in H264 format, and include timestamp information for each video frame.
[0116] The associated format-specific collection function 131a picks up the raw data over IP (LAN network) on a configurable predefined TCP port. Audio data associated with every Teams® conference participant is associated with a separate port. The collection function 131 then uses timestamps from the audio signal (50 Hz) to downsample the video data to a fixed 25 Hz output signal before storing the video streams 220 in their respective individual buffers 250.
[0117] As mentioned above, common protocol 240 stores data in a raw binary format. It can be designed to process raw bits and bytes of video / audio data at a very low level. In a preferred embodiment, data is stored in common protocol 240 as a simple byte array or corresponding data structure (such as a slice). This means that the data does not need to be placed in a traditional video container at all (common protocol 240 does not constitute such a traditional container in this context). Also, video encoding and decoding is computationally intensive, causing delays and requiring expensive hardware. Furthermore, this problem scales with the number of participants.
[0118] The common protocol 240 allows memory to be reserved within the collection function 131 for the primary video stream 210 associated with each Teams® conference participant 122 and any external video sources 300, and the amount of allocated memory can be changed on the fly during the process. In this way, the number of input streams can be changed, thereby keeping each buffer valid. For example, information such as resolution and frame rate is variable, but is stored as metadata in the common protocol 240, so this information can be used to quickly change the size of each buffer as needed.
[0119] The following is an example of the specification of this type of common protocol 240:
[0120] [Table 1]
[0121] In the above table, the "Detected event in, if any" data is included as part of the specification of common protocol 260. However, in some embodiments, this information (regarding detected events) may instead be placed in a separate memory buffer.
[0122] In some embodiments, the at least one additional portion of the digital video information 220, which may be an overlay or effect, is also stored in each individual buffer 250 as an individual frame or sequence of frames, each associated with a corresponding timestamp that is in turn associated with the common time reference 260.
[0123] As illustrated above, the event detection step S502 may include using a common protocol 240 to store metadata 242 describing the detected event 211 in association with the primary digital video stream 210, 301 in which the event 211 was detected.
[0124] Event detection can be performed in different ways. In some embodiments performed by the AI component 132a, the event detection step S502 includes a first trained neural network or other machine learning component individually analyzing at least one, e.g., some, or all, of the multiple primary digital video streams 210, 301 to automatically detect any of the events 211. This may involve the AI component 132a classifying the data of the primary video streams 210, 301 into a predefined set of events in a supervised classification and / or into a dynamically determined set of events in an unsupervised classification.
[0125] In some embodiments, the detected event 211 is a change in a presentation slide of a presentation that is or is included in the primary video stream 210, 301.
[0126] For example, if a presenter of a presentation decides to change the slide in the presentation that they are currently making to the audience, this means that what is interesting to a given viewer may change. The newly displayed slide may simply be a general-level image that is best viewed for a short time in so-called "butterfly" mode (e.g., displaying the slide side-by-side with the presenter's video in output video stream 230). Or the slide may contain a lot of detail, text in a small font size, etc. In the latter case, the slide will be displayed full screen and for a somewhat longer time than would normally be the case. In this case, the slide may be more interesting to the viewer than the presenter's face, so butterfly mode may not be as appropriate.
[0127] In practice, the event detection step S502 consists of at least one of the following:
[0128] First, the event 211 may be detected based on an image analysis of the difference between a first image of a detected slide and a subsequent second image of the detected slide. The nature of the primary video stream 220, 301 as being indicative of a slide may be determined automatically using conventional digital image processing, such as using motion detection in combination with OCR (Optical Character Recognition).
[0129] This may involve using automated computer image processing techniques to check whether the detected slide has changed significantly enough to be classified as a true slide change. This can be done by checking the delta between the current and previous slides in terms of RGB color values. For example, one can evaluate how much the RGB values have changed globally in the screen area covered by the slide in question, and simultaneously evaluate whether it is possible to find groups of adjacent pixels that change in concert with this. This allows relevant slide changes to be detected while filtering out irrelevant changes, such as computer mouse movements across the screen. This approach allows for complete configurability. For example, it may be desirable to be able to capture computer mouse movements, for example, if a presenter wants to present something in detail while pointing at different things with the computer mouse.
[0130] Second, the event 211 may be detected based on image analysis of the information complexity of the second image itself to determine the type of event with greater specificity.
[0131] This might involve, for example, assessing the amount of textual information on the slide in question and the associated font size, which can be done using traditional OCR methods, including deep learning-based character recognition techniques.
[0132] Note that because the raw binary format of the evaluated video streams 210, 301 is known, this may be performed directly in the binary domain without first decoding or re-encoding the video data. For example, the event detection function 132 may invoke an associated format-specific collection function for an image interpretation service, or the event detection function 132 itself may include functionality for evaluating image information, such as to the individual pixel level, for a number of different supported raw binary video data formats.
[0133] In another example, the detected event 211 is a loss of a communication connection of a participant client 121 to the digital video communication service 110. In this case, the detecting step S502 may include detecting that the participant client 121 has lost the communication connection based on image analysis of a series of subsequent video frames 213 of the primary digital video stream 210 corresponding to that participant client 121.
[0134] Because participant clients 121 may be associated with different physical locations and different Internet connections, it may occur that someone loses connection to the video communication service 110 or the central server 130. In such a situation, it is desirable to avoid a black or blank screen appearing in the generated output video stream 230.
[0135] Alternatively, such loss of connection can be detected as an event by the event detection function 132, for example, by applying a two-class classification algorithm where the two classes used are connected / not connected (no data). In this case, "no data" is understood to be different from the presenter intentionally sending a black screen. Because a short black screen, such as just one or two frames, may not be noticeable in the final generated stream 230, the two-class classification algorithm can be applied over time to create a time series. A threshold specifying the minimum length of a connection interruption can then be used to determine whether a connection has been lost.
[0136] In another example, an event is the detection of the presence or movement of a participating human user within one or more images of said primary digital video stream 210, 301. In another example, an event is the detection of movement (rotation, zoom, pan, etc.) of a camera used to generate said primary digital video stream 210, 301, possibly including information about the general and / or noisy motion components of such movement. The noisy motion component may be due, for example, to the camera being manually moved. Such human user presence / motion detection and / or camera movement detection may be achieved using per se conventional digital image processing techniques, for example as exemplified above.
[0137] As will be explained below, detected events of the types exemplified above may be used by pattern detection function 134 to take various responses, as appropriate and desired.
[0138] As noted above, the individual primary video streams 210, 301 are each associated with a common time reference 260 and can be time-synchronized relative to one another by the synchronization function 133.
[0139] In some embodiments, the common time reference 260 is based on or comprises a common audio signal 111 (see Figures 1 to 3), which, as described above, is common to a shared digital video communication service 110 participating in at least two remotely connected participant clients 121, each providing a respective one of the primary digital video streams 210.
[0140] In the Microsoft® Teams® example described above, a common audio signal may be generated and captured by the central server 130 via the auto-join client 140 and / or via the API 112. In this and other examples, such common audio signal may be used as a heartbeat signal to time-synchronize the individual primary video streams 220 by combining the individual primary video streams 220 at specific points in time based on the heartbeat signal. Such a common audio signal may be provided as a separate signal (relative to each of the other primary video streams 210), such that each of the other primary video streams 210 may be individually time-correlated to the common audio signal based on audio contained in the other primary video streams 210 or based on image information contained therein (e.g., using automatic image processing-based lip-sync techniques).
[0141] In other words, to handle variable and / or different latencies associated with the individual primary video streams 210 and to achieve time synchronization of the combined video output stream 230, such a common audio signal may be used as a heartbeat for all primary video streams 210 within the central server 130 (but perhaps not the external primary video stream 301). In other words, all other signals may be mapped to this common audio time heartbeat to ensure they are all time synchronized.
[0142] In another example, time synchronization is achieved using a time synchronization element 231 that is introduced into the output digital video stream 230 and detected by a respective local time synchronization software function 125 provided as part of one or more individual ones of the participant clients 121. The local software function 125 is configured to detect the time of arrival of the time synchronization element 231 in the output video stream 230. As will be appreciated, in such an embodiment, the output video stream 230 is fed back to the video communication service 110 or otherwise made available to each participant client 121 and its local software function 125.
[0143] For example, the time synchronization elements 231 may be visual markers, such as pixels that change color in a predetermined order or manner, that are placed or updated in the output video 230 at regular time intervals; a visual clock that is updated and displayed in the output video 230; or an audio signal (which may be designed to be inaudible to the participants 122, for example, by having a sufficiently low amplitude and / or a sufficiently high frequency) that is added to the audio that forms part of the output video stream 230. The local software function 125 is configured to automatically detect the arrival time of each of the time synchronization elements 231 using appropriate image and / or audio processing.
[0144] The common time reference 260 may then be determined, at least in part, based on the detected arrival times. For example, each of the local software functions 125 may communicate respective information indicative of the detected arrival times to the central server 130.
[0145] Such communication may occur via a direct communication link between the participant client 121 and the central server 130. However, communication may also occur via a primary video stream 210 associated with the participant client 121. For example, the participant client 121 may introduce a visual or audible code, such as the type described above, into the primary video stream 210 generated by the participant client 121 for automatic detection by the central server 130 and use to determine the common time reference 260.
[0146] In still additional embodiments, each participant client 121 may perform image detection on a common video stream viewable by all participant clients 121 for the video communication service 110 and relay the results of such image detection to the central server 130 in a manner corresponding to that described above, where they may be used to determine the respective offsets of each participant client 121 relative to one another over time. In this manner, a common time reference 260 may be determined as a set of individual relative offsets. For example, a selected reference pixel of the commonly available video stream may be monitored by some or all participant clients 121, such as by local software functions 125, and the current color of that pixel may be communicated to the central server 130. The central server 130 may generate an estimated set of relative time offsets across the different participant clients 121 by calculating respective time series based on such color values received consecutively from each of many (or all) of the participant clients 121 and performing cross-correlation.
[0147] In practice, the output video stream 230 provided to the video communication service 110 may be included as part of the shared screen to all participant clients of that video communication and may therefore be used to evaluate such time offsets associated with the participant clients 121. In particular, the output video stream 230 provided to the video communication service 110 may be made available back to the central server via the auto-join client 140 and / or API 112.
[0148] In some embodiments, the common time reference 260 may be determined based at least in part on a detected discrepancy between the audio portion 214 of a first one of the multiple primary digital video streams 210, 301 and the image portion 215 of said first one of the multiple primary digital video streams 210, 301. Such discrepancy may be based, for example, on a digital lip-sync video image analysis of the speaking participant 122 viewed in said first primary digital video stream 210, 301. Such lip-sync analysis may be conventional per se and may, for example, use a trained neural network. The analysis may be performed by the synchronization function 133 for each primary video stream 210, 301 with respect to available common audio information, and the relative offset across the individual primary video streams 210, 301 may be determined based on this information.
[0149] In some embodiments, the synchronization step S503 includes intentionally introducing a delay (in this context, "delay" and "latency" are intended to mean the same thing) of up to 30 seconds, e.g., up to 5 seconds, e.g., up to 1 second, e.g., up to 0.5 seconds, but more than 0 seconds, to provide at least that delay in the output digital video stream 230. Whatever the length, the intentionally introduced delay is at least a plurality of video frames, such as at least three, or at least five, or even ten, such as this number of frames (or individual images) stored after any resampling in the acquisition step S501. As used herein, the term "intentionally" means that the delay is introduced regardless of the need to introduce such a delay based on synchronization issues or the like. In other words, the intentionally introduced delay is in addition to the delay introduced as part of the synchronization of the multiple primary video streams 210, 301, in order to time-synchronize the multiple primary video streams 210, 301 with each other. The intentionally introduced delay may be predetermined, fixed, or variable relative to the common time reference 260. The delay time may be measured relative to the least latent one of the multiple primary video streams 210, 301, and as a result of the time synchronization, the more latent ones of these streams 210, 301 may be associated with a relatively small intentionally added delay.
[0150] In some embodiments, a relatively small delay, such as 0.5 seconds or less, is introduced, which delay is barely noticeable to participants in the video communication service 110 using the output video stream 230. In other embodiments, a larger delay may be introduced, such as when the output video stream 230 is not used in an interactive context but instead is published in a one-way communication to an external consumer 150.
[0151] This intentionally introduced delay may be sufficient to allow the synchronization function 133 sufficient time to map the collected video frames of the individual primary streams 210, 301 to the correct timestamps 261 of the common time reference 260. It may also be sufficient to provide sufficient time to perform the event detection described above to detect missing primary stream 210, 301 signals, slide changes, resolution changes, etc. Furthermore, the intentionally introduced delay may be sufficient to improve the pattern detection function 134, as described below.
[0152] It will be understood that introducing a delay involves buffering 250 each of the collected and time-synchronized multiple primary video streams 210, 301 before publishing the output video stream 230 using that buffered frame 213. In other words, the video and / or audio data of at least one, some, or all of the multiple primary video streams 210, 301 may be present in the central server 130 in a buffered manner, for the reasons discussed above, particularly for use by the pattern detection function 134, rather than being used like a cache but intended to be able to handle varying bandwidth situations (as in a traditional cache buffer).
[0153] Thus, in some embodiments, the pattern detection step S504 involves considering specific information of at least one, e.g., some, e.g., at least four, or all, of the multiple primary digital video streams 210, 301, which specific information resides in a frame 213 that follows a frame 213 of the time-synchronized primary digital video stream 210 that has not yet been used in generating the output digital video stream 230. Thus, a newly added frame 213 resides in the buffer 250 for a specific waiting time before forming part of (or the basis for) the output video stream 230. During this period, the information of that frame 213 constitutes "future" information with respect to the frame currently being used to generate the current frame of the output video stream 230. When the timeline of the output video stream 230 reaches that frame 213, that frame is used to generate the corresponding frame of the output video stream 230 and may thereafter be discarded.
[0154] In other words, the pattern detection function 134 has at its disposal a set of video / audio frames 213 that have not yet been used to generate the output video stream 230, and uses this data to detect the above patterns.
[0155] Pattern detection can be performed in different ways: In some embodiments performed by the AI component 134a, the pattern detection step S504 includes a second trained neural network or other machine learning component analyzing in concert at least two, e.g., at least three, e.g., at least four, or all of the multiple primary digital video streams 120, 301 to automatically detect the pattern 212.
[0156] In some embodiments, the detected pattern 212 comprises a speech pattern including at least two, e.g., at least three, e.g., at least four different speaking participants 122, each associated with a respective participant client 121, for the shared video communication service 110, and each of these speaking participants 122 may be visually viewed in a respective one of the multiple primary digital video streams 210, 301.
[0157] Preferably, the generating step S505 includes determining, tracking, and updating the current generation state of the output video stream 230. For example, such state can dictate which participants 122 (if any) are visible in the output video stream 230 and where on-screen they are visible; which external video streams 300 are visible in the output video stream 230 and where on-screen they are visible; whether any slides or shared screens are displayed in full-screen mode or in combination with any live video streams; etc. Furthermore, such state can dictate the cropping or virtual panning / zooming of any one of the primary digital video streams 210, 301 used at any one instance. Thus, the generating function 135 can be viewed as a state machine for the generated output video stream 230.
[0158] In order to generate the output video stream 230 as a combined video experience to be viewed, for example, by the end consumer 150, it is advantageous for the central server 130 to be able to understand what is happening at a deeper level than simply detecting individual events associated with the individual primary video streams 210, 301.
[0159] In a first example, the presenting participant client 121 changes the currently displayed slide. This slide change is detected by the event detection function 132 as described above, and metadata 242 is added to the frame indicating that a slide change has occurred. This happens multiple times as the presenting participant client 121 is found to be skipping forward a number of slides in rapid succession, resulting in a series of "slide change" events that are also detected by the detection function 132 and stored along with the corresponding metadata 242 in a separate buffer 250 of the primary video stream 210. In practice, each such rapidly skipped forward slide may only be displayed for a few seconds.
[0160] The pattern detection function 134 looks at the information in the buffer 250 across these detected slide changes and detects a pattern that corresponds to one single slide change rather than multiple or rapidly executed slide changes (i.e., a single slide change to the last slide in a forward skip, with that last slide remaining visible once the fast skip ends). In other words, the pattern detection function 134 notes, for example, that there were 10 slide changes in a very short period of time, and so treats them as a detected pattern representing one single slide change. As a result, the generation function 135 has access to the pattern detected by the pattern detection function 134 and, because it determines that this last slide is potentially important in the state machine, it can choose to display that last slide in full-screen mode for a few seconds in the output video stream 230. It can also choose not to display intermediately viewed slides at all in the output stream 230.
[0161] Detection of patterns with multiple rapid slide changes may be detected by simple rule-based algorithms, but alternatively may be detected using neural networks designed and trained to detect such patterns in video images by classification.
[0162] In another example, it may be desirable to quickly switch visual attention between current speakers while still providing a relevant viewing experience for the consumer 150 by generating and presenting a calm and smooth output video stream 230, as may be useful, for example, when the video communication is a talk show, panel discussion, or the like. In this case, the event detection function 132 may continuously analyze each primary video stream 210, 301 to determine at any given time whether the person being viewed in that particular primary video stream 210, 301 is currently speaking. This may be performed, for example, as described above, using conventional image processing tools. The pattern detection function 134 may then be operable to detect certain overall patterns involving multiple primary video streams 210, 301, which patterns are useful for generating a smooth output video stream 230. For example, the pattern detection function 134 may detect patterns of very frequent switches between current speakers and / or patterns involving multiple simultaneous speakers.
[0163] The generation functionality 135 can then take such detected patterns into account when making automated decisions related to the generation state, such as, for example, not automatically switching visual focus to a speaker who speaks for only half a second and then goes silent again, or switching to a state where multiple speakers are displayed side by side during a period where they are alternating or speaking simultaneously. This state determination process can itself be performed using time-series pattern recognition techniques or using trained neural networks, but can also be based at least in part on a predetermined set of rules.
[0164] In some embodiments, there may be multiple patterns detected in parallel and form input to the state machine of the generation function 135. Such multiple patterns may be used by the generation function 135 in different AI components, computer vision detection algorithms, etc. As an example, a permanent slide change may be detected while simultaneously detecting unstable connections for some participant clients 121, and other patterns may detect the current main speaking participant 122. Using all such available pattern data, a classifier neural network may be trained and / or a set of rules may be developed to analyze the time series of such pattern data. Such classification may be supervised, at least in part, e.g., completely, to yield determined desired state changes used in the generation. For example, different such predetermined classifiers may be generated that are specifically configured to automatically generate the output video stream 230 according to a variety of different generation styles and desires. Training may be based on known generation state change sequences as the desired output and known pattern time series data as training data. In some embodiments, a Bayesian model may be used to generate such classifiers. In a concrete example, a priori information can be obtained from experienced producers, who can provide input such as "In talk shows, we never switch directly from speaker A to speaker B, but always give an overview first before focusing on other speakers unless they are very dominant and loud." This generation logic is expressed as a Bayesian model of the general form "If X is true, then | given the fact that Y is true, | do Z." The actual detection (e.g., whether someone is speaking loudly) can be done using classifiers or threshold-based rules.
[0165] Given large datasets (of pattern time series data), deep learning techniques can be used to develop correct and attractive generative formats for use in the automatic generation of video streams.
[0166] In some embodiments, the generation functionality 135 may include information regarding what objects or human participants to display in the output video stream 230. For example, in a particular setting, one or more participants 122 may not be desired or permitted to be visible to consumers of the output video stream 230. In this case, the generation functionality 135 may make a generation decision to generate the output video stream 230 without displaying such one or more object or human participants based on digital image processing techniques to recognize the object or person, automatically crop the primary video stream 210, 301 to not include the object or person before it is added as part of the output video stream 230, or not include in the output video stream 230 a primary video stream 210, 301 that currently displays the object or person.
[0167] In some embodiments, the generation functionality 130 may be configured to introduce externally provided information, such as in the form of the primary digital video stream 300 or other types of externally provided data, in response to a pattern detected in one or more of the multiple primary digital video streams 210, 301. For example, the generation functionality 130 may be configured to automatically detect a discussion topic or a predetermined trigger event or pattern (such as a predetermined trigger phrase) through digital processing of images and / or audio contained in the primary video streams 210, 301. In a specific example, this may include automatically introducing updated textual or graphical information from a remote source into the output video stream 230 regarding a topic currently being discussed by participating users 122. In general, detection of such a trigger event or pattern may cause the generation functionality 130 to alter the currently used generation state in any manner as a function of the type or characteristics of the detected trigger event or pattern.
[0168] In summary, by using a combination of event detection based on multiple individual primary video streams 210, 301; intentionally introduced delay; pattern detection based on multiple time-synchronized primary video streams 210, 301 and detected events; and a generation process based on the detected patterns, it is possible to achieve automatic generation of the output digital video stream 230 according to a wide range of possible tastes and styles. This result is valid across a wide range of possible neural network and / or rule-based analysis techniques used by the event detection function 132, pattern detection function 134, and generation function 135.
[0169] As illustrated above, the generating step S505 may include generating the output digital video stream 230 based on a set of predetermined and / or dynamically variable parameters relating to the visibility of each of the plurality of primary digital video streams 210, 301 in the output digital video stream 230, the arrangement of visual and / or audio video content, the visual or audio effects used, and / or the output mode of the output digital video stream 230. Such parameters may be determined automatically by a state machine in the generating functionality 135, and / or set by an operator controlling the generation (semi-automated), and / or predetermined based on some a priori organizational desire (such as a minimum time between layout changes of the output video stream 230 or state changes of the types illustrated above).
[0170] In a practical example, the state machine may support a set of predefined standard layouts that can be applied to the output video stream 230, such as a full-screen presenter view (showing the currently speaking participant 122 in full screen); a slide view (showing the currently shared presentation slide in full screen); a "butterfly view" (showing both the currently speaking participant 122 and the currently shared presentation slide in a side-by-side view); or a multi-speaker view (showing all or a selected subset of the participants 122 in a side-by-side or matrix layout). Various available production formats can be defined by a set of available states (such as the set of standard layouts described above) along with a set of state machine state change rules (such as those illustrated above). For example, one such production format could be "panel discussion," another "presentation," etc. By selecting a particular production format via a GUI or other interface to the central server 130, an operator of the system 100 can quickly select one of a set of such predefined production formats and then enable the central server 130 to fully automatically generate the output video stream 230 in accordance with that production format based on available information such as those described above.
[0171] Additionally, during production, a respective in-memory buffer may be created and maintained for each conference participant client 121 or external video source 300, as described above. These buffers can be easily deleted, added, and modified on the fly. The central server 130 may then be configured to receive information regarding added / dropped-off participant clients 121 and participants 122 scheduled to speak during production of the output video stream 230; planned or unexpected pauses / resumes of the presentation; desired changes to the currently used production format; and the like. Such information may be provided to the central server 130, for example, via an operator GUI or interface, as described above.
[0172] As illustrated above, in some embodiments, at least one of the multiple primary digital video streams 210, 301 may be provided to a digital video communication service 110, and then the publishing step S506 may include providing the output digital video stream 230 to that same communication service 110. For example, the output video stream 230 may be provided to a participant client 121 of the video communication service 110, or may be provided as an external video stream to the video communication service 110 via the API 112. In this manner, the output video stream 230 may be made available to several or all of the participants in the video communication event currently being facilitated by the video communication service 110.
[0173] As also mentioned above, the output video stream 230 may additionally or alternatively be provided to one or more external consumers 150 .
[0174] Generally, the generating step S505 is performed by the central server 130, which can provide the output digital video stream 230 as a live video stream via the API 137 to one or more concurrent consumers.
[0175] Figure 8 illustrates a method according to a first aspect of the present invention for providing an output digital video stream, which method will now be described with reference to the above. Thus, in the method illustrated in Figure 8, all of the mechanisms and principles described above for collection, event detection, synchronization, pattern detection, generation and publishing of digital video streams can be applied.
[0176] 9 and 10 are respective simplified diagrams of a system 100 configured to perform the method illustrated in FIG.
[0177] In FIG. 9, three different central servers 130′, 130″, 130′″ are shown. These central servers 130′, 130″, 130′″ may be a single integrated central server of the type described above, or may be multiple separate such central servers. They may or may not run on the same physical or virtual hardware. In any case, they are configured to communicate with each other.
[0178] In some embodiments, the central servers 130′ and 130″ may be configured to run on the same piece of physical hardware 402 (illustrated by the dotted rectangle in FIG. 9 ), e.g., in the form of separate hardware appliances, such as conventional computing devices. In some embodiments, such separate hardware appliances 402 are computing devices 402′ (see FIG. 10 ) located in or physically connected to a conference room and specifically configured for conducting digital video conferences in that conference room. In other embodiments, the separate hardware appliances are personal computers 402″, 402′″, such as laptop computers, used by individual human conference participants 122, 122″, 122′″ for such digital video conferences, where participants 122, 122″, 122′″ are present in the room or remotely.
[0179] Each of the central servers 130', 130", 130'" includes a respective collection functionality 131', 131", 131'" as generally described above. Collection functionality 131' is configured to collect digital video streams 401 from digital cameras (such as video cameras 123 of the type generally described above). Such digital cameras may be integral parts of the individual hardware appliances 402 described above, or may be separate cameras connected to the hardware appliances 402 using appropriate wired or wireless digital communications channels. In either case, the cameras are preferably located locally relative to the hardware appliances 402.
[0180] Each of the collection functions 131'', 131''' can collect a digital video signal corresponding to the digital video stream 401 either directly from a digital camera or from the collection function 131'.
[0181] Each of the central servers 130', 130", 130'" may also include a respective generating function 135', 135", 135'". Each such generating function 135', 135", 135'" corresponds to the generating function 135 described above, and what was described above with respect to generating function 135 applies equally to generating functions 135', 135", 135'". Also, there may be more than three generating functions, depending on the detailed configuration of the central servers 130', 130", 130'". Various digital communications between generating functions 135', 135", 135'" and other entities may occur via appropriate APIs.
[0182] Additionally, each of the central servers 130', 130", 130'" may include a respective publishing function 136', 136", 136'". Each such publishing function 136', 136", 136'" corresponds to the publishing function 136 described above, and what was described above with respect to the publishing function 136 applies equally to the publishing function 136', 136", 136'". The publishing functions 136', 136", 136'" may be separate functions, multiple functions may be co-located in a single logical function, or there may be more than three publishing functions, depending on the detailed configuration of the central servers 130', 130", 130'". The publishing functions 136', 136", 136'" may in some cases be different functional aspects of the same publishing function 136.
[0183] Publishing functions 136" and 136'" are optional and may be configured to output a different (possibly more elaborate) video stream than the video stream output by publishing function 136', whereas publishing function 136' is configured to output an output digital video stream in accordance with the present invention. Generally speaking, generation functions 135" and 135'" are configured to process the input video stream to generate generation control parameters used by generation function 135' to generate said output video stream in accordance with the present invention.
[0184] FIG. 9 also illustrates three external consumers 150′, 150″, 150′″, each corresponding to the external consumer 150 described above. It is understood that there may be fewer or more than three such external consumers 150′, 150″, 150′″. For example, two or more of the publishing functions 136′, 136″, 136′″ may output to the same external consumer 150′, 150″, 150′″, or one of the publishing functions 136′, 136″, 136′″ may output to two or more of the external consumers 150′, 150″, 150′″. It should also be noted that at least the publishing function 136′ may publish the generated video stream back to the collection function 131′. Furthermore, each of the publishing functions 136′, 136″, 136′″ may be configured to publish its respective generated video stream to participant clients 121 of the general type described above.
[0185] It will be appreciated that the consumer 150′ may also be a participant client 121 that also comprises the central server 130′, for example, by having a laptop computer 402″, 402′″ configured with the functionality of the central server 130′ (and possibly also the central server 130″) to provide a corresponding human user 122″, 122′″ with an enhanced real-time output video stream on the screen of said laptop computer 402″, 402′″ as part of a video communication service in which the human user 122″, 122′″ participates.
[0186] Additionally, FIG. 9 illustrates three external information sources 300′, 300″, 300′″, each corresponding to the external information source 300 described above and providing information to a respective one of the collection functions 131′, 131″, 131′″. There may be fewer or more than three such external information sources 300′, 300″, 300′″. For example, one such external information source 300′, 300″, 300′″ may feed two or more collection functions 131′, 131″, 131′″, and each collection function 131′, 131″, 131′″ may be fed by two or more external information sources 300′, 300″, 300′″.
[0187] 9 does not show the video communication service 110 for reasons of simplicity, it will be understood that video communication services of the general type described above may be used with the central server 130′, 130″, 130′″, such as using the central server 130′, 130″, 130′″ to provide a shared video communication service to participant clients 121 in the manner described above. In some embodiments, the central server 130′″ constitutes, comprises, or is included in the video communication service 110.
[0188] 10 illustrates three different exemplary hardware appliances: a conference room hardware appliance 402′ that in turn includes a digital camera 401′ configured to capture images showing one or more human conference participants 122″, 122′″, 122′″ at a conference location, room, or venue; and two laptop computers 402″, 402′″ that include respective digital webcams 401″, 401′″ configured to use the laptop computers 402″, 402′″ to capture images showing each of the human conference participants 122″, 122′″. It should be understood that FIG. 10 illustrates one of many different configurations for purposes of illustrating the principles of the present invention, and that other types of configurations are possible. For example, only a portion of participants 122", 122'", 122'" may be visible to camera 401'; additional participant users (not shown in FIG. 10) may remotely join the video communication service; external information sources may be used; etc., as exemplified herein. As used herein, the term "remote" means not "local." Two entities, one located "remotely" with respect to the other, are preferably configured to communicate over the open Internet (WAN).
[0189] An exemplary one of the participant users 122'''' is invisible to cameras 401'', 401''', but is visible only to camera 401'.
[0190] Each of the hardware devices 402′, 402″, 402′″ may correspond to the device 402 shown in more detail in FIG. 9 , and each of the hardware devices 402′, 402″, 402′″ may be configured to communicate with the video communication service 110 via the Internet 10 or another digital communication network. Note that the video communication service 110 may then provide a shared video communication service using the devices 402′, 402″, 402′″ as participant users 121 of the general type described above. In some embodiments, the video communication service 110, such as the central server 130′″, is remote with respect to the central server 130′, 130″.
[0191] Returning to FIG. 8, the method begins at a first step S800.
[0192] In a subsequent collection step S801, one or more real-time first primary digital video streams 210, 301 are continuously collected. In the case shown in Figures 9 and 10, the video streams 210, 301 are streams 401 continuously collected from one of the cameras 401', 401'', 401''' by the collection function 131' (and / or 131'', 131''', in the case of collection of an external information source 300).
[0193] The first primary digital video stream is a "real-time" stream, meaning that it is provided from the capture camera to the collection function 131 without delay and without time-consuming image processing before it reaches the collection function 131. For example, any data processing and / or communication between the capture of an image frame on the camera sensor and the corresponding digital video stream being stored in the collection function 131 may take less than 0.1 seconds, e.g., less than 0.05 seconds.
[0194] As described above, the first primary digital video stream 210, 301 may be continuously captured by a camera 410', 401'', 401''' located locally to the participant client 121 consuming the output video stream, such as the device 402, 402', 402'', 402''' itself.
[0195] Furthermore, the first primary digital video stream 210, 301 may be continuously captured by a camera 401', 401'', 401''' positioned to capture images showing the participant user 122'', 122'''' of the device 402'', 402''' (participant client 121).
[0196] In these and other cases, the first primary digital video stream 210, 301 may be continuously captured by a camera 401', 401'', 401''' located physically locally (as "locally" defined above) relative to the computing device 402', 402'', 402''' that performs the application of the first generation control parameters (see below).
[0197] In a subsequent generating step S804, a first digital image analysis is performed on at least one of the collected first primary digital video stream or streams 210, 301. This digital image analysis results in the detection of at least one first event 211 or pattern 212 of the general type described above in the first primary digital video stream 210, 301. In particular, this first digital image analysis results in the subsequent step S805 of establishing first generation control parameters based on the detection of the first event 211 or pattern 212.
[0198] The digital image analysis itself can be carried out in any suitable manner, as is well known per se in the art.
[0199] The detection of the event 211 or pattern 212 may have several different purposes. In general, it may be desirable to affect the video or videos displayed to participant users 122 of the video communication service 110. Such affecting may be configured to dynamically select virtual cropping and / or panning and / or zooming and / or tilting of the captured primary digital video stream to highlight or follow a physical object shown in the primary digital video stream and / or to highlight a current speaker or physical object shown in the primary digital video stream that is currently being discussed. Such affecting may also include adding additional information, such as metadata or externally provided information, to the primary digital video stream based on the content currently being discussed by participant users 122 shown in the primary digital video stream, as heard in an audio capture from the conference venue and determined using digital audio processing steps including natural language interpretation.
[0200] The first generation control parameters may thus comprise the visual position of, or visual tracking information relating to, a stationary or moving object or participant user 122 in the first primary digital video stream 210, 301, where the position or tracking information is automatically detected using digital image processing. Thus, in this case, the detected event 211 or pattern 212 is the position or movement of such object or user in the primary digital video stream 210, 301. Such visual tracking information may comprise information regarding virtual cropping, panning, or zooming to be applied to the primary digital video stream 210, 301 to achieve an image showing the object or participant user 122 as a subpart of that primary digital video stream 210, 301. Thus, the first generation control parameters typically do not comprise information regarding the physical movement or the camera capturing that primary digital video stream 210, 301, but rather instructions regarding how to modify the first primary digital video stream 210, 301 to show the object or participant user 122. This is generally true in the sense that the production control parameters described herein preferably do not include instructions regarding the physical movement of hardware equipment belonging to system 100, but rather only include instructions regarding the digital post-capture processing of images captured by one or more cameras 401', 401'', 401'''.
[0201] The first generation control parameters may further include discrete generation commands that are automatically generated based on the detection of a predetermined event 211 or pattern 212 and / or based on a predetermined or variable generation schedule. For example, the event 211 may be the automatic detection of a particular predetermined object of interest in the primary digital video stream 210, 301, and the discrete generation command may then be to display a brief instructional video about that predetermined object in the output digital video stream. The "discrete" nature of a generation command means that the generation command is applied only once, at a discrete point in time, and not over time. For example, the generation command may be to launch such an instructional video.
[0202] The first generation control parameters may further include virtual cropping, panning, tilting, and / or zooming instructions, e.g., parameter data describing such cropping / panning / tilting / zooming to follow or emphasize a participant user 122 or object of the type described above. The virtual cropping, panning, tilting, and / or zooming instructions may be static or may change dynamically along the time axis of the primary digital video stream, such as between individual frames.
[0203] The first generation control parameters may further include camera stabilization information that is automatically generated based on the detection of movement of the camera 401′, 401″, 401′″ providing the primary digital video stream. Thus, in this case, the digital video analysis aims to dynamically detect, in a manner that may be conventional per se, an event 211 or pattern 212 in the form of shaking or other movement of the camera 401′, 401″, 401′″, such as due to the camera 401′, 401″, 401′″ being manually operated by one of the participant users 122″, 122′″, 122′″. The first generation control parameters may be pan / rotate / tilt / zoom commands for the primary digital video stream 210, 301 that, when applied, aim to at least partially offset the detected movement and / or stabilize the primary digital video stream 210, 301 over time. The movement of the cameras 401', 401'', 401''' can itself be detected entirely based on image processing of the primary digital video stream using conventional image processing techniques, for example pixel correlation techniques to detect image transformations between frames.
[0204] In all of these examples, the first digital image analysis takes a certain amount of time to perform due to the calculations involved. As a result, the first generation control parameter is established with a first time delay relative to the time of occurrence of the first event 211 or pattern 212 in the first primary digital video stream 210, 301. This first time delay may be large enough to cause a noticeable audio delay if the primary digital video stream with the first time delay is not time-synchronized with the corresponding captured digital sound stream, and / or the first time delay may be small enough not to cause interaction difficulties for participant users 122'', 122''', 122'''' interacting with each other using the digital video communication service 110. Specifically, the first time delay may be 0.1 seconds or greater. In this and other embodiments, the first time delay may be less than 1 second, e.g., less than 0.5 seconds, e.g., less than 0.3 seconds.
[0205] In a subsequent generating step S806, the first generation control parameters are applied to the real-time first primary digital video stream 210, 301 as part of generating the digital output video stream. As a result of applying the first generation control parameters, the first primary digital video stream 210, 301 is modified based on the first generation control parameters to generate a first generated digital video stream. However, this modification occurs without the first digital video stream itself being delayed by the first time delay.
[0206] Thus, even though the first generation control parameter is determined only after the first time delay (because it takes some time to establish the first generation control parameter), it is the non-delayed first primary digital video stream 210, 301 that is affected by the application of the first generation control parameter. In other words, the first generation control parameter is applied to the first primary digital video stream 210, 301 at a point along the timeline of the first primary digital video stream 210, 301 that is at least the first time delay later than the point along the timeline at which the event 211 or pattern 212 was detected. Thus, for example, if the event 211 or pattern 212 is the detected movement of an object in frame x of the first primary digital video stream 210, 301, the first generation control parameter may translate the virtual cropping of the first primary digital video stream 210, 301 to follow the object to its new position. However, the first generation control parameter is established only at a first time delay, such as frame y (e.g., y=10) of the first primary digital video stream 210, 301. Thus, when the first generation control parameter is applied to effect the above-mentioned shift in cropping, this shift in cropping of the first primary digital video signal 210, 301 is relative to frame x+1y of the first primary digital video signal 210, 301.
[0207] In some embodiments, the first digital image analysis is performed by a computing device 402 that is also configured to generate and publish the output digital video stream 230. In particular, it may be a generation function 135'' of the central server 130'' that is part of the same physical computing device 402 that also constitutes the central server 130' that performs the first digital image analysis and establishes the first production control parameters. In that case, it may be the central server 130' (via generation function 135') that applies the first production control parameters.
[0208] This minimizes the first time delay by eliminating the need for communication with an external or peripheral computer device to establish and apply the first generation control parameters.
[0209] This also enables any camera-enabled hardware device 402′ already installed in the conference room or venue, and / or any laptop 402″, 402′″, or similar used by any conference participant 122, to be used to create an enhanced shared digital video conferencing service that can be accessed and used by other conference participants 122 as part of the same interactive digital video communication session in the context of the first primary digital video stream 210, 301 originating.
[0210] In other embodiments, multiple such hardware apparatuses 402′, 402″, 402′′ can each be a device 402 of the type illustrated in FIG. 9 , each simultaneously generating and publishing a respective enhanced output digital video stream 230 for consumption by other participant users 121 or for use in combination by a common central video communication service 110 to further refine a common digital video communication service experience. Because such output digital video streams 230 are provided in real time, without a first time delay, the central video communication service 110's generation can result in a relatively low-latency video communication, even if the central video communication service 110 intentionally adds delay as described above for production purposes.
[0211] For example, in a classroom, a fixed computing device 402′ with a wide-angle webcam 401′, which may be the type of computing device 402 shown in FIG. 9, captures images showing all of the students in the classroom, or at least some of the students. At the same time, one or more of the students may operate their own computing device 402 in the form of a webcam-enabled laptop 402″, 402′″ to capture images showing only that student. All devices 402′, 402″, 402′′ can then generate and publish their own enhanced real-time output digital video streams 230 that can be viewed on their respective devices 402′, 402″, 402′′, thus forming a first group of participants 121 associated with extremely low latency and / or a second group of participants 121 associated with greater latency that can also be used by the central video communications service 110 to generate a more elaborate common video communications experience with little latency that can be consumed by external participants or viewers.
[0212] In such cases, it should be noted that all of the collection, analysis, and generation can be configured to occur in a fully automatic manner based on parameter inputs to the system 100, thereby resulting in generation steps that are automatically and dynamically applied.
[0213] The analysis of the first digital image and the publishing of the output digital video stream 230 may be performed on the same physical computer device 402, but at least the analysis of the first digital image and the application of the first production control parameters to produce the output digital video stream 230 use computer software operating at least in part in separate processes or threads.
[0214] To provide a high-quality, low-latency output digital video stream 230, some embodiments may perform processor throttling for the first digital image analysis as a function of the current processor load of the computing device 402 performing the first digital image analysis. This is particularly true when the central servers 130′, 130″ are executed on one and the same central processor unit. Such processor throttling may be performed such that the provision and publishing of the output digital video stream 230 (central server 130′) is given processor priority over the first digital image analysis (central server 130″). In other words, under limited CPU conditions, CPU priority is given to the generation function 135′ and the publishing function 136′ compared to the generation function 135″. For example, processor throttling of the first digital image analysis may be performed by limiting the first digital image analysis to only a portion of all video frames of the first primary digital video stream 210, 301, such as performing the first digital image analysis only on every other or every third frame, or performing the first digital image analysis only on one frame per given unit of time, such as only one frame per 0.1 seconds, or less frequently. In many cases, this may provide sufficiently accurate event 211 or pattern 212 detection while maintaining high video quality in the output digital video stream 230. Throttling may be used all the time, or may be switched on or off as needed. In the latter case, throttling may be switched on as a result of detecting that available CPU capacity is too low to provide and publish the output digital video stream 230 at a desired minimum video quality.
[0215] In a subsequent publishing step S807, the output digital video stream 230 is subsequently provided (published) by the publishing function 136′ to at least one participant client 121. It should be noted that the participant client 121 may be the same computing device 402 that performs the generation of the output digital video stream 230 and may provide the enhanced, low-latency video stream for viewing to a user 122 of the computing device 402. The participant client 121 may also be another computing device 402′, 402″, 402′″, or may be a video communication service 110 that uses the generated and published output digital video stream 230 to generate a secondary output digital video stream, as described above.
[0216] The output digital video stream 230 is provided and published in the form of or based on the first generated digital video stream generated by the generation function 135'. As noted above, the output digital video stream 230 may be generated and made available in multiple layers or stages. For example, the output digital video stream 230 may be generated based on both the first primary digital video stream 210, 301 and the first generated digital video stream by the same or different hardware device 402 that generates the first generated digital video stream.
[0217] In the following step S808, the method ends, although an iterative process typically occurs, as shown in FIG.
[0218] The generation performed by the generating step 135' simply applies the generation control parameters established by the generating step 135'' to the first primary digital video stream 210, 301 and outputs the resulting first generated digital video stream continuously and in real time, and should be as fast as possible. As explained above, any information regarding how to crop or otherwise adjust the primary digital video stream 210, 301, including any add-on (e.g., in the form of externally provided video material or metadata), is received from the generating function 135'', resulting in the above-mentioned time discrepancy between the application of the first generation control parameters and the content of the first primary video stream 210, 301 based on which the values of the first generation control parameters were originally established.
[0219] The first resulting digital video stream may then be provided as a real-time enhanced primary video stream input to any acquisition function 131, 131', 131'', 131'''.
[0220] It is understood that application of the first generation control parameters, such as performing a crop or zoom of a video frame in the digital domain, can be very fast and result in minimal time delay, for example, application of the first generation control parameters can result in a delay of the first generated digital video stream (relative to the first primary digital video stream) of at most 0.2 seconds.
[0221] The collection functionality 131′ may then further comprise continuously collecting or capturing a first digital audio stream in addition to or as part of the first primary digital video stream 210, 301, where the first digital audio stream is associated with the first primary digital video stream 210, 301. Once collected / captured, the first digital audio stream may, in effect, be time-synchronized with the first resulting video stream 210, 301 by slightly delaying the first primary digital audio stream, and the time-synchronized first digital audio stream may then be provided to at least one participant client 121 or collection functionality 131 along with or as part of the first resulting digital video stream and / or output digital video stream 230.
[0222] In this way, both video and corresponding audio can be input into the collection function 131' and an expanded bundle of synchronized video and corresponding audio can be output by the publishing function 136' after being expanded, but with minimal time delay. In this way, the central server 130', assisted by the central server 130'', can be thought of as a "virtual video cable", operating with only minimal time delay, but capable of outputting at the distal end an expanded version of the input video / audio information input at the proximal end.
[0223] The generation function 135'' performs video analysis to generate the augmentation, resulting in the establishment of first generation control parameters as described above. This may include detecting specific people in the image, selecting cropping or virtual zooming to track and / or focus on specific moving people in the image, identifying offsetting virtual camera movements to stabilize the camera to achieve soft camera panning / movement, etc. This analysis results in a number of decisions / commands for the generation function 135', resulting in a small time delay before such decisions / commands are applied. This small time delay has generally been found to not adversely affect the experience of the user 122 consuming the output digital video stream 230, even in the case of camera tracking. Even when a human camera operator performs camera tracking on a moving person or object, the operator has a certain minimum reaction time. The generation function 135'' may also be configured to detect events 211 or patterns 212 while ignoring small movements of the person or object being tracked, and to perform virtual camera panning or movement only in response to larger movements. As described above, multi-threaded software implementation on the same physical computer hardware allows this to be achieved using a single hardware appliance 402 across a wide array of possible hardware platforms in a way that the complexity of the production does not affect, for example, the image or sound quality of the first produced digital video stream.
[0224] It is particularly noted that the first generation control parameters are configured solely to control the digital conversion of the first primary digital video stream 210, 301. It may also be possible for the first generation function 135' to provide instructions to the camera to perform physical panning, zooming or the like, although this would be outside the scope of the present invention.
[0225] 8, the method may further comprise a step S802 of performing a second digital image analysis of the first primary digital video stream 210, 301. In some embodiments, the method comprises a step S802 of performing a second digital audio analysis of a continuously captured digital audio stream associated with the first primary digital video stream 210, 301.
[0226] A second image / audio analysis is performed to identify at least one second event 211 or pattern 212 in the first primary digital video stream 210, 301 and / or said digital audio stream and in a subsequent step S803 to establish second generation control parameters, which may be generally similar to the first generation control parameters described above.
[0227] In a similar manner to the first image analysis, the second digital image / audio analysis takes a certain amount of time to perform and establishes the second generation control parameters after a second time delay relative to the time of occurrence of the second event 211 or pattern 212 in the first primary digital video stream 210, 301. However, the second time delay is longer than the first time delay.
[0228] Next, the second generation control parameters are applied to the real-time first primary digital video stream 210, 301, such that the first primary digital video stream 210, 301 is modified based on the second generation control parameters without being delayed by the second time delay to generate the first generated digital video stream.
[0229] In practice, the second generation control parameter may be applied directly by the generation functionality 135′ to the first primary digital video stream 210, 301 in a manner corresponding to the application of the first generation control parameter. In other embodiments, the second generation control parameter may be established with greater latency than the first generation control parameter and thus only indirectly applied to the first primary digital video stream 210, 301, e.g., by the second generation control parameter influencing the value of the first generation control parameter before the first generation control parameter is applied to the first primary digital video stream 210, 301. For example, the second generation control parameter may relate to a broader generation aspect, such as the currently displayed camera angle or presentation slide selection, while the first generation control parameter may relate to a more detailed generation aspect, such as the virtual panning or cropping that is strictly currently being applied.
[0230] The second digital image analysis may be performed by a computing device that is remote from the computing device that performs the application of the first production control parameter and / or a computing device that is remote from the computing device that performs the application of the second production control parameter. In the example shown in Figure 9, computing device 402 applies the first production control parameter and possibly the second production control parameter, and computing device 402 is remote from the computing device on which central server 130''' executes.
[0231] In some embodiments, the second generation control parameters constitute input to the first digital image analysis. For example, the second generation control parameters may comprise information identifying a particular person or object viewed in the first primary digital video stream 210, 301, and this identification information may be used as input to the first digital image analysis to cause tracking (by virtual panning and / or zooming) of the identified person or object as it moves through the image frames making up the first primary digital video stream 210, 301.
[0232] Thus, the second generation control parameter may include instructions as to whether to display in the first generated digital video stream a particular participant user that appears in an unaffected captured image, the participant user being automatically identified based on digital image processing.
[0233] In this and other embodiments, the second production control parameters may comprise a second primary video stream 210, 301, for example, the second primary video stream 210, 301 is embedded in the first produced digital video stream.
[0234] The second time delay being greater than the first time delay may be due to the second image processing and / or audio processing taking more time than the first image processing and / or due to communication taking time between the central server 130', 130'' on the one hand and the central server 130'''' on the other hand.
[0235] In general, the second digital image processing, which results in the establishment of second processing control parameters, may constitute a load-relief from the first digital image processing, such as when the hardware device 402 is overloaded. Offloading part of the first image analysis to the central server 130''' may be a reasonable trade-off in that it accepts a larger time delay for application of the second generation control parameters, while at the same time obtaining a more complete image analysis.
[0236] However, in general, central server 130''' may be a cloud resource or other external computing resource with greater processing power and / or quicker access to a larger database of external data than central server 130'' (which may be locally located as described above). Thus, the second image analysis and / or audio analysis preferably comprises more sophisticated and therefore more processing-intensive tasks than the first image analysis.
[0237] For example, the second image processing may comprise advanced facial recognition based on a database of potential persons to be detected and their facial features. In contrast, the first image processing may consist of an algorithm that finds and tracks already identified faces of such identified persons in the first primary digital video stream 210, 301.
[0238] The second image or audio processing may include correlating automatically interpreted information (displayed or discussed) contained in the analyzed images and / or audio with external data. For example, if the automatic audio processing is comprised of natural language detection, parsing, and interpretation components, it may conclude that the speaker heard in the audio is talking about a particular type of flower. In this case, the second generation control parameters may include information for incorporating a particular image showing a flower of that type of flower into the first generated digital video stream. Correspondingly, if the second image processing determines that the first primary digital video stream displays a logo of a particular rock band, the second generation control parameters may include instructions for displaying an image of the rock band in the first generated digital video stream.
[0239] It should be noted that the establishment of the first and second generation control parameters may be performed based on an available space of possible such generation control parameters and associated values, which may in some cases be defined by configuration parameters definable by a user of the video communication service 110. For example, in a video communication service 110 used to deliver a lecture on geography, the generation function 135''' may be instructed via such configuration parameters to monitor countries that have been or will be mentioned (e.g., in text on a displayed whiteboard) and automatically generate a second generation control parameter that specifies showing the corresponding flag in a particular predetermined substream of the first generated digital video stream.
[0240] In some embodiments, only the digital audio stream is sent to the generating function 135''' of the central server 130''', rather than the first primary digital video stream 210, 301. This significantly reduces the amount of data that needs to be communicated to the central server 130''' (which, as noted above, may be remotely located).
[0241] In a specific example, second generation control parameters are subsequently established to instruct generation function 135' to select and virtually pan / crop / zoom one or more primary digital video streams based on a second digital image process that performs face recognition, such that the output digital video stream shows the teacher but none of the students. In this way, an output digital video stream is produced that does not risk showing the faces of the students. As explained above, the virtual pan / crop / zoom operations may actually be performed by the first image process via the first generation control parameters using input in the form of the second generation control parameters.
[0242] In this case, the associated digital audio stream associated with the first generated digital video stream may also be affected by the first and / or second generation control parameters. This may include selecting from among a set of multiple available primary digital audio streams, such as those captured by different microphones of participant users of communication service 110, or may include digitally suppressing or enhancing certain sounds. In the teacher-student example, questions posed by the students may be muted in the first generated digital video stream, while answers from the teacher may be unaffected.
[0243] The invention also relates to a computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out a method for providing an output digital video stream 230 according to any of the preceding claims.
[0244] The computer program product may be embodied by a non-transitory computer-readable medium encoding instructions that cause one or more hardware processors located in at least one of the computer hardware devices in the system to perform the method steps described herein.
[0245] The present invention also relates to the system 100, comprising: one or more collection functions 131, 131', 131'', 131''' configured to continuously collect a real-time first primary digital video stream 210, 301; one or more production functions 135, 135', 135'', 135''' configured to perform first digital image analysis and establish first production control parameters as described herein; and one or more publishing functions 136, 136', 136'', 136''' configured to continuously provide an output digital video stream 230 as described herein.
[0246] The system 100 may further include multiple cameras, each configured to capture a respective non-delayed primary digital video stream 210, 301. In this case, the generation functionality 135′ may be configured to generate a non-delayed first generated digital video stream based on each of the captured primary digital video streams 210, 301.
[0247] Although preferred embodiments have been described above, it will be apparent to those skilled in the art that many modifications can be made to the disclosed embodiments without departing from the essential concepts of the invention.
[0248] For example, many additional features not described herein may be provided as part of the system 100 described herein. In general, the solutions described herein provide a framework upon which detailed functionality and features can be built to accommodate a wide variety of specific applications in which streams of video data are used for communication.
[0249] As an example, the output digital video stream may form the input primary digital video stream for another automatic digital video generation method or system of this disclosure or of a different type.
[0250] While the first and second events / patterns are specifically illustrated above in connection with Figures 8-10, it should be understood that many other types of events and patterns are possible. Possible event types and patterns are described throughout this specification, and those skilled in the art will understand that these examples and discussions are for illustrative purposes only and are not intended to constitute an exclusive list.
[0251] In general, anything said about the present method is also applicable to the present system and computer program product, and vice versa.
[0252] Therefore, the invention is not limited to the described embodiments, but can be modified within the scope of the appended claims.
Claims
1. 1. A method for providing an output digital video stream (230), the method comprising: continuously collecting a real-time first primary digital video stream (210, 301); performing a first digital image analysis of the first primary digital video stream (210, 301) to identify at least one first event (211) or pattern (212) within the first primary digital video stream (210, 301), wherein the first digital image analysis results in establishing first generation control parameters based on the detection of the first event (211) or pattern (212), and the first digital image analysis takes a certain amount of time to perform such that the first generation control parameters are established after a first time delay related to the time of occurrence of the first event (211) or pattern (212) within the first primary digital video stream (210, 301); applying the first generation control parameters to the real-time first primary digital video stream (210, 301), wherein application of the first generation control parameters results in the first primary digital video stream (210, 301) being modified based on the first generation control parameters to generate a first generated digital video stream without being delayed by the first time delay; and The output digital video stream (230) is continuously provided to at least one participant client (121), wherein the output digital video stream (230) is provided in the form of or based on the first generated digital video stream.
2. 10. The method of claim 1, further comprising: The output digital video stream (230) is generated based on both the first primary digital video stream (210, 301) and the first generated digital video stream.
3. 3. The method of claim 2, said collecting further comprising continuously capturing a first digital audio stream, wherein said first digital audio stream is captured in association with said first primary digital video stream (210, 301); The method further comprises: time-synchronizing the first digital audio stream with the first resulting video stream; and The time-synchronized first digital audio stream is provided to the at least one participant client (121) together with or as part of the output digital video stream (230).
4. 4. The method according to any one of claims 1 to 3, The method, wherein the first primary digital video stream (210, 301) is continuously captured by a camera located locally to the participant client (121).
5. 5. The method of claim 4, The method, wherein the first primary digital video stream (210, 301) is continuously captured by a camera to display a participant user of the participant client (121) within the first primary digital video stream (210, 301).
6. 6. The method of claim 4 or 5, The method, wherein the first primary digital video stream (210, 301) is continuously captured by a camera located locally relative to a computing device that performs the application of the first production control parameters.
7. 7. The method of any one of claims 1 to 6, The first production control parameters comprise one or more of the following: a) location or tracking information of stationary or moving objects or people in said first primary digital video stream (210, 301), wherein said location or tracking information is automatically detected using digital image processing; b) Discrete generation commands that are automatically generated based on the detection of predetermined events (211) or patterns (212) and / or that are automatically generated based on a predetermined or variable generation schedule; c) virtual panning and / or zooming commands; and d) Camera stabilization information that is automatically generated based on camera motion detection.
8. 8. The method according to any one of claims 1 to 7, The method, wherein the first digital image analysis is performed by a computing device that is also configured to provide the output digital video stream (230) to the at least one participant client.
9. 9. The method of claim 8, The method, wherein the analyzing the first digital image and providing the output digital video stream (230) are performed in separate processes or threads.
10. 10. The method of claim 8 or 9, further comprising: The first digital image analysis is processor-throttled as a function of a current processor load of a computing device performing the first digital image analysis, such that providing the output digital video stream (230) has processor priority over the first digital image analysis.
11. 11. The method of claim 10, 10. A method according to claim 9, wherein the processor throttling of the first digital image analysis is performed by limiting the first digital image analysis to only a portion of all video frames of the first primary digital video stream (210, 301).
12. 12. The method of any one of claims 1 to 11, further comprising: performing a second digital image analysis of the first primary digital video stream (210, 301) and / or a second digital audio analysis of a digital audio stream captured consecutively with and associated with the first primary digital video stream (210, 301) to identify at least one second event (211) or pattern (212) in the first primary digital video stream (210, 301) and / or the digital audio stream, wherein the second digital image or audio analysis takes a certain period of time to perform so as to establish second generation control parameters after a second time delay related to the time of occurrence of the second event (211) or pattern (212) in the first primary digital video stream (210, 301), the second time delay being longer than the first time delay; and applying the second generation control parameters to the real-time first primary digital video stream (210, 301), wherein as a result of applying the second generation control parameters, the first primary digital video stream (210, 301) is modified based on the second generation control parameters to generate the first generated digital video stream without being delayed by the second time delay.
13. 13. The method of claim 12, The method, wherein the second digital image analysis is performed by a computing device that is remote from a computing device that performs the applying of the second production control parameters.
14. 14. The method of claim 12 or 13, The method, wherein the second generation control parameters constitute input to the first digital image analysis.
15. 15. The method of any one of claims 12 to 14, The second production control parameters comprise one or more of the following: a) a second primary video stream (210, 301); and b) instructions as to whether to display a particular participant user within said first generated digital video stream, wherein said participant user is automatically identified based on digital image processing;
16. A computer program product comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 15 for providing an output digital video stream (230).
17. A system (100) for providing an output digital video stream (230), the system (100) comprising: an acquisition function (131, 131', 131'', 131''') configured to continuously acquire a real-time first primary digital video stream (210, 301); a generation function (135, 135', 135'', 135''') configured to perform a first digital image analysis of the first primary digital video stream (210, 301) to identify at least one first event (211) or pattern (212) within the first primary digital video stream (210, 301), wherein the first digital image analysis results in establishing first generation control parameters based on detection of the first event (211) or pattern (212), and the first digital image analysis identifies the first event (211) or pattern (212) within the first primary digital video stream (210, 301). the generating function (135, 135', 135'', 135''') is further configured to apply the first generating control parameters to the real-time first primary digital video stream (210, 301), wherein application of the first generating control parameters results in the first primary digital video stream (210, 301) being modified based on the first generating control parameters to generate a first generated digital video stream without being delayed by the first time delay; and a publishing function (136, 136', 136'', 136''') configured to continuously provide the output digital video stream (230) to at least one participant client (121), wherein the output digital video stream (230) is provided in the form of or based on the first generated digital video stream.
18. 18. A system (100) according to claim 17, comprising: The system (100) further comprises a plurality of cameras, each configured to capture a respective non-delayed primary digital video stream (210, 301); and The generating functionality (135, 135', 135'', 135''') is configured to generate the non-delayed first generated digital video stream based on each of the captured primary digital video streams (210, 301).