System and method for processing speech audio stream interruptions

CN116830559BActive Publication Date: 2026-05-29QUALCOMM INC

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
QUALCOMM INC
Filing Date
2021-12-09
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

During online meetings, network issues can cause audio frames to be lost, leading users to miss important information and have to guess or request repetition, resulting in a poor user experience.

Method used

The system processes audio streams by receiving them on the device and converting them into text streams, generating text or synthesizing audio streams to replace interrupted audio streams, displaying voice emotion and intonation using metadata streams, and providing text or virtual avatars on the display to indicate the model training status.

Benefits of technology

It effectively restored the information flow in online meetings, reduced the need for users to guess, improved the user experience, and solved the problem of interrupted audio streams by combining text and synthesized speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116830559B_ABST
    Figure CN116830559B_ABST
Patent Text Reader

Abstract

An apparatus for communication includes one or more processors configured to receive, during an online meeting, a speech audio stream representing speech of a first user. The one or more processors are also configured to receive a text stream representing the speech of the first user. The one or more processors are also configured to selectively generate output based on the text stream in response to an interruption in the speech audio stream.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority requirements

[0002] This application claims priority to jointly owned U.S. nonprovisional patent application No. 17 / 166,250, filed February 3, 2021, the contents of which are expressly incorporated herein by reference in their entirety. Technical Field

[0003] This disclosure generally relates to systems and methods for handling interruptions in voice audio streams.

[0004] Description of related technologies

[0005] Technological advancements have led to smaller and more powerful computing devices. For example, a variety of portable personal computing devices exist today, including cordless phones such as mobile and smartphone phones, and small, lightweight tablets and laptops that are easy for users to carry. These devices can transmit voice and data packets over wireless networks. Furthermore, many of these devices incorporate additional features such as digital still cameras, digital video cameras, digital recorders, and audio file players. Moreover, such devices can process executable instructions, including software applications that can be used to access the Internet, such as web browser applications. Thus, these devices can include significant computing power.

[0006] Such computing devices typically incorporate the ability to receive audio signals from one or more microphones. For example, the audio signal could represent user speech captured by a microphone, external sounds captured by a microphone, or a combination thereof. Such devices could include communication equipment used for online meetings or calls. Network problems during an online meeting between a first user and a second user can cause frame loss, resulting in some audio and video frames sent by the first user's first device not being received by the second user's second device. Frame loss due to network problems can lead to irrecoverable information loss during an online meeting. For example, the second user might have to guess what they missed or ask the first user to repeat the missed content, resulting in a poor user experience. Summary of the Invention

[0007] According to one implementation of this disclosure, a communication device includes one or more processors configured to receive an audio stream representing the voice of a first user during an online meeting. The one or more processors are also configured to receive a text stream representing the voice of the first user. The one or more processors are further configured to selectively generate output based on the text stream in response to an interruption in the audio stream.

[0008] According to another implementation of this disclosure, a communication method includes receiving, at a device, a voice audio stream representing the voice of a first user during an online meeting. The method also includes receiving, at the device, a text stream representing the voice of the first user. Furthermore, the method includes selectively generating output at the device based on the text stream in response to an interruption in the voice audio stream.

[0009] According to another implementation of this disclosure, a non-transitory computer-readable medium includes instructions that, when executed by one or more processors, cause the one or more processors to receive an audio stream representing the speech of a first user during an online meeting. When executed by the one or more processors, the instructions also cause the one or more processors to receive a text stream representing the speech of the first user. When executed by the one or more processors, the instructions also cause the one or more processors to selectively generate output based on the text stream in response to an interruption in the audio stream.

[0010] According to another implementation of this disclosure, an apparatus includes components for receiving an audio stream representing the speech of a first user during an online meeting. The apparatus also includes components for receiving a text stream representing the speech of the first user. Furthermore, the apparatus includes components for selectively generating output based on the text stream in response to an interruption in the audio stream.

[0011] Other aspects, advantages, and features of this disclosure will become apparent upon reading the entire application, including the following sections: description of the drawings, detailed description, and claims. Attached Figure Description

[0012] Figure 1 This is a block diagram illustrating specific aspects of a system operable to handle interruptions in a speech audio stream, based on some examples of this disclosure.

[0013] Figure 2 This is an illustrative diagram of an operable system for handling interruptions in a speech audio stream, based on some examples of this disclosure.

[0014] Figure 3A Based on some examples of this disclosure Figure 1 The system or Figure 2 An illustration of the system-generated descriptive graphical user interface (GUI).

[0015] Figure 3B Based on some examples of this disclosure Figure 1 The system or Figure 2 A diagram illustrating the descriptive GUI generated by the system.

[0016] Figure 3C Based on some examples of this disclosure Figure 1The system or Figure 2 A diagram illustrating the descriptive GUI generated by the system.

[0017] Figure 4A Based on some examples of this disclosure Figure 1 The system or Figure 2 The system is illustrated with diagrams illustrating its operation.

[0018] Figure 4B Based on some examples of this disclosure Figure 1 The system or Figure 2 The system is illustrated with diagrams illustrating its operation.

[0019] Figure 5 This is an illustrative diagram of an operable system for handling interruptions in a speech audio stream, based on some examples of this disclosure.

[0020] Figure 6A Based on some examples of this disclosure Figure 5 An illustration of the system-generated descriptive graphical user interface (GUI).

[0021] Figure 6B Based on some examples of this disclosure Figure 5 A diagram illustrating the descriptive GUI generated by the system.

[0022] Figure 6C Based on some examples of this disclosure Figure 5 A diagram illustrating the descriptive GUI generated by the system.

[0023] Figure 7A Based on some examples of this disclosure Figure 5 The system is illustrated with diagrams illustrating its operation.

[0024] Figure 7B Based on some examples of this disclosure Figure 5 The system is illustrated with diagrams illustrating its operation.

[0025] Figure 8 Based on some examples of this disclosure, it is possible to... Figure 1 , Figure 2 or Figure 5 A diagram illustrating a specific implementation of a method for handling interruptions to the voice audio stream, executed by any entity in the system.

[0026] Figure 9 Examples of integrated circuits operable to handle interruptions in a voice audio stream according to some examples of this disclosure are illustrated.

[0027] Figure 10 This is an illustration of a mobile device capable of handling interruptions to a voice audio stream according to some examples of this disclosure.

[0028] Figure 11 This is an illustration of a headset that can handle interruptions in a speech audio stream according to some examples of this disclosure.

[0029] Figure 12 This is an illustration of a wearable electronic device that can be used to handle interruptions in a voice audio stream, based on some examples of this disclosure.

[0030] Figure 13 This is a diagram illustrating a voice-controlled speaker system operable to handle interruptions in a speech audio stream, based on some examples of this disclosure.

[0031] Figure 14 This is an illustration of a camera that can handle interruptions in a speech audio stream according to some examples of this disclosure.

[0032] Figure 15 This is an illustration of a headset (such as a virtual reality or augmented reality headset) that can handle interruptions in the voice audio stream according to some examples of this disclosure.

[0033] Figure 16 This is a diagram illustrating a first example of a vehicle capable of handling interruptions to a voice audio stream according to some examples of this disclosure.

[0034] Figure 17 This is a diagram illustrating a second example of a vehicle capable of handling interruptions to a voice audio stream, based on some examples of the present disclosure.

[0035] Figure 18 This is a block diagram of a particular illustrative example of a device operable to handle interruptions in a voice audio stream according to some examples of this disclosure. Detailed Implementation

[0036] Missing a portion of an online meeting or call can negatively impact the user experience. For example, during an online meeting between a first user and a second user, if some audio frames sent by the first user's first device are not received by the second user's second device, the second user may miss a portion of the first user's voice. The second user must then either guess what the first user is saying or ask the first user to repeat the missed content. This can lead to communication errors, interruptions in the conversation, and wasted time.

[0037] Systems and methods for handling interruptions to voice audio streams are disclosed. For example, each device includes a conference manager configured to establish online conferences or calls between the device and one or more other devices. An interruption manager (at the device or server) is configured to handle interruptions to the voice audio stream.

[0038] During an online meeting between a first user's first device and a second user's second device, the meeting manager on the first device sends a media stream to the second device. This media stream includes an audio stream, a video stream, or both. The audio stream corresponds to the first user's voice during the meeting.

[0039] A stream manager (at the first device or server) generates a text stream by performing a speech-to-text conversion on the audio stream and forwards the text stream to the second device. In a first operating mode (e.g., a conference manager at the first device or server), the stream manager concurrently forwards the text stream with the media stream throughout the online conference. In an alternative example, in a second operating mode (e.g., an interrupted data transmission mode), the stream manager (e.g., an interruption manager at the first device or server) forwards the text stream to the second device in response to detecting network problems (e.g., low bandwidth, packet loss, etc.) in transmitting the media stream to the second device.

[0040] In some examples, network problems cause an interruption in receiving the media stream at the second device, but not in receiving the text stream. In some examples, the second device, in a first operating mode (e.g., display caption data mode), provides the text stream to the display, regardless of the detected network problem. In other examples, the second device, in a second operating mode (e.g., display interrupted data mode), displays the text stream in response to the detection of an interruption in the media stream.

[0041] In certain examples, the stream manager (e.g., a meeting manager or interruption manager) forwards a metadata stream in addition to the text data. The metadata indicates the emotion, tone, and other attributes of the first user's voice. In certain examples, the second device displays a metadata stream in addition to the text stream. For example, the text stream is annotated based on the metadata stream.

[0042] In a particular example, the second device performs a text-to-speech conversion on the text stream to generate a synthesized speech audio stream, and outputs the synthesized speech audio stream (e.g., to replace an interrupted speech audio stream). In a particular example, the text-to-speech conversion is based at least in part on the metadata stream.

[0043] In a specific example, the second device displays a virtual avatar during the output of the synthesized speech audio stream (e.g., to replace an interrupted video stream). In a specific example, the text-to-speech conversion is based on a general speech model. For example, a first general speech model might be used for one user, and a second general speech model for another user, so that listeners can distinguish the speech corresponding to different users. In another specific example, the text-to-speech conversion is based on a user speech model generated from the first user's speech. In a specific example, the user speech model is generated before the online meeting. In a specific example, the user speech model is generated (or updated) during the online meeting. In a specific example, the user speech model is initialized from a general speech model and updated based on the first user's speech.

[0044] In a specific example, the avatar indicates that a speech model is being trained. For instance, the avatar is initialized to red to indicate that a generic speech model is being used (or the user's speech model is not yet ready), and the avatar transitions from red to green over time to indicate that a speech model is being trained. A green avatar indicates that the user's speech model has been trained (or the user's speech model is ready).

[0045] Online meetings can be held between more than two users. If the first device is experiencing network problems but the third user's third device in the online meeting is not experiencing network problems, the second device can output the first user's synthesized audio stream while simultaneously outputting a second media stream received from the third device that corresponds to the third user's voice, video, or both.

[0046] Specific aspects of this disclosure are described below with reference to the accompanying drawings. In the specification, common features are indicated by common reference numerals. As used herein, various terms are used only for the purpose of describing particular implementations and are not intended to limit the implementations. For example, the singular forms “an,” “a,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. Furthermore, some features described herein are singular in some implementations and plural in others. For illustration, Figure 1 It describes a system that includes one or more processors ( Figure 1 The device 104 (of which the “processor” 160) indicates that in some implementations, the device 104 includes a single processor 160, while in other implementations, the device 104 includes multiple processors 160.

[0047] As used herein, the terms “comprise,” “comprises,” and “comprising” are used interchangeably with “include,” “includes,” or “including,” and the term “wherein” is used interchangeably with “where.” As used herein, “exemplary” indicates an example, implementation, and / or aspect, and should not be construed as limiting or indicating a preference or preferred implementation. As used herein, ordinal terms used to modify elements such as structures, components, operations, etc. (e.g., “first,” “second,” “third,” etc.) do not themselves indicate any priority or order of that element relative to another element, but merely distinguish that element from another element with the same name (but using ordinal terms). As used herein, the term “set” refers to one or more specific elements, and the term “multiple” refers to multiple (e.g., two or more) specific elements.

[0048] As used herein, “coupling” can include “communicationally coupled,” “electrically coupled,” or “physically coupled,” and may (or alternatively) include any combination thereof. Two devices (or components) may be coupled directly or indirectly (e.g., communicationally coupled, electrically coupled, or physically coupled) via one or more other devices, components, wires, buses, networks (e.g., wired networks, wireless networks, or combinations thereof). As an illustrative, non-limiting example, two electrically coupled devices (or components) may be included in the same device or different devices and may be connected via electronics, one or more connectors, or inductive coupling. In some implementations, two communicationally coupled (e.g., electrically coupled) devices (or components) may directly or indirectly send and receive signals (e.g., digital or analog signals) via one or more wires, buses, networks, etc. As used herein, “direct coupling” can include two devices coupled without intermediate components (e.g., communicationally coupled, electrically coupled, or physically coupled).

[0049] In this disclosure, terms such as “determine,” “calculate,” “estimate,” “shift,” and “adjust” are used to describe how one or more operations are performed. It should be noted that these terms should not be construed as restrictive, and similar operations can be performed using other techniques. Furthermore, as mentioned herein, “generate,” “calculate,” “estimate,” “use,” “select,” “access,” and “determine” are used interchangeably. For example, “generate,” “calculate,” “estimate,” or “determine” a parameter (or signal) can refer to actively generating, estimating, calculating, or determining a parameter (or signal), or it can refer to using, selecting, or accessing a parameter (or signal) that has already been generated, such as by another component or device.

[0050] refer to Figure 1 This document discloses specific illustrative aspects of a system configured to handle interruptions in a voice audio stream, and generally designates it as 100. System 100 includes a device 102 coupled to device 104 via a network 106. Network 106 includes a wired network, a wireless network, or both. Device 102 is coupled to a camera 150, a microphone 152, or both. Device 104 is coupled to a speaker 154, a display device 156, or both.

[0051] Device 104 includes one or more processors 160 coupled to memory 132. The one or more processors 160 include a conference manager 162 coupled to an interrupt manager 164. Conference manager 162 and interrupt manager 164 are coupled to a graphical user interface (GUI) generator 168. Interrupt manager 164 includes a text-to-speech converter 166. Device 102 includes one or more processors 120, which includes a conference manager 122 coupled to interrupt manager 124. Conference manager 122 and conference manager 162 are configured to establish online conferences (e.g., audio calls, video calls, conference calls, etc.). In a particular example, conference manager 122 and conference manager 162 correspond to clients of a communication application (e.g., an online conferencing application). Interrupt manager 124 and interrupt manager 164 are configured to handle voice audio interrupts.

[0052] In some implementations, conference managers 122 and 162 ignore (e.g., are unaware of) any voice or audio interruptions managed by interrupt managers 124 and 164. In some implementations, conference managers 122 and 162 correspond to higher layers (e.g., the application layer) of the network protocol stacks (e.g., the Open Systems Interconnection (OSI) model) of devices 102 and 104, respectively. In some implementations, interrupt managers 124 and 164 correspond to lower layers (e.g., the transport layer) of the network protocol stacks of devices 102 and 104, respectively.

[0053] In some implementations, device 102, device 104, or both correspond to or are included in various types of devices. In illustrative examples, one or more processors 120, one or more processors 160, or combinations thereof are integrated into a headset device, as referenced. Figure 11 Further described. In other examples, one or more processors 120, one or more processors 160, or combinations thereof are integrated as described in reference 1. Figure 10 The described mobile phone or tablet computer device, as shown in the reference Figure 12 The wearable electronic devices described, as in the reference Figure 13 The described voice-controlled speaker system, as in the reference Figure 14 The camera equipment described, or as referenced Figure 15 At least one of the described virtual reality headset, augmented reality headset, or mixed reality headset. In another illustrative example, one or more processors 120, one or more processors 160, or combinations thereof are integrated into a vehicle, as referenced. Figure 16 and Figure 17 Further description.

[0054] During operation, conference managers 122 and 162 establish online meetings (e.g., audio calls, video calls, conference calls, or combinations thereof) between devices 102 and 104. For example, the online meeting is between user 142 on device 102 and user 144 on device 104. Microphone 152 captures user 142's speech while user 142 is speaking and provides audio input 153 representing that speech to device 102. In a particular aspect, camera 150 (e.g., a still camera, a video camera, or both) captures one or more images (e.g., still images or videos) of user 142 and provides video input 151 representing those images to device 102. In a particular aspect, camera 150 provides video input 151 to device 102 while microphone 152 provides audio input 153 to device 102.

[0055] Conference manager 122 generates a media stream 109 of media frames based on audio input 153, video input 151, or both. For example, media stream 109 may include an audio stream 111, a video stream 113, or both. In a particular aspect, conference manager 122 transmits media stream 109 to device 104 in real time via network 106. For example, conference manager 122 generates media frames of media stream 109 upon receiving video input 151, audio input 153, or both, and transmits (e.g., initiates transmission) media stream 109 of media frames upon generation of media frames.

[0056] In a particular implementation, during a first operating mode of device 102 (e.g., sending caption data mode), conference manager 122 generates a text stream 121, a metadata stream 123, or both, based on audio input 153. For example, conference manager 122 performs a speech-to-text conversion on audio input 153 to generate text stream 121. Text stream 121 indicates text corresponding to the speech detected in audio input 153. In a particular aspect, conference manager 122 performs speech intonation analysis on audio input 153 to generate metadata stream 123. For example, metadata stream 123 indicates the intonation (e.g., emotion, pitch, tone, or a combination thereof) of the speech detected in audio input 153. In the first operating mode of device 102 (e.g., sending caption data mode), conference manager 122 sends text stream 121, metadata stream 123, or both (e.g., as closed caption data) along with media stream 109 to device 104 (e.g., independently of network problems or audio interruptions). Alternatively, during a second operating mode of device 102 (e.g., a data interruption transmission mode), the conference manager 122 avoids generating text stream 121 and metadata stream 123 in response to determining that no audio interruption has been detected.

[0057] Device 104 receives a media stream 109 of media frames from device 102 via network 106. In a particular implementation, device 104 receives a collection (e.g., a burst) of media frames from media stream 109. In an alternative implementation, device 104 receives a media frame at a time from media stream 109. Conference manager 162 plays the media frames from media stream 109. For example, conference manager 162 generates audio output 143 based on voice audio stream 111 and plays the audio output 143 via speaker 154 (e.g., as streaming audio content). In a particular aspect, GUI generator 168 generates GUI 145 based on media stream 109, as referenced. Figure 3A Further described. For example, GUI generator 168 generates (or updates) GUI 145 to display the video content of video stream 113 and provides GUI 145 (e.g., streaming video content) to display device 156. User 144 can view the image of user 142 on display device 156 while listening to the audio voice of user 142 via speaker 154.

[0058] In one implementation, conference manager 162 stores media frames of media stream 109 in a buffer before playback. For example, conference manager 162 adds a delay between receiving media frames and playing them back at a first playback time to increase the likelihood that subsequent media frames will have a corresponding playback time (e.g., a second playback time) in the buffer. In another aspect, conference manager 162 plays media stream 109 in real time. For example, conference manager 162 retrieves media frames of media stream 109 from the buffer to play audio output 143, video content of GUI 145, or both, while subsequent media frames of media stream 109 are being received (or expected to be received) by device 104.

[0059] In a first operating mode of device 104 (e.g., a caption data display mode), conference manager 162 plays text stream 121 along with media stream 109 (e.g., independently of detecting interruptions in audio stream 111). In a particular aspect, conference manager 162, for example, receives text stream 121, metadata stream 123, or both, along with media stream 109 during the first operating mode of device 102 (e.g., a caption data transmission mode). Alternatively, conference manager 162, for example, does not receive text stream 121, metadata stream 123, or both, during a second operating mode of device 102 (e.g., an interruption data transmission mode), and generates text stream 121, metadata stream 123, or both based on audio stream 111, video stream 113, or both. For example, conference manager 162 performs speech-to-text conversion on audio stream 111 to generate text stream 121 and performs intonation analysis on audio stream 111 to generate metadata stream 123.

[0060] During a first operating mode of device 104 (e.g., display caption data mode), conference manager 162 provides text stream 121 as output to display device 156. For example, conference manager 162 uses GUI 145 along with the video content of display video stream 113 to provide audio output 143 to speaker 154, or both, while simultaneously displaying the text content of text stream 121 (e.g., as closed captions). For illustration, conference manager 162 provides text stream 121 to GUI generator 168, while simultaneously providing video stream 113 to GUI generator 168. GUI generator 168 updates GUI 145 to display text stream 121, video stream 113, or both. GUI generator 168 provides the updated GUI 145 to display device 156, while conference manager 162 provides audio stream 111 as audio output 143 to speaker 154.

[0061] In a specific example, the meeting manager 162 generates annotated text stream 137 based on text stream 121 and metadata stream 123. Specifically, the meeting manager 162 generates the annotated text stream 137 by adding annotations to text stream 121 based on metadata stream 123. The meeting manager 162 provides the annotated text stream 137 as output to display device 156. For example, the meeting manager 162 plays the annotated text stream 137 along with media stream 109. To illustrate, the meeting manager 162 uses GUI 145 to display the annotated text content of the annotated text stream 137 (e.g., as closed captions with intonation indicators) while displaying the video content of video stream 113, providing audio output 143 to speaker 154, or both.

[0062] In a particular implementation, the conference manager 162 avoids playing text stream 121 (e.g., comment text stream 137) in a second operating mode of device 104 (e.g., display interrupt data mode or closed captions disabled mode). For example, the conference manager 162 does not receive text stream 121 (e.g., during the second operating mode of device 102) and does not generate text stream 121 in the second operating mode (e.g., display interrupt data mode or closed captions disabled mode). As another example, the conference manager 162 receives text stream 121 and avoids playing text stream 121 (e.g., comment text stream 137) in response to detecting a second operating mode of device 104 (e.g., display interrupt data mode or closed captions disabled mode). In a particular aspect, the interrupt manager 164 avoids playing text stream 121 (e.g., comment text stream 137) in the second operating mode of device 104 (e.g., display interrupt data mode) in response to determining that no interruption has been detected in media stream 109 (e.g., a portion of media stream 109 corresponding to text stream 121 has been received).

[0063] In one aspect, the interrupt manager 164 initializes a speech model 131, such as an artificial neural network, based on a generic speech model before or near the start of an online meeting. In another aspect, the interrupt manager 164 selects a generic speech model from multiple generic speech models based on determining that the generic speech model matches (e.g., is associated with) demographic data of user 142 (e.g., age, location, gender, or combinations thereof). In another aspect, the interrupt manager 164 predicts demographic data based on user 142's contact information (e.g., name, location, phone number, address, or combinations thereof) before the online meeting (e.g., a scheduled meeting). In another aspect, the interrupt manager 164 estimates demographic data based on speech audio stream 111, video stream 113, or both during the initial portion of the online meeting. For example, the interrupt manager 164 analyzes speech audio stream 111, video stream 113, or both to estimate user 142's age, regional accent, gender, or combinations thereof. In a particular aspect, interrupt manager 164 retrieves a voice model 131 (e.g., previously generated) associated with user 142 (e.g., a user identifier matching user 102).

[0064] In a specific aspect, the interrupt manager 164 trains (e.g., generates or updates) the speech model 131 based on speech detected in the speech audio stream 111 during an online meeting (e.g., before an interruption in the speech audio stream 111). For illustration, the text-to-speech converter 166 is configured to perform text-to-speech conversion using the speech model 131. In a specific aspect, the interrupt manager 164 receives (e.g., during a first operating mode of device 102) or generates (e.g., during a second operating mode of device 102) a text stream 121, a metadata stream 123, or both corresponding to the speech audio stream 111. The text-to-speech converter 166 uses the speech model 131 to generate a synthesized speech audio stream 133 by performing text-to-speech conversion on the text stream 121, the metadata stream 123, or both. The interrupt manager 164 updates the speech model 131 using training techniques based on a comparison of the speech audio stream 111 and the synthesized speech audio stream 133. In an illustrative example where the speech model 131 includes an artificial neural network, the interrupt manager 164 uses backpropagation to update the weights and biases of the speech model 131. Depending on several aspects, the speech model 131 is updated such that subsequent text-to-speech conversions using the speech model 131 are more likely to generate synthesized speech that more closely matches the speech characteristics of the user 142.

[0065] In a particular aspect, interrupt manager 164 generates a virtual avatar 135 (e.g., a visual representation) for user 142. In a particular aspect, virtual avatar 135 includes or corresponds to a training indicator that indicates the training level of speech model 131, as referenced... Figures 3A-3CFurther described. For example, in response to determining that a first training criterion is not met, interrupt manager 164 initializes virtual avatar 135 to a first visual representation indicating that speech model 131 has not been trained. During an online meeting, in response to determining that the first training criterion is met but the second training criterion is not met, interrupt manager 164 updates virtual avatar 135 from the first visual representation to a second visual representation to indicate that training of speech model 131 is in progress. In response to determining that the second training criterion is met, interrupt manager 164 updates virtual avatar 135 to a third visual representation to indicate that training of speech model 131 is complete.

[0066] Training criteria may be based on the count of audio samples used to train speech model 131, the playback duration of the audio samples used to train speech model 131, the coverage of the audio samples used to train speech model 131, a success metric for speech model 131, or a combination thereof. In a particular aspect, the coverage of the audio samples used to train speech model 131 corresponds to different sounds represented by the audio samples (e.g., vowels, consonants, etc.). In a particular aspect, the success metric is based on a comparison (e.g., matching) between the audio samples used to train speech model 131 and synthesized speech generated based on speech model 131.

[0067] According to some implementations, the first color, first shadow, first size, first animation, or a combination thereof of the virtual avatar 135 indicate that the speech model 131 has not been trained. The second color, second shadow, second size, second animation, or a combination thereof of the virtual avatar 135 indicate that the speech model 131 has been partially trained. The training of the third color, third shadow, third size, third animation, or a combination thereof of the virtual avatar 135 indicates that the speech model 131 has been completed. In a particular aspect, the GUI generator 168 generates (or updates) the GUI 145 to indicate the visual representation of the virtual avatar 135.

[0068] In a specific aspect, interrupt manager 124 detects a network problem (e.g., reduced bandwidth) in the communication link of device 104. In response to the detected network problem, interrupt manager 124 sends an interrupt notification 119 to device 104 indicating an interruption in audio stream 111, preventing the transmission of subsequent media frames of media stream 109 to device 104 (e.g., stopping transmission) until the detected network problem is resolved or both. For example, interrupt manager 124, in response to the detected network problem, prevents the transmission (e.g., stops transmission) of audio stream 111, video stream 113, or both to device 104 until the interruption ends.

[0069] Interruption manager 124 sends a text stream 121, metadata stream 123, or both corresponding to subsequent media frames. For example, in a first operating mode of device 102 (e.g., sending caption data mode), interrupt manager 124 continues to send a text stream 121, metadata stream 123, or both corresponding to subsequent media frames. For illustration, in the first operating mode (e.g., sending caption data mode), conference manager 122 generates a media stream 109, text stream 121, metadata stream 123, or a combination thereof. In response to detecting a network problem in the first operating mode (e.g., sending caption data mode), interrupt manager 124 stops sending subsequent media frames of media stream 109 and continues to send a text stream 121, metadata stream 123, or both corresponding to subsequent media frames to device 104. Alternatively, in response to detecting a network problem in a second operating mode of device 102 (e.g., sending interrupt data mode), interrupt manager 124 generates a text stream 121, metadata stream 123, or both based on audio input 153 corresponding to subsequent media frames. For illustration, in the second operating mode (e.g., interrupted data transmission mode), the conference manager 122 generates media stream 109 but not text stream 121, metadata stream 123, or both. In response to detecting a network problem in the second operating mode of device 102 (e.g., interrupted data transmission mode), the interrupt manager 124 stops the transmission of subsequent media frames of media stream 109 and initiates the transmission of text stream 121, metadata stream 123, or both corresponding to the subsequent media frames to device 104. In a particular aspect, in the second operating mode of device 102 (e.g., interrupted data transmission mode), sending text stream 121, metadata stream 123, or both to device 104 corresponds to sending interruption notification 119 to device 104.

[0070] In one aspect, interrupt manager 164 detects an interrupt in voice audio stream 111 in response to receiving interrupt notification 119 from device 102. In another aspect, when device 102 is operating in a second operating mode (e.g., sending interrupt data mode), interrupt manager 164 detects an interrupt in voice audio stream 111 in response to receiving text stream 121, metadata stream 123, or both.

[0071] In a particular aspect, interrupt manager 164 detects an interruption in voice audio stream 111 in response to determining that no audio frame of voice audio stream 111 has been received within a threshold duration of the last received audio frame of voice audio stream 111. For example, the last received audio frame of voice audio stream 111 is received at a first reception time of device 104. Interrupt manager 164 detects an interruption in response to determining that no audio frame of voice audio stream 111 has been received within a threshold duration of the first reception time. In a particular aspect, interrupt manager 164 sends an interruption notification to device 102. In a particular aspect, interrupt manager 124 detects a network problem in response to receiving an interruption notification from device 104. As described above, interrupt manager 124, in response to detecting a network problem, sends text stream 121, metadata stream 123, or both to device 104 (e.g., instead of sending subsequent media frames of media stream 109).

[0072] In response to an interrupt detection, interrupt manager 164 selectively generates output based on text stream 121. For example, in response to an interrupt, interrupt manager 164 provides text stream 121, metadata stream 123, annotation text stream 137, or a combination thereof to text-to-speech converter 166. Text-to-speech converter 166 generates synthesized speech audio stream 133 by performing text-to-speech conversion based on text stream 121, metadata stream 123, annotation text stream 137, or a combination thereof using speech model 131. For example, the synthesized speech audio stream 133 based on text stream 121 and independent of metadata stream 123 corresponds to the speech indicated by text stream 121, which has neutral speech characteristics of user 142 represented by speech model 131. As another example, the synthesized speech audio stream 133 based on the annotated text stream 137 (e.g., text stream 121 and metadata stream 123) corresponds to the speech indicated by the text stream 121, which has the speech characteristics of user 142 represented by the speech model 131, which have the intonation indicated by the metadata stream 123. Performing the text-to-speech conversion using the speech model 131, which is at least partially trained on user 142's speech (e.g., speech audio stream 111), allows the synthesized speech audio stream 133 to more closely match the speech characteristics of user 142. In response to an interruption, the interrupt manager 164 provides the synthesized speech audio stream 133 as audio output 143 to the speaker 154, stops the playback of the speech audio stream 111, stops the playback of the video stream 113, or a combination thereof.

[0073] In a specific aspect, interrupt manager 164 selectively displays virtual avatar 135 while providing synthesized speech audio stream 133 as audio output 143 to speaker 154. For example, interrupt manager 164 avoids displaying virtual avatar 135 when speech audio stream 111 is provided as audio output 143 to speaker 154. As another example, interrupt manager 164 displays virtual avatar 135 while providing synthesized speech audio stream 133 as audio output 143 to speaker 154. For illustration, GUI generator 168 updates GUI 145 to display virtual avatar 135 instead of video stream 113, while synthesized speech audio stream 133 is output as audio output 143 for playback by speaker 154. In a specific aspect, interrupt manager 164 displays a first representation of virtual avatar 135 while providing speech audio stream 111 as audio output 143 to speaker 154, and a second representation of virtual avatar 135 while providing synthesized speech audio stream 133 as audio output 143 to speaker 154. For example, the first indicates that the virtual avatar 135 is being or has been trained (e.g., a training indicator for the speech model 131), and the second indicates that the virtual avatar 135 is speaking (e.g., the speech model 131 is being used to generate synthetic speech), as referenced. Figure 3C Further description.

[0074] In a particular implementation, interrupt manager 164 selectively provides text stream 121, comment text stream 137, or both as output to display device 156. For example, in response to an interruption during a second operating mode of device 104 (e.g., display interrupt data mode), interrupt manager 164 provides text stream 121, comment text stream 137, or both to GUI generator 168 to update GUI 145 to display text stream 121, comment text stream 137, or both. In an alternative implementation, during a first operating mode of device 104 (e.g., display caption data mode), interrupt manager 164 continues (e.g., independently of the interrupt) to provide text stream 121, comment text stream 137, or both as output to display device 156. In a particular aspect, interrupt manager 164 provides text stream 121, comment text stream 137, and both to display device 156 while providing synthesized speech audio stream 133 as audio output 143 to speaker 154.

[0075] In a particular implementation, interrupt manager 164 outputs one or more of a synthesized speech audio stream 133, a text stream 121, or a comment text stream 137 based on and in response to an interrupt configuration setting. For example, in response to an interrupt and determining that the interrupt configuration setting has a first value (e.g., 0 or "audio and text"), interrupt manager 164 provides text stream 121, comment text stream 137, or both to display device 156, while providing synthesized speech audio stream 133 as audio output 143 to speaker 154. In response to an interrupt and determining that the interrupt configuration setting has a second value (e.g., 1 or "text only"), interrupt manager 164 provides text stream 121, comment text stream 137, or both to display device 156, and avoids providing audio output 143 to speaker 154. Interrupt manager 164 responds to an interrupt and determines that the interrupt configuration setting has a third value (e.g., 2 or "audio only"), preventing text stream 121, comment text stream 137, or both from being provided to display device 156, and providing synthesized speech audio stream 133 as audio output 143 to speaker 154. In certain aspects, the interrupt configuration setting is based on default data, user input, or both.

[0076] In a specific aspect, interrupt manager 124 detects that an interrupt has ended and sends an interrupt end notification to device 104. For example, interrupt manager 124 detects that an interrupt has ended in response to determining that the available communication bandwidth of the communication link with device 104 is greater than a threshold. In a specific aspect, interrupt manager 164 detects that an interrupt has ended in response to receiving an interrupt end notification from device 102.

[0077] In another specific aspect, interrupt manager 164 detects that an interrupt has ended and sends an interrupt end notification to device 102. For example, interrupt manager 164 detects that an interrupt has ended in response to determining that the available communication bandwidth of the communication link with device 102 is greater than a threshold. In another specific aspect, interrupt manager 124 detects that an interrupt has ended in response to receiving an interrupt end notification from device 104.

[0078] In response to detecting that an interruption has ended, the conference manager 122 resumes sending audio stream 111, video stream 113, or both to the device 104. In a particular aspect, the transmission of audio stream 111, video stream 113, or both corresponds to the transmission of an interruption end notification. In response to detecting that an interruption has ended during a second operating mode of the device 102 (e.g., interrupted data transmission mode), the interruption manager 124 avoids sending text stream 121, metadata stream 123, or both to the device 104.

[0079] In response to detecting that an interruption has ended, the conference manager 162 avoids generating a synthesized speech audio stream 133 based on the text stream 121, avoids providing the synthesized speech audio stream 133 as audio output 143 (e.g., stops) to the speaker 154, and resumes playback of the speech audio stream 111 as audio output 143 (e.g., provides) to the speaker 154. In response to detecting that an interruption has ended, the conference manager 162 resumes providing the video stream 113 to the display device 156. For example, the conference manager 162 provides the video stream 113 to the GUI generator 168 to update the GUI 145 to display the video stream 113.

[0080] In one aspect, in response to detecting that an interruption has ended, the interruption manager 164 sends a first request to the GUI generator 168 to update the GUI 145, thereby indicating that the speech model 131 has not been used to output synthesized speech audio (e.g., the virtual avatar 135 is not speaking). In response to receiving the first request, the GUI generator 168 updates the GUI 145 to display a first representation of the virtual avatar 135, indicating that the speech model 131 is being or has been trained and that the speech model 131 is not being used to output synthesized speech audio (e.g., the virtual avatar 135 is not speaking). Alternatively, in response to detecting that an interruption has ended, the interruption manager 164 sends a second request to the GUI generator 168 to stop displaying the virtual avatar 135. For example, in response to receiving the second request, the GUI generator 168 updates the GUI 145 to avoid displaying the virtual avatar 135.

[0081] In a specific aspect, interrupt manager 164, in response to detecting that an interrupt has ended during a second operating mode (e.g., a mode that displays more interrupt data or no caption data), avoids providing text stream 121, comment text stream 137, or both to display device 156. For example, GUI generator 168 updates GUI 145 to avoid displaying text stream 121, comment text stream 137, or both.

[0082] System 100 thus reduces (e.g., eliminates) information loss during interruptions of the audio stream 111 during online meetings. For example, even if network problems prevent the audio stream 111 from being received by device 104, user 144 continues to receive audio (e.g., synthesized audio stream 133), text (e.g., text stream 121, annotated text stream 137, or both) corresponding to the speech of user 142, while text can be received by device 104, even if the text stream 111 is not received by device 104.

[0083] Although camera 150 and microphone 152 are illustrated as coupled to device 102, in other implementations, camera 150, microphone 152, or both may be integrated into device 102. Although speaker 154 and display device 156 are illustrated as coupled to device 104, in other implementations, speaker 154, display device 156, or both may be integrated into device 104. Although a microphone and a speaker are illustrated, in other implementations, one or more additional microphones configured to capture user speech, one or more additional speakers configured to output speech audio, or combinations thereof may be included.

[0084] It should be understood that, for ease of illustration, device 102 is described as a transmitting device and device 104 as a receiving device. During a call, the roles of devices 102 and 104 can switch when user 144 begins to speak. For example, device 104 can be a transmitting device, while device 102 can be a receiving device. For illustration, device 104 may include a microphone and a camera to capture audio and video of user 144, and device 102 may include or be coupled to a speaker and a display to play audio and video to user 142. In certain aspects, for example, when both user 142 and user 144 are speaking simultaneously or at overlapping times, each of device 102 and device 104 may be both a transmitting and receiving device.

[0085] In certain aspects, the conference manager 122 is also configured to perform one or more operations as described with reference to conference manager 162, and vice versa. In certain aspects, the interrupt manager 124 is also configured to perform one or more operations as described with reference to interrupt manager 164, and vice versa. Although the GUI generator 168 is described differently from conference manager 162 and interrupt manager 164, in other implementations, the GUI generator 168 is integrated into conference manager 162, interrupt manager 164, or both. For illustration, in some examples, conference manager 162, interrupt manager 164, or both are configured to perform some of the operations described with reference to GUI generator 168.

[0086] refer to Figure 2 This illustrates a system operable to handle interruptions in a voice audio stream, and it is generally specified as 200. In certain aspects, Figure 1 System 100 includes one or more components of system 200.

[0087] System 200 includes a server 204 coupled to devices 102 and 104 via network 106. Server 204 includes a conference manager 122 and an interrupt manager 124. Server 204 is configured to forward online conference data from device 102 to device 104 and vice versa. For example, conference manager 122 is configured to establish an online conference between device 102 and device 104.

[0088] Device 102 includes a conference manager 222. During an online meeting, conference manager 222 sends media stream 109 (e.g., audio stream 111, video stream 113, or both) to server 204. Conference manager 122 of server 204 receives media stream 109 (e.g., audio stream 111, video stream 113, or both) from device 102. In a particular implementation, device 102 sends text stream 121, metadata stream 123, or both simultaneously with sending media stream 109 to server 204.

[0089] In certain aspects, such as reference Figure 1 As described, server 204 is used instead of device 102 to perform subsequent operations. For example, conference manager 122 (on server 204 instead of device 102) Figure 1 (The operation of device 102 in the reference) Figure 1 Media stream 109, text stream 121, metadata stream 123, or a combination thereof, are sent to device 104 in a manner similar to those described. For example, during a first operating mode of server 204 (e.g., sending caption data mode), conference manager 122 sends text stream 121, metadata stream 123, or both. In a particular implementation, conference manager 122 forwards text stream 121, metadata stream 123, or both received from device 102 to device 104. In some implementations, conference manager 122 generates metadata stream 123 based on text stream 121, media stream 109, or a combination thereof. In these implementations, conference manager 122 forwards text stream 121 received from device 102 to device 104 and sends metadata stream 123 generated at server 204 to device 104, or both. In some implementations, conference manager 122 generates text stream 121, metadata stream 123, or both based on media stream 109 and forwards text stream 121, metadata stream 123, or both to device 104. Alternatively, during a second operating mode of server 204 (e.g., interrupted data transmission mode), conference manager 122 avoids transmitting text stream 121, metadata stream 123, or both in response to determining that no interruption has been detected. Device 104 receives media stream 109, text stream 121, comment text stream 137, or a combination thereof from server 204 via network 106. Conference manager 162 plays media frames of media stream 109, text stream 121, comment text stream 137, or a combination thereof, as referenced. Figure 1 As described. Interrupt manager 164 trains speech model 131, displays virtual avatar 135, or both, as referenced. Figure 1 As described.

[0090] In a specific aspect, interrupt manager 124, in response to detecting a network problem, sends an interrupt notification 119 indicating an interruption in the audio stream 111 to device 104, preventing the transmission of subsequent media frames of media stream 109 to device 104 (e.g., stopping transmission) until the network problem is detected to be resolved (e.g., the interruption has ended) or both. Interruption manager 124 then sends a text stream 121, a metadata stream 123, or both corresponding to subsequent media frames to device 104, as referenced. Figure 1 As described. For example, interrupt manager 124 forwards text stream 121, metadata stream 123, or both received from device 102 to device 104. In some examples, interrupt manager 124 sends metadata stream 123, text stream 121, or both generated at server 204 to device 104. In a particular aspect, interrupt manager 124 selectively generates metadata stream 123, text stream 121, or both in response to the detection of an interruption in voice audio stream 111 during a second operating mode of server 204 (e.g., send interrupted data mode).

[0091] In certain aspects, the interrupt manager 164 is in accordance with the reference Figure 1 In a similar manner to the described one, in response to receiving an interruption notification 119 from interrupt manager 124 (e.g., at server 204), when server 204 is operating in a second operating mode (e.g., sending interrupted data mode) and receiving text stream 121, metadata stream 123, or both, it determines that no audio frame of voice audio stream 111, or a combination thereof, has been received within a threshold duration of the last received audio frame of voice audio stream 111, and thus detects an interruption in voice audio stream 111. In a particular aspect, interrupt manager 164 sends an interruption notification to server 204. In a particular aspect, interrupt manager 124 detects a network problem in response to receiving an interruption notification from device 104. Interruption manager 124 sends text stream 121, metadata stream 123, or both corresponding to subsequent media frames to device 104, as referenced. Figure 1 As described.

[0092] In response to an interrupt detection, interrupt manager 164 provides text stream 121, metadata stream 123, annotation text stream 137, or a combination thereof to text-to-speech converter 166. Text-to-speech converter 166 generates a synthesized speech audio stream 133 by performing text-to-speech conversion based on text stream 121, metadata stream 123, annotation text stream 137, or a combination thereof using speech model 131, as referenced. Figure 1As described. Interrupt manager 164, in response to an interrupt, provides synthesized speech audio stream 133 as audio output 143 to speaker 154, stops playback of speech audio stream 111, stops playback of video stream 113, displays virtual avatar 135, displays a specific representation of virtual avatar 135, displays text stream 121, displays comment text stream 137, or combinations thereof, as described in the reference. Figure 1 As described.

[0093] In response to the detection that an interruption has ended, the conference manager 122 resumes sending audio stream 111, video stream 113, or both to device 104. In a particular aspect, in response to the detection that an interruption has ended during a second operating mode of server 204 (e.g., sending interrupted data mode), the interruption manager 124 avoids sending (e.g., stops sending) text stream 121, metadata stream 123, or both to device 104.

[0094] In response to detecting that an interruption has ended, the meeting manager 162 avoids generating a synthesized speech audio stream 133 based on the text stream 121, avoids providing the synthesized speech audio stream 133 as audio output 143 (e.g., stops) to the speaker 154, resumes playback of the speech audio stream 111 as audio output 143 to the speaker 154, resumes providing the video stream 113 to the display device 156, stops or adjusts the display of the virtual avatar 135, avoids providing the text stream 121 to the display device 156, avoids providing the comment text stream 137 to the display device 156, or a combination thereof.

[0095] Therefore, system 200 reduces (e.g., eliminates) information loss during interruptions of the voice audio stream 111 during online meetings with conventional devices (e.g., device 102 excluding the interrupt manager). For example, even if network problems prevent the voice audio stream 111 from being received by device 104, user 144 continues to receive audio (e.g., synthesized voice audio stream 133), text (e.g., text stream 121, annotated text stream 137, or both) corresponding to the voice of user 142, even if the text can be received by device 104.

[0096] In certain aspects, server 204 may also be closer to device 104 (e.g., fewer network hops), and sending text stream 121, metadata stream 123, or both from server 204 (e.g., rather than from device 102) can save all network resources. In certain aspects, server 204 may have access to network information that can be used to successfully send text stream 121, metadata stream 123, or both to device 104. For example, server 204 initially sends media stream 109 via a first network link. Server 204 detects a network problem and, at least in part based on the determination that the first network link is unavailable or ineffective, uses a second network link that appears capable of accommodating text transmission to send text stream 121, metadata stream 123, or both.

[0097] refer to Figure 3A An example of GUI 145 is shown. In certain aspects, GUI 145 is composed of... Figure 1 System 100 Figure 2 The system is generated by system 200 or both.

[0098] GUI 145 includes video display 306, virtual avatar 135, and training indicator (TI) 304. For example, GUI generator 168 generates GUI 145 during the start of an online meeting. Video stream 113 (e.g., an image of user 142, such as Jill Pratt) is displayed via video display 306.

[0099] Training indicator 304 indicates the training level of speech model 131 (e.g., 0% or not trained). For example, training indicator 304 indicates that speech model 131 has not yet been custom-trained. In a particular aspect, the representation of virtual avatar 135 (e.g., solid color) also indicates the training level. In a particular aspect, the representation of virtual avatar 135 indicates that synthesized speech has not been output. For example, GUI 145 does not include a synthesized speech indicator, as referenced... Figure 3C Further description.

[0100] In a particular implementation, if an interruption is generated before the custom-trained speech model 131, and the text-to-speech converter 166 uses the speech model 131 (e.g., a non-customized general speech model) to generate a synthesized speech audio stream 133, then the synthesized speech audio stream 133 corresponds to audio speech with general speech characteristics that may differ from those of the user 142. In a particular aspect, the speech model 131 is initialized using a general speech model associated with the user 142's demographic data. In this aspect, the synthesized speech audio stream 133 corresponds to general speech features that match the user 142's demographic data (e.g., age, gender, regional accent, etc.).

[0101] refer to Figure 3BAn example of GUI 145 is shown. In certain aspects, GUI 145 is composed of... Figure 1 System 100 Figure 2 The system is generated by system 200 or both.

[0102] In a specific example, GUI generator 168 updates GUI 145 during an online meeting. Training indicator 304 indicates a second training level (e.g., 20% or partially trained) for the speech model 131. For example, training indicator 304 indicates that the speech model 131 is being custom-trained or has been partially custom-trained. In a specific aspect, the representation of virtual avatar 135 (e.g., partially colored) also indicates the second training level. In a specific aspect, the representation of virtual avatar 135 indicates that synthesized speech is not output. For example, GUI 145 does not include a synthesized speech indicator.

[0103] In a particular implementation, if an interrupt is generated after partially customized training of the speech model 131 and the text-to-speech converter 166 uses the speech model 131 (e.g., a partially customized speech model) to generate a synthesized speech audio stream 133, then the synthesized speech audio stream 133 corresponds to audio speech with some similarities to the speech characteristics of the user 142.

[0104] refer to Figure 3C An example of GUI 145 is shown. In certain aspects, GUI 145 is composed of... Figure 1 System 100 Figure 2 The system is generated by system 200 or both.

[0105] In a specific example, GUI generator 168 updates GUI 145 in response to an interrupt. Training indicator 304 indicates a third training level (e.g., 100% or training complete) for speech model 131. For example, training indicator 304 indicates that speech model 131 is custom-trained or that custom training has been completed (e.g., a threshold level has been reached). In a specific aspect, the representation of avatar 135 (e.g., fully colored) also indicates the third training level. In a specific aspect, the representation of avatar 135 indicates that synthesized speech is being output. For example, GUI 145 includes a synthesized speech indicator 398 displayed as part of or with avatar 135 to indicate that the speech being played is synthesized speech.

[0106] Because in Figure 3C In the example, the interruption occurs after custom training of the speech model 131, and the text-to-speech converter 166 uses the speech model 131 (e.g., a custom speech model) to generate a synthetic speech audio stream 133, so the synthetic speech audio stream 133 corresponds to audio speech with speech characteristics similar to those of the user 142.

[0107] In response to an interrupt, interrupt manager 164 stops the output of video stream 113. For example, video display 306 indicates that the output of video stream 113 has stopped due to an interrupt (e.g., a network problem). GUI 145 includes text display 396. For example, interrupt manager 164 outputs text stream 121 via text display 396 in response to an interrupt.

[0108] In a specific aspect, text stream 121 is displayed in real time, allowing user 144 to continue participating in the conversation. For example, user 144 can respond to user 142 after reading the text displayed 396 spoken by user 142. In a specific aspect, if a network problem prevents the audio stream corresponding to user 144's speech from being received by device 102, interrupt manager 124 can display a text stream corresponding to user 144's speech at device 102. One or more participants in the online meeting can thus receive text streams or audio streams corresponding to the speech of other participants.

[0109] refer to Figure 4A , showed Figure 1 System 100 or Figure 2 The diagram illustrates the operation of system 200 and is generally represented by 400. Figure 4A The timings and operations shown are for illustrative purposes only and not as limitations. Additional or fewer operations may be performed, and the timings may differ.

[0110] Figure 400 illustrates the timing of media frame transmission from media stream 109 of device 102. In a particular aspect, media frames of media stream 109 are transmitted from device 102 to device 104, as shown in the reference. Figure 1 As described. Alternatively, media frames of media stream 109 are sent from device 102 to server 204 and from server 204 to device 102, as described in reference [reference needed]. Figure 2 As described.

[0111] Device 102 transmits media frames (FRs) 410 of media stream 109 at a first transmission time. Device 104 receives media frames 410 at a first reception time and provides media frames 410 for playback at a first playback time. In a particular example, conference manager 162 stores media frames 410 in a buffer during a first buffering interval between the first reception time and the first playback time. In a particular aspect, media frame 410 includes a first portion of video stream 113 and a first portion of audio stream 111. At the first playback time, conference manager 162 outputs the first portion of audio stream 111 as the first portion of audio output 143 to speaker 154 and outputs the first portion of video stream 113 to display device 156.

[0112] Device 102 (or server 204) is expected to transmit media frame 411 at a second expected transmission time. Device 104 is expected to receive media frame 411 at a second expected reception time. In response to determining that no media frame of media stream 109 has been received within a reception threshold duration of the first reception time, interrupt manager 164 of device 104 detects an interruption in voice audio stream 111. For example, interrupt manager 164 determines a second time based on the first reception time and the reception threshold duration (e.g., second time = first reception time + reception threshold duration). In response to determining that no media frame of media stream 109 has been received between the first reception time and the second time, interrupt manager 164 detects an interruption in voice audio stream 111. The second time is after the second expected reception time of media frame 411 and before the expected playback time of media frame 411. For example, the second time is during the expected buffering interval of media frame 411.

[0113] Device 102 (or server 204) detects an interruption in the voice audio stream 111, as shown in the reference. Figure 1-2 As described. In response to an interruption in the audio stream 111, the interrupt manager 124 (of device 102 or server 204) sends a text stream 121 corresponding to subsequent media frames (e.g., a set of media frames 491) to device 104 until the interruption ends. In a particular aspect, media frame 411 includes a second portion of video stream 113 and a second portion of audio stream 111. The interrupt manager 124 (or conference manager 122) generates text 451 of text stream 121 by performing a speech-to-text conversion on the second portion of audio stream 111 and sends text 451 to device 104.

[0114] Device 104 receives text 451 from text stream 121 from device 102 or server 204, as shown in the reference. Figure 1-2 As described. In response to an interrupt, interrupt manager 164 initiates playback of text stream 121 corresponding to subsequent media frames until the interrupt ends. For example, interrupt manager 164 provides text 451 to display device 156 at a second playback time. In a particular aspect, the second playback time is based on (e.g., the same as) the expected playback time of media frame 411.

[0115] In certain respects, Figure 2 The conference manager 222 is unaware of the interruption and sends media frame 413 of media stream 109 to server 204. In a particular aspect, ( Figure 1 Device 102 or Figure 2In response to the interrupt, the interrupt manager 124 of server 204 stops the transmission of media frame 413 to device 104. In a specific aspect, media frame 413 includes a third portion of video stream 113 and a third portion of audio stream 111. The interrupt manager 124 generates text 453 based on the third portion of audio stream 111. The interrupt manager 124 sends text 453 to device 104.

[0116] Device 104 receives text 453. In response to an interrupt, interrupt manager 164 provides text 453 to display device 156 at a third playback time. In a particular aspect, the third playback time is based on (for example, the same as) the expected playback time of media frame 413.

[0117] In response to the end of an interrupt, the interrupt manager 124 of device 102 or server 204 resumes the transmission of subsequent media frames of media stream 109 (e.g., the next media frame 493) to device 104, as per reference. Figure 1-2 For example, conference manager 122 sends media frame 415 to device 104. In response to the end of an interruption, interrupt manager 164 resumes playback of media stream 109 and stops playback of text stream 121. In a particular aspect, media frame 415 includes a fourth portion of video stream 113 and a fourth portion of audio stream 111. During the fourth playback time, conference manager 162 outputs the fourth portion of audio stream 111 as part of audio output 143 to speaker 154 and outputs the fourth portion of video stream 113 to display device 156.

[0118] As another example, conference manager 122 sends media frame 417 to device 104. In a specific aspect, media frame 417 includes a fifth portion of video stream 113 and a fifth portion of audio stream 111. At the fifth playback time, conference manager 162 outputs the fifth portion of audio stream 111 as part of audio output 143 to speaker 154 and outputs the fifth portion of video stream 113 to display device 156.

[0119] Therefore, device 104 prevents information loss by replaying text stream 121 during an interruption in media stream 109. When the interruption ends, playback of media stream 109 resumes.

[0120] refer to Figure 4B , showed Figure 1 System 100 or Figure 2 The diagram illustrates the operation of system 200 and is generally represented as 490. Figure 4B The timings and operations shown are for illustrative purposes only and not as limitations. Additional or fewer operations may be performed, and the timings may differ.

[0121] Figure 490 illustrates the timing of the transmission of media frames from media stream 109 of device 102. Figure 1 The GUI generator 168 generates a GUI 145 indicating the training level of the virtual avatar 135. For example, GUI 145 indicates that the virtual avatar 135 (e.g., speech model 131) is untrained or partially trained. Device 104 receives media frames 410 including a first portion of video stream 113 and a first portion of audio stream 111. The conference manager 162 outputs the first portion of audio stream 111 as the first portion of audio output 143 to speaker 154 during the first playback time, and outputs the first portion of video stream 113 to display device 156, as referenced. Figure 4A As described. The interrupt manager 164 trains a speech model 131 based on media frame 410 (e.g., the first part of the speech audio stream 111), as referenced. Figure 1 As described, the GUI generator 168 updates the training level (e.g., partially or fully trained) of the virtual avatar 135 in the GUI 145.

[0122] Device 104 receives text 451 from text stream 121 from device 102 or server 204, as shown in the reference. Figure 4A As described. In response to the interruption, interrupt manager 164 stops playback of media stream 109, stops training of speech model 131, and starts playback of synthesized speech audio stream 133. For example, interrupt manager 164 generates a synthesized speech frame 471 of synthesized speech audio stream 133 based on text 451. For illustration, interrupt manager 164 provides text 451 to text-to-speech converter 166. Text-to-speech converter 166 performs text-to-speech conversion on text 451 using speech model 131 to generate synthesized speech frame (SFR) 471. Interrupt manager 164 provides synthesized speech frame 471 as a second part of audio output 143 during a second playback time. GUI generator 168 updates GUI 145 to include a synthesized speech indicator 398 indicating that synthesized speech is being output. For example, GUI 145 indicates that avatar 135 is speaking.

[0123] Device 104 receives text 453, as per reference. Figure 4A As described. In response to the interrupt, interrupt manager 164 generates a synthesized speech frame 473 of the synthesized speech audio stream 133 based on text 453. Interrupt manager 164 provides the synthesized speech frame 473 as the third part of the audio output 143 at the third playback time.

[0124] In response to the end of an interrupt, the interrupt manager 124 of device 102 or server 204 resumes the transmission of subsequent media frames of media stream 109 (e.g., the next media frame 493) to device 104, as per reference. Figure 4A As described above. For example, conference manager 122 sends media frame 415 to device 104. In response to the end of the interrupt, interrupt manager 164 resumes playback of media stream 109, stops playback of synthesized speech audio stream 133, and resumes training of speech model 131. GUI generator 168 updates GUI 145 to remove synthesized speech indicator 398, thereby indicating that synthesized speech is not being output.

[0125] In a specific example, conference manager 162 plays media frames 415 and 417. For illustration, media frame 415 includes a fourth portion of video stream 113 and a fourth portion of audio stream 111. During the fourth playback, conference manager 162 outputs the fourth portion of audio stream 111 as the fourth portion of audio output 143 to speaker 154 and outputs the fourth portion of video stream 113 to display device 156. In a specific aspect, during the fifth playback, conference manager 162 outputs the fifth portion of audio stream 111 as the fifth portion of audio output 143 to speaker 154 and outputs the fifth portion of video stream 113 to display device 156.

[0126] Therefore, device 104 prevents information loss by replaying the synthesized speech audio stream 133 during the interruption of media stream 109. When the interruption ends, the playback of media stream 109 resumes.

[0127] refer to Figure 5 This illustrates a system operable to handle interruptions in a voice audio stream, and it is generally specified as 500. In certain aspects, Figure 1 System 100 includes one or more components of system 500.

[0128] System 500 includes device 502 coupled to device 104 via network 106. During operation, conference manager 162 establishes online conferences with multiple devices (e.g., device 102 and device 502). For example, conference manager 162 establishes an online conference between user 144 and user 142 of device 102 and user 542 of device 502. Device 104 receives media streams 109 (e.g., audio stream 111, video stream 113, or both) representing user 142's voice, images, or both from device 102 or server 204, as referenced. Figure 1-2 As described. Similarly, device 104 receives from device 502 or a server (e.g., server 204 or another server) a media stream 509 representing the voice, images, or both of user 542 (e.g., a second voice audio stream 511, a second video stream 513, or both).

[0129] For reference Figure 6AFurther described, the conference manager 162 plays media stream 109 while playing media stream 509. For example, the conference manager 162 provides video stream 113 to display device 156 simultaneously with a second video stream 513. To illustrate, user 144 can simultaneously view the images of user 142 and user 542 during an online meeting. As another example, the conference manager 162 provides audio stream 111, a second audio stream 511, or both as audio output 143 to speaker 154. To illustrate, user 144 can hear the voice of user 142, the voice of user 542, or both. In a particular aspect, the interrupt manager 164 trains a speech model 131 based on audio stream 111, as referenced... Figure 1 As described. Similarly, the interrupt manager 164 trains a second speech model for user 542 based on the second speech audio stream 511.

[0130] In a specific example, device 104 continues to receive media stream 509 during an interruption of speech audio stream 111. Interrupt manager 164 plays media stream 509 while playing synthesized speech audio stream 133, text stream 121, comment text stream 137, or a combination thereof, as referenced. Figure 6C Further described. For example, interrupt manager 164 provides a second audio stream 511 while generating and providing the synthesized speech audio stream 133 to speaker 154. As another example, interrupt manager 164 provides a second video stream 513 to display device 156 while generating an update to GUI 145, including text stream 121 or comment text stream 137, and providing the update to display device 156. User 144 can thus follow the conversation between user 142 and user 542 during the interruption of audio stream 111.

[0131] In a specific aspect, the interruption in media stream 509 overlaps with the interruption in speech audio stream 111. Interruption manager 164 receives a second text stream, a second metadata stream, or both corresponding to the second speech audio stream 511. In a specific aspect, interruption manager 164 generates a second annotated text stream based on the second text stream, the second metadata stream, or both. Interruption manager 164 generates a second synthesized speech audio stream by performing text-to-speech conversion based on the second text stream, the second metadata stream, the second annotated text stream, or a combination thereof using a second speech model. Interruption manager 164 plays the second speech audio stream 511 to speaker 154 while playing synthesized speech audio stream 133. In a specific aspect, interruption manager 164 plays text stream 121, annotated text stream 137, or both while playing the second text stream, the second annotated text stream, or both to display device 156. Therefore, during the interruption of speech audio stream 111 and the second speech audio stream 511, user 144 can follow the dialogue between user 142 and user 542.

[0132] Therefore, system 500 reduces (e.g., eliminates) information loss during interruptions of one or more audio streams (e.g., audio stream 111, second audio stream 511, or both) during online meetings with multiple users. For example, even if network problems prevent one or more audio streams from being received by device 104, user 144 continues to receive audio, text, or a combination thereof corresponding to the voice of user 142 and the voice of user 542, provided that text can be received by device 104.

[0133] refer to Figure 6A An example of GUI 145 is shown. In certain aspects, GUI 145 is composed of... Figure 5 The system 500 was generated.

[0134] GUI 145 includes video displays, virtual avatars, training indicators, or combinations thereof for multiple participants in an online meeting. For example, GUI 145 includes a video display 306, a virtual avatar 135, a training indicator 304, or a combination thereof for user 142, as referenced. Figure 3A As described. GUI 145 also includes a video display 606 for user 542, a virtual avatar 635, a training indicator (TI) 604, or a combination thereof. For example, GUI generator 168 generates GUI 145 during the start of an online meeting. A second video stream 513 of media stream 509 (e.g., an image of user 542 (e.g., Emily F.)) is displayed via video display 606, while video stream 113 (e.g., an image of user 142 (e.g., Jill P.)) is displayed via video display 306.

[0135] Training indicator 304 indicates the training level of speech model 131 (e.g., 0% or untrained), and training indicator 604 indicates the training level of the second speech model (e.g., 10% or partially trained). The training level of the speech models may differ if one user speaks more than another, or if one user's speech includes more types of sounds (e.g., higher model coverage).

[0136] In certain aspects, the representations of virtual avatar 135 (e.g., solid color) and virtual avatar 635 (e.g., partial coloring) also indicate the training level of the respective speech models. In certain aspects, the representations of virtual avatar 135 and virtual avatar 635 indicate that synthesized speech was not output. For example, GUI 145 does not include any synthesized speech indicators.

[0137] In one implementation, if an interrupt is generated when receiving media stream 109, text-to-speech converter 166 uses speech model 131 (e.g., a non-customized general speech model) to generate synthesized speech audio stream 133. If an interrupt is generated during receiving media stream 509, text-to-speech converter 166 uses a second speech model (e.g., a partially customized speech model) to generate a second synthesized speech audio stream. In one aspect, interrupt manager 164 initializes the second speech model based on a second general speech model different from the first general speech model used to initialize speech model 131, such that if an interrupt is generated before training (or full training) of speech model 131 and the second speech model, the synthesized speech of user 142 can be distinguished from the synthesized speech of user 542. In one aspect, speech model 131 is initialized using a first general speech model associated with the demographics of user 142, and the second speech model is initialized using a second general speech model associated with the demographics of user 542.

[0138] refer to Figure 6B An example of GUI 145 is shown. In certain aspects, GUI 145 is composed of... Figure 5 The system 500 was generated.

[0139] In a specific example, GUI generator 168 updates GUI 145 during an online meeting. For example, training indicator 304 indicates a second training level for speech model 131 (e.g., 20% or partially trained) and a second training level for the second speech model (e.g., 100% or fully trained).

[0140] refer to Figure 6C An example of GUI 145 is shown. In certain aspects, GUI 145 is composed of... Figure 5 The system 500 was generated.

[0141] In a particular example, GUI generator 168 updates GUI 145 in response to an interruption in received media stream 109. Training indicator 304 indicates a third training level (e.g., 55% or partially trained) for speech model 131, and training indicator 604 indicates a third training level (e.g., 100% or fully trained) for the second speech model. In a particular aspect, representations of virtual avatar 135 indicate that synthesized speech is being output. For example, GUI 145 includes a synthesized speech indicator 398. Representations of virtual avatar 635 indicate that synthesized speech is not being output to user 542. For example, GUI 145 does not include a synthesized speech indicator associated with virtual avatar 635.

[0142] In response to an interrupt, interrupt manager 164 stops the output of video stream 113. For example, video display 306 indicates that the output of video stream 113 has stopped due to an interrupt (e.g., a network problem). In response to this interrupt, interrupt manager 164 outputs text stream 121 via text display 396.

[0143] In certain aspects, text stream 121 is displayed in real time, allowing user 144 to continue following and participating in the conversation. For example, user 144 may hear from synthesized speech audio stream 133, read on text display 396, or both, and user 142 may make a first statement (e.g., “I hope you have something similar to celebrate”). User 144 may hear a response from user 542 in a second speech audio stream of media stream 509 output from speaker 154. User 144 may hear from synthesized speech audio stream 133, read on text display 396, or both, and user 142 may make a second statement (e.g., “That was so much fun! I’m glad you had fun”). User 144 can thus listen to audio from synthesized speech audio stream, read text from text stream, or both for one or more other participants in the online meeting while receiving media streams from one or more other participants in the online meeting.

[0144] refer to Figure 7A , showed Figure 5 The diagram illustrates the operation of system 500 and is generally represented by 700. Figure 7A The timings and operations shown are for illustrative purposes only and not as limitations. Additional or fewer operations may be performed, and the timings may differ.

[0145] Figure 700 illustrates the timing of the transmission of media stream 109 from device 102 and media frames of media stream 509 from device 502. In a particular aspect, media frames of media stream 109 are transmitted from device 102 or server 204 to device 104, as referenced. Figure 1-2As described. Similarly, media frames of media stream 509 are sent from device 502 or a server (e.g., server 204 or another server) to device 104.

[0146] Device 104 receives media frame 410 of media stream 109 and media frame 710 of media stream 509, and provides media frame 410 and media frame 710 for playback. For example, conference manager 162 outputs a first portion of audio stream 111 (e.g., indicated by media frame 410) and a first portion of a second audio stream (e.g., indicated by media frame 710) as audio output 143 to speaker 154, outputs a first portion of video stream 113 (e.g., indicated by media frame 410) via video display 306, and outputs a first portion of a second video stream (e.g., indicated by media frame 710) via video display 606, as referenced. Figure 6A As described.

[0147] During an interruption of media stream 109, device 104 receives text 451 (corresponding to media frame 411) from text stream 121, as referenced. Figure 4A As described. Device 104 receives media frame 711 of media stream 509. In response to this interrupt, interrupt manager 164 initiates playback of text stream 121 corresponding to subsequent media frames of media stream 109 simultaneously with the playback of media stream 509, until the interrupt ends. For example, interrupt manager 164 provides text 451 (e.g., indicated by media frame 411) to display device 156 while providing media frame 711 for playback.

[0148] During an interruption of media stream 109, device 104 receives text 453 (corresponding to media frame 413) from text stream 121, as referenced. Figure 4A As described. Device 104 receives media frame 713 of media stream 509. Interrupt manager 164 provides text 453 to display device 156 while providing media frame 713 for playback.

[0149] Interrupt manager 164, in response to the end of an interrupt, resumes playback of media stream 109 and stops playback of text stream 121, as per reference. Figure 4A The conference manager 162 receives and plays back media frames 415 and 715. Similarly, the conference manager 162 receives and plays back media frames 417 and 717.

[0150] Therefore, device 104 prevents information loss by playing back text stream 121 simultaneously with the playback of media stream 509 during an interruption in media stream 109. When the interruption ends, playback of media stream 109 resumes.

[0151] refer to Figure 7B , showed Figure 5The diagram illustrates the illustrative aspects of the operation of system 500, and is typically designated as 790. Figure 7B The timings and operations shown are for illustrative purposes only and not as limitations. Additional or fewer operations may be performed, and the timings may differ.

[0152] Figure 790 illustrates the transmission timing of media stream 109 from device 102 and media frames of media stream 509 from device 502. Figure 1 The GUI generator 168 generates a GUI 145 indicating the training level of virtual avatar 135 and virtual avatar 635. For example, GUI 145 indicates that virtual avatar 135 (e.g., speech model 131) has not been trained, while virtual avatar 635 (e.g., a second speech model) has been partially trained. Device 104 receives and plays back media frames 410 and 710. Interrupt manager 164 trains speech model 131 based on media frame 410, as referenced. Figure 4B The second speech model is trained based on media frame 710. GUI generator 168 updates GUI 145 to indicate the updated training level (e.g., partially trained) of virtual avatar 135 and the updated training level (e.g., fully trained) of virtual avatar 635.

[0153] Device 104 receives text 451 and media frame 711 from text stream 121. Interrupt manager 164 generates synthesized speech frame 471 based on text 451, as shown in the reference. Figure 4B The interrupt manager 164 replays the synthesized speech frame 471 and media frame 711. The GUI generator 168 updates the GUI 145 to include a synthesized speech indicator 398 indicating that synthesized speech is being output for user 142. For example, GUI 145 indicates that avatar 135 is speaking. GUI 145 does not include a synthesized speech indicator for user 542 (e.g., avatar 635 is not indicated as speaking).

[0154] Device 104 receives text 453 and media frame 713. Interrupt manager 164 generates synthesized speech frame 473 based on text 453, as shown in the reference. Figure 4B The interrupt manager 164 plays back the synthesized audio frame 473 and media frame 417.

[0155] Interrupt manager 164, in response to the end of an interrupt, resumes playback of media stream 109, stops playback of synthesized speech audio stream 133, and resumes training of speech model 131, as per reference. Figure 4B The GUI generator 168 updates the GUI 145 to remove the synthesized speech indicator 398, thereby indicating that the synthesized speech has not been output.

[0156] In a specific example, conference manager 162 receives and plays media frames 415 and 715. As another example, conference manager 162 receives and plays media frames 417 and 717.

[0157] Therefore, device 104 prevents information loss by playing back the synthesized speech audio stream 133 while playing media stream 509 during an interruption in media stream 109. When the interruption ends, playback of media stream 109 resumes.

[0158] refer to Figure 8 This illustrates a specific implementation of a method 800 for handling interruptions to a voice audio stream. In a particular aspect, one or more operations of method 800 are performed by a conference manager 162, an interrupt manager 164, one or more processors 160, and a device 104. Figure 1 The system 100 or a combination thereof is used to execute.

[0159] Method 800 includes, at 802, receiving a speech audio stream representing the speech of a first user during an online meeting. For example, Figure 1 During an online meeting, device 104 receives an audio stream 111 representing the voice of user 142, as shown in the reference. Figure 1 As described.

[0160] Method 800 also includes, at 804, receiving a text stream representing the speech of the first user. For example, Figure 1 The device 104 receives a text stream 121 representing the voice of the user 142, as shown in the reference. Figure 1 As described.

[0161] Method 800 also includes, at 806, selectively generating output based on the text stream in response to an interruption in the speech audio stream. For example, Figure 1 In response to an interrupt in the speech audio stream 111, the interrupt manager 164 selectively generates a synthesized speech audio stream 133 based on the text stream 121, as shown in the reference. Figure 1 As described. In a particular implementation, interrupt manager 164 selectively outputs text stream 121, comment text stream 137, or both, in response to an interrupt in the voice audio stream 111, as described in the reference. Figure 1 As described.

[0162] Therefore, method 800 improves upon the reduction (e.g., elimination) of information loss during interruptions to the audio stream 111 during online meetings. For example, even if network problems prevent the audio stream 111 from being received by device 104, user 144 continues to receive audio (e.g., synthesized audio stream 133), text (e.g., text stream 121, annotated text stream 137, or both) corresponding to the speech of user 142, while text can be received by device 104, even if the text stream 111 is not received by device 104.

[0163] Figure 8 Method 800 can be implemented by a field-programmable gate array (FPGA) device, an application-specific integrated circuit (ASIC), a processing unit such as a central processing unit (CPU), a DSP, a controller, another hardware device, a firmware device, or any combination thereof. As an example, Figure 8 Method 800 can be executed by a processor that runs instructions, such as reference... Figure 18 As described.

[0164] Figure 9 An implementation 900 of device 104 is depicted as an integrated circuit 902 including one or more processors 160. Integrated circuit 902 also includes inputs 904 (e.g., one or more bus interfaces) to enable receiving input data 928 (e.g., voice audio stream 111, video stream 113, media stream 109, interrupt notification 119, text stream 121, metadata stream 123, media stream 509, or combinations thereof) for processing. Integrated circuit 902 also includes outputs 906 (e.g., bus interfaces) to enable sending output signals (e.g., voice audio stream 111, synthesized voice audio stream 133, audio output 143, video stream 113, text stream 121, annotation text stream 137, GUI 145, or combinations thereof). Integrated circuit 902 enables implementations that handle voice audio stream interrupts to be used as components in systems such as... Figure 10 The mobile phone or tablet computer depicted in the image Figure 11 The headphones depicted in the text Figure 12 Wearable electronic devices depicted in the text Figure 13 The voice-controlled speaker system depicted in the text Figure 14 The camera depicted in the text Figure 15 The virtual reality headset or augmented reality headset depicted in the text Figure 16 or Figure 17 The vehicles depicted in the text.

[0165] Figure 10An implementation 1000 is depicted, wherein device 104 includes a mobile device 1002, such as a telephone or tablet, as an illustrative, non-limiting example. Mobile device 1002 includes a microphone 1010, a speaker 154, and a display screen 1004. Components of one or more processors 160, including a conference manager 162, an interrupt manager 164, a GUI generator 168, or combinations thereof, are integrated into mobile device 1002 and are shown using dashed lines to indicate internal components generally not visible to the user of mobile device 1002. In a particular example, conference manager 162 outputs a voice audio stream 111 or interrupt manager 164 outputs a synthesized voice audio stream 133, which is then processed to perform one or more operations at mobile device 1002, such as launching a graphical user interface or otherwise displaying additional information associated with the user's voice at display screen 1004 (e.g., via an integrated "smart assistant" application).

[0166] Figure 11 An implementation 1100 of which device 104 includes a headset device 1102 is depicted. The headset device 1102 includes a speaker 154, a microphone 1110, or both. Components of one or more processors 160 (including a conference manager 162, an interrupt manager 164, or both) are integrated into the headset device 1102. In a particular example, the conference manager 162 outputs a speech audio stream 111, or the interrupt manager 164 outputs a synthesized speech audio stream 133, which allows the headset device 1102 to perform one or more operations at the headset device 1102 to transmit audio data corresponding to a user's speech to a second device (not shown) for further processing.

[0167] Figure 12An implementation 1200 of a wearable electronic device 1202, illustrated as a "smartwatch," is depicted, comprising device 104. A conference manager 162, an interrupt manager 164, a GUI generator 168, a speaker 154, a microphone 1210, or a combination thereof, are integrated into the wearable electronic device 1202. In a particular example, the conference manager 162 outputs a voice audio stream 111, or the interrupt manager 164 outputs a synthesized voice audio stream 133, which is then processed to perform one or more operations at the wearable electronic device 1202, such as launching a GUI 145 or otherwise displaying additional information associated with the user's voice on a display screen 1204 of the wearable electronic device 1202. For illustration, the wearable electronic device 1202 may include a display screen configured to display notifications based on user voice detected by the wearable electronic device 1202. In a particular example, the wearable electronic device 1202 includes a haptic device that provides haptic notifications (e.g., vibration) in response to the detection of user voice. For example, haptic notifications can allow a user to view a notification displayed on the wearable electronic device 1202 indicating that a keyword spoken by the user has been detected. The wearable electronic device 1202 can thus alert users with hearing impairments or those wearing headphones that their voice has been detected.

[0168] Figure 13 This is an implementation 1300 of device 104 including a wireless speaker and voice-activated device 1302. The wireless speaker and voice-activated device 1302 may have wireless network connectivity and be configured to perform auxiliary operations. One or more processors 160, including a conference manager 162, an interrupt manager 164 or both, a speaker 154, a microphone 1310 or a combination thereof, are included in the wireless speaker and voice-activated device 1302. During operation, in response to a spoken command identified as a user's voice received in the voice audio stream 111 output by the conference manager 162 or the synthesized voice audio stream 133 output by the interrupt manager 164, the wireless speaker and voice-activated device 1302 may perform an assistant operation, such as via the operation of a voice-activated system (e.g., an integrated assistant application). Auxiliary operations may include creating calendar events, adjusting the temperature, playing music, turning on lights, etc. For example, an assistant operation is performed in response to receiving a command following a keyword or key phrase (e.g., "Hello Assistant").

[0169] Figure 14Implementation 1400 is depicted, wherein device 104 includes a portable electronic device corresponding to camera device 1402. A conference manager 162, an interrupt manager 164, a GUI generator 168, a speaker 154, a microphone 1410, or a combination thereof, are included in camera device 1402. During operation, as an illustrative example, in response to a verbal command identified as user speech received in the speech audio stream 111 output by conference manager 162 or the synthesized speech audio stream 133 output by interrupt manager 164, camera device 1402 can perform operations in response to the verbal user command, such as adjusting image or video capture settings, image or video playback settings, or image or video capture instructions.

[0170] Figure 15 An implementation 1500 is depicted, wherein device 104 includes a portable electronic device corresponding to a virtual reality, augmented reality, or mixed reality headset 1502. A conference manager 162, an interrupt manager 164, a GUI generator 168, a speaker 154, a microphone 1510, or a combination thereof, are integrated into the headset 1502. User voice detection can be performed based on the speech audio stream 111 output by the conference manager 162 or the synthesized speech audio stream 133 output by the interrupt manager 164. A visual interface device is positioned in front of the user to display augmented reality or virtual reality images or scenes to the user when the headset 1502 is worn. In a particular example, the visual interface device is configured to display a notification indicating user voice detected in the audio stream. In another example, the visual interface device is configured to display a GUI 145.

[0171] Figure 16 An implementation 1600 is depicted, in which device 104 corresponds to or is integrated within a vehicle 1602, illustrated as a manned or unmanned aerial device (e.g., a package delivery drone). A conference manager 162, an interrupt manager 164, a GUI generator 168, a speaker 154, a microphone 1610, or a combination thereof, are integrated into the vehicle 1602. User voice detection can be performed based on the speech audio stream 111 output by the conference manager 162 or the synthesized speech audio stream 133 output by the interrupt manager 164, such as delivery instructions for an authorized user of the vehicle 1602.

[0172] Figure 17Another implementation 1700 is depicted, in which device 104 corresponds to or is integrated within a vehicle 1702, illustrated as an automobile. Vehicle 1702 includes one or more processors 160, which include a conference manager 162, an interrupt manager 164, a GUI generator 168, or a combination thereof. Vehicle 1702 also includes a speaker 154, a microphone 1710, or both. User voice detection can be performed based on the speech audio stream 111 output by conference manager 162 or the synthesized speech audio stream 133 output by interrupt manager 164. For example, user voice detection can be used to detect voice commands (e.g., starting the engine or heating) from an authorized user of vehicle 1702. In a particular implementation, in response to a spoken command recognized as a user's voice received in the voice audio stream 111 output by the conference manager 162 or the synthesized voice audio stream 133 output by the interrupt manager 164, the voice activation system of the vehicle 1702 initiates one or more operations of the vehicle 1702 based on one or more keywords detected in the voice audio stream 111 or the synthesized voice audio stream 133 (e.g., "unlock", "start engine", "play music", "display weather forecast", or another voice command), such as providing feedback or information via the display 1720 or one or more speakers (e.g., speaker 154). In a particular implementation, the GUI generator 168 provides the display 1720 with information about an online meeting (e.g., a call). For example, the GUI generator 168 provides the display 1720 with a GUI 145.

[0173] refer to Figure 18 This describes a block diagram depicting a specific illustrative implementation of the device and generally designates it as 1800. In various implementations, device 1800 may have... Figure 18 The number of components may be more or less. In an illustrative implementation, device 1800 may correspond to device 104. In an illustrative implementation, device 1800 may perform the reference... Figure 1-17 One or more operations described.

[0174] In a particular implementation, device 1800 includes a processor 1806 (e.g., a central processing unit (CPU)). Device 1800 may include one or more additional processors 1810 (e.g., one or more DSPs). In a particular aspect, Figure 1One or more processors 160 correspond to processors 1806, 1810, or combinations thereof. Processor 1810 may include a speech and music encoder-decoder (CODEC, codec) 1808, which includes a speech codec (“vocoder”) encoder 1836, a vocoder decoder 1838, a conference manager 162, an interrupt manager 164, a GUI generator 168, or combinations thereof. In certain aspects, Figure 1 One or more processors 160 include processor 1806, processor 1810, or a combination thereof.

[0175] Device 1800 may include memory 1886 and CODEC 1334. Memory 1886 may include instructions 1856 that can be executed by one or more additional processors 1810 (or processor 1806) to implement the functions described by reference to conference manager 162, interrupt manager 164, GUI generator 168, or combinations thereof. In a particular aspect, memory 1886 stores program data 1858 used or generated by conference manager 162, interrupt manager 164, GUI generator 168, or combinations thereof. In a particular aspect, memory 1886 includes... Figure 1 The memory 132. The device 1800 may include a modem 1840 coupled to the antenna 1842 via a transceiver 1850.

[0176] Device 1800 may include display device 156 coupled to display controller 1826. Speaker 154 and one or more microphones 1832 may be coupled to CODEC 1334. CODEC 1834 may include digital-to-analog converter (DAC) 1802, analog-to-digital converter (ADC) 1804, or both. In a particular implementation, CODEC 1834 may receive analog signals from one or more microphones 1832, convert the analog signals to digital signals using ADC 1804, and provide the digital signals to voice and music codec 1808. Voice and music codec 1808 may process digital signals, and the digital signals may also be processed by conference manager 162, interrupt manager 164, or both. In a particular implementation, voice and music codec 1808 may provide digital signals to CODEC 1334. CODEC 1834 may use ADC 1802 to convert the digital signals to analog signals and may provide the analog signals to speaker 154.

[0177] In a particular implementation, device 1800 may be included in a system-in-package or system-on-a-chip device 1822. In a particular implementation, memory 1886, processor 1806, processor 1810, display controller 1826, CODEC 1334, modem 1840, and transceiver 1850 are included in a system-in-package or system-on-a-chip device 1822. In a particular implementation, input device 1830 and power supply 1844 are coupled to system-on-a-chip device 1822. Furthermore, in a particular implementation, such as Figure 18 As illustrated, display device 156, input device 1830, speaker 154, one or more microphones 1832, antenna 1842, and power supply 1844 are external to system-on-chip device 1822. In a particular implementation, each of display device 156, input device 1830, speaker 154, one or more microphones 1832, antenna 1842, and power supply 1844 may be coupled to components of system-on-chip device 1822, such as interfaces or controllers.

[0178] Device 1800 may include virtual assistants, home appliances, smart devices, Internet of Things (IoT) devices, communication devices, headsets, vehicles, computers, display devices, televisions, game consoles, music players, radios, video players, entertainment units, personal media players, digital video players, cameras, navigation devices, smart speakers, speaker sticks, mobile communication devices, smartphones, cellular phones, laptops, tablets, personal digital assistants, digital video disc (DVD) players, tuners, augmented reality headsets, virtual reality headsets, aircraft, home automation systems, voice-activated devices, wireless speakers and voice-activated devices, portable electronic devices, automobiles, computing devices, virtual reality (VR) devices, base stations, mobile devices, or any combination thereof.

[0179] In accordance with the described implementation, the apparatus includes components for receiving an audio stream representing the voice of a first user during an online meeting. For example, the components for receiving the audio stream may correspond to a meeting manager 162, an interrupt manager 164, one or more processors 160, or a device 104. Figure 1 System 100, Conference Manager 122, Server 204 Figure 2 The system 200, one or more processors 1810, processor 1806, voice and music codecs 1808, modem 1840, transceiver 1850, antenna 1842, device 1800, and one or more other circuits or components or any combination thereof configured to receive voice audio streams during online meetings.

[0180] The device also includes components for receiving a text stream representing the voice of a first user. For example, the components for receiving the text stream may correspond to a conference manager 162, an interrupt manager 164, a text-to-speech converter 166, one or more processors 160, or a device 104. Figure 1 System 100, Conference Manager 122, Interrupt Manager 124, Server 204 Figure 2 The system 200, one or more processors 1810, processor 1806, voice and music codecs 1808, modem 1840, transceiver 1850, antenna 1842, device 1800, one or more other circuits or components or any combination thereof configured to receive a text stream.

[0181] The apparatus also includes components for selectively generating output based on a text stream in response to an interruption in the speech audio stream. For example, the components for selectively generating output may correspond to an interrupt manager 164, a text-to-speech converter 166, a GUI generator 168, one or more processors 160, or a device 104. Figure 1 System 100, Interrupt Manager 124, Server 204 Figure 2 The system 200, one or more processors 1810, processor 1806, voice and music codecs 1808, device 1800, one or more other circuits or components or any combination thereof configured to selectively generate output.

[0182] In some implementations, a non-transitory computer-readable medium (e.g., a computer-readable storage device, such as memory 1886) includes instructions (e.g., instruction 1856) that, when executed by one or more processors (e.g., one or more processors 1810 or processor 1806), cause the one or more processors to receive, during an online meeting, a voice audio stream (e.g., voice audio stream 111) representing a first user (e.g., user 142). When executed by the one or more processors, these instructions also cause the one or more processors to receive a text stream (e.g., text stream 121) representing the voice of the first user (e.g., user 142). When executed by the one or more processors, the instructions also cause the one or more processors to selectively generate output based on the text stream in response to an interruption in the voice audio stream (e.g., synthesized voice audio stream 133, annotated text stream 137, or both).

[0183] The following specific aspects of this disclosure are described in the first set of relevant clauses:

[0184] According to Clause 1, an apparatus for communication includes: one or more processors configured to: receive an audio stream representing the speech of a first user during an online meeting; receive a text stream representing the speech of the first user; and selectively generate output based on the text stream in response to an interruption in the audio stream.

[0185] Clause 2 includes the device of Clause 1, wherein the one or more processors are configured to detect the interruption in response to determining that no audio frame of the voice audio stream has been received within a threshold duration of the last received audio frame of the voice audio stream.

[0186] Clause 3 includes the device of Clause 1, wherein the one or more processors are configured to detect the interruption in response to receiving the text stream.

[0187] Clause 4 includes the device of Clause 1, wherein the one or more processors are configured to detect an interrupt in response to receiving an interrupt notification.

[0188] Clause 5 includes a device as described in any of Clauses 1 to 4, wherein the one or more processors are configured to provide the text stream as output to a display.

[0189] Clause 6 includes the device of any one of Clauses 1 to 5, wherein the one or more processors are further configured to: receive a metadata stream indicating the intonation of the first user's voice; and annotate the text stream based on the metadata stream.

[0190] Clause 7 includes the device of any one of Clauses 1 to 6, wherein the one or more processors are further configured to: perform text-to-speech conversion on the text stream to generate a synthesized speech audio stream; and provide the synthesized speech audio stream as output to a speaker.

[0191] Clause 8 includes the device of Clause 7, wherein the one or more processors are further configured to receive a metadata stream indicative of the intonation of the first user’s speech, wherein the text-to-speech conversion is based on the metadata stream.

[0192] Clause 9 includes the device of Clause 7, wherein the one or more processors are further configured to display the virtual avatar while providing the synthesized speech audio stream to the speaker.

[0193] Clause 10 includes the device of Clause 9, wherein the one or more processors are configured to receive media streams during an online meeting, the media streams including a first user’s voice audio stream and a video stream.

[0194] Clause 11 includes the device of Clause 10, wherein the one or more processors are configured to, in response to the interrupt, stop playback of the audio stream and stop playback of the video stream.

[0195] Clause 12 includes the device of Clause 10, wherein the one or more processors are configured to, in response to the termination of the interrupt,: prevent the supply of the synthesized speech audio stream to the speaker; prevent the display of the virtual avatar; resume playback of the video stream; and resume playback of the speech audio stream.

[0196] Clause 13 includes the devices in Clause 7, wherein text-to-speech conversion is performed based on a speech model.

[0197] Clause 14 includes the device of Clause 13, wherein the speech model corresponds to the general speech model.

[0198] Clause 15 includes the device of Clause 13 or Clause 14, wherein the one or more processors are configured to update the speech model based on the speech audio stream prior to the interruption.

[0199] Clause 16 includes a device of any one of Clauses 1 to 15, wherein the one or more processors are configured to: receive a second voice audio stream representing the voice of a second user during the online meeting; and provide the second voice audio stream to a speaker while generating the output.

[0200] Clause 17 includes a device of any one of Clauses 1 to 16, wherein the one or more processors are configured to: stop playback of the voice audio stream in response to an interruption in the voice audio stream; and in response to the end of the interruption: avoid generating the output based on the text stream; and resume playback of the voice audio stream.

[0201] The following specific aspects of this disclosure are described in the second set of relevant clauses:

[0202] According to Clause 18, a communication method includes: receiving, at a device, a voice audio stream representing the voice of a first user during an online meeting; receiving, at the device, a text stream representing the voice of the first user; and, in response to an interruption in the voice audio stream, selectively generating output at the device based on the text stream.

[0203] Clause 19 includes the method of Clause 18, and further includes detecting an interruption in response to determining that no audio frame of the speech audio stream has been received within a threshold duration of the last received audio frame of the speech audio stream.

[0204] Clause 20 includes the methods of Clause 18, and also includes detecting an interruption in response to receiving the text stream.

[0205] Clause 21 includes the methods of Clause 18, and also includes detecting an interruption in response to receiving an interruption notification.

[0206] Clause 22 includes the methods described in any of Clauses 18 to 21, and also includes providing the text stream as output to a display.

[0207] Clause 23 includes the method of any of Clauses 18 to 22, and further includes: receiving a metadata stream indicating the tone of the first user's voice; and annotating the text stream based on the metadata stream.

[0208] Specific aspects of this disclosure are described in the relevant provisions of this disclosure:

[0209] According to Clause 24, a non-transitory computer-readable storage medium stores instructions that, when executed by one or more processors, cause the one or more processors to: receive an audio stream representing the speech of a first user during an online meeting; receive a text stream representing the speech of the first user; and, in response to an interrupt in the audio stream, selectively generate output based on the text stream.

[0210] Clause 25 includes the non-transitory computer-readable storage medium of Clause 24, wherein the instructions, when executed by one or more processors, cause one or more processors to: perform a text-to-speech conversion on the text stream to generate a synthesized speech audio stream; and provide the synthesized speech audio stream as output to a speaker.

[0211] Clause 26 includes the non-transitory computer-readable storage medium of Clause 25, wherein the instructions, when executed by one or more processors, cause one or more processors to receive a metadata stream indicative of the intonation of a first user's speech, wherein the text-to-speech conversion is based on the metadata stream.

[0212] Clause 27 includes the non-transitory computer-readable storage medium of Clause 25 or Clause 26, wherein the instructions, when executed by one or more processors, cause one or more processors to display the virtual avatar while providing a stream of synthesized speech audio to a speaker.

[0213] Clause 28 includes a non-transitory computer-readable storage medium of any of Clauses 25 to 27, wherein the instructions, when executed by one or more processors, cause one or more processors to update a speech model based on a speech audio stream prior to an interruption, and wherein the text-to-speech conversion is performed based on the speech model.

[0214] Specific aspects of this disclosure are described below in the fourth set of relevant clauses:

[0215] According to Clause 29, an apparatus includes: components for receiving an audio stream of speech representing the speech of a first user during an online meeting; components for receiving a text stream representing the speech of the first user; and components for selectively generating output based on the text stream in response to an interruption in the audio stream.

[0216] Clause 30 includes the apparatus of Clause 29, wherein the components for receiving a voice audio stream, the components for receiving a text stream, and the components for selectively generating output are integrated into at least one of a virtual assistant, home appliance, smart device, Internet of Things (IoT) device, communication device, headset, vehicle, computer, display device, television, game console, music player, radio, video player, entertainment unit, personal media player, digital video player, camera, or navigation device.

[0217] Those skilled in the art will further understand that the various illustrative logic blocks, configurations, modules, circuits, and algorithmic steps described in conjunction with the implementations disclosed herein can be implemented as electronic hardware, computer software executed by a processor, or a combination of both. The various illustrative components, blocks, configurations, modules, circuits, and steps have been described above in terms of functionality. Whether this functionality is implemented as hardware or processor-executable instructions depends on the specific application and the design constraints imposed on the system as a whole. Those skilled in the art can implement the described functionality in different ways for each specific application, and these implementation decisions should not be construed as departing from the scope of this disclosure.

[0218] The steps of the methods or algorithms described in conjunction with the implementations disclosed herein can be implemented directly in hardware, as a software module executed by a processor, or as a combination of both. The software module can reside in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, optical disc read-only memory (CD-ROM), or any other form of non-transitory storage medium known in the art. An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Alternatively, the storage medium can be integrated with the processor. The processor and storage medium can reside in an application-specific integrated circuit (ASIC). The ASIC can reside in a computing device or user terminal. Alternatively, the processor and storage medium can reside as discrete components in a computing device or user terminal.

[0219] The prior description of the disclosed aspects is provided to enable those skilled in the art to make or use the disclosed aspects. Various modifications to these aspects will readily be apparent to those skilled in the art, and the principles defined herein can be applied to other aspects without departing from the scope of the invention. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but should be given the broadest possible scope consistent with the principles and novel features as defined by the appended claims.

Claims

1. A device for communication, comprising: One or more processors are configured as follows: During an online meeting, a voice audio stream representing the voice of the first user is received from a device corresponding to the first user, based on audio input; Receive a text stream representing the voice of the first user based on the audio input from the device corresponding to the first user; Receive metadata stream based on the audio input from the device corresponding to the first user; and In response to an interruption in the audio stream, output is selectively generated based on the text stream and the metadata stream.

2. The device according to claim 1, wherein, The one or more processors are configured to detect the interruption in response to determining that no audio frame of the voice audio stream has been received within a threshold duration of the last received audio frame of the voice audio stream.

3. The device according to claim 1, wherein, The one or more processors are configured to detect the interrupt in response to receiving the text stream.

4. The device according to claim 1, wherein, The one or more processors are configured to detect the interrupt in response to receiving an interrupt notification.

5. The device according to claim 1, wherein, The metadata stream indicates the tone of the first user's voice.

6. The device according to claim 1, wherein, The one or more processors are further configured to: The text stream is annotated based on the metadata stream to generate an annotated text stream; and The annotated text stream is provided to the display as the output.

7. The device according to claim 1, wherein, The one or more processors are further configured to: Perform text-to-speech conversion on the text stream and generate a synthesized speech audio stream based on the metadata stream; and The synthesized speech audio stream is provided as output to the speaker.

8. The device according to claim 7, wherein, The one or more processors are also configured to display a virtual avatar while providing the synthesized speech audio stream to the speaker.

9. The device according to claim 8, wherein, The virtual avatar includes a first representation based on performing the text-to-speech conversion using an untrained speech model, and wherein the virtual avatar includes a second representation different from the first representation based on performing the text-to-speech conversion using a trained speech model.

10. The device according to claim 9, wherein, The one or more processors are configured to receive media streams during the online meeting, the media streams including the first user's voice audio stream and video stream.

11. The device according to claim 10, wherein, The one or more processors are configured to respond to the interrupt: Stop the playback of the audio stream; and Stop the playback of the video stream.

12. The device according to claim 10, wherein, The one or more processors are configured to respond to the termination of the interrupt: Avoid providing the synthesized speech audio stream to the speaker; Avoid displaying the virtual avatar; Restore the playback of the video stream; and Resume playback of the audio stream.

13. The device according to claim 7, wherein, The text-to-speech conversion is performed based on a speech model.

14. The device according to claim 13, wherein, The speech model corresponds to the general speech model.

15. The device according to claim 13, wherein, The one or more processors are configured to update the speech model based on the speech audio stream prior to the interruption.

16. The device according to claim 1, wherein, The one or more processors are configured to: During the online meeting, a second voice audio stream representing the voice of a second user is received; and The second voice audio stream is provided to the speaker while the output is being generated.

17. The device according to claim 1, wherein, The one or more processors are configured to: The playback of the audio stream is stopped in response to an interruption in the audio stream; and The interrupt ends in response to the following: Avoid generating the output based on the text stream; and Resume playback of the audio stream.

18. A communication method, comprising: During an online meeting, a voice audio stream representing the voice of the first user is received from a device corresponding to the first user, based on audio input; Receive a text stream representing the voice of the first user based on the audio input from the device corresponding to the first user; Receive metadata stream based on the audio input from the device corresponding to the first user; and Output is selectively generated based on the text stream and the metadata stream in response to an interruption in the audio stream.

19. The method of claim 18, further comprising detecting the interruption in response to determining that no audio frame of the speech audio stream has been received within a threshold duration of the last received audio frame of the speech audio stream.

20. The method of claim 18, further comprising detecting the interruption in response to receiving the text stream.

21. The method of claim 18, further comprising detecting the interruption in response to receiving an interruption notification.

22. The method of claim 18, further comprising annotating the text stream based on the metadata stream to generate an annotated text stream.

23. The method of claim 22, further comprising providing an annotated text stream as said output to a display.

24. A non-transitory computer-readable storage medium that stores instructions, which, when executed by one or more processors, cause the one or more processors to: During an online meeting, a voice audio stream representing the voice of the first user is received from a device corresponding to the first user, based on audio input; Receive a text stream representing the voice of the first user based on the audio input from the device corresponding to the first user; Receive metadata stream based on the audio input from the device corresponding to the first user; and In response to an interruption in the audio stream, output is selectively generated based on the text stream and the metadata stream.

25. The non-transitory computer-readable storage medium according to claim 24, wherein, The instruction, when executed by the one or more processors, causes the one or more processors to: Perform text-to-speech conversion on the text stream and generate a synthesized speech audio stream based on the metadata stream; and The synthesized speech audio stream is provided as output to the speaker.

26. The non-transitory computer-readable storage medium according to claim 25, wherein, When the instruction is executed by the one or more processors, it causes the one or more processors to display the virtual avatar while providing the synthesized speech audio stream to the speaker.

27. The non-transitory computer-readable storage medium according to claim 26, wherein, The virtual avatar includes a first representation based on performing the text-to-speech conversion using an untrained speech model, and wherein the virtual avatar includes a second representation different from the first representation based on performing the text-to-speech conversion using a trained speech model.

28. The non-transitory computer-readable storage medium according to claim 25, wherein, The instructions, when executed by the one or more processors, cause the one or more processors to update the speech model based on the speech audio stream prior to the interrupt, wherein the text-to-speech conversion is performed based on the speech model.

29. An apparatus comprising: Components for receiving a voice audio stream from a device corresponding to a first user during an online meeting, the voice audio stream representing the first user's voice based on audio input; A component for receiving, from the device corresponding to the first user, a text stream representing the voice of the first user based on the audio input; A component for receiving a metadata stream based on the audio input from the device corresponding to the first user; and A component for selectively generating output based on the text stream and the metadata stream in response to an interruption in the audio stream.

30. The apparatus according to claim 29, wherein, The components for receiving voice audio streams, the components for receiving text streams, the components for receiving metadata streams, and the components for selectively generating outputs are integrated into virtual assistants, home appliances, smart devices, Internet of Things (IoT) devices, communication devices, headsets, vehicles, computers, display devices, televisions, game consoles, music players, radios, video players, entertainment units, personal media players, digital video players, cameras, or navigation devices.