Systems and methods for handling voice audio stream interruptions

The system addresses frame loss in online conferences by generating a text stream from voice audio and using text-to-speech conversion to maintain communication continuity, improving user experience.

JP7798901B2Active Publication Date: 2026-01-14QUALCOMM INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023546311
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-02-03
Filing Date
2021-12-09
Publication Date
2026-01-14
Estimated Expiration
2041-12-09

AI Technical Summary

Technical Problem

Network issues during online conferences cause frame loss, leading to irrecoverable information loss and negatively impacting user experience, as users must guess or ask for repetitions, disrupting the conversation flow.

Method used

A system that generates a text stream through speech-to-text conversion on the voice audio stream and selectively outputs this text, along with metadata, to ensure continuous communication during interruptions, using text-to-speech conversion and avatars to replace interrupted audio.

Benefits of technology

Ensures uninterrupted communication by providing a text-based alternative to audio streams, reducing misunderstandings and time wastage due to frame loss, and enhancing user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007798901000001
    Figure 0007798901000001
  • Figure 0007798901000002
    Figure 0007798901000002
  • Figure 0007798901000003
    Figure 0007798901000003
Patent Text Reader

Abstract

The device for communications includes one or more processors configured to receive a voice audio stream representing a speech of a first user during an online conference. The one or more processors are also configured to receive a text stream representing the speech of the first user. The one or more processors are further configured to selectively generate an output based on the text stream in response to an interruption in the voice audio stream.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Priority claims This application claims the benefit of priority to commonly owned U.S. Non-Provisional Patent Application No. 17 / 166,250, filed February 3, 2021, the entire contents of which are expressly incorporated herein by reference.

[0002] TECHNICAL FIELD This disclosure relates generally to systems and methods for handling voice audio stream interruptions. [Background technology]

[0003] Advances in technology have resulted in smaller and more powerful computing devices. For example, there are now a variety of portable personal computing devices, including wireless telephones such as mobile phones and smartphones, tablet and laptop computers, that are small, lightweight, and easily carried by users. These devices can communicate voice and data packets over wireless networks. Furthermore, many such devices incorporate additional functionality, such as digital still cameras, digital video cameras, digital recorders, and audio file players. Such devices can also process executable instructions, including software applications, such as web browser applications that may be used to access the Internet. Thus, these devices can contain significant computing power.

[0004] Such computing devices often incorporate the capability to receive audio signals from one or more microphones. For example, the audio signals may represent a user's voice captured by the microphone, external sounds captured by the microphone, or a combination thereof. Such devices may include communication devices used for online conferences or calls. Network issues during an online conference between a first user and a second user may cause frame loss, such that some audio and video frames sent by the first user's first device are not received by the second user's second device. Frame loss due to network issues may lead to irrecoverable information loss during the online conference. For example, the second user may have to guess what they missed or ask the first user to repeat what they missed, which negatively impacts the user experience. Summary of the Invention [Means for solving the problem]

[0005] According to one embodiment of the present disclosure, a communication device includes one or more processors configured to receive an audio stream representing a speech of a first user during an online conference. The one or more processors are also configured to receive a text stream representing the speech of the first user. The one or more processors are further configured to selectively generate an output based on the text stream in response to an interruption of the audio stream.

[0006] According to another implementation of the present disclosure, a method of communication includes receiving, at a device, a voice audio stream representing a voice of a first user during an online conference. The method also includes receiving, at the device, a text stream representing the voice of the first user. The method further includes, at the device, selectively generating, in response to a break in the voice audio stream, an output based on the text stream.

[0007] According to another implementation of the present disclosure, a non-transitory computer-readable medium includes instructions that, when executed by one or more processors, cause the one or more processors to receive an audio stream representing a voice of a first user during an online conference. The instructions, when executed by the one or more processors, also cause the one or more processors to receive a text stream representing the voice of the first user. The instructions, when executed by the one or more processors, further cause the one or more processors to selectively generate an output based on the text stream in response to a break in the audio stream.

[0008] According to another implementation of the present disclosure, an apparatus includes means for receiving an audio stream during an online conference, the audio stream representing a speech of a first user. The apparatus also includes means for receiving a text stream representing the speech of the first user. The apparatus further includes means for selectively generating an output based on the text stream in response to an interruption of the audio stream.

[0009] Other aspects, advantages, and features of the present disclosure will become apparent after review of the entire disclosure, including the following sections: Brief Description of the Drawings, Detailed Description, and Claims. [Brief explanation of the drawings]

[0010] [Figure 1]1 is a block diagram of a particular illustrative aspect of a system operable to handle voice audio stream interruptions, in accordance with certain examples of the present disclosure. [Figure 2] FIG. 1 is a diagram of an example aspect of a system operable to handle voice audio stream interruptions, according to some examples of the present disclosure. [Figure 3A] 3 is a diagram of an exemplary graphical user interface (GUI) generated by the system of FIG. 1 or the system of FIG. 2, according to some examples of the present disclosure. [Figure 3B] 3A-3C are diagrams of exemplary GUIs generated by the system of FIG. 1 or the system of FIG. 2, according to some examples of the present disclosure. [Figure 3C] 3A-3C are diagrams of exemplary GUIs generated by the system of FIG. 1 or the system of FIG. 2, according to some examples of the present disclosure. [Figure 4A] 3A-3C are diagrams of exemplary aspects of operation of the system of FIG. 1 or the system of FIG. 2, according to some examples of the present disclosure. [Figure 4B] 3A-3C are diagrams of exemplary aspects of operation of the system of FIG. 1 or the system of FIG. 2, according to some examples of the present disclosure. [Figure 5] FIG. 1 is a diagram of an example aspect of a system operable to handle voice audio stream interruptions, according to some examples of the present disclosure. [Figure 6A] 6 is a diagram of an exemplary graphical user interface (GUI) generated by the system of FIG. 5, according to some examples of the present disclosure. [Figure 6B] 6 is a diagram of an example GUI generated by the system of FIG. 5, according to some examples of the present disclosure. [Figure 6C] 6 is a diagram of an example GUI generated by the system of FIG. 5, according to some examples of the present disclosure. [Figure 7A] 6 is a diagram of an exemplary aspect of the operation of the system of FIG. 5, according to some examples of the present disclosure. [Figure 7B] 6 is a diagram of an exemplary aspect of the operation of the system of FIG. 5, according to some examples of the present disclosure. [Figure 8] 6 is a diagram of a specific implementation of a method for handling voice audio stream interruptions that may be performed by any of the systems of FIG. 1, FIG. 2, or FIG. 5, according to some examples of the present disclosure. [Figure 9] FIG. 1 illustrates an example of an integrated circuit operable to handle voice audio stream interruptions, according to some examples of the present disclosure. [Figure 10] FIG. 1 is a diagram of a mobile device operable to handle voice audio stream interruptions, according to some examples of the present disclosure. [Figure 11] FIG. 1 is a diagram of a headset operable to handle voice audio stream interruptions, according to some examples of the present disclosure. [Figure 12] FIG. 1 is a diagram of a wearable electronic device operable to handle voice audio stream interruptions, according to some examples of the present disclosure. [Figure 13] FIG. 1 is a diagram of a voice-controlled speaker system operable to handle voice audio stream interruptions, according to some examples of the present disclosure. [Figure 14] 1 is a diagram of a camera operable to handle voice audio stream interruptions, according to some examples of the present disclosure. [Figure 15] FIG. 1 is a diagram of a headset, such as a virtual reality headset or an augmented reality headset, operable to handle voice audio stream interruptions, in accordance with some examples of the present disclosure. [Figure 16] 1 is a diagram of a first example of a vehicle operable to handle voice audio stream interruptions, according to some examples of the present disclosure. [Figure 17] FIG. 10 is a diagram of a second example vehicle operable to handle voice audio stream interruptions, according to some examples of the present disclosure. [Figure 18] FIG. 1 is a block diagram of a particular illustrative example of a device operable to handle voice audio stream interruptions, in accordance with some examples of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0011] Missing portions of an online conference or call can negatively impact a user experience. For example, during an online conference between a first user and a second user, if some audio frames sent by the first user's first device are not received by the second user's second device, the second user may miss portions of the first user's speech. The second user must either guess what the first user said or ask the first user to repeat what they missed. This can lead to misunderstandings, disrupt the flow of conversation, and waste time.

[0012] A system and method for handling voice audio stream interruptions is disclosed. For example, each device includes a conference manager configured to establish an online conference or call between the device and one or more other devices. The interruption manager (in the device or in a server) is configured to handle voice audio stream interruptions.

[0013] During an online conference between a first device of a first user and a second device of a second user, a conference manager of the first device sends a media stream to the second device. The media stream includes a voice audio stream, a video stream, or both. The voice audio stream corresponds to the first user's voice during the conference.

[0014] A stream manager (at the first device or at the server) generates a text stream by performing speech-to-text conversion on the voice audio stream and forwards the text stream to the second device. In a first operating mode (e.g., a caption data transmission mode), the stream manager (e.g., a conference manager at the first device or at the server) forwards the text stream simultaneously with the media stream throughout the online conference. In an alternative example, in a second operating mode (e.g., an interruption manager at the first device or at the server), in response to detecting a network problem (e.g., low bandwidth, packet loss, etc.), forwards the text stream to the second device simultaneously with sending the media stream to the second device.

[0015] In some examples, the network problem causes an interruption in reception of the media stream at the second device without an interruption in reception of the text stream. In some examples, the second device, in a first operating mode (e.g., a caption data display mode), provides the text stream to a display regardless of detecting a network problem. In other examples, the second device, in a second operating mode (e.g., an interrupted data display mode), displays the text stream in response to detecting an interruption in the media stream.

[0016] In certain examples, a stream manager (e.g., a conference manager or an interruption manager) transfers a metadata stream in addition to the text data. The metadata indicates emotion, intonation, and other attributes of the first user's voice. In certain examples, the second device displays the metadata stream in addition to the text stream. For example, the text stream is annotated based on the metadata stream.

[0017] In certain examples, the second device performs text-to-speech conversion on the text stream to generate a synthetic speech audio stream and outputs the synthetic speech audio stream (e.g., to replace the interrupted speech audio stream). In certain examples, the text-to-speech conversion is based at least in part on the metadata stream.

[0018] In a particular example, the second device displays an avatar (e.g., to replace an interrupted video stream) while outputting the synthesized speech audio stream. In a particular example, the text-to-speech conversion is based on a generic speech model. For example, a first generic speech model may be used for one user and a second generic speech model may be used for another user so that a listener can distinguish between speech corresponding to different users. In another particular example, the text-to-speech conversion is based on a user speech model generated based on the speech of the first user. In a particular example, the user speech model is generated prior to the online conference. In a particular example, the user speech model is generated (or updated) during the online conference. In a particular example, the user speech model is initialized from the generic speech model and updated based on the speech of the first user.

[0019] In a particular example, the avatar indicates that a voice model is being trained. For example, the avatar is initialized as red, indicating that a generic voice model is being used (or that a user voice model is not ready), and over time the avatar transitions from red to green, indicating that a voice model is being trained. A green avatar indicates that a user voice model has been trained (or that a user voice model is ready).

[0020] An online conference can be between two or more users. In a situation where a first device is experiencing network problems but a third device of a third user in the online conference is not experiencing network problems, the second device can output a synthesized voice audio stream for the first user while simultaneously outputting a second media stream received from the third device corresponding to the third user's audio, video, or both.

[0021] Certain aspects of the present disclosure are described below with reference to the drawings. In this description, common features are designated by common reference numerals. Various terms used herein are used only for the purpose of describing particular implementations and are not intended to limit the implementations. For example, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise. Furthermore, some features described herein are singular in some implementations and plural in other implementations. To illustrate, FIG. 1 illustrates a device 104 that includes one or more processors ("processor(s)" 160 in FIG. 1), indicating that in some implementations the device 104 includes a single processor 160 and in other implementations the device 104 includes multiple processors 160.

[0022] As used herein, the terms “comprise,” “comprises,” and “comprising” may be used interchangeably with “include,” “includes,” or “including.” Additionally, the term “wherein” may be used interchangeably with “where.” As used herein, “exemplary” indicates an example, an implementation, and / or an aspect and should not be construed as limiting or indicating a preferred or preferred implementation. As used herein, ordinal terms (e.g., “first,” “second,” “third,” etc.) used to modify an element such as a structure, component, or operation do not in themselves indicate any priority or order of that element relative to another element, but rather merely distinguish that element from another element having the same name (apart from the use of ordinal terms). As used herein, the term "set" refers to one or more of a particular element, and the term "plurality" refers to a plurality (e.g., two or more) of a particular element.

[0023] As used herein, “coupled” may include “communicatively coupled,” “electrically coupled,” or “physically coupled,” and also (or alternatively) may include any combination thereof. Two devices (or components) may be directly or indirectly coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) via one or more other devices, components, wires, buses, networks (e.g., wired networks, wireless networks, or combinations thereof), etc. Two devices (or components) that are electrically coupled may be included in the same device or in different devices and may be connected via electronic circuits, one or more connectors, or inductive coupling, as illustrative and non-limiting examples. In some implementations, two devices (or components) that are communicatively coupled, such as in electrical communication, may send and receive signals (e.g., digital or analog signals) directly or indirectly via one or more wires, buses, networks, etc. As used herein, "directly coupled" may include two devices coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) with no intervening components.

[0024] In this disclosure, terms such as “determining,” “calculating,” “estimating,” “shifting,” “adjusting,” and the like may be used to describe how one or more operations are performed. It should be noted that such terms should not be construed as limiting, and other techniques may be utilized to perform similar operations. Additionally, the terms “generating,” “calculating,” “estimating,” “using,” “selecting,” “accessing,” and “determining” referred to herein may be used interchangeably. For example, “generating,” “calculating,” “estimating,” or “determining” a parameter (or signal) may refer to actively generating, estimating, calculating, or determining a parameter (or signal), or may refer to using, selecting, or accessing a parameter (or signal) that has already been generated by another component, device, or the like.

[0025] 1, a particular exemplary embodiment of a system configured to handle voice audio stream interruptions is disclosed, generally designated 100. System 100 includes a device 102 coupled to a device 104 via a network 106. Network 106 includes a wired network, a wireless network, or both. Device 102 is coupled to a camera 150, a microphone 152, or both. Device 104 is coupled to a speaker 154, a display device 156, or both.

[0026] The device 104 includes one or more processors 160 coupled to the memory 132. The one or more processors 160 include a conference manager 162 coupled to an interruption manager 164. The conference manager 162 and the interruption manager 164 are coupled to a graphical user interface (GUI) generator 168. The interruption manager 164 includes a text-to-speech converter 166. The device 102 includes one or more processors 120 including a conference manager 122 coupled to an interruption manager 124. The conference manager 122 and the conference manager 162 are configured to establish online conferences (e.g., audio calls, video calls, conference calls, etc.). In a particular example, the conference manager 122 and the conference manager 162 correspond to clients of a communication application (e.g., an online conferencing application). The interruption manager 124 and the interruption manager 164 are configured to process voice-audio interruptions.

[0027] In some implementations, conference manager 122 and conference manager 162 are blind to (e.g., unaware of) any voice audio interruptions managed by interrupt manager 124 and interrupt manager 164. In some implementations, conference manager 122 and conference manager 162 correspond to an upper layer (e.g., application layer) of a network protocol stack (e.g., Open Systems Interconnection (OSI) model) of device 102 and device 104, respectively. In some implementations, interrupt manager 124 and interrupt manager 164 correspond to a lower level (e.g., transport layer) of a network protocol stack of device 102 and device 104, respectively.

[0028] In some implementations, device 102, device 104, or both correspond to or are included in various types of devices. In an illustrative example, one or more processors 120, one or more processors 160, or a combination thereof are integrated into a headset device, such as that further described with reference to FIG. 11. In other examples, one or more processors 120, one or more processors 160, or a combination thereof are integrated into at least one of a mobile phone or tablet computing device as described with reference to FIG. 10, a wearable electronic device as described with reference to FIG. 12, a voice-controlled speaker system as described with reference to FIG. 13, a camera device as described with reference to FIG. 14, or a virtual reality headset, an augmented reality headset, or a mixed reality headset as described with reference to FIG. 15. In another illustrative example, one or more processors 120, one or more processors 160, or a combination thereof are integrated into a vehicle, such as that further described with reference to FIGS. 16 and 17.

[0029] During operation, conference manager 122 and conference manager 162 establish an online conference (e.g., an audio call, a video call, a conference call, or a combination thereof) between device 102 and device 104. For example, the online conference is between user 142 of device 102 and user 144 of device 104. Microphone 152 captures the voice of user 142 while user 142 is speaking and provides audio input 153 representing the voice to device 102. In certain aspects, camera 150 (e.g., a still camera, a video camera, or both) captures one or more images (e.g., still images or videos) of user 142 and provides video input 151 representing the one or more images to device 102. In certain aspects, camera 150 provides video input 151 to device 102 at the same time that microphone 152 provides audio input 153 to device 102.

[0030] Conference manager 122 generates media stream 109 of media frames based on audio input 153, video input 151, or both. For example, media stream 109 includes voice audio stream 111, video stream 113, or both. In certain aspects, conference manager 122 sends media stream 109 to device 104 in real time over network 106. For example, conference manager 122 generates media frames for media stream 109 as video input 151, audio input 153, or both are being received, and sends (e.g., begins transmitting) media stream 109 of media frames once the media frames are generated.

[0031] In particular implementations, conference manager 122 generates text stream 121, metadata stream 123, or both based on audio input 153 during a first operating mode (e.g., a caption data transmission mode) of device 102. For example, conference manager 122 performs speech-to-text conversion on audio input 153 to generate text stream 121. Text stream 121 indicates text corresponding to speech detected in audio input 153. In particular aspects, conference manager 122 performs speech intonation analysis on audio input 153 to generate metadata stream 123. For example, metadata stream 123 indicates the intonation (e.g., emotion, pitch, tone, or a combination thereof) of speech detected in audio input 153. In a first operating mode of device 102 (e.g., a caption data transmission mode), conference manager 122 sends text stream 121, metadata stream 123, or both (e.g., as subtitling data) along with media stream 109 to device 104 (e.g., regardless of network issues or voice audio interruptions). Alternatively, during a second operating mode of device 102 (e.g., a interrupted data transmission mode), conference manager 122 refrains from generating text stream 121 and metadata stream 123 in response to determining that no voice audio interruptions are detected.

[0032] Device 104 receives media stream 109 of media frames from device 102 over network 106. In particular implementations, device 104 receives a set (e.g., a burst) of media frames of media stream 109. In alternative implementations, device 104 receives one media frame of media stream 109 at a time. Conference manager 162 plays out the media frames of media stream 109. For example, conference manager 162 generates audio output 143 based on voice audio stream 111 and plays out audio output 143 through speaker 154 (e.g., as streaming audio content). In particular aspects, GUI generator 168 generates GUI 145 based on media stream 109, as further described with reference to FIG. 3A . For example, GUI generator 168 generates (or updates) GUI 145 to display video content of video stream 113 and provides GUI 145 to display device 156 (e.g., streaming video content). User 144 can view an image of user 142 on display device 156 while hearing the audio of user 142 through speaker 154 .

[0033] In particular implementations, conference manager 162 stores media frames of media stream 109 in a buffer prior to playout. For example, conference manager 162 adds a delay between receiving a media frame and playing the media frame at a first playout time to increase the likelihood that a subsequent media frame will be available in the buffer at the corresponding playout time (e.g., the second playout time). In particular aspects, conference manager 162 plays out media stream 109 in real time. For example, conference manager 162 retrieves media frames of media stream 109 from the buffer to play out audio output 143, video content of GUI 145, or both, while subsequent media frames of media stream 109 are being received (or are expected to be received) by device 104.

[0034] Conference manager 162 plays out text stream 121 along with media stream 109 (e.g., regardless of detecting an interruption in audio audio stream 111) in a first operating mode (e.g., caption data display mode) of device 104. In particular aspects, conference manager 162 receives text stream 121, metadata stream 123, or both along with media stream 109, for example, during a first operating mode (e.g., caption data transmission mode) of device 102. In alternative aspects, conference manager 162 does not receive text stream 121, metadata stream 123, or both, and generates text stream 121, metadata stream 123, or both, based on audio audio stream 111, video stream 113, or both, for example, during a second operating mode (e.g., interrupted data transmission mode) of device 102. For example, conference manager 162 performs speech-to-text conversion on speech audio stream 111 to generate text stream 121 and performs intonation analysis on speech audio stream 111 to generate metadata stream 123 .

[0035] During a first operational mode (e.g., a caption data display mode) of device 104, conference manager 162 provides text stream 121 as an output to display device 156. For example, conference manager 162 displays the text content of text stream 121 (e.g., as subtitles) using GUI 145, simultaneously with displaying the video content of video stream 113, providing audio output 143 to speaker 154, or both. Illustratively, conference manager 162 provides text stream 121 to GUI generator 168 simultaneously with providing video stream 113 to GUI generator 168. GUI generator 168 updates GUI 145 to display text stream 121, video stream 113, or both. GUI generator 168 provides updates of GUI 145 to display device 156 simultaneously with conference manager 162 providing voice audio stream 111 as audio output 143 to speaker 154.

[0036] In particular examples, conference manager 162 generates annotated text stream 137 based on text stream 121 and metadata stream 123. In particular aspects, conference manager 162 generates annotated text stream 137 by adding annotations to text stream 121 based on metadata stream 123. Conference manager 162 provides annotated text stream 137 as an output to display device 156. For example, conference manager 162 plays out annotated text stream 137 along with media stream 109. Illustratively, conference manager 162 displays the annotated text content of annotated text stream 137 using GUI 145 (e.g., as subtitles with intonation indication) while simultaneously displaying video content in video stream 113, providing audio output 143 to speaker 154, or both.

[0037] In particular implementations, conference manager 162 refrains from playing out text stream 121 (e.g., annotated text stream 137) in the second operating mode (e.g., interrupted data display mode or subtitle-disabled mode) of device 104. For example, conference manager 162 does not receive text stream 121 (e.g., during the second operating mode of device 102) and does not generate text stream 121 in the second operating mode (e.g., interrupted data display mode or subtitle-disabled mode). As another example, conference manager 162 receives text stream 121 and refrains from playing out text stream 121 (e.g., annotated text stream 137) in response to detecting the second operating mode (e.g., interrupted data display mode or subtitle-disabled mode) of device 104. In certain aspects, in a second operating mode (e.g., an interrupted data display mode) of the device 104, the interruption manager 164 refrains from playing out the text stream 121 (e.g., the annotated text stream 137) in response to determining that no interruption has been detected in the media stream 109 (e.g., a portion of the media stream 109 corresponding to the text stream 121 has been received).

[0038] In certain aspects, the interruption manager 164 initializes the voice model 131, such as an artificial neural network, based on a generic voice model prior to or near the start of the online conference. In certain aspects, the interruption manager 164 selects a generic voice model from a plurality of generic voice models based on a determination that the generic voice model matches (e.g., is associated with) demographic data of the user 142, such as the user's age, location, gender, or a combination thereof. In certain aspects, the interruption manager 164 predicts demographic data based on the user's 142 contact information (e.g., name, location, phone number, address, or a combination thereof) prior to the online conference (e.g., a scheduled conference). In certain aspects, the interruption manager 164 estimates demographic data based on the voice audio stream 111, the video stream 113, or both during the beginning portion of the online conference. For example, the interruption manager 164 analyzes the voice audio stream 111, the video stream 113, or both to estimate the user's 142 age, regional accent, gender, or a combination thereof. In certain aspects, the interruption manager 164 retrieves a (eg, previously generated) speech model 131 associated with the user 142 (eg, matching the user identifier of the user 142).

[0039] In particular aspects, the interruption manager 164 trains (e.g., generates or updates) the speech model 131 based on speech detected in the speech audio stream 111 during the online conference (e.g., prior to interrupting the speech audio stream 111). Illustratively, the text-to-speech converter 166 is configured to use the speech model 131 to perform text-to-speech conversion. In particular aspects, the interruption manager 164 receives (e.g., during a first operating mode of the device 102) or generates (e.g., during a second operating mode of the device 102) the text stream 121, the metadata stream 123, or both, corresponding to the speech audio stream 111. The text-to-speech converter 166 uses the speech model 131 to generate the synthetic speech audio stream 133 by performing text-to-speech conversion on the text stream 121, the metadata stream 123, or both. The interruption manager 164 uses training techniques to update the speech model 131 based on a comparison of the speech audio stream 111 and the synthetic speech audio stream 133. In illustrative examples where the speech model 131 includes an artificial neural network, the interruption manager 164 uses backpropagation to update the weights and biases of the speech model 131. According to some aspects, the speech model 131 is updated so that subsequent text-to-speech conversions using the speech model 131 are more likely to produce synthesized speech that closely matches the voice characteristics of the user 142.

[0040] In certain aspects, the interruption manager 164 generates an avatar 135 (e.g., a visual representation) of the user 142. In certain aspects, the avatar 135 includes or corresponds to a training indicator that indicates a level of training of the voice model 131, as further described with reference to FIGS. 3A-3C . For example, in response to determining that a first training criterion has not been met, the interruption manager 164 initializes the avatar 135 to a first visual representation that indicates that the voice model 131 has not been trained. During an online meeting, in response to determining that the first training criterion has been met and the second training criterion has not been met, the interruption manager 164 updates the avatar 135 from the first visual representation to a second visual representation that indicates that training of the voice model 131 is in progress. In response to determining that the second training criterion has been met, the interruption manager 164 updates the avatar 135 to a third visual representation that indicates that training of the voice model 131 has been completed.

[0041] The training criteria may be based on a count of audio samples used to train the speech model 131, a playback duration of the audio samples used to train the speech model 131, a coverage of the audio samples used to train the speech model 131, a success metric of the speech model 131, or a combination thereof. In particular aspects, the coverage of the audio samples used to train the speech model 131 corresponds to the distinct sounds represented by the audio samples (e.g., vowels, consonants, etc.). In particular aspects, the success metric is based on a comparison of the audio samples used to train the speech model 131 and the synthesized speech generated based on the speech model 131 (e.g., a match between the audio samples and the synthesized speech).

[0042] According to some implementations, a first color, a first shading, a first size, a first animation, or a combination thereof, of avatar 135 indicates that voice model 131 is not trained. A second color, a second shading, a second size, a second animation, or a combination thereof, of avatar 135 indicates that voice model 131 is partially trained. A third color, a third shading, a third size, a third animation, or a combination thereof, of avatar 135 indicates that training of voice model 131 is complete. In particular aspects, GUI generator 168 generates (or updates) GUI 145 to show a visual representation of avatar 135.

[0043] In particular aspects, the interruption manager 124 detects a network problem (e.g., reduced bandwidth) in the communication link to the device 104. In response to detecting the network problem, the interruption manager 124 sends an interruption notification 119 to the device 104 indicating the interruption of the voice audio stream 111, refrains from sending (e.g., stops transmitting) subsequent media frames of the media stream 109 to the device 104 until it detects that the network problem has been resolved, or both. For example, in response to detecting the network problem, the interruption manager 124 refrains from sending (e.g., stops transmitting) the voice audio stream 111, the video stream 113, or both to the device 104 until the interruption ends.

[0044] The interruption manager 124 sends the text stream 121, the metadata stream 123, or both corresponding to the subsequent media frames. For example, the interruption manager 124 continues to send the text stream 121, the metadata stream 123, or both corresponding to the subsequent media frames in a first operating mode (e.g., a caption data transmission mode) of the device 102. Illustratively, in the first operating mode (e.g., a caption data transmission mode), the conference manager 122 generates the media stream 109, the text stream 121, the metadata stream 123, or a combination thereof. In response to detecting a network problem in the first operating mode (e.g., a caption data transmission mode), the interruption manager 124 stops sending subsequent media frames of the media stream 109 and continues sending the text stream 121, the metadata stream 123, or both corresponding to the subsequent media frames to the device 104. Alternatively, in response to detecting a network problem in the second operating mode (e.g., the interrupted data transmission mode) of device 102, interrupt manager 124 generates text stream 121, metadata stream 123, or both, corresponding to subsequent media frames based on audio input 153. Illustratively, in the second operating mode (e.g., the interrupted data transmission mode), conference manager 122 generates media stream 109 and does not generate text stream 121, metadata stream 123, or both. In response to detecting a network problem in the second operating mode (e.g., the interrupted data transmission mode) of device 102, interrupt manager 124 stops transmitting subsequent media frames of media stream 109 and starts transmitting text stream 121, metadata stream 123, or both, corresponding to the subsequent media frames, to device 104. In certain aspects, in a second operating mode of device 102 (e.g., an interrupted data transmission mode), sending text stream 121, metadata stream 123, or both to device 104 corresponds to sending interruption notification 119 to device 104.

[0045] In certain aspects, the interruption manager 164 detects the interruption in the audio audio stream 111 in response to receiving an interruption notification 119 from the device 102. In certain aspects, when the device 102 is operating in the second operating mode (e.g., the interrupted data transmission mode), the interruption manager 164 detects the interruption in the audio audio stream 111 in response to receiving the text stream 121, the metadata stream 123, or both.

[0046] In certain aspects, the interruption manager 164 detects an interruption in the audio audio stream 111 in response to determining that an audio frame of the audio audio stream 111 was not received within a threshold duration of a last received audio frame of the audio audio stream 111. For example, the last received audio frame of the audio audio stream 111 is received at the device 104 at a first reception time. The interruption manager 164 detects the interruption in response to determining that an audio frame of the audio audio stream 111 was not received within a threshold duration of the first reception time. In certain aspects, the interruption manager 164 sends an interruption notification to the device 102. In certain aspects, the interruption manager 124 detects a network problem in response to receiving an interruption notification from the device 104. In response to detecting the network problem, the interruption manager 124 sends the text stream 121, the metadata stream 123, or both, to the device 104 (e.g., instead of sending a subsequent media frame of the media stream 109), as described above.

[0047] In response to detecting the interruption, the interruption manager 164 selectively generates an output based on the text stream 121. For example, in response to the interruption, the interruption manager 164 provides the text stream 121, the metadata stream 123, the annotated text stream 137, or a combination thereof, to the text-to-speech converter 166. The text-to-speech converter 166 generates the synthetic speech audio stream 133 by using the speech model 131 to perform text-to-speech conversion based on the text stream 121, the metadata stream 123, the annotated text stream 137, or a combination thereof. For example, the synthetic speech audio stream 133, which is based on the text stream 121 and independent of the metadata stream 123, corresponds to the speech indicated by the text stream 121 with the neural speech characteristics of the user 142, as represented by the speech model 131. As another example, the synthesized speech audio stream 133 based on the annotated text stream 137 (e.g., text stream 121 and metadata stream 123) corresponds to the speech indicated by the text stream 121 with the speech characteristics of the user 142, as represented by the speech model 131 with the intonation indicated by the metadata stream 123. Using the speech model 131 trained at least in part on the voice of the user 142 (e.g., the speech audio stream 111) to perform the text-to-speech conversion allows the synthesized speech audio stream 133 to better match the speech characteristics of the user 142. In response to the interruption, the interruption manager 164 provides the synthesized speech audio stream 133 as audio output 143 to the speaker 154, stops playback of the speech audio stream 111, stops playback of the video stream 113, or a combination thereof.

[0048] In certain aspects, the interruption manager 164 selectively displays the avatar 135 while simultaneously providing the synthesized speech audio stream 133 to the speaker 154 as the audio output 143. For example, the interruption manager 164 refrains from displaying the avatar 135 while providing the speech audio stream 111 to the speaker 154 as the audio output 143. As another example, the interruption manager 164 displays the avatar 135 while providing the synthesized speech audio stream 133 to the speaker 154 as the audio output 143. Illustratively, the GUI generator 168 updates the GUI 145 to display the avatar 135 instead of the video stream 113 while the synthesized speech audio stream 133 is output as the audio output 143 for playout by the speaker 154. In certain aspects, the interruption manager 164 displays a first representation of the avatar 135 while simultaneously providing the voice audio stream 111 as audio output 143 to the speaker 154 and displays a second representation of the avatar 135 while simultaneously providing the synthetic voice audio stream 133 as audio output 143 to the speaker 154. For example, as further described with reference to FIG. 3C , the first representation indicates that the avatar 135 is being trained or has been trained (e.g., a training indicator for the voice model 131), and the second representation indicates that the avatar 135 is speaking (e.g., the voice model 131 is being used to generate the synthetic voice).

[0049] In particular implementations, interruption manager 164 selectively provides text stream 121, annotated text stream 137, or both as output to display device 156. For example, interruption manager 164 provides text stream 121, annotated text stream 137, or both to GUI generator 168 for updating GUI 145 to display text stream 121, annotated text stream 137, or both, in response to an interruption during a second operating mode (e.g., an interrupted data display mode) of device 104. In an alternative implementation, interruption manager 164 continues to provide text stream 121, annotated text stream 137, or both as output to display device 156 (e.g., regardless of an interruption) during a first operating mode (e.g., a caption data display mode) of device 104. In certain aspects, the interruption manager 164 provides the text stream 121, the annotated text stream 137, or both to the display device 156 while simultaneously providing the synthesized speech audio stream 133 as audio output 143 to the speaker 154.

[0050] In particular implementations, interruption manager 164 outputs one or more of synthesized speech audio stream 133, text stream 121, or annotated text stream 137 based on an interrupt configuration setting and in response to an interruption. For example, interruption manager 164 provides synthesized speech audio stream 133 as audio output 143 to speaker 154 while simultaneously providing text stream 121, annotated text stream 137, or both, to display device 156 in response to an interruption and in response to determining that the interrupt configuration setting has a first value (e.g., 0 or “audio and text”). Interrupt manager 164 provides text stream 121, annotated text stream 137, or both, to display device 156 and refrains from providing audio output 143 to speaker 154 in response to an interruption and in response to determining that the interrupt configuration setting has a second value (e.g., 1 or “text only”). In response to the interruption and in response to determining that the interruption configuration setting has a third value (e.g., 2 or “audio only”), interruption manager 164 refrains from providing text stream 121, annotated text stream 137, or both, to display device 156 and provides synthesized speech audio stream 133 as audio output 143 to speaker 154. In certain aspects, the interruption configuration setting is based on default data, user input, or both.

[0051] In certain aspects, the interrupt manager 124 detects that the interrupt has ended and sends an interruption end notification to the device 104. For example, the interrupt manager 124 detects that the interruption has ended in response to determining that the available communication bandwidth of the communication link with the device 104 is greater than a threshold. In certain aspects, the interrupt manager 164 detects that the interruption has ended in response to receiving an interruption end notification from the device 102.

[0052] In another particular aspect, the interrupt manager 164 detects that the interrupt has ended and sends an interruption end notification to the device 102. For example, the interrupt manager 164 detects that the interruption has ended in response to determining that the available communication bandwidth of the communication link with the device 102 is greater than a threshold. In a particular aspect, the interrupt manager 164 detects that the interruption has ended in response to receiving an interruption end notification from the device 104.

[0053] In response to detecting that the interruption has ended, conference manager 122 resumes transmitting voice audio stream 111, video stream 113, or both, to device 104. In certain aspects, transmitting voice audio stream 111, video stream 113, or both, corresponds to transmitting an interruption end notification. In response to detecting that the interruption has ended during a second operating mode (e.g., an interrupted data transmission mode) of device 102, interrupt manager 124 refrains from sending text stream 121, metadata stream 123, or both, to device 104.

[0054] In response to detecting that the interruption has ended, conference manager 162 refrains from generating synthesized speech audio stream 133 based on text stream 121, refrains from providing (e.g., stops providing) synthesized speech audio stream 133 as audio output 143 to speaker 154, and resumes playing (e.g., providing) speech audio stream 111 as audio output 143 to speaker 154. In response to detecting that the interruption has ended, conference manager 162 resumes providing video stream 113 to display device 156. For example, conference manager 162 provides video stream 113 to GUI generator 168 for updating GUI 145 to display video stream 113.

[0055] In a particular aspect, in response to detecting that the interruption has ended, the interruption manager 164 sends a first request to the GUI generator 168 to update the GUI 145 to indicate that the voice model 131 is not being used to output synthetic voice audio (e.g., the avatar 135 is not speaking). In response to receiving the first request, the GUI generator 168 updates the GUI 145 to display a first representation of the avatar 135 indicating that the voice model 131 is being trained or has been trained and that the voice model 131 is not being used to output synthetic voice audio (e.g., the avatar 135 is not speaking). In an alternative aspect, in response to detecting that the interruption has ended, the interruption manager 164 sends a second request to the GUI generator 168 to stop displaying the avatar 135. For example, in response to receiving the second request, the GUI generator 168 updates the GUI 145 to refrain from displaying the avatar 135.

[0056] In particular aspects, in response to detecting that the interruption has ended during the second operating mode (e.g., the interrupted data display mode or the no captioned data mode), interruption manager 164 refrains from providing text stream 121, annotated text stream 137, or both, to display device 156. For example, GUI generator 168 updates GUI 145 to refrain from displaying text stream 121, annotated text stream 137, or both.

[0057] In this way, system 100 reduces (e.g., eliminates) information loss during interruptions in voice audio stream 111 during an online conference. For example, if a network problem prevents voice audio stream 111 from being received by device 104 but text can be received by device 104, user 144 continues to receive audio corresponding to user 142's voice (e.g., synthesized voice audio stream 133), text (e.g., text stream 121, annotated text stream 137, or both), or a combination thereof.

[0058] Although camera 150 and microphone 152 are shown as being coupled to device 102, in other implementations, camera 150, microphone 152, or both may be integrated into device 102. Although speaker 154 and display device 156 are shown as being coupled to device 104, in other implementations, speaker 154, display device 156, or both may be integrated into device 104. While one microphone and one speaker are shown, other implementations may include one or more additional microphones configured to capture user voice, one or more additional speakers configured to output voice audio, or a combination thereof.

[0059] It should be understood that for ease of explanation, device 102 will be described as a transmitting device and device 104 will be described as a receiving device. During a call, the roles of device 102 and device 104 can switch when user 144 begins speaking. For example, device 104 may be the transmitting device and device 102 may be the receiving device. By way of example, device 104 may include a microphone and a camera for capturing audio and video of user 144, and device 102 may include or be coupled to a speaker and a display for playing out audio and video to user 142. In certain aspects, for example, when both user 142 and user 144 are speaking simultaneously or at overlapping times, device 102 and device 104 may each be a transmitting device and a receiving device.

[0060] In certain aspects, conference manager 122 is also configured to perform one or more of the operations described with reference to conference manager 162, and vice versa. In certain aspects, interruption manager 124 is also configured to perform one or more of the operations described with reference to interruption manager 164, and vice versa. Although GUI generator 168 is described as separate from conference manager 162 and interruption manager 164, in other implementations, GUI generator 168 is integrated into conference manager 162, interruption manager 164, or both. To illustrate, in some examples, conference manager 162, interruption manager 164, or both are configured to perform some of the operations described with reference to GUI generator 168.

[0061] 2, a system operable to handle voice audio stream interruptions is shown and generally designated 200. In certain aspects, system 100 of FIG.

[0062] The system 200 includes a server 204 coupled to the device 102 and the device 104 via the network 106. The server 204 includes a conference manager 122 and an interruption manager 124. The server 204 is configured to transfer online conference data from the device 102 to the device 104 and vice versa. For example, the conference manager 122 is configured to establish an online conference between the device 102 and the device 104.

[0063] The device 102 includes a conference manager 222. During an online conference, the conference manager 222 sends the media stream 109 (e.g., the voice audio stream 111, the video stream 113, or both) to the server 204. The conference manager 122 of the server 204 receives the media stream 109 (e.g., the voice audio stream 111, the video stream 113, or both) from the device 102. In a particular implementation, the device 102 sends the text stream 121, the metadata stream 123, or both, simultaneously with sending the media stream 109 to the server 204.

[0064] 1, with server 204 replacing device 102. For example, conference manager 122 (operating at server 204 instead of at device 102 as in FIG. 1) sends media stream 109, text stream 121, metadata stream 123, or a combination thereof, to device 104 in a manner similar to that described with reference to FIG. 1. For example, conference manager 122 sends text stream 121, metadata stream 123, or both, during a first operating mode (e.g., a captioned data transmission mode) of server 204. In certain implementations, conference manager 122 forwards text stream 121, metadata stream 123, or both received from device 102 to device 104. In some implementations, conference manager 122 generates metadata stream 123 based on text stream 121, media stream 109, or a combination thereof. In these implementations, conference manager 122 forwards text stream 121 received from device 102 to device 104, sends metadata stream 123 generated at server 204 to device 104, or both. In some implementations, conference manager 122 generates text stream 121, metadata stream 123, or both based on media stream 109 and forwards text stream 121, metadata stream 123, or both to device 104. Alternatively, conference manager 122 refrains from sending text stream 121, metadata stream 123, or both in response to determining that no interruption is detected during a second operating mode (e.g., an interrupted data transmission mode) of server 204. Device 104 receives media stream 109, text stream 121, annotated text stream 137, or a combination thereof from server 204 over network 106.The conference manager 162 plays out media frames from the media stream 109, the text stream 121, the annotated text stream 137, or a combination thereof, as described with reference to Figure 1. The interruption manager 164 trains the speech model 131, displays the avatar 135, or both, as described with reference to Figure 1.

[0065] In particular aspects, in response to detecting a network problem, the interruption manager 124 sends an interruption notification 119 to the device 104 indicating the interruption of the voice audio stream 111, refrains from sending (e.g., stops transmitting) subsequent media frames of the media stream 109 to the device 104 until it detects that the network problem has been resolved (e.g., the interruption has ended), or both. The interruption manager 124 sends the text stream 121, the metadata stream 123, or both corresponding to the subsequent media frames to the device 104, as described with reference to FIG. 1 . For example, the interruption manager 124 forwards the text stream 121, the metadata stream 123, or both received from the device 102 to the device 104. In some examples, the interruption manager 124 sends the metadata stream 123, the text stream 121, or both generated at the server 204 to the device 104. In certain aspects, the interruption manager 124 selectively generates the metadata stream 123, the text stream 121, or both, in response to detecting an interruption in the voice audio stream 111 during a second operating mode (e.g., an interrupted data transmission mode) of the server 204.

[0066] In certain aspects, the interruption manager 164 detects an interruption in the voice audio stream 111 in a manner similar to that described with reference to FIG. 1 in response to receiving an interruption notification 119 from the interruption manager 124 (e.g., at the server 204), receiving the text stream 121, the metadata stream 123, or both when the server 204 is operating in a second operating mode (e.g., an interrupted data transmission mode), determining that an audio frame of the voice audio stream 111 is not received within a threshold duration of the last received audio frame of the voice audio stream 111, or a combination thereof. In certain aspects, the interruption manager 164 sends the interruption notification to the server 204. In certain aspects, the interruption manager 124 detects a network problem in response to receiving an interruption notification from the device 104. The interruption manager 124 sends the text stream 121, the metadata stream 123, or both, corresponding to a subsequent media frame, to the device 104, as described with reference to FIG. 1.

[0067] In response to detecting an interruption, the interruption manager 164 provides the text stream 121, the metadata stream 123, the annotated text stream 137, or a combination thereof, to the text-to-speech converter 166. The text-to-speech converter 166 generates a synthetic speech audio stream 133 by using the speech model 131 to perform text-to-speech conversion based on the text stream 121, the metadata stream 123, the annotated text stream 137, or a combination thereof, as described with reference to Figure 1. In response to an interruption, the interruption manager 164 provides the synthetic speech audio stream 133 as audio output 143 to the speaker 154, stops playing the speech audio stream 111, stops playing the video stream 113, displays the avatar 135, displays a particular representation of the avatar 135, displays the text stream 121, displays the annotated text stream 137, or a combination thereof, as described with reference to Figure 1.

[0068] In response to detecting that the interruption has ended, the conference manager 122 resumes transmitting the voice audio stream 111, the video stream 113, or both, to the device 104. In particular aspects, in response to detecting that the interruption has ended during the second operating mode of the server 204 (e.g., the interrupted data transmission mode), the interruption manager 124 refrains from sending (e.g., ceases transmitting) the text stream 121, the metadata stream 123, or both, to the device 104.

[0069] In response to detecting that the interruption has ended, the conference manager 162 refrains from generating the synthesized speech audio stream 133 based on the text stream 121, refrains from providing (e.g., stops) the synthesized speech audio stream 133 as audio output 143 to the speaker 154, resumes playing the speech audio stream 111 as audio output 143 to the speaker 154, resumes providing the video stream 113 to the display device 156, stops or adjusts the display of the avatar 135, refrains from providing the text stream 121 to the display device 156, refrains from providing the annotated text stream 137 to the display device 156, or any combination thereof.

[0070] In this way, system 200 reduces (e.g., eliminates) information loss during interruptions in voice audio stream 111 during online conferences with legacy devices (e.g., device 102 that does not include an interruption manager). For example, if network issues prevent voice audio stream 111 from being received by device 104 but text can be received by device 104, user 144 continues to receive audio (e.g., synthesized voice audio stream 133), text (e.g., text stream 121, annotated text stream 137, or both), or a combination thereof, corresponding to user 142's voice.

[0071] In certain aspects, server 204 may be closer (e.g., fewer network hops) to device 104, and sending text stream 121, metadata stream 123, or both from server 204 (e.g., instead of from device 102) may conserve overall network resources. In certain aspects, server 204 may have access to network information that may be useful in successfully sending text stream 121, metadata stream 123, or both to device 104. As an example, server 204 initially sends media stream 109 over a first network link. Server 204 detects a network problem and, based at least in part on determining that the first network link is unavailable or not functioning, sends text stream 121, metadata stream 123, or both using a second network link that appears available to accept text transmissions.

[0072] 3A, there is shown an example of GUI 145. In certain embodiments, GUI 145 is generated by system 100 of FIG. 1, system 200 of FIG. 2, or both.

[0073] The GUI 145 includes a video display 306, an avatar 135, and a training indicator (TI) 304. For example, the GUI generator 168 generates the GUI 145 during the initiation of an online conference. A video stream 113 (e.g., an image of a user 142 (e.g., Jill Pratt)) is displayed via the video display 306.

[0074] The training indicator 304 indicates the training level of the voice model 131 (e.g., 0% or not trained). For example, the training indicator 304 indicates that the voice model 131 has not been custom trained. In certain aspects, the representation of the avatar 135 (e.g., no color) also indicates the training level. In certain aspects, the representation of the avatar 135 indicates that no synthesized speech is being output. For example, the GUI 145 does not include a synthesized speech indicator such as that further described with reference to FIG. 3C.

[0075] In particular implementations, if an interruption occurs prior to custom training of the voice model 131 and the text-to-speech converter 166 uses the voice model 131 (e.g., a non-customized generic voice model) to generate the synthetic voice audio stream 133, the synthetic voice audio stream 133 corresponds to audio speech having generic voice characteristics that may differ from the voice characteristics of the user 142. In particular aspects, the voice model 131 is initialized using a generic voice model associated with demographic data of the user 142. In this aspect, the synthetic voice audio stream 133 corresponds to generic voice characteristics that match the demographic data of the user 142 (e.g., age, gender, regional accent, etc.).

[0076] 3B, there is shown an example of GUI 145. In certain embodiments, GUI 145 is generated by system 100 of FIG. 1, system 200 of FIG. 2, or both.

[0077] In particular examples, GUI generator 168 updates GUI 145 during the online conference. Training indicator 304 indicates a second training level (e.g., 20% or partially trained) of voice model 131. For example, training indicator 304 indicates that voice model 131 is custom trained or partially custom trained. In particular aspects, the (e.g., partially colored) representation of avatar 135 also indicates the second training level. In particular aspects, the representation of avatar 135 indicates that synthetic speech is not being output. For example, GUI 145 does not include a synthetic speech indicator.

[0078] In certain implementations, when an interruption occurs after partial custom training of the voice model 131 and the text-to-speech converter 166 uses the voice model 131 (e.g., the partially customized voice model) to generate the synthetic voice audio stream 133, the synthetic voice audio stream 133 corresponds to an audio voice having voice characteristics that have some similarity to the voice characteristics of the user 142.

[0079] 3C, there is shown an example of GUI 145. In certain embodiments, GUI 145 is generated by system 100 of FIG. 1, system 200 of FIG. 2, or both.

[0080] In particular examples, GUI generator 168 updates GUI 145 in response to the interruption. Training indicator 304 indicates a third training level of voice model 131 (e.g., 100% or training complete). For example, training indicator 304 indicates that voice model 131 is being custom trained or that custom training is complete (e.g., a threshold level has been reached). In particular aspects, the (e.g., fully colored) representation of avatar 135 also indicates the third training level. In particular aspects, the representation of avatar 135 indicates that synthetic speech is being output. For example, GUI 145 includes a synthetic speech indicator 398 displayed as part of or with avatar 135 that indicates that the speech being played out is synthetic speech.

[0081] In the example of FIG. 3C , an interruption occurs after custom training of the voice model 131, and the text-to-speech converter 166 uses the voice model 131 (e.g., a customized voice model) to generate the synthetic voice audio stream 133, so that the synthetic voice audio stream 133 corresponds to an audio voice having voice characteristics similar to those of the user 142.

[0082] In response to the interruption, the interruption manager 164 stops outputting the video stream 113. For example, the video display 306 indicates that the output of the video stream 113 has been stopped due to an interruption (e.g., a network problem). The GUI 145 includes a text display 396. For example, the interruption manager 164 outputs the text stream 121 via the text display 396 in response to the interruption.

[0083] In certain aspects, the text stream 121 is displayed in real time so that the user 144 can continue to participate in the conversation. For example, the user 144 can read what the user 142 said on the text display 396 and then speak a response to the user 142. In certain aspects, if network issues prevent the voice audio stream corresponding to the voice of the user 144 from being received by the device 102, the interruption manager 124 can display the text stream corresponding to the voice of the user 144 at the device 102. In this way, one or more participants in an online conference can receive the text stream or voice audio stream corresponding to the voice of the other participants.

[0084] Referring to Figure 4A, a diagram of an exemplary embodiment of the operation of system 100 of Figure 1 or system 200 of Figure 2 is shown and is generally designated 400. The timing and operations shown in Figure 4A are for illustration purposes and not limitation. In other embodiments, additional or fewer operations may be performed, and the timing may differ.

[0085] Diagram 400 illustrates the timing of transmission of media frames of media stream 109 from device 102. In a particular aspect, media frames of media stream 109 are transmitted from device 102 to device 104 as described with reference to Figure 1. In an alternative aspect, media frames of media stream 109 are transmitted from device 102 to server 204 and from server 204 to device 102 as described with reference to Figure 2.

[0086] Device 102 transmits media frame (FR) 410 of media stream 109 at a first transmit time. Device 104 receives media frame 410 at a first receive time and provides media frame 410 for playback at a first play time. In a particular example, conference manager 162 buffers media frame 410 during a first buffer interval between the first receive time and the first play time. In a particular aspect, media frame 410 includes a first portion of video stream 113 and a first portion of voice audio stream 111. Conference manager 162 outputs the first portion of voice audio stream 111 to speaker 154 as a first portion of audio output 143 and the first portion of video stream 113 to display device 156 at the first play time.

[0087] The device 102 (or the server 204) is expected to transmit the media frame 411 at a second expected transmission time. The device 104 is expected to receive the media frame 411 at a second expected reception time. The interruption manager 164 of the device 104 detects a break in the voice audio stream 111 in response to determining that a media frame of the media stream 109 has not been received within a reception threshold duration of the first reception time. For example, the interruption manager 164 determines the second time based on the first reception time and the reception threshold duration (e.g., the second time = first reception time + reception threshold duration). The interruption manager 164 detects a break in the voice audio stream 111 in response to determining that a media frame of the media stream 109 has not been received between the first reception time and the second time. The second time is after the second expected reception time of the media frame 411 and prior to the expected playout time of the media frame 411. For example, the second time is during the expected buffer interval of the media frame 411.

[0088] The device 102 (or the server 204) detects the interruption in the voice audio stream 111, as described with reference to FIGS. 1-2. In response to the interruption in the voice audio stream 111, the interruption manager 124 (of the device 102 or the server 204) sends the text stream 121 corresponding to subsequent media frames (e.g., set of media frames 491) to the device 104 until the interruption ends. In a particular aspect, the media frames 411 include a second portion of the video stream 113 and a second portion of the voice audio stream 111. The interruption manager 124 (or the conference manager 122) generates text 451 for the text stream 121 by performing speech-to-text conversion on the second portion of the voice audio stream 111 and sends the text 451 to the device 104.

[0089] Device 104 receives text 451 of text stream 121 from device 102 or server 204, as described with reference to FIGS. 1-2. In response to the interruption, interruption manager 164 begins playing text stream 121 corresponding to subsequent media frames until the interruption ends. For example, interruption manager 164 provides text 451 to display device 156 at a second playback time. In certain aspects, the second playback time is based on (e.g., is the same as) the expected playback time of media frame 411.

[0090] In certain aspects, conference manager 222 of FIG. 2 is unaware of the interruption and sends media frames 413 of media stream 109 to server 204. In certain aspects, interruption manager 124 (of device 102 of FIG. 1 or server 204 of FIG. 2) stops sending media frames 413 to device 104 in response to the interruption. In certain aspects, media frames 413 include a third portion of video stream 113 and a third portion of voice audio stream 111. Interrupt manager 124 generates text 453 based on the third portion of voice audio stream 111. Interrupt manager 124 sends text 453 to device 104.

[0091] Device 104 receives text 453. Interruption manager 164, in response to the interruption, provides text 453 to display device 156 at a third playback time. In certain aspects, the third playback time is based on (e.g., is the same as) the expected playback time of media frame 413.

[0092] In response to the interruption ending, interruption manager 124 (of device 102 or server 204) resumes transmission of subsequent media frames (e.g., next media frame 493) of media stream 109 to device 104, as described with reference to FIGS. 1-2 . For example, conference manager 122 transmits media frame 415 to device 104. In response to the interruption ending, interruption manager 164 resumes playback of media stream 109 and stops playback of text stream 121. In a particular aspect, media frame 415 includes a fourth portion of video stream 113 and a fourth portion of voice audio stream 111. Conference manager 162 outputs the fourth portion of voice audio stream 111 to speaker 154 as part of audio output 143 and the fourth portion of video stream 113 to display device 156 during the fourth playback time.

[0093] As another example, conference manager 122 transmits media frame 417 to device 104. In a particular aspect, media frame 417 includes a fifth portion of video stream 113 and a fifth portion of voice audio stream 111. Conference manager 162 outputs the fifth portion of voice audio stream 111 to speaker 154 as part of audio output 143 and the fifth portion of video stream 113 to display device 156 at a fifth playback time.

[0094] In this way, device 104 prevents information loss by playing text stream 121 during interruptions in media stream 109. Playback of media stream 109 resumes when the interruption ends.

[0095] Referring to Figure 4B, a diagram of an exemplary embodiment of the operation of system 100 of Figure 1 or system 200 of Figure 2 is shown and is generally designated 490. The timing and operations shown in Figure 4B are for illustration purposes and not limitation. In other embodiments, additional or fewer operations may be performed, and the timing may differ.

[0096] Diagram 490 illustrates the timing of transmission of media frames of media stream 109 from device 102. GUI generator 168 of FIG. 1 generates GUI 145 indicating the training level of avatar 135. For example, GUI 145 may indicate that avatar 135 (e.g., voice model 131) is untrained or partially trained. Device 104 receives media frames 410 including a first portion of video stream 113 and a first portion of voice audio stream 111. Conference manager 162 outputs the first portion of voice audio stream 111 to speaker 154 as a first portion of audio output 143 and the first portion of video stream 113 to display device 156 at a first playback time, as described with reference to FIG. 4A . Interruption manager 164 trains voice model 131 based on media frames 410 (e.g., the first portion of voice audio stream 111), as described with reference to FIG. 1 . GUI generator 168 updates GUI 145 to indicate the updated training level of avatar 135 (eg, partially trained or fully trained).

[0097] The device 104 receives text 451 of the text stream 121 from the device 102 or the server 204, as described with reference to FIG. 4A . In response to the interruption, the interruption manager 164 stops playing the media stream 109, stops training the speech model 131, and starts playing the synthetic speech audio stream 133. For example, the interruption manager 164 generates synthetic speech frames 471 of the synthetic speech audio stream 133 based on the text 451. Illustratively, the interruption manager 164 provides the text 451 to the text-to-speech converter 166. The text-to-speech converter 166 uses the speech model 131 to perform text-to-speech conversion on the text 451 to generate synthetic speech frames (SFRs) 471. The interruption manager 164 provides the synthetic speech frames 471 as a second portion of the audio output 143 at a second playback time. The GUI generator 168 updates the GUI 145 to include a synthetic speech indicator 398, indicating that synthetic speech is being output. For example, GUI 145 shows avatar 135 speaking.

[0098] Device 104 receives text 453, as described with reference to Figure 4A. In response to the interruption, interruption manager 164 generates synthesized speech frame 473 of synthesized speech audio stream 133 based on text 453. Interruption manager 164 provides synthesized speech frame 473 as a third portion of audio output 143 at a third playback time.

[0099] 4A , the interruption manager 124 (of the device 102 or the server 204) resumes sending subsequent media frames (e.g., next media frame 493) of the media stream 109 to the device 104 in response to the interruption ending. For example, the conference manager 122 sends media frame 415 to the device 104. In response to the interruption ending, the interruption manager 164 resumes playing the media stream 109, stops playing the synthetic speech audio stream 133, and resumes training of the speech model 131. The GUI generator 168 updates the GUI 145 to remove the synthetic speech indicator 398, which indicates that synthetic speech is not being output.

[0100] In a particular example, conference manager 162 plays out media frame 415 and media frame 417. Illustratively, media frame 415 includes a fourth portion of video stream 113 and a fourth portion of voice audio stream 111. Conference manager 162 outputs the fourth portion of voice audio stream 111 as a fourth portion of audio output 143 to speaker 154 and the fourth portion of video stream 113 to display device 156 at a fourth playback time. In a particular aspect, conference manager 162 outputs the fifth portion of voice audio stream 111 as a fifth portion of audio output 143 to speaker 154 and the fifth portion of video stream 113 to display device 156 at a fifth playback time.

[0101] In this way, device 104 prevents information loss by playing synthesized speech audio stream 133 during interruptions in media stream 109. Playback of media stream 109 resumes when the interruption ends.

[0102] 5, a system operable to handle voice audio stream interruptions is shown and generally designated 500. In certain aspects, system 100 of FIG. 1 includes one or more components of system 500.

[0103] System 500 includes device 502 coupled to device 104 via network 106. During operation, conference manager 162 establishes an online conference with multiple devices (e.g., device 102 and device 502). For example, conference manager 162 establishes an online conference for user 144 with user 142 of device 102 and user 542 of device 502. Device 104 receives media stream 109 (e.g., voice audio stream 111, video stream 113, or both) representing the voice, image, or both of user 142 from device 102 or server 204, as described with reference to FIGS. 1-2 . Similarly, device 104 receives media stream 509 (e.g., second voice audio stream 511, second video stream 513, or both) representing the voice, image, or both of user 542 from device 502 or a server (e.g., server 204 or another server).

[0104] Conference manager 162 plays out media stream 109 simultaneously with playing out media stream 509, as further described with reference to FIG. 6A . For example, conference manager 162 provides video stream 113 to display device 156 simultaneously with providing second video stream 513 to display device 156. Illustratively, user 144 can view an image of user 142 simultaneously with viewing an image of user 542 during an online conference. As another example, conference manager 162 provides voice audio stream 111, second voice audio stream 511, or both, as audio output 143 to speaker 154. Illustratively, user 144 can hear the voice of user 142, the voice of user 542, or both. In certain aspects, interruption manager 164 trains voice model 131 based on voice audio stream 111, as described with reference to FIG. 1 . Similarly, the interruption manager 164 trains a second voice model for the user 542 based on the second voice audio stream 511 .

[0105] In a particular example, device 104 continues to receive media stream 509 during an interruption in voice audio stream 111. Interruption manager 164 plays out media stream 509 simultaneously with playing out synthetic voice audio stream 133, text stream 121, annotated text stream 137, or a combination thereof, as further described with reference to FIG. 6C . For example, interruption manager 164 generates synthetic voice audio stream 133 and provides secondary voice audio stream 511 simultaneously with providing synthetic voice audio stream 133 to speaker 154. As another example, interruption manager 164 generates updates to GUI 145, including text stream 121 or annotated text stream 137, and provides secondary video stream 513 to display device 156 simultaneously with providing updates to GUI 145 to display device 156. In this manner, user 144 can keep up with the conversation between user 142 and user 542 during an interruption in voice audio stream 111.

[0106] In certain aspects, the interruption of media stream 509 overlaps with the interruption of voice audio stream 111. Interruption manager 164 receives a second text stream, a second metadata stream, or both, corresponding to second voice audio stream 511. In certain aspects, interruption manager 164 generates a second annotated text stream based on the second text stream, the second metadata stream, or both. Interruption manager 164 generates a second synthetic voice audio stream by using a second voice model to perform text-to-speech conversion based on the second text stream, the second metadata stream, the second annotated text stream, or a combination thereof. Interruption manager 164 plays out second voice audio stream 511 to speaker 154 simultaneously with playing out synthetic voice audio stream 133. In certain aspects, interruption manager 164 plays out text stream 121, annotated text stream 137, or both, simultaneously with playing out second text stream, second annotated text stream, or both to display device 156. In this manner, user 144 can keep up with the conversation between user 142 and user 542 during interruptions in voice audio stream 111 and second voice audio stream 511.

[0107] In this way, system 500 reduces (e.g., eliminates) information loss during interruptions in one or more voice audio streams (e.g., voice audio stream 111, second voice audio stream 511, or both) during an online conference with multiple users. For example, if a network problem prevents one or more voice audio streams from being received by device 104, but text can be received by device 104, user 144 continues to receive the voice of user 142 and the audio, text, or a combination thereof corresponding to the voice of user 542.

[0108] 6A, there is shown an example of GUI 145. In certain embodiments, GUI 145 is generated by system 500 of FIG.

[0109] GUI 145 includes video displays, avatars, training indicators, or combinations thereof for multiple participants in the online conference. For example, GUI 145 includes video display 306, avatar 135, training indicator 304, or combinations thereof for user 142, as described with reference to FIG. 3A . GUI 145 also includes video display 606, avatar 635, training indicator (TI) 604, or combinations thereof for user 542. For example, GUI generator 168 generates GUI 145 during the initiation of the online conference. Simultaneously with the display of video stream 113 (e.g., an image of user 142 (e.g., Jill P.)) via video display 306, second video stream 513 of media stream 509 (e.g., an image of user 542 (e.g., Emily F.)) is displayed via video display 606.

[0110] Training indicator 304 indicates the training level of voice model 131 (e.g., 0% or untrained), and training indicator 604 indicates the training level of a second voice model (e.g., 10% or partially trained). The training levels of the voice models may be different if one user speaks more than the other user or if one user's voice contains a wider variety of sounds (e.g., higher model coverage).

[0111] In certain embodiments, the representations of avatar 135 (e.g., colorless) and avatar 635 (e.g., partially colored) also indicate the training level of the respective voice models. In certain embodiments, the representations of avatar 135 and avatar 635 indicate that no synthesized voice is being output. For example, GUI 145 does not include any synthesized voice indicator.

[0112] In particular implementations, when an interruption occurs in receiving media stream 109, text-to-speech converter 166 generates synthetic speech audio stream 133 using voice model 131 (e.g., a non-customized generic voice model). When an interruption occurs in receiving media stream 509, text-to-speech converter 166 generates a second synthetic speech audio stream using a second voice model (e.g., a partially customized voice model). In particular aspects, interruption manager 164 initializes the second voice model based on a second generic voice model that is distinct from the first generic voice model used to initialize voice model 131, such that when an interruption occurs prior to training (or full training) of voice model 131 and the second voice model, the synthesized voice of user 142 is distinguishable from the synthesized voice of user 542. In particular aspects, voice model 131 is initialized using a first generic voice model associated with demographic data of user 142, and the second voice model is initialized using a second generic voice model associated with demographic data of user 542.

[0113] 6B, there is shown an example of GUI 145. In certain embodiments, GUI 145 is generated by system 500 of FIG.

[0114] In a particular example, the GUI generator 168 updates the GUI 145 during the online conference. For example, the training indicator 304 indicates a second training level of the speech model 131 (e.g., 20% or partially trained) and a second training level of the second speech model (e.g., 100% or fully trained).

[0115] 6C, there is shown an example of GUI 145. In certain embodiments, GUI 145 is generated by system 500 of FIG.

[0116] In a particular example, GUI generator 168 updates GUI 145 in response to an interruption in reception of media stream 109. Training indicator 304 indicates a third training level of voice model 131 (e.g., 55% or partially trained), and training indicator 604 indicates a third training level of second voice model (e.g., 100% or fully trained). In a particular aspect, a representation of avatar 135 indicates that synthetic voice is being output. For example, GUI 145 includes synthetic voice indicator 398. A representation of avatar 635 indicates that synthetic voice is not being output to user 542. For example, GUI 145 does not include a synthetic voice indicator associated with avatar 635.

[0117] In response to the interruption, the interruption manager 164 stops outputting the video stream 113. For example, the video display 306 indicates that the output of the video stream 113 has been stopped due to an interruption (e.g., a network problem). In response to the interruption, the interruption manager 164 outputs the text stream 121 via the text display 396.

[0118] In certain aspects, text stream 121 is displayed in real time so that user 144 can keep up with and participate in the conversation. For example, user 144 can hear from synthesized speech audio stream 133, read on text display 396, or both, that user 142 has made a first utterance (e.g., "I hope you have something to celebrate"). User 144 can hear a response from user 542 in a second speech audio stream of media stream 509 output by speaker 154. User 144 can hear from synthesized speech audio stream 133, read on text display 396, or both, that user 142 has made a second utterance (e.g., "That's very interesting. I'm glad you enjoyed it"). In this way, user 144 can hear audio from a synthesized speech audio stream, read text in a text stream, or both, for one or more participants in an online conference while receiving a media stream for one or more other participants in the online conference.

[0119] Referring to Figure 7A, a diagram of an exemplary embodiment of the operation of system 500 of Figure 5 is shown and is generally designated 700. The timing and operations shown in Figure 7A are for illustration purposes and not limitation. In other embodiments, additional or fewer operations may be performed, and the timing may differ.

[0120] Diagram 700 illustrates the timing of transmission of media frames of media stream 109 from device 102 and media stream 509 from device 502. In certain aspects, media frames of media stream 109 are transmitted from device 102 or server 204 to device 104, as described with reference to Figures 1-2. Similarly, media frames of media stream 509 are transmitted from device 502 or a server (e.g., server 204 or another server) to device 104.

[0121] Device 104 receives media frames 410 of media stream 109 and media frames 710 of media stream 509 and provides media frames 410 and 710 for playback. For example, conference manager 162 outputs a first portion of voice audio stream 111 (e.g., as indicated by media frame 410) and a first portion of the second voice audio stream (e.g., as indicated by media frame 710) to speaker 154 as audio output 143, a first portion of video stream 113 (e.g., as indicated by media frame 410) via video display 306, and a first portion of the second video stream (e.g., as indicated by media frame 710) via video display 606, as described with reference to FIG.

[0122] 4A , device 104 receives text 451 (corresponding to media frame 411) of text stream 121 during a break in media stream 109. Device 104 receives media frame 711 of media stream 509. In response to the break, interruption manager 164 begins playing text stream 121 corresponding to subsequent media frames of media stream 109 simultaneously with the playback of media stream 509 until the break ends. For example, interruption manager 164 provides text 451 (e.g., indicated by media frame 411) to display device 156 simultaneously with providing media frame 711 for playback.

[0123] Device 104 receives text 453 (corresponding to media frame 413) of text stream 121 during a break in media stream 109, as described with reference to Figure 4A. Device 104 receives media frame 713 of media stream 509. Break manager 164 provides text 453 to display device 156 simultaneously with providing media frame 713 for playback.

[0124] 4A, in response to the interruption ending, interruption manager 164 resumes playback of media stream 109 and stops playback of text stream 121. Conference manager 162 receives and plays media frame 415 and media frame 715. Similarly, conference manager 162 receives and plays media frame 417 and media frame 717.

[0125] In this way, device 104 prevents information loss by playing text stream 121 during interruptions in media stream 109 simultaneously with the playback of media stream 509. Playback of media stream 109 resumes when the interruption ends.

[0126] Referring to Figure 7B, a diagram of an exemplary embodiment of the operation of system 500 of Figure 5 is shown and is generally designated 790. The timing and operations shown in Figure 7B are for illustration purposes only and are not limiting. In other embodiments, additional or fewer operations may be performed, and the timing may differ.

[0127] Diagram 790 illustrates the timing of transmission of media frames of media stream 109 from device 102 and media stream 509 from device 502. GUI generator 168 of FIG. 1 generates GUI 145 indicating the training level of avatar 135 and the training level of avatar 635. For example, GUI 145 indicates that avatar 135 (e.g., voice model 131) is untrained and that avatar 635 (e.g., second voice model) is partially trained. Device 104 receives and plays media frames 410 and 710. Interruption manager 164 trains voice model 131 based on media frames 410 and trains the second voice model based on media frames 710, as described with reference to FIG. 4B . GUI generator 168 updates GUI 145 to indicate the updated training level of avatar 135 (e.g., partially trained) and the updated training level of avatar 635 (e.g., fully trained).

[0128] Device 104 receives text 451 and media frames 711 of text stream 121. Interruption manager 164 generates synthesized speech frames 471 based on text 451, as described with reference to FIG. 4B . Interruption manager 164 plays synthesized speech frames 471 and media frames 711. GUI generator 168 updates GUI 145 to include synthesized speech indicator 398, which indicates to user 142 that synthesized speech is being output. For example, GUI 145 indicates that avatar 135 is speaking. GUI 145 does not include a synthesized speech indicator for user 542 (e.g., avatar 635 is not shown as speaking).

[0129] The device 104 receives the text 453 and the media frames 713. The interruption manager 164 generates the synthesized speech frames 473 based on the text 453, as described with reference to Figure 4B. The interruption manager 164 plays the synthesized speech frames 473 and the media frames 713.

[0130] 4B, in response to the interruption ending, the interruption manager 164 resumes playback of the media stream 109, stops playback of the synthetic speech audio stream 133, and resumes training of the speech model 131. The GUI generator 168 updates the GUI 145 to remove the synthetic speech indicator 398, which indicates that synthetic speech is not being output.

[0131] In a particular example, conference manager 162 receives and plays out media frame 415 and media frame 715. As another example, conference manager 162 receives and plays out media frame 417 and media frame 717.

[0132] In this way, device 104 prevents information loss by simultaneously playing out media stream 509 and playing synthesized speech audio stream 133 during interruptions in media stream 109. Playback of media stream 109 resumes when the interruption ends.

[0133] 8, a particular implementation of a method 800 for handling voice audio stream interruptions is shown. In a particular aspect, one or more operations of the method 800 are performed by the conference manager 162, the interruption manager 164, one or more processors 160, the device 104, the system 100, or a combination thereof of FIG.

[0134] Method 800 includes receiving a voice audio stream representing the voice of a first user during an online conference, at 802. For example, device 104 of FIG. 1 receives voice audio stream 111 representing the voice of user 142 during the online conference, as described with reference to FIG.

[0135] Method 800 also includes receiving 804 a text stream representing the speech of the first user. For example, device 104 of FIG. 1 receives text stream 121 representing the speech of user 142, as described with reference to FIG.

[0136] Method 800 further includes selectively generating an output based on the text stream in response to an interruption in the voice audio stream, at 806. For example, interruption manager 164 of Figure 1 selectively generates synthesized voice audio stream 133 based on text stream 121 in response to an interruption in voice audio stream 111, as described with reference to Figure 1. In particular implementations, interruption manager 164 selectively outputs text stream 121, annotated text stream 137, or both in response to an interruption in voice audio stream 111, as described with reference to Figure 1.

[0137] Method 800 improves upon, and thus reduces (e.g., eliminates), information loss during interruptions in voice audio stream 111 during an online conference. For example, if a network problem prevents voice audio stream 111 from being received by device 104 but text can be received by device 104, user 144 continues to receive audio corresponding to user 142's voice (e.g., synthesized voice audio stream 133), text (e.g., text stream 121, annotated text stream 137, or both), or a combination thereof.

[0138] 8 may be implemented by a field programmable gate array (FPGA) device, an application specific integrated circuit (ASIC), a processing unit such as a central processing unit (CPU), a DSP, a controller, another hardware device, a firmware device, or any combination thereof. As an example, the method 800 of FIG. 8 may be performed by a processor executing instructions such as those described with reference to FIG.

[0139] 9 shows one implementation 900 of device 104 as an integrated circuit 902 that includes one or more processors 160. The integrated circuit 902 also includes inputs 904 (e.g., one or more bus interfaces) to allow input data 928 (e.g., voice audio stream 111, video stream 113, media stream 109, interrupt notification 119, text stream 121, metadata stream 123, media stream 509, or a combination thereof) to be received for processing. The integrated circuit 902 also includes outputs 906 (e.g., bus interfaces) to allow transmission of output signals (e.g., voice audio stream 111, synthesized voice audio stream 133, audio output 143, video stream 113, text stream 121, annotated text stream 137, GUI 145, or a combination thereof). The integrated circuit 902 enables implementations that handle voice audio stream interruptions as components in a system, such as the mobile phone or tablet shown in FIG. 10, the headset shown in FIG. 11, the wearable electronic device shown in FIG. 12, the voice-controlled speaker system shown in FIG. 13, the camera shown in FIG. 14, the virtual reality or augmented reality headset shown in FIG. 15, or the vehicle shown in FIG. 16 or FIG. 17.

[0140] 10 shows an implementation 1000 in which the device 104 includes a mobile device 1002, such as a phone or tablet, as an illustrative, non-limiting example. The mobile device 1002 includes a microphone 1010, a speaker 154, and a display screen 1004. Components of one or more processors 160, including a conference manager 162, an interruption manager 164, a GUI generator 168, or a combination thereof, are integrated into the mobile device 1002 and are shown using dashed lines to indicate internal components that are generally not visible to a user of the mobile device 1002. In a particular example, the conference manager 162 outputs a voice audio stream 111 or the interruption manager 164 outputs a synthesized voice audio stream 133, which is then processed to perform one or more operations on the mobile device 1002, such as to launch a graphical user interface (e.g., via an integrated “smart assistant” application) or otherwise display other information associated with the user's voice on the display screen 1004.

[0141] 11 shows an implementation 1100 in which device 104 includes a headset device 1102. Headset device 1102 includes speaker 154, a microphone 1110, or both. One or more components of processor 160, including conference manager 162, interruption manager 164, or both, are integrated into headset device 1102. In a particular example, conference manager 162 outputs voice audio stream 111, or interruption manager 164 outputs synthesized voice audio stream 133, and voice audio stream 111 or synthesized voice audio stream 133 can cause headset device 1102 to perform one or more operations in headset device 1102 to transmit audio data corresponding to user voice to a second device (not shown) for further processing.

[0142] 12 shows an implementation 1200 in which the device 104 includes a wearable electronic device 1202 depicted as a “smart watch.” The conference manager 162, the interruption manager 164, the GUI generator 168, the speaker 154, the microphone 1210, or a combination thereof, are integrated into the wearable electronic device 1202. In a particular example, the conference manager 162 outputs the voice audio stream 111 or the interruption manager 164 outputs the synthesized voice audio stream 133, and the voice audio stream 111 or the synthesized voice audio stream 133 is then processed to perform one or more operations on the wearable electronic device 1202, such as to launch the GUI 145 or otherwise display other information associated with the user's voice on a display screen 1204 of the wearable electronic device 1202. By way of example, the wearable electronic device 1202 may include a display screen configured to display notifications based on the user's voice detected by the wearable electronic device 1202. In particular examples, the wearable electronic device 1202 includes a haptic device that provides a tactile notification (e.g., vibrates) in response to the detection of a user's voice. For example, the tactile notification may cause the user to look at the wearable electronic device 1202 to see a displayed notification indicating the detection of a keyword spoken by the user. In this manner, the wearable electronic device 1202 can alert a user who is hearing impaired or wearing a headset that the user's voice has been detected.

[0143] 13 is an implementation 1300 in which the device 104 includes a wireless speaker and voice-activated device 1302. The wireless speaker and voice-activated device 1302 can have wireless network connectivity and is configured to perform assistant operations. One or more processors 160 including a conference manager 162, an interruption manager 164, or both, a speaker 154, a microphone 1310, or a combination thereof are included in the wireless speaker and voice-activated device 1302. During operation, in response to receiving a verbal command identified as user speech in the voice audio stream 111 output by the conference manager 162 or in the synthesized voice audio stream 133 output by the interruption manager 164, the wireless speaker and voice-activated device 1302 can perform an assistant operation, such as through execution of a voice activation system (e.g., an integrated assistant application). The assistant operation can include creating a calendar event, adjusting the temperature, playing music, turning on lights, etc. For example, an assistant action is performed in response to receiving a command after a keyword or key phrase (e.g., "Hello, assistant").

[0144] 14 shows an implementation 1400 in which device 104 includes a portable electronic device corresponding to a camera device 1402. Conference manager 162, interruption manager 164, GUI generator 168, speaker 154, microphone 1410, or a combination thereof, are included in camera device 1402. In operation, in response to receiving a verbal command identified as user speech in voice audio stream 111 output by conference manager 162 or in synthesized voice audio stream 133 output by interruption manager 164, camera device 1402 can perform operations responsive to the verbal user command, such as, as illustrative examples, to adjust image or video capture settings, image or video playback settings, or image or video capture instructions.

[0145] 15 shows an implementation 1500 in which the device 104 includes a portable electronic device corresponding to a virtual reality, augmented reality, or mixed reality headset 1502. The conference manager 162, the interruption manager 164, the GUI generator 168, the speaker 154, the microphone 1510, or a combination thereof, are integrated into the headset 1502. User voice detection may be performed based on the voice audio stream 111 output by the conference manager 162 or the synthesized voice audio stream 133 output by the interruption manager 164. The visual interface device is placed in front of the user's eyes to enable augmented reality or virtual reality images or scenes to be displayed to the user while the headset 1502 is worn. In a particular example, the visual interface device is configured to display a notification indicating user voice detected in the audio stream. In another example, the visual interface device is configured to display a GUI 145.

[0146] 16 shows an implementation 1600 in which the device 104 corresponds to or is integrated within a vehicle 1602, shown as a manned or unmanned aerial device (e.g., a delivery drone). The conference manager 162, the interruption manager 164, the GUI generator 168, the speaker 154, the microphone 1610, or a combination thereof, are integrated into the vehicle 1602. User voice detection may be performed based on the voice audio stream 111 output by the conference manager 162 or the synthesized voice audio stream 133 output by the interruption manager 164, such as for delivery instructions from an authorized user of the vehicle 1602.

[0147] 17 shows another implementation 1700 in which the device 104 corresponds to or is integrated within a vehicle 1702, shown as an automobile. The vehicle 1702 includes one or more processors 160, including a conference manager 162, an interruption manager 164, a GUI generator 168, or a combination thereof. The vehicle 1702 also includes a speaker 154, a microphone 1710, or both. User voice detection may be performed based on the voice audio stream 111 output by the conference manager 162 or the synthesized voice audio stream 133 output by the interruption manager 164. For example, the user voice detection may be used to detect a voice command from an authorized user of the vehicle 1702 (e.g., to start the engine or the heating). In particular implementations, in response to receiving a verbal command identified as user speech in the voice audio stream 111 output by the conference manager 162 or in the synthesized voice audio stream 133 output by the interruption manager 164, the voice-activated system of the vehicle 1702 initiates one or more actions of the vehicle 1702 based on one or more keywords (e.g., “unlock,” “start the engine,” “play music,” “show weather,” or another voice command) detected in the voice audio stream 111 or the synthesized voice audio stream 133, such as by providing feedback or information via the display 1720 or one or more speakers (e.g., speaker 154). In particular implementations, the GUI generator 168 provides information about the online conference (e.g., the call) to the display 1720. For example, the GUI generator 168 provides the GUI 145 to the display 1720.

[0148] 18, a block diagram of a particular example implementation of a device is shown, generally designated 1800. In various implementations, device 1800 may have more or fewer components than shown in FIG. 18. In an example implementation, device 1800 may correspond to device 104. In an example implementation, device 1800 may perform one or more of the operations described with reference to FIGS. 1-17.

[0149] In particular implementations, device 1800 includes a processor 1806 (e.g., a central processing unit (CPU)). Device 1800 may include one or more additional processors 1810 (e.g., one or more DSPs). In particular aspects, one or more processors 160 of FIG. 1 correspond to processor 1806, processor 1810, or a combination thereof. Processor 1810 may include a voice and music coder-decoder (codec) 1808, including a speech coder (“vocoder”) encoder 1836, a vocoder decoder 1838, a conference manager 162, an interruption manager 164, a GUI generator 168, or a combination thereof. In particular aspects, one or more processors 160 of FIG. 1 include processor 1806, processor 1810, or a combination thereof.

[0150] The device 1800 may include a memory 1886 and a codec 1834. The memory 1886 may include instructions 1856 executable by one or more additional processors 1810 (or processor 1806) to implement the functionality described with reference to the conference manager 162, the interruption manager 164, the GUI generator 168, or a combination thereof. In particular aspects, the memory 1886 stores program data 1858 used or generated by the conference manager 162, the interruption manager 164, the GUI generator 168, or a combination thereof. In particular aspects, the memory 1886 includes the memory 132 of FIG. 1. The device 1800 may include a modem 1840 coupled to an antenna 1842 via a transceiver 1850.

[0151] The device 1800 may include a display device 156 coupled to a display controller 1826. The speaker 154 and one or more microphones 1832 may be coupled to a codec 1834. The codec 1834 may include a digital-to-analog converter (DAC) 1802, an analog-to-digital converter (ADC) 1804, or both. In particular implementations, the codec 1834 can receive analog signals from the one or more microphones 1832, convert the analog signals to digital signals using the analog-to-digital converter 1804, and provide the digital signals to the voice and music codec 1808. The voice and music codec 1808 can process the digital signals, which may be further processed by the conference manager 162, the interruption manager 164, or both. In particular implementations, the voice and music codec 1808 can provide the digital signals to the codec 1834. The codec 1834 can convert the digital signals to analog signals using the digital-to-analog converter 1802 and provide the analog signals to the speaker 154.

[0152] In particular implementations, the device 1800 may be included in a system-in-package or system-on-chip device 1822. In particular implementations, the memory 1886, the processor 1806, the processor 1810, the display controller 1826, the codec 1834, the modem 1840, and the transceiver 1850 are included in the system-in-package or system-on-chip device 1822. In particular implementations, the input device 1830 and the power supply 1844 are coupled to the system-on-chip device 1822. Additionally, in particular implementations, the display device 156, the input device 1830, the speaker 154, the one or more microphones 1832, the antenna 1842, and the power supply 1844 are external to the system-on-chip device 1822, as shown in FIG. In particular implementations, each of the display device 156, the input device 1830, the speaker 154, the one or more microphones 1832, the antenna 1842, and the power supply 1844 may be coupled to a component of the system-on-chip device 1822, such as an interface or a controller.

[0153] The device 1800 may include a virtual assistant, a home appliance, a smart device, an Internet of Things (IoT) device, a communication device, a headset, a vehicle, a computer, a display device, a television, a game console, a music player, a radio, a video player, an entertainment unit, a personal media player, a digital video player, a camera, a navigation device, a smart speaker, a speaker bar, a mobile communication device, a smartphone, a cellular phone, a laptop computer, a tablet, a personal digital assistant, a digital video disc (DVD) player, a tuner, an augmented reality headset, a virtual reality headset, an aviation vehicle, a home automation system, a voice-activated device, a wireless speaker and a voice-activated device, a portable electronic device, an automobile, a computing device, a virtual reality (VR) device, a base station, a mobile device, or any combination thereof.

[0154] In accordance with the described implementation, the apparatus includes means for receiving a voice audio stream during the online conference, the voice audio stream representing the voice of the first user. For example, the means for receiving the voice audio stream can correspond to conference manager 162, interruption manager 164, one or more processors 160, device 104, system 100, conference manager 122, server 204, system 200, one or more processors 1810, processor 1806, voice and music codec 1808, modem 1840, transceiver 1850, antenna 1842, device 1800, one or more other circuits or components configured to receive a voice audio stream during the online conference, or any combination thereof.

[0155] The apparatus also includes means for receiving a text stream representing the speech of the first user. For example, the means for receiving the text stream may correspond to conference manager 162, interruption manager 164, text-to-speech converter 166, one or more processors 160, device 104, system 100, conference manager 122, interruption manager 124, server 204, system 200, one or more processors 1810, processor 1806, voice and music codec 1808, modem 1840, transceiver 1850, antenna 1842, device 1800, one or more other circuits or components configured to receive the text stream, or any combination thereof.

[0156] The apparatus further includes means for selectively generating an output based on the text stream in response to an interruption in the voice audio stream. For example, the means for selectively generating an output may correspond to interruption manager 164, text-to-speech converter 166, GUI generator 168, one or more processors 160, device 104, system 100, interruption manager 124, server 204, system 200, one or more processors 1810, processor 1806, voice and music codec 1808, device 1800, one or more other circuits or components configured to selectively generate an output, or any combination thereof.

[0157] In some implementations, a non-transitory computer-readable medium (e.g., a computer-readable storage device such as memory 1886) includes instructions (e.g., instructions 1856) that, when executed by one or more processors (e.g., one or more processors 1810 or processor 1806), cause the one or more processors to receive a voice audio stream (e.g., voice audio stream 111) representing the voice of a first user (e.g., user 142) during an online conference. The instructions, when executed by the one or more processors, also cause the one or more processors to receive a text stream (e.g., text stream 121) representing the voice of the first user (e.g., user 142). The instructions, when executed by the one or more processors, further cause the one or more processors to selectively generate an output (e.g., synthesized voice audio stream 133, annotated text stream 137, or both) based on the text stream in response to an interruption of the voice audio stream.

[0158] Certain aspects of the disclosure are described below in a first set of interrelated clauses.

[0159] According to clause 1, a device for communications includes one or more processors, the one or more processors configured to receive an audio stream representing a voice of a first user during an online conference, receive a text stream representing the voice of the first user, and selectively generate an output based on the text stream in response to an interruption of the audio stream.

[0160] Clause 2 includes the device of clause 1, wherein the one or more processors are configured to detect the interruption in response to determining that an audio frame of the voice audio stream has not been received within a threshold duration of a last-received audio frame of the voice audio stream.

[0161] Clause 3 includes the device of clause 1, wherein the one or more processors are configured to detect the interruption in response to receiving the text stream.

[0162] Clause 4 includes the device of clause 1, wherein the one or more processors are configured to detect the interruption in response to receiving the interruption notification.

[0163] Clause 5 includes the device of any of clauses 1-4, wherein the one or more processors are configured to provide the text stream as output to a display.

[0164] Clause 6 includes the device of any of clauses 1 to 5, wherein the one or more processors are further configured to receive a metadata stream indicating an intonation of the first user's voice, and annotate the text stream based on the metadata stream.

[0165] Clause 7 includes the device of any of clauses 1-6, wherein the one or more processors are further configured to perform text-to-speech conversion on the text stream to generate a synthetic speech audio stream, and to provide the synthetic speech audio stream as output to a speaker.

[0166] Clause 8 includes the device of clause 7, wherein the one or more processors are further configured to receive a metadata stream indicating an intonation of the first user's voice, and the text-to-speech conversion is based on the metadata stream.

[0167] Clause 9 includes the device of clause 7, wherein the one or more processors are further configured to display an avatar concurrently with providing the synthesized speech audio stream to the speaker.

[0168] Clause 10 includes the device of clause 9, wherein the one or more processors are configured to receive media streams during the online conference, the media streams including a voice audio stream and a video stream of the first user.

[0169] Clause 11 includes the device of clause 10, wherein the one or more processors are configured to, in response to the interruption, stop playing the voice audio stream and stop playing the video stream.

[0170] Clause 12 includes the device of clause 10, wherein the one or more processors are configured to, in response to the interruption terminating, refrain from providing the synthesized speech audio stream to the speaker, refrain from displaying the avatar, resume playing the video stream, and resume playing the speech audio stream.

[0171] Clause 13 includes the device of clause 7, wherein the text-to-speech conversion is performed based on a speech model.

[0172] Clause 14 includes the device of clause 13, wherein the speech model corresponds to a generic speech model.

[0173] Clause 15 includes the device of clause 13 or clause 14, wherein the one or more processors are configured to update the speech model based on the speech audio stream prior to the interruption.

[0174] Clause 16 includes the device of any of clauses 1-15, wherein the one or more processors are configured to receive a second voice audio stream representing a voice of a second user during the online conference, and to provide the second voice audio stream to a speaker simultaneously with generating the output.

[0175] Clause 17 includes the device of any of clauses 1-16, wherein the one or more processors are configured to: in response to an interruption in the audio stream, stop playing the audio stream; and in response to the interruption ending, refrain from generating output based on the text stream and resume playing the audio audio stream.

[0176] Certain aspects of the disclosure are described below in a second set of interrelated clauses.

[0177] According to clause 18, a method of communication includes receiving, at a device, a voice audio stream representing a voice of a first user during an online conference; receiving, at the device, a text stream representing the voice of the first user; and selectively generating, at the device, an output based on the text stream in response to an interruption of the voice audio stream.

[0178] Clause 19 includes the method of clause 18, further including detecting an interruption in response to determining that an audio frame of the voice audio stream has not been received within a threshold duration of a last received audio frame of the voice audio stream.

[0179] Clause 20 includes the method of clause 18, further including the step of detecting a break in response to receiving the text stream.

[0180] Clause 21 includes the method of clause 18, further including the step of detecting the interruption in response to receiving an interruption notification.

[0181] Clause 22 includes the method of any of clauses 18-21, further including providing the text stream as output to a display.

[0182] Clause 23 includes the method of any of clauses 18-22, further including receiving a metadata stream indicative of the first user's vocal intonation, and annotating the text stream based on the metadata stream.

[0183] Certain aspects of the disclosure are described below in a third set of interrelated clauses.

[0184] According to clause 24, a non-transitory computer-readable storage medium stores instructions that, when executed by one or more processors, cause the one or more processors to receive an audio stream representing a voice of a first user during an online conference, receive a text stream representing the voice of the first user, and, in response to a disruption of the audio stream, selectively generate an output based on the text stream.

[0185] Clause 25 includes the non-transitory computer-readable storage medium of clause 24, the instructions, when executed by one or more processors, cause the one or more processors to perform text-to-speech conversion on the text stream to generate a synthetic speech audio stream, and provide the synthetic speech audio stream as output to a speaker.

[0186] Clause 26 includes the non-transitory computer-readable storage medium of clause 25, the instructions, when executed by one or more processors, cause the one or more processors to receive a metadata stream indicative of an intonation of the first user's voice, and the text-to-speech conversion is based on the metadata stream.

[0187] Clause 27 includes the non-transitory computer-readable storage medium of clause 25 or clause 26, wherein the instructions, when executed by one or more processors, cause the one or more processors to display an avatar simultaneously with providing a synthesized voice audio stream to a speaker.

[0188] Clause 28 includes the non-transitory computer-readable storage medium of any of clauses 25-27, wherein the instructions, when executed by one or more processors, cause the one or more processors to update a speech model based on the speech audio stream prior to the interruption, and wherein the text-to-speech conversion is performed based on the speech model.

[0189] Certain aspects of the present disclosure are described below in a fourth set of interrelated clauses.

[0190] According to clause 29, the apparatus includes means for receiving an audio stream during an online conference, the audio stream representing a speech of a first user, means for receiving a text stream representing the speech of the first user, and means for selectively generating an output based on the text stream in response to an interruption of the audio stream.

[0191] Clause 30 includes the apparatus of clause 29, wherein the means for receiving the voice audio stream, the means for receiving the text stream, and the means for selectively generating the output are integrated into at least one of a virtual assistant, a home appliance, a smart device, an Internet of Things (IoT) device, a communications device, a headset, a vehicle, a computer, a display device, a television, a game console, a music player, a radio, a video player, an entertainment unit, a personal media player, a digital video player, a camera, or a navigation device.

[0192] Those skilled in the art will further appreciate that the various illustrative logical blocks, configurations, modules, circuits, and algorithm steps described in connection with the implementations disclosed herein may be implemented as electronic hardware, computer software executed by a processor, or a combination of both. Various illustrative components, blocks, configurations, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or as processor-executable instructions depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in varying ways for each particular application, and such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.

[0193] The steps of a method or algorithm described in connection with the implementations disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. The software module may reside in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, a hard disk, a removable disk, a compact disk read-only memory (CD-ROM), or any other form of non-transitory storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from, and write information to, the storage medium. Alternatively, the storage medium may be integral to the processor. The processor and the storage medium may reside in an application-specific integrated circuit (ASIC). The ASIC may reside in a computing device or a user terminal. Alternatively, the processor and the storage medium may reside as discrete components in a computing device or a user terminal.

[0194] The previous description of the disclosed embodiments is provided to enable any person skilled in the art to make or use the disclosed embodiments. Various modifications of these embodiments will be readily apparent to those skilled in the art, and the principles defined herein may be applied to other embodiments without departing from the scope of the present disclosure. Thus, the present disclosure is not intended to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features defined by the following claims. [Explanation of symbols]

[0195] 100 systems 102 devices 104 devices 106 Network 109 Media Streams 111 Voice Audio Stream 113 Video Stream 119 Suspension Notice 120 processors 121 Text Stream 122 Conference Manager 123 Metadata Stream 124 Suspension Manager 131 Voice Models 132 memory 133 Synthetic Speech Audio Stream 135 Avatar 137 Annotated Text Streams 142 users 143 Audio Output 144 users 145 GUI 150 cameras 151 Video Input 152 microphones 153 Audio Input 154 speakers 156 Display Devices 160 processors 162 Conference Manager 164 Suspension Manager 166 Text to Speech Converter 168 Graphical User Interface (GUI) Generator, GUI Generator 200 systems 204 Server 222 Conference Manager 304 Training Indicator (TI), Training Indicator 306 Video Display 396 Text Display 398 Synthetic Speech Indicator 400 Figures 410 Media Frame (FR), Media Frame 411 Media Frame 413 Media Frame 415 Media Frame 417 Media Frame 451 Text 453 Text 471 Synthetic Speech Frames 473 Synthetic Speech Frames 491 Set of Media Frames 493 Next Media Frame 500 Systems 502 devices 509 Media Stream 511 Secondary Audio Stream 513 Secondary Video Stream 542 users 604 Training Indicator (TI), Training Indicator 606 Video Display 635 Avatar 700 Figures 710 Media Frame 711 Media Frame 713 Media Frame 715 Media Frame 717 Media Frame 790 Figures 800 ways 900 Implementation 902 Integrated Circuits 904 Input section 906 Output section 928 input data 1000 Implementation 1002 mobile devices 1004 display screen 1010 Microphone 1100 Mounting Form 1102 Headset Device 1110 Microphone 1200 Implementation 1202 Wearable Electronic Devices 1204 display screen 1210 Microphone 1300 Implementation 1302 Wireless Speakers and Voice-Activated Devices 1310 Microphone 1400 Implementation 1402 Camera Device 1410 Microphone 1500 Implementation 1502 Virtual reality, augmented reality, or mixed reality headsets; headsets 1510 Microphone 1600 Implementation 1602 Vehicle 1610 Microphone 1700 Mounting Form 1702 Vehicle 1710 Microphone 1720 display 1800 devices 1802 Digital-to-analog converter (DAC), digital-to-analog converter 1804 Analog-to-Digital Converter (ADC), Analog-to-Digital Converter 1806 processor 1808 Speech and Music Codec Decoder (Codec) 1810 processor 1822 System in Package or System on Chip Device, System on Chip Device 1826 Display Controller 1830 Input Devices 1832 Microphone 1834 codec 1836 Speech Coder ("Vocoder") Encoder 1838 Vocoder Decoder 1840 modem 1842 Antenna 1844 Power supply 1850 transceiver 1856 command 1858 Program Data 1886 Memory

Claims

1. A communication device, one or more processors, wherein the one or more processors: receiving a voice audio stream representing a voice of a first user during an online conference; receiving a text stream representing the speech of the first user; receiving a metadata stream indicative of the intonation of the voice of the first user; detecting an interruption in the voice audio stream based on determining that an audio frame of the voice audio stream has not been received within a threshold duration of a last received audio frame of the voice audio stream; generating an output based on the text stream in response to detecting the interruption in the voice audio stream; annotating the text stream based on the metadata stream; configured to: device.

2. The device of claim 1 , wherein the one or more processors are further configured to detect the interruption in response to receiving the text stream.

3. the one or more processors: responsive to receiving a notification of the interruption, detecting the interruption; and providing the text stream as the output to a display; The device of claim 1 configured to:

4. the one or more processors: performing text-to-speech conversion on the text stream to generate a synthetic speech audio stream; providing said synthesized speech audio stream as said output to a speaker; The device of claim 1 , further configured to:

5. The device described in claim 4, wherein the text-to-speech conversion is based on the metadata stream.

6. The device of claim 4 , wherein the one or more processors are further configured to display an avatar simultaneously with providing the synthesized speech audio stream to the speaker.

7. 7. The device of claim 6, wherein the one or more processors are configured to receive media streams during the online conference, the media streams including the voice audio stream and video stream of the first user.

8. the one or more processors, in response to the interruption, stopping playback of the voice audio stream; stopping the playback of said video stream; The device of claim 7 configured to:

9. the one or more processors, in response to the interruption being terminated, refraining from providing the synthesized speech audio stream to the speaker; refraining from displaying said avatar; resuming playback of said video stream; resuming playback of the voice audio stream; The device of claim 7 configured to:

10. The device of claim 4 , wherein the text-to-speech conversion is performed based on a speech model.

11. The device of claim 10 , wherein the speech model corresponds to a generic speech model.

12. The device of claim 10 , wherein the one or more processors are configured to update the speech model based on the voice audio stream prior to the interruption.

13. the one or more processors: receiving a second voice audio stream representing a voice of a second user during the online conference; providing the second voice audio stream to a speaker simultaneously with generating the output; and / or ceasing playback of the voice audio stream in response to the interruption of the voice audio stream; and In response to the interruption being terminated, refraining from generating the output based on the text stream; resuming playback of the voice audio stream; The device of claim 1 configured to:

14. A method of communication comprising: receiving, at the device, a voice audio stream representing a voice of a first user during an online conference; receiving, at the device, a text stream representing the speech of the first user; receiving, at the device, a metadata stream indicative of the intonation of the voice of the first user; detecting an interruption at the device in response to determining that an audio frame of the voice audio stream has not been received within a threshold duration of a last received audio frame of the voice audio stream; generating, at the device, an output based on the text stream in response to detecting the interruption in the voice audio stream; annotating the text stream based on the metadata stream; Including, method.

15. A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to: receiving a voice audio stream representing a voice of a first user during an online conference; receiving a text stream representing the speech of the first user; receiving a metadata stream indicative of the intonation of the voice of the first user; detecting an interruption in response to determining that an audio frame of the voice audio stream has not been received within a threshold duration of a last received audio frame of the voice audio stream; generating an output based on the text stream in response to detecting the interruption in the voice audio stream; annotating the text stream based on the metadata stream; to carry out A non-transitory computer-readable storage medium.

Citation Information

Patent Citations

  • Communication system and its method, communication service server and communication terminal

    JP2001230801A

  • Portable telephone apparatus with translation function, method for translating voice data, voice data translation program, and program recording medium

    JP2008021058A

  • How to maintain voice communication in congested communication channels

    JP2016529839A

  • Artificially generated speech for a communication session

    US20180218727A1

  • Context-based cognitive speech to text engine

    US20180226073A1