Dynamic presentation of audio transcription for electronic voice messaging
By implementing a live voicemail system and learning from local voice data on user devices, the system automatically answers and provides caller voice transcription, solving the problem of real-time processing of calls from unknown numbers, improving user experience and transcription accuracy, while protecting user privacy.
Patent Information
- Application Number
- CN202480037384.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-17
- Filing Date
- 2024-05-31
- Publication Date
- 2026-01-02
AI Technical Summary
In existing technologies, calls from unknown numbers require users to listen to voicemails and call back based on the content, lacking real-time transcription and simplified processing methods, resulting in a poor user experience.
By implementing a live voicemail system on user devices, calls are automatically answered and caller voice transcripts are provided, allowing users to read and decide whether to answer while maintaining a connection. The system combines confidence assessment and local voice data learning to improve transcription accuracy.
It improves the ease of managing calls from unknown numbers, enhances the user experience, improves transcription accuracy through local voice data, and protects user privacy.
Smart Images

Figure CN121264031A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present description relates generally to audio transcription, and more specifically to dynamic presentation of audio transcription, e.g., for electronic voice messaging. BACKGROUND
[0002] Voice mail is a widely used communication feature that allows a caller to leave a recorded audio message for a called party when the called party is unable to answer a telephone call. It serves as a means of relaying information, delivering messages, and facilitating communication when immediate interaction is not possible. BRIEF DESCRIPTION OF DRAWINGS
[0003] The specific features of the subject technology are set forth in the appended claims. However, for purposes of explanation, several implementations of the subject technology are set forth in the following drawings.
[0004] Figure 1 An example network environment in which electronic voice messaging with dynamic presentation of transcription can be implemented is illustrated in accordance with one or more implementations.
[0005] Figure 2 A schematic diagram of an electronic device for providing dynamic presentation of transcription during an electronic voice messaging session is illustrated in accordance with one or more implementations.
[0006] Figure 3 A schematic diagram illustrating example user interface views in which transcription is dynamically displayed on an electronic device during an electronic voice messaging session is illustrated in accordance with one or more implementations.
[0007] Figure 4 A flow diagram of an example process for providing dynamic presentation of transcription during an electronic voice messaging session is illustrated in accordance with one or more implementations.
[0008] Figure 5 An electronic system that can be used for implementing one or more implementations of the subject technology is illustrated. DETAILED DESCRIPTION
[0009] The detailed description set forth below is intended as a description of various configurations of the subject technology and is not intended to represent the only configurations in which the subject technology can be practiced. The appended drawings are incorporated herein and constitute a part of the detailed description. The detailed description includes specific details for the purpose of providing a thorough understanding of the subject technology. However, the subject technology is not limited to the specific details set forth herein and can be practiced with one or more other implementations. In one or more implementations, structures and components are shown in block diagram form in order to avoid obscuring the concepts of the subject technology.
[0010] Embodiments of the subject technology in this disclosure provide for the generation of a live audio transcription of an ongoing voicemail, allowing a user of an electronic device to respond to a received phone call (e.g., which was forwarded to voicemail) based on a displayed live audio transcription. The live audio transcription can provide a real-time textual representation of the caller’s speech, facilitating quick and efficient communication. By utilizing this feature, the user can easily determine the caller’s identity and the purpose of the call, enhancing the overall user experience. In one or more implementations, the user can answer the call while the voicemail is being recorded, e.g., based on the content of the transcription.
[0011] Embodiments of the subject technology in this disclosure also provide for the handling of calls from unknown numbers. In existing methods, an incoming call from an unknown number is directed to the carrier’s voicemail. Subsequently, the user must listen to the voicemail and then dial back the call based on the content of the voicemail message as needed. However, the present subject technology provides for a different handling of incoming calls from unknown numbers, providing a more streamlined approach. When a call is received from an unknown number, it can be selectively directed to a live voicemail system on the user device, rather than to the carrier’s voicemail. Thus, the user device can automatically answer the call and play a personal recorded greeting or a default greeting using synthesized speech. The default greeting can inform the caller to leave a message, which can be seen by someone and picked up.
[0012] Substantially concurrently, the user device can display a transcription of the caller’s speech, allowing the user to read the incoming message while keeping the connection. Based on the transcription, the user can decide to pick up the call and engage in a live conversation, utilize the information provided in the transcription, or let the call go to voicemail. Upon receiving the voicemail, they can listen to it and dial back the call as needed based on the content of the voicemail message. This feature enhances the user experience by providing improved call management and convenience. If the user chooses to silence calls from unknown numbers, unknown callers can be quickly directed to the live voicemail system, which the user can access if needed. In some aspects, the user device can not provide a notification when a call from an unknown number is received, but can provide a haptic response to indicate the availability of the recording of the voice message. Furthermore, the live voicemail system can incorporate a confidence assessment into the transcription of the voice message content. In instances where the live voicemail system has low confidence in the accuracy of the transcription, the content can not be displayed on the user device. In one or more implementations, the present subject system can forward any calls deemed to be spam, e.g., based on a directory of known spam numbers, to the carrier voicemail.
[0013] Embodiments of the subject technology in this disclosure also provide for intercepting an audio stream received at a user device and routing the audio stream through a transcription service, such as an on-device speech recognition model. Subsequently, the transcription service processes the audio and generates a set of text / conversations that can be dynamically displayed on the screen of the user device, continually updated as new conversations are received. Each conversation can be assigned a confidence score. If the confidence score is not high enough, it can be visually indicated with an underline, highlighting the uncertainty.
[0014] Generating a transcription at an electronic device that receives an audio input (e.g., as opposed to transmitting an audio stream to a server or other external transcription service for transcription) can be advantageous because local speech data of a speaker corresponding to the audio input is available to the electronic device that receives the audio input, learned, and / or stored, and used to improve audio transcription. Because this local speech data is maintained and privacy protected locally at the electronic device, the privacy of the user, the speaker, to whom the audio input belongs, can be maintained while the local speech data of the user is utilized to improve the ability of the electronic device to generate accurate and / or complete transcriptions.
[0015] Figure 1 An example network environment 100 in which electronic voice messaging with dynamic presentation of transcriptions can be implemented in accordance with one or more implementations is illustrated. However, not all of the depicted components can be used in all implementations, and one or more implementations can include additional or different components than those shown in the figure. Variations in the arrangement and type of the components can be made without departing from the spirit or ambit of the claims as set forth herein. Additional components, different components, or fewer components can be provided.
[0016] The network environment 100 includes electronic devices 110, 115, 117, 119, a server 120, and a server 130. A network 106 can communicatively couple the electronic devices 110, 115, 117, 119, the server 120, and / or the server 130. In one or more implementations, the network 106 can be an interconnected network of devices that can include or be communicatively coupled to the Internet. For purposes of explanation, the network environment 100 is illustrated in Figure 1 the figure as including the electronic devices 110, 115, 117, 119, the server 120, and the server 130; however, the network environment 100 can include any number of electronic devices and / or any number of servers communicatively coupled to each other, directly or via the network 106.
[0017] The server 160 can form all or part of a computer network or server group 170, such as in an access network implementation. For example, the server 170 stores data and software and includes specific hardware (e.g., processors and other specialized or customized processors) for providing access to Internet Protocol (IP) services, such as the Internet, an intranet, streaming services, cellular services, and / or other IP services. In one implementation, the server group 170 can be used as part of a cellular service that provides wireless communication to the electronic device 110, the electronic device 115, the electronic device 117, and / or the electronic device 119. The network 150 can communicatively couple the server group 170 to the electronic device 110, the electronic device 115, the electronic device 117, the electronic device 119, the server 120, and / or the server 130 via the network 106. Although the network 106 and the network 150 are depicted as separate networks, in other implementations these networks can form and / or can include all or part of a common network.
[0018] Any of the electronic device 110, the electronic device 115, the electronic device 117, or the electronic device 119 may, for example, be: a desktop computer; a portable computing device, such as a laptop computer, a smartphone, a peripheral device (e.g., a digital camera, headphones), a tablet device, a stand-alone voice messaging hardware, a wearable device (such as a watch, a wristband, etc.); or any other appropriate device that includes, for example, one or more wireless interfaces, such as a WLAN radio, a cellular radio, a Bluetooth radio, a Zigbee radio, a near-field communication (NFC) radio, and / or other radios. Any of the electronic device 110, the electronic device 115, the electronic device 117, and / or the electronic device 119 can be and / or can include all or part of the electronic system discussed below with respect to Figure 5
[0019] In Figure 1 In the example, the electronic device 110 is depicted as a desktop computer, the electronic devices 115 and 117 are depicted as tablet devices, and the electronic device 119 is depicted as a smartphone. In one or more implementations, the electronic device 110, the electronic device 115, the electronic device 117, and / or the electronic device 119 can include a voice messaging application and / or a transcription service installed and / or accessible at the electronic device. In one or more implementations, the electronic device 110, the electronic device 115, the electronic device 117, and / or the electronic device 119 can include a camera and / or a microphone, and can provide a voice messaging application for exchanging audio streams, video streams, and / or transcriptions (such as exchanged with a corresponding voice messaging application installed and accessible at one or more other electronic devices, e.g., the electronic device 110, the electronic device 115, the electronic device 117, and / or the electronic device 119) over the network 106.
[0020] In one or more implementations, one or more of the electronic device 110, the electronic device 115, the electronic device 117, and / or the electronic device 119 can have a voice messaging application installed and accessible at the electronic device, and can not have a transcription service available at the electronic device. In one or more implementations, one or more of the electronic device 110, the electronic device 115, the electronic device 117, and / or the electronic device 119 can not have a voice messaging application installed and available at the electronic device, but can be able to access an electronic voice messaging session without a voice messaging application, such as via a web-based voice messaging application provided at least in part by one or more servers.
[0021] In one or more implementations, one or more servers, such as server 120 and / or server 130, can perform operations for managing the secure exchange of audio and / or video streams between various electronic devices, such as electronic device 110, electronic device 115, electronic device 117, and / or electronic device 119, such as during an electronic voice messaging session (e.g., an audio voice messaging session or a video voice messaging session). In one or more implementations, server 120 can store account information associated with electronic device 110, electronic device 115, electronic device 117, electronic device 119, and / or users of these devices. In one or more implementations, one or more servers, such as server 130, can provide resources (e.g., web-based application resources) to manage connections to electronic voice messaging sessions and / or communications within electronic voice messaging sessions. In one or more implementations, one or more servers, such as server 130, can store information indicating one or more capabilities of electronic devices that are participants in an electronic voice messaging session, such as device transcription capabilities and / or other device capability information of the participant devices.
[0022] Figure 2 A schematic diagram of electronic device 119 for providing dynamic presentation of transcriptions during an electronic voice messaging session is illustrated in accordance with one or more implementations. In Figure 2 The rectangular boxes are used to indicate hardware components, and the trapezoidal boxes are used to indicate software processes that can be performed by one or more processors of the electronic device.
[0023] As shown in Figure 2 , electronic devices, such as electronic device 119, can include one or more microphones, such as microphone 202, and an output component 204 (e.g., a display and / or one or more speakers). Figure 2 A voice messaging application 208 and a transcription service 210 that can be installed and / or run at electronic device 119 are also illustrated. In Figure 2 the example of FIG. 2B, transcription service 210 is shown separately from voice messaging application 208 (e.g., as a system process at electronic device 119). However, in other implementations, transcription service 210 can be provided as part of voice messaging application 208.
[0024] Figure 2 A phone application 214 that is run at electronic device 119 is also illustrated. In Figure 2In the example, voice messaging application 208 is shown separately from telephone application 214. However, in other specific implementations, voice messaging application 208 may be provided as part of telephone application 214. Telephone application 214 may refer to the ability of electronic device 119 to make and receive voice calls, send and receive text messages, and access various communication features. Telephone application 214 may utilize the built-in cellular network connectivity of electronic device 119, thereby allowing users to establish real-time voice communication with other devices via conventional telephone calls or IP services such as Voice over IP (VoIP). Additionally, telephone application 214 may enable text-based communication via SMS (Short Message Service) or other messaging applications. Telephone application 214 may provide audio output corresponding to an audio stream and / or video output corresponding to a video stream for output via output component 204 (e.g., audio streams are output via output devices such as one or more speakers of electronic device 119 or one or more speakers connected to the electronic device), and / or video streams are output via a display device of electronic device 119.
[0025] like Figure 2 As shown, local input (e.g., audio input to microphone 202) can be received by a telephone application 214 running on electronic device 119. For example, a user of electronic device 119 can speak into microphone 202.
[0026] In one or more embodiments, voice messaging application 208 can act as a virtual answering service by accepting incoming calls on behalf of electronic device 119. When a call comes in, voice messaging application 208 can take over and prompt the caller to leave a voice message. For example, voice messaging application 208 can intercept incoming calls received at electronic device 119. This can be achieved by configuring settings on electronic device 119 to direct calls to voice messaging application 208. In one or more embodiments, voice messaging application 208 intercepts incoming calls before the telephone service provider's voicemail service establishes a voice messaging session with the incoming call.
[0027] Using intuitive prompts or automated instructions, callers can be guided through the process of recording their voice messages. For example, when a call is received at electronic device 119, voice messaging application 208 can control and present the caller with pre-recorded or synthesized prompts. These prompts can inform the caller that the called party (e.g., the user of electronic device 119) is currently unavailable and instruct them to leave a voice message. Voice messaging application 208 can capture audio input, convert the audio input into a digital audio file, and store the digital audio file as local voice data 212 for later retrieval by voice messaging application 208 and / or transcription service 210. This service ensures that callers can leave voice messages directly on electronic device 119, even when the user of electronic device 119 is unavailable or unable to answer the call.
[0028] like Figure 2 As shown, voice messaging application 208 can receive remote content input (e.g., remote audio content and / or remote video content) from one or more other electronic devices (such as electronic device 110, electronic device 115 and / or electronic device 117) during an electronic voice messaging session. Figure 2 Examples are also illustrated of how, in some operational scenarios of one or more embodiments, remote content (e.g., an audio stream from one or more other electronic devices, such as electronic device 110, electronic device 115, and / or electronic device 117, during an electronic voice messaging session) can be provided to transcription service 210. In one or more embodiments, transcription service 210 may generate a transcription of the audio portion of the remote content input and provide the transcription of the audio portion of the remote content input for display by output component 204 (e.g., dynamically displayed on a display device of electronic device 119 during at least a portion of the electronic voice messaging session or throughout the entire electronic voice messaging session). In one or more embodiments, the transcription may be provided to output component 204 for transmission to one or more electronic devices among electronic devices 110, electronic device 115, and / or electronic device 117, where it may be displayed and / or stored in a privacy-preserving manner. In one or more other embodiments, the voice messaging application 208 may use an indicator flag to indicate that the transcription is suppressed at the receiving device, based on the type of device to which the transcription is relayed (e.g., if one of the electronic devices in network environment 100 that serves as the receiving device is an electronic device integrated into a vehicle).
[0029] In one or more implementations, the transcription service 210 can also provide the transcription to the voice messaging application 208. In one or more other implementations, the transcription can be generated by the voice messaging application 208 (e.g., the transcription service 210 can be implemented as an integral part of the voice messaging application 208). In one or more implementations, the transcription can be generated and sent in segments, such that each segment of the transcription can be displayed at the electronic device 119 as the corresponding audio input is provided to the electronic device 119. The transcription service 210 or the voice messaging application 208 can generate time information for the transcription. The time information can be used to synchronize the transcription with the remote content input audio / video when the remote content input audio / video and the transcription are rendered at the electronic device 119. For example, the time at which the transcription (or a segment thereof) is generated or the time at which the transcribed audio input (or a segment thereof) is provided to be displayed at the electronic device 119 can be provided together with the time at which the audio stream (or a segment thereof) of the remote content input is received at the electronic device 119, and the times corresponding to the transcription and the times corresponding to the audio input can be used to synchronize the transcription and the corresponding audio stream of the user speaking the words in the transcription.
[0030] As Figure 2 illustrated, in one or more implementations, the transcription service 210 can use the local speech data 212 to help generate the transcription of the audio portion of the remote content input. For example, the local speech data 212 can include one or more stored and / or learned properties of the speech of the user of the electronic device 119 and / or the users of the electronic device 110, the electronic device 115, and / or the electronic device 117 (e.g., frequency characteristics, commonly used words or phrases, and / or a speech model that has been trained at the electronic device 119 with respect to speech input from the user of the electronic device 119 and / or the users of the electronic device 110, the electronic device 115, and / or the electronic device 117). In this way, the transcription service 210 at the electronic device 119 can leverage its own pre-existing knowledge of the user of the electronic device 119 and / or the users of the electronic device 110, the electronic device 115, and / or the electronic device 117 to generate a transcription of the spoken input of that user that is of higher quality than what a general transcription service of generic speech (e.g., a transcription service provided by a server or another device of another user) might otherwise have.
[0031] As Figure 2As shown, the transcription service 210 can generate a confidence of the transcription (e.g., a confidence score) (e.g., for a segment of the transcription, such as for a set of words spoken during a particular time period during the electronic voice messaging session). In one or more implementations, the transcription service 210 can determine whether the confidence score exceeds a confidence threshold. In some aspects, the confidence threshold is a predefined value. In other aspects, the confidence threshold can be a user-configured value. Accordingly, the electronic device 119 can display the transcription based on the transcription service 210 determining that the confidence score exceeds the confidence threshold. In one or more other implementations, the voice messaging application 208 can compare the confidence score to the confidence threshold to determine whether the transcription should be displayed on the electronic device 119. In one or more implementations, when the confidence is below the threshold, the transcription service 210 can generate an updated transcription with an updated confidence score. For example, if the electronic device 119 (e.g., the voice messaging application 208 or the transcription service 210) determines that the updated confidence score is greater than the confidence score of the previously generated transcription, the electronic device 119 can display the updated transcription on the electronic device 119.
[0032] In one or more implementations, the electronic voice messaging session can refer to a voicemail interaction conducted via the network 106 and / or the network 150 through a temporary connection between the voice messaging application 208 at the electronic device 119 and another electronic device (e.g., the electronic device 110, the electronic device 115, or the electronic device 117) acting as a caller. During the electronic voice messaging session, the caller can record a voice message, which is then stored as local voice data 212 at the electronic device 119 and made available for dynamic display on the electronic device 119 for a user of the electronic device 119.
[0033] In one or more implementations, the voice messaging application 208 can store voicemail messages in a privacy-preserving manner on a cloud network and synchronize the voicemail messages, enabling a user to access their voicemail messages on multiple devices. By utilizing a cloud network, the voicemail messages can be securely stored in the cloud network and synchronized across various devices associated with the user’s account. The synchronization process ensures that the voicemail messages are consistently updated and available for retrieval, providing a seamless and unified voicemail experience across multiple devices.
[0034] Figure 3An input component 216 for receiving input from a user of electronic device 119 is also shown. For example, the electronic device may provide input options such as telephone call handover options (e.g., for switching from an electronic voice messaging session to a voice communication session via telephone application 214). When a user of electronic device 119 selects a telephone call handover option via input component 216, a call whose audio stream, intercepted by voice messaging application 208 to capture and record voice messages, can be transferred to telephone application 214, for example, via inter-process communication, and the audio stream of the call can be recovered via output component 204 through audio output, allowing the user of electronic device 119 to alternatively conduct live communication with the caller in a voice communication session.
[0035] Figure 3 A schematic diagram illustrating an exemplary user interface view according to one or more specific embodiments is shown, wherein a transcription is dynamically displayed on electronic device 119 during an electronic voice messaging session using a voice messaging application, such as voice messaging application 208 running at electronic device 119. Figure 3 In the example, for illustrative purposes, the dynamic presentation of the transcription is represented as scrolling text on the display of electronic device 119. For example... Figure 3 As shown, during an electronic voice messaging session, the voice messaging application can provide a scrolling transcript 350 for display.
[0036] exist Figure 3 In the example, the user background view 320 essentially covers the entire display of the electronic device 119, while the scrolling transcription 350 covers a portion of that display. However, this is merely illustrative and other arrangements of the user background view 320 and the scrolling transcription 350 are possible (e.g., two video stream views of the same size arranged side-by-side or vertically).
[0037] like Figure 3As shown, electronic device 119 may also provide input options such as telephone call handover option 340 (e.g., for switching from an electronic voice messaging session to a voice communication session via telephone application 214). When the user of electronic device 119 selects telephone call handover option 340, the call, whose audio stream is intercepted by voice messaging application 208 to capture and record voice messages, can be transferred to telephone application 214, and the audio stream of the call can be recovered via output component 204 through audio output, allowing the user of electronic device 119 to alternatively communicate live with the caller in a voice communication session. In one or more other specific implementations, the handover between an electronic voice messaging session and a voice communication session can be a two-step process: 1) wherein the voice messaging application 208 may first transmit a segment of the audio stream of the recorded voice message to the output unit 204 in response to the user of the electronic device 119 selecting the telephone call handover option 340, so that the user can determine the tone or context of the message before deciding whether to answer the call, thereby enabling the use of the electronic device 119 to make informed decisions about call acceptance based on audio transcription; and 2) the voice messaging application 208 facilitates the transition to a voice communication session with the telephone application 214 in response to the user confirming the selection of the telephone call handover option 340.
[0038] like Figure 3 As shown, during an electronic voice messaging session, the rolling transcription 350 can be displayed by the voice messaging application. Figure 2 In the example, rolling transcription 350 is the transcription of audio content received as input to an electronic device (e.g., electronic device 119) of a user of one of the electronic devices 110, 115, or 117. The transcription can be running transcription, which includes the text of segments (e.g., sentences, phrases, words, idioms, etc.) corresponding to the audio input of electronic device 119 (e.g., words spoken by another user into a microphone associated with one of the electronic devices 110, 115, or 117), and the text of each segment of the audio input displayed when (e.g., synchronized with) that other user speaks the segment during an electronic voice messaging session.
[0039] As further described in detail herein (e.g., in conjunction with) Figure 4 and Figure 4Electronic device 119 can also receive and display updates to the rolling transcript 350 during an electronic voice messaging session. For example, while a fragment of the transcript is still displayed in the rolling transcript 350, the electronic device that generated the transcript (e.g., electronic device 119) can generate an update to that fragment of the transcript (e.g., a correction of the fragment of the transcript based on an improved confidence level of the update, such as using words or other contextual improvements received after receiving the audio corresponding to the fragment), and provide the update for display on electronic device 119. For example, electronic device 119 can modify the currently displayed fragment of the transcript in the rolling transcript 350 according to the update. The update may change one or more words in the fragment to updated words that are more meaningful in the overall transcript of the fragment.
[0040] In various examples, transcriptions may be generated in response to a reduction in bandwidth of a voice communication session via telephone application 214. For example, one or more electronic devices and / or servers (e.g., server group 160) relaying information about a voice communication session may determine that the bandwidth of one or more electronic devices has become too low to exchange audio and / or video data, and may provide transcriptions in place of the audio and / or video data (e.g., until an increase in bandwidth is detected).
[0041] Figure 1 A flowchart illustrating an example process 400 for providing dynamic presentation of transcription during an electronic voice messaging session, according to one or more specific implementations, is provided. For illustrative purposes, this document primarily refers to... Figure 1 The process 400 is described by the components (especially reference electronic device 117), which can be described by Figure 2 The process 400 is executed by one or more processors of electronic device 117. However, process 400 is not limited to electronic device 117, and one or more blocks (or operations) of process 400 may be executed by other suitable devices (such as one or more electronic devices among electronic devices 110, 115, and 119) and / or one or more other components of servers (such as server 120 and / or server 130). Further for illustrative purposes, the blocks of process 400 are described herein as occurring sequentially or linearly. However, multiple blocks of process 400 may occur in parallel. Moreover, the blocks of process 400 need not be executed in the order shown, and / or one or more blocks of process 400 need not be executed and / or may be replaced by other operations.
[0042] In the example process 400, at block 402, during an electronic voice messaging session between a first device (e.g., electronic device 119) and a second device (e.g., one of electronic device 110, electronic device 115, or electronic device 117), the first device receives audio input corresponding to audio generated at the second device. For example, the first device can receive audio input that can correspond to a user of the second device speaking into a microphone of the second device (or connected to the second device). For example, the electronic voice messaging session can be an audio voice messaging session, such as a call intercepted by a voice messaging application (e.g., as described above in connection with voice messaging application 208) and prompting a user of the second device to record a voice message. In one or more implementations, during the electronic voice messaging session between the first device and the second device, the first device can determine whether the audio input corresponds to an unknown user of the second device. In some aspects, a transcription of the audio input can be generated in response to determining that the audio input corresponds to an unknown user of the second device. In some aspects, the audio input is received from the second device over a wireless network. For example, the wireless network can be a cellular network. Figure 2 In one or more implementations, during the electronic voice messaging session between the first device and the second device, the first device can determine whether the audio input corresponds to an unknown user of the second device. In some aspects, a transcription of the audio input can be generated in response to determining that the audio input corresponds to an unknown user of the second device. In some aspects, the audio input is received from the second device over a wireless network. For example, the wireless network can be a cellular network.
[0043] At block 404, during the electronic voice messaging session between the first device and the second device, the first device can generate a transcription of the audio input. For example, the first device can use a transcription service (e.g., as described above in connection with transcription service 210) at the first device to generate the transcription of the audio input. In one or more implementations, the transcription is associated with a confidence score. In some aspects, the confidence score can indicate a likelihood that the transcription represents all of the content in the audio input. In one or more other implementations, the first device can determine whether the confidence score exceeds a confidence threshold during the electronic voice messaging session between the first device and the second device. In some aspects, the transcription is provided for display on the first device based on determining that the confidence score exceeds the confidence threshold. Figure 5
[0044] At block 406, during the electronic voice messaging session between the first device and the second device, the first device can provide the transcription for display on the first device. In one or more implementations, the first device can transmit the transcription of the audio input and an audio stream corresponding to the audio input to a third device associated with a user of the first device during the electronic voice messaging session to display or store the transcription and the audio stream at the third device. In one or more other implementations, the first device can tag the transcription with an indication that display of the transcription is suppressed at the third device based on a device type of the third device.
[0045] In one or more implementations, the first device can receive user input during the electronic voice messaging session between the first device and the second device and in response to the transcription being displayed on the first device, the user input indicating a request to transition from the electronic voice messaging session to a voice communication session with the second device. In one or more other implementations, the first device can provide, during the electronic voice messaging session between the first device and the second device and in response to a request to transition from the electronic voice messaging session to a voice communication session with the second device, an audio stream corresponding to at least a portion of the audio input for output on the first device prior to the transition. The first device can receive user input in response to the audio stream being provided to an output device of the first device, the user input indicating confirmation of the request to transition to the voice communication session with the second device.
[0046] Generating a transcription at an electronic device that receives audio input (e.g., as opposed to transmitting an audio stream for transcription at a server or other external transcription service) can be advantageous because local speech data corresponding to a speaker of the audio input can be obtained, learned, and / or stored by the electronic device that receives the audio input and used to improve audio transcription. Because this local speech data is maintained and protected for privacy at the electronic device, the privacy of the user, speaker, to whom the audio input belongs can be maintained while the local speech data of the user is utilized to improve the ability of the electronic device to generate accurate and / or complete transcriptions.
[0047] As described herein, aspects of the subject technology can include the collection and processing of privacy-sensitive data on a user's computing device. The present disclosure contemplates that in some instances, the collected data can include personal information data that uniquely identifies or can be linked to a specific person. Such personal information data can include demographic data, location-based data, online identifiers, telephone numbers, email addresses, voice recordings, audio recordings, video recordings, home addresses, images, data or records relating to a user's health or fitness level, date of birth, or any other personal information.
[0048] The present disclosure recognizes that the use of such personal information data, in the present technology, can be used to the benefit of users. For example, personal information data can be used to provide transcriptions for video voice messaging sessions. Further, other uses for personal information data that benefit the user are also contemplated. For example, health and fitness data can be used in accordance with a user's preferences to provide insights into their overall well-being, or can be used as positive feedback to individuals using technology to pursue health goals.
[0049] The present disclosure contemplates that the entities responsible for the collection, analysis, disclosure, transfer, storage, or other use of such personal information data will comply with well-established privacy policies and / or privacy practices. In particular, such entities should implement and consistently apply privacy practices that are generally recognized as meeting or exceeding industry or governmental requirements for maintaining privacy and security, which are not less than industry standard privacy policies and procedures. Such information should be prominently disclosed, in a way that is easily accessible and understood, and updated as necessary, to users, consumers, or visitors of the entity. Users should have the opportunity to opt in or opt out of data collection, sharing, use, or storage, and should be able to review, modify, or delete their personal information data as desired. In addition, such entities should take any necessary steps to protect and secure such personal information data consistent with the maintenance and protection of their own confidentiality, privacy, and security policies and / or procedures, including implementing appropriate technical, administrative and physical procedures to protect personal information data from loss, misuse, and / or unauthorized alteration or destruction.
[0050] Notwithstanding the foregoing, the present disclosure also contemplates implementations in which users selectively block the use of, or access to, personal information data. That is, the present disclosure contemplates that hardware and / or software elements can be provided to prevent or block access to such personal information data. For example, in the case of transcribed video voice messaging, the present technology can be configured to allow users to select to "opt in" or "opt out" of participation in the collection of personal information data during registration for services or anytime thereafter. In addition to providing the "opt in" and "opt out" option, the present disclosure contemplates providing notifications relating to the access or use of personal information. For instance, a user can be notified upon download of an application that their personal information data will be accessed, and then again notified prior to the application accessing the personal information data.
[0051] Moreover, the present disclosure is intended to apply to any sort of personal information data, such as market segmentation and other demographic information, geo-location data, email addresses, and transaction history. As such, various embodiments can be used in the context of online or offline businesses, websites, mobile applications, etc., that collect personal information data from users. For example, businesses or websites can implement one or more of the various embodiments described herein to better understand users and / or for security purposes.
[0052] Accordingly, while the present disclosure broadly covers techniques employing personal information data, the present disclosure also contemplates embodiments in which user data is not used in some manner, or is minimally used. In particular, various embodiments contemplate the use of generic data, such as generic demographic information, generic geographic information, and / or the like in the context of the various embodiments.
[0053] Figure 1 An electronic system 500 is illustrated that can be used to implement one or more specific implementations of the subject technology. The electronic system 500 can be Figure 5 The electronic system 500 can be, and / or can include, the electronic device 110, the electronic device 115, the electronic device 117, the electronic device 119, the server 120, and / or the server 130 shown in FIG. 1, and / or can be part of the electronic device 110, the electronic device 115, the electronic device 117, the electronic device 119, the server 120, and / or the server 130. The electronic system 500 can include various types of computer readable media and interfaces for various other types of computer readable media. The electronic system 500 includes a bus 508, one or more processing unit(s) 512, a system memory 504 (and / or a buffer), a ROM 510, a permanent storage device 502, an input device interface 514, an output device interface 506, and one or more network interfaces 516, or a subset or variation thereof.
[0054] The bus 508 generally represents any communication fabric, and in one or more implementations, the bus 508 communicates internal system devices. In one or more implementations, the bus 508 communicatively connects the one or more processing units 512 with the ROM 510, the system memory 504, and the permanent storage device 502. The one or more processing units 512 retrieve instructions from these various memory units and process data to execute processes of the subject disclosure. In different implementations, the one or more processing units 512 can be a single processor or multiple processors.
[0055] ROM 510 stores static data and instructions that are needed by the one or more processing units 512 and other modules of the electronic system 500. Permanent storage device 502, on the other hand, can be a read-and-write memory device. Permanent storage device 502 can be a non-volatile memory unit that even when electronic system 500 is off, it stores instructions and data for the electronic system 500. In one or more implementations, a mass storage device (such as a magnetic or optical disk and its corresponding disk drive) can be used as the permanent storage device 502.
[0056] In one or more implementations, a removable storage device (such as a floppy disk drive and its corresponding floppy disk or CD-ROM drive and its corresponding CD-ROM) can be used as the permanent storage device 502. Like permanent storage device 502, system memory 504 can be a read-and-write memory device. However, unlike permanent storage device 502, system memory 504 can be a volatile read-and-write memory, such as a random access memory. System memory 504 can store any of the instructions and data that one or more processing units 512 can need at runtime. In one or more implementations, the processes of the subject disclosure are stored in the system memory 504, the permanent storage device 502, and / or the ROM 510. From these various memory units, processing unit(s) 512 retrieves instructions to execute and data to process in order to execute the processes of one or more implementations.
[0057] Bus 508 also connects to input and output devices 514 and 506. Input device interface 514 enables the user to communicate information and select commands to the electronic system 500. Input devices that can be used with input device interface 514 include, for example, alphanumeric keyboards and pointing devices (also referred to as “cursor control devices”). Output device interface 506 can, for example, enable a display of images generated by the electronic system 500. Output devices that can be used with output device interface 506 include, for example, printers and display devices, such as a liquid crystal display (LCD), a light emitting diode (LED) display, an organic light emitting diode (OLED) display, a flexible display, a flat-panel display, a solid-state display, a projector, or any other device for outputting information. One or more implementations can include devices that function as both input and output devices, such as a touchscreen. In these implementations, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0058] Finally, as Figure 1 shown in FIG. 5, bus 508 also couples electronic system 500 to one or more networks and / or one or more network nodes, such as network 520, through one or more network interfaces 516. The electronic system 500 can be part of a computer network, such as a LAN, a wide-area network ("WAN"), or an Intranet, or can be part of one of the networks, such as the Internet. Any or all components of the electronic system 500 can be used in connection with the subject disclosure.
[0059] According to various aspects of the subject disclosure, an apparatus is provided that includes a memory and one or more processors configured to, during an electronic voice messaging session between at least a first device and a second device: receive, by the electronic device, a first audio input; generate a first transcription of the first audio input; and transmit the first transcription from the electronic device to the other device; and during the electronic voice messaging session and after transmitting the first transcription: receive a second audio input; generate a second transcription of the second audio input; and transmit the second transcription to the other device.
[0060] According to various aspects of the subject disclosure, a non-transitory computer-readable medium including instructions that, when executed by one or more processors, cause the one or more processors to perform operations including: during an electronic voice messaging session between at least a first device and a second device: receiving, by the first device, a first audio input; generating, by the first device, a first transcription of the first audio input; and transmitting the first transcription from the first device to the second device; and during the electronic voice messaging session and after transmitting the first transcription: receiving, by the first device, a second audio input; generating, by the first device, a second transcription of the second audio input; and transmitting the second transcription from the first device to the second device.
[0061] According to various aspects of the subject disclosure, a method is provided that includes: during an electronic voice messaging session between at least a first device and a second device: receiving, by the first device, a first audio input; generating, by the first device, a first transcription of the first audio input; and transmitting the first transcription from the first device to the second device; and during the electronic voice messaging session and after transmitting the first transcription: receiving, by the first device, a second audio input; generating, by the first device, a second transcription of the second audio input; and transmitting the second transcription from the first device to the second device.
[0062] Implementations within the scope of the disclosure can be realized, in part, by tangible computer-readable storage media (or multiple tangible computer-readable storage media of one or more types) encoding one or more instructions. Tangible computer-readable storage media substantially retains data for a period of time (e.g., at least several seconds, at least a minute, at least an hour, at least a day, or at least a week), and thus is non-transitory.
[0063] A computer-readable storage medium can be any storage medium that can be read, written, or otherwise accessed by a general or special purpose computing device, including any processing electronics and / or processing circuitry capable of executing instructions. For example, without limitation, a computer-readable medium can include any volatile semiconductor memory, such as RAM, DRAM, SRAM, T-RAM, Z-RAM, and TTRAM. A computer-readable medium can also include any non-volatile semiconductor memory, such as ROM, PROM, EPROM, EEPROM, NVRAM, flash, nvSRAM, FeRAM, FeTRAM, MRAM, PRAM, CBRAM, SONOS, RRAM, NRAM, racetrack memory, FJG, and Millipede memory.
[0064] Furthermore, a computer-readable storage medium can include any non- semiconductor memory, such as optical disk storage, magnetic disk storage, magnetic tape, other magnetic storage devices, or any other storage medium that can store one or more instructions. In one or more implementations, a tangible computer-readable storage medium can be directly coupled to a computing device, while in other implementations, a tangible computer-readable storage medium can be indirectly coupled to a computing device, e.g., via one or more wired connections, one or more wireless connections, or any combination thereof.
[0065] Instructions can be directly executable, or can be used to develop executable instructions. For example, instructions can be implemented as executable or non-executable machine code, or as high-level language instructions that can be compiled to produce executable or non-executable machine code. Furthermore, instructions can also be implemented as data, or can include data. Computer-executable instructions can also be organized in any format, including routines, subroutines, programs, data structures, objects, modules, applications, applets, functions, etc. As one of skill in the art will recognize, details including, but not limited to, the number, structure, sequence, and organization of instructions can vary significantly without changing the underlying logic, function, processing, and output.
[0066] While the above discussion primarily refers to microprocessor or multi-core processors that execute software, one or more implementations are performed by one or more integrated circuits such as ASICs or FPGAs. In one or more implementations, such integrated circuits execute instructions stored on the integrated circuits themselves.
[0067] Those skilled in the art will recognize that the various illustrative blocks, modules, elements, components, methods, and algorithms described herein can be implemented as electronic hardware, computer software, or combinations of both. To illustrate the interchangeability of hardware and software, various illustrative blocks, modules, elements, components, methods, and algorithms have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans can implement the described functionality in varying ways for each particular application. Various components and blocks can be arranged differently or eliminated from the
[0068] It should be understood that any particular order or hierarchy of blocks within the processes disclosed is an example of an illustrative method. Other specific orders and hierarchies of blocks can be implemented depending on the design choices of a given implementation. For example, blocks can be reordered or eliminated, and other blocks can be added. Any of the blocks can be performed simultaneously or in different orders than as described. In one or more embodiments, tasks and processes described herein can be performed in parallel or in different orders. Furthermore, the division of various system components described above should not be understood as requiring such division in all implementations, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0069] As used in this specification and any claims of this application, the terms "base station," "receiver," "computer," "server," "processor," and "memory" all refer to electronic or other technological devices. These terms exclude people or groups of people. For purposes of this specification, the term "display" or "displaying" means displaying on an electronic device.
[0070] As used herein, the phrase "at least one of a list of items" following the term "comprising" or "including" or other variations such as "containing," "satisfying," and "including," presents an example of a list of items that can be present in a composition, a method, a process, a machine, a manufacture, or an article of manufacture, without limitation to only those items recited. The phrase "at least one of a list of items" does not mean "only one of the items in the list can be employed." It means one or more items of the list of items will be employed. Each item can be employed by itself or in a combination with one or more other items. For example, "A and B" is an example of a list of items that can be present and means A or B or any combination of A and B. "At least one of A and B" thus means, A1B, A, B, or A and B.
[0071] The predicate words “configured to,” “operable to,” and “programmed to” do not imply any specific tangible or intangible modification of a subject, but, rather, are intended to be synonymous with the phrase “caused to,” which means that the subject is caused to operate in a certain manner, either by a device or a method. In one or more specific embodiments, a processor configured to monitor and control operations or components can also mean the processor is programmed to monitor and control the operations or the processor is operable to monitor and control the operations. Likewise, a processor configured to execute code can be interpreted as a processor programmed to execute code or a processor operable to execute code.
[0072] The phrases “one or more of the following aspects,” “some of the aspects,” “one or more aspects,” “one or more of the aspects,” “some aspects,” “one or more of the embodiments,” “some embodiments,” “one or more embodiments,” “one or more of the implementations,” “some implementations,” “one or more implementations,” “aspects,” “the aspect,” “another aspect,” “some aspects,” “one or more aspects,” “the implementation,” “the implementations,” “another implementation,” “some implementations,” “one or more implementations,” “the embodiment,” “the embodiments,” “another embodiment,” “some embodiments,” “one or more embodiments,” “configuration,” “the configuration,” “another configuration,” “some configurations,” “one or more configurations,” “subject technology,” “disclosure,” “the disclosure,” “other variations thereof,” and the like are used herein to facilitate reading the specification from a first paragraph to a last paragraph. The disclosure involved with such phrases can apply to all configurations or one or more configurations. The disclosure involved with such phrases can provide one or more examples. The phrase one or some aspects can refer to one or more aspects and vice versa, and similarly for other aforementioned phrases.
[0073] The word “exemplary” is used herein to mean “serving as an example, instance, or illustration.” Any implementation described herein as “exemplary” or as an “example” is not necessarily to be construed as preferred or advantageous over other implementations. Likewise, the term “includes” and its variants are used inclusively and therefore do not exclude additional, unrecited elements or method steps. The term “coupled” as used herein is intended to mean two or more elements interacting in some manner.
[0074] All structural and functional equivalents to the elements of the various aspects described throughout this disclosure that are known or later come to be known to those of ordinary skill in the art are expressly incorporated herein by reference and intended to be encompassed by the claims. Moreover, nothing disclosed herein is intended to be dedicated to the public regardless of whether these disclosures are explicitly recited in the claims. No claim element is to be construed under the provisions of 35 U.S.C. § 112(f) unless the element is expressly recited in the claim as “means plus function.” The terms “a” or “an,” as used herein, mean “one or more” unless otherwise explicitly recited in the claim.
[0075] The preceding description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other aspects. Thus, the claims are not intended to be limited to the aspects shown herein, but is to be accorded the full scope consistent with the language claims, wherein reference to an element in the singular is not intended to mean "one and only one" unless specifically so stated, and wherein advantages offered herein can be used alone or in combination. Unless otherwise specifically noted, the term "some" refers to one or more. Pronouns in the masculine (his) include the feminine (her) and the neuter (its) and vice versa, and the use of "one" includes "more than one." The use of the term "about" accompanying an entity or entities modifies that entity or entities to encompass roughly or approximately the value attributed to the entity or entities. Headings and subheadings, if any, are used for convenience only and do not limit the subject disclosure.
Claims
1. A method, the method comprising: During the electronic voice messaging session between the first and second devices: The first device receives an audio input corresponding to the audio generated at the second device; The first device generates a transcription of the audio input; as well as The transcription is provided for display on the first device.
2. The method according to claim 1, further comprising: During the electronic voice messaging session between the first device and the second device, the first device determines whether the audio input corresponds to an unknown user of the second device, wherein the transcription of the audio input is generated in response to determining that the audio input corresponds to an unknown user of the second device.
3. The method of claim 1, wherein the audio input is received from the second device via a wireless network.
4. The method of claim 3, wherein the wireless network is a cellular network.
5. The method according to claim 1, further comprising: During the electronic voice messaging session, the transcription of the audio input and the audio stream corresponding to the audio input are transmitted from the first device to a third device associated with the user of the first device, so as to display or store the transcription and the audio stream at the third device.
6. The method according to claim 5, further comprising: The transcription is labeled with an indicator that makes the transcription appear suppressed at the third device, based on the device type of the third device.
7. The method of claim 1, wherein the transcription is associated with a confidence score, the confidence score indicating the likelihood that the transcription represents the full content of the audio input, the method further comprising: During the electronic voice messaging session between the first device and the second device, the first device determines whether the confidence score exceeds a confidence threshold, wherein the transcription is provided for display on the first device based on the determination that the confidence score exceeds the confidence threshold.
8. The method according to claim 1, further comprising: During the electronic voice messaging session between the first device and the second device, the first device receives user input in response to the transcription being displayed on the first device, the user input indicating a request to switch from the electronic voice messaging session to a voice communication session with the second device.
9. The method according to claim 8, further comprising: During the electronic voice messaging session between the first device and the second device: In response to the request to switch from the electronic voice messaging session to the voice communication session with the second device, an audio stream corresponding to at least a portion of the audio input prior to the switch is provided to the output device of the first device; as well as The first device receives user input in response to the audio stream being provided to the output device of the first device, the user input indicating confirmation of the request to switch to the voice communication session with the second device.
10. An electronic device, the electronic device comprising: Memory; and One or more processors, said one or more processors being configured to: During the electronic voice messaging session between the first and second devices: The first device receives an audio input corresponding to the audio generated at the second device; The first device generates a transcription of the audio input; as well as The transcription is provided for display on the first device.
11. The electronic device of claim 10, wherein the one or more processors are further configured to: during the electronic voice messaging session between the first device and the second device, determine by the first device whether the audio input corresponds to an unknown user of the second device, wherein the transcription of the audio input is generated in response to determining that the audio input corresponds to an unknown user of the second device.
12. The electronic device of claim 10, wherein the audio input is received from the second device via a wireless network.
13. The electronic device of claim 12, wherein the wireless network is a cellular network.
14. The electronic device of claim 10, wherein the one or more processors are further configured to: during the electronic voice messaging session, transmit the transcription of the audio input and an audio stream corresponding to the audio input from the first device to a third device associated with a user of the first device, so as to display or store the transcription and the audio stream at the third device.
15. The electronic device of claim 14, wherein the one or more processors are further configured to: mark the transcription with an indication that causes the transcription to be suppressed at the third device based on the device type of the third device.
16. The electronic device of claim 10, wherein the transcription is associated with a confidence score indicating the likelihood that the transcription represents the full content of the audio input, wherein the one or more processors are further configured to: during the electronic voice messaging session between the first device and the second device, determine by the first device whether the confidence score exceeds a confidence threshold, wherein the transcription is provided for display on the first device based on the determination that the confidence score exceeds the confidence threshold.
17. The electronic device of claim 10, wherein the one or more processors are further configured to: during the electronic voice messaging session between the first device and the second device, receive user input by the first device in response to the transcription being displayed on the first device, the user input indicating a request to switch from the electronic voice messaging session to a voice communication session with the second device.
18. The electronic device of claim 17, wherein the one or more processors are further configured to: during the electronic voice messaging session between the first device and the second device: In response to the request to switch from the electronic voice messaging session to the voice communication session with the second device, an audio stream corresponding to at least a portion of the audio input prior to the switch is provided; and The first device receives user input indicating confirmation of the request to switch to the voice communication session with the second device.
19. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations including: During the electronic voice messaging session between the first and second devices: The first device receives an audio input corresponding to the audio generated at the second device; The first device generates a transcription of the audio input; as well as The transcription is provided for display on the first device.