Audio presentation of a conversation thread

By converting electronic communication into audible presentation, personal assistant devices solve the distraction problem of visual presentation, providing a safe and efficient solution to handle communications while performing other tasks.

CN113950698BActive Publication Date: 2025-07-22MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202080042680.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-06-10
Filing Date
2020-04-27
Publication Date
2025-07-22
Estimated Expiration
2040-04-27

AI Technical Summary

Technical Problem

In the prior art, visual presentation of text-based electronic communications can distract users in some cases, especially when performing other tasks, such as driving a vehicle, resulting in safety hazards.

Method used

By converting electronic communication into an audible presentation, the unviewed electronic communication is outputted in audio form with a personal assistant device, including summary and detailed content of conversation threads, providing controls to enhance the user experience, enabling the user to process communications while performing other tasks.

Benefits of technology

It realizes processing electronic communication without distraction, improves user task execution efficiency and security, and provides an experience similar to listening to podcasts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113950698B_ABST
    Figure CN113950698B_ABST
Patent Text Reader

Abstract

In one example, a computing system receives instructions to initiate an audio rendering of an electronic communication for a recipient. In response to the instructions, the computing system audibly outputs each unviewed electronic communication in the most recent conversation thread, the most recent conversation thread including a set of the most recent unviewed, reply-linked electronic communications for the recipient. Each unviewed electronic communication in the most recent conversation thread may be audibly output in chronological order, starting with the oldest unviewed electronic communication and continuing to the most recent unviewed electronic communication. In response to completing the audible output of the most recent unviewed electronic communications from the conversation thread, the computing device audibly outputs each unviewed electronic communication in the next most recent conversation thread, the next most recent conversation thread including a set of the next most recent unviewed, reply-linked electronic communications for the recipient. Each unviewed electronic communication in the next most recent conversation thread may be audibly output in chronological order, starting with the oldest unviewed electronic communication.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND OF THE INVENTION

[0001] Communication networks support a wide range of electronic communication between users. Electronic communication based on text can take many different forms, including email, text / SMS messages, real-time / instant messages, multimedia messages, social network messages, messages in multi-player video games, and so on. Users can read and type responses to these forms of electronic communication via a personal electronic device, such as a mobile device or a desktop computer. SUMMARY OF THE INVENTION

[0002] The present invention content is provided to introduce, in a simplified form, a selection of design concepts further described below in the detailed description. The present invention content is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Additionally, the claimed subject matter is not limited to implementations that solve any or all of the disadvantages noted in any part of the present disclosure.

[0003] In one example, a computing system receives instructions to initiate an audio presentation of an electronic communication to a recipient. In response to the instructions, the computing system audibly outputs each unread electronic communication in the most recent conversation thread, including a set of the most recent unread, reply-linked electronic communications to the recipient. Each unread electronic communication in the most recent conversation thread can be audibly output in chronological order, starting with the oldest unread electronic communication and continuing to the most recent unread electronic communication. In response to completing the audible output of the most recent unread electronic communications from the conversation thread, the computing device audibly outputs each unread electronic communication in the next most recent conversation thread, including a set of the next most recent unread, reply-linked electronic communications to the recipient. Each unread electronic communication in the next most recent conversation thread can be audibly output in chronological order, starting with the oldest unread electronic communication and continuing to the most recent unread electronic communication. BRIEF DESCRIPTION OF THE DRAWINGS

[0004] Figure 1 An example interaction between a user and a personal assistant device is depicted.

[0005] Figure 2 An example computing system is schematically depicted.

[0006] Figure 3 An example electronic communication is schematically depicted.

[0007] Figure 4 An example timeline of an audio presentation output by a personal assistant device is depicted.

[0008] Figure 5 An example timeline of an electronic communication to a recipient is depicted.

[0009] Figure 6 depicts Figure 5 an example timeline of the presentation order of electronic communications.

[0010] Figure 7A depicts a flowchart of an example method for organizing and presenting a conversation thread.

[0011] Figure 7B depicts a flowchart of an example method for presenting a conversation thread.

[0012] Figure 8 depicts a flowchart of an example method for presenting time identification information of a conversation thread.

[0013] Figures 9A - 9E depicts tables providing example audible outputs of a personal assistant device for a series of conditions.

[0014] Figures 10A - 10P depicts an example conversation between a user and a personal assistant device. DETAILED DESCRIPTION

[0015] The use of text-based electronic communications (such as email, text messaging, and instant messaging) has evolved into a primary means of communication in modern society. Mobile computing devices enable people to receive their electronic communications at almost any time and location. When people are in their daily lives, they may often be interrupted by the need or desire to view new electronic communications. The visual consumption of text and multimedia content through a graphical user interface may distract people from performing other tasks simultaneously or may prevent people from performing tasks until the electronic communication is visually viewed. For example, when operating a vehicle, a person may not be able to visually view new text-based communications, or it is dangerous.

[0016] According to one aspect of the present disclosure, the use of a graphical user interface to present the text and multimedia content of electronic communications can be enhanced or replaced by an audible presentation of the electronic communications in a manner that provides context for the presentation experience to the user and control over the audible presentation. Such an audible presentation can provide a user experience that is commensurate with or improved over the visual consumption of electronic communications, while enabling the user to perform tasks that are difficult or impossible to perform while using a graphical user interface. In essence, the disclosed audible presentation can transform text-based communications into an experience similar to listening to a podcast.

[0017] Figure 1Depicts an example interaction 100 between a user 110 and a personal assistant device 120. In this example, the user 110 is commuting to and from work by bike while interacting with the personal assistant device 120 via the user's voice 130. The personal assistant device 120 in this example takes the form of a mobile computing device. In response to an oral command of the user's voice 130, the personal assistant device 120 can output audio information to the user 110 as device voice 140. This is an example of a hands-free, display-free interaction that enables the user to handle electronic communications while engaged in a task (such as commuting to and from work by bike).

[0018] As shown in the user's voice 130, the user 110 starts a conversation with the personal assistant device 120 by saying the command "Read messages". In response to the user's oral command, in the device voice 140, the personal assistant device 120 outputs audio information that includes: "Hi Sam! I've got 6 conversations for you. This'll take about 5 minutes". In this part of the device voice 140, the personal assistant device 120 outputs audio information in the form of natural language that greets the user 110 by the user's name (i.e., "Sam"), identifies a number (i.e., "6") of conversation threads that contain the user's unread electronic communications, and identifies the duration (i.e., "about 5 minutes") of the user's audible output of the content of the electronic communications to view the conversation threads. Thus, before proceeding with the audio presentation, the personal assistant device 120 notifies the user 110 of the expected duration of the audio presentation of the unread electronic communications, enabling the user to make an informed decision about whether to view or skip a particular electronic communication.

[0019] Continuing Figure 1 the example conversation, the personal assistant device 120 continues by outputting a summary of the first conversation thread to the user 110 that identifies the number and / or duration (i.e., "long") of the unread electronic communications of the conversation thread and identifies the topic of the conversation thread (i.e., "World Cup office pool"). Thus, before proceeding with the audio presentation of the first conversation thread, the user 110 is informed of the topic and estimated time for viewing the unread electronic communications of the first conversation thread. Additionally, the personal assistant device 120 indicates to the user 110 that the user "can interrupt at will", which notifies the user that the user's oral command can be used to advance or interrupt the audio presentation of the first conversation thread.

[0020] Next, the personal assistant device 120 outputs a summary of the first electronic communication of the first conversation thread to the user 110, which identifies the relative time when the first electronic communication was received (i.e., "a few hours ago"), identifies the sender of the first electronic communication (i.e., "Greg"), identifies the type of the first electronic communication (i.e., "email"), identifies a certain number of other recipients or audiences of the first electronic communication (i.e., "a large group"), identifies the presence of an attachment to the first electronic communication (i.e., "with attachment"), and identifies at least a part of the text content of the message of the first electronic communication (e.g., "Goal! Can you believe it's already World Cup time?...").

[0021] In this example, upon hearing a part of the text content of the first electronic communication, in the user speech 130, the user 110 says the command "next conversation". In response to this oral command of the user 110, the personal assistant device 120 advances the audio presentation of the unviewed electronic communications to the second conversation thread, thus skipping the audio presentation of the remaining unviewed electronic communications of the first conversation thread. For example, the personal assistant device 120 responds to the user 110 by outputting a summary of the second conversation thread, which identifies the number of unviewed electronic communications of the second conversation thread (i.e., "3"), identifies the type of the electronic communications of the second conversation thread (i.e., "email"), and identifies the topic of the second conversation thread (i.e., "human resources event").

[0022] The personal assistant device 120 can progress through the conversation threads in the above manner until the user 110 has viewed all unviewed electronic communications or the user pre-selects to stop the conversation. By summarizing the conversation threads and their content by the personal assistant device 120, enough information is provided to the user 110 to make an informed decision on whether a particular conversation thread or electronic communication should be viewed by the user in the current session. In an example where the user 110 does not advance or interrupt the audio presentation of the unviewed electronic communications, the audio presentation of the personal assistant device 120 will end within approximately the duration identified by the personal assistant device (e.g., "5 minutes"). However, by advancing the audio presentation, the user 110 can view the electronic communications within a shorter time period.

[0023] Figure 2 An example computing system 200 is schematically depicted, which includes a computing device 210. As an example, the computing device 210 can take the form of a mobile computing device, a wearable computing device, a computing device integrated with a vehicle, a desktop computing device, a household appliance computing device, or other suitable devices. Figure 1 The personal assistant device 120 is an example of the computing device 210. The computing device 210 includes a logic subsystem 212, a storage subsystem 214, an input / output (I / O) subsystem 216, and / or Figure 2Other suitable components not shown.

[0024] The logical subsystem 212 includes one or more physical devices (e.g., processors) configured to execute instructions. The storage subsystem 214 includes one or more physical devices (e.g., memories) configured to store data 220, the data including instructions 222 executable by the logical subsystem 212 to implement the methods and operations described herein. Additional aspects of the logical subsystem 212 and the storage subsystem 214 are described below.

[0025] As Figure 2 shown, the logical subsystem 212 and the storage subsystem 214 can cooperate to instantiate one or more functional components, such as a personal assistant 230, a voice outputter 232, a voice inputter 234, one or more communication applications 236, and / or other suitable components. As used herein, the term "machine" can be used to collectively refer to a combination of instructions 222 (e.g., firmware, software, etc.) and hardware and / or other suitable components that cooperate to provide the described functionality. Although the personal assistant 230, the voice outputter 232, the voice inputter 234, and / or the communication application 236 are described as being instantiated through the cooperation of the logical subsystem 212 and the storage subsystem 214, in at least some examples, one or more of the personal assistant 230, the voice outputter 232, the voice inputter 234, and / or the communication application 236 can be instantiated in whole or in part by a remote computing device or system (e.g., a server system 260). Accordingly, the methods or operations described herein can be executed locally at the computing device 210, remotely at the server system 260, or can be distributed between one or more computing devices 210 and / or one or more server systems 260.

[0026] The personal assistant 230 can communicate with a user by receiving and processing the user's spoken commands to perform tasks, including outputting information to the user. As an example, the personal assistant 230 can output an audio presentation of multiple conversation threads and / or electronic communications for a recipient in a presented order. The personal assistant 230 can include natural language processing, thereby supporting a natural language interface through which a user can interact with the computing device 210. A device implementing the personal assistant 230, such as the computing device 210, can be referred to as a personal assistant device.

[0027] The voice output machine 232 receives data (such as machine-readable data and / or text-based data) from the personal assistant machine 230 for output to the user, and converts this data into audio data containing speech with natural language components. In one example, the voice output machine 232 can provide text-to-speech conversion. For example, the personal assistant machine 230 can provide a selected portion of the text content of an electronic communication to the voice output machine 232 to convert the text content into an audible output of the text content for the user's audible consumption. For example, in Figure 1 , the personal assistant device 120 outputs "GOAL! Can you believe it’s already World Cup time?", which is an audible output of the text content of an electronic communication of which the user 110 is the recipient.

[0028] The voice input machine 234 receives audio data representing human speech and converts the audio data into machine-readable data and / or text data that can be used by the personal assistant machine 230 or other suitable components of the computing device 210. In one example, the voice input machine 232 can provide speech-to-text conversion. For example, in Figure 1 , the personal assistant device receives and processes the spoken commands of the user 110 via the voice input machine 234, including "Read messages" and "Next conversation".

[0029] One or more communication applications 236 can support the sending and receiving of electronic communications 238, where electronic communication 240 is an example. The communication application can support one or more types of electronic communications, including email, text / SMS messages, real-time / instant messages, multimedia messages, social network messages, messages in multiplayer video games, and / or any other type of electronic communication. The personal assistant machine 230 can interface with the communication application 236 to enable the personal assistant machine to receive, process, and send one or more different types of electronic communications on behalf of the user.

[0030] The I / O subsystem 216 can include one or more of the following: an audio input interface 250, an audio output interface 252, a display interface 254, a communication interface 256, and / or other suitable interfaces.

[0031] The computing device 210 receives audio data representing audio captured via the audio input interface 250. The audio input interface 250 can include one or more integrated audio microphones and / or can interface with one or more peripheral audio microphones. For example, the computing device 210 can receive audio representing the user's speech captured via the audio input interface 250 (such as Figure 1The audio data of the user speech 130). The audio data from the audio input interface 250 can be provided to the speech input machine 234 and / or the personal assistant machine 230 for processing. The audio input interface 250 can be omitted in at least some examples.

[0032] The computing device 210 outputs audio representing the audio data via the audio output interface 252. The audio output interface 252 can include one or more integrated audio speakers and / or can dock with one or more peripheral audio speakers. For example, the computing device 210 can output an audio representation of speech with natural language components via the audio output interface 252, such as Figure 1 the device speech 140. The audio data can be provided by the speech output machine 232, the personal assistant machine 230, or other suitable components of the computing device 210 to the audio output interface 252 for output as an audible output of the audio data. The audio output interface 252 can be omitted in at least some examples.

[0033] The computing device 210 can output graphical content representing the graphical data via the display interface 254. The display interface 254 can include one or more integrated display devices and / or can dock with one or more peripheral display devices. The display interface 254 can be omitted in at least some examples.

[0034] The computing device 210 can communicate with other devices (such as the server system 260 and / or other computing devices 270) via the communication interface 256, enabling the computing device 210 to send electronic communications to other devices and / or receive electronic communications from other devices. The communication interface 256 can include one or more integrated transceivers and associated communication hardware that support wireless and / or wired communication according to any suitable communication protocol. For example, the communication interface 256 can be configured to communicate via a wireless or wired telephone network and / or a wireless or wired personal area network, local area network, and / or wide area network (such as the Internet, a cellular network, or a portion thereof) through the communication network 280. The communication interface 256 can be omitted in at least some examples.

[0035] The I / O subsystem 216 can also include one or more additional input devices and / or output devices in integrated and / or peripheral form. Additional examples of input devices include user input devices (such as keyboards, mice, touchscreens, touchpads, game controllers, etc.), and / or inertial sensors, global positioning sensors, cameras, optical sensors, etc. Additional examples of output devices include vibration motors and light indicators.

[0036] The computing system 200 may also include a server system 260 of one or more server computing devices. The computing system 200 may also include a plurality of other computing devices 270, of which the computing device 272 is an example. The server system 260 may host a communication service 262 that receives, processes, and transmits electronic communications between and among senders and receivers addressed by the electronic communications. For example, a user may operate the computing devices 210 and 270 to send or receive electronic communications via the communication service 262. The communication service 262 is depicted as including a plurality of electronic communications 264, of which the electronic communication 266 is an example. In one example, an electronic communication 266 may be received from the computing device 272 via the network 280 for processing and / or transmitted to the computing device 210 via the network 280. One or more communication applications 236 may be configured to operate in coordination with the communication service 262 such that electronic communications can be sent, received, and / or processed for senders and receivers who are users of the computing devices 210 and 270.

[0037] Figure 3 An example electronic communication 300 is schematically depicted. Figure 2 The electronic communications 240 and 266 are examples of the electronic communication 300. In one example, the electronic communication 300 takes the form of data that includes or identifies a sender 310, one or more receivers 312, a timestamp 314 indicating the timing of receipt or transmission of the electronic communication (e.g., clock time and date of transmission or receipt), a subject 316 that may include text content 318, a message 320 that may include text content 322 and / or media content 324, one or more attachments 326, calendar data 328, a communication type 330, and / or other data 332. The electronic communication 300 is provided as a non-limiting example. The present disclosure is compatible with almost any type of electronic communication, regardless of the content of the electronic communication that may be specific to that type of electronic communication. Thus, various aspects of the electronic communication may optionally be omitted, and / or various aspects not shown may be included.

[0038] In an example, a user acting as the sender of the electronic communication 300 may define one or more of the following via user input: the receiver 312, the subject 316 that includes the text content 318, the message 320 that includes the text content 322 and / or media content 324, the attachment 326, the calendar data 328, and / or other data 332 of the electronic communication 300. The timestamp 314 may be assigned by the communication application or communication service as the timing of transmission or receipt of the electronic communication 300. The communication type 330 may depend on the communication application or service used by the sender, or may be defined or otherwise selected by the user input of the sender in the case of a communication application or service that supports multiple communication types.

[0039] Figure 4 depicts an example timeline 400 of an audio presentation output by a personal assistant device (such as Figure 1 device 120 or Figure 2 computing device 210) as disclosed herein. Within timeline 400, time progresses from the left - hand side of the figure to the right - hand side of the figure. Timeline 400 can be instantiated from a predefined template that can be implemented by the personal assistant devices disclosed herein. Thus, in other examples, the audible outputs described for timeline 400 can be omitted, repeated, or presented in a different order. Additionally, additional audible outputs can be included instead of or between the audible outputs of timeline 400.

[0040] At 410, a greeting can be presented as an audible output. In one example, the greeting can be presented in response to an instruction 412 received by the personal assistant device to initiate the presentation of unviewed electronic communications to a recipient. Instruction 412 can take the form of a user's spoken command or other types of user input received by the personal assistant device. For example, in Figure 1 , user 110 provides the instruction "Read messages" as a spoken command, and personal assistant device 120 responds by presenting the greeting "Hi Sam!"

[0041] At 414, a roadmap presentation can be presented as an audible output. The roadmap presentation can identify one or more of the following: the number of conversation threads including one or more unviewed electronic communications for the recipient, the number of unviewed electronic communications, an estimated time for presenting an audio presentation of the conversation threads including unviewed electronic communications, an estimated length of the unviewed electronic communications, one or more highlighted items, and / or other suitable information.

[0042] At 416, a barge - in notice can be presented as an audible output. The barge - in notice can be used to inform the user that the user can provide a spoken command to perform an action on the audio presentation or its content. Referring to Figure 1 for an example, the personal assistant device can present the audible output "Feel free to interrupt" as an example of the barge - in notice presented at 416.

[0043] At 418, one or more changes to the user's day can be presented as an audible output. Changes to the day can include updates to the user's calendar and optionally can be derived from the calendar data of one or more unviewed electronic communications.

[0044] As referenced in Figure 5- As further described in FIG. 7, the recipient's electronic communications can be organized into conversation threads, where each conversation thread includes two or more electronically communicated reply links. By organizing electronic communications into conversation threads, a user listening to the audio presentation of the electronic communications may be able to better understand or track the conversation between or among the senders and recipients of the electronic communications that form part of the same conversation thread. In contrast, presenting electronic communications only in chronological order without considering the context of the conversation may make it more difficult for the user to understand or track the conversation between or among the senders and recipients, especially in the context of an audio presentation of such communications.

[0045] A first conversation thread including one or more unviewed electronic communications of the user can be presented at 470, which includes a conversation thread summary 420 of the first conversation thread, a communication summary 422 of each unviewed electronic communication of the first conversation thread, and a message content 424 of each unviewed electronic communication of the first conversation thread.

[0046] At 420, the conversation thread summary of the first conversation thread can be presented as an audible output. The conversation thread summary can identify one or more of the following: the topic of the conversation thread identified from the electronic communications of the conversation thread, the type of the electronic communications of the conversation thread, the number of unviewed electronic communications of the conversation thread, the recipient and / or audience of the conversation thread identified from the electronic communications of the conversation thread (e.g., the number, identity, and / or the number / identity of added or removed recipients related to the previously replied link communications), an estimated time for presenting a portion of the audio presentation of the unviewed electronic communications of the conversation thread, an estimated length of the unviewed electronic communications of the conversation thread, and / or other suitable information.

[0047] Reference Figure 9C Examples of the output of the personal assistant device regarding the number of unviewed electronic communications of the conversation thread are described in more detail. Reference Figure 9A and Figure 9E Examples of the output of the personal assistant device regarding the time and / or length of the conversation thread and / or electronic communications are described in more detail. In the example, the time and / or length estimate of the conversation thread summary can include a length warning. Reference Figure 1 In an example, the personal assistant device can present an audible output "long conversation" as an example of the length warning.

[0048] At 422, a communication summary of a first unviewed electronic communication of a first conversation thread can be presented as an audible output. The communication summary can identify one or more of the following: the subject of the electronic communication, the type of the electronic communication, the timing of the electronic communication based on the timestamp of the electronic communication, the sender, recipient, and / or audience of the electronic communication, an estimate of the time for presenting a portion of the audio presentation of the electronic communication, an estimate of the length of the electronic communication, an indication of whether one or more attachments are included in the electronic communication, and / or other suitable information. Refer to Figure 9B Example outputs of the personal assistant device regarding the recipient and / or audience of the conversation thread are described in more detail.

[0049] At 424, the message content of the first unviewed electronic communication of the first conversation thread can be presented as an audible output. For example, at 424, an audible output of the text content of the message of the first unviewed electronic communication can be presented in part or in full. For example, in Figure 1 , the personal assistant device 120 outputs an audible output of the text content of the electronic communication as "GOAL! Can you believe it’s already World Cup time?". In at least some examples, the personal assistant device can select one or more portions of the text content to include in and / or exclude from the audible output. For example, the personal assistant device can avoid audibly outputting the text content of the signature block at the end of the message or the network domain address included in the message. In some examples, the text content can be audibly output as an audible reproduction of its text to provide a literal reading of the text content. In other examples, the text content can be intelligently edited by the personal assistant device to provide an enhanced listening experience for the user, including correcting spelling / grammar errors in the text content, reordering the text components of the text content, and / or summarizing the text content in the audible output.

[0050] After presenting the first unviewed electronic communication, the audio presentation can proceed to the second unviewed electronic communication of the first conversation thread. For example, at 426, a communication summary of the second unviewed electronic communication of the first conversation thread can be presented as an audible output. At 428, the message content of the second unviewed electronic communication of the first conversation thread can be presented as an audible output. The audio presentation can proceed sequentially through each unviewed electronic communication of the first conversation thread. In at least some examples, the unviewed electronic communications of the conversation thread can be presented in chronological order based on the respective timestamps of the unviewed electronic communications, starting with the oldest unviewed electronic communication of the conversation thread and continuing to the newest unviewed electronic communication of the conversation thread.

[0051] At 430, a guided notification can be presented as an audible output. The guided notification can be used to ask the user if they want to perform an action for the first conversation thread. For example, the guided notification can provide a general notification to the user such as "perform an action or proceed to the next conversation?" or can provide a targeted notification such as "would you like to reply to this conversation?" At 432, a silent period can be provided to enable the user to provide instructions or otherwise act on the conversation thread before proceeding to the next conversation thread for audio presentation.

[0052] After presenting the first conversation thread at 470, the audio presentation can continue to present a second conversation thread at 472, which includes one or more unviewed electronic communications for the recipient. The presentation of the second conversation thread can similarly include the presentation of the following: a thread summary of the second conversation thread at 440, a communication summary of the first unviewed electronic communication of the second conversation thread at 442, the message content of the first unviewed electronic communication of the second conversation thread at 444, a communication summary of the second unviewed electronic communication of the second conversation thread at 446, the message content of the second unviewed electronic communication of the second conversation thread at 448, and so on, until each unviewed electronic communication of the second conversation thread has been presented as an audible output.

[0053] The audio presentation can continue for each conversation thread that includes one or more unviewed electronic communications for the recipient, as described previously with reference to the presentation of the first conversation thread at 470. After the presentation of a conversation thread that includes one or more unviewed electronic communications, additional information determined to be potentially relevant to the user by the personal assistant device can be presented as an audible output at 460. At 462, the user can end the audio presentation session via the personal assistant device.

[0054] Continue Figure 4 For an example timeline, the user can provide instructions to the personal assistant device to navigate within the audio presentation or between conversation threads and their electronic communications. For example, in response to instruction 480, the personal assistant device can advance the audio presentation from presenting a communication summary at 422 to presenting a thread summary of the second conversation thread at 440, enabling the user to skip some or all of the presentation of the first conversation thread. In Figure 1In this case, user 110 provides the spoken command "Next conversation", as an example of instruction 480. For example, in response to instruction 480, the personal assistant device can advance the audio presentation from presenting a communication summary of a first unviewed electronic communication at 422 to presenting a communication summary of a second unviewed electronic communication at 426, enabling the user to skip the presentation of some or all of the first unviewed electronic communications.

[0055] By organizing electronic communications into conversation threads, the user can perform actions on the electronic communications for that conversation thread. For example, as described above, the user can skip the audio presentation of a specific conversation thread (including the unviewed electronic communications of that conversation thread) by providing a spoken command such as Figure 1 "Next conversation". As another example, the user can delete the electronic communications of a conversation thread or mark such electronic communications as important by providing a spoken command (such as instruction 496) during the silent period 452. Thus, the personal assistant device can apply actions to each of the multiple electronic communications of a conversation thread in response to the user's spoken command.

[0056] In at least some examples, an audible indicator can be presented by the personal assistant device as an audible output to notify the user of transitions between portions of the audio presentation. For example, an audible indicator 482 can be presented between the presentation of the change of day at 418 and the thread summary at 420, audible indicators 484 and 490 can be presented between electronic communications, audible indicators 486 and 492 can be presented between the guided notification and the silent period, and audible indicators 488 and 494 can be presented between the silent period and the subsequent conversation thread and the additional information presented at 460 or the closing statement presented at 462. The audible indicator can take the form of an audible tone or any suitable sound. Audible indicators with distinguishable sounds can be presented at different portions of the audio presentation. For example, the audible indicator 484 that identifies a transition between electronic communications can be different from the audible indicator 488 that identifies a transition between conversation threads. Such audible indicators can help the user easily understand whether the personal assistant device has started or completed a specific portion of the audio presentation, whether the personal assistant device has completed a specific action in accordance with the user's instructions, or whether the personal assistant device is currently listening for an instruction to be provided by the user.

[0057] A personal assistant device can support various presentation modes, including a continuous presentation mode and a guided presentation mode. In the continuous presentation mode, the personal assistant device can continue with audio presentation without instructions from the user. In the guided presentation mode, the personal assistant device can pause the audio presentation at a transition point to wait for instructions from the user to continue. For example, in the guided presentation mode, the personal assistant device can pause the audio presentation and output a query after presenting a conversation summary: "Would you like to hear this conversation thread".

[0058] Figure 5 Depicts an example timeline 500 of electronic communications. Within the timeline 500, time progresses from the left - hand side of the figure to the right - hand side of the figure. Figure 5 The time of each electronic communication within can correspond to the corresponding timestamp of that electronic communication, such as described by the timestamp 314 of reference Figure 3 as described.

[0059] The timeline 500 is divided into multiple conversation threads 510 - 520, each conversation thread including one or more electronic communications of a recipient. In this example, conversation thread 510 includes electronic communications 530 - 540, conversation thread 512 includes electronic communications 550 - 558, conversation thread 514 includes electronic communications 560 - 564, conversation thread 516 includes electronic communication 570, conversation thread 518 includes electronic communication 580, and conversation thread 520 includes electronic communications 590 - 594.

[0060] The multiple electronic communications of a conversation thread can be referred to as reply - linked electronic communications, where one or more of the electronic communications are a reply to an original electronic communication, thus linking these electronic communications to each other through a common conversation thread. The first electronic communication is a reply to an earlier second electronic communication, which in turn is a reply to an even earlier third electronic communication. The first electronic communication can be considered to reply - link to both the second and third electronic communications, thus forming a common conversation thread. For example, electronic communication 534 is a reply to electronic communication 532, and electronic communication 532 is a reply to electronic communication 530. Thus, each of the electronic communications 530, 532, and 534 forms part of conversation thread 510. For certain types of electronic communications, such as a collaborative messaging platform or a multi - player game platform, electronic communications associated with a particular channel (e.g., a particular collaborative project or multi - player game) can be identified as being reply - linked to each other.

[0061] In addition, in this example, electronic communications 530-540, 554-558, 560-564, 570, and 594 are unviewed electronic communications of the recipient. In contrast, electronic communications 550, 552, 580, and 590 are previously viewed electronic communications of the recipient. In one example, an electronic communication can be referred to as an unviewed electronic communication if the message of the electronic communication (e.g., Figure 3 message 320) has not been presented to the recipient user in any of a visual, auditory, or other (e.g., Braille) presentation mode. For example, in the context of email, individual email messages can be marked as "read" or "unread," which can correspond to previously viewed or unviewed electronic communications. In Figure 5 the example, electronic communication 592 corresponds to the recipient's reply to a previous electronic communication 590.

[0062] As described in the example conversation between user 110 and personal assistant device 120 of reference Figure 1 , multiple conversation threads can be presented according to a specific presentation order. In at least some examples, the presentation order for presenting two or more conversation threads can be based on the timing of unviewed electronic communications of each conversation thread. In Figure 5 the example, each of the electronic communications 530-540 of conversation thread 510 is received after each of the electronic communications 550-558 of conversation thread 512, while the electronic communications 560-564 of conversation thread 514 are interspersed in time between the electronic communications of conversation threads 510 and 512.

[0063] In a first example presentation order, conversation threads can be presented according to a reverse chronological order based on the latest unviewed electronic communication of each conversation thread. In Figure 5 the example timeline, conversation thread 510 can be presented before conversation threads 512, 514, 516, and 520 because conversation thread 510 includes the latest unviewed electronic communication 540, which has a timing after the latest unviewed electronic communications 558, 564, 570, and 594 of conversation threads 512, 514, 516, and 520, respectively. This first example presentation order can be used to prioritize conversation threads with the latest activity according to unviewed electronic communications received by the recipient. In this example, conversation thread 518 may not be presented because conversation thread 518 does not include any unviewed electronic communications.

[0064] Figure 6 depicts the situation where, in the absence of user instructions to advance or interrupt the presentation of conversation threads, as described above for Figure 5A first example of the electronic communication description presents an example timeline 600 of the order. Within the timeline 600, time advances from the left - hand side of the figure to the right - hand side of the figure. The conversation threads 510 - 516 and 520 are presented in Figure 6 in reverse chronological order based on the most recent unviewed electronic communication of each conversation thread. In each conversation thread, the unviewed electronic communications can be presented in chronological order, starting from the earliest unviewed electronic communication of the conversation thread and continuing to the most recent unviewed electronic communication of that conversation thread, and the same is true in the absence of a user instruction to advance or interrupt the presentation of the conversation thread. For example, according to Figure 6 the first example presentation order depicted in Figure 5 the unviewed electronic communications received are: 560, 554, 594, 556, 558, 562, 530, 532, 570, 534, 564, 536, 538, 540 and are presented in the following order: the electronic communications 530 - 540 of conversation thread 510, the electronic communications 560 - 564 of conversation thread 514, the electronic communication 516 of conversation thread 570, the electronic communications 554 - 558 of conversation thread 512, and the electronic conversation 594 of conversation thread 520.

[0065] Returning to Figure 5 , in a second example presentation order, the conversation threads can be presented in chronological order based on the most recent unviewed electronic communication of each conversation thread. Compared with the above reverse chronological order, this will result in the opposite sorting of the conversation threads. For example, in the example timeline of Figure 5 , conversation thread 512 can be presented before conversation threads 510 and 514 because conversation thread 512 includes the most recent unviewed electronic communication 558, which has a timing before the most recent unviewed electronic communications 540 and 564 of conversation threads 510 and 514 respectively.

[0066] In a third example presentation order, the conversation threads can be presented in reverse chronological order based on the timing of the earliest unviewed electronic communication of each conversation thread. In the example timeline of Figure 5 , conversation thread 510 can be presented before conversation threads 512 and 514 because conversation thread 510 includes the earliest unviewed electronic communication 530, which has a timing after the earliest unviewed electronic communications 554 and 560 of conversation threads 512 and 514 respectively.

[0067] In a fourth example presentation order, the conversation threads can be presented in chronological order based on the timing of the earliest unviewed electronic communication of each conversation thread. In Figure 5In the example timeline, conversation thread 514 can be presented before conversation threads 510 and 512 because conversation thread 514 includes the earliest unviewed electronic communication 560, which has a timing prior to the earliest unviewed electronic communications 530 and 554 of conversation threads 510 and 512, respectively.

[0068] In a fifth example presentation order, a conversation thread that includes a recipient's reply at some point within the thread can be prioritized in the presentation order over a conversation thread that does not include the recipient's reply. In Figure 5 the example timeline, the unviewed electronic communication 594 of conversation thread 520 can be presented before the electronic communications of conversation threads 510 - 516 because conversation thread 520 includes the recipient's reply electronic communication 592. The presence of the reply electronic communication 592 in conversation thread 520 can indicate an increased importance of conversation thread 520 compared to other conversation threads. Among multiple conversation threads each including a recipient's reply, the presentation order of the unviewed electronic communications can utilize any one of the first, second, third, or fourth example presentation orders discussed above for presenting conversation threads that include a recipient's reply before presenting the unviewed electronic communications of conversation threads that do not include the recipient's reply.

[0069] In a sixth example presentation order, the prioritization of a conversation thread having a recipient's reply, such as described above for the fifth example presentation order, can consider only such replies of the recipient: the unviewed electronic communication is a reply directly to that reply of the recipient. This presentation order can be used to prioritize conversation threads that include unviewed electronic communications that are direct replies linked to the recipient's reply over other conversation threads.

[0070] In a seventh example presentation order, conversation threads can be prioritized based on one or more factors, including: the subject of the electronic communication, the content of the message or attachment, the sender of the electronic communication, the number of electronic communications in each conversation thread, the frequency of electronic communications in each conversation thread, the presence of an importance indicator (e.g., flag) associated with the electronic communication, and so on. In one example, conversation threads can be ranked according to one or more factors and presented in an order based on the ranking of the conversation threads. Such ranking can be based on any desired heuristic, machine learning algorithm, or other ranking method.

[0071] Figure 7A Depicts a flowchart of an example method 700 for organizing and presenting conversation threads. Method 700 or portions thereof can be executed by one or more computing devices of a computing system. For example, method 700 can be executed by Figure 2 the computing device 210 or by a computing system including the computing device 210 in combination with Figure 2 the server system 260.

[0072] At 710, an electronic communication is obtained for a recipient. In an example, the electronic communication can be obtained from a remote server system via a communication network at a user's computing device. The electronic communication obtained for the recipient at 710 can span one or more types of electronic communications and can be collected from one or more communication services and / or applications. Additionally, the electronic communication obtained at 710 can refer to a subset of all of the recipient's electronic communications. For example, the electronic communication obtained at 710 can include the recipient's primary or focal inbox or folder and can exclude other inboxes or folders, such as spam, promotions, etc.

[0073] At 712, unviewed electronic communications are identified among the electronic communications obtained for the recipient at 710. As previously referenced Figure 5 as described, if a message of an electronic communication (e.g., Figure 3 message 320) has not been presented to the recipient user via any of a visual, auditory, or tactile (e.g., braille) presentation mode, the electronic communication can be referred to as an unviewed electronic communication. In one example, an identifier indicating whether the electronic communication is viewed or unviewed can be stored as metadata of the electronic communication. In another example, the identifier can be stored at the communication application or service from which the electronic communication is obtained and can be reported by the application or service along with the electronic communication.

[0074] At 714, the electronic communications obtained at 710 are organized according to a pattern. The pattern can be programmatically defined by one or more of a communication application of the user's computing device, a communication service of the server system, or a personal assistant machine, depending on the implementation. For example, some communication services or applications can organize or partially organize electronic communications into conversation threads, while other communication services or applications may not support the use of conversation threads.

[0075] At 716, the electronic communications obtained at 710 are grouped into multiple conversation threads of electronic communications that include two or more reply links. As previously described, two or more electronic communications are reply-linked if an electronic communication is a reply to an earlier electronic communication, and the electronic communication can be reply-linked to the earlier electronic communication via one or more intermediate reply-linked electronic communications. After operation 716, each conversation thread includes two or more electronic communications for the recipient that are reply-linked to each other. However, it will be understood that at least some conversation threads can include individual electronic communications. At 718, data representing the grouping of the electronic communications can be stored for each conversation thread. For example, data representing the grouping from operation 716 can be stored in a storage subsystem of the computing device, including at the user's local computing device and / or at a remote server system.

[0076] At 720, electronic communications of each conversation thread can be sorted chronologically according to timestamps indicating the timing of each electronic communication. At 722, data representing the sorting of the electronic communications can be stored for each conversation thread. For example, the data representing the sorting from operation 722 can be stored in a storage subsystem of a computing device, including at a computing device local to the user and / or at a remote server system.

[0077] At 724, the conversation threads can be sorted based on rules to obtain a presentation order among the conversation threads. As described in the previous example of the presentation order Figure 5 Multiple different presentation orders can be supported among the conversation threads. According to the first example presentation order described in further detail with reference Figure 6 The rules applied at operation 724 can include identifying the most recent unviewed electronic communication of each conversation thread and sorting the conversation threads in reverse chronological order based on the timing of the most recent unviewed electronic communication of the conversation thread. The rules applied at operation 724 can be defined to provide any of the example presentation orders described herein. At 726, data representing the sorting of the conversation threads can be stored. For example, the data representing the sorting from operation 724 can be stored in a storage subsystem of a computing device, including at a computing device local to the user and / or at a remote server system.

[0078] At 728, an instruction for initiating an audio presentation of an electronic communication to a recipient is received. The instruction can be in the form of a user dictated command, such as the previous reference Figure 1As described, where the user speech 130 includes "Read messages". In at least some examples, the dictated commands for initiating the audio presentation can be one or more keywords predefined at the personal assistant device and recognizable by the personal assistant device, such as "Messages", "Play messages", "Read messages", "Hear messages", "Get mail", "tell me about my emails", "What emails do I have?", "Did anyone email me?", "Do I have any new emails?", and so on. In at least some examples, the intention of the user to initiate an audio presentation through a specific dictated utterance can be inferred from the context and / or can be learned from previous interactions with the user. For example, the personal assistant device can ask the user whether the user wants to initiate an audio presentation of unviewed electronic communications, and the user can respond by saying "yes" or "please". The instruction received at 728 can also include non-verbal commands, such as user input provided via any input device or the interface of the user's computing device. Additionally, in some examples, the audio presentation of unviewed electronic communications can be initiated by the personal assistant device in certain contexts without receiving an instruction. For example, the personal assistant device can initiate an audio presentation in response to specific operating conditions, such as a scheduled time, the user picking up the personal assistant device, receiving a new unviewed electronic communication, and so on.

[0079] At 730, in response to the instruction received at 728, an audio presentation of the conversation thread is output according to the presentation order obtained at operation 724. The presentation order can be defined by one or more of the following: the grouping of electronic communications at 716, the sorting of electronic communications at 720, and the sorting of the conversation thread at 724, and can be based on the data stored at 718, 722, and 726.

[0080] In one example, the audio presentation includes unviewed electronic communications of each conversation thread in chronological order, starting from the oldest unviewed electronic communication of the conversation thread and continuing to the newest unviewed electronic communication of the conversation thread. The conversation thread is before another conversation thread among the multiple conversation threads, and the other conversation thread includes unviewed electronic communications that are interspersed in time between the oldest unviewed electronic communication and the newest unviewed electronic communication of that conversation thread. For example, at 732, two or more unviewed electronic communications of the first conversation thread are audibly output in chronological order before the unviewed electronic communications of the second conversation thread at 734.

[0081] In addition, in one example, the presentation order of the conversation threads can be in reverse chronological order based on the newest unviewed electronic communication of each conversation thread among the multiple conversation threads, such that the first conversation thread with the first newest unviewed electronic communication is presented before the second conversation thread at 732, and the second conversation thread has a second newest unviewed electronic communication that is older than the first newest unviewed electronic communication of the multiple conversation threads. Reference Figure 6 describes an example of such reverse chronological order.

[0082] The audio presentation output at 730 can include, for each unviewed electronic communication, at least a portion of the text content of the message of the unviewed electronic communication presented as an audible output. In one example, all of the text content of the message of the unviewed electronic communication can be presented as an audible output. In addition, in at least some examples, the audio presentation further includes: for each conversation thread among the multiple conversation threads, a thread summary of the conversation thread presented as an audible output before the text content of the conversation thread. Reference Figure 4 describes an example of the thread summary presented before the message content.

[0083] At 740, a second instruction for advancing the audio presentation can be received. The instruction received at 740 can be in the form of a user's spoken command, such as previously referenced Figure 1 described, where the user voice 130 includes "Next conversation". However, the instruction received at 740 can include a non-verbal command, such as user input provided via any input device or the interface of the user's computing device.

[0084] At 742, in response to a second instruction, the audio presentation of multiple conversation threads can be advanced from the current conversation thread to a subsequent conversation thread in the presentation order. It should be understood that the personal assistant device can support other forms of navigation within the audio presentation, including ending the audio presentation, restarting the audio presentation, skipping to the next conversation thread, skipping to a specific conversation thread identified by the user, skipping the next unviewed electronic communication, skipping to a specific unviewed electronic communication identified by the user, and so on.

[0085] The action of advancing the audio presentation for a conversation thread is one of the multiple actions that the personal assistant device can support. For example, operation 740 can alternatively include instructions for performing different actions, such as replying, forwarding to another recipient, storing or deleting the conversation thread, or marking the conversation thread as important (e.g., marking the conversation thread or its electronic communication). For at least some types of actions, in response to an instruction to perform the action, at 742, the personal assistant device can apply the action to each electronic communication of the conversation thread. The spoken commands used by the personal assistant device to initiate a specific action can include: one or more keywords predefined at the personal assistant device and recognizable by the personal assistant device, or the intent of the spoken utterance can be inferred by the personal assistant device from the context, such as previously referring to the instructions received at 728.

[0086] Figure 7B A flowchart depicting an example method 750 for presenting conversation threads is shown. Method 750 can be executed in conjunction with Figure 7A method 700. For example, method 750 or a portion thereof can form part of operation 730 of method 700. Method 750 or a portion thereof can be executed by one or more computing devices of a computing system. For example, method 700 can be executed by Figure 2 computing device 210 or by a computing system including computing device 210 in conjunction with Figure 2 server system 260.

[0087] At 752, an instruction can be received. For example, the instruction received at 752 can correspond to the instruction received at Figure 7A 728. In response to the instruction, the method at 752 includes: audibly outputting each unviewed electronic communication in the latest conversation thread, which includes a set of the latest unviewed, reply-linked electronic communications for the recipient. For example, at 754, the personal assistant device audibly outputs the next latest conversation thread. As part of audibly outputting the next latest conversation thread at 754, at 756, the personal assistant device can audibly output a thread summary. However, in other examples, the thread summary may not be audibly output.

[0088] At 758, each unviewed electronic communication in the latest conversation thread can be aurally output in chronological order, starting at 760 with the oldest unviewed electronic communication. Aurally outputting the oldest unviewed electronic communication at 760 can include: aurally outputting a communication summary at 762, and aurally outputting some or all of the text content of the message at 764. However, in other examples the communication summary may not be aurally output.

[0089] At 766, if there are still unviewed electronic communications in the conversation thread, the method returns to 760, where the oldest unviewed electronic communication is aurally output. Thus, the method continues to the latest unviewed electronic communication, as described, for example, in the order of presentation of the examples previously referenced Figure 6 above.

[0090] At 766, if there are no longer unviewed electronic communications in the conversation thread, the method proceeds to 768. At 768, if there are more conversation threads that include unviewed electronic communications, the method can return to 754, where the next latest conversation thread is aurally output at 754. Thus, in response to completing the aural output of the latest unviewed electronic communications from a conversation thread, the method includes: aurally outputting each unviewed electronic communication in the next latest conversation thread, including a next latest set of unviewed, reply-link electronic communications for the recipient. Each unviewed electronic communication in the next latest conversation thread is aurally output in chronological order at 758, starting with the oldest unviewed electronic communication and continuing to the latest unviewed electronic communication.

[0091] For example, as described with reference to Figures 4 - 6 at least one unviewed electronic communication from the next latest communication thread can be in chronological order between two unviewed electronic communications from the latest communication thread, and all unviewed electronic communications from the latest conversation thread can be aurally output before any unviewed electronic communication from the next latest communication thread is aurally output by using method 750.

[0092] Figure 8 FIG. depicts a flowchart of an example method 800 for presenting time identification information of a conversation thread. Method 800 or portions thereof can be executed by one or more computing devices of a computing system. For example, method 800 can be executed by Figure 2 computing device 210 or by a computing system including computing device 210 in combination with Figure 2 server system 260.

[0093] At 810, the method includes: receiving an instruction to initiate an aural presentation of an electronic communication for a recipient. As described previously with reference to operation 728 of FIG. 7, the instruction can include an oral command of a user.

[0094] At 812, an electronic communication for a recipient is obtained. As described previously with reference to operation 710 of FIG. 7, the electronic communication for the recipient can be obtained from a remote server system via a communication network at the user's computing device.

[0095] At 814, unviewed electronic communications for the recipient are identified. As previously referenced Figure 5 described, if the message of the electronic communication (e.g., Figure 3 message 320) has not been presented to the recipient user in any of the visual, auditory, or other (e.g., Braille) presentation modes, the electronic communication can be referred to as an unviewed electronic communication. In one example, an identifier indicating whether the electronic communication is viewed or unviewed can be stored as metadata of the electronic communication. In another example, the identifier can be stored at the communication application or service from which the electronic communication is obtained and can be reported by the application or service together with the electronic communication.

[0096] At 816, an estimated time for presenting a portion of an audio presentation is determined, where the portion includes an audible output of the text content of the unviewed electronic communications for the recipient. The text content can include the text content of the message of each unviewed electronic communication. As an example, the estimated time is determined based on characteristics of the text content of multiple unviewed electronic communications. For example, the characteristics of the text content can include the word count or character count of the text content; and the time estimate can be calculated algorithmically based on the word count or character count (e.g., 0.7 seconds per word). As another example, the method can further include: converting the text content of multiple unviewed electronic communications into audio data representing the audible output of the text content, and determining an estimated time for presenting a subsequent portion of the audio presentation based on characteristics of the audio data. For example, the characteristics of the audio data can include the amount of audio data (e.g., byte count), or the duration of the audio data at a target presentation rate.

[0097] The estimated time can be determined based on other information included in the audio presentation that will be audibly output by the personal assistant device in the subsequent portion. For example, in the case where the audio presentation includes a thread summary for each conversation thread, the estimated time can also be determined based on the duration of the thread summary within the subsequent portion of the audio presentation.

[0098] In at least some examples, the estimated time identified by the presentation roadmap can be in the form of a generalized time estimate. Figure 9A An example of a generalized time estimate is depicted. In the case of a generalized time estimate, operation 816 can further include: determining an initial value of the estimated time and selecting a generalized time estimate from multiple hierarchical generalized time estimates based on the initial value of the estimated time. Figure 9AThe example of the generalized time estimate depicted refers to the session duration representing the initial value of the estimated time. In at least some examples, the estimated time can be rounded to a generalized time estimate, e.g., as Figure 9A depicted.

[0099] At 818, in response to an instruction, an audio presentation is output. The output audio presentation includes: an initial portion of the output audio presentation that includes a presentation roadmap 820, and a subsequent portion that includes an audible output of the text content of multiple unviewed electronic communications of the recipient. In one example, the presentation roadmap output at 820 identifies an estimated time for presenting a subsequent portion of the audio presentation output at operation 822, which corresponds to the portion for which an estimated time was determined at operation 816.

[0100] The presentation roadmap output at 818 can identify other features of the audio presentation, such as those previously referenced Figure 4 described. As an example, the presentation roadmap can also identify the number of unviewed electronic communications and / or the number of conversation threads for the unviewed electronic communications.

[0101] Aspects of method 800 can be similarly performed to present an estimated time in a thread summary of a conversation thread of an electronic communication that includes one or more reply links or in a communication summary for an individual electronic communication, e.g., as referenced Figure 4 described.

[0102] Figures 9A - 9E Some tables are depicted where example audible outputs of a personal assistant device are provided for a series of conditions. Figures 9A - 9E The audible output depicted in can be used as part of a conversation with a user, including, for example, as part of a presentation roadmap, thread summary, and communication summary.

[0103] Figure 9A Various example natural language responses of a personal assistant device based on an estimated time or duration of an audio presentation or a portion thereof are depicted.

[0104] Figure 9B Various example natural language responses of a personal assistant device based on the recipient of an electronic communication or conversation thread are depicted.

[0105] Figure 9C Various example natural language responses of a personal assistant device based on the number of unviewed electronic communications in a conversation thread are depicted.

[0106] Figure 9D Various example natural language responses of a personal assistant device based on changes in the recipient of an electronic communication within a conversation thread are depicted.

[0107] Figure 9EDepicts various example natural language responses for estimating the audio presentation duration of message-based text content by a personal assistant device.

[0108] Figures 10A - 10P Depicts an example conversation between a user and a personal assistant device as described above. Figures 10A - 10P The portion of the example conversation corresponding to the personal assistant device represented by "Assistant" can take the form of the audible output of the personal assistant device, and the portion of the conversation corresponding to the user represented by "User" can take the form of the user's spoken words.

[0109] In at least some examples, the personal assistant device can utilize one or more conversation templates configured to implement the logic of method 700. For example, Figure 4 The timeline can represent a conversation instantiated from a conversation template that starts with a greeting 410, progresses to presenting a roadmap 414, changes to a date 418, and then cycles through each unviewed conversation thread according to method 750 before ending with a guided notification 450, additional information 460, and a closing statement 462. It should be understood that different templates presenting information in different orders can be used. Such templates can be configured to branch to different conversation sequences in response to user instructions.

[0110] Figures 10A - 10C Depicts an example conversation. In Figure 10A the personal assistant device audibly outputs a presented roadmap such as that previously referenced Figure 1 followed by the audible output of additional conversation threads. In Figure 10B and Figure 10C the user provides instructions to perform additional actions on the conversation thread, including marking an electronic communication as important. For example, in Figure 10B when the personal assistant device is audibly outputting the text content of a message from sender "Satya", the user uses an interruptive spoken command in the form of "flag that". Additionally, in Figure 10B after the conversation thread for the topic "Pizza party" is audibly output by the personal assistant device, the user provides the spoken command "flag that" during a silent period provided by the personal assistant device (e.g., Figure 4 the silent period 432 of Figure 10C the personal assistant device audibly outputs "You’ve got a package from Company XYZ on its way" as Figure 4An example of additional information 460, and “That’s all for now” as an audible indication of the closing statement 462 to end the audio presentation of an electronic communication. Figure 4 to end the audio presentation of an electronic communication.

[0111] Figure 10D and Figure 10E depicts an example dialog for an inbox query. In Figure 10D , the personal assistant device uses a guided presentation mode, where the personal assistant device asks the user “Which sender do you wanna hear more about?” after the presentation roadmap is audibly output, and the roadmap identifies specific senders “Jade”, “Ruby”, and “Trent” as well as other roadmap information. Such an inquiry by the personal assistant device can take the form of Figure 4 an interruption notification 416. In response to the user saying “Jade”, the personal assistant device presents a thread summary for the unviewed electronic communication from Jade, which again identifies the sender “Jade”, the subject “Touching letter..”, and the time / length estimate “it’s a long one”. After the thread summary, the personal assistant device uses the guided presentation mode to ask the user “Wanna hear it?”, and in response to the user providing the verbal command “yes”, the personal assistant device audibly outputs at least a portion of the text content of the message.

[0112] In Figure 10E , the personal assistant device highlights three unviewed electronic communications that the user might want to hear from a total of 10 unviewed electronic communications.

[0113] Figure 10F depicts an example conversation for a person-based query.

[0114] Figure 10G depicts an example conversation where the personal assistant device highlights a specific sender of an electronic communication in the presentation roadmap.

[0115] Figure 10H depicts an example conversation for an inbox query where the personal assistant device determines that the unviewed electronic communication is not important.

[0116] Figure 10I depicts an example dialog for an inbox query where there are no unviewed electronic communications for the recipient.

[0117] Figure 10JAn example dialogue is depicted in which a personal assistant device prepares and sends electronic communications on behalf of a user in response to spoken commands.

[0118] Figure 10K An example conversation is depicted in which a personal assistant device replies to electronic communications on behalf of a user in response to spoken commands.

[0119] Figure 10L An example conversation is depicted in which a personal assistant device responds to an electronic communication with multiple recipients on behalf of a user in response to spoken commands.

[0120] Figure 10M An example conversation is depicted in which a personal assistant device forwards an electronic communication to another recipient identified by a user via a spoken command.

[0121] Figure 10N Depicted is an example conversation in which a personal assistant device saves a draft of a response on behalf of a user.

[0122] Figure 10O An example conversation is depicted in which a user selects a particular electronic communication to be audibly output by a personal assistant device.

[0123] Figure 10P An example dialogue is depicted in which a personal assistant device audibly outputs calendar data of an electronic communication and performs actions on the calendar data in response to a user's spoken commands. For example, the personal assistant device outputs "Would you like to accept this meeting?", to which the user responds "Yes", in response to which the personal assistant device sends a meeting confirmation reply to the sender of the meeting request (i.e., "Nicki").

[0124] In at least some examples, the methods and processes described herein may be bound to a computing system of one or more computing devices. Specifically, these methods and processes may be implemented as computer applications or services, application programming interfaces (APIs), libraries, and / or other computer program products.

[0125] Reference again Figure 2 , computing system 200 is an example computing system that can perform one or more of the methods and operations described herein. Computing system 200 is shown in simplified form. Computing system 200 can take the form of one or more mobile computing devices, wearable computing devices, computing devices integrated with vehicles, desktop computing devices, home appliance computing devices, personal computers, server computers, tablet computers, home entertainment computers, network computing devices, gaming devices, mobile computing devices, mobile communication devices (e.g., smartphones), Internet of Things (IoT) devices, embedded computing devices, and / or other computing devices.

[0126] The logic subsystem 212 may include one or more processors configured to execute software instructions. Additionally or alternatively, the logic subsystem may include one or more hardware or firmware logic circuits configured to execute hardware or firmware instructions. The processors of the logic subsystem may be single-core or multi-core, and the instructions executed thereon may be configured for sequential, parallel, and / or distributed processing. Optionally, the various components of the logic subsystem may be distributed among two or more separate devices, which may be remotely located and / or configured for cooperative processing. Some aspects of the logic subsystem may be virtualized and executed by remotely accessible networked computing devices configured in a cloud computing configuration.

[0127] The storage subsystem 214 may include removable and / or built-in devices. The storage subsystem 214 may include optical memory (e.g., CD, DVD, HD-DVD, Blu-ray Disc, etc.), semiconductor memory (e.g., RAM, EPROM, EEPROM, etc.), and / or magnetic memory (e.g., hard disk drive, floppy disk drive, tape drive, MRAM, etc.), among others. The storage subsystem 214 may include volatile, non-volatile, dynamic, static, read / write, read-only, random access, sequential access, location-addressable, file-addressable, and / or content-addressable devices. It should be understood that the storage subsystem 214 includes one or more physical devices and is not merely electromagnetic signals, optical signals, etc. that are not physically stored by a physical device for a limited duration.

[0128] Aspects of the logic subsystem 212 and the storage subsystem 214 may be integrated together into one or more hardware logic components. For example, such hardware logic components may include field-programmable gate arrays (FPGAs), program-specific and application-specific integrated circuits (PASIC / ASICs), program-specific and application-specific standard products (PSSP / ASSPs), systems-on-a-chip (SoCs), and complex programmable logic devices (CPLDs).

[0129] When the methods and operations described herein are implemented by the logic subsystem 212 and the storage subsystem 214, the state of the storage subsystem 214 may be transformed—e.g., to store different data. For example, the logic subsystem 212 may be configured to execute instructions 222 that are part of one or more applications, services, programs, routines, libraries, objects, components, data structures, or other logical constructs. Such instructions may be implemented to perform tasks, implement data types, transform the state of one or more components, achieve a technical effect, or otherwise achieve a desired result.

[0130] The logical subsystem and the storage subsystem can cooperate to instantiate one or more logical machines, such as those previously described with reference to the personal assistant machine 230, the voice output machine 232, and the voice input machine 234. It will be understood that the "machines" described herein (e.g., with reference to Figure 2 ) are never abstract ideas and always have a tangible form. Instructions 222 that provide functionality in combination with hardware to a particular machine may optionally be saved as unexecuted modules on a suitable storage device and may be sent via network communication and / or transfer of the physical storage device that holds such modules.

[0131] Machines can be implemented using any suitable combination of state-of-the-art and / or future machine learning (ML), artificial intelligence (AI), and / or natural language processing (NLP) techniques. Non-limiting examples of techniques that can be incorporated into the implementation of one or more machines include support vector machines, multi-layer neural networks, convolutional neural networks (e.g., including spatial convolutional networks for processing images and / or videos, temporal convolutional neural networks for processing audio signals and / or natural language sentences, and / or any other suitable convolutional neural network configured to convolve and pool features across one or more temporal and / or spatial dimensions), recurrent neural networks (e.g., long short-term memory networks), associative memories (e.g., lookup tables, hash tables, Bloom filters, neural Turing machines, and / or neural random access memories), word embedding models (e.g., GloVe or Word2Vec), unsupervised spatial and / or clustering methods (e.g., nearest neighbor algorithms, topological data analysis, and / or k-means clustering), graphical models (e.g., (hidden) Markov models, Markov random fields, (hidden) conditional random fields, and / or artificial intelligence knowledge bases), and / or natural language processing techniques (e.g., tokenization, stemming, constituency and / or dependency parsing, and / or intent recognition, segmentation models, and / or hyper-segmentation models (e.g., hidden dynamic models)).

[0132] In some examples, one or more differentiable functions can be used to implement the methods and processes described herein, where the gradient of the differentiable function can be calculated and / or estimated for the input and / or output of the differentiable function (e.g., for training data and / or for an objective function). Such methods and processes can be determined at least in part by a set of trainable parameters. Thus, the trainable parameters used for a particular method or process can be adjusted by any suitable training process to continuously improve the functionality of the method or process.

[0133] Non-limiting examples of training processes for adjusting trainable parameters include supervised training (e.g., using gradient descent or any other suitable optimization method), zero-shot, few-shot, unsupervised learning methods (e.g., classification based on categories derived from unsupervised clustering methods), reinforcement learning (e.g., deep Q-learning based on feedback), and / or generative adversarial neural network training methods, belief propagation, RANSAC (random sample consensus), contextual bandit methods, maximum likelihood methods, and / or expectation maximization. In some examples, components of multiple methods, processes, and / or systems described herein can be trained simultaneously for an objective function that measures the performance of the collective functionality of multiple components (e.g., for enhanced feedback and / or for labeled training data). Training multiple methods, processes, and / or components simultaneously can improve this collective functionality. In some examples, one or more methods, processes, and / or components can be trained independently of other components (e.g., offline training for historical data).

[0134] A language model can utilize lexical features to guide sampling / searching for words to identify speech. For example, a language model can be defined at least in part by the statistical distribution of words or other lexical features. For example, a language model can be defined by the statistical distribution of n-grams, which define the transition probabilities between candidate words based on lexical statistics. A language model can also be based on any other suitable statistical features and / or the results of processing statistical features using one or more machine learning and / or statistical algorithms (e.g., confidence values produced by such processing). In some examples, a statistical model can restrict the words that can be identified for an audio signal, e.g., based on the assumption that the words in the audio signal come from a particular vocabulary.

[0135] Alternatively or additionally, a language model can be based on one or more neural networks that have been previously trained to represent audio inputs and words in a shared latent space, e.g., a vector space learned by one or more audio and / or word models (e.g., wav2letter and / or word2vec). Thus, finding candidate words can include searching the shared latent space based on the vector encoded by the audio model for the audio input in order to find candidate word vectors to be decoded by the word model. The shared latent space can be used to evaluate the confidence that a candidate word plays an important role in the speech audio for one or more candidate words.

[0136] The language model can be used in combination with an acoustic model that is configured to evaluate the confidence that a candidate word is included in the speech audio of an audio signal based on acoustic features of the word (e.g., Mel-frequency cepstral coefficients, formants, etc.) for the candidate word and the audio signal. Optionally, in some examples, the language model can be incorporated into the acoustic model (e.g., the evaluation and / or training of the language model can be based on the acoustic model). The acoustic model, for example based on labeled speech audio, defines a mapping between acoustic signals and basic sound units such as phonemes. The acoustic model can be based on state-of-the-art techniques or any suitable combination of future machine learning (ML) and / or artificial intelligence (AI) models, such as: deep neural networks (e.g., long short-term memory, temporal convolutional neural network, restricted Boltzmann machine, deep belief network), hidden Markov model (HMM), conditional random field (CRF), and / or Markov random field, Gaussian mixture model, and / or other graphical models (e.g., deep Bayesian network). The audio signal to be processed by the acoustic model can be preprocessed in any suitable way, e.g., encoded at any suitable sampling rate, Fourier transform, band-pass filter, etc. The acoustic model can be trained to identify the mapping between acoustic signals and sound units based on training on labeled audio data. For example, the acoustic model can be trained based on labeled audio data including speech audio and corrected text in order to learn the mapping between the speech audio signal and the sound units represented by the corrected text. Thus, the acoustic model can be continuously improved to improve its utility in correctly identifying speech audio.

[0137] In some examples, in addition to statistical models, neural networks, and / or acoustic models, the language model can also incorporate any suitable graphical model, such as, for example, a hidden Markov model (HMM) or a conditional random field (CRF). Given the speech audio and / or other words recognized so far, the graphical model can utilize statistical features (e.g., transition probabilities) and / or confidence values to determine the likelihood of the recognized words. Thus, the graphical model can utilize statistical features, previously trained machine learning models, and / or acoustic models to define the transition probabilities between the states represented in the graphical model.

[0138] In at least some examples, the I / O subsystem 216 can include or interface with a selected portion of natural user input (NUI) elements. Such element portions can be integrated or peripheral, and the conversion and / or processing of input actions can be on-board or off-board processing. Example NUI element portions can include: microphones for speech and / or sound recognition; infrared, color, stereo, and / or depth cameras for machine vision and / or pose recognition; head trackers, eye trackers, accelerometers, and / or gyroscopes for motion detection and / or intent recognition; and electric field sensing element portions for evaluating brain activity.

[0139] It will be understood that the "service" as used herein is an application program executable across multiple user sessions. The service can be used by one or more system components, programs, and / or other services. In some implementations, the service can run on one or more server computing devices.

[0140] According to an example of the present disclosure, a method performed by a computing system includes: receiving an instruction to initiate an audio presentation of an electronic communication to a recipient; in response to the instruction, audibly outputting each unviewed electronic communication in the most recent conversation thread, which includes a set of the most recent unviewed, reply-link electronic communications to the recipient, wherein each unviewed electronic communication in the most recent conversation thread is audibly output in chronological order, starting from the oldest unviewed electronic communication and continuing to the most recent unviewed electronic communication; and in response to completing the audible output of the most recent unviewed electronic communications from the conversation thread, audibly outputting each unviewed electronic communication in the next most recent conversation thread, which includes a set of the next most recent unviewed, reply-link electronic communications to the recipient, wherein each unviewed electronic communication in the next most recent conversation thread is audibly output in chronological order, starting from the oldest unviewed electronic communication and continuing to the most recent unviewed electronic communication. In this example or any other example disclosed herein, at least one unviewed electronic communication from the next most recent communication thread is chronologically between two unviewed electronic communications from the most recent conversation thread, and all unviewed electronic communications from the most recent conversation thread are audibly output before any unviewed electronic communication from the next most recent communication thread is audibly output.

[0141] According to another example of the present disclosure, a method performed by a computing system includes: receiving an instruction to initiate an audio presentation of an electronic communication for a recipient; and in response to the instruction, outputting an audio presentation of a plurality of conversation threads according to a presentation order, wherein each conversation thread includes two or more unviewed electronic communications for the recipient that are linked to each other in reply, the audio presentation includes two or more unviewed electronic communications in chronological order in each conversation thread, starting from the oldest unviewed electronic communication and continuing to the latest unviewed electronic communication of the conversation thread, and the conversation thread is before another conversation thread among the plurality of conversation threads, and the other conversation thread includes unviewed electronic communications that are interspersed in time between the oldest unviewed electronic communication and the latest unviewed electronic communication of the conversation thread. In this example or any other example disclosed herein, the presentation order is based on the reverse chronological order of the latest unviewed electronic communication of each conversation thread among the plurality of conversation threads, such that: a first conversation thread having a first latest unviewed electronic communication is presented before a second conversation thread having a second latest unviewed electronic communication, and the second latest unviewed electronic communication is older than the first latest unviewed electronic communication of the plurality of conversation threads. In this example or any other example disclosed herein, the method further includes: receiving a second instruction to advance the audio presentation of the plurality of conversation threads; and in response to the second instruction, advancing the audio presentation of the plurality of conversation threads from the current conversation thread to a subsequent conversation thread in the presentation order. In this example or any other example disclosed herein, the method further includes: receiving a second instruction to perform an action related to a conversation thread among the plurality of conversation threads; and in response to the second instruction, applying the action to each electronic conversation of the conversation thread. In this example or any other example disclosed herein, the audio presentation includes: for each unviewed electronic communication, at least a part of the text content of the message of the unviewed electronic communication presented as an audible output. In this example or any other example disclosed herein, the audio presentation further includes: for each conversation thread among the plurality of conversation threads, a thread summary of the conversation thread presented as an audible output before the text content of the conversation thread. In this example or any other example disclosed herein, the thread summary identifies the number of unviewed electronic communications of the conversation thread. In this example or any other example disclosed herein, the thread summary identifies an estimated time for presenting the conversation thread. In this example or any other example disclosed herein, the thread summary identifies the number of recipients of the unviewed electronic communications of the conversation thread. In this example or any other example disclosed herein, the thread summary identifies the topic of the conversation thread.

[0142] According to another example of the present disclosure, a computing system includes: an audio output interface for outputting audio via one or more audio speakers; a logic subsystem; and a storage subsystem having instructions stored thereon that are executable by the logic subsystem for: receiving instructions to initiate an audio presentation of an electronic communication for a recipient; and in response to the instructions, outputting an audio presentation of a plurality of conversation threads via the audio interface in a presentation order, wherein each conversation thread includes two or more unviewed electronic communications for the recipient that are linked to each other in reply, the audio presentation includes two or more unviewed electronic communications of each conversation thread in chronological order, starting from the oldest unviewed electronic communication and continuing to the most recent unviewed electronic communication of the conversation thread, the conversation thread is before another conversation thread among the plurality of conversation threads, and the other conversation thread includes unviewed electronic communications that are interleaved in time between the oldest unviewed electronic communication and the most recent unviewed electronic communication of the conversation thread. In this example or any other example disclosed herein, the presentation order is based on the reverse chronological order of the most recent unviewed electronic communication of each conversation thread among the plurality of conversation threads, such that: a first conversation thread having a first most recent unviewed electronic communication is presented before a second conversation thread having a second most recent unviewed electronic communication, and the second most recent unviewed electronic communication is older than the first most recent unviewed electronic communication of the plurality of conversation threads. In this example or any other example disclosed herein, the instructions may further be executable by the logic subsystem for: receiving a second instruction to advance the audio presentation of the plurality of conversation threads; and in response to the second instruction, advancing the audio presentation of the plurality of conversation threads from the current conversation thread to a subsequent conversation thread in the presentation order. In this example or any other example disclosed herein, the instructions may further be executable by the logic subsystem for: receiving a second instruction to perform an action related to a conversation thread among the plurality of conversation threads; and in response to the second instruction, applying the action to each electronic conversation of the conversation thread. In this example or any other example disclosed herein, the audio presentation further includes: for each conversation thread among the plurality of conversation threads, a thread summary of the conversation thread that is presented as an audible output before the text content of the conversation thread. In this example or any other example disclosed herein, the thread summary identifies the number of unviewed electronic communications of the conversation thread. In this example or any other example disclosed herein, the thread summary identifies an estimated time for presenting the conversation thread. In this example or any other example disclosed herein, the thread summary identifies the topic of the conversation thread.

[0143] It will be understood that the configurations and / or methods described herein are exemplary in nature, and these specific embodiments or examples should not be taken in a limiting sense, as many variations are possible. The specific routines or methods described herein may represent one or more of any number of processing strategies. Accordingly, the various acts illustrated and / or described may be performed in the illustrated and / or described order, in other orders, in parallel, or may be omitted. Likewise, the order of the above-described processes may be changed.

[0144] The subject matter of the present disclosure includes all novel and non-obvious combinations and sub-combinations of the various processes, systems, and configurations, as well as other features, functions, acts, and / or properties disclosed herein and any and all equivalents thereof.

Claims

1. A method performed by a computing system, the method comprising: Receiving an instruction to initiate an audio presentation of an electronic communication to a recipient; In response to the instruction, audibly outputting each unviewed electronic communication in the most recent conversation thread, the most recent conversation thread including a set of the most recent unviewed, reply-linked electronic communications to the recipient, wherein each unviewed electronic communication in the most recent conversation thread is audibly output in chronological order, starting from the oldest unviewed electronic communication and continuing to the most recent unviewed electronic communication; and In response to completing the audible output of the most recent unviewed electronic communications from the conversation thread, audibly outputting each unviewed electronic communication in the next most recent conversation thread, the next most recent conversation thread including a set of the next most recent unviewed, reply-linked electronic communications to the recipient, wherein each unviewed electronic communication in the next most recent conversation thread is audibly output in chronological order, starting from the oldest unviewed electronic communication and continuing to the most recent unviewed electronic communication, wherein at least one unviewed electronic communication from the next most recent conversation thread is chronologically between two unviewed electronic communications from the most recent conversation thread, and wherein all unviewed electronic communications from the most recent conversation thread are audibly output before any unviewed electronic communication from the next most recent conversation thread is audibly output.

2. The method according to claim 1, further comprising: Receiving a second instruction to advance the audible output of the unviewed electronic communications; And In response to the second instruction, advancing the audible output to a subsequent unviewed electronic communication.

3. The method according to claim 1, further comprising: Receiving a second instruction to perform an action related to the conversation thread; And In response to the second instruction, applying the action to each electronic communication of the conversation thread.

4. A computing system, comprising: An audio output interface for outputting audio via one or more audio speakers; A logic subsystem; And A storage subsystem having stored thereon instructions executable by the logic subsystem to: Receive an instruction to initiate an audio presentation of an electronic communication to a recipient; And In response to the instruction, output an audio presentation of a plurality of conversation threads via the audio interface according to a presentation order, wherein each conversation thread includes two or more unviewed electronic communications to the recipient that are reply-linked to each other, The audio presentation including the two or more unviewed electronic communications in chronological order in each conversation thread, starting from the oldest unviewed electronic communication and continuing to the most recent unviewed electronic communication of the conversation thread, the conversation thread being before another conversation thread in the plurality of conversation threads, the another conversation thread including unviewed electronic communications that are interspersed in time between the oldest unviewed electronic communication and the most recent unviewed electronic communication of the conversation thread.

5. The computing system according to claim 4, wherein, The presentation order is based on the reverse chronological order of the most recent unviewed electronic communications in each of the plurality of conversation threads, such that: A first conversation thread having a first most recent unviewed electronic communication is presented before a second conversation thread having a second most recent unviewed electronic communication, the second most recent unviewed electronic communication being older than the first most recent unviewed electronic communication of the plurality of conversation threads.

6. The computing system according to claim 4, wherein, The instructions may also be executed by the logic subsystem for: Receiving a second instruction for advancing the audio presentation of the plurality of conversation threads; And In response to the second instruction, advancing the audio presentation of the plurality of conversation threads from a current conversation thread to a subsequent conversation thread in the presentation order.

7. The computing system according to claim 4, wherein, The instructions may also be executed by the logic subsystem for: Receiving a second instruction for performing an action associated with a conversation thread among the plurality of conversation threads; and In response to the second instruction, applying the action to each electronic conversation of the conversation thread.

8. The computing system according to claim 4, wherein, The audio presentation further includes: For each of the plurality of conversation threads, a thread summary of the conversation thread that is presented as an audible output before the text content of the conversation thread.

9. The computing system according to claim 8, wherein The thread summary identifies the number of unviewed electronic communications of the conversation thread.

10. The computing system according to claim 8, wherein, The thread summary identifies an estimated time for presenting the conversation thread.

11. The computing system according to claim 8, wherein, The thread summary identifies the subject of the conversation thread.

Citation Information

Patent Citations

  • Short message broadcasting method and system

    CN103533519A

  • Method to set the flag as replied or forwarded to all replied or forwarded voice messages

    US20100322397A1

  • Smart communications assistant with audio interface

    WO2019079079A1