Drill back to original audio clip in virtual assistant initiated list and reminder

By storing and replaying the original voice input and context, the system addresses virtual assistant misinterpretations, improving task accuracy and user verification.

JP2025183215APending Publication Date: 2025-12-16ORACLE INT CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025134671
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-01-04
Filing Date
2025-08-13
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

Virtual assistants often misinterpret voice commands, leading to inaccurate task execution and incomplete information provision, necessitating a mechanism to review and correct these errors.

Method used

The system stores and replays the original voice input associated with a task, along with context information, allowing users to verify and correct misinterpretations by the virtual assistant.

Benefits of technology

Enables users to accurately retrieve and review the original voice command and context, enhancing the accuracy of task execution and information provision by virtual assistants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025183215000001_ABST
    Figure 2025183215000001_ABST
Patent Text Reader

Abstract

To provide a non-transitory computer-readable medium, method and system for drilling back to an original audio clip in virtual assistant initiated lists and reminders.SOLUTION: A method includes: receiving audio input comprising a first request; based on the first request, scheduling an action to be performed by a virtual assistant platform; storing at least a portion of the audio input and a mapping between the action and at least the portion of the audio input; performing the action; subsequent to performing the action, receiving a second request for audio playback of the first request corresponding to the action; retrieving at least the portion of the audio input based on the mapping between the action and at least the portion of the audio input; and playing at least the portion of the audio input comprising the first request.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to virtual assistants. In particular, the present disclosure relates to providing original audio recordings of voice inputs that caused a virtual assistant to create content intended for future review. [Background technology]

[0002] background A virtual assistant is a software agent used to perform tasks. A virtual assistant may accept instructions from a user via voice commands and / or text commands. Voice commands may be received by a smart speaker. Alternatively, a virtual assistant may receive commands from a user via text commands typed into a chat interface. Generally, a virtual assistant performs simple tasks in response to a request. For example, a virtual assistant may read today's weather forecast in response to a voice command such as, "What's the weather like today?"

[0003] The virtual assistant may also be used to create content such as lists and reminders that the user intends to review in the future. For example, a user may add items to a shopping list with the intention of reviewing the list content at the supermarket. As another example, a user may give the voice command "Remind me to call John tomorrow at noon," with the intention of reviewing the reminder when it is presented the next day at noon.

[0004] A virtual assistant may use a specific application or module to perform a specific task. As an example, a virtual assistant may call a stand-alone application to find directions, check the weather, or update a calendar. A virtual assistant may determine a user's intent to identify a task to perform. A virtual assistant may determine this intent using sample utterances. As an example, a virtual assistant may call an application called lookupBalance based on sample utterances such as "What's the balance in my savings account?" The application is invoked.

[0005] The approaches described in this paragraph are approaches that could be pursued, but not necessarily approaches that have been previously conceived or pursued, and thus, unless otherwise indicated, it should not be assumed that any of the approaches described in this paragraph qualify as prior art merely by virtue of their inclusion in this paragraph.

[0006] Embodiments are illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings. It should be noted that references to "an" or "one" embodiment in this disclosure do not necessarily refer to the same embodiment, but rather to at least one. [Brief explanation of the drawings]

[0007] [Figure 1] FIG. 1 illustrates a system according to one or more embodiments. [Figure 2] FIG. 1 illustrates a set of example actions for drilling back into an audio recording using a virtual assistant, according to one or more embodiments. [Figure 3] FIG. 1 illustrates an example device using a virtual assistant platform, according to one or more embodiments. [Figure 4] FIG. 1 is a block diagram illustrating a computer system according to one or more embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0008] Detailed Description In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding. One or more embodiments may be practiced without these specific details. Features described in one embodiment may be combined with features described in another embodiment. In some instances, well-known structures and devices are shown in block diagram form in order to avoid unnecessarily obscuring the present invention.

[0009] 1. Overview 2. Virtual Assistant System 3. Drill back into voice recordings in virtual assistants 4. Exemplary Embodiments 5. Other: Extension examples 6. Hardware Overview 1. Overview One or more embodiments play previously received voice input used to configure or schedule a task. In one example, a virtual assistant receives an initial command from a user via voice input to perform a task. The task may include, for example, setting a reminder to do something at a specific time. The virtual assistant associates the voice input with information corresponding to the task and stores it. After performing the task, the virtual assistant receives a request to provide additional information corresponding to the initial command. The virtual assistant identifies the stored voice input based on a stored mapping between the information corresponding to the task and the voice input. The virtual assistant then plays the voice input received from the user to the user.

[0010] Storing and playing back voice inputs can be useful when tasks performed by a virtual assistant are insufficient and / or inaccurate. In one example, the initial command uttered by the user is "Remind me to call Joe at 5 p.m." The virtual assistant misinterprets the initial command and instead plays a reminder at 5 PM stating, "This is a reminder to call Mo." The user may not recognize this reminder and may request that the initial command be played back. When the initial command received from the user is played back to the user, the user may realize that the reminder is to call "Joe" instead of "Mo." The initial command may also specify details that were not included in the reminder played by the virtual assistant. In one example, the initial command may be, "Remind me to call Larry to talk about selling a house." "Please call me," but the reminder played by the virtual assistant simply said, "This is a reminder to call Larry." Therefore, storing and replaying the initial command by the virtual assistant can help the user obtain additional information related to the reminder task.

[0011] In another example, an initial command spoken by a user includes "Add a treat to my grocery list." The virtual assistant incorrectly interprets this initial command. interprets the error and adds "wheat" to the user's grocery list. When the user reviews the grocery list and sees "wheat" on the grocery list, The user requests additional information about the "wheat" entry in the grocery list. The Assistant will search for the stored commands between the "wheat" entry and the stored initial command. Based on the mapping, the system identifies the initial command for playback. The system plays the user's initial command, "Add a treat to my grocery list." This helps the user identify the correct item to purchase.

[0012] One or more embodiments play or present any context-related information associated with receiving an initial command from a user. Instead of or in addition to playing the initial command, the virtual assistant may present (a) the geolocation of the user and / or the virtual assistant when the initial command was received, or (b) the time when the initial command was received. The virtual assistant may present device information corresponding to the time when the initial command was received. The device information may include, for example, the set of applications running when the initial command was received, the last application the user accessed before receiving the initial command, or the configuration of the device when the initial command was received.

[0013] One or more embodiments described herein and / or claimed below may not be included in this Summary section.

[0014] 2. Virtual Assistant System FIG. 1 illustrates a system 100 according to one or more embodiments. As shown in FIG. 1, the system 100 includes a query system 102, a user communication device 118, and a data repository 126. In one or more embodiments, the system 100 may include more or fewer components than those illustrated in FIG. 1. The components illustrated in FIG. 1 may be local to one another or remote from one another. The components illustrated in FIG. 1 may be implemented in software and / or hardware. Each component may be distributed across multiple applications and / or machines. Multiple components may be combined into one application and / or machine. Operations described with respect to one component may instead be performed by another component.

[0015] In one or more embodiments, the system 100 performs tasks based on input from the user 124. Exemplary tasks include making travel arrangements, providing directions, displaying requested images, setting reminders, creating shopping lists, and adding items to those lists. One or more steps within a task may be performed based on an interaction with the user 124. The interaction may include input received from the user 124 and output generated by the system 100. The interaction may include an initial request from the user 124. The interaction may include a response from the system 100 that resolves the user request. The interaction may include a request generated by the system 100 for additional information from the user 124.

[0016] In one or more embodiments, the user communication device 118 includes hardware and / or software configured to facilitate communication with the user 124. The user communication device 118 may receive information from the user 124. The user communication device 118 may transmit information to the user 124. The user communication device may facilitate communication with the user 124 via an audio interface 120 and / or a visual interface 122. The user communication device 118 is communicatively coupled to the query system 102.

[0017] In one embodiment, the user communication device 118 is implemented on one or more digital devices. The term "digital device" generally refers to any hardware device that includes a processor. A digital device may refer to a physical device that runs an application or a virtual machine. Examples of digital devices include computers, tablets, laptops, and the like. This includes laptops, desktops, netbooks, servers, web servers, network policy servers, proxy servers, general-purpose machines, specific function hardware devices, hardware routers, hardware switches, hardware firewalls, hardware firewalls, hardware network address translators (NATs), hardware load balancers, mainframes, televisions, content receivers, set-top boxes, printers, mobile phones, smartphones, personal digital assistants (PDAs), wireless receivers and / or transmitters, base stations, communications management devices, routers, switches, controllers, access points, and / or client devices.

[0018] In one embodiment, the user communication device 118 is a smart speaker. The smart speaker receives voice data from the user 124. The smart speaker plays the voice. The smart speaker exchanges information with the query system 102. The smart speaker can be implemented as a stand-alone device or as part of a smart device such as a smartphone, tablet, or computer.

[0019] In one or more embodiments, audio interface 120 refers to hardware and / or software configured to facilitate audio communication between user 124 and user communication device 118. Audio interface 120 may include a speaker for playing audio. The played audio may include spoken questions and answers, including conversations. Audio interface 120 may include a microphone for receiving audio. The received audio may include requests and other information received from user 124.

[0020] In one or more embodiments, visual interface 122 refers to hardware and / or software configured to facilitate visual communication between a user and user communication device 118. Visual interface 122 renders user interface elements and receives input through such user interface elements. Examples of visual interfaces include graphical user interfaces (GUIs) and command line interfaces (CLIs). Examples of user interface elements include check boxes, radio buttons, drop-down lists, list boxes, buttons, toggles, text fields, date and time selectors, command lines, sliders, pages, and forms.

[0021] The visual interface 122 may present a messaging interface. The messaging interface may be used to accept typed input from the user (e.g., via a keyboard coupled to the user communication device 118 and / or a soft keyboard displayed via the visual interface). The messaging interface may also be used to display text to the user 124. The visual interface 122 may include functionality for displaying images such as maps and photographs. The visual interface 122 may include functionality for uploading images. In one example, the user 124 uploads a photo of an animal with the text, "What is this?"

[0022] In an embodiment, query system 102 is a system for performing one or more tasks in response to input from a user 124 via a user communication device 118. Query system 102 may receive voice input, text input, and / or images from user communication device 118.

[0023] The query system 102 may include a speech recognition component 108 for converting voice input to text using speech recognition technology. The speech recognition component 108 may digitize and / or filter the received voice input. The speech recognition component 108 may compare the voice input to stored template sound samples to identify words or phrases. The speech recognition component 108 may separate the voice input into components for comparing it to sounds used in a particular language.

[0024] In one embodiment, one or more of the speech recognition components 108 employ a machine learning engine 110. In particular, the machine learning engine 110 may be used to recognize speech and determine its associated meaning. Machine learning encompasses a variety of techniques in the field of artificial intelligence that address computer-implemented, user-independent processes for solving problems involving varying inputs.

[0025] In some embodiments, the machine learning engine 110 trains the machine learning model 112 to perform one or more operations. When training the machine learning model 112, training data is used to generate a function that calculates a corresponding output given one or more inputs to the machine learning model 112. The output may correspond to a prediction based on traditional machine learning. In one embodiment, the output includes a label, classification, and / or categorization assigned to the provided input. The machine learning model 112 corresponds to a trained model for performing a desired operation (e.g., labeling, classifying, and / or categorizing the input). As a particular example, training may include having an individual speaker read specific text or isolated vocabulary into the system. The system analyzes the speaker's specific voice and uses it to fine-tune recognition of that person's speech.

[0026] In one embodiment, the machine learning engine 110 may use supervised learning, semi-supervised learning, unsupervised learning, reinforcement learning, and / or another training method, or a combination thereof. In supervised learning, labeled training data includes input / output pairs, where each input is labeled with a desired output (e.g., a label, classification, and / or categorization), also referred to as a supervised signal. In semi-supervised learning, some inputs are associated with supervised signals, and other inputs are not. In unsupervised learning, the training data does not include supervised signals. Reinforcement learning uses a feedback system in which the machine learning engine 110 receives positive and / or negative reinforcement in the process of attempting to solve a particular problem (e.g., to optimize performance in a particular scenario according to one or more predefined performance criteria). In one embodiment, the machine learning engine 110 initially trains the machine learning model 112 using supervised learning and then continuously updates the machine learning model 112 using unsupervised learning.

[0027] In one embodiment, the machine learning engine 110 may label, classify, and / or categorize inputs using many different techniques. The machine learning engine 110 may convert inputs into feature vectors that describe one or more characteristics (“features”) of the inputs. The machine learning engine 110 may label, classify, and / or categorize the inputs based on the feature vectors. Alternatively or additionally, the machine learning engine 110 may use clustering (also referred to as cluster analysis) to identify commonalities in the inputs. The machine learning engine 110 may group (i.e., cluster) the inputs based on their commonalities. The machine learning engine 110 may use hierarchical clustering, k-means clustering, and / or another clustering method, or a combination thereof. For example, the machine learning engine 110 may receive one or more parsed query terms as inputs and may identify one or more additional parsed query terms to include in the search based on the commonalities among the received parsed query terms. In one embodiment, the machine learning engine 110 includes an artificial neural network. An artificial neural network includes multiple nodes (also called artificial neurons) and edges between the nodes. Each edge may be associated with a corresponding weight representing the strength of the connection between the nodes, which the machine learning engine 110 adjusts as the machine learning progresses. Alternatively or additionally, the machine learning engine 110 may include a support vector machine. A support vector machine represents inputs as vectors. The machine learning engine 110 may label, classify, and / or categorize the inputs based on the vectors. Alternatively or additionally, the machine learning engine 110 may use a Naive Bayes classifier to label, classify, and / or categorize the inputs. Alternatively or additionally, given a particular input, the machine learning model may apply a decision tree to predict an output for a given input. Alternatively or additionally, the machine learning engine 110 may apply fuzzy logic in situations where it is impossible or impractical to label, classify, and / or categorize the inputs among a fixed set of mutually exclusive options. The above-described machine learning models 112 and techniques are described for illustrative purposes only and should not be construed as limiting one or more embodiments.

[0028] As a particular example, the machine learning engine 110 may be based on a Hidden Markov Model (HMM). An HMM may output a sequence of symbols or quantities. HMMs can be used in speech recognition because the signal can be viewed as a piecewise stationary signal or a short-term stationary signal. On short time scales (e.g., 10 milliseconds), speech can be approximated as a stationary process. Speech can be thought of as a Markov model for many probabilistic purposes.

[0029] HMMs can be trained automatically and are simple and computationally feasible to use. In speech recognition, an HMM may output a sequence of multidimensional real-valued vectors. In this case, the vector output is periodic (e.g., every 10 milliseconds). Each vector may contain coefficients obtained by applying an algorithm (e.g., a Fourier transform) to a short-time window of speech and using the most significant coefficients. Each word or phoneme may have a different output distribution. A hidden Markov model for a series of words or phonemes is created by concatenating individual trained hidden Markov models for separate words or phonemes. Speech decoding may use an algorithm (e.g., the Viterbi algorithm) to find the best path.

[0030] Neural networks may be used in speech recognition, and in particular have been used in many aspects of speech recognition, such as phoneme classification, phoneme classification by multi-objective evolutionary algorithms, and isolated word recognition.

[0031] Neural networks can be used to estimate the probabilities of speech feature segments and enable discriminative training in a natural and efficient manner. In particular, neural networks can be used in preprocessing, feature transformation, or dimensionality reduction, steps prior to HMM-based recognition. Alternatively, recurrent neural networks (RNNs) and / or time-delay neural networks (TDNNs) can be used.

[0032] In one embodiment, as the machine learning engine 110 applies various inputs to the machine learning model 112, the corresponding output may not always be accurate. As an example, the machine learning engine 110 may train the machine learning model 112 using supervised learning. After training the machine learning model 112, if the subsequent inputs are the same as the inputs included in the labeled training data and the output is the same as the supervised signals in the training data, then the output is certain to be accurate. If the inputs are different from the inputs included in the labeled training data, the machine learning engine 110 may generate a corresponding output that is inaccurate or of uncertain accuracy. Generating a particular output for a given input In addition, the machine learning engine 110 may be configured to generate an indicator that represents confidence (or lack thereof) in the accuracy of the output. The confidence indicator may include a numeric score, a Boolean value, and / or any other type of indicator that corresponds to confidence (or lack thereof) in the accuracy of the output.

[0033] The speech recognition component 108 may use natural language processing to identify one or more executable commands based on the text generated from the voice input, and the query system 102 may parse the text to determine one or more relevant portions of the text.

[0034] In particular, the speech recognition component can parse the text to identify the wake word or phrase (e.g., a word used to indicate that the user intends to issue a command to the virtual assistant) and the location of the command word or phrase that follows the wake word. The query system 102 can identify keywords in the command. The query system 102 can compare the command with template language associated with tasks that can be performed by the query system 102.

[0035] In an embodiment, the query system 102 is implemented remotely from the user communication device 118. The query system 102 may run on a cloud network. Alternatively, the query system 102 may run locally to the user communication device 118. The query system 102 may perform tasks or retrieve information from one or more external servers. As an example, the query system 102 may retrieve traffic data from a third-party map application.

[0036] In one embodiment, the query system 102 includes a context information collector 104 that determines context information associated with the user communication device 118. The context information collector 104 may receive data from the user communication device 118. As an example, the query system may use Global Positioning System (GPS) capabilities. or other geolocation services. As another example, the query system 102 may query the user communication device 118 to determine information regarding the state of the user communication device 118. In particular, the query system may request the set of applications running when the initial command was received, the last application the user accessed before receiving the initial command, and / or the configuration of the device when the initial command was received. Additionally or alternatively, the query system 102 may query the user communication device 118 to determine the time the initial command was received.

[0037] The query system may rely at least in part on the user input history 106 when determining what action to take. In one embodiment, the user input history 106 is a record of user inputs over the course of one or more interactions. The user input history 106 may include information about the sequence of a series of voice inputs. The user input history 106 may categorize user inputs by type. For example, a user creates a shopping list at 6:00 PM on Thursday.

[0038] The query system 102 may include a command execution component 116 for executing commands based on the received voice input. The command execution component 115 may include executing commands and / or scheduling commands to be executed in the future.

[0039] The query system 102 may include a voice input storage and retrieval component 114. The voice input storage and retrieval component 114 may store at least a portion of the voice input in a voice record. The voice input storage and retrieval component 114 may communicate with the data repository 126 for storage as voice recording data. In some embodiments, the voice input storage and retrieval component 114 stores voice recording data associated with each received voice input. In other embodiments, the voice input storage and retrieval component 114 stores voice recording data associated with the voice input based on characteristics of the received voice input. For example, the voice input storage and retrieval component 114 may store voice recording data in response to a determination that the speech recognition component 108 was unable to determine one or more phonemes of the voice input or that a confidence level associated with the determination of one or more phonemes is below a confidence threshold. As another example, the voice input storage and retrieval component 114 may store voice recording data in response to a determination that the command corresponds to an action to be scheduled for future execution (e.g., a reminder) and / or adding an item to a list.

[0040] In one or more embodiments, data repository 126 is any type of storage unit and / or device for storing data (e.g., a file system, a database, a collection of tables, or any other storage mechanism). Furthermore, data repository 126 may include multiple different storage units and / or devices. The multiple different storage units and / or devices may or may not be of the same type, or may or may not be located at the same physical site. Furthermore, data repository 126 may be implemented or executed on the same computing system as query system 102 and / or user communication device 118. Alternatively or additionally, data repository 126 may be implemented or executed on a storage system and / or computing system separate from query system 102 and / or user communication device 118. Data repository 126 may be communicatively coupled to query system 102 and / or user communication device 118 by a direct connection or via a network.

[0041] The voice recording data 128 includes at least a portion of a voice recording received as a user command from a user communication device. For example, the voice recording data may include the entirety of the received voice data (e.g., the wake word and the command). As another example, the voice recording data may not include the entirety of the received voice data (e.g., the command without the wake word). The voice recording data may be stored in any format playable on the user communication device.

[0042] The context information 130 includes context information related to the user communication device 118 and / or the user 124 at the time the user issued the command (e.g., the time the voice recording was sent from the user communication device 118 to the query system 102). By way of example, the context information may include location data indicating the location of the user communication device at the time the command was issued, application information indicating other applications running at the time the command was issued and / or the application last accessed before receiving the command, configuration information indicating the configuration of the device when the initial command was received, timing information indicating the time and / or date the command was received, and / or other information indicative of the state of the user communication device at the time the command was received.

[0043] Scheduled action information 132 may include information associated with an action scheduled to be performed at a future time based on the received voice input. For example, the scheduled action information may include a command identifier associated with the scheduled action, a time at which the action should be performed, or other information associated with the scheduled action.

[0044] Mapping information 134 may include information that associates voice recording data and / or context information with particular commands issued by a user. In some embodiments, mapping information associates particular voice recording data and / or particular context data with particular scheduled action information. This data may be used to retrieve stored voice recording data and / or context information in response to a request for the performance of a particular action and / or information associated with a particular command.

[0045] 3. Drill back into voice recordings in virtual assistants 2 shows an example set of operations for drilling back to an original audio recording using a virtual assistant according to one or more embodiments. One or more of the operations shown in FIG. 2 may be modified, rearranged, or omitted altogether. Therefore, the particular sequence of operations shown in FIG. 2 should not be construed as limiting the scope of one or more embodiments.

[0046] In one embodiment, the query system may receive voice input including a request to perform an action (operation 202). The voice input is received via a microphone system input (e.g., via a user communication device as described above). In some embodiments, the voice input includes a command (one or more words that cause the system to perform an action) as well as a wake word (a word that prepares the system to receive a command). In some embodiments, the command may be a type of command that causes the system to perform an action but does not provide specific immediate feedback to the user.

[0047] The system may parse the speech input and divide the speech input into sections. For example, the system may divide the speech input into a wake word section and a command section. In some embodiments, the system may convert the received speech input into text. The system may process one or more (e.g., each) sections of the speech input using natural language processing to generate a transcript of the section. In some embodiments, the system may determine, for each transcribed section, a confidence score associated with the transcript of the section.

[0048] In one embodiment, the system schedules an action to be performed (operation 204). In some embodiments, scheduling an action to be performed may include determining an action to be performed based on the received voice input. In some embodiments, the command may be a type of command that causes the system to perform an action at a future time. In particular, the system may receive the command, "Remind me to call John next Thursday at noon." In response to the command, the system may schedule a reminder that causes the system to prompt the user with the phrase "Call John" at 12:00 PM next Thursday. In some embodiments, the system may provide feedback to the user such as "Reminder set" or "OK, remind me," but may not provide specific feedback about what form the reminder text will take.

[0049] As another example, the command may be a type of command that adds an entry to a list. In particular, the system may receive the command "Add milk to my grocery list." In response, the system may add the list item "milk" to the list. In some embodiments, the system may provide feedback to the user, such as "Okay, I'll add that to your list," without specifying what the list item is.

[0050] The system may store information (e.g., at least a portion of the voice input and / or contextual information associated with the user and / or user device) in association with the scheduled action (operation 206). The system may store at least a portion of the voice input. For example, the system may store the entire voice input, the portion of the voice input that corresponds to the command, the portion of the voice input that cannot be translated by a speech-to-text process running on the system, and / or any other portion of the voice input. Alternatively or additionally, the system may store contextual information. For example, the contextual information may include location data indicating the location of the user communication device at the time the command was issued, application information indicating other applications running at the time the command was issued and / or the application last accessed before receiving the command, configuration information indicating the configuration of the device when the initial command was received, timing information indicating the time and / or date the command was received, and / or other information useful for providing context for the command.

[0051] In some embodiments, the system stores a mapping between commands and stored voice inputs and / or stored context information. For example, each command may be associated with a particular identifier, and the identifier may be stored in combination with the voice input and / or context information.

[0052] In some embodiments, the system may store the voice input and / or the context information based on one or more characteristics of the voice input. For example, the system may determine a command type associated with the voice input. The system may determine whether to store the voice input and / or the context information based on the command type.

[0053] In some embodiments, the system may store the audio input and / or the contextual information based on a confidence score of a transcript of the audio input. For example, the system may store the audio input and / or the contextual information based on a determination that a confidence score associated with at least one section of the audio input does not exceed a certain threshold.

[0054] In some embodiments, storing at least a portion of the speech input includes storing the entire speech input, hi some embodiments, storing at least a portion of the speech input includes parsing the speech input to determine a particular portion of the speech input associated with a command and storing only the portion of the speech input associated with the command.

[0055] The system may perform the action requested by the user (operation 208). That is, the system may perform the action according to the command portion of the voice input. In some embodiments, the system may perform the action approximately simultaneously upon receiving the voice input. For example, in response to receiving a voice input including the command, "Add milk to my grocery list," the system may add the item "milk" to the grocery list. In some embodiments, the system may perform the action at a time after receiving the voice input. For example, the system may receive a voice input including the command, "Remind me to call John next Thursday at noon." The system may refrain from performing the action (e.g., providing a reminder) until the time specified in the command (e.g., Thursday at noon). At the specified time, the system performs the action to remind the user to call John. For example, the system may present the user with a visual reminder and / or an audio reminder.

[0056] In some embodiments, performing the action includes indicating that one or more of the voice input or contextual information associated with the action is stored. For example, the system may display an icon indicating that the contextual information and / or voice input data associated with the action is stored.

[0057] In response to performing the action, the system may receive a request to present the stored context information and / or voice input data (operation 210). In some embodiments, the request may include clicking or activating a displayed icon. Alternatively or additionally, the request may include a subsequent voice request to present the stored information. As an example, after the system performs the action, the user may issue the voice command, "Play my original audio."

[0058] In some embodiments, upon receiving a request to present stored context information and / or speech input data, the system may mark at least the stored portion of the speech input for use in training data for the system. That is, the system may determine, based on the request, that there is an error in the transcription. The speech input may be used in training to help improve the accuracy of future transcriptions.

[0059] In response to a request to present the stored context information and / or voice input data, the system may retrieve the stored context information and / or voice input data (operation 212). The system may determine the most recent prior action and an identifier associated with that action. The system may determine the context information and / or voice input data associated with the action based on the stored mapping information. Based on this association, the system may retrieve the context information and / or voice input data by known means, such as by reading the context information and / or voice input data from a data repository.

[0060] The system may present the retrieved context information and / or the voice input data (operation 214). Presenting the context information and / or the voice input data includes playing back at least a portion of the retrieved voice data. In some embodiments, the system may play back the entirety of the retrieved voice data. Alternatively, the system may play back a portion of the retrieved voice data that corresponds to the command. Additionally or alternatively, presenting the context information and / or the voice input data includes presenting the retrieved context information. The retrieved context information may be presented visually on a display, played audibly (e.g., using a text-to-speech algorithm), and / or provided via a message (e.g., email, text message, log file, or other message).

[0061] 4. Exemplary Embodiments Detailed examples are provided below for clarity. The components and / or operations described below should be understood as examples that may not be applicable to a particular embodiment. Thus, the components and / or operations described below should not be construed as limiting the scope of any of the claims.

[0062] 3 shows an example device 300 using a virtual assistant platform. The user previously provided a voice input including the command "Remind me to call John next Thursday at noon."

[0063] The virtual assistant platform 304 performs speech-to-text translation for voice input. A transcription error in the speech-to-text process caused the virtual assistant platform to translate the command received by the user as "Remind me to call Paul Fawn next Thursday at noon." So the virtual assistant platform would send a reminder for Thursday at noon containing the text "Paul Fawn." The system also stores the voice input and context information associated with the device at the time the voice input was received.

[0064] At noon on Thursday, the virtual assistant platform 304 causes the device 300 to present a scheduled reminder 306 on the display 302. The reminder 306 includes the incorrectly transcribed text "Paul Fawn." The reminder 306 includes an icon 308. The icon 308 , indicating that the virtual assistant platform stored the voice input that caused the virtual assistant platform to schedule the reminder. Reminder 306 includes icon 310. Icon 310 indicates that the virtual assistant platform stored context information associated with the device at the time the voice input was received.

[0065] The user may request that the platform play the voice input that caused the reminder to be scheduled by clicking or activating icon 308 and / or issuing a voice command such as "play original audio." In response to such a command, virtual assistant 304 may retrieve and play the stored voice input.

[0066] The user may request that context information associated with the device be presented at the time the voice input is received by clicking or activating icon 310 and / or issuing a voice command such as "Show me context information." In response to such a command, virtual assistant 304 may retrieve and display stored context information.

[0067] 5. Other: Extension examples Some embodiments are directed to systems having one or more devices including a hardware processor, the one or more devices configured to perform any of the operations described herein and / or recited in any of the appended claims.

[0068] In an embodiment, a non-transitory computer-readable storage medium includes instructions that, when executed by one or more hardware processors, cause the processor to perform any of the operations described herein and / or recited in any of the appended claims.

[0069] Any combination of the features and functions described herein may be used in accordance with one or more embodiments. In the foregoing specification, the embodiments have been described with reference to numerous specific details that may vary from implementation to implementation. Accordingly, the specification and accompanying drawings should be regarded in an illustrative rather than a limiting sense. The sole and exclusive indication of the scope of the invention, and what is intended by the applicant to be the scope of the invention, is the literal and equivalent scope of the set of claims issuing from this application in the specific form in which such claims arise, including any subsequent amendments.

[0070] 6. Hardware Overview According to one embodiment, the techniques described herein involve one or more dedicated computing devices. These special-purpose computing devices may be hardwired to perform these techniques, or may include one or more application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or network processing units (NPUs) or other digital electronic devices that are permanently programmed to perform these techniques, or may be implemented using firmware. The special-purpose computing device may include one or more general-purpose hardware processors programmed to perform these techniques according to program instructions in software, memory, other storage, or a combination thereof. Such special-purpose computing devices may also combine custom hardwired logic, ASICs, FPGAs, or NPUs with custom programming to achieve these techniques. The special-purpose computing device may be a desktop computer system, a portable computer system, a handheld device, a networking device, or any other device incorporating hardwired logic and / or program logic to implement these techniques.

[0071] For example, Figure 4 is a block diagram illustrating a computer system 400 upon which one embodiment of the present invention may be implemented. Computer system 400 includes a bus 402 or other communication mechanism for communicating information, and a hardware processor 404 coupled with bus 402 for processing information. Hardware processor 404 may be, for example, a general-purpose microprocessor.

[0072] Computer system 400 also includes a main memory 406, such as a random access memory (RAM) or other dynamic storage device, coupled to bus 402 for storing instructions and information to be executed by processor 404. Main memory 406 may also be used to store temporary variables or other intermediate information during execution of instructions to be executed by processor 404. Such instructions, when stored on a non-transitory storage medium accessible to processor 404, render computer system 400 a special-purpose machine customized to perform the operations specified in the instructions.

[0073] Computer system 400 further includes a read only memory (ROM) 408 or other static storage device coupled to bus 402 for storing static information and instructions for processor 404. A storage device 410, such as a magnetic disk or optical disk, is provided and coupled to bus 402 for storing information and instructions.

[0074] Computer system 400 may be coupled via bus 402 to a display 412, such as a cathode ray tube (CRT), for displaying information to a computer user. An input device 414, including alphanumeric and other keys, is coupled to bus 402 for communicating information and command selections to processor 404. Another type of user input device is a cursor control device 416, such as a mouse, trackball, or cursor direction keys, for communicating directional information and command selections to processor 404 and for controlling cursor movement on display 412. The input device typically has two degrees of freedom in two axes, a first axis (e.g., x) and a second axis (e.g., y), allowing the device to specify a position in a plane.

[0075] The computer system 400 may be configured with a computer system that, when combined with the computer system, makes the computer system 400 a dedicated machine or programs it to be a dedicated machine. The techniques described herein may be implemented using customized hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic. According to one embodiment, the techniques herein are performed by computer system 400 in response to processor 404 executing one or more sequences of one or more instructions contained in main memory 406. Such instructions may be read into main memory 406 from another storage medium, such as storage device 410. Execution of the sequences of instructions contained in main memory 406 causes processor 404 to perform the process steps described herein. In alternative embodiments, hardwired circuitry may be used in place of or in combination with software instructions.

[0076] The term "storage medium" as used herein refers to any non-transitory medium that stores data and / or instructions that cause a machine to operate in a specific manner. Such storage media may include non-volatile media and / or volatile media. Non-volatile media include, for example, optical or magnetic disks, such as storage device 410. Volatile media include dynamic memory, such as main memory 406. Common forms of storage media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tape, or any other magnetic data storage medium, CD-ROMs, any other optical data storage medium, any physical medium with a pattern of holes, RAM, PROM, EPROM, FLASH-EPROM, NVRAM, any other memory chip or cartridge, content-addressable memory (CAM), and ternary content-addressable memory (TCAM).

[0077] Storage media are distinct from but may be used in conjunction with transmission media. Transmission media involves transferring information between storage media. For example, transmission media include coaxial cables, copper wire and fiber optics, including the wires that comprise bus 402. Transmission media can also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications.

[0078] Various forms of media may be involved in carrying one or more sequences of one or more instructions to processor 404 for execution. For example, the instructions may initially be carried on a magnetic disk or solid state drive of a remote computer. The remote computer may load the instructions into its dynamic memory and send the instructions over a telephone line using a modem. A modem local to computer system 400 may receive the data on the telephone line and use an infrared transmitter to convert the data to an infrared signal. An infrared detector may receive the data carried in the infrared signal and appropriate circuitry may place the data on bus 402. Bus 402 carries the data to main memory 406, from which processor 404 retrieves and executes the instructions. The instructions received by main memory 406 may optionally be stored on storage device 410 either before or after execution by processor 404.

[0079] Computer system 400 also includes a communication interface 418 coupled to bus 402. Communication interface 418 provides a two-way data communication coupling to a network link 420 that is connected to a local network 422. For example, communication interface 418 may be coupled to an integrated services digital network (ISDN) or other network. The communication interface 418 may be an ISDN (Instrumental Services Digital Network) card, cable modem, satellite modem, or modem to provide a data communication connection to a corresponding type of telephone line. As another example, communication interface 418 may be a local area network (LAN) card to provide a data communication connection to a compatible LAN. A wireless link may also be implemented. In any such implementation, communication interface 418 sends and receives electrical, electromagnetic or optical signals that carry digital data streams representing various types of information.

[0080] Network link 420 typically provides data communication through one or more networks to other data devices. For example, network link 420 may provide data communication through local network 422 to a host computer 424 or to a data facility operated by an Internet Service Provider (ISP) 426. ISP 426 in turn provides data communication services through the worldwide packet data communication network now commonly referred to as the "Internet" 428. Local network 422 and Internet 428 both use electrical, electromagnetic or optical signals that carry digital data streams. The signals through the various networks and the signals on network link 420 and through communication interface 418, which carry the digital data to and from computer system 400, are exemplary forms of transmission media.

[0081] Computer system 400 can send messages and receive data, including program code, through the network(s), network link 420 and communication interface 418. In the Internet example, a server 430 might transmit a requested code for an application program through Internet 428, ISP 426, local network 422 and communication interface 418.

[0082] The received code may be executed by processor 404 as it is received, and / or stored in storage device 410, or other non-volatile storage for later execution.

[0083] In the foregoing specification, embodiments of the invention have been described with reference to numerous specific details that may vary depending on the implementation. Accordingly, the specification and accompanying drawings should be considered in an illustrative rather than a restrictive sense. The sole and exclusive indication of the scope of the invention, and what is intended by the applicant to be the scope of the invention, is the literal and equivalent scope of the set of claims issuing from this application in the specific form in which such claims arise, including any subsequent amendments.

Claims

1. A non-transitory computer-readable medium containing instructions that, when executed by one or more hardware processors, cause operations to be performed, the operations including: An operation in which the virtual assistant platform receives a voice input including a first request; An operation of scheduling an action to be performed by the virtual assistant platform based on the first request; (a) storing at least a portion of the speech input; and (b) storing a mapping between the action and at least a portion of the speech input. The virtual assistant platform executes the action; After the operation of performing the action, the virtual assistant platform receives a second request for audio playback of the first request corresponding to the action performed by the virtual assistant platform; retrieving at least a portion of the speech input based on the mapping between the action and at least a portion of the speech input; A non-transitory computer-readable medium, comprising: an operation in which the virtual assistant platform plays back at least a portion of the voice input including the first request.

2. The operation further comprises: (a) storing context information at a time when the first request was received; and (b) storing a mapping between the action and a user location, wherein the context information includes one or more of: (i) a user location at a time when the first request was received; (ii) a time when the first request was received; or (iii) a list of applications running at a time when the first request was received; and the operation further comprises: retrieving the context information at the time the first request was received based on the mapping between the action and the context information at the time the first request was received; The medium of claim 1, further comprising: an operation in which the virtual assistant platform presents at least a portion of the context information at the time the first request was received.

3. the act of storing at least a portion of the speech input and a mapping between the action and at least a portion of the speech input is performed in response to evaluating one or more characteristics of the speech input; The operation further comprises: The virtual assistant platform receives a second voice input including a second request; An operation of scheduling a second action to be performed by the virtual assistant platform based on the second request; evaluating one or more characteristics of the second speech input; and refraining from storing any portion of the second voice input.

4. The operation further comprises: determining a confidence level associated with scheduling the action to be performed based on the first request; and in response to the determined confidence level being below a confidence threshold, prompting a user to confirm the action.

5. The act of storing at least a portion of the voice input may include the act of storing all of the voice input. The medium of claim 1 , comprising:

6. The medium of claim 1 , wherein the act of playing back at least a portion of the audio input comprises an act of playing back all of the audio input.

7. The operation further comprises: Parsing the speech input to divide the input into sections; determining a particular section of the plurality of sections of the speech input that includes the first request; The medium of claim 1 , wherein storing at least a portion of the audio input comprises storing the particular section.

8. The operation further comprises: Parsing the speech input to divide the input into sections; generating, for each of the plurality of sections of the audio input, a transcript of the audio contained in the section; and determining, for each of the plurality of sections of the speech input, a likelihood that a transcript of the section will contain a transcription error; The medium of claim 1 , wherein storing at least a portion of the speech input comprises storing at least a section where the transcript of the section is most likely to include a transcription error.

9. The operation further comprises:

10. The medium of claim 1, further comprising: an operation of marking at least a portion of the speech input to be included in a training dataset in response to an operation of receiving the second request for speech playback of the first request corresponding to the action performed by a virtual assistant platform.

10. A method comprising the operations of any one of claims 1 to 9.

11. A system comprising means for performing the operations of any one of claims 1 to 9.

12. 1. A system comprising: at least one device including a hardware processor; The system is configured to perform the operations of any one of claims 1 to 9.