Method, system, and program for synchronizing virtual assistant voice responses with user activities
The method synchronizes voice responses from virtual assistants with user activities by processing user tasks and sub-activities, addressing inefficiencies in existing systems and improving user experience.
Patent Information
- Application Number
- JP2023515698
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-09-24
- Filing Date
- 2021-08-02
- Publication Date
- 2025-05-19
- Estimated Expiration
- 2041-08-02
AI Technical Summary
Existing virtual assistants struggle to synchronize voice responses with user activities effectively, leading to inefficiencies and unnecessary processing resources being used.
A method and system that utilize one or more processors to identify user tasks corresponding to voice queries, generate sequences of sub-activities, determine completion status, and synchronize voice responses with the user's activities based on activity data from computing devices.
This approach improves efficiency by reducing unnecessary voice responses and processing resources, while also enhancing user experience by providing timely and relevant voice assistance.
Smart Images

Figure 0007679144000001 
Figure 0007679144000002 
Figure 0007679144000003
Abstract
Description
Technical Field
[0001] The present invention generally relates to the field of Internet of Things, and more particularly to synchronizing voice responses from virtual assistants based on user activities.
Background Art
[0002] In recent years, the development and growth of Internet of Things (IoT)-enabled devices have created rich opportunities to advance the ability to integrate systems. The Internet of Things (IoT) is the networking of physical devices (also referred to as "connected devices" and "smart devices"), vehicles, buildings, and other products embedded with electronic devices, software, sensors, actuators, and network connectivity, enabling these objects to collect and exchange data. The IoT enables the detection or control, or both, of objects remotely across existing network infrastructures, creating opportunities to more directly incorporate the physical world into computer-based systems, reducing human intervention and increasing efficiency, accuracy, and economic benefits. Each object is uniquely identifiable through its embedded computing system but can interoperate within existing Internet infrastructures.
[0003] A virtual assistant, also known as an artificial intelligence (AI) assistant or digital assistant, is an application program that understands natural language voice commands and completes tasks for the user. The user can ask the assistant questions, control home automation devices and media playback via voice, and manage other basic tasks such as email, to-do lists, and calendars with verbal commands. The capabilities and uses of virtual assistants are expanding rapidly, with new products entering the market and placing great emphasis on both email and voice user interfaces. SUMMARY OF THE INVENTION
[0004] Aspects of the present invention disclose a method, computer program, and system for synchronizing a voice response of an artificial intelligence (AI)-based voice assistant with user actions based on user activities. The method includes one or more processors identifying a user task corresponding to a user voice query. The method further includes one or more processors generating a sequence of sub-activities of the task corresponding to the user voice query. The method further includes one or more processors determining the completion status of each sub-activity of the sequence of sub-activities of the task corresponding to the user voice query based at least in part on activity data received from one or more computing devices within the user's operating environment. The method further includes one or more processors synchronizing a voice response of the computing device with the sequence of sub-activities of the task based at least in part on the completion status of each sub-activity of the sequence of sub-activities of the task. BRIEF DESCRIPTION OF THE DRAWINGS
[0005]
Figure 1
Figure 2
Figure 3
Embodiments for Carrying Out the Invention
[0006] Embodiments of the present invention enable synchronizing a voice response from an artificial intelligence (AI) voice assistant based on user activities. Embodiments of the present invention use multiple data sources to identify a user's past activities related to a user's query. Further embodiments of the present invention identify a user's activities related to a user's query. Embodiments of the present invention estimate the time used to execute each identified activity to determine whether the user is executing an essential subset of the activities related to the user's query. Further embodiments of the present invention identify one or more stages of an activity related to a user's query at which a voice response is appropriate. Further embodiments of the present invention synchronize a voice response to a user's query with user activities. Other embodiments of the present invention anticipate future activities (e.g., mistakes, activity patterns, etc.) by the user and generate additional voice responses based on the future activities.
[0007] Some embodiments of the present invention recognize that there are problems regarding a virtual assistant synchronizing a voice response with activities performed by a user. For example, while receiving a voice response from a virtual assistant, a user may wish to synchronize the voice response with an ongoing activity or a future activity the user wishes to perform. In this example, the user may have performed some steps of the activity and may wish to complete the remaining steps of the activity according to voice-based guidance. Embodiments of the present invention determine the steps of the activity at which the voice response should start and the timing of the voice response such that each voice response synchronizes with the steps of the user's activity.
[0008] Embodiments of the present invention can operate to improve the efficiency of a computer system by reducing the amount of processing resources used by the computer system by reducing the number of tasks executed in response to generating unnecessary voice responses. As a result, embodiments of the present invention reduce the power consumption associated with transmitting unnecessary voice responses.
[0009] Embodiments of the present invention can be implemented in various forms, and the details of exemplary implementations will be described below with reference to the drawings.
[0010] Here, the present invention will be described in detail with reference to the drawings. FIG. 1 is a functional block diagram illustrating a distributed data processing environment, indicated generally at 100, according to an embodiment of the present invention. FIG. 1 is merely an example of one implementation and does not represent any limitation as to the environments in which different embodiments can be implemented. Those skilled in the art can make many modifications to the illustrated environment without departing from the scope of the present invention as recited in the claims.
[0011] The present invention includes a database 144 and a sensor 126, which may include personal data, content, or information that the user desires not to be processed. 1-Ncan include various accessible data sources such as etc. Personal data includes information that identifies an individual, or confidential personal information, as well as user information such as tracking information or geographical location information. Processing refers to any automatic or non-automatic action or set of actions performed on personal data, including collection, recording, organization, structuring, storage, adaptation, alteration, retrieval, reference, use, disclosure by transmission, dissemination, or otherwise making available, combination, restriction, erasure, or destruction. Response program 200 enables the approved and secure processing of personal data. Response program 200 notifies and provides an explanation and consent for the collection of personal data, enabling the user to opt-in or opt-out of the processing of personal data. Consent can take several forms. Opt-in consent can impose on the user the requirement to take a positive action before personal data is processed. Alternatively, opt-out consent can impose on the user the requirement to take a positive action to prevent the processing of personal data before personal data is processed. Response program 200 provides information regarding the nature of the personal data and the processing (e.g., type, scope, purpose, duration, etc.). Response program 200 provides the user with a copy of the stored personal data. Response program 200 enables the correction or completion of inaccurate or incomplete personal data. Response program 200 enables the immediate deletion of personal data.
[0012] The distributed data processing environment 100 includes a server 140 and a client device 120 1 ~client device 120 Nincluding these, all of which are interconnected through network 110. Network 110 can be, for example, a telecommunications network, a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN) such as the Internet, or a combination of these three, and can include wired, wireless, or optical fiber connections. Network 110 can include one or more wired or wireless or both networks capable of receiving and transmitting data, voice, or video signals, or a combination thereof, including multimedia signals including voice, data, and video information. Generally, network 110 is for communication between servers 140, client devices 120 1-N in a distributed data processing environment 100 and other computing devices (not shown), and can be any combination of connections and protocols.
[0013] Client device 120 1-N (that is, client device 120 1 ~client device 120 N ) can be one or more of a laptop computer, a tablet computer, a smart phone, a smart watch, a smart speaker, a virtual assistant, an Internet of Things (IoT)-enabled device, or any programmable electronic device capable of communicating with various components and devices in a distributed data processing environment 100 via network 110. Generally, client device 120 1-N represents one or more programmable electronic devices or a combination of programmable electronic devices capable of executing machine-readable program instructions and communicating with other computing devices (not shown) in a distributed data processing environment 100 via a network such as network 110. Client device 120 1-N can include components that will be illustrated and described in more detail with respect to FIG. 3 according to an embodiment of the present invention.
[0014] Client device 120 1-N includes a user interface 122 1-N an application 124 1-N and a sensor 126 1-N and each instance can include an instance of each, and each instance corresponds to each instance of the client device and performs equivalent functions in each instance of the client device. In various embodiments of the present invention, the user interface is a program that provides an interface between the user of the device and a plurality of applications present in the client device. User interface 122 1 A user interface such as 122 refers to information (such as graphics, text, and sound) presented by the program to the user, as well as control sequences used by the user to control the program. There are various types of user interfaces. In one embodiment, user interface 122 1 is a graphical user interface. A graphical user interface (GUI) is a type of user interface that allows a user to interact with an electronic device such as a computer keyboard and mouse through graphical icons and visual indicators such as secondary notation, as opposed to a text-based interface, typed command labels, or text navigation. In computing, the GUI was introduced in response to the recognition of the steep learning curve of the command-line interface where commands had to be typed on the keyboard. Operations in the GUI are often performed through direct manipulation of graphical elements. In another embodiment, user interface 122 1 is a script or application programming interface (API). In one embodiment, user interface 122 1is a voice user interface. A voice user interface (VUI) enables spoken human interaction with computers using speech recognition to understand spoken commands and answer questions, and generally text to speech to play back responses.
[0015] Application 124 1 is a computer program designed to be executed on a client device 120 1 An application often functions to provide users with services similar to those accessed on a personal computer (such as a web browser, music playback, an email program, or other media). In one embodiment, Application 124 1 is mobile application software. For example, mobile application software, i.e., an "app", is a computer program designed to be executed on a smartphone, tablet computer, and other mobile devices. In another embodiment, Application 124 1 is a web user interface (WUI) that can display text, documents, web browser windows, user options, application interfaces, and operation instructions, and can include information (such as graphics, text, and sound) presented by the program to the user, as well as control sequences used by the user to control the program. In another embodiment, Application 124 1 is a client-side application of the response program 200.
[0016] Sensor 126 1is a device, module, machine, or subsystem that detects events or changes in an operating environment and often transmits that information to other electronic devices, often computer processors. Generally, sensor 126 1 represents various sensors of client device 120 1 that collect and provide various types of data (e.g., sound, images, movement, etc.). In one embodiment, client device 120 1 transmits data of sensor 126 1 to server 140 via network 110. For example, sensor 126 1 may be a camera for capturing an image of an environment including a user by client device 120 1 and that image is transmitted to a remote server (e.g., server 140).
[0017] In various embodiments of the present invention, server 140 may be a desktop computer, a computer server, or any other computer system known in the art. Generally, server 140 represents any electronic device or combination of electronic devices capable of executing computer-readable program instructions. Server 140 can include components that are illustrated and described in more detail with respect to FIG. 3 according to an embodiment of the present invention.
[0018] Server 140 may be a stand-alone computing device, a management server, a web server, a mobile computing device, or any other electronic device or computing system capable of receiving, transmitting, and processing data. In one embodiment, server 140 can represent a server computing system that uses multiple computers as a server system in a cloud computing environment or the like. In another embodiment, server 140 is a laptop computer, a tablet computer, a netbook computer, a personal computer (PC), a desktop computer, a personal digital assistant (PDA), a smart phone, or a client device 120 within a distributed data processing environment 100 1-N and any programmable electronic device that can communicate with other computer devices (not shown). In another embodiment, server 140 represents a computing system that uses clustered computers and components (e.g., database server computers, application server computers, etc.) that operate as a single pool of seamless resources when accessed within distributed data processing environment 100.
[0019] Server 140 includes a storage device 142, a database 144, and a response program 200. Storage device 142 is any type of storage device, e.g., client device 120 1-NAnd a persistent storage 305, a hard disk drive, or a flash memory that can store data accessed and used by a server 140 such as a database server. In one embodiment, the storage device 142 can represent a plurality of storage devices within the server 140. In various embodiments of the present invention, the storage device 142 stores many types of data that may include a database 144. The database 144 can represent one or more organized collections of data stored and accessed from the server 140. For example, the database 144 includes user activities, activity steps, voice commands, historical activity data, and the like. In one embodiment, the data processing environment 100 can include an additional server (not shown) that hosts additional information accessible via the network 110.
[0020] Generally, the response program 200 may enable a user to perform activities based on a synchronized voice response from the voice assistant. In one embodiment, the response program 200 identifies a user's voice query and identifies the user's past activities related to the voice query from a plurality of data sources (e.g., the client device 120 1-N ). In addition, the response program 200 generates a sequence of activities (e.g., a set of steps) corresponding to the task of the voice query and estimates the time taken for each of the identified past activities. The response program 200 uses the estimated values or data or both of the IoT devices (e.g., the client device 120 1-N ) to determine whether the user has performed one or more steps of the voice query task. Also, the response program 200 verifies whether the user has completed past activities and is predicting the user's future activities. Further, the response program 200 identifies one or more stages in the sequence of activities where an additional voice response is appropriate and synchronizes the additional voice response with the user's identified or future activities.
[0021] Figure 2 is a flowchart showing the operation steps of a response program 200 according to an embodiment of the present invention, that is, a program for synchronizing a voice response from an AI voice assistant based on user activities. In one embodiment, the response program 200 starts in response to a user connecting a client device 120 1 to the response program 200 through the network 110. For example, the response program 200 starts in response to a user registering (e.g., opting in) a virtual assistant (e.g., client device 120 1 ) to the response program 200 via a WLAN (e.g., network 110). In another embodiment, the response program 200 is a background application that continuously monitors the client device 120 1 . For example, the response program 200 is a client-side application (e.g., application 124) that starts when a user's virtual assistant (e.g., client device 120 1 ) is launched and monitors the virtual assistant for received voice commands.
[0022] In step 202, the response program 200 identifies one or more activities of the user corresponding to the user's voice query. In various embodiments of the present invention, the response program 200 uses the client device 120 1-N to collect the user's historical data and generate a collection of data in a database 144 corresponding to the user. For example, the response program 200 captures the user's activity data (e.g., body movement, biometric data, etc.) from wearable devices and IoT devices (e.g., client device 120 1-N ) to determine the activities performed by the user. In addition, the response program 200 can upload a user-specific activity log including the identified past activities (e.g., activity data) to generate a corpus (e.g., database 144) corresponding to the user.
[0023] In one embodiment, the response program 200 identifies a voice command from a user's client device 120 1 . For example, the response program 200 uses natural language processing (NLP) techniques (e.g., lexical semantics, relational semantics, term extraction, etc.) to determine the topic of the user's voice query received by the virtual assistant (e.g., client device 120 1 ). In addition, the response program 200 determines a task corresponding to the user's voice query.
[0024] In another embodiment, the response program 200 determines whether one or more activities of the user (e.g., past or currently executing) correspond to the user's voice command. For example, the response program 200 uses a corpus corresponding to the user (e.g., database 144) to identify one or more activities that the user has historically performed related to a task corresponding to the topic of the user's voice query. In this example, the response program 200 uses IoT devices (e.g., wearable devices, mobile devices, client device 120 2-N etc.) within the user's operating environment to track the user's movement behavior, gestures, and biometric data to determine the activity the user is performing. The collection of activity data (e.g., personal data) is an opt-in service, and the user gives permission for the collection, processing, and use of the activity data. In addition, the response program 200 determines whether the task of the voice query corresponds to the user's current activity. Also, the response program 200 can use the user's biometric data (e.g., voiceprint) to determine whether one or more additional voice queries are related to the user's voice query.
[0025] In an alternative example, the response program 200 is an IoT device (client device 120 2-NUsing the activity data in (), it is determined whether the voice query is related to the user's activity. Also, in a situation where the activity data of the IoT device cannot be accessed due to privacy issues or availability, the response program 200 can use a question and answer framework (for example, an AI-enabled chatbot, dialogue management, etc.) to obtain activity information or confirmation information from the user.
[0026] In step 204, the response program 200 generates a sequence of activities corresponding to the user's voice query. In one embodiment, the response program 200 determines one or more activities corresponding to the user's voice command. For example, the response program 200 uses a corpus of the user's historical data (for example, database 144) to determine the task of the user's voice query. In this example, the response program 200 identifies one or more activities in the corpus corresponding to the execution of previous tasks by the user and generates a list of sub-activities (for example, steps) required to complete the task. In an alternative, the response program 200 accesses the Internet to data corresponding to the topic and task of the voice query and identifies one or more activities (for example, steps) required to complete the task of the voice query.
[0027] In step 206, the response program 200 estimates the execution time of each step of the sequence of activities corresponding to the voice query. In one embodiment, the response program 200 estimates the time used by the user to perform one or more activities to complete the task corresponding to the voice command. For example, the response program 200 uses the corpus of the user's historical data (e.g., database 144) to estimate the execution time of each sub-activity (e.g., step) required to complete the task of the user's voice query. In this example, the response program 200 estimates the execution time of each sub-activity of the task using the previous execution time of the activity corresponding to the sub-activity of the task. In addition, the response program 200 can access from the Internet a remote server having all the timing parameters corresponding to the task of the voice query, and estimate the execution time of each sub-activity of the task with respect to the number of sub-activities of the task.
[0028] In decision step 208, the response program 200 determines whether the user is performing one or more steps of the sequence of activities. In one embodiment, the response program 200 uses the activity data of the client device 120 1-N to determine whether the user is performing one or more activities corresponding to the task of the user's voice command. For example, the response program 200 uses IoT devices (e.g., wearable devices, mobile devices, client device 120) within the user's operating environment 1-NUsing, for example, sensors (such as accelerometers, gyroscopes, etc.), the user's activity data (such as movement behavior, gestures, biometric data, etc.) is tracked to determine the user's current activity. In this example, the response program 200 uses the corpus corresponding to the user (such as the database 144) and the activity data to determine whether the determined user activity corresponds to one or more sub-activities of the task of the user's voice query. Alternatively, the response program 200 can use a question-and-answer framework (such as an AI-enabled chatbot, dialogue management, etc.) to determine whether the user is performing one or more sub-activities of the task of the user's voice query.
[0029] In another embodiment, if the response program 200 determines that the user is not performing one or more activities corresponding to the task of the user's voice command (the "no" branch of decision step 208), the response program 200 monitors the activity data of the client device 120 1-N to determine whether one or more activities of the user correspond to one or more activities corresponding to the task of the user's voice command. For example, if the response program 200 determines that the user's current activity is not related to one or more sub-activities (such as those generated in step 204) of the task of the user's voice query, the response program 200, as described above in step 202, uses IoT devices (such as wearable devices, mobile devices, client device 120 1-N etc.) within the user's operating environment to track the user's movement behavior, gestures, and biometric data.
[0030] In another embodiment, when the response program 200 determines that the user is performing one or more activities corresponding to the task of the user's voice command (the "yes" branch in decision step 208), the response program 200 verifies the occurrence of one or more activities to complete the task corresponding to the voice command, as described in step 210. For example, when the response program 200 determines that the user's current activity corresponds to one or more sub-activities of the task of the user's voice query, the response program 200 checks each of the one or more sub-activities of the voice query task performed by the user.
[0031] In step 210, the response program 200 verifies the completion of each step of the activity sequence. In one embodiment, the response program 200 is the client device 120 1 to verify the completion of each step of one or more activities to complete the task corresponding to the voice command. For example, the response program 200 determines whether the user has completed one or more sub-activities of the voice query task received by the virtual assistant (e.g., the client device 120 1 ). In this example, the response program 200 is an IoT device (e.g., the client device 120 2-NUsing the data in ( ), it is identified that the user is performing each of one or more sub - activities of the voice query task. In addition, the response program 200 determines the stage of completion of the task based on the identified one or more sub - activities of the task. Also, the response program 200 can identify the sub - activities of the task that the user has omitted or has not yet performed using the identified one or more sub - activities of the task. Further, the response program 200 determines whether the user is performing the sub - activities of the task (i.e., determines whether the step is not completed) and can generate a voice response to assist the user in completing the uncompleted steps. In addition, the response program 200 can also determine whether the user has completed the sub - activity using the estimated time of each sub - activity as described in step 206.
[0032] In another embodiment, the response program 200 uses the client device 120 1 to verify whether the user is performing each step of one or more activities to complete the task corresponding to the voice command of the client device 120 1 For example, if the response program 200 determines that it cannot access the activity data of the IoT device due to privacy issues or availability, the response program 200 uses a question - answering framework (e.g., NLP, chatbot, dialogue management, etc.) to obtain the status (e.g., completed, in progress, not executed) of each of one or more sub - activities of the voice query task from the user.
[0033] In step 212, the response program 200 identifies one or more stages of the sequence of activities for which one or more voice responses to the voice query are appropriate. In one embodiment, the response program 200 sends one or more voice responses for completing the task corresponding to the voice command to the client device 120 1Identify one or more stages of one or more sub - activities that are appropriate for transmission. For example, the response program 200 identifies one or more sub - activities (e.g., unexecuted sub - activities) of the task of a voice query from a user for which one or more voice responses are appropriate, and the content of one or more responses. In this example, the response program 200 generates a voice response for one or more future sub - activities of one or more sub - activities that the user may perform so that the response program 200 can intervene based on historical data or biometric data before the user performs one or more future sub - activities (e.g., incorrect / harmful steps).
[0034] In another embodiment, the response program 200 uses the activity data of the client device 120 1-N to determine whether the user desires to correlate a voice response with one or more stages of one or more activities to complete a task corresponding to a voice command. In various embodiments of the present invention, the response program 200 uses the user's historical activity data to identify whether the user needs assistance regarding steps of a task corresponding to a voice command. For example, the response program 200 tracks the user's activity patterns as well as related behaviors and biometric data (e.g., blood pressure, heart rate, stress level, etc.) to determine the user's comfort level regarding sub - activities of the task of the user's voice query. In this example, the response program 200 can use the user's comfort level to identify one sub - activity among one or more sub - activities that the user has not yet performed and that corresponds to a decrease in the user's comfort level.
[0035] In step 214, the response program 200 synchronizes one or more voice responses with the user's activities. In one embodiment, the response program 200 correlates the voice responses with one or more stages of the activity to complete the task corresponding to the voice command. For example, the response program 200 receives an IoT feed from one or more IoT devices (e.g., client device 120 1-N ) that includes activity data and determines whether to ignore one or more sub-activities of the voice query task. In this example, the response program 200 determines whether the user has completed one of the one or more sub-activities based on the activity data collected before the transmission of the voice query. In addition, the response program 200 identifies which voice responses of the one or more sub-activities should be provided to the user based on the completed activity. As a result, the response program 200 does not repeat or provide voice responses corresponding to activities previously performed by the user.
[0036] In another example, the response program 200 uses a history corpus corresponding to the user (e.g., database 144) to determine the reason (e.g., next step, verification information, additional information request, etc.) for which the user sends a voice query or related follow-up voice command corresponding to the task. In this example, the response program 200 identifies a set of conditions indicating that the user is feeling stressed or facing difficulties in performing the sub-activities corresponding to the task. In addition, the response program 200 correlates this set of conditions with the fact that the user is sending the voice query to the virtual assistant (e.g., client device 120 1 ) to determine the reason for which the user is sending the voice query.
[0037] In another embodiment, the response program 200 is the client device 120 1-NBy using it to monitor the user's activities, additional voice responses are generated, and the additional voice responses and the user's physical activities are synchronized based on the user's activity data. For example, while sending a voice response, the response program 200 uses the user's behavior (e.g., activity data) and IoT-enabled devices (e.g., client device 120 1-N ) in the user's operating environment to predict the speed at which the user may complete one or more sub-activities of a task corresponding to the user's voice query, and dynamically synchronize future voice responses with the user's physical activities. Additionally, the response program 200 can modify the transmission speed of the voice response based on the predicted speed. In this example, the response program 200 uses the user's behavior and IoT data to predict the user's future activities (e.g., mistakes, activity patterns, etc.) based on historical activity data, and can generate new voice responses corresponding to the future activities (i.e., the response program 200 understands that the user is wrong or may make a mistake, and generates correction steps and additional voice responses synchronized with the current sub-activities the user is performing).
[0038] In another example, the response program 200 uses the question-and-answer framework (e.g., NLP, chatbot, dialogue management, etc.) of a virtual assistant (e.g., client device 120 1 ) to determine the reason why the user sends a voice query corresponding to a task or a related follow-up voice query. In this example, the response program 200 provides an audible prompt to the user to provide the reason for a follow-up voice query or completion status of one or more sub-activities of the voice query task, and synchronizes the voice response with the user's activities regarding completing the task (i.e., correlates the voice response with the current stage of the user's activity / step).
[0039] In various embodiments of the present invention, the response program 200 can provide a voice response to the user through various communication channels based on the activity type. For example, the response program 200 can be embedded in a display system (e.g., a computing device) to provide a graphical display of one or more sub-activities of a complex activity, and assist the user in understanding the content of the voice response to the voice command executed by the user through a visual explanation. In another embodiment, the response program 200 determines a method of transmitting a voice response corresponding to one or more stages of one or more activities for completing a task corresponding to a voice command. For example, the response program 200 uses IoT data of the user's wearable computing device (e.g., the client device 120 2 ) to determine the time taken for the user to complete one or more sub-activities of the task of the voice query, and determines the length of time for which the voice response corresponding to the sub-activity of the task should continue.
[0040] In another example, the response program 200 determines the command type (e.g., shell, social, web) of the user's voice query (e.g., voice command) from the metadata of the voice query. A shell command can be related to a directory-based system and can be used to identify the location (e.g., path) in the directories of any file, folder, and application on the computing device. A social command can be related to a request-response system (e.g., an interactive system) and can be used for "what" type questions. A web command can be related to a web-based command system and can be used to access a Uniform Resource Locator (URL) using the default web browser (e.g., application 124). In this example, the response program 200 uses the command type to determine the communication channel for providing a voice response to the user. In one situation, if the response program 200 determines that the voice query is a social command and the user has confirmed that one or more sub-activities of the task have been previously executed, the response program 200 can use natural language generation to generate a summary (e.g., text, audio, multimedia, etc.) of the previously executed sub-activities and send a voice response including the summarized sub-activities. Alternatively, if the response program 200 determines that one or more sub-activities of the task have been previously executed by the user, the response program 200 can ignore the voice response corresponding to the previously executed sub-activities and send a voice response from the current stage of the activity being performed by the user.
[0041] In another example, the response program 200 is an IoT device (e.g., client device 120 1-N) Use the activity data (e.g., engagement level, distraction level, etc.) to determine the number of hours to play the voice response (e.g., continuous reminder). In this example, the response program 200 uses the verification status in step 210 and the corpus of the user's past executions (e.g., database 144) to predict one or more sub-activities of the task that the user may execute incorrectly.
[0042] In decision step 216, the response program 200 determines whether the sequence of activities is complete. In one embodiment, the response program 200 uses the activity data of the client device 120 1-N to determine whether the user has completed each of one or more activities corresponding to the task of the user's voice command. For example, the response program 200 uses IoT devices (e.g., wearable devices, mobile devices, client device 120 1-N etc.) in the user's operating environment to track the user's activity data (e.g., movement behavior, gestures, biometric data, etc.) and determine whether the user has completed each of one or more sub-activities of the task of the user's voice query. In this example, the response program 200 can use the verification status of each of the one or more sub-activities described in step 210 to determine whether the user has completed the task.
[0043] In another embodiment, when the response program 200 determines that the user has not completed one or more activities corresponding to the task of the user's voice command (the "no" branch of decision step 216), the response program 200 correlates the voice response with the stage of one or more activities to complete the task corresponding to the voice command. For example, when the response program 200 determines that the verification status of one or more sub-activities of the task of the user's voice query is "not completed", the response program 200 resynchronizes the voice response, provides a status update question, or infers the status update from the IoT or wearable feed (e.g., client device 120 1-N ), and can send the voice response to the user from the stage where the user cannot perform the sub-activity of the task. In addition, the response program 200 can continuously monitor the user's physical activity to estimate the speed at which the user performs one or more sub-activities of the task and determine whether the user is likely to make a mistake. As a result, the response program 200 identifies various channels to provide repeated voice responses to the user. Alternatively, the response program 200 can continue to provide voice responses in synchronization with the user's activities.
[0044] In another embodiment, when the response program 200 determines that the user has completed each of one or more activities corresponding to the task of the user's voice command (the "yes" branch of decision step 216), the response program 200 collects the activity data corresponding to the user from the client device 120 1-N and updates the database 144. For example, when the response program 200 determines that the verification status of each of one or more sub-activities of the task of the user's voice query is "completed", the response program 200 logs the user's activities for future reference as described below in step 218.
[0045] In step 218, the response program 200 updates the user's activity log. In one embodiment, the response program 200 collects activity data corresponding to the user from the client device 120 1-N and updates the database 144. For example, when the task of the voice query by the user is completed, the response program 200 sends a partial update of the activity log (for example, the database 144) to enable self-improvement via automatic or manual or both feedback. In this example, the response program 200 collects the context of the voice query (for example, time, location, surrounding activities, etc.) and the specific patterns identified while the user is performing the sub-activities corresponding to the task. In addition, the response program 200 logs a summary of the user's physical activities and saves and updates the activity log for future reference.
[0046] FIG. 3 is a block diagram of the components of the client device 120 1-N and the server 140 according to an exemplary embodiment of the present invention. It should be understood that FIG. 3 is merely an illustration of one implementation form and does not represent any limitation regarding the environment in which different embodiments can be implemented. Many modifications can be made to the illustrated environment.
[0047] FIG. 3 includes a processor 301, a cache 303, a memory 302, a persistent storage 305, a communication unit 307, an input / output (I / O) interface 306, and a communication fabric 304. The communication fabric 304 provides communication between the cache 303, the memory 302, the persistent storage 305, the communication unit 307, and the input / output (I / O) interface 306. The communication fabric 304 may be implemented by any architecture designed to pass data or control information or both between a processor (such as a microprocessor, a communication and network processor, etc.), a system memory, peripheral devices, and any other hardware component within the system. For example, the communication fabric 304 may be implemented by one or more buses or a crossbar switch.
[0048] Memory 302 and persistent storage 305 are computer-readable storage media. In this embodiment, memory 302 includes random access memory (RAM). In general, memory 302 can include any suitable volatile or non-volatile computer-readable storage media. Cache 303 is a high-speed memory that improves the performance of processor 301 by holding recently accessed data from memory 302 and data near the recently accessed data.
[0049] The program instructions and data (e.g., software and data 310) used to implement embodiments of the present invention may be stored in persistent storage 305 and memory 302 for execution by one or more of the respective processors 301 via cache 303. In an embodiment, persistent storage 305 includes a magnetic hard disk drive. Instead of or in addition to a magnetic hard disk drive, persistent storage 305 can include a solid state hard drive, a semiconductor storage device, a read only memory (ROM), an erasable programmable read only memory (EPROM), a flash memory, or any other computer-readable storage media capable of storing program instructions or digital information.
[0050] The media used by persistent storage 305 may be removable. For example, a removable hard drive may be used for persistent storage 305. Other examples include optical and magnetic disks, thumb drives, and smart cards that are inserted into a drive for transfer to another computer-readable storage media that is also part of persistent storage 305. Software and data 310 may be stored in persistent storage 305 for access or execution or both by one or more of the respective processors 301 via cache 303. Client device 120 1-NRegarding the software and data 310, they include the data of the user interface 122 1-N , the application 124 1-N , and the sensor 126 1-N . Regarding the server 140, the software and data 310 include the data of the storage device 142 and the response program 200.
[0051] The communication unit 307 in these examples provides communication with other data processing systems or devices. In these examples, the communication unit 307 includes one or more network interface cards. The communication unit 307 can provide communication by using either or both of a physical communication link and a wireless communication link. The program instructions and data (e.g., the software and data 310) used to implement the embodiments of the present invention may be downloaded to the persistent storage 305 through the communication unit 307.
[0052] The I / O interface 306 enables the input and output of data by other devices connected to each computer system. For example, the I / O interface 306 can provide a connection to an external device 308 such as a keyboard, keypad, touch screen, or other suitable input device, or a combination thereof. The external device 308 may include, for example, a portable computer-readable storage medium such as a thumb drive, portable optical disk or magnetic disk, and memory card. The program instructions and data (e.g., the software and data 310) used to implement the embodiments of the present invention may be stored in such a portable computer-readable storage medium and loaded into the persistent storage 305 via the I / O interface 306. The I / O interface 306 is also connected to the display 309.
[0053] The display 309 provides a mechanism for displaying data to the user and may be, for example, a computer monitor.
[0054] The programs described in this specification are identified based on the applications in which they are implemented in particular embodiments of the present invention. However, since the naming of any particular program in this specification is used merely for convenience, it should be understood that the present invention should not be limited to use only in any particular application in which such naming is used for identification or implication or both.
[0055] The present invention may be a system, method, or computer program, or a combination thereof, at any possible integrated technical detail level. The computer program may include a computer-readable storage medium having computer-readable program instructions for causing a processor to execute aspects of the present invention.
[0056] A computer-readable storage medium may be a tangible device that holds and stores instructions for use by an instruction execution device. The computer-readable storage medium may be, for example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof, but is not limited thereto. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy (R) disk, mechanically encoded devices such as punch cards or raised structures in grooves in which instructions are recorded, and any suitable combination thereof. As used herein, a computer-readable storage medium should not be construed to be a transitory signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.
[0057] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or can be downloaded from an external computer or an external storage device via a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network can include a copper transmission cable, an optical transmission fiber, a wireless transmission, a router, a firewall, a switch, a gateway computer, or an edge server, or a combination thereof. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers those computer-readable program instructions for storage in a computer-readable storage medium within each respective computing / processing device.
[0058] The computer-readable program instructions for carrying out the operations of the present invention may be source code or object code written in any combination of one or more programming languages, including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or any combination of object-oriented programming languages such as Smalltalk(R), C++, and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter case, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, an electronic circuit, including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), can execute the computer-readable program instructions by personalizing the electronic circuit using the state information of the computer-readable program instructions.
[0059] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer programs according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0060] These computer-readable program instructions, when executed via the processor of a computer or other programmable data processing apparatus, are provided to the processor of the general purpose computer, special purpose computer, or other programmable data processing apparatus to create means for implementing the functions / acts specified in one or more blocks of the flowchart, block diagram, or both, and thereby create a machine. These computer-readable program instructions may be stored in a computer-readable storage medium that includes a product that contains instructions for implementing the aspects of the functions / acts specified in one or more blocks of the flowchart, block diagram, or both, and that can direct a computer, programmable data processing apparatus, or other device, or combinations thereof, to function in a particular manner.
[0061] The computer-readable program instructions may be loaded onto a computer, other programmable apparatus, or other device to produce a process that is executed by the computer to cause the instructions executed on the computer, other programmable apparatus, or other device to implement the functions / acts specified in one or more blocks of the flowchart, block diagram, or both, and thereby perform a series of operational steps on the computer, other programmable apparatus, or other device.
[0062] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer programs according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, segment, or portion of instructions that includes one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks can be performed in an order different than that illustrated. For example, two blocks shown in succession can in fact be executed substantially simultaneously, or the blocks can sometimes be executed in the reverse order depending on the functions involved. It should also be noted that each block of the block diagrams or flowchart diagrams, or combinations of blocks in the block diagrams or flowchart diagrams or both, can be implemented by a dedicated hardware-based system that performs the specified function or operation, or a combination of dedicated hardware instructions and computer instructions.
[0063] The descriptions of the various embodiments of the present invention have been presented for purposes of illustration, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the present invention. The terms used herein were chosen in order to best explain the principles of the embodiments, the practical application, or technical improvement in the marketplace, or to enable others of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A method for synchronizing a voice response of a virtual assistant by information processing of a computer having one or more processors, the processor comprising: identifying a task of the user corresponding to the user's voice query; generating a sequence of sub-activities of the task corresponding to the voice query of the user; determining a completion status of each sub-activity of the sequence of sub-activities of the task responsive to the voice query of the user based at least in part on activity data received from one or more computing devices within the operating environment of the user; synchronizing a voice response corresponding to the voice query of the user with the sequence of sub-activities of the task based at least in part on the completion status of each sub-activity in the sequence of sub-activities of the task; using the activity data to determine whether to ignore one or more sub-activities of the task of the voice query, determining whether the user has completed one of the one or more sub-activities based on activity data collected prior to sending a voice query, identifying which voice response of the one or more sub-activities should be provided to the user based on the completed activity, and determining whether to ignore the one or more sub-activities as a result of the identification without repeating or providing a voice response corresponding to an activity previously performed by the user. How to do it.
2. The method of claim 1 , further comprising determining a respective execution time for each sub-activity of the sequence of sub-activities of the task corresponding to the voice query of the user.
3. 2. The method of claim 1 , further comprising determining one or more stages of the sequence of sub-activities of the task that correspond to one or more voice responses of the computing device, each stage of the one or more stages corresponding to one or more sub-activities of the sequence of sub-activities of the task.
4. collecting a context for the voice query; collecting execution patterns of the user corresponding to sequences of sub-activities of the task corresponding to the voice queries of the user; updating an activity log corresponding to the user and the voice query, the activity log including the context of the voice query and the execution patterns of the user; The method of claim 1 further comprising:
5. determining that the activity data of the one or more computing devices in the operating environment of the user is inaccessible; and determining a completion status of each sub-activity of the sequence of sub-activities of the task responsive to the voice query of the user based at least in part on a question and answer framework; determining whether the user has completed the sequence of sub-activities of the task corresponding to the voice query of the user; The method of claim 1 further comprising:
6. collecting the activity data for one or more computing devices within the operating environment of the user; identifying a sub-activity of the sequence of sub-activities of the task corresponding to the voice query of the user based at least in part on the activity data and a corpus of historical data corresponding to the user; The method of claim 1 further comprising:
7. The method of claim 6, further comprising: synchronizing the voice response corresponding to the voice query of the user with the sequence of sub-activities of the task based at least in part on the completion status of each sub-activity in the sequence of sub-activities of the task. determining a stage of a sequence of sub-activities of the task that corresponds to a sub-activity of the task that the user is performing; providing a response to the voice query based at least in part on the sub-activity of the task that the user is performing; and The method of claim 1 further comprising:
8. A computer program product for causing a computer to execute the method according to any one of claims 1 to 7.
9. 9. A computer readable storage medium having the computer program of claim 8 stored therein.
10. A computer system for synchronizing voice responses of a virtual assistant, comprising: one or more computer processors; one or more computer readable storage media; program instructions stored on the computer-readable storage medium for execution by at least one of the one or more processors; wherein the program instructions include: program instructions for identifying a task of a user corresponding to a voice query of the user; program instructions for generating a sequence of sub-activities of the task responsive to the voice query of the user; program instructions for determining a completion status of each sub-activity of the sequence of sub-activities of the task responsive to the voice query of the user based at least in part on activity data received from one or more computing devices within the operating environment of the user; program instructions for synchronizing a voice response corresponding to the voice query of the user with the sequence of sub-activities of the task based at least in part on the completion status of each sub-activity in the sequence of sub-activities of the task; and program instructions for using the activity data to determine whether to ignore one or more sub-activities of the task of the voice query, the program instructions including: determining whether the user has completed one of the one or more sub-activities based on activity data collected prior to sending a voice query; identifying which voice response of the one or more sub-activities should be provided to the user based on the completed activity; and determining whether to ignore the one or more sub-activities as a result of the identification, without repeating or providing a voice response corresponding to an activity previously performed by the user.
1. A computer system comprising:
11. 11. The computer system of claim 10, further comprising program instructions stored on the one or more computer readable storage media for execution by at least one of the one or more processors for determining a respective execution time for each sub-activity in the sequence of sub-activities of the task responsive to the voice query of the user.
12. 11. The computer system of claim 10, further comprising: program instructions stored on the one or more computer-readable storage media for execution by at least one of the one or more processors for determining one or more stages of the sequence of sub-activities of the task that correspond to one or more voice responses of the computing device, each stage of the one or more stages corresponding to one or more sub-activities of the sequence of sub-activities of the task.
13. stored on the one or more computer-readable storage media for execution by at least one of the one or more processors; program instructions for collecting context for the voice query; program instructions for collecting execution patterns of the user corresponding to the sequences of sub-activities of the task corresponding to the voice queries of the user; program instructions for updating an activity log corresponding to the user and the voice query, the activity log including the context of the voice query and the execution patterns of the user; 11. The computer system of claim 10, further comprising:
14. stored on the one or more computer-readable storage media for execution by at least one of the one or more processors; program instructions for determining that the activity data of the one or more computing devices in the operating environment of the user is inaccessible; program instructions for determining a completion status of each sub-activity of the sequence of sub-activities of the task responsive to the voice query of the user based at least in part on a question and answer framework; program instructions for determining whether the user has completed the sequence of sub-activities of the task corresponding to the voice query of the user; 11. The computer system of claim 10, further comprising:
15. A method for synchronizing a voice response corresponding to the voice query of the user with the sequence of sub-activities of the task based at least in part on the completion status of each sub-activity in the sequence of sub-activities of the task, the method comprising: program instructions for determining a stage of a sequence of the sub-activity of the task that corresponds to a sub-activity of the task that the user is performing; program instructions for providing a response to the voice query based at least in part on the sub-activity of the task that the user is performing; and 11. The computer system of claim 10, further comprising:
Citation Information
Patent Citations
Living task support system
JP2008015585A
Non-deterministic task initiation by the personal assistant module
JP2019523907A
Non-deterministic task initiation by the personal assistant module
KR1020190016552A
Information processing device, information processing method, and program
WO2020026799A1