System
A voice-command-based system converts and automates operations across different applications using generative AI, addressing the challenge of diverse UIs by enhancing user experience and efficiency.
Patent Information
- Application Number
- JP2024115265
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-18
- Publication Date
- 2026-01-29
AI Technical Summary
Modern users face challenges in efficiently operating multiple applications with different user interfaces (UIs), particularly when UIs change, leading to increased learning costs and reduced operational efficiency due to the need to remember unique functions and perform manual operations.
A system that converts voice commands into text, analyzes them using generative AI to generate specific operating procedures, and automates these procedures across different applications, allowing seamless operation without manual intervention.
This system significantly reduces learning costs and stress associated with UI changes by enabling intuitive and efficient operation of various applications through voice commands, improving user experience and operational efficiency.
Smart Images

Figure 2026014268000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] It is common for modern users to use many different applications. Each application has a different user interface (UI), so users must learn how to use each new application, which takes time to get used to. In particular, when the UI changes due to an update, users must re-learn how to operate the application, which can be stressful. Furthermore, remembering different operating procedures for each application is cumbersome and hinders efficient operation.
[0005] Furthermore, many applications each have their own unique functions, making it difficult to operate them in a unified manner. It is particularly difficult to seamlessly operate these different applications using voice commands. This situation degrades the quality of the user experience (UX) and reduces operational efficiency. [Means for solving the problem]
[0006] In order to solve the above problems, the present invention provides the following means: First, a means for a user to input a voice command; Next, a means for acquiring the voice command and converting it into text; Then, a means for transmitting the text data to a server, which analyzes the text data and uses generative artificial intelligence to generate specific operating procedures.
[0007] It also includes a means for transmitting the generated operation procedure to the terminal, and the terminal has a means for actually performing the operation based on the operation procedure. This series of steps enables the user to seamlessly operate multiple different applications with a single voice command, significantly reducing the learning cost and stress caused by differences in UI.
[0008] Furthermore, because the generative AI understands and generates natural language, it is possible for users to instruct operations in natural language that they normally use. Furthermore, because the generated operation procedures include automatic operation of the user interface, users no longer need to perform complex operations manually, resulting in improved operation efficiency. In this way, the present invention provides a system that allows users to operate different applications in a unified manner, thereby improving the user experience and operation efficiency.
[0009] A "voice command" is an input means by which a user gives instructions to a terminal via voice.
[0010] "Text data" is a data format in which a voice command is converted into a character string.
[0011] A "server" is a remote computer system that receives and processes requests from terminals via a network.
[0012] "Generative AI" is an AI technology that has the ability to analyze text data and generate specific operating procedures based on specified tasks.
[0013] An "operational procedure" is a series of steps in operating a user interface to perform a particular task.
[0014] A "terminal" is a computing device into which a user inputs voice commands and which performs operations based on operating instructions received from a server.
[0015] "Natural language understanding" is a technology that analyzes the natural language used by users and understands its meaning.
[0016] "Automatic operation" means that the terminal automatically operates the user interface without manual operation by the user.
[0017] "User interface" is a general term for the screens and operating means through which a user interacts with an application.
[0018] "User experience" refers to the overall satisfaction and ease of use that a user feels when using a system or application. [Brief explanation of the drawings]
[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4]FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0021] First, the terms used in the following description will be explained.
[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0027] [First embodiment]
[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0040] The present invention relates to a system that allows users to easily and intuitively operate multiple different applications using voice commands. This system automates the entire process from receiving a voice command to executing the operation, improving the user experience.
[0041] Overall system configuration
[0042] The system consists of the following main components:
[0043] 1. A device that receives voice commands
[0044] 2. A voice recognition engine within the device that converts voice commands into text data
[0045] 3. A means of communication to send the converted text data to the server
[0046] 4. Generative AI in the server to analyze text data and generate specific operating procedures
[0047] 5. Response method from the server that sends the generated operation procedure to the terminal
[0048] 6. An automated operation system for executing actual operations based on operating procedures at the terminal
[0049] Program processing
[0050] Acquiring and translating voice commands
[0051] The user inputs voice commands into the terminal. The terminal is equipped with a voice recognition engine that converts the user's voice into text data in real time. This text data is then sent to the server via the server communication means described below.
[0052] Sending and analyzing text data
[0053] When text data is sent from a device to a server, the server first receives the data. The server then analyzes the received text data and uses generative artificial intelligence to understand the user's intent. Specifically, it uses natural language understanding technology to analyze the content of the text and generate the necessary operating procedures.
[0054] Generate and send operating instructions
[0055] The generative AI then creates specific operational steps based on the analysis results. For example, if a command such as "Add a meeting tomorrow at 3 PM" is entered, the generative AI will analyze this and generate a series of operational steps in the calendar app, such as "Add a new event," "Set the date," "Set the time," and "Enter a title."
[0056] Execute operations on the device
[0057] The server generates a sequence of operations and sends it to the device, which then automatically executes the sequence of operations, eliminating the need for manual intervention by the user and allowing them to complete the desired task efficiently and intuitively.
[0058] Specific examples
[0059] Below is a specific example of adding an event to a calendar app.
[0060] 1. The user types the voice command "Add a meeting tomorrow at 3 PM."
[0061] 2. The device uses a voice recognition engine to convert this voice into text data.
[0062] 3. The device sends the converted text data to the server.
[0063] 4. The server receives the text data and uses generative artificial intelligence to generate operational steps such as "Open the calendar app," "Add a new event," "Set the date," "Set the time," and "Enter a title."
[0064] 5. The server sends the generated operation procedure to the terminal.
[0065] 6. The device automatically performs the necessary operations within the calendar app based on the received instructions and adds the meeting.
[0066] In this way, the present invention provides a system that allows users to seamlessly operate a variety of applications through voice commands, achieving an efficient and intuitive operating experience while maintaining a consistent UI.
[0067] The processing flow will be explained below.
[0068] Step 1:
[0069] The user enters a voice command into the device, for example, "Add a meeting for tomorrow at 3 PM."
[0070] Step 2:
[0071] The device activates its built-in voice recognition engine and captures voice data in real time.
[0072] Step 3:
[0073] The device converts the voice data into text data, and the speech recognition engine generates the text "Add a meeting tomorrow at 3 PM."
[0074] Step 4:
[0075] The terminal includes the converted text data in the body of the HTTP request and prepares to send it to the server.
[0076] Step 5:
[0077] The device sends text data to the server. An HTTP request is sent to the server.
[0078] Step 6:
[0079] The server receives the HTTP request and extracts the text data from the request body.
[0080] Step 7:
[0081] The server analyzes the text data using a Natural Language Understanding (NLU) engine, and extracts the following information as the analysis result: "Date = tomorrow", "Time = 3 PM", "Action = Add meeting".
[0082] Step 8:
[0083] The server calls the API of a generative artificial intelligence (e.g., GPT-3) and generates specific operating procedures based on the analysis results.
[0084] Step 9:
[0085] The server formats the generated operating instructions and creates an HTTP response to send to the terminal.
[0086] Step 10:
[0087] The server sends an HTTP response to the device.
[0088] Step 11:
[0089] The terminal receives the HTTP response and analyzes the operation procedure from the response.
[0090] Step 12:
[0091] The device launches the calendar app. Specifically, it calls the calendar app based on the application ID.
[0092] Step 13:
[0093] Based on the operating procedure, the device will automatically perform a series of operations within the calendar app, such as "add a new event," "set the date," "set the time," and "enter a title."
[0094] Step 14:
[0095] The terminal notifies the user that the operation is complete, for example by displaying a message such as "Conference added."
[0096] Example 1
[0097] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0098] Current voice-command-based operation systems have the problem of being difficult to seamlessly operate multiple different applications, limiting the user experience. Furthermore, manual operation is required, reducing efficiency and making intuitive operation difficult. Therefore, there is a need for a system that allows users to easily and intuitively operate multiple applications through voice commands.
[0099] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0100] In this invention, the server includes means for receiving a voice command, means for converting the voice command into text data, means for transmitting the text data to the server, means for analyzing the text data and using generative artificial intelligence to generate specific operating procedures, means for transmitting the operating procedures to the terminal, means for executing an operation on the terminal based on the operating procedures, means for generating specific operating procedures including an operation for adding an event to a calendar application, and means for transmitting the operating procedures from the server to the terminal and automatically executing them, thereby enabling efficient and intuitive operation of various applications through voice commands.
[0101] A "voice command" is an instruction or request that a user communicates to a system through speech.
[0102] "Text data" refers to data obtained by converting voice commands into text using a voice recognition engine.
[0103] A "server" is a computer system that receives and analyzes text data, and generates and transmits operating procedures.
[0104] "Generative AI" refers to an AI technology that analyzes text data and generates specific operating procedures based on the user's intentions.
[0105] An "operation procedure" refers to a series of operations performed within an application.
[0106] "Terminal" means a device that inputs voice commands, converts voice to text, and receives and executes operating instructions.
[0107] A "voice recognition engine" is software or hardware that converts voice commands into text data.
[0108] "Communication means" refers to the communication technology for transmitting text data from the terminal to the server and transmitting operating procedures from the server to the terminal.
[0109] "Natural language understanding" refers to the technology for analyzing text written in natural language and understanding its content and intent.
[0110] "Automatic operation of a user interface" means automatically operating an application without user assistance, according to a generated operating procedure.
[0111] The present invention relates to a system that allows users to easily and intuitively operate multiple different applications using voice commands. The system aims to improve the user experience by automating the entire process from receiving a voice command to executing the operation.
[0112] Overall system configuration
[0113] The system consists of the following main components:
[0114] 1. A device that receives voice commands
[0115] 2. A voice recognition engine within the device that converts voice commands into text data
[0116] 3. A means of communication to send the converted text data to the server
[0117] 4. Generative AI in the server to analyze text data and generate specific operating procedures
[0118] 5. Response method from the server that sends the generated operation procedure to the terminal
[0119] 6. An automated operation system for executing actual operations based on operating procedures at the terminal
[0120] System Operation
[0121] Acquiring voice commands
[0122] The user inputs voice commands into the terminal. The terminal uses a voice recognition engine (e.g., general voice recognition software) to convert the user's voice into text data in real time. This text data is sent to the server via the communication means described below.
[0123] Sending and analyzing text data
[0124] After the device converts the speech into text, it sends the converted text data to a server. The server immediately analyzes the received text data and uses generative artificial intelligence (e.g., a general generative AI model) to understand the user's intent. Specifically, it uses natural language understanding technology to analyze the content of the text and generate the necessary operating instructions.
[0125] Generate and execute operating procedures
[0126] The server creates specific operational steps based on the analysis results. For example, if a command such as "Add a meeting tomorrow at 3 PM" is entered, the generative AI analyzes this and generates a series of operational steps in a calendar app (e.g., general calendar software) such as "Add a new event," "Set the date," "Set the time," and "Enter a title."
[0127] Sending and executing operating instructions
[0128] The server generates and sends the operation instructions to the terminal, which then automatically executes a series of operations according to the instructions, eliminating the need for the user to perform manual operations and allowing the user to complete the desired task efficiently and intuitively.
[0129] Specific examples
[0130] Example of adding an event in the calendar app
[0131] 1. The user types the voice command "Add a meeting tomorrow at 3 PM."
[0132] 2. The device uses a voice recognition engine to convert this voice into text data.
[0133] 3. The device sends the converted text data to the server.
[0134] 4. The server receives the text data and uses generative artificial intelligence to generate operational steps such as "Open the calendar app," "Add a new event," "Set the date," "Set the time," and "Enter a title."
[0135] 5. The server sends the generated operation procedure to the terminal.
[0136] 6. The device automatically performs the necessary operations within the calendar app based on the received instructions and adds the meeting.
[0137] Example prompts for generative AI models
[0138] The voice command "Add a meeting tomorrow at 3 PM" has been converted into text data. Please analyze this text and generate instructions for the calendar app.
[0139] In this way, the present invention provides a system that allows users to seamlessly operate a variety of applications through voice commands, achieving an efficient and intuitive operating experience while maintaining a consistent UI.
[0140] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0141] Step 1:
[0142] The user inputs a voice command into the terminal.
[0143] What happens: A user says, "Add a meeting tomorrow at 3 PM."
[0144] Input: User's voice
[0145] Output: Audio data
[0146] Step 2:
[0147] The device uses its built-in voice recognition engine to convert the voice data into text data.
[0148] What it does: The speech recognition engine analyzes the user's speech and converts the audio waveform into text.
[0149] Input: Audio data
[0150] Output: Text data "Add a meeting tomorrow at 3 PM"
[0151] Step 3:
[0152] The terminal transmits the text data to the server.
[0153] Specific behavior: Sends text data to the server using an HTTP POST request.
[0154] Input: Text data
[0155] Output: Request message (including text data)
[0156] Step 4:
[0157] The server receives the text data and analyzes it using a natural language processing engine.
[0158] Specific operation: The server receives the transmitted text data and inputs it into the generative AI model.
[0159] Input: Request message
[0160] Output: Intention analysis results of text data
[0161] Step 5:
[0162] Generative artificial intelligence generates specific operating procedures based on the analysis results.
[0163] Specific actions: Based on the command "Add a meeting tomorrow at 3 PM," it generates operational steps such as "Open the calendar app," "Add a new event," "Set the date," "Set the time," and "Enter a title."
[0164] Input: Intention analysis results of text data
[0165] Output: Specific operating instructions
[0166] Step 6:
[0167] The server sends the generated operation procedure to the terminal.
[0168] Specific operation: Sends operation instructions to the terminal as an HTTP response.
[0169] Input: Specific operating instructions
[0170] Output: Response message (including operation instructions)
[0171] Step 7:
[0172] The device will automatically perform operations within the calendar app based on the operating instructions.
[0173] Specific operation: In the calendar app, automatically "add a new event," "set the date," "set the time," and "enter a title."
[0174] Input: Response message (operation procedure)
[0175] Output: New event added to the Calendar app
[0176] In this way, a system is constructed in which a user's voice command is converted into a specific action via the terminal and the server, and is then automatically executed.
[0177] (Application example 1)
[0178] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0179] In today's diverse applications, it is cumbersome for users to perform multiple operations individually, and it is particularly time-consuming to efficiently complete an order in food delivery services. To solve this problem, a system is needed that automates a series of operations through voice commands and allows users to complete an order intuitively and efficiently.
[0180] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0181] In this invention, the server includes means for receiving a voice command, means for converting the voice command into text data, means for transmitting the text data to the server, means for analyzing the text data and using generative artificial intelligence to generate specific operation procedures, means for transmitting the operation procedures to a terminal, means for executing an operation on the terminal based on the operation procedures, and means for automating an ordering process to provide a meal delivery service, thereby enabling a user to efficiently and intuitively complete a meal order using voice commands, thereby improving the user experience.
[0182] A "voice command" is a voice instruction given by a user to a terminal.
[0183] "Text data" is voice commands converted into text information.
[0184] A "server" is a computer system that receives text data, analyzes it, and generates specific operating procedures.
[0185] "Generative AI" is an algorithm that understands and generates natural language, generating specific operating procedures from text data.
[0186] A "terminal" is a device through which a user inputs voice commands and executes generated operating procedures.
[0187] A "meal delivery service" is a service in which a user orders a meal and the order is delivered to a location specified by the user.
[0188] "Automating the ordering process" refers to the process of completing an order in a food delivery service based on voice commands without the need for manual intervention.
[0189] This invention relates to a system that allows users to easily and intuitively order food delivery services using voice commands. This system automates a series of operations from receiving a voice command to completing an order, improving the user experience.
[0190] Overall system configuration
[0191] The system consists of the following main components:
[0192] 1. A device that receives voice commands
[0193] 2. A voice recognition engine within the device that converts voice commands into text data
[0194] 3. A means of communication to send the converted text data to the server
[0195] 4. Generative AI in the server to analyze text data and generate specific operating procedures
[0196] 5. Response method from the server that sends the generated operation procedure to the terminal
[0197] 6. An automated operation system for executing actual operations based on operating procedures at the terminal
[0198] Program processing
[0199] Acquiring and translating voice commands:
[0200] The user inputs a voice command into the terminal. For example, the voice command is "Order a Margherita pizza." This terminal is equipped with a speech recognition engine (for example, Google Speech Recognition API), which converts the user's voice into text data in real time. This text data is sent to the server by a communication means to the server, which will be explained next.
[0201] Text data transmission and analysis:
[0202] When text data is sent from a device to a server, the server first receives the data. The server then analyzes the received text data and uses generative artificial intelligence (e.g., GPT-4) to understand the user's intent. Specifically, it uses natural language understanding technology to analyze the content of the text and generate the necessary operating procedures.
[0203] Generate and send operating instructions:
[0204] The generative AI then creates specific operational steps based on the analysis results. For example, if a command such as "Order a Margherita pizza" is entered, the generative AI will analyze this and generate a series of operational steps in a food delivery app, such as "Select a category," "Select a Margherita pizza," "Specify the quantity," and "Confirm the order."
[0205] To perform an action on the device:
[0206] The server generates instructions and sends them to the device, which then automatically performs the necessary operations within the food delivery app to complete the order. This eliminates the need for manual operations on the part of the user, allowing them to complete their meal order efficiently and intuitively.
[0207] Specific examples
[0208] For example, if a user enters the voice command "Order a Margherita pizza," the system automatically performs the following steps:
[0209] 1. Acquire voice command: The user enters a voice command.
[0210] 2. Speech recognition and text conversion: The device's speech recognition engine converts speech into text data.
[0211] 3. Send to server: The converted text data is sent to the server.
[0212] 4. Analysis by generative AI: The server's generative AI analyzes the text data and generates operating procedures.
[0213] 5. Sending operation instructions: The server sends the generated operation instructions to the terminal.
[0214] 6. Automated operation execution: The device follows the operation steps and automatically completes the order in the food delivery app.
[0215] An example prompt is:
[0216] When the voice command "Order a Margherita pizza" is entered, the system performs the following actions in the food delivery app:
[0217] 1. Select a category
[0218] 2. Choose a Margherita pizza
[0219] 3. Specify the quantity
[0220] 4. Confirm your order
[0221] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0222] Step 1:
[0223] Acquiring voice commands
[0224] The user inputs a voice command into the terminal. The terminal is equipped with a microphone that captures the user's voice in real time. For example, the user says, "Order a Margherita pizza." The input here is the user's voice command, and the output is voice data.
[0225] Step 2:
[0226] Speech recognition and text conversion
[0227] The device's speech recognition engine (for example, Google Speech Recognition API) converts the voice data into text data. This voice data is then analyzed and converted into text information. The input to this step is the voice data acquired in step 1, and the output is text data.
[0228] Step 3:
[0229] Sending to the server
[0230] The terminal sends the converted text data to the server. For this purpose, a network communication means is used. The input is the text data, and the output is the transmission result to the server. If the communication is successful, the text data reaches the server.
[0231] Step 4:
[0232] Text data analysis
[0233] The server analyzes the received text data. It uses generative artificial intelligence (e.g., GPT-4) and applies natural language understanding techniques to understand the user's intent. The input is the text data sent to the server, and the output is the extracted operation procedure.
[0234] Step 5:
[0235] Generate operating instructions
[0236] Generative AI generates specific operational steps based on the analysis results. For example, based on the text data "Order a Margherita pizza," it generates a series of operational steps such as "Select a category," "Select a Margherita pizza," "Specify quantity," and "Confirm the order." The input is the analysis results, and the output is the generated series of operational steps.
[0237] Step 6:
[0238] Sending operating instructions
[0239] The server sends the generated operation procedure to the terminal. To return this operation procedure to the terminal, a network communication means is used again. The input is the generated operation procedure, and the output is the transmission result to the terminal. If the communication is successful, the operation procedure arrives at the terminal.
[0240] Step 7:
[0241] Performing automated operations
[0242] The device automatically performs the necessary operations within the food delivery app based on the received operation instructions. For example, it automatically performs operations such as "select a category," "select Margherita pizza," "specify the quantity," and "confirm the order." The input is the operation instructions sent from the server, and the output is the final order completion. This eliminates the need for the user to perform each operation manually.
[0243] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0244] The present invention relates to a system that allows users to operate multiple different applications easily and intuitively using voice commands. Furthermore, the present invention provides a more personalized user experience by combining an emotion engine that recognizes the user's emotions.
[0245] Overall system configuration
[0246] The system consists of the following main components:
[0247] 1. A device that receives voice commands
[0248] 2. A voice recognition engine within the device that converts voice commands into text data
[0249] 3. A method for analyzing user emotions using an emotion engine
[0250] 4. A communication method for sending the converted text data and emotion data to the server
[0251] 5. Generative AI in the server that analyzes text data and emotion data and generates specific operating procedures
[0252] 6. Means for sending generated operating procedures and appropriate feedback to the terminal
[0253] 7. An automated operation system for executing actual operations based on operating procedures at the terminal
[0254] Program processing
[0255] Voice command capture and emotion analysis
[0256] The user inputs voice commands into the device. The device is equipped with a speech recognition engine and an emotion engine, which not only converts the user's voice into text data in real time, but also analyzes the user's emotional state using the emotion engine. For example, the device can analyze the user's emotions based on the tone, speed, and intonation of the voice and generate corresponding data.
[0257] Sending and analyzing text and emotion data
[0258] Text data and emotion data are sent from the device to a server. The server receives and analyzes this data. Specifically, it analyzes the text data using a Natural Language Understanding (NLU) engine and uses generative artificial intelligence to generate operating procedures based on the user's instructions and emotional state.
[0259] Generate and send operating instructions
[0260] The server generates specific operating procedures based on the analysis results. For example, if a user instructs, "Add a meeting tomorrow at 3 p.m.", the generative AI analyzes this and generates operating procedures that take into account the user's emotional state. If the user's emotional state is tense, the server selects concise and reliable operating procedures and adds reassuring feedback.
[0261] Execute operations on the device
[0262] The generated operation instructions and feedback are sent to the device. The device then automatically performs a series of operations based on this. In the case of a calendar app, the device performs operations such as "add a new event," "set the date," "set the time," and "enter a title," and finally notifies the user that the operation is complete. An appropriate message based on the user's emotions (e.g., "Is everything OK?") is also displayed.
[0263] Specific examples
[0264] Below is a specific example of adding an event to a calendar app.
[0265] Example 1: Adding an event in the calendar app
[0266] 1. The user types the voice command "Add a meeting tomorrow at 3 PM."
[0267] 2. The device uses a speech recognition engine to convert the voice into text data, and an emotion engine analyzes the user's emotional state from the voice (e.g., nervousness).
[0268] 3. The device sends the text data and emotion data to the server.
[0269] 4. The server analyzes the text data and uses generative artificial intelligence to generate specific operating instructions, including a reassuring message that takes into account the user's emotional state.
[0270] 5. The server sends the generated operating instructions and feedback to the terminal.
[0271] 6. The device automatically performs the necessary operations in the calendar app based on the user's instructions to add the meeting, and displays a message to the user such as "Meeting added successfully. Is there anything else we can help you with?"
[0272] In this way, by combining voice commands and an emotion engine, the present invention provides a personalized operating experience that adapts to the user's emotional state, enabling efficient and intuitive system operation.
[0273] The processing flow will be explained below.
[0274] Step 1:
[0275] The user enters a voice command into the device, for example, "Add a meeting for tomorrow at 3:00 PM."
[0276] Step 2:
[0277] The device activates a voice recognition engine and captures voice data in real time.
[0278] Step 3:
[0279] The device converts the voice data into text data, and the speech recognition engine generates the text "Add a meeting tomorrow at 3 PM."
[0280] Step 4:
[0281] The device activates an emotion engine and analyzes the user's emotional state from the voice data. For example, it determines that the user is in a tense state based on the tone and speed of the voice.
[0282] Step 5:
[0283] The device includes the text data and emotion data in the body of an HTTP request and prepares to send it to the server.
[0284] Step 6:
[0285] The device sends text data and emotion data to the server, which then sends an HTTP request to the server.
[0286] Step 7:
[0287] The server receives the HTTP request and extracts text data and emotion data from the request body.
[0288] Step 8:
[0289] The server analyzes the text data using a Natural Language Understanding (NLU) engine, extracting information such as "Date = tomorrow," "Time = 3 PM," and "Action = Add meeting."
[0290] Step 9:
[0291] The server calls the API of a generative artificial intelligence (e.g., GPT-3) and generates specific operating procedures based on the analysis results and emotional data. For users in a tense situation, the server generates concise operating procedures that provide a sense of security.
[0292] Step 10:
[0293] The server formats the generated operation instructions and feedback messages and creates an HTTP response to send to the device.
[0294] Step 11:
[0295] The server sends an HTTP response to the device.
[0296] Step 12:
[0297] The terminal receives the HTTP response and analyzes the operation procedure and feedback message from the response.
[0298] Step 13:
[0299] The device launches the calendar app. Specifically, it calls the calendar app based on the application ID.
[0300] Step 14:
[0301] Based on the operating procedure, the device will automatically perform a series of operations within the calendar app, such as "add a new event," "set the date," "set the time," and "enter a title."
[0302] Step 15:
[0303] The device notifies the user when the operation is complete, for example by displaying a reassuring message such as "Your meeting has been successfully added. You seem nervous. Is there anything else I can help you with?"
[0304] This process allows users to enjoy a smooth and personalized experience based on voice commands and emotional data.
[0305] Example 2
[0306] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0307] Conventional voice-operated systems execute commands without considering the user's emotional state, resulting in a uniform user experience. Furthermore, device operation may not always match the user's intuition, resulting in a lack of operability. Furthermore, the lack of proper feedback makes it difficult for users to confirm whether their commands were executed correctly.
[0308] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for acquiring a voice command, means for converting the voice command into text data, means for analyzing the emotional state when acquiring the voice command, means for transmitting the text data and emotional data to the server, means for analyzing the text data and using generative artificial intelligence to generate specific operating procedures based on the user's emotional state, means for transmitting the operating procedures and feedback messages to the terminal, and means for automatically executing operations on the terminal based on the operating procedures. This makes it possible to provide personalized operating procedures according to the user's emotional state, improving operability and providing a sense of security through feedback.
[0309] A "voice command" is an operation instruction given by voice to a terminal by a user.
[0310] "Text data" refers to data obtained by converting a voice command into character information using a voice recognition engine.
[0311] "Emotional state" refers to emotional information analyzed from the user's voice, and is obtained from tone, speed, and intonation.
[0312] An "emotion engine" is a software or hardware system for analyzing the emotional state of a user's voice.
[0313] A "server" is a central processing unit that receives data sent from a terminal and performs analysis and processing.
[0314] "Generative AI" is artificial intelligence software that analyzes received data and generates operating procedures based on the user's instructions and emotional state.
[0315] An "operation procedure" is a series of steps for operating a specific device or application that is executed based on a voice command.
[0316] A "feedback message" is a message provided to a user after an operation procedure is performed, and includes information indicating that the operation was successful and prompting the user to give the next instruction.
[0317] A "terminal" is an electronic device that allows a user to input voice commands and transmit data to a server, and includes smartphones, tablets, and the like.
[0318] The present invention relates to a system that allows users to easily and intuitively operate multiple different applications using voice commands. Furthermore, the present invention provides a more personalized user experience by combining an emotion engine that recognizes the user's emotions.
[0319] System configuration
[0320] This system consists of the following main components:
[0321] 1. A device that receives voice commands
[0322] 2. A voice recognition engine within the device that converts voice commands into text data
[0323] 3. A method for analyzing user emotions using an emotion engine
[0324] 4. A communication method for sending the converted text data and emotion data to the server
[0325] 5. Generative AI in the server that analyzes text data and emotion data and generates specific operating procedures
[0326] 6. Means for sending generated operating procedures and appropriate feedback to the terminal
[0327] 7. An automated operation system for executing actual operations based on operating procedures at the terminal
[0328] Voice command capture and emotion analysis
[0329] The user inputs voice commands into the device. The device is equipped with a speech recognition engine (e.g., Google Speech-to-Text) and an emotion engine (e.g., Affectiva) that converts the user's voice into text data in real time. The emotion engine also analyzes the user's emotional state from the voice. For example, the device generates data representing the user's emotion based on the tone, speed, and intonation of the voice.
[0330] Sending and analyzing text and emotion data
[0331] The device uses a communication method (Wi-Fi, 4G / 5G, etc.) to send the converted text data and emotion data to the server. The server receives and analyzes this data. Specifically, it analyzes the text data using a Natural Language Understanding (NLU) engine (e.g., IBM Watson NLU) and generates operational procedures based on the user's instructions and emotional state using generative artificial intelligence (e.g., OpenAI GPT-4).
[0332] Generate and send operating instructions
[0333] The server generates specific operating procedures based on the analysis results. For example, if a user instructs, "Add a meeting tomorrow at 3 p.m.", the generative AI analyzes this and generates operating procedures that take into account the user's emotional state. If the user's emotional state is tense, the server selects concise and reliable operating procedures and also generates a feedback message that provides a sense of security.
[0334] Execute operations on the device
[0335] The generated operation instructions and feedback are sent to the device, which then automatically performs a series of operations based on them. For example, when using a calendar app, the user performs operations such as "add a new event," "set the date," "set the time," and "enter a title." Finally, the device notifies the user that the operation is complete and displays an appropriate emotion-based message (e.g., "The meeting has been successfully added. Is there anything else we can help you with?").
[0336] Examples and prompts
[0337] Below is a specific example of adding an event to a calendar app.
[0338] Example 1: Adding an event in the calendar app
[0339] The user enters a voice command such as "Add a meeting tomorrow at 3pm." The device converts the voice to text data and uses an emotion engine to analyze the user's emotional state (e.g., nervousness). The device sends the text data and emotion data to the server. The server analyzes the text data and generates specific operating instructions using generative artificial intelligence. A reassuring message is also generated taking the emotional state into account. The server sends the generated operating instructions and feedback to the device. The device automatically performs the necessary operations in the calendar app based on the operating instructions and adds the meeting. The user is then shown a message such as "The meeting has been added successfully. Is there anything else we can help you with?"
[0340] Prompt Sentence Examples
[0341] "Please explain in detail how the system works when a user adds an event to a calendar app using voice commands."
[0342] In this way, by combining voice commands and an emotion engine, the present invention provides a personalized operating experience that adapts to the user's emotional state, enabling efficient and intuitive system operation.
[0343] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0344] Step 1:
[0345] The user inputs a voice command into the device. The input is the user's voice, and includes instructions such as "Add a meeting tomorrow at 3:00 PM." The device's microphone picks up the user's voice, and the voice data is input.
[0346] Step 2:
[0347] The device uses a speech recognition engine (e.g., speech recognition software) to convert voice data into text data. The input is the user's voice data, and the output is text instructions. Specifically, the voice waveform is analyzed and the spoken content is converted into text.
[0348] Step 3:
[0349] The device uses an emotion engine (e.g., emotion analysis algorithm) to analyze the user's emotional state. The input is voice data and text data, and the output is data indicating the user's emotional state (e.g., tension). Specifically, the tone, intonation, and speed of the voice are analyzed to generate emotion data.
[0350] Step 4:
[0351] The device sends text data and emotion data to the server. The input is text data and emotion data, and the output is data transmission to the server. Specifically, the data is sent via a network using a communication module (e.g., Wi-Fi or 4G / 5G).
[0352] Step 5:
[0353] The server analyzes the received text data and emotional data. The input is text data and emotional data, and the output is the analysis results. Specifically, it analyzes the text data using a Natural Language Understanding (NLU) engine (e.g., natural language analysis software), and generates operating procedures based on the user's instructions and emotional state using generative artificial intelligence (e.g., AI generation model).
[0354] Step 6:
[0355] The server sends the generated operation procedures and feedback messages to the terminal. The input is the operation procedures and feedback messages, and the output is data transmission to the terminal. Specifically, the server converts the data into packets and transmits them to the terminal through the communication module.
[0356] Step 7:
[0357] The device automatically executes the actual operation based on the operation instructions received. The input is the operation instructions, and the output is the result of the operation. Specifically, using automatic operation software (e.g., automation tool), a specified application (e.g., calendar app) is opened and operations such as "add a new event," "set the date," "set the time," and "enter a title" are performed.
[0358] Step 8:
[0359] The terminal notifies the user that the operation is complete. The input is the result of the operation and a feedback message, and the output is a notification to the user. Specifically, the screen displays "The meeting has been successfully added. Is there anything else we can help you with?"
[0360] (Application example 2)
[0361] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0362] Conventional systems using speech recognition technology simply convert voice commands into text data and perform specific operations without considering the user's emotional state, resulting in a uniform user experience that cannot fully meet the needs of individual users. Additionally, there is a lack of a mechanism for recommending content appropriate for a specific emotional state, which can lead to reduced user satisfaction.
[0363] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring a voice command, means for converting the voice command into text data, means for analyzing a user's emotions, means for transmitting the text data and emotional data to the server, means for analyzing the text data and emotional data and using generative artificial intelligence to generate a specific recommendation list, means for transmitting the recommendation list and feedback to a terminal, and means for providing content on the terminal based on the recommendation list. This enables personalized content recommendation that is adapted to the user's emotional state.
[0364] Definitions of important words
[0365] A "voice command" is an instruction given by a user to a device by voice.
[0366] "Text data" refers to data obtained by converting a voice command into a string of characters.
[0367] "User emotion" refers to the emotional state analyzed from the user's voice and language.
[0368] "Emotion data" refers to data that expresses the analysis results of a user's emotions as numerical values or categories.
[0369] A "server" is a computer system that receives, analyzes, and processes voice data and emotional data.
[0370] "Generative AI" refers to an AI technology that outputs content generated based on input data.
[0371] A "recommendation list" is a list of content and operating procedures generated based on the user's emotions and voice commands.
[0372] "Feedback" refers to the response or information that a system provides to a user.
[0373] "Content" refers to digital media and information such as music, movies, articles, etc.
[0374] MODE FOR CARRYING OUT THE INVENTION
[0375] The present invention is a system that allows users to operate multiple applications using voice commands and provides personalized operation and recommendations based on emotion analysis. The system receives voice commands, converts them into text data, analyzes the user's emotions, and uses generative artificial intelligence to generate appropriate operation procedures and recommendation lists.
[0376] Overall system configuration
[0377] The system consists of the following main components:
[0378] 1. Device receiving voice commands:
[0379] A device that allows users to input voice commands, such as a smartphone or smart speaker.
[0380] 2. Speech Recognition Engine:
[0381] The device's built-in voice recognition software converts the user's voice commands into text data, using, for example, Google's voice recognition API.
[0382] 3. Emotion Engine:
[0383] This is software or algorithm that analyzes the emotional state of a user's voice. It determines emotions based on the tone and intonation of the voice. For example, an emotion analysis API can be used.
[0384] 4. Means of communication:
[0385] A communication module for transmitting the converted text data and emotion data to a server, for example, by transmitting the data over an internet connection.
[0386] 5. Generative AI in the server:
[0387] The received text data and emotion data are analyzed to generate specific operation procedures and recommendation lists in response to user instructions. Examples of this include natural language understanding engines and generative AI models.
[0388] 6. Operating instructions and means of sending recommendation lists:
[0389] This technology allows the server to generate operating procedures and recommendation lists and send them to the terminal. The data is transmitted in real time via the Internet.
[0390] 7. How to perform operations on the terminal:
[0391] This function allows the device to automatically perform actual operations or provide feedback to the user based on the operating procedures and recommendation list sent from the server.
[0392] Program processing
[0393] 1. Voice command capture and sentiment analysis:
[0394] The user inputs a voice command, and the smartphone or smart speaker picks up the voice and converts it into text data using a speech recognition engine.
[0395] The emotion engine analyzes the tone, speed, and intonation of the voice to generate emotion data for the user. For example, if a user says "I'm tired," the emotion engine analyzes the voice to determine "fatigue."
[0396] 2. Transmission and analysis of text and emotion data:
[0397] Text and emotion data are sent from the device to the server, which analyzes the data and uses a Natural Language Understanding (NLU) engine to understand the user's intent.
[0398] 3. Generate and send operating instructions:
[0399] The server uses generative artificial intelligence to generate specific operation procedures and recommendation lists based on the user's voice commands and emotional data. For example, if the user requests, "Tell me about some uplifting movies," it generates a list of "uplifting movies" recommendations.
[0400] The generated operation procedures and recommendation lists are optimized to also respond to the user's emotions.
[0401] 4. Execute operations on the device:
[0402] The generated operation instructions and recommendation list are sent to the device, which then automatically performs specific operations based on them. For example, a movie recommendation app may list "Movie A" and "Movie B" as "uplifting movies" and display a message asking, "How about these movies?"
[0403] Specific examples
[0404] Specific examples of movie recommendations
[0405] When a user inputs a voice command such as "Tell me some uplifting movies," the device converts the voice into text data and analyzes the emotions using an emotion engine. The text data and emotion data are sent to a server, which uses generative artificial intelligence to generate a list of appropriate movie recommendations. The data is finally sent to the device and displayed as "Recommended movies are: Movie A, Movie B." In this way, by combining voice commands and emotion analysis technology, the present invention provides a personalized operating experience that adapts to the user's emotional state.
[0406] Prompt Sentence Examples
[0407] "Tell me some uplifting movies"
[0408] "Play some relaxing music"
[0409] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0410] Program processing steps
[0411] Step 1: Get voice commands
[0412] The user inputs a voice command into the device. Specifically, the user speaks to a smartphone or smart speaker, saying, "Tell me about an uplifting movie."
[0413] Input: User's voice command
[0414] Output: Captured audio data
[0415] Step 2: Voice Recognition
[0416] The device converts the audio data into text using a speech recognition engine, which uses software such as Google's speech recognition API to convert the audio to text.
[0417] Input: Audio data
[0418] Output: Text data
[0419] Step 3: Sentiment Analysis
[0420] The device uses an emotion engine to generate emotion data from the user's voice, analyzing the tone, speed, and intonation of the voice.
[0421] Input: Audio data
[0422] Output: Emotion data
[0423] Step 4: Send data
[0424] The terminal transmits the text data and emotion data to the server, and the data is sent to the server via the Internet using a communication means.
[0425] Input: Text data, emotion data
[0426] Output: Data sent to the server
[0427] Step 5: Data analysis
[0428] The server analyzes the received text and emotion data and uses a Natural Language Understanding (NLU) engine to understand the user's intent and emotions.
[0429] Input: Text data, emotion data
[0430] Output: Analysis results (information based on user intent and emotions)
[0431] Step 6: Generate operating instructions and recommendation list
[0432] The server uses generative artificial intelligence to generate specific operating procedures and recommendation lists based on the user's intentions and emotional data, such as a list of uplifting movies.
[0433] Input: Analysis results
[0434] Output: Generated instructions and recommendation list
[0435] Step 7: Submit instructions and list
[0436] The server transmits the generated operation procedure and recommendation list to the terminal, and transmits the data in real time using a transmission means.
[0437] Input: Operation procedure, recommendation list
[0438] Output: Data sent to the terminal
[0439] Step 8: Operation execution and feedback
[0440] Based on the operation instructions and recommendation list received from the server, the device automatically performs specific operations and provides feedback to the user, such as displaying "Recommended movies are: Movie A, Movie B."
[0441] Input: Operation procedure, recommendation list
[0442] Output: Displayed movie recommendation list, feedback message
[0443] This allows for personalized operation and recommendations based on the user's voice commands and emotional state.
[0444] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0445] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0446] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0447] [Second embodiment]
[0448] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0449] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0450] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0451] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0452] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0453] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0454] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0455] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0456] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0457] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0458] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0459] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0460] The present invention relates to a system that allows users to easily and intuitively operate multiple different applications using voice commands. This system automates the entire process from receiving a voice command to executing the operation, improving the user experience.
[0461] Overall system configuration
[0462] The system consists of the following main components:
[0463] 1. A device that receives voice commands
[0464] 2. A voice recognition engine within the device that converts voice commands into text data
[0465] 3. A means of communication to send the converted text data to the server
[0466] 4. Generative AI in the server to analyze text data and generate specific operating procedures
[0467] 5. Response method from the server that sends the generated operation procedure to the terminal
[0468] 6. An automated operation system for executing actual operations based on operating procedures at the terminal
[0469] Program processing
[0470] Acquiring and translating voice commands
[0471] The user inputs voice commands into the terminal. The terminal is equipped with a voice recognition engine that converts the user's voice into text data in real time. This text data is then sent to the server via the server communication means described below.
[0472] Sending and analyzing text data
[0473] When text data is sent from a device to a server, the server first receives the data. The server then analyzes the received text data and uses generative artificial intelligence to understand the user's intent. Specifically, it uses natural language understanding technology to analyze the content of the text and generate the necessary operating procedures.
[0474] Generate and send operating instructions
[0475] The generative AI then creates specific operational steps based on the analysis results. For example, if a command such as "Add a meeting tomorrow at 3 PM" is entered, the generative AI will analyze this and generate a series of operational steps in the calendar app, such as "Add a new event," "Set the date," "Set the time," and "Enter a title."
[0476] Execute operations on the device
[0477] The server generates a sequence of operations and sends it to the device, which then automatically executes the sequence of operations, eliminating the need for manual intervention by the user and allowing them to complete the desired task efficiently and intuitively.
[0478] Specific examples
[0479] Below is a specific example of adding an event to a calendar app.
[0480] 1. The user types the voice command "Add a meeting tomorrow at 3 PM."
[0481] 2. The device uses a voice recognition engine to convert this voice into text data.
[0482] 3. The device sends the converted text data to the server.
[0483] 4. The server receives the text data and uses generative artificial intelligence to generate operational steps such as "Open the calendar app," "Add a new event," "Set the date," "Set the time," and "Enter a title."
[0484] 5. The server sends the generated operation procedure to the terminal.
[0485] 6. The device automatically performs the necessary operations within the calendar app based on the received instructions and adds the meeting.
[0486] In this way, the present invention provides a system that allows users to seamlessly operate a variety of applications through voice commands, achieving an efficient and intuitive operating experience while maintaining a consistent UI.
[0487] The processing flow will be explained below.
[0488] Step 1:
[0489] The user enters a voice command into the device, for example, "Add a meeting for tomorrow at 3 PM."
[0490] Step 2:
[0491] The device activates its built-in voice recognition engine and captures voice data in real time.
[0492] Step 3:
[0493] The device converts the voice data into text data, and the speech recognition engine generates the text "Add a meeting tomorrow at 3 PM."
[0494] Step 4:
[0495] The terminal includes the converted text data in the body of the HTTP request and prepares to send it to the server.
[0496] Step 5:
[0497] The device sends text data to the server. An HTTP request is sent to the server.
[0498] Step 6:
[0499] The server receives the HTTP request and extracts the text data from the request body.
[0500] Step 7:
[0501] The server analyzes the text data using a Natural Language Understanding (NLU) engine, and extracts the following information as the analysis result: "Date = tomorrow", "Time = 3 PM", "Action = Add meeting".
[0502] Step 8:
[0503] The server calls the API of a generative artificial intelligence (e.g., GPT-3) and generates specific operating procedures based on the analysis results.
[0504] Step 9:
[0505] The server formats the generated operating instructions and creates an HTTP response to send to the terminal.
[0506] Step 10:
[0507] The server sends an HTTP response to the device.
[0508] Step 11:
[0509] The terminal receives the HTTP response and analyzes the operation procedure from the response.
[0510] Step 12:
[0511] The device launches the calendar app. Specifically, it calls the calendar app based on the application ID.
[0512] Step 13:
[0513] Based on the operating procedure, the device will automatically perform a series of operations within the calendar app, such as "add a new event," "set the date," "set the time," and "enter a title."
[0514] Step 14:
[0515] The terminal notifies the user that the operation is complete, for example by displaying a message such as "Conference added."
[0516] Example 1
[0517] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0518] Current voice-command-based operation systems have the problem of being difficult to seamlessly operate multiple different applications, limiting the user experience. Furthermore, manual operation is required, reducing efficiency and making intuitive operation difficult. Therefore, there is a need for a system that allows users to easily and intuitively operate multiple applications through voice commands.
[0519] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0520] In this invention, the server includes means for receiving a voice command, means for converting the voice command into text data, means for transmitting the text data to the server, means for analyzing the text data and using generative artificial intelligence to generate specific operating procedures, means for transmitting the operating procedures to the terminal, means for executing an operation on the terminal based on the operating procedures, means for generating specific operating procedures including an operation for adding an event to a calendar application, and means for transmitting the operating procedures from the server to the terminal and automatically executing them, thereby enabling efficient and intuitive operation of various applications through voice commands.
[0521] A "voice command" is an instruction or request that a user communicates to a system through speech.
[0522] "Text data" refers to data obtained by converting voice commands into text using a voice recognition engine.
[0523] A "server" is a computer system that receives and analyzes text data, and generates and transmits operating procedures.
[0524] "Generative AI" refers to an AI technology that analyzes text data and generates specific operating procedures based on the user's intentions.
[0525] An "operation procedure" refers to a series of operations performed within an application.
[0526] "Terminal" means a device that inputs voice commands, converts voice to text, and receives and executes operating instructions.
[0527] A "voice recognition engine" is software or hardware that converts voice commands into text data.
[0528] "Communication means" refers to the communication technology for transmitting text data from the terminal to the server and transmitting operating procedures from the server to the terminal.
[0529] "Natural language understanding" refers to the technology for analyzing text written in natural language and understanding its content and intent.
[0530] "Automatic operation of a user interface" means automatically operating an application without user assistance, according to a generated operating procedure.
[0531] The present invention relates to a system that allows users to easily and intuitively operate multiple different applications using voice commands. The system aims to improve the user experience by automating the entire process from receiving a voice command to executing the operation.
[0532] Overall system configuration
[0533] The system consists of the following main components:
[0534] 1. A device that receives voice commands
[0535] 2. A voice recognition engine within the device that converts voice commands into text data
[0536] 3. A means of communication to send the converted text data to the server
[0537] 4. Generative AI in the server to analyze text data and generate specific operating procedures
[0538] 5. Response method from the server that sends the generated operation procedure to the terminal
[0539] 6. An automated operation system for executing actual operations based on operating procedures at the terminal
[0540] System Operation
[0541] Acquiring voice commands
[0542] The user inputs voice commands into the terminal. The terminal uses a voice recognition engine (e.g., general voice recognition software) to convert the user's voice into text data in real time. This text data is sent to the server via the communication means described below.
[0543] Sending and analyzing text data
[0544] After the device converts the speech into text, it sends the converted text data to a server. The server immediately analyzes the received text data and uses generative artificial intelligence (e.g., a general generative AI model) to understand the user's intent. Specifically, it uses natural language understanding technology to analyze the content of the text and generate the necessary operating instructions.
[0545] Generate and execute operating procedures
[0546] The server creates specific operational steps based on the analysis results. For example, if a command such as "Add a meeting tomorrow at 3 PM" is entered, the generative AI analyzes this and generates a series of operational steps in a calendar app (e.g., general calendar software) such as "Add a new event," "Set the date," "Set the time," and "Enter a title."
[0547] Sending and executing operating instructions
[0548] The server generates and sends the operation instructions to the terminal, which then automatically executes a series of operations according to the instructions, eliminating the need for the user to perform manual operations and allowing the user to complete the desired task efficiently and intuitively.
[0549] Specific examples
[0550] Example of adding an event in the calendar app
[0551] 1. The user types the voice command "Add a meeting tomorrow at 3 PM."
[0552] 2. The device uses a voice recognition engine to convert this voice into text data.
[0553] 3. The device sends the converted text data to the server.
[0554] 4. The server receives the text data and uses generative artificial intelligence to generate operational steps such as "Open the calendar app," "Add a new event," "Set the date," "Set the time," and "Enter a title."
[0555] 5. The server sends the generated operation procedure to the terminal.
[0556] 6. The device automatically performs the necessary operations within the calendar app based on the received instructions and adds the meeting.
[0557] Example prompts for generative AI models
[0558] The voice command "Add a meeting tomorrow at 3 PM" has been converted into text data. Please analyze this text and generate instructions for the calendar app.
[0559] In this way, the present invention provides a system that allows users to seamlessly operate a variety of applications through voice commands, achieving an efficient and intuitive operating experience while maintaining a consistent UI.
[0560] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0561] Step 1:
[0562] The user inputs a voice command into the terminal.
[0563] What happens: A user says, "Add a meeting tomorrow at 3 PM."
[0564] Input: User's voice
[0565] Output: Audio data
[0566] Step 2:
[0567] The device uses its built-in voice recognition engine to convert the voice data into text data.
[0568] What it does: The speech recognition engine analyzes the user's speech and converts the audio waveform into text.
[0569] Input: Audio data
[0570] Output: Text data "Add a meeting tomorrow at 3 PM"
[0571] Step 3:
[0572] The terminal transmits the text data to the server.
[0573] Specific behavior: Sends text data to the server using an HTTP POST request.
[0574] Input: Text data
[0575] Output: Request message (including text data)
[0576] Step 4:
[0577] The server receives the text data and analyzes it using a natural language processing engine.
[0578] Specific operation: The server receives the transmitted text data and inputs it into the generative AI model.
[0579] Input: Request message
[0580] Output: Intention analysis results of text data
[0581] Step 5:
[0582] Generative artificial intelligence generates specific operating procedures based on the analysis results.
[0583] Specific actions: Based on the command "Add a meeting tomorrow at 3 PM," it generates operational steps such as "Open the calendar app," "Add a new event," "Set the date," "Set the time," and "Enter a title."
[0584] Input: Intention analysis results of text data
[0585] Output: Specific operating instructions
[0586] Step 6:
[0587] The server sends the generated operation procedure to the terminal.
[0588] Specific operation: Sends operation instructions to the terminal as an HTTP response.
[0589] Input: Specific operating instructions
[0590] Output: Response message (including operation instructions)
[0591] Step 7:
[0592] The device will automatically perform operations within the calendar app based on the operating instructions.
[0593] Specific operation: In the calendar app, automatically "add a new event," "set the date," "set the time," and "enter a title."
[0594] Input: Response message (operation procedure)
[0595] Output: New event added to the Calendar app
[0596] In this way, a system is constructed in which a user's voice command is converted into a specific action via the terminal and the server, and is then automatically executed.
[0597] (Application example 1)
[0598] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0599] In today's diverse applications, it is cumbersome for users to perform multiple operations individually, and it is particularly time-consuming to efficiently complete an order in food delivery services. To solve this problem, a system is needed that automates a series of operations through voice commands and allows users to complete an order intuitively and efficiently.
[0600] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0601] In this invention, the server includes means for receiving a voice command, means for converting the voice command into text data, means for transmitting the text data to the server, means for analyzing the text data and using generative artificial intelligence to generate specific operation procedures, means for transmitting the operation procedures to a terminal, means for executing an operation on the terminal based on the operation procedures, and means for automating an ordering process to provide a meal delivery service, thereby enabling a user to efficiently and intuitively complete a meal order using voice commands, thereby improving the user experience.
[0602] A "voice command" is a voice instruction given by a user to a terminal.
[0603] "Text data" is voice commands converted into text information.
[0604] A "server" is a computer system that receives text data, analyzes it, and generates specific operating procedures.
[0605] "Generative AI" is an algorithm that understands and generates natural language, generating specific operating procedures from text data.
[0606] A "terminal" is a device through which a user inputs voice commands and executes generated operating procedures.
[0607] A "meal delivery service" is a service in which a user orders a meal and the order is delivered to a location specified by the user.
[0608] "Automating the ordering process" refers to the process of completing an order in a food delivery service based on voice commands without the need for manual intervention.
[0609] This invention relates to a system that allows users to easily and intuitively order food delivery services using voice commands. This system automates a series of operations from receiving a voice command to completing an order, improving the user experience.
[0610] Overall system configuration
[0611] The system consists of the following main components:
[0612] 1. A device that receives voice commands
[0613] 2. A voice recognition engine within the device that converts voice commands into text data
[0614] 3. A means of communication to send the converted text data to the server
[0615] 4. Generative AI in the server to analyze text data and generate specific operating procedures
[0616] 5. Response method from the server that sends the generated operation procedure to the terminal
[0617] 6. An automated operation system for executing actual operations based on operating procedures at the terminal
[0618] Program processing
[0619] Acquiring and translating voice commands:
[0620] The user inputs a voice command into the terminal. For example, the voice command is "Order a Margherita pizza." This terminal is equipped with a speech recognition engine (for example, Google Speech Recognition API), which converts the user's voice into text data in real time. This text data is sent to the server by a communication means to the server, which will be explained next.
[0621] Text data transmission and analysis:
[0622] When text data is sent from a device to a server, the server first receives the data. The server then analyzes the received text data and uses generative artificial intelligence (e.g., GPT-4) to understand the user's intent. Specifically, it uses natural language understanding technology to analyze the content of the text and generate the necessary operating procedures.
[0623] Generate and send operating instructions:
[0624] The generative AI then creates specific operational steps based on the analysis results. For example, if a command such as "Order a Margherita pizza" is entered, the generative AI will analyze this and generate a series of operational steps in a food delivery app, such as "Select a category," "Select a Margherita pizza," "Specify the quantity," and "Confirm the order."
[0625] To perform an action on the device:
[0626] The server generates instructions and sends them to the device, which then automatically performs the necessary operations within the food delivery app to complete the order. This eliminates the need for manual operations on the part of the user, allowing them to complete their meal order efficiently and intuitively.
[0627] Specific examples
[0628] For example, if a user enters the voice command "Order a Margherita pizza," the system automatically performs the following steps:
[0629] 1. Acquire voice command: The user enters a voice command.
[0630] 2. Speech recognition and text conversion: The device's speech recognition engine converts speech into text data.
[0631] 3. Send to server: The converted text data is sent to the server.
[0632] 4. Analysis by generative AI: The server's generative AI analyzes the text data and generates operating procedures.
[0633] 5. Sending operation instructions: The server sends the generated operation instructions to the terminal.
[0634] 6. Automated operation execution: The device follows the operation steps and automatically completes the order in the food delivery app.
[0635] An example prompt is:
[0636] When the voice command "Order a Margherita pizza" is entered, the system performs the following actions in the food delivery app:
[0637] 1. Select a category
[0638] 2. Choose a Margherita pizza
[0639] 3. Specify the quantity
[0640] 4. Confirm your order
[0641] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0642] Step 1:
[0643] Acquiring voice commands
[0644] The user inputs a voice command into the terminal. The terminal is equipped with a microphone that captures the user's voice in real time. For example, the user says, "Order a Margherita pizza." The input here is the user's voice command, and the output is voice data.
[0645] Step 2:
[0646] Speech recognition and text conversion
[0647] The device's speech recognition engine (for example, Google Speech Recognition API) converts the voice data into text data. This voice data is then analyzed and converted into text information. The input to this step is the voice data acquired in step 1, and the output is text data.
[0648] Step 3:
[0649] Sending to the server
[0650] The terminal sends the converted text data to the server. For this purpose, a network communication means is used. The input is the text data, and the output is the transmission result to the server. If the communication is successful, the text data reaches the server.
[0651] Step 4:
[0652] Text data analysis
[0653] The server analyzes the received text data. It uses generative artificial intelligence (e.g., GPT-4) and applies natural language understanding techniques to understand the user's intent. The input is the text data sent to the server, and the output is the extracted operation procedure.
[0654] Step 5:
[0655] Generate operating instructions
[0656] Generative AI generates specific operational steps based on the analysis results. For example, based on the text data "Order a Margherita pizza," it generates a series of operational steps such as "Select a category," "Select a Margherita pizza," "Specify quantity," and "Confirm the order." The input is the analysis results, and the output is the generated series of operational steps.
[0657] Step 6:
[0658] Sending operating instructions
[0659] The server sends the generated operation procedure to the terminal. To return this operation procedure to the terminal, a network communication means is used again. The input is the generated operation procedure, and the output is the transmission result to the terminal. If the communication is successful, the operation procedure arrives at the terminal.
[0660] Step 7:
[0661] Performing automated operations
[0662] The device automatically performs the necessary operations within the food delivery app based on the received operation instructions. For example, it automatically performs operations such as "select a category," "select Margherita pizza," "specify the quantity," and "confirm the order." The input is the operation instructions sent from the server, and the output is the final order completion. This eliminates the need for the user to perform each operation manually.
[0663] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0664] The present invention relates to a system that allows users to operate multiple different applications easily and intuitively using voice commands. Furthermore, the present invention provides a more personalized user experience by combining an emotion engine that recognizes the user's emotions.
[0665] Overall system configuration
[0666] The system consists of the following main components:
[0667] 1. A device that receives voice commands
[0668] 2. A voice recognition engine within the device that converts voice commands into text data
[0669] 3. A method for analyzing user emotions using an emotion engine
[0670] 4. A communication method for sending the converted text data and emotion data to the server
[0671] 5. Generative AI in the server that analyzes text data and emotion data and generates specific operating procedures
[0672] 6. Means for sending generated operating procedures and appropriate feedback to the terminal
[0673] 7. An automated operation system for executing actual operations based on operating procedures at the terminal
[0674] Program processing
[0675] Voice command capture and emotion analysis
[0676] The user inputs voice commands into the device. The device is equipped with a speech recognition engine and an emotion engine, which not only converts the user's voice into text data in real time, but also analyzes the user's emotional state using the emotion engine. For example, the device can analyze the user's emotions based on the tone, speed, and intonation of the voice and generate corresponding data.
[0677] Sending and analyzing text and emotion data
[0678] Text data and emotion data are sent from the device to a server. The server receives and analyzes this data. Specifically, it analyzes the text data using a Natural Language Understanding (NLU) engine and uses generative artificial intelligence to generate operating procedures based on the user's instructions and emotional state.
[0679] Generate and send operating instructions
[0680] The server generates specific operating procedures based on the analysis results. For example, if a user instructs, "Add a meeting tomorrow at 3 p.m.", the generative AI analyzes this and generates operating procedures that take into account the user's emotional state. If the user's emotional state is tense, the server selects concise and reliable operating procedures and adds reassuring feedback.
[0681] Execute operations on the device
[0682] The generated operation instructions and feedback are sent to the device. The device then automatically performs a series of operations based on this. In the case of a calendar app, the device performs operations such as "add a new event," "set the date," "set the time," and "enter a title," and finally notifies the user that the operation is complete. An appropriate message based on the user's emotions (e.g., "Is everything OK?") is also displayed.
[0683] Specific examples
[0684] Below is a specific example of adding an event to a calendar app.
[0685] Example 1: Adding an event in the calendar app
[0686] 1. The user types the voice command "Add a meeting tomorrow at 3 PM."
[0687] 2. The device uses a speech recognition engine to convert the voice into text data, and an emotion engine analyzes the user's emotional state from the voice (e.g., nervousness).
[0688] 3. The device sends the text data and emotion data to the server.
[0689] 4. The server analyzes the text data and uses generative artificial intelligence to generate specific operating instructions, including a reassuring message that takes into account the user's emotional state.
[0690] 5. The server sends the generated operating instructions and feedback to the terminal.
[0691] 6. The device automatically performs the necessary operations in the calendar app based on the user's instructions to add the meeting, and displays a message to the user such as "Meeting added successfully. Is there anything else we can help you with?"
[0692] In this way, by combining voice commands and an emotion engine, the present invention provides a personalized operating experience that adapts to the user's emotional state, enabling efficient and intuitive system operation.
[0693] The processing flow will be explained below.
[0694] Step 1:
[0695] The user enters a voice command into the device, for example, "Add a meeting for tomorrow at 3:00 PM."
[0696] Step 2:
[0697] The device activates a voice recognition engine and captures voice data in real time.
[0698] Step 3:
[0699] The device converts the voice data into text data, and the speech recognition engine generates the text "Add a meeting tomorrow at 3 PM."
[0700] Step 4:
[0701] The device activates an emotion engine and analyzes the user's emotional state from the voice data. For example, it determines that the user is in a tense state based on the tone and speed of the voice.
[0702] Step 5:
[0703] The device includes the text data and emotion data in the body of an HTTP request and prepares to send it to the server.
[0704] Step 6:
[0705] The device sends text data and emotion data to the server, which then sends an HTTP request to the server.
[0706] Step 7:
[0707] The server receives the HTTP request and extracts text data and emotion data from the request body.
[0708] Step 8:
[0709] The server analyzes the text data using a Natural Language Understanding (NLU) engine, extracting information such as "Date = tomorrow," "Time = 3 PM," and "Action = Add meeting."
[0710] Step 9:
[0711] The server calls the API of a generative artificial intelligence (e.g., GPT-3) and generates specific operating procedures based on the analysis results and emotional data. For users in a tense situation, the server generates concise operating procedures that provide a sense of security.
[0712] Step 10:
[0713] The server formats the generated operation instructions and feedback messages and creates an HTTP response to send to the device.
[0714] Step 11:
[0715] The server sends an HTTP response to the device.
[0716] Step 12:
[0717] The terminal receives the HTTP response and analyzes the operation procedure and feedback message from the response.
[0718] Step 13:
[0719] The device launches the calendar app. Specifically, it calls the calendar app based on the application ID.
[0720] Step 14:
[0721] Based on the operating procedure, the device will automatically perform a series of operations within the calendar app, such as "add a new event," "set the date," "set the time," and "enter a title."
[0722] Step 15:
[0723] The device notifies the user when the operation is complete, for example by displaying a reassuring message such as "Your meeting has been successfully added. You seem nervous. Is there anything else I can help you with?"
[0724] This process allows users to enjoy a smooth and personalized experience based on voice commands and emotional data.
[0725] Example 2
[0726] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0727] Conventional voice-operated systems execute commands without considering the user's emotional state, resulting in a uniform user experience. Furthermore, device operation may not always match the user's intuition, resulting in a lack of operability. Furthermore, the lack of proper feedback makes it difficult for users to confirm whether their commands were executed correctly.
[0728] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for acquiring a voice command, means for converting the voice command into text data, means for analyzing the emotional state when acquiring the voice command, means for transmitting the text data and emotional data to the server, means for analyzing the text data and using generative artificial intelligence to generate specific operating procedures based on the user's emotional state, means for transmitting the operating procedures and feedback messages to the terminal, and means for automatically executing operations on the terminal based on the operating procedures. This makes it possible to provide personalized operating procedures according to the user's emotional state, improving operability and providing a sense of security through feedback.
[0729] A "voice command" is an operation instruction given by voice to a terminal by a user.
[0730] "Text data" refers to data obtained by converting a voice command into character information using a voice recognition engine.
[0731] "Emotional state" refers to emotional information analyzed from the user's voice, and is obtained from tone, speed, and intonation.
[0732] An "emotion engine" is a software or hardware system for analyzing the emotional state of a user's voice.
[0733] A "server" is a central processing unit that receives data sent from a terminal and performs analysis and processing.
[0734] "Generative AI" is artificial intelligence software that analyzes received data and generates operating procedures based on the user's instructions and emotional state.
[0735] An "operation procedure" is a series of steps for operating a specific device or application that is executed based on a voice command.
[0736] A "feedback message" is a message provided to a user after an operation procedure is performed, and includes information indicating that the operation was successful and prompting the user to give the next instruction.
[0737] A "terminal" is an electronic device that allows a user to input voice commands and transmit data to a server, and includes smartphones, tablets, and the like.
[0738] The present invention relates to a system that allows users to easily and intuitively operate multiple different applications using voice commands. Furthermore, the present invention provides a more personalized user experience by combining an emotion engine that recognizes the user's emotions.
[0739] System configuration
[0740] This system consists of the following main components:
[0741] 1. A device that receives voice commands
[0742] 2. A voice recognition engine within the device that converts voice commands into text data
[0743] 3. A method for analyzing user emotions using an emotion engine
[0744] 4. A communication method for sending the converted text data and emotion data to the server
[0745] 5. Generative AI in the server that analyzes text data and emotion data and generates specific operating procedures
[0746] 6. Means for sending generated operating procedures and appropriate feedback to the terminal
[0747] 7. An automated operation system for executing actual operations based on operating procedures at the terminal
[0748] Voice command capture and emotion analysis
[0749] The user inputs voice commands into the device. The device is equipped with a speech recognition engine (e.g., Google Speech-to-Text) and an emotion engine (e.g., Affectiva) that converts the user's voice into text data in real time. The emotion engine also analyzes the user's emotional state from the voice. For example, the device generates data representing the user's emotion based on the tone, speed, and intonation of the voice.
[0750] Sending and analyzing text and emotion data
[0751] The device uses a communication method (Wi-Fi, 4G / 5G, etc.) to send the converted text data and emotion data to the server. The server receives and analyzes this data. Specifically, it analyzes the text data using a Natural Language Understanding (NLU) engine (e.g., IBM Watson NLU) and generates operational procedures based on the user's instructions and emotional state using generative artificial intelligence (e.g., OpenAI GPT-4).
[0752] Generate and send operating instructions
[0753] The server generates specific operating procedures based on the analysis results. For example, if a user instructs, "Add a meeting tomorrow at 3 p.m.", the generative AI analyzes this and generates operating procedures that take into account the user's emotional state. If the user's emotional state is tense, the server selects concise and reliable operating procedures and also generates a feedback message that provides a sense of security.
[0754] Execute operations on the device
[0755] The generated operation instructions and feedback are sent to the device, which then automatically performs a series of operations based on them. For example, when using a calendar app, the user performs operations such as "add a new event," "set the date," "set the time," and "enter a title." Finally, the device notifies the user that the operation is complete and displays an appropriate emotion-based message (e.g., "The meeting has been successfully added. Is there anything else we can help you with?").
[0756] Examples and prompts
[0757] Below is a specific example of adding an event to a calendar app.
[0758] Example 1: Adding an event in the calendar app
[0759] The user enters a voice command such as "Add a meeting tomorrow at 3pm." The device converts the voice to text data and uses an emotion engine to analyze the user's emotional state (e.g., nervousness). The device sends the text data and emotion data to the server. The server analyzes the text data and generates specific operating instructions using generative artificial intelligence. A reassuring message is also generated taking the emotional state into account. The server sends the generated operating instructions and feedback to the device. The device automatically performs the necessary operations in the calendar app based on the operating instructions and adds the meeting. The user is then shown a message such as "The meeting has been added successfully. Is there anything else we can help you with?"
[0760] Prompt Sentence Examples
[0761] "Please explain in detail how the system works when a user adds an event to a calendar app using voice commands."
[0762] In this way, by combining voice commands and an emotion engine, the present invention provides a personalized operating experience that adapts to the user's emotional state, enabling efficient and intuitive system operation.
[0763] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0764] Step 1:
[0765] The user inputs a voice command into the device. The input is the user's voice, and includes instructions such as "Add a meeting tomorrow at 3:00 PM." The device's microphone picks up the user's voice, and the voice data is input.
[0766] Step 2:
[0767] The device uses a speech recognition engine (e.g., speech recognition software) to convert voice data into text data. The input is the user's voice data, and the output is text instructions. Specifically, the voice waveform is analyzed and the spoken content is converted into text.
[0768] Step 3:
[0769] The device uses an emotion engine (e.g., emotion analysis algorithm) to analyze the user's emotional state. The input is voice data and text data, and the output is data indicating the user's emotional state (e.g., tension). Specifically, the tone, intonation, and speed of the voice are analyzed to generate emotion data.
[0770] Step 4:
[0771] The device sends text data and emotion data to the server. The input is text data and emotion data, and the output is data transmission to the server. Specifically, the data is sent via a network using a communication module (e.g., Wi-Fi or 4G / 5G).
[0772] Step 5:
[0773] The server analyzes the received text data and emotional data. The input is text data and emotional data, and the output is the analysis results. Specifically, it analyzes the text data using a Natural Language Understanding (NLU) engine (e.g., natural language analysis software), and generates operating procedures based on the user's instructions and emotional state using generative artificial intelligence (e.g., AI generation model).
[0774] Step 6:
[0775] The server sends the generated operation procedures and feedback messages to the terminal. The input is the operation procedures and feedback messages, and the output is data transmission to the terminal. Specifically, the server converts the data into packets and transmits them to the terminal through the communication module.
[0776] Step 7:
[0777] The device automatically executes the actual operation based on the operation instructions received. The input is the operation instructions, and the output is the result of the operation. Specifically, using automatic operation software (e.g., automation tool), a specified application (e.g., calendar app) is opened and operations such as "add a new event," "set the date," "set the time," and "enter a title" are performed.
[0778] Step 8:
[0779] The terminal notifies the user that the operation is complete. The input is the result of the operation and a feedback message, and the output is a notification to the user. Specifically, the screen displays "The meeting has been successfully added. Is there anything else we can help you with?"
[0780] (Application example 2)
[0781] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0782] Conventional systems using speech recognition technology simply convert voice commands into text data and perform specific operations without considering the user's emotional state, resulting in a uniform user experience that cannot fully meet the needs of individual users. Additionally, there is a lack of a mechanism for recommending content appropriate for a specific emotional state, which can lead to reduced user satisfaction.
[0783] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring a voice command, means for converting the voice command into text data, means for analyzing a user's emotions, means for transmitting the text data and emotional data to the server, means for analyzing the text data and emotional data and using generative artificial intelligence to generate a specific recommendation list, means for transmitting the recommendation list and feedback to a terminal, and means for providing content on the terminal based on the recommendation list. This enables personalized content recommendation that is adapted to the user's emotional state.
[0784] Definitions of important words
[0785] A "voice command" is an instruction given by a user to a device by voice.
[0786] "Text data" refers to data obtained by converting a voice command into a string of characters.
[0787] "User emotion" refers to the emotional state analyzed from the user's voice and language.
[0788] "Emotion data" refers to data that expresses the analysis results of a user's emotions as numerical values or categories.
[0789] A "server" is a computer system that receives, analyzes, and processes voice data and emotional data.
[0790] "Generative AI" refers to an AI technology that outputs content generated based on input data.
[0791] A "recommendation list" is a list of content and operating procedures generated based on the user's emotions and voice commands.
[0792] "Feedback" refers to the response or information that a system provides to a user.
[0793] "Content" refers to digital media and information such as music, movies, articles, etc.
[0794] MODE FOR CARRYING OUT THE INVENTION
[0795] The present invention is a system that allows users to operate multiple applications using voice commands and provides personalized operation and recommendations based on emotion analysis. The system receives voice commands, converts them into text data, analyzes the user's emotions, and uses generative artificial intelligence to generate appropriate operation procedures and recommendation lists.
[0796] Overall system configuration
[0797] The system consists of the following main components:
[0798] 1. Device receiving voice commands:
[0799] A device that allows users to input voice commands, such as a smartphone or smart speaker.
[0800] 2. Speech Recognition Engine:
[0801] The device's built-in voice recognition software converts the user's voice commands into text data, using, for example, Google's voice recognition API.
[0802] 3. Emotion Engine:
[0803] This is software or algorithm that analyzes the emotional state of a user's voice. It determines emotions based on the tone and intonation of the voice. For example, an emotion analysis API can be used.
[0804] 4. Means of communication:
[0805] A communication module for transmitting the converted text data and emotion data to a server, for example, by transmitting the data over an internet connection.
[0806] 5. Generative AI in the server:
[0807] The received text data and emotion data are analyzed to generate specific operation procedures and recommendation lists in response to user instructions. Examples of this include natural language understanding engines and generative AI models.
[0808] 6. Operating instructions and means of sending recommendation lists:
[0809] This technology allows the server to generate operating procedures and recommendation lists and send them to the terminal. The data is transmitted in real time via the Internet.
[0810] 7. How to perform operations on the terminal:
[0811] This function allows the device to automatically perform actual operations or provide feedback to the user based on the operating procedures and recommendation list sent from the server.
[0812] Program processing
[0813] 1. Voice command capture and sentiment analysis:
[0814] The user inputs a voice command, and the smartphone or smart speaker picks up the voice and converts it into text data using a speech recognition engine.
[0815] The emotion engine analyzes the tone, speed, and intonation of the voice to generate emotion data for the user. For example, if a user says "I'm tired," the emotion engine analyzes the voice to determine "fatigue."
[0816] 2. Transmission and analysis of text and emotion data:
[0817] Text and emotion data are sent from the device to the server, which analyzes the data and uses a Natural Language Understanding (NLU) engine to understand the user's intent.
[0818] 3. Generate and send operating instructions:
[0819] The server uses generative artificial intelligence to generate specific operation procedures and recommendation lists based on the user's voice commands and emotional data. For example, if the user requests, "Tell me about some uplifting movies," it generates a list of "uplifting movies" recommendations.
[0820] The generated operation procedures and recommendation lists are optimized to also respond to the user's emotions.
[0821] 4. Execute operations on the device:
[0822] The generated operation instructions and recommendation list are sent to the device, which then automatically performs specific operations based on them. For example, a movie recommendation app may list "Movie A" and "Movie B" as "uplifting movies" and display a message asking, "How about these movies?"
[0823] Specific examples
[0824] Specific examples of movie recommendations
[0825] When a user inputs a voice command such as "Tell me some uplifting movies," the device converts the voice into text data and analyzes the emotions using an emotion engine. The text data and emotion data are sent to a server, which uses generative artificial intelligence to generate a list of appropriate movie recommendations. The data is finally sent to the device and displayed as "Recommended movies are: Movie A, Movie B." In this way, by combining voice commands and emotion analysis technology, the present invention provides a personalized operating experience that adapts to the user's emotional state.
[0826] Prompt Sentence Examples
[0827] "Tell me some uplifting movies"
[0828] "Play some relaxing music"
[0829] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0830] Program processing steps
[0831] Step 1: Get voice commands
[0832] The user inputs a voice command into the device. Specifically, the user speaks to a smartphone or smart speaker, saying, "Tell me about an uplifting movie."
[0833] Input: User's voice command
[0834] Output: Captured audio data
[0835] Step 2: Voice Recognition
[0836] The device converts the audio data into text using a speech recognition engine, which uses software such as Google's speech recognition API to convert the audio to text.
[0837] Input: Audio data
[0838] Output: Text data
[0839] Step 3: Sentiment Analysis
[0840] The device uses an emotion engine to generate emotion data from the user's voice, analyzing the tone, speed, and intonation of the voice.
[0841] Input: Audio data
[0842] Output: Emotion data
[0843] Step 4: Send data
[0844] The terminal transmits the text data and emotion data to the server, and the data is sent to the server via the Internet using a communication means.
[0845] Input: Text data, emotion data
[0846] Output: Data sent to the server
[0847] Step 5: Data analysis
[0848] The server analyzes the received text and emotion data and uses a Natural Language Understanding (NLU) engine to understand the user's intent and emotions.
[0849] Input: Text data, emotion data
[0850] Output: Analysis results (information based on user intent and emotions)
[0851] Step 6: Generate operating instructions and recommendation list
[0852] The server uses generative artificial intelligence to generate specific operating procedures and recommendation lists based on the user's intentions and emotional data, such as a list of uplifting movies.
[0853] Input: Analysis results
[0854] Output: Generated instructions and recommendation list
[0855] Step 7: Submit instructions and list
[0856] The server transmits the generated operation procedure and recommendation list to the terminal, and transmits the data in real time using a transmission means.
[0857] Input: Operation procedure, recommendation list
[0858] Output: Data sent to the terminal
[0859] Step 8: Operation execution and feedback
[0860] Based on the operation instructions and recommendation list received from the server, the device automatically performs specific operations and provides feedback to the user, such as displaying "Recommended movies are: Movie A, Movie B."
[0861] Input: Operation procedure, recommendation list
[0862] Output: Displayed movie recommendation list, feedback message
[0863] This allows for personalized operation and recommendations based on the user's voice commands and emotional state.
[0864] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0865] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0866] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0867] [Third embodiment]
[0868] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0869] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0870] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0871] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0872] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0873] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0874] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0875] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0876] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0877] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0878] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0879] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0880] The present invention relates to a system that allows users to easily and intuitively operate multiple different applications using voice commands. This system automates the entire process from receiving a voice command to executing the operation, improving the user experience.
[0881] Overall system configuration
[0882] The system consists of the following main components:
[0883] 1. A device that receives voice commands
[0884] 2. A voice recognition engine within the device that converts voice commands into text data
[0885] 3. A means of communication to send the converted text data to the server
[0886] 4. Generative AI in the server to analyze text data and generate specific operating procedures
[0887] 5. Response method from the server that sends the generated operation procedure to the terminal
[0888] 6. An automated operation system for executing actual operations based on operating procedures at the terminal
[0889] Program processing
[0890] Acquiring and translating voice commands
[0891] The user inputs voice commands into the terminal. The terminal is equipped with a voice recognition engine that converts the user's voice into text data in real time. This text data is then sent to the server via the server communication means described below.
[0892] Sending and analyzing text data
[0893] When text data is sent from a device to a server, the server first receives the data. The server then analyzes the received text data and uses generative artificial intelligence to understand the user's intent. Specifically, it uses natural language understanding technology to analyze the content of the text and generate the necessary operating procedures.
[0894] Generate and send operating instructions
[0895] The generative AI then creates specific operational steps based on the analysis results. For example, if a command such as "Add a meeting tomorrow at 3 PM" is entered, the generative AI will analyze this and generate a series of operational steps in the calendar app, such as "Add a new event," "Set the date," "Set the time," and "Enter a title."
[0896] Execute operations on the device
[0897] The server generates a sequence of operations and sends it to the device, which then automatically executes the sequence of operations, eliminating the need for manual intervention by the user and allowing them to complete the desired task efficiently and intuitively.
[0898] Specific examples
[0899] Below is a specific example of adding an event to a calendar app.
[0900] 1. The user types the voice command "Add a meeting tomorrow at 3 PM."
[0901] 2. The device uses a voice recognition engine to convert this voice into text data.
[0902] 3. The device sends the converted text data to the server.
[0903] 4. The server receives the text data and uses generative artificial intelligence to generate operational steps such as "Open the calendar app," "Add a new event," "Set the date," "Set the time," and "Enter a title."
[0904] 5. The server sends the generated operation procedure to the terminal.
[0905] 6. The device automatically performs the necessary operations within the calendar app based on the received instructions and adds the meeting.
[0906] In this way, the present invention provides a system that allows users to seamlessly operate a variety of applications through voice commands, achieving an efficient and intuitive operating experience while maintaining a consistent UI.
[0907] The processing flow will be explained below.
[0908] Step 1:
[0909] The user enters a voice command into the device, for example, "Add a meeting for tomorrow at 3 PM."
[0910] Step 2:
[0911] The device activates its built-in voice recognition engine and captures voice data in real time.
[0912] Step 3:
[0913] The device converts the voice data into text data, and the speech recognition engine generates the text "Add a meeting tomorrow at 3 PM."
[0914] Step 4:
[0915] The terminal includes the converted text data in the body of the HTTP request and prepares to send it to the server.
[0916] Step 5:
[0917] The device sends text data to the server. An HTTP request is sent to the server.
[0918] Step 6:
[0919] The server receives the HTTP request and extracts the text data from the request body.
[0920] Step 7:
[0921] The server analyzes the text data using a Natural Language Understanding (NLU) engine, and extracts the following information as the analysis result: "Date = tomorrow", "Time = 3 PM", "Action = Add meeting".
[0922] Step 8:
[0923] The server calls the API of a generative artificial intelligence (e.g., GPT-3) and generates specific operating procedures based on the analysis results.
[0924] Step 9:
[0925] The server formats the generated operating instructions and creates an HTTP response to send to the terminal.
[0926] Step 10:
[0927] The server sends an HTTP response to the device.
[0928] Step 11:
[0929] The terminal receives the HTTP response and analyzes the operation procedure from the response.
[0930] Step 12:
[0931] The device launches the calendar app. Specifically, it calls the calendar app based on the application ID.
[0932] Step 13:
[0933] Based on the operating procedure, the device will automatically perform a series of operations within the calendar app, such as "add a new event," "set the date," "set the time," and "enter a title."
[0934] Step 14:
[0935] The terminal notifies the user that the operation is complete, for example by displaying a message such as "Conference added."
[0936] Example 1
[0937] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0938] Current voice-command-based operation systems have the problem of being difficult to seamlessly operate multiple different applications, limiting the user experience. Furthermore, manual operation is required, reducing efficiency and making intuitive operation difficult. Therefore, there is a need for a system that allows users to easily and intuitively operate multiple applications through voice commands.
[0939] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0940] In this invention, the server includes means for receiving a voice command, means for converting the voice command into text data, means for transmitting the text data to the server, means for analyzing the text data and using generative artificial intelligence to generate specific operating procedures, means for transmitting the operating procedures to the terminal, means for executing an operation on the terminal based on the operating procedures, means for generating specific operating procedures including an operation for adding an event to a calendar application, and means for transmitting the operating procedures from the server to the terminal and automatically executing them, thereby enabling efficient and intuitive operation of various applications through voice commands.
[0941] A "voice command" is an instruction or request that a user communicates to a system through speech.
[0942] "Text data" refers to data obtained by converting voice commands into text using a voice recognition engine.
[0943] A "server" is a computer system that receives and analyzes text data, and generates and transmits operating procedures.
[0944] "Generative AI" refers to an AI technology that analyzes text data and generates specific operating procedures based on the user's intentions.
[0945] An "operation procedure" refers to a series of operations performed within an application.
[0946] "Terminal" means a device that inputs voice commands, converts voice to text, and receives and executes operating instructions.
[0947] A "voice recognition engine" is software or hardware that converts voice commands into text data.
[0948] "Communication means" refers to the communication technology for transmitting text data from the terminal to the server and transmitting operating procedures from the server to the terminal.
[0949] "Natural language understanding" refers to the technology for analyzing text written in natural language and understanding its content and intent.
[0950] "Automatic operation of a user interface" means automatically operating an application without user assistance, according to a generated operating procedure.
[0951] The present invention relates to a system that allows users to easily and intuitively operate multiple different applications using voice commands. The system aims to improve the user experience by automating the entire process from receiving a voice command to executing the operation.
[0952] Overall system configuration
[0953] The system consists of the following main components:
[0954] 1. A device that receives voice commands
[0955] 2. A voice recognition engine within the device that converts voice commands into text data
[0956] 3. A means of communication to send the converted text data to the server
[0957] 4. Generative AI in the server to analyze text data and generate specific operating procedures
[0958] 5. Response method from the server that sends the generated operation procedure to the terminal
[0959] 6. An automated operation system for executing actual operations based on operating procedures at the terminal
[0960] System Operation
[0961] Acquiring voice commands
[0962] The user inputs voice commands into the terminal. The terminal uses a voice recognition engine (e.g., general voice recognition software) to convert the user's voice into text data in real time. This text data is sent to the server via the communication means described below.
[0963] Sending and analyzing text data
[0964] After the device converts the speech into text, it sends the converted text data to a server. The server immediately analyzes the received text data and uses generative artificial intelligence (e.g., a general generative AI model) to understand the user's intent. Specifically, it uses natural language understanding technology to analyze the content of the text and generate the necessary operating instructions.
[0965] Generate and execute operating procedures
[0966] The server creates specific operational steps based on the analysis results. For example, if a command such as "Add a meeting tomorrow at 3 PM" is entered, the generative AI analyzes this and generates a series of operational steps in a calendar app (e.g., general calendar software) such as "Add a new event," "Set the date," "Set the time," and "Enter a title."
[0967] Sending and executing operating instructions
[0968] The server generates and sends the operation instructions to the terminal, which then automatically executes a series of operations according to the instructions, eliminating the need for the user to perform manual operations and allowing the user to complete the desired task efficiently and intuitively.
[0969] Specific examples
[0970] Example of adding an event in the calendar app
[0971] 1. The user types the voice command "Add a meeting tomorrow at 3 PM."
[0972] 2. The device uses a voice recognition engine to convert this voice into text data.
[0973] 3. The device sends the converted text data to the server.
[0974] 4. The server receives the text data and uses generative artificial intelligence to generate operational steps such as "Open the calendar app," "Add a new event," "Set the date," "Set the time," and "Enter a title."
[0975] 5. The server sends the generated operation procedure to the terminal.
[0976] 6. The device automatically performs the necessary operations within the calendar app based on the received instructions and adds the meeting.
[0977] Example prompts for generative AI models
[0978] The voice command "Add a meeting tomorrow at 3 PM" has been converted into text data. Please analyze this text and generate instructions for the calendar app.
[0979] In this way, the present invention provides a system that allows users to seamlessly operate a variety of applications through voice commands, achieving an efficient and intuitive operating experience while maintaining a consistent UI.
[0980] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0981] Step 1:
[0982] The user inputs a voice command into the terminal.
[0983] What happens: A user says, "Add a meeting tomorrow at 3 PM."
[0984] Input: User's voice
[0985] Output: Audio data
[0986] Step 2:
[0987] The device uses its built-in voice recognition engine to convert the voice data into text data.
[0988] What it does: The speech recognition engine analyzes the user's speech and converts the audio waveform into text.
[0989] Input: Audio data
[0990] Output: Text data "Add a meeting tomorrow at 3 PM"
[0991] Step 3:
[0992] The terminal transmits the text data to the server.
[0993] Specific behavior: Sends text data to the server using an HTTP POST request.
[0994] Input: Text data
[0995] Output: Request message (including text data)
[0996] Step 4:
[0997] The server receives the text data and analyzes it using a natural language processing engine.
[0998] Specific operation: The server receives the transmitted text data and inputs it into the generative AI model.
[0999] Input: Request message
[1000] Output: Intention analysis results of text data
[1001] Step 5:
[1002] Generative artificial intelligence generates specific operating procedures based on the analysis results.
[1003] Specific actions: Based on the command "Add a meeting tomorrow at 3 PM," it generates operational steps such as "Open the calendar app," "Add a new event," "Set the date," "Set the time," and "Enter a title."
[1004] Input: Intention analysis results of text data
[1005] Output: Specific operating instructions
[1006] Step 6:
[1007] The server sends the generated operation procedure to the terminal.
[1008] Specific operation: Sends operation instructions to the terminal as an HTTP response.
[1009] Input: Specific operating instructions
[1010] Output: Response message (including operation instructions)
[1011] Step 7:
[1012] The device will automatically perform operations within the calendar app based on the operating instructions.
[1013] Specific operation: In the calendar app, automatically "add a new event," "set the date," "set the time," and "enter a title."
[1014] Input: Response message (operation procedure)
[1015] Output: New event added to the Calendar app
[1016] In this way, a system is constructed in which a user's voice command is converted into a specific action via the terminal and the server, and is then automatically executed.
[1017] (Application example 1)
[1018] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1019] In today's diverse applications, it is cumbersome for users to perform multiple operations individually, and it is particularly time-consuming to efficiently complete an order in food delivery services. To solve this problem, a system is needed that automates a series of operations through voice commands and allows users to complete an order intuitively and efficiently.
[1020] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1021] In this invention, the server includes means for receiving a voice command, means for converting the voice command into text data, means for transmitting the text data to the server, means for analyzing the text data and using generative artificial intelligence to generate specific operation procedures, means for transmitting the operation procedures to a terminal, means for executing an operation on the terminal based on the operation procedures, and means for automating an ordering process to provide a meal delivery service, thereby enabling a user to efficiently and intuitively complete a meal order using voice commands, thereby improving the user experience.
[1022] A "voice command" is a voice instruction given by a user to a terminal.
[1023] "Text data" is voice commands converted into text information.
[1024] A "server" is a computer system that receives text data, analyzes it, and generates specific operating procedures.
[1025] "Generative AI" is an algorithm that understands and generates natural language, generating specific operating procedures from text data.
[1026] A "terminal" is a device through which a user inputs voice commands and executes generated operating procedures.
[1027] A "meal delivery service" is a service in which a user orders a meal and the order is delivered to a location specified by the user.
[1028] "Automating the ordering process" refers to the process of completing an order in a food delivery service based on voice commands without the need for manual intervention.
[1029] This invention relates to a system that allows users to easily and intuitively order food delivery services using voice commands. This system automates a series of operations from receiving a voice command to completing an order, improving the user experience.
[1030] Overall system configuration
[1031] The system consists of the following main components:
[1032] 1. A device that receives voice commands
[1033] 2. A voice recognition engine within the device that converts voice commands into text data
[1034] 3. A means of communication to send the converted text data to the server
[1035] 4. Generative AI in the server to analyze text data and generate specific operating procedures
[1036] 5. Response method from the server that sends the generated operation procedure to the terminal
[1037] 6. An automated operation system for executing actual operations based on operating procedures at the terminal
[1038] Program processing
[1039] Acquiring and translating voice commands:
[1040] The user inputs a voice command into the terminal. For example, the voice command is "Order a Margherita pizza." This terminal is equipped with a speech recognition engine (for example, Google Speech Recognition API), which converts the user's voice into text data in real time. This text data is sent to the server by a communication means to the server, which will be explained next.
[1041] Text data transmission and analysis:
[1042] When text data is sent from a device to a server, the server first receives the data. The server then analyzes the received text data and uses generative artificial intelligence (e.g., GPT-4) to understand the user's intent. Specifically, it uses natural language understanding technology to analyze the content of the text and generate the necessary operating procedures.
[1043] Generate and send operating instructions:
[1044] The generative AI then creates specific operational steps based on the analysis results. For example, if a command such as "Order a Margherita pizza" is entered, the generative AI will analyze this and generate a series of operational steps in a food delivery app, such as "Select a category," "Select a Margherita pizza," "Specify the quantity," and "Confirm the order."
[1045] To perform an action on the device:
[1046] The server generates instructions and sends them to the device, which then automatically performs the necessary operations within the food delivery app to complete the order. This eliminates the need for manual operations on the part of the user, allowing them to complete their meal order efficiently and intuitively.
[1047] Specific examples
[1048] For example, if a user enters the voice command "Order a Margherita pizza," the system automatically performs the following steps:
[1049] 1. Acquire voice command: The user enters a voice command.
[1050] 2. Speech recognition and text conversion: The device's speech recognition engine converts speech into text data.
[1051] 3. Send to server: The converted text data is sent to the server.
[1052] 4. Analysis by generative AI: The server's generative AI analyzes the text data and generates operating procedures.
[1053] 5. Sending operation instructions: The server sends the generated operation instructions to the terminal.
[1054] 6. Automated operation execution: The device follows the operation steps and automatically completes the order in the food delivery app.
[1055] An example prompt is:
[1056] When the voice command "Order a Margherita pizza" is entered, the system performs the following actions in the food delivery app:
[1057] 1. Select a category
[1058] 2. Choose a Margherita pizza
[1059] 3. Specify the quantity
[1060] 4. Confirm your order
[1061] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1062] Step 1:
[1063] Acquiring voice commands
[1064] The user inputs a voice command into the terminal. The terminal is equipped with a microphone that captures the user's voice in real time. For example, the user says, "Order a Margherita pizza." The input here is the user's voice command, and the output is voice data.
[1065] Step 2:
[1066] Speech recognition and text conversion
[1067] The device's speech recognition engine (for example, Google Speech Recognition API) converts the voice data into text data. This voice data is then analyzed and converted into text information. The input to this step is the voice data acquired in step 1, and the output is text data.
[1068] Step 3:
[1069] Sending to the server
[1070] The terminal sends the converted text data to the server. For this purpose, a network communication means is used. The input is the text data, and the output is the transmission result to the server. If the communication is successful, the text data reaches the server.
[1071] Step 4:
[1072] Text data analysis
[1073] The server analyzes the received text data. It uses generative artificial intelligence (e.g., GPT-4) and applies natural language understanding techniques to understand the user's intent. The input is the text data sent to the server, and the output is the extracted operation procedure.
[1074] Step 5:
[1075] Generate operating instructions
[1076] Generative AI generates specific operational steps based on the analysis results. For example, based on the text data "Order a Margherita pizza," it generates a series of operational steps such as "Select a category," "Select a Margherita pizza," "Specify quantity," and "Confirm the order." The input is the analysis results, and the output is the generated series of operational steps.
[1077] Step 6:
[1078] Sending operating instructions
[1079] The server sends the generated operation procedure to the terminal. To return this operation procedure to the terminal, a network communication means is used again. The input is the generated operation procedure, and the output is the transmission result to the terminal. If the communication is successful, the operation procedure arrives at the terminal.
[1080] Step 7:
[1081] Performing automated operations
[1082] The device automatically performs the necessary operations within the food delivery app based on the received operation instructions. For example, it automatically performs operations such as "select a category," "select Margherita pizza," "specify the quantity," and "confirm the order." The input is the operation instructions sent from the server, and the output is the final order completion. This eliminates the need for the user to perform each operation manually.
[1083] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1084] The present invention relates to a system that allows users to operate multiple different applications easily and intuitively using voice commands. Furthermore, the present invention provides a more personalized user experience by combining an emotion engine that recognizes the user's emotions.
[1085] Overall system configuration
[1086] The system consists of the following main components:
[1087] 1. A device that receives voice commands
[1088] 2. A voice recognition engine within the device that converts voice commands into text data
[1089] 3. A method for analyzing user emotions using an emotion engine
[1090] 4. A communication method for sending the converted text data and emotion data to the server
[1091] 5. Generative AI in the server that analyzes text data and emotion data and generates specific operating procedures
[1092] 6. Means for sending generated operating procedures and appropriate feedback to the terminal
[1093] 7. An automated operation system for executing actual operations based on operating procedures at the terminal
[1094] Program processing
[1095] Voice command capture and emotion analysis
[1096] The user inputs voice commands into the device. The device is equipped with a speech recognition engine and an emotion engine, which not only converts the user's voice into text data in real time, but also analyzes the user's emotional state using the emotion engine. For example, the device can analyze the user's emotions based on the tone, speed, and intonation of the voice and generate corresponding data.
[1097] Sending and analyzing text and emotion data
[1098] Text data and emotion data are sent from the device to a server. The server receives and analyzes this data. Specifically, it analyzes the text data using a Natural Language Understanding (NLU) engine and uses generative artificial intelligence to generate operating procedures based on the user's instructions and emotional state.
[1099] Generate and send operating instructions
[1100] The server generates specific operating procedures based on the analysis results. For example, if a user instructs, "Add a meeting tomorrow at 3 p.m.", the generative AI analyzes this and generates operating procedures that take into account the user's emotional state. If the user's emotional state is tense, the server selects concise and reliable operating procedures and adds reassuring feedback.
[1101] Execute operations on the device
[1102] The generated operation instructions and feedback are sent to the device. The device then automatically performs a series of operations based on this. In the case of a calendar app, the device performs operations such as "add a new event," "set the date," "set the time," and "enter a title," and finally notifies the user that the operation is complete. An appropriate message based on the user's emotions (e.g., "Is everything OK?") is also displayed.
[1103] Specific examples
[1104] Below is a specific example of adding an event to a calendar app.
[1105] Example 1: Adding an event in the calendar app
[1106] 1. The user types the voice command "Add a meeting tomorrow at 3 PM."
[1107] 2. The device uses a speech recognition engine to convert the voice into text data, and an emotion engine analyzes the user's emotional state from the voice (e.g., nervousness).
[1108] 3. The device sends the text data and emotion data to the server.
[1109] 4. The server analyzes the text data and uses generative artificial intelligence to generate specific operating instructions, including a reassuring message that takes into account the user's emotional state.
[1110] 5. The server sends the generated operating instructions and feedback to the terminal.
[1111] 6. The device automatically performs the necessary operations in the calendar app based on the user's instructions to add the meeting, and displays a message to the user such as "Meeting added successfully. Is there anything else we can help you with?"
[1112] In this way, by combining voice commands and an emotion engine, the present invention provides a personalized operating experience that adapts to the user's emotional state, enabling efficient and intuitive system operation.
[1113] The processing flow will be explained below.
[1114] Step 1:
[1115] The user enters a voice command into the device, for example, "Add a meeting for tomorrow at 3:00 PM."
[1116] Step 2:
[1117] The device activates a voice recognition engine and captures voice data in real time.
[1118] Step 3:
[1119] The device converts the voice data into text data, and the speech recognition engine generates the text "Add a meeting tomorrow at 3 PM."
[1120] Step 4:
[1121] The device activates an emotion engine and analyzes the user's emotional state from the voice data. For example, it determines that the user is in a tense state based on the tone and speed of the voice.
[1122] Step 5:
[1123] The device includes the text data and emotion data in the body of an HTTP request and prepares to send it to the server.
[1124] Step 6:
[1125] The device sends text data and emotion data to the server, which then sends an HTTP request to the server.
[1126] Step 7:
[1127] The server receives the HTTP request and extracts text data and emotion data from the request body.
[1128] Step 8:
[1129] The server analyzes the text data using a Natural Language Understanding (NLU) engine, extracting information such as "Date = tomorrow," "Time = 3 PM," and "Action = Add meeting."
[1130] Step 9:
[1131] The server calls the API of a generative artificial intelligence (e.g., GPT-3) and generates specific operating procedures based on the analysis results and emotional data. For users in a tense situation, the server generates concise operating procedures that provide a sense of security.
[1132] Step 10:
[1133] The server formats the generated operation instructions and feedback messages and creates an HTTP response to send to the device.
[1134] Step 11:
[1135] The server sends an HTTP response to the device.
[1136] Step 12:
[1137] The terminal receives the HTTP response and analyzes the operation procedure and feedback message from the response.
[1138] Step 13:
[1139] The device launches the calendar app. Specifically, it calls the calendar app based on the application ID.
[1140] Step 14:
[1141] Based on the operating procedure, the device will automatically perform a series of operations within the calendar app, such as "add a new event," "set the date," "set the time," and "enter a title."
[1142] Step 15:
[1143] The device notifies the user when the operation is complete, for example by displaying a reassuring message such as "Your meeting has been successfully added. You seem nervous. Is there anything else I can help you with?"
[1144] This process allows users to enjoy a smooth and personalized experience based on voice commands and emotional data.
[1145] Example 2
[1146] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1147] Conventional voice-operated systems execute commands without considering the user's emotional state, resulting in a uniform user experience. Furthermore, device operation may not always match the user's intuition, resulting in a lack of operability. Furthermore, the lack of proper feedback makes it difficult for users to confirm whether their commands were executed correctly.
[1148] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for acquiring a voice command, means for converting the voice command into text data, means for analyzing the emotional state when acquiring the voice command, means for transmitting the text data and emotional data to the server, means for analyzing the text data and using generative artificial intelligence to generate specific operating procedures based on the user's emotional state, means for transmitting the operating procedures and feedback messages to the terminal, and means for automatically executing operations on the terminal based on the operating procedures. This makes it possible to provide personalized operating procedures according to the user's emotional state, improving operability and providing a sense of security through feedback.
[1149] A "voice command" is an operation instruction given by voice to a terminal by a user.
[1150] "Text data" refers to data obtained by converting a voice command into character information using a voice recognition engine.
[1151] "Emotional state" refers to emotional information analyzed from the user's voice, and is obtained from tone, speed, and intonation.
[1152] An "emotion engine" is a software or hardware system for analyzing the emotional state of a user's voice.
[1153] A "server" is a central processing unit that receives data sent from a terminal and performs analysis and processing.
[1154] "Generative AI" is artificial intelligence software that analyzes received data and generates operating procedures based on the user's instructions and emotional state.
[1155] An "operation procedure" is a series of steps for operating a specific device or application that is executed based on a voice command.
[1156] A "feedback message" is a message provided to a user after an operation procedure is performed, and includes information indicating that the operation was successful and prompting the user to give the next instruction.
[1157] A "terminal" is an electronic device that allows a user to input voice commands and transmit data to a server, and includes smartphones, tablets, and the like.
[1158] The present invention relates to a system that allows users to easily and intuitively operate multiple different applications using voice commands. Furthermore, the present invention provides a more personalized user experience by combining an emotion engine that recognizes the user's emotions.
[1159] System configuration
[1160] This system consists of the following main components:
[1161] 1. A device that receives voice commands
[1162] 2. A voice recognition engine within the device that converts voice commands into text data
[1163] 3. A method for analyzing user emotions using an emotion engine
[1164] 4. A communication method for sending the converted text data and emotion data to the server
[1165] 5. Generative AI in the server that analyzes text data and emotion data and generates specific operating procedures
[1166] 6. Means for sending generated operating procedures and appropriate feedback to the terminal
[1167] 7. An automated operation system for executing actual operations based on operating procedures at the terminal
[1168] Voice command capture and emotion analysis
[1169] The user inputs voice commands into the device. The device is equipped with a speech recognition engine (e.g., Google Speech-to-Text) and an emotion engine (e.g., Affectiva) that converts the user's voice into text data in real time. The emotion engine also analyzes the user's emotional state from the voice. For example, the device generates data representing the user's emotion based on the tone, speed, and intonation of the voice.
[1170] Sending and analyzing text and emotion data
[1171] The device uses a communication method (Wi-Fi, 4G / 5G, etc.) to send the converted text data and emotion data to the server. The server receives and analyzes this data. Specifically, it analyzes the text data using a Natural Language Understanding (NLU) engine (e.g., IBM Watson NLU) and generates operational procedures based on the user's instructions and emotional state using generative artificial intelligence (e.g., OpenAI GPT-4).
[1172] Generate and send operating instructions
[1173] The server generates specific operating procedures based on the analysis results. For example, if a user instructs, "Add a meeting tomorrow at 3 p.m.", the generative AI analyzes this and generates operating procedures that take into account the user's emotional state. If the user's emotional state is tense, the server selects concise and reliable operating procedures and also generates a feedback message that provides a sense of security.
[1174] Execute operations on the device
[1175] The generated operation instructions and feedback are sent to the device, which then automatically performs a series of operations based on them. For example, when using a calendar app, the user performs operations such as "add a new event," "set the date," "set the time," and "enter a title." Finally, the device notifies the user that the operation is complete and displays an appropriate emotion-based message (e.g., "The meeting has been successfully added. Is there anything else we can help you with?").
[1176] Examples and prompts
[1177] Below is a specific example of adding an event to a calendar app.
[1178] Example 1: Adding an event in the calendar app
[1179] The user enters a voice command such as "Add a meeting tomorrow at 3pm." The device converts the voice to text data and uses an emotion engine to analyze the user's emotional state (e.g., nervousness). The device sends the text data and emotion data to the server. The server analyzes the text data and generates specific operating instructions using generative artificial intelligence. A reassuring message is also generated taking the emotional state into account. The server sends the generated operating instructions and feedback to the device. The device automatically performs the necessary operations in the calendar app based on the operating instructions and adds the meeting. The user is then shown a message such as "The meeting has been added successfully. Is there anything else we can help you with?"
[1180] Prompt Sentence Examples
[1181] "Please explain in detail how the system works when a user adds an event to a calendar app using voice commands."
[1182] In this way, by combining voice commands and an emotion engine, the present invention provides a personalized operating experience that adapts to the user's emotional state, enabling efficient and intuitive system operation.
[1183] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1184] Step 1:
[1185] The user inputs a voice command into the device. The input is the user's voice, and includes instructions such as "Add a meeting tomorrow at 3:00 PM." The device's microphone picks up the user's voice, and the voice data is input.
[1186] Step 2:
[1187] The device uses a speech recognition engine (e.g., speech recognition software) to convert voice data into text data. The input is the user's voice data, and the output is text instructions. Specifically, the voice waveform is analyzed and the spoken content is converted into text.
[1188] Step 3:
[1189] The device uses an emotion engine (e.g., emotion analysis algorithm) to analyze the user's emotional state. The input is voice data and text data, and the output is data indicating the user's emotional state (e.g., tension). Specifically, the tone, intonation, and speed of the voice are analyzed to generate emotion data.
[1190] Step 4:
[1191] The device sends text data and emotion data to the server. The input is text data and emotion data, and the output is data transmission to the server. Specifically, the data is sent via a network using a communication module (e.g., Wi-Fi or 4G / 5G).
[1192] Step 5:
[1193] The server analyzes the received text data and emotional data. The input is text data and emotional data, and the output is the analysis results. Specifically, it analyzes the text data using a Natural Language Understanding (NLU) engine (e.g., natural language analysis software), and generates operating procedures based on the user's instructions and emotional state using generative artificial intelligence (e.g., AI generation model).
[1194] Step 6:
[1195] The server sends the generated operation procedures and feedback messages to the terminal. The input is the operation procedures and feedback messages, and the output is data transmission to the terminal. Specifically, the server converts the data into packets and transmits them to the terminal through the communication module.
[1196] Step 7:
[1197] The device automatically executes the actual operation based on the operation instructions received. The input is the operation instructions, and the output is the result of the operation. Specifically, using automatic operation software (e.g., automation tool), a specified application (e.g., calendar app) is opened and operations such as "add a new event," "set the date," "set the time," and "enter a title" are performed.
[1198] Step 8:
[1199] The terminal notifies the user that the operation is complete. The input is the result of the operation and a feedback message, and the output is a notification to the user. Specifically, the screen displays "The meeting has been successfully added. Is there anything else we can help you with?"
[1200] (Application example 2)
[1201] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1202] Conventional systems using speech recognition technology simply convert voice commands into text data and perform specific operations without considering the user's emotional state, resulting in a uniform user experience that cannot fully meet the needs of individual users. Additionally, there is a lack of a mechanism for recommending content appropriate for a specific emotional state, which can lead to reduced user satisfaction.
[1203] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring a voice command, means for converting the voice command into text data, means for analyzing a user's emotions, means for transmitting the text data and emotional data to the server, means for analyzing the text data and emotional data and using generative artificial intelligence to generate a specific recommendation list, means for transmitting the recommendation list and feedback to a terminal, and means for providing content on the terminal based on the recommendation list. This enables personalized content recommendation that is adapted to the user's emotional state.
[1204] Definitions of important words
[1205] A "voice command" is an instruction given by a user to a device by voice.
[1206] "Text data" refers to data obtained by converting a voice command into a string of characters.
[1207] "User emotion" refers to the emotional state analyzed from the user's voice and language.
[1208] "Emotion data" refers to data that expresses the analysis results of a user's emotions as numerical values or categories.
[1209] A "server" is a computer system that receives, analyzes, and processes voice data and emotional data.
[1210] "Generative AI" refers to an AI technology that outputs content generated based on input data.
[1211] A "recommendation list" is a list of content and operating procedures generated based on the user's emotions and voice commands.
[1212] "Feedback" refers to the response or information that a system provides to a user.
[1213] "Content" refers to digital media and information such as music, movies, articles, etc.
[1214] MODE FOR CARRYING OUT THE INVENTION
[1215] The present invention is a system that allows users to operate multiple applications using voice commands and provides personalized operation and recommendations based on emotion analysis. The system receives voice commands, converts them into text data, analyzes the user's emotions, and uses generative artificial intelligence to generate appropriate operation procedures and recommendation lists.
[1216] Overall system configuration
[1217] The system consists of the following main components:
[1218] 1. Device receiving voice commands:
[1219] A device that allows users to input voice commands, such as a smartphone or smart speaker.
[1220] 2. Speech Recognition Engine:
[1221] The device's built-in voice recognition software converts the user's voice commands into text data, using, for example, Google's voice recognition API.
[1222] 3. Emotion Engine:
[1223] This is software or algorithm that analyzes the emotional state of a user's voice. It determines emotions based on the tone and intonation of the voice. For example, an emotion analysis API can be used.
[1224] 4. Means of communication:
[1225] A communication module for transmitting the converted text data and emotion data to a server, for example, by transmitting the data over an internet connection.
[1226] 5. Generative AI in the server:
[1227] The received text data and emotion data are analyzed to generate specific operation procedures and recommendation lists in response to user instructions. Examples of this include natural language understanding engines and generative AI models.
[1228] 6. Operating instructions and means of sending recommendation lists:
[1229] This technology allows the server to generate operating procedures and recommendation lists and send them to the terminal. The data is transmitted in real time via the Internet.
[1230] 7. How to perform operations on the terminal:
[1231] This function allows the device to automatically perform actual operations or provide feedback to the user based on the operating procedures and recommendation list sent from the server.
[1232] Program processing
[1233] 1. Voice command capture and sentiment analysis:
[1234] The user inputs a voice command, and the smartphone or smart speaker picks up the voice and converts it into text data using a speech recognition engine.
[1235] The emotion engine analyzes the tone, speed, and intonation of the voice to generate emotion data for the user. For example, if a user says "I'm tired," the emotion engine analyzes the voice to determine "fatigue."
[1236] 2. Transmission and analysis of text and emotion data:
[1237] Text and emotion data are sent from the device to the server, which analyzes the data and uses a Natural Language Understanding (NLU) engine to understand the user's intent.
[1238] 3. Generate and send operating instructions:
[1239] The server uses generative artificial intelligence to generate specific operation procedures and recommendation lists based on the user's voice commands and emotional data. For example, if the user requests, "Tell me about some uplifting movies," it generates a list of "uplifting movies" recommendations.
[1240] The generated operation procedures and recommendation lists are optimized to also respond to the user's emotions.
[1241] 4. Execute operations on the device:
[1242] The generated operation instructions and recommendation list are sent to the device, which then automatically performs specific operations based on them. For example, a movie recommendation app may list "Movie A" and "Movie B" as "uplifting movies" and display a message asking, "How about these movies?"
[1243] Specific examples
[1244] Specific examples of movie recommendations
[1245] When a user inputs a voice command such as "Tell me some uplifting movies," the device converts the voice into text data and analyzes the emotions using an emotion engine. The text data and emotion data are sent to a server, which uses generative artificial intelligence to generate a list of appropriate movie recommendations. The data is finally sent to the device and displayed as "Recommended movies are: Movie A, Movie B." In this way, by combining voice commands and emotion analysis technology, the present invention provides a personalized operating experience that adapts to the user's emotional state.
[1246] Prompt Sentence Examples
[1247] "Tell me some uplifting movies"
[1248] "Play some relaxing music"
[1249] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1250] Program processing steps
[1251] Step 1: Get voice commands
[1252] The user inputs a voice command into the device. Specifically, the user speaks to a smartphone or smart speaker, saying, "Tell me about an uplifting movie."
[1253] Input: User's voice command
[1254] Output: Captured audio data
[1255] Step 2: Voice Recognition
[1256] The device converts the audio data into text using a speech recognition engine, which uses software such as Google's speech recognition API to convert the audio to text.
[1257] Input: Audio data
[1258] Output: Text data
[1259] Step 3: Sentiment Analysis
[1260] The device uses an emotion engine to generate emotion data from the user's voice, analyzing the tone, speed, and intonation of the voice.
[1261] Input: Audio data
[1262] Output: Emotion data
[1263] Step 4: Send data
[1264] The terminal transmits the text data and emotion data to the server, and the data is sent to the server via the Internet using a communication means.
[1265] Input: Text data, emotion data
[1266] Output: Data sent to the server
[1267] Step 5: Data analysis
[1268] The server analyzes the received text and emotion data and uses a Natural Language Understanding (NLU) engine to understand the user's intent and emotions.
[1269] Input: Text data, emotion data
[1270] Output: Analysis results (information based on user intent and emotions)
[1271] Step 6: Generate operating instructions and recommendation list
[1272] The server uses generative artificial intelligence to generate specific operating procedures and recommendation lists based on the user's intentions and emotional data, such as a list of uplifting movies.
[1273] Input: Analysis results
[1274] Output: Generated instructions and recommendation list
[1275] Step 7: Submit instructions and list
[1276] The server transmits the generated operation procedure and recommendation list to the terminal, and transmits the data in real time using a transmission means.
[1277] Input: Operation procedure, recommendation list
[1278] Output: Data sent to the terminal
[1279] Step 8: Operation execution and feedback
[1280] Based on the operation instructions and recommendation list received from the server, the device automatically performs specific operations and provides feedback to the user, such as displaying "Recommended movies are: Movie A, Movie B."
[1281] Input: Operation procedure, recommendation list
[1282] Output: Displayed movie recommendation list, feedback message
[1283] This allows for personalized operation and recommendations based on the user's voice commands and emotional state.
[1284] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1285] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1286] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1287] [Fourth embodiment]
[1288] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1289] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1290] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1291] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1292] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1293] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1294] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1295] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1296] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1297] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1298] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1299] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1300] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1301] The present invention relates to a system that allows users to easily and intuitively operate multiple different applications using voice commands. This system automates the entire process from receiving a voice command to executing the operation, improving the user experience.
[1302] Overall system configuration
[1303] The system consists of the following main components:
[1304] 1. A device that receives voice commands
[1305] 2. A voice recognition engine within the device that converts voice commands into text data
[1306] 3. A means of communication to send the converted text data to the server
[1307] 4. Generative AI in the server to analyze text data and generate specific operating procedures
[1308] 5. Response method from the server that sends the generated operation procedure to the terminal
[1309] 6. An automated operation system for executing actual operations based on operating procedures at the terminal
[1310] Program processing
[1311] Acquiring and translating voice commands
[1312] The user inputs voice commands into the terminal. The terminal is equipped with a voice recognition engine that converts the user's voice into text data in real time. This text data is then sent to the server via the server communication means described below.
[1313] Sending and analyzing text data
[1314] When text data is sent from a device to a server, the server first receives the data. The server then analyzes the received text data and uses generative artificial intelligence to understand the user's intent. Specifically, it uses natural language understanding technology to analyze the content of the text and generate the necessary operating procedures.
[1315] Generate and send operating instructions
[1316] The generative AI then creates specific operational steps based on the analysis results. For example, if a command such as "Add a meeting tomorrow at 3 PM" is entered, the generative AI will analyze this and generate a series of operational steps in the calendar app, such as "Add a new event," "Set the date," "Set the time," and "Enter a title."
[1317] Execute operations on the device
[1318] The server generates a sequence of operations and sends it to the device, which then automatically executes the sequence of operations, eliminating the need for manual intervention by the user and allowing them to complete the desired task efficiently and intuitively.
[1319] Specific examples
[1320] Below is a specific example of adding an event to a calendar app.
[1321] 1. The user types the voice command "Add a meeting tomorrow at 3 PM."
[1322] 2. The device uses a voice recognition engine to convert this voice into text data.
[1323] 3. The device sends the converted text data to the server.
[1324] 4. The server receives the text data and uses generative artificial intelligence to generate operational steps such as "Open the calendar app," "Add a new event," "Set the date," "Set the time," and "Enter a title."
[1325] 5. The server sends the generated operation procedure to the terminal.
[1326] 6. The device automatically performs the necessary operations within the calendar app based on the received instructions and adds the meeting.
[1327] In this way, the present invention provides a system that allows users to seamlessly operate a variety of applications through voice commands, achieving an efficient and intuitive operating experience while maintaining a consistent UI.
[1328] The processing flow will be explained below.
[1329] Step 1:
[1330] The user enters a voice command into the device, for example, "Add a meeting for tomorrow at 3 PM."
[1331] Step 2:
[1332] The device activates its built-in voice recognition engine and captures voice data in real time.
[1333] Step 3:
[1334] The device converts the voice data into text data, and the speech recognition engine generates the text "Add a meeting tomorrow at 3 PM."
[1335] Step 4:
[1336] The terminal includes the converted text data in the body of the HTTP request and prepares to send it to the server.
[1337] Step 5:
[1338] The device sends text data to the server. An HTTP request is sent to the server.
[1339] Step 6:
[1340] The server receives the HTTP request and extracts the text data from the request body.
[1341] Step 7:
[1342] The server analyzes the text data using a Natural Language Understanding (NLU) engine, and extracts the following information as the analysis result: "Date = tomorrow", "Time = 3 PM", "Action = Add meeting".
[1343] Step 8:
[1344] The server calls the API of a generative artificial intelligence (e.g., GPT-3) and generates specific operating procedures based on the analysis results.
[1345] Step 9:
[1346] The server formats the generated operating instructions and creates an HTTP response to send to the terminal.
[1347] Step 10:
[1348] The server sends an HTTP response to the device.
[1349] Step 11:
[1350] The terminal receives the HTTP response and analyzes the operation procedure from the response.
[1351] Step 12:
[1352] The device launches the calendar app. Specifically, it calls the calendar app based on the application ID.
[1353] Step 13:
[1354] Based on the operating procedure, the device will automatically perform a series of operations within the calendar app, such as "add a new event," "set the date," "set the time," and "enter a title."
[1355] Step 14:
[1356] The terminal notifies the user that the operation is complete, for example by displaying a message such as "Conference added."
[1357] Example 1
[1358] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1359] Current voice-command-based operation systems have the problem of being difficult to seamlessly operate multiple different applications, limiting the user experience. Furthermore, manual operation is required, reducing efficiency and making intuitive operation difficult. Therefore, there is a need for a system that allows users to easily and intuitively operate multiple applications through voice commands.
[1360] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1361] In this invention, the server includes means for receiving a voice command, means for converting the voice command into text data, means for transmitting the text data to the server, means for analyzing the text data and using generative artificial intelligence to generate specific operating procedures, means for transmitting the operating procedures to the terminal, means for executing an operation on the terminal based on the operating procedures, means for generating specific operating procedures including an operation for adding an event to a calendar application, and means for transmitting the operating procedures from the server to the terminal and automatically executing them, thereby enabling efficient and intuitive operation of various applications through voice commands.
[1362] A "voice command" is an instruction or request that a user communicates to a system through speech.
[1363] "Text data" refers to data obtained by converting voice commands into text using a voice recognition engine.
[1364] A "server" is a computer system that receives and analyzes text data, and generates and transmits operating procedures.
[1365] "Generative AI" refers to an AI technology that analyzes text data and generates specific operating procedures based on the user's intentions.
[1366] An "operation procedure" refers to a series of operations performed within an application.
[1367] "Terminal" means a device that inputs voice commands, converts voice to text, and receives and executes operating instructions.
[1368] A "voice recognition engine" is software or hardware that converts voice commands into text data.
[1369] "Communication means" refers to the communication technology for transmitting text data from the terminal to the server and transmitting operating procedures from the server to the terminal.
[1370] "Natural language understanding" refers to the technology for analyzing text written in natural language and understanding its content and intent.
[1371] "Automatic operation of a user interface" means automatically operating an application without user assistance, according to a generated operating procedure.
[1372] The present invention relates to a system that allows users to easily and intuitively operate multiple different applications using voice commands. The system aims to improve the user experience by automating the entire process from receiving a voice command to executing the operation.
[1373] Overall system configuration
[1374] The system consists of the following main components:
[1375] 1. A device that receives voice commands
[1376] 2. A voice recognition engine within the device that converts voice commands into text data
[1377] 3. A means of communication to send the converted text data to the server
[1378] 4. Generative AI in the server to analyze text data and generate specific operating procedures
[1379] 5. Response method from the server that sends the generated operation procedure to the terminal
[1380] 6. An automated operation system for executing actual operations based on operating procedures at the terminal
[1381] System Operation
[1382] Acquiring voice commands
[1383] The user inputs voice commands into the terminal. The terminal uses a voice recognition engine (e.g., general voice recognition software) to convert the user's voice into text data in real time. This text data is sent to the server via the communication means described below.
[1384] Sending and analyzing text data
[1385] After the device converts the speech into text, it sends the converted text data to a server. The server immediately analyzes the received text data and uses generative artificial intelligence (e.g., a general generative AI model) to understand the user's intent. Specifically, it uses natural language understanding technology to analyze the content of the text and generate the necessary operating instructions.
[1386] Generate and execute operating procedures
[1387] The server creates specific operational steps based on the analysis results. For example, if a command such as "Add a meeting tomorrow at 3 PM" is entered, the generative AI analyzes this and generates a series of operational steps in a calendar app (e.g., general calendar software) such as "Add a new event," "Set the date," "Set the time," and "Enter a title."
[1388] Sending and executing operating instructions
[1389] The server generates and sends the operation instructions to the terminal, which then automatically executes a series of operations according to the instructions, eliminating the need for the user to perform manual operations and allowing the user to complete the desired task efficiently and intuitively.
[1390] Specific examples
[1391] Example of adding an event in the calendar app
[1392] 1. The user types the voice command "Add a meeting tomorrow at 3 PM."
[1393] 2. The device uses a voice recognition engine to convert this voice into text data.
[1394] 3. The device sends the converted text data to the server.
[1395] 4. The server receives the text data and uses generative artificial intelligence to generate operational steps such as "Open the calendar app," "Add a new event," "Set the date," "Set the time," and "Enter a title."
[1396] 5. The server sends the generated operation procedure to the terminal.
[1397] 6. The device automatically performs the necessary operations within the calendar app based on the received instructions and adds the meeting.
[1398] Example prompts for generative AI models
[1399] The voice command "Add a meeting tomorrow at 3 PM" has been converted into text data. Please analyze this text and generate instructions for the calendar app.
[1400] In this way, the present invention provides a system that allows users to seamlessly operate a variety of applications through voice commands, achieving an efficient and intuitive operating experience while maintaining a consistent UI.
[1401] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1402] Step 1:
[1403] The user inputs a voice command into the terminal.
[1404] What happens: A user says, "Add a meeting tomorrow at 3 PM."
[1405] Input: User's voice
[1406] Output: Audio data
[1407] Step 2:
[1408] The device uses its built-in voice recognition engine to convert the voice data into text data.
[1409] What it does: The speech recognition engine analyzes the user's speech and converts the audio waveform into text.
[1410] Input: Audio data
[1411] Output: Text data "Add a meeting tomorrow at 3 PM"
[1412] Step 3:
[1413] The terminal transmits the text data to the server.
[1414] Specific behavior: Sends text data to the server using an HTTP POST request.
[1415] Input: Text data
[1416] Output: Request message (including text data)
[1417] Step 4:
[1418] The server receives the text data and analyzes it using a natural language processing engine.
[1419] Specific operation: The server receives the transmitted text data and inputs it into the generative AI model.
[1420] Input: Request message
[1421] Output: Intention analysis results of text data
[1422] Step 5:
[1423] Generative artificial intelligence generates specific operating procedures based on the analysis results.
[1424] Specific actions: Based on the command "Add a meeting tomorrow at 3 PM," it generates operational steps such as "Open the calendar app," "Add a new event," "Set the date," "Set the time," and "Enter a title."
[1425] Input: Intention analysis results of text data
[1426] Output: Specific operating instructions
[1427] Step 6:
[1428] The server sends the generated operation procedure to the terminal.
[1429] Specific operation: Sends operation instructions to the terminal as an HTTP response.
[1430] Input: Specific operating instructions
[1431] Output: Response message (including operation instructions)
[1432] Step 7:
[1433] The device will automatically perform operations within the calendar app based on the operating instructions.
[1434] Specific operation: In the calendar app, automatically "add a new event," "set the date," "set the time," and "enter a title."
[1435] Input: Response message (operation procedure)
[1436] Output: New event added to the Calendar app
[1437] In this way, a system is constructed in which a user's voice command is converted into a specific action via the terminal and the server, and is then automatically executed.
[1438] (Application example 1)
[1439] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1440] In today's diverse applications, it is cumbersome for users to perform multiple operations individually, and it is particularly time-consuming to efficiently complete an order in food delivery services. To solve this problem, a system is needed that automates a series of operations through voice commands and allows users to complete an order intuitively and efficiently.
[1441] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1442] In this invention, the server includes means for receiving a voice command, means for converting the voice command into text data, means for transmitting the text data to the server, means for analyzing the text data and using generative artificial intelligence to generate specific operation procedures, means for transmitting the operation procedures to a terminal, means for executing an operation on the terminal based on the operation procedures, and means for automating an ordering process to provide a meal delivery service, thereby enabling a user to efficiently and intuitively complete a meal order using voice commands, thereby improving the user experience.
[1443] A "voice command" is a voice instruction given by a user to a terminal.
[1444] "Text data" is voice commands converted into text information.
[1445] A "server" is a computer system that receives text data, analyzes it, and generates specific operating procedures.
[1446] "Generative AI" is an algorithm that understands and generates natural language, generating specific operating procedures from text data.
[1447] A "terminal" is a device through which a user inputs voice commands and executes generated operating procedures.
[1448] A "meal delivery service" is a service in which a user orders a meal and the order is delivered to a location specified by the user.
[1449] "Automating the ordering process" refers to the process of completing an order in a food delivery service based on voice commands without the need for manual intervention.
[1450] This invention relates to a system that allows users to easily and intuitively order food delivery services using voice commands. This system automates a series of operations from receiving a voice command to completing an order, improving the user experience.
[1451] Overall system configuration
[1452] The system consists of the following main components:
[1453] 1. A device that receives voice commands
[1454] 2. A voice recognition engine within the device that converts voice commands into text data
[1455] 3. A means of communication to send the converted text data to the server
[1456] 4. Generative AI in the server to analyze text data and generate specific operating procedures
[1457] 5. Response method from the server that sends the generated operation procedure to the terminal
[1458] 6. An automated operation system for executing actual operations based on operating procedures at the terminal
[1459] Program processing
[1460] Acquiring and translating voice commands:
[1461] The user inputs a voice command into the terminal. For example, the voice command is "Order a Margherita pizza." This terminal is equipped with a speech recognition engine (for example, Google Speech Recognition API), which converts the user's voice into text data in real time. This text data is sent to the server by a communication means to the server, which will be explained next.
[1462] Text data transmission and analysis:
[1463] When text data is sent from a device to a server, the server first receives the data. The server then analyzes the received text data and uses generative artificial intelligence (e.g., GPT-4) to understand the user's intent. Specifically, it uses natural language understanding technology to analyze the content of the text and generate the necessary operating procedures.
[1464] Generate and send operating instructions:
[1465] The generative AI then creates specific operational steps based on the analysis results. For example, if a command such as "Order a Margherita pizza" is entered, the generative AI will analyze this and generate a series of operational steps in a food delivery app, such as "Select a category," "Select a Margherita pizza," "Specify the quantity," and "Confirm the order."
[1466] To perform an action on the device:
[1467] The server generates instructions and sends them to the device, which then automatically performs the necessary operations within the food delivery app to complete the order. This eliminates the need for manual operations on the part of the user, allowing them to complete their meal order efficiently and intuitively.
[1468] Specific examples
[1469] For example, if a user enters the voice command "Order a Margherita pizza," the system automatically performs the following steps:
[1470] 1. Acquire voice command: The user enters a voice command.
[1471] 2. Speech recognition and text conversion: The device's speech recognition engine converts speech into text data.
[1472] 3. Send to server: The converted text data is sent to the server.
[1473] 4. Analysis by generative AI: The server's generative AI analyzes the text data and generates operating procedures.
[1474] 5. Sending operation instructions: The server sends the generated operation instructions to the terminal.
[1475] 6. Automated operation execution: The device follows the operation steps and automatically completes the order in the food delivery app.
[1476] An example prompt is:
[1477] When the voice command "Order a Margherita pizza" is entered, the system performs the following actions in the food delivery app:
[1478] 1. Select a category
[1479] 2. Choose a Margherita pizza
[1480] 3. Specify the quantity
[1481] 4. Confirm your order
[1482] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1483] Step 1:
[1484] Acquiring voice commands
[1485] The user inputs a voice command into the terminal. The terminal is equipped with a microphone that captures the user's voice in real time. For example, the user says, "Order a Margherita pizza." The input here is the user's voice command, and the output is voice data.
[1486] Step 2:
[1487] Speech recognition and text conversion
[1488] The device's speech recognition engine (for example, Google Speech Recognition API) converts the voice data into text data. This voice data is then analyzed and converted into text information. The input to this step is the voice data acquired in step 1, and the output is text data.
[1489] Step 3:
[1490] Sending to the server
[1491] The terminal sends the converted text data to the server. For this purpose, a network communication means is used. The input is the text data, and the output is the transmission result to the server. If the communication is successful, the text data reaches the server.
[1492] Step 4:
[1493] Text data analysis
[1494] The server analyzes the received text data. It uses generative artificial intelligence (e.g., GPT-4) and applies natural language understanding techniques to understand the user's intent. The input is the text data sent to the server, and the output is the extracted operation procedure.
[1495] Step 5:
[1496] Generate operating instructions
[1497] Generative AI generates specific operational steps based on the analysis results. For example, based on the text data "Order a Margherita pizza," it generates a series of operational steps such as "Select a category," "Select a Margherita pizza," "Specify quantity," and "Confirm the order." The input is the analysis results, and the output is the generated series of operational steps.
[1498] Step 6:
[1499] Sending operating instructions
[1500] The server sends the generated operation procedure to the terminal. To return this operation procedure to the terminal, a network communication means is used again. The input is the generated operation procedure, and the output is the transmission result to the terminal. If the communication is successful, the operation procedure arrives at the terminal.
[1501] Step 7:
[1502] Performing automated operations
[1503] The device automatically performs the necessary operations within the food delivery app based on the received operation instructions. For example, it automatically performs operations such as "select a category," "select Margherita pizza," "specify the quantity," and "confirm the order." The input is the operation instructions sent from the server, and the output is the final order completion. This eliminates the need for the user to perform each operation manually.
[1504] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1505] The present invention relates to a system that allows users to operate multiple different applications easily and intuitively using voice commands. Furthermore, the present invention provides a more personalized user experience by combining an emotion engine that recognizes the user's emotions.
[1506] Overall system configuration
[1507] The system consists of the following main components:
[1508] 1. A device that receives voice commands
[1509] 2. A voice recognition engine within the device that converts voice commands into text data
[1510] 3. A method for analyzing user emotions using an emotion engine
[1511] 4. A communication method for sending the converted text data and emotion data to the server
[1512] 5. Generative AI in the server that analyzes text data and emotion data and generates specific operating procedures
[1513] 6. Means for sending generated operating procedures and appropriate feedback to the terminal
[1514] 7. An automated operation system for executing actual operations based on operating procedures at the terminal
[1515] Program processing
[1516] Voice command capture and emotion analysis
[1517] The user inputs voice commands into the device. The device is equipped with a speech recognition engine and an emotion engine, which not only converts the user's voice into text data in real time, but also analyzes the user's emotional state using the emotion engine. For example, the device can analyze the user's emotions based on the tone, speed, and intonation of the voice and generate corresponding data.
[1518] Sending and analyzing text and emotion data
[1519] Text data and emotion data are sent from the device to a server. The server receives and analyzes this data. Specifically, it analyzes the text data using a Natural Language Understanding (NLU) engine and uses generative artificial intelligence to generate operating procedures based on the user's instructions and emotional state.
[1520] Generate and send operating instructions
[1521] The server generates specific operating procedures based on the analysis results. For example, if a user instructs, "Add a meeting tomorrow at 3 p.m.", the generative AI analyzes this and generates operating procedures that take into account the user's emotional state. If the user's emotional state is tense, the server selects concise and reliable operating procedures and adds reassuring feedback.
[1522] Execute operations on the device
[1523] The generated operation instructions and feedback are sent to the device. The device then automatically performs a series of operations based on this. In the case of a calendar app, the device performs operations such as "add a new event," "set the date," "set the time," and "enter a title," and finally notifies the user that the operation is complete. An appropriate message based on the user's emotions (e.g., "Is everything OK?") is also displayed.
[1524] Specific examples
[1525] Below is a specific example of adding an event to a calendar app.
[1526] Example 1: Adding an event in the calendar app
[1527] 1. The user types the voice command "Add a meeting tomorrow at 3 PM."
[1528] 2. The device uses a speech recognition engine to convert the voice into text data, and an emotion engine analyzes the user's emotional state from the voice (e.g., nervousness).
[1529] 3. The device sends the text data and emotion data to the server.
[1530] 4. The server analyzes the text data and uses generative artificial intelligence to generate specific operating instructions, including a reassuring message that takes into account the user's emotional state.
[1531] 5. The server sends the generated operating instructions and feedback to the terminal.
[1532] 6. The device automatically performs the necessary operations in the calendar app based on the user's instructions to add the meeting, and displays a message to the user such as "Meeting added successfully. Is there anything else we can help you with?"
[1533] In this way, by combining voice commands and an emotion engine, the present invention provides a personalized operating experience that adapts to the user's emotional state, enabling efficient and intuitive system operation.
[1534] The processing flow will be explained below.
[1535] Step 1:
[1536] The user enters a voice command into the device, for example, "Add a meeting for tomorrow at 3:00 PM."
[1537] Step 2:
[1538] The device activates a voice recognition engine and captures voice data in real time.
[1539] Step 3:
[1540] The device converts the voice data into text data, and the speech recognition engine generates the text "Add a meeting tomorrow at 3 PM."
[1541] Step 4:
[1542] The device activates an emotion engine and analyzes the user's emotional state from the voice data. For example, it determines that the user is in a tense state based on the tone and speed of the voice.
[1543] Step 5:
[1544] The device includes the text data and emotion data in the body of an HTTP request and prepares to send it to the server.
[1545] Step 6:
[1546] The device sends text data and emotion data to the server, which then sends an HTTP request to the server.
[1547] Step 7:
[1548] The server receives the HTTP request and extracts text data and emotion data from the request body.
[1549] Step 8:
[1550] The server analyzes the text data using a Natural Language Understanding (NLU) engine, extracting information such as "Date = tomorrow," "Time = 3 PM," and "Action = Add meeting."
[1551] Step 9:
[1552] The server calls the API of a generative artificial intelligence (e.g., GPT-3) and generates specific operating procedures based on the analysis results and emotional data. For users in a tense situation, the server generates concise operating procedures that provide a sense of security.
[1553] Step 10:
[1554] The server formats the generated operation instructions and feedback messages and creates an HTTP response to send to the device.
[1555] Step 11:
[1556] The server sends an HTTP response to the device.
[1557] Step 12:
[1558] The terminal receives the HTTP response and analyzes the operation procedure and feedback message from the response.
[1559] Step 13:
[1560] The device launches the calendar app. Specifically, it calls the calendar app based on the application ID.
[1561] Step 14:
[1562] Based on the operating procedure, the device will automatically perform a series of operations within the calendar app, such as "add a new event," "set the date," "set the time," and "enter a title."
[1563] Step 15:
[1564] The device notifies the user when the operation is complete, for example by displaying a reassuring message such as "Your meeting has been successfully added. You seem nervous. Is there anything else I can help you with?"
[1565] This process allows users to enjoy a smooth and personalized experience based on voice commands and emotional data.
[1566] Example 2
[1567] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1568] Conventional voice-operated systems execute commands without considering the user's emotional state, resulting in a uniform user experience. Furthermore, device operation may not always match the user's intuition, resulting in a lack of operability. Furthermore, the lack of proper feedback makes it difficult for users to confirm whether their commands were executed correctly.
[1569] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for acquiring a voice command, means for converting the voice command into text data, means for analyzing the emotional state when acquiring the voice command, means for transmitting the text data and emotional data to the server, means for analyzing the text data and using generative artificial intelligence to generate specific operating procedures based on the user's emotional state, means for transmitting the operating procedures and feedback messages to the terminal, and means for automatically executing operations on the terminal based on the operating procedures. This makes it possible to provide personalized operating procedures according to the user's emotional state, improving operability and providing a sense of security through feedback.
[1570] A "voice command" is an operation instruction given by voice to a terminal by a user.
[1571] "Text data" refers to data obtained by converting a voice command into character information using a voice recognition engine.
[1572] "Emotional state" refers to emotional information analyzed from the user's voice, and is obtained from tone, speed, and intonation.
[1573] An "emotion engine" is a software or hardware system for analyzing the emotional state of a user's voice.
[1574] A "server" is a central processing unit that receives data sent from a terminal and performs analysis and processing.
[1575] "Generative AI" is artificial intelligence software that analyzes received data and generates operating procedures based on the user's instructions and emotional state.
[1576] An "operation procedure" is a series of steps for operating a specific device or application that is executed based on a voice command.
[1577] A "feedback message" is a message provided to a user after an operation procedure is performed, and includes information indicating that the operation was successful and prompting the user to give the next instruction.
[1578] A "terminal" is an electronic device that allows a user to input voice commands and transmit data to a server, and includes smartphones, tablets, and the like.
[1579] The present invention relates to a system that allows users to easily and intuitively operate multiple different applications using voice commands. Furthermore, the present invention provides a more personalized user experience by combining an emotion engine that recognizes the user's emotions.
[1580] System configuration
[1581] This system consists of the following main components:
[1582] 1. A device that receives voice commands
[1583] 2. A voice recognition engine within the device that converts voice commands into text data
[1584] 3. A method for analyzing user emotions using an emotion engine
[1585] 4. A communication method for sending the converted text data and emotion data to the server
[1586] 5. Generative AI in the server that analyzes text data and emotion data and generates specific operating procedures
[1587] 6. Means for sending generated operating procedures and appropriate feedback to the terminal
[1588] 7. An automated operation system for executing actual operations based on operating procedures at the terminal
[1589] Voice command capture and emotion analysis
[1590] The user inputs voice commands into the device. The device is equipped with a speech recognition engine (e.g., Google Speech-to-Text) and an emotion engine (e.g., Affectiva) that converts the user's voice into text data in real time. The emotion engine also analyzes the user's emotional state from the voice. For example, the device generates data representing the user's emotion based on the tone, speed, and intonation of the voice.
[1591] Sending and analyzing text and emotion data
[1592] The device uses a communication method (Wi-Fi, 4G / 5G, etc.) to send the converted text data and emotion data to the server. The server receives and analyzes this data. Specifically, it analyzes the text data using a Natural Language Understanding (NLU) engine (e.g., IBM Watson NLU) and generates operational procedures based on the user's instructions and emotional state using generative artificial intelligence (e.g., OpenAI GPT-4).
[1593] Generate and send operating instructions
[1594] The server generates specific operating procedures based on the analysis results. For example, if a user instructs, "Add a meeting tomorrow at 3 p.m.", the generative AI analyzes this and generates operating procedures that take into account the user's emotional state. If the user's emotional state is tense, the server selects concise and reliable operating procedures and also generates a feedback message that provides a sense of security.
[1595] Execute operations on the device
[1596] The generated operation instructions and feedback are sent to the device, which then automatically performs a series of operations based on them. For example, when using a calendar app, the user performs operations such as "add a new event," "set the date," "set the time," and "enter a title." Finally, the device notifies the user that the operation is complete and displays an appropriate emotion-based message (e.g., "The meeting has been successfully added. Is there anything else we can help you with?").
[1597] Examples and prompts
[1598] Below is a specific example of adding an event to a calendar app.
[1599] Example 1: Adding an event in the calendar app
[1600] The user enters a voice command such as "Add a meeting tomorrow at 3pm." The device converts the voice to text data and uses an emotion engine to analyze the user's emotional state (e.g., nervousness). The device sends the text data and emotion data to the server. The server analyzes the text data and generates specific operating instructions using generative artificial intelligence. A reassuring message is also generated taking the emotional state into account. The server sends the generated operating instructions and feedback to the device. The device automatically performs the necessary operations in the calendar app based on the operating instructions and adds the meeting. The user is then shown a message such as "The meeting has been added successfully. Is there anything else we can help you with?"
[1601] Prompt Sentence Examples
[1602] "Please explain in detail how the system works when a user adds an event to a calendar app using voice commands."
[1603] In this way, by combining voice commands and an emotion engine, the present invention provides a personalized operating experience that adapts to the user's emotional state, enabling efficient and intuitive system operation.
[1604] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1605] Step 1:
[1606] The user inputs a voice command into the device. The input is the user's voice, and includes instructions such as "Add a meeting tomorrow at 3:00 PM." The device's microphone picks up the user's voice, and the voice data is input.
[1607] Step 2:
[1608] The device uses a speech recognition engine (e.g., speech recognition software) to convert voice data into text data. The input is the user's voice data, and the output is text instructions. Specifically, the voice waveform is analyzed and the spoken content is converted into text.
[1609] Step 3:
[1610] The device uses an emotion engine (e.g., emotion analysis algorithm) to analyze the user's emotional state. The input is voice data and text data, and the output is data indicating the user's emotional state (e.g., tension). Specifically, the tone, intonation, and speed of the voice are analyzed to generate emotion data.
[1611] Step 4:
[1612] The device sends text data and emotion data to the server. The input is text data and emotion data, and the output is data transmission to the server. Specifically, the data is sent via a network using a communication module (e.g., Wi-Fi or 4G / 5G).
[1613] Step 5:
[1614] The server analyzes the received text data and emotional data. The input is text data and emotional data, and the output is the analysis results. Specifically, it analyzes the text data using a Natural Language Understanding (NLU) engine (e.g., natural language analysis software), and generates operating procedures based on the user's instructions and emotional state using generative artificial intelligence (e.g., AI generation model).
[1615] Step 6:
[1616] The server sends the generated operation procedures and feedback messages to the terminal. The input is the operation procedures and feedback messages, and the output is data transmission to the terminal. Specifically, the server converts the data into packets and transmits them to the terminal through the communication module.
[1617] Step 7:
[1618] The device automatically executes the actual operation based on the operation instructions received. The input is the operation instructions, and the output is the result of the operation. Specifically, using automatic operation software (e.g., automation tool), a specified application (e.g., calendar app) is opened and operations such as "add a new event," "set the date," "set the time," and "enter a title" are performed.
[1619] Step 8:
[1620] The terminal notifies the user that the operation is complete. The input is the result of the operation and a feedback message, and the output is a notification to the user. Specifically, the screen displays "The meeting has been successfully added. Is there anything else we can help you with?"
[1621] (Application example 2)
[1622] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1623] Conventional systems using speech recognition technology simply convert voice commands into text data and perform specific operations without considering the user's emotional state, resulting in a uniform user experience that cannot fully meet the needs of individual users. Additionally, there is a lack of a mechanism for recommending content appropriate for a specific emotional state, which can lead to reduced user satisfaction.
[1624] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring a voice command, means for converting the voice command into text data, means for analyzing a user's emotions, means for transmitting the text data and emotional data to the server, means for analyzing the text data and emotional data and using generative artificial intelligence to generate a specific recommendation list, means for transmitting the recommendation list and feedback to a terminal, and means for providing content on the terminal based on the recommendation list. This enables personalized content recommendation that is adapted to the user's emotional state.
[1625] Definitions of important words
[1626] A "voice command" is an instruction given by a user to a device by voice.
[1627] "Text data" refers to data obtained by converting a voice command into a string of characters.
[1628] "User emotion" refers to the emotional state analyzed from the user's voice and language.
[1629] "Emotion data" refers to data that expresses the analysis results of a user's emotions as numerical values or categories.
[1630] A "server" is a computer system that receives, analyzes, and processes voice data and emotional data.
[1631] "Generative AI" refers to an AI technology that outputs content generated based on input data.
[1632] A "recommendation list" is a list of content and operating procedures generated based on the user's emotions and voice commands.
[1633] "Feedback" refers to the response or information that a system provides to a user.
[1634] "Content" refers to digital media and information such as music, movies, articles, etc.
[1635] MODE FOR CARRYING OUT THE INVENTION
[1636] The present invention is a system that allows users to operate multiple applications using voice commands and provides personalized operation and recommendations based on emotion analysis. The system receives voice commands, converts them into text data, analyzes the user's emotions, and uses generative artificial intelligence to generate appropriate operation procedures and recommendation lists.
[1637] Overall system configuration
[1638] The system consists of the following main components:
[1639] 1. Device receiving voice commands:
[1640] A device that allows users to input voice commands, such as a smartphone or smart speaker.
[1641] 2. Speech Recognition Engine:
[1642] The device's built-in voice recognition software converts the user's voice commands into text data, using, for example, Google's voice recognition API.
[1643] 3. Emotion Engine:
[1644] This is software or algorithm that analyzes the emotional state of a user's voice. It determines emotions based on the tone and intonation of the voice. For example, an emotion analysis API can be used.
[1645] 4. Means of communication:
[1646] A communication module for transmitting the converted text data and emotion data to a server, for example, by transmitting the data over an internet connection.
[1647] 5. Generative AI in the server:
[1648] The received text data and emotion data are analyzed to generate specific operation procedures and recommendation lists in response to user instructions. Examples of this include natural language understanding engines and generative AI models.
[1649] 6. Operating instructions and means of sending recommendation lists:
[1650] This technology allows the server to generate operating procedures and recommendation lists and send them to the terminal. The data is transmitted in real time via the Internet.
[1651] 7. How to perform operations on the terminal:
[1652] This function allows the device to automatically perform actual operations or provide feedback to the user based on the operating procedures and recommendation list sent from the server.
[1653] Program processing
[1654] 1. Voice command capture and sentiment analysis:
[1655] The user inputs a voice command, and the smartphone or smart speaker picks up the voice and converts it into text data using a speech recognition engine.
[1656] The emotion engine analyzes the tone, speed, and intonation of the voice to generate emotion data for the user. For example, if a user says "I'm tired," the emotion engine analyzes the voice to determine "fatigue."
[1657] 2. Transmission and analysis of text and emotion data:
[1658] Text and emotion data are sent from the device to the server, which analyzes the data and uses a Natural Language Understanding (NLU) engine to understand the user's intent.
[1659] 3. Generate and send operating instructions:
[1660] The server uses generative artificial intelligence to generate specific operation procedures and recommendation lists based on the user's voice commands and emotional data. For example, if the user requests, "Tell me about some uplifting movies," it generates a list of "uplifting movies" recommendations.
[1661] The generated operation procedures and recommendation lists are optimized to also respond to the user's emotions.
[1662] 4. Execute operations on the device:
[1663] The generated operation instructions and recommendation list are sent to the device, which then automatically performs specific operations based on them. For example, a movie recommendation app may list "Movie A" and "Movie B" as "uplifting movies" and display a message asking, "How about these movies?"
[1664] Specific examples
[1665] Specific examples of movie recommendations
[1666] When a user inputs a voice command such as "Tell me some uplifting movies," the device converts the voice into text data and analyzes the emotions using an emotion engine. The text data and emotion data are sent to a server, which uses generative artificial intelligence to generate a list of appropriate movie recommendations. The data is finally sent to the device and displayed as "Recommended movies are: Movie A, Movie B." In this way, by combining voice commands and emotion analysis technology, the present invention provides a personalized operating experience that adapts to the user's emotional state.
[1667] Prompt Sentence Examples
[1668] "Tell me some uplifting movies"
[1669] "Play some relaxing music"
[1670] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1671] Program processing steps
[1672] Step 1: Get voice commands
[1673] The user inputs a voice command into the device. Specifically, the user speaks to a smartphone or smart speaker, saying, "Tell me about an uplifting movie."
[1674] Input: User's voice command
[1675] Output: Captured audio data
[1676] Step 2: Voice Recognition
[1677] The device converts the audio data into text using a speech recognition engine, which uses software such as Google's speech recognition API to convert the audio to text.
[1678] Input: Audio data
[1679] Output: Text data
[1680] Step 3: Sentiment Analysis
[1681] The device uses an emotion engine to generate emotion data from the user's voice, analyzing the tone, speed, and intonation of the voice.
[1682] Input: Audio data
[1683] Output: Emotion data
[1684] Step 4: Send data
[1685] The terminal transmits the text data and emotion data to the server, and the data is sent to the server via the Internet using a communication means.
[1686] Input: Text data, emotion data
[1687] Output: Data sent to the server
[1688] Step 5: Data analysis
[1689] The server analyzes the received text and emotion data and uses a Natural Language Understanding (NLU) engine to understand the user's intent and emotions.
[1690] Input: Text data, emotion data
[1691] Output: Analysis results (information based on user intent and emotions)
[1692] Step 6: Generate operating instructions and recommendation list
[1693] The server uses generative artificial intelligence to generate specific operating procedures and recommendation lists based on the user's intentions and emotional data, such as a list of uplifting movies.
[1694] Input: Analysis results
[1695] Output: Generated instructions and recommendation list
[1696] Step 7: Submit instructions and list
[1697] The server transmits the generated operation procedure and recommendation list to the terminal, and transmits the data in real time using a transmission means.
[1698] Input: Operation procedure, recommendation list
[1699] Output: Data sent to the terminal
[1700] Step 8: Operation execution and feedback
[1701] Based on the operation instructions and recommendation list received from the server, the device automatically performs specific operations and provides feedback to the user, such as displaying "Recommended movies are: Movie A, Movie B."
[1702] Input: Operation procedure, recommendation list
[1703] Output: Displayed movie recommendation list, feedback message
[1704] This allows for personalized operation and recommendations based on the user's voice commands and emotional state.
[1705] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1706] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1707] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1708] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1709] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1710] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1711] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1712] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1713] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1714] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1715] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1716] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1717] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1718] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1719] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1720] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1721] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1722] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1723] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1724] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1725] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1726] The following is further disclosed regarding the above embodiment.
[1727] (Claim 1)
[1728] a means for obtaining voice commands;
[1729] means for converting the voice command into text data;
[1730] means for transmitting the text data to a server;
[1731] A means for analyzing the text data and using generative artificial intelligence to generate specific operating procedures;
[1732] means for transmitting the operation procedure to a terminal;
[1733] A system including means for executing an operation based on the operating procedure at the terminal.
[1734] (Claim 2)
[1735] The system according to claim 1, wherein the generative artificial intelligence performs natural language understanding and generation.
[1736] (Claim 3)
[1737] 2. The system of claim 1, wherein the operational procedure includes automatic operation of a user interface.
[1738] "Example 1"
[1739] (Claim 1)
[1740] a means for obtaining voice commands;
[1741] means for converting the voice command into text data;
[1742] means for transmitting the text data to a server;
[1743] A means for analyzing the text data and using generative artificial intelligence to generate specific operating procedures;
[1744] means for transmitting the operation procedure to a terminal;
[1745] means for executing an operation based on the operation procedure at the terminal;
[1746] a means for generating a specific operation procedure including an event addition operation of a calendar application;
[1747] A system including a means for transmitting the operating procedure from a server to a terminal and automatically executing the procedure.
[1748] (Claim 2)
[1749] The system of claim 1, wherein the generative artificial intelligence performs natural language understanding and generation, and generates operating procedures for multiple applications including a calendar application.
[1750] (Claim 3)
[1751] The system according to claim 1, wherein the operation procedure includes automatic operation of a user interface, and realizes automatic operation based on voice commands by utilizing a voice recognition engine, communication means, and generative artificial intelligence.
[1752] "Application Example 1"
[1753] Rewriting of claims
[1754] (Claim 1)
[1755] a means for obtaining voice commands;
[1756] means for converting the voice command into text data;
[1757] means for transmitting the text data to a server;
[1758] A means for analyzing the text data and using generative artificial intelligence to generate specific operating procedures;
[1759] means for transmitting the operation procedure to a terminal;
[1760] means for executing an operation based on the operation procedure at the terminal;
[1761] A system including means for automating an ordering process for providing a meal delivery service.
[1762] (Claim 2)
[1763] The system according to claim 1, wherein the generative artificial intelligence performs natural language understanding and generation.
[1764] (Claim 3)
[1765] 10. The system of claim 1, wherein the operating procedure automatically operates a user interface of a meal delivery application.
[1766] "Example 2: Combining Emotion Engines"
[1767] (Claim 1)
[1768] a means for obtaining voice commands;
[1769] means for converting the voice command into text data;
[1770] means for analyzing an emotional state when receiving the voice command;
[1771] means for transmitting the text data and emotion data to a server;
[1772] means for analyzing the text data and using generative artificial intelligence to generate specific operating procedures based on the emotional state of the user;
[1773] means for transmitting the operation procedure and feedback message to a terminal;
[1774] A system including means for automatically executing an operation based on the operating procedure at the terminal.
[1775] (Claim 2)
[1776] 2. The system according to claim 1, wherein the generative artificial intelligence performs natural language understanding and generation, and generates a message adapted to the user's emotional state.
[1777] (Claim 3)
[1778] 10. The system of claim 1, wherein the operational procedure includes automatic operation of a user interface and also provides feedback messages to the user.
[1779] "Application example 2 when combining emotion engines"
[1780] Rewritten claims
[1781] (Claim 1)
[1782] a means for obtaining voice commands;
[1783] means for converting the voice command into text data;
[1784] means for analyzing user emotions;
[1785] means for transmitting the text data and emotion data to a server;
[1786] means for analyzing the text data and emotion data and utilizing generative artificial intelligence to generate a specific recommendation list;
[1787] means for transmitting the recommendation list and feedback to a terminal;
[1788] The system includes means for providing content based on the recommendation list at the terminal.
[1789] (Claim 2)
[1790] The system according to claim 1, wherein the generative artificial intelligence performs natural language understanding and generation.
[1791] (Claim 3)
[1792] 10. The system of claim 1, wherein the recommendation list includes personalized content. [Explanation of symbols]
[1793] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a means for obtaining voice commands; means for converting the voice command into text data; means for transmitting the text data to a server; A means for analyzing the text data and using generative artificial intelligence to generate specific operating procedures; means for transmitting the operation procedure to a terminal; A system including means for executing an operation based on the operating procedure at the terminal.
2. The system according to claim 1 , wherein the generative artificial intelligence performs natural language understanding and generation.
3. 2. The system of claim 1, wherein the operational procedure includes automatic operation of a user interface.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A