system
A voice-based system for electronic devices facilitates easy operation and troubleshooting by converting voice inputs to text, analyzing, and generating voice responses, addressing the challenges of text-based interfaces for users with text input difficulties or visual impairments.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-01
- Publication Date
- 2026-04-13
AI Technical Summary
Modern electronic devices and software interfaces are difficult for users who are not good at text input or visually impaired users, lacking means for easy operation and troubleshooting.
A system that allows users to input questions and operating instructions by voice, converting voice data to text, analyzing the text to generate responses, and converting back to voice for playback, using natural language processing and speech synthesis technologies.
Enables users to perform operations and troubleshooting easily without text input, providing quick and accurate voice responses tailored to user needs.
Smart Images

Figure 2026063796000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Many modern electronic devices and software provide text-based interfaces for users to operate, but these interfaces are difficult to use for users who are not good at text input or visually impaired users. For such users, there is a lack of means to assist in operation procedures and troubleshooting. Therefore, there is a need for a system that enables a wide range of users, including those who are not good at text input and visually impaired users, to perform operations and problem-solving more easily.
Means for Solving the Problems
[0005] This invention provides a system in which a user inputs questions and operating instructions by voice, and receives an answer in voice. Specifically, the system configuration includes a voice input means, a means for transmitting voice data to a server, a means for converting voice data into text data, a means for analyzing the text data to generate an answer, a means for converting the generated answer into voice data, a means for transmitting the voice data, and a means for playing back the voice data. This allows the user to receive assistance with operating procedures and troubleshooting without having to input text. Furthermore, this system generates answers using data from instruction manuals and overview booklets, and provides appropriate and prompt answers to the user by performing advanced text analysis using a natural language processing model.
[0006] "Voice input means" refers to a device or system that captures the user's voice as digital voice data.
[0007] "Means for sending audio data to a server" refers to a device or system for sending captured audio data to a server via a network.
[0008] "Means for converting audio data to text data" refers to a device or system that uses speech recognition technology on a server or other processing device to convert audio data to corresponding text data.
[0009] "Means for analyzing text data and generating responses" refers to a device or system that uses natural language processing technology to analyze text data and generate appropriate responses.
[0010] "Means for converting generated responses into audio data" refers to a device or system that converts responses generated as text data into audio data using speech synthesis technology.
[0011] "Means for transmitting audio data to a terminal" refers to a device or system for transmitting audio data generated from a server to a terminal via a network.
[0012] "Means for playing audio data" refers to a device or system for playing audio data to a user through an audio output device.
[0013] "Instruction manuals and overview booklets" refer to documents and data that contain information on how to use a product or system and how to troubleshoot it.
[0014] A "natural language processing model" refers to an algorithm or data model that constitutes part of artificial intelligence technology used to analyze text data and generate responses. [Brief explanation of the drawing]
[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11]It is a sequence diagram showing the processing flow of the data processing system in Embodiment 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.
Mode for Carrying Out the Invention
[0016] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described according to the attached drawings.
[0017] First, the language used in the following description will be explained.
[0018] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0019] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0020] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0023] [First Embodiment]
[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0036] The present invention will now describe embodiments for carrying out this invention. The present invention is a system in which a user can give operating instructions or ask questions using their voice and receive answers in voice. The system comprises a voice input means, a means for transmitting voice data to a server, a means for converting voice data into text data, a means for analyzing the text data to generate answers, a means for converting the generated answers into voice data, a means for transmitting voice data to a terminal, and a means for playing back the voice data.
[0037] System Configuration
[0038] 1. Voice input means
[0039] The user provides operating instructions and asks questions via voice to the terminal. This voice input method uses voice input devices such as microphones. In addition, voice recognition software is incorporated to capture the voice as digital data.
[0040] 2. Means for sending audio data to the server
[0041] The device sends the captured audio data to a cloud server or local server. This transmission takes place via an internet connection or local network.
[0042] 3. Means for converting audio data into text data
[0043] The server uses a speech recognition API (e.g., a cloud-based speech recognition service) to convert the received audio data into text data. This ensures that instructions and questions entered by voice are represented as text.
[0044] 4. Means for analyzing text data to generate responses
[0045] The server analyzes text data using a natural language processing model (e.g., pre-trained models such as BERT or GPT) to understand the intent behind the user's questions and instructions. It then refers to a specific database (e.g., instruction manuals or overview booklets) to generate appropriate answers to the questions.
[0046] 5. Means for converting the generated response into audio data
[0047] The generated responses are converted into audio data using a speech synthesis API (e.g., a cloud-based speech synthesis service). This provides the user with an audio response.
[0048] 6. Means for transmitting audio data to a terminal
[0049] The server sends the generated audio data to the terminal. This transmission takes place via an internet connection or a local network.
[0050] 7. Means for playing audio data
[0051] The device plays the received audio data. This allows the user to hear the response in audio. Audio output devices such as speakers or headphones are used.
[0052] Specific example
[0053] For example, if a user asks "The printer ink isn't coming out," the following process will be performed.
[0054] The user speaks into the voice input device and says, "The printer ink isn't coming out."
[0055] The device captures this audio data and sends it to the server.
[0056] The server converts the audio data into text data using a speech recognition API.
[0057] The text data is analyzed using a natural language processing model to search for relevant sections in the instruction manual related to the problem "the printer ink is not coming out."
[0058] The server generates the appropriate response from the instruction manual, for example, "You need to replace the ink cartridge. Open the printer cover, remove the old cartridge, and insert the new cartridge."
[0059] Convert this answer into speech data using a speech synthesis API.
[0060] Audio data is sent to the terminal, which then plays it back to the user through an audio output device.
[0061] In this way, users can ask questions and receive answers by voice without having to type. This provides an easy-to-use support system for users who have difficulty typing or who are visually impaired.
[0062] The following describes the processing flow.
[0063] Step 1:
[0064] The user inputs operating instructions or questions by voice into the terminal. For example, the user might say, "The printer ink isn't coming out."
[0065] Step 2:
[0066] The device uses its built-in microphone to capture the user's voice and temporarily stores it as digital audio data.
[0067] Step 3:
[0068] The device sends the captured audio data to the server. This transmission takes place via an internet connection.
[0069] Step 4:
[0070] The server sends the received audio data to a speech recognition API. Typically, the speech recognition API uses a cloud-based speech recognition service.
[0071] Step 5:
[0072] The server processes the text data received from the speech recognition API. For example, the text "The printer ink isn't coming out" is generated.
[0073] Step 6:
[0074] The server uses a natural language processing model to analyze text data. This model performs analysis to understand the user's intent and find relevant information.
[0075] Step 7:
[0076] The server searches a database of instruction manuals and summary booklets to obtain appropriate answers to user questions. For example, this might include questions such as "Providing instructions on how to replace ink cartridges."
[0077] Step 8:
[0078] The server generates answers in a user-friendly format based on the relevant section of the instruction manual. For example, it might generate text such as, "To replace the ink cartridge, first open the printer cover, remove the old cartridge, and then insert the new cartridge."
[0079] Step 9:
[0080] The server generates text responses, which are then sent to a speech synthesis API to be converted into audio data. This speech synthesis API also often uses a cloud-based service.
[0081] Step 10:
[0082] The server sends the generated audio data to the terminal. This transmission also takes place via the internet connection.
[0083] Step 11:
[0084] The device plays the audio data it received from the server. Using the device's speaker or a headset or other audio output device, it provides the user with an audio response. The message played is: "To replace the ink cartridge, first open the printer cover, remove the old cartridge, and then insert the new cartridge."
[0085] Through the steps described above, this system responds to user voice input with voice commands. Therefore, it is easy to use for users who have difficulty with text input or those with visual impairments.
[0086] (Example 1)
[0087] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0088] In recent years, voice-based interfaces have attracted attention, becoming a particularly important technology for users with visual impairments or those who have difficulty with text input. However, current voice response systems face several challenges. For example, accurately converting voice input data into text and then quickly analyzing it to generate appropriate responses is difficult. Furthermore, there is a need to convert the generated responses into natural and easy-to-understand speech. In addition, optimization is required to perform these processes in real time. This invention aims to solve these problems and provide quick and accurate voice responses to user questions and operational instructions.
[0089] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0090] In this invention, the server includes means for understanding the intent of questions and operational instructions by analyzing voice data into text data and the text data; means for generating appropriate answers from a database based on the analysis results; and means for converting the generated answer text into voice data. This makes it possible to quickly and accurately perform text analysis and answer generation for questions and instructions entered by the user via voice, and to provide those answers as natural and easy-to-understand voice.
[0091] "Voice input means" refers to a device that allows users to ask questions or give instructions using their voice.
[0092] "Means for sending audio data to a server" refers to the technology for sending captured audio data to a server via a network.
[0093] "Means of converting audio data to text data" refers to technology that converts captured audio data into text format.
[0094] "Methods for analyzing text data and generating responses" refers to technologies that analyze data expressed as text and generate appropriate responses to user questions or instructions.
[0095] "Means of converting generated responses into audio data" refers to technology that converts text-based responses into audio data.
[0096] "Means for transmitting audio data to a terminal" refers to the technology for transmitting generated audio data to a user's terminal via a network.
[0097] "Means of playing audio data" refers to the devices and software used to play audio data on the user's device.
[0098] "Means of converting to digital audio data in real time" refers to technology that converts user voice input into digital data in real time.
[0099] "Text data and means for analyzing that text data" refers to technologies for analyzing converted text data and understanding its intent and content.
[0100] "Means for generating appropriate answers from a database" refers to technologies that refer to a database based on analysis results and extract appropriate answers to user questions and operational instructions.
[0101] "Methods for converting response text into audio data" refers to technologies that convert generated responses into natural and easy-to-understand audio data.
[0102] This invention is a system that allows users to give operating instructions and ask questions using their voice and receive answers in voice. This system combines a voice input device, a network, a server, a natural language processing model, and speech synthesis technology.
[0103] Voice input method
[0104] The user uses a microphone to speak into the device, asking questions or giving instructions. The device is equipped with speech recognition software such as Google® Speech-to-Text API, which converts the voice into digital audio data in real time. For example, if the user says "The printer ink isn't coming out," this voice is captured as digital data.
[0105] Sending audio data
[0106] The device transmits the captured audio data to the server via an internet connection or local network. This transmission process uses encryption technologies such as TLS to ensure data security.
[0107] Converting audio data to text
[0108] The server receives the incoming audio data and uses the Amazon Transcribe API to convert the audio into text data. For example, the user's audio data, "The printer ink isn't coming out," is converted into the text format "The printer ink isn't coming out."
[0109] Text data analysis
[0110] The server analyzes the converted text data using natural language processing models (e.g., GPT-3®, BERT) to understand the intent behind questions and instructions. For example, the server analyzes the text data "The printer ink isn't coming out" and understands its meaning. This process includes generating text tokens and interpreting intent.
[0111] Answer generation
[0112] Based on the analysis results, the server extracts and generates the appropriate answer from a database (e.g., a printer instruction manual database). For example, it might generate an answer such as, "You need to replace the ink cartridge. Open the printer cover, remove the old cartridge, and insert the new cartridge."
[0113] Voice conversion of the answer
[0114] The generated response text is converted into speech data using speech synthesis technology such as the Google Text-to-Speech API. This ensures that the user's requested response is produced as natural-sounding speech data.
[0115] Sending audio data
[0116] The server then transmits the generated audio data back to the terminal via the internet or local network. Security is ensured during this process by encrypting the data using TLS or similar protocols.
[0117] Playback of audio data
[0118] The device plays the received audio data through an audio output device such as a speaker or headphones. The user can then listen to this audio and take the necessary actions or responses. For example, the device's speaker might play an audio message saying, "You need to replace the ink cartridge. Open the printer cover, remove the old cartridge, and insert the new cartridge."
[0119] Specific example
[0120] For example, if a user asks "The printer ink isn't coming out," the following process will be performed.
[0121] The user speaks into the voice input device and says, "The printer ink isn't coming out."
[0122] The device captures this audio data and sends it to the server.
[0123] The server converts the audio data into text data using a speech recognition API.
[0124] The text data is analyzed using a natural language processing model to search for relevant sections in the instruction manual related to the problem "the printer ink is not coming out."
[0125] The server generates the appropriate response from the instruction manual, for example, "You need to replace the ink cartridge. Open the printer cover, remove the old cartridge, and insert the new cartridge."
[0126] Convert this answer into speech data using a speech synthesis API.
[0127] Audio data is sent to the terminal, which then plays it back to the user through an audio output device.
[0128] Example of a prompt
[0129] User input: "The printer ink isn't coming out."
[0130] System response: "You need to replace the ink cartridge. Open the printer cover, remove the old cartridge, and insert the new one."
[0131] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0132] Step 1: Capture voice input
[0133] The user speaks into a voice input device (e.g., a microphone) to ask a question or give an instruction. The device converts this voice into digital speech data using the Google Speech-to-Text API. The input is the user's utterance, "The printer ink isn't coming out," and the output is the text data of this utterance. Specifically, the device captures the user's voice in real time and converts it into a digital signal.
[0134] Step 2: Sending the audio data
[0135] The terminal sends the generated digital audio data to the cloud server using encryption technology such as TLS. The input is digital audio data, and the output is the transmission of encrypted data. Specifically, the terminal sends the data to the server via an internet connection or a local network.
[0136] Step 3: Converting audio data to text
[0137] The server converts received audio data into text data using the Amazon Transcribe API. The input is digital audio data, and the output is the corresponding text data. Specifically, the server receives audio data and makes an API request to convert it into text data.
[0138] Step 4: Analyzing Text Data
[0139] The server receives text data and performs analysis using a natural language processing model (e.g., BERT, GPT-3). The input is text data, and the output is semantic understanding data resulting from the analysis. Specifically, the server tokenizes the text and uses the model to analyze the intent of questions and instructions.
[0140] Step 5: Generating the answer
[0141] Based on the analysis results, the server extracts and generates appropriate answers from a database (e.g., an instruction manual database). The input is semantic comprehension data, and the output is the generated answer text. Specifically, the server searches the queried database, extracts matching answers, and assembles them.
[0142] Step 6: Voice conversion of the answer
[0143] The generated response text is converted into speech data using the Google Text-to-Speech API. The input is the response text, and the output is the corresponding speech data. Specifically, the server makes an API request to convert the text to speech.
[0144] Step 7: Sending audio data
[0145] The server re-encrypts the generated audio data using TLS or similar methods and sends it to the terminal. The input is the audio data, and the output is the transmission of the encrypted data. Specifically, the server sends the data back to the terminal via the network.
[0146] Step 8: Play back audio data
[0147] The device decodes the received audio data and plays it back through an audio output device such as a speaker or headphones. The input is audio data, and the output is the audio that the user hears. Specifically, the device decodes the encryption and outputs the audio using an audio playback device.
[0148] (Application Example 1)
[0149] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0150] Traditional food delivery systems require users to place orders via a screen, which presents challenges for users with dirty hands or those with visual impairments. Furthermore, checking order status and estimated delivery times also requires manual operation, resulting in a lack of convenience.
[0151] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0152] In this invention, the server includes a voice input means, a means for transmitting voice data to the server, a means for converting voice data into text data, a means for analyzing the text data to generate a response corresponding to the order, a means for converting the generated response into voice data, a means for transmitting voice data to a terminal, and a means for playing back the voice data. This allows users to order food and check the order status using only their voice, improving convenience for people with visual impairments and users whose hands are occupied.
[0153] A "voice input method" refers to a device or system that allows a user to give instructions by voice.
[0154] "Means for sending audio data to a server" refers to the mechanisms and protocols for sending captured audio data to a server via the internet.
[0155] "Means for converting audio data to text data" refers to technologies for converting audio data captured using speech recognition technology into text format.
[0156] "A means of analyzing text data to generate responses corresponding to orders" refers to a system that uses natural language processing technology to analyze text data and generate appropriate responses from the obtained information.
[0157] "Means for converting generated responses into audio data" refers to a mechanism or system that converts text-based responses into audio data using speech synthesis technology.
[0158] "Means for transmitting audio data to a terminal" refers to the mechanisms and protocols for transmitting generated audio data to a terminal used by the user.
[0159] "Means for playing audio data" refers to devices or systems that play transmitted audio data and allow the user to listen to it.
[0160] "Data related to food and beverage orders" refers to data including menus, prices, and restaurant information handled within the food delivery system.
[0161] A "natural language processing model" is a machine learning technique or algorithm used to analyze user text data, understand user intent, and generate appropriate responses.
[0162] "Responses regarding orders and delivery status" refers to responses that include information about the products ordered by the user and their delivery status.
[0163] The system for carrying out this invention allows users to place food delivery orders and check the order status using voice commands. The system includes the following main components:
[0164] 1. Voice input means
[0165] Users input voice commands, such as ordering food or checking delivery status, by speaking into the microphone of their device (smartphone or smart glasses). This voice input method utilizes the built-in microphone of the smartphone or smart glasses, and captures the voice as digital data using the Google Speech-to-Text API.
[0166] 2. Means for sending audio data to the server
[0167] The device sends the captured audio data to a cloud server via an internet connection. HTTPS communication is used for data transmission.
[0168] 3. Means for converting audio data into text data
[0169] The server converts the received audio data into text data using speech recognition technology. Specifically, it uses the Google Speech-to-Text API to convert the audio data into text format.
[0170] 4. A means of analyzing text data to generate responses corresponding to orders.
[0171] The server analyzes the data, which has been converted to text format, using a natural language processing model. It utilizes advanced pre-trained models such as GPT-4® to understand user intent and generate appropriate responses. During this process, it references a food delivery database, providing data such as menu information, pricing information, and estimated delivery times.
[0172] 5. Means for converting the generated response into audio data
[0173] The server converts the generated text-based responses into speech data using speech synthesis technology. This conversion utilizes the Google Text-to-Speech API.
[0174] 6. Means for transmitting audio data to a terminal
[0175] The server then sends the generated audio data back to the terminal. This is also done via HTTPS communication over the internet connection.
[0176] 7. Means for playing audio data
[0177] Finally, the device plays back the received audio data and provides the user with an answer. A smartphone, smart glasses speaker, or headphones are used as the playback device.
[0178] Adding specific examples
[0179] For example, if a user uses voice input to say "I want to order a pizza," the following process will occur.
[0180] The user says, "I want to order a pizza."
[0181] The device captures this audio and sends it to the server.
[0182] The server uses the Google Speech-to-Text API to convert the audio data into text data.
[0183] A natural language processing model (e.g., GPT-4) is used to analyze text data such as "I want to order a pizza" and provide information on appropriate restaurants and menus.
[0184] The server generates an answer such as "Which restaurant would you order pizza from?", and this answer is converted into speech data using the Google Text-to-Speech API.
[0185] The server generates audio data and sends it to the terminal, which then plays the audio.
[0186] Example of a prompt
[0187] Analyze user comments and suggest corresponding food delivery restaurants and menus: {User comments}
[0188] Follow user instructions, confirm the order, and communicate with the backend to check its status.
[0189] In this way, users can perform a series of operations, from ordering food delivery to checking the delivery status, using only their voice, without using a visual interface. This significantly improves convenience for users whose hands are full or who have visual impairments.
[0190] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0191] Step 1:
[0192] The user speaks into the device's voice input and says, "I want to order a pizza." This captures the audio data. The input is what the user said, and the output is the captured audio data.
[0193] Step 2:
[0194] The device sends the captured audio data to the cloud server via the internet connection. The input is the captured audio data, and the output is the audio data received on the server side. Specifically, the operation involves sending the audio data using HTTPS communication.
[0195] Step 3:
[0196] The server converts the received audio data into text data using the Google Speech-to-Text API. The input is the received audio data, and the output is text data. Specifically, it analyzes the audio data using speech recognition technology and converts it into the corresponding text.
[0197] Step 4:
[0198] The server analyzes the converted text data using a natural language processing model (e.g., GPT-4). The input is text data, and the output is the user's intent or question content as a result of the analysis. Specifically, it analyzes the text using a generative AI model and understands the content of the corresponding instructions or questions.
[0199] Step 5:
[0200] The server generates appropriate restaurant and menu information based on the analysis results and creates a text-based response. The input is the analysis results, and the output is the generated text-based response. Specifically, it refers to an internal restaurant database, extracts appropriate information to meet the user's request, and generates a response.
[0201] Step 6:
[0202] The server converts the generated text-based response into audio data using the Google Text-to-Speech API. The input is a text-based response, and the output is audio data. Specifically, it uses speech synthesis technology to convert text into speech.
[0203] Step 7:
[0204] The server sends the generated audio data back to the terminal. The input is the generated audio data, and the output is the audio data received by the terminal. Specifically, the operation involves sending the audio data to the terminal using HTTPS communication.
[0205] Step 8:
[0206] The device plays the received audio data and provides a response to the user. The input is the received audio data, and the output is the audio information to be heard by the user. Specifically, the device plays the audio data using its speaker or headphones.
[0207] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0208] The embodiments for carrying out this invention will be described in detail. The present invention combines an emotion recognition function with a system in which a user inputs operation instructions or questions by voice and receives answers by voice. The system comprises a voice input means, a means for transmitting voice data to a server, a means for converting voice data into text data, a means for analyzing the text data to generate answers, a means for converting the generated answers into voice data, a means for transmitting voice data to a terminal, a means for playing back the voice data, and an emotion engine that recognizes the user's emotions.
[0209] System Configuration
[0210] 1. Voice input means
[0211] The user inputs operating instructions or questions by voice into the terminal. For example, the user might say, "The printer ink isn't coming out." This voice input method uses a microphone to capture the voice data and saves it as digital audio data.
[0212] 2. Means for sending audio data to the server
[0213] The device sends audio data captured via voice input to a cloud server or local server. This transmission takes place via an internet connection.
[0214] 3. Means for converting audio data into text data
[0215] The server uses a speech recognition API to convert the received audio data into text data. The speech recognition API generates the text "The printer ink isn't coming out."
[0216] 4. Means for analyzing text data to generate responses
[0217] The server uses a natural language processing model to analyze text data and understand the intent behind user questions and instructions. It then consults a database of instruction manuals and overview booklets to generate appropriate answers to the questions.
[0218] 5. Means for converting the generated response into audio data
[0219] The server converts the generated response into audio data using a speech synthesis API. This speech synthesis API then prepares the response for the user as audio.
[0220] 6. Means for transmitting audio data to a terminal
[0221] The server sends the generated audio data to the terminal. This transmission also takes place via the internet connection.
[0222] 7. Means for playing audio data
[0223] The device plays the received audio data. It then provides the user with the answer via voice using a speaker or headphones.
[0224] 8. Emotional Engine
[0225] The emotion engine analyzes and recognizes the user's emotional state using voice data. This emotion analysis is based on features such as voice tone, speed, and intonation.
[0226] Specific example
[0227] For example, the sequence of operations when a user asks "The printer ink isn't coming out" is as follows:
[0228] The user speaks into the voice input device and says, "The printer ink isn't coming out."
[0229] The device captures this audio and sends the audio data to the server.
[0230] The server converts the audio data into text data using a speech recognition API.
[0231] The server uses a natural language processing model to analyze text data and search for information related to the problem "the printer ink isn't coming out."
[0232] The server generates appropriate responses to the user based on the instruction manual. For example, it might generate text such as, "To replace the ink cartridge, first open the printer cover, remove the old cartridge, and then insert the new cartridge."
[0233] The text response generated by the server is converted into speech data using a speech synthesis API.
[0234] The server sends the audio data to the terminal.
[0235] The device plays audio data and provides the user with an answer.
[0236] The emotion engine analyzes the user's original voice data, and if it determines that the user is confused, the server modifies the response to be more helpful and detailed. For example, it might adjust the tone and content of the response to something like, "Don't worry, first try opening the printer cover. Next, carefully remove the old ink cartridge. Then, firmly insert the new cartridge."
[0237] Through these steps, the system can provide flexible responses tailored to the user's emotional state. This allows users to use the system more comfortably and effectively.
[0238] The following describes the processing flow.
[0239] Step 1:
[0240] The user inputs operating instructions or questions by voice into the terminal. For example, the user might say, "The printer ink isn't coming out."
[0241] Step 2:
[0242] The device uses its built-in microphone to capture the user's voice and temporarily stores it as digital audio data.
[0243] Step 3:
[0244] The device sends the captured audio data to the server. This transmission takes place via an internet connection.
[0245] Step 4:
[0246] The server sends audio data to a speech recognition API, which then converts the audio data into text data. For example, a cloud-based speech recognition service could be used.
[0247] Step 5:
[0248] The server processes the text data "Printer ink is not coming out" received from the speech recognition API.
[0249] Step 6:
[0250] The server uses a natural language processing model to analyze the text data. This analysis helps the server understand what the user wants. For example, it might determine that the user wants to solve a printer problem.
[0251] Step 7:
[0252] The server searches a database of instruction manuals and summary booklets to obtain appropriate answers to user questions. For example, it might find information such as, "The ink cartridge needs to be replaced."
[0253] Step 8:
[0254] The server generates appropriate responses for the user based on the instruction manual. For example, it might generate text such as, "To replace the ink cartridge, first open the printer cover, remove the old cartridge, and then insert the new cartridge."
[0255] Step 9:
[0256] The server sends the generated text response to a speech synthesis API, which converts it into audio data. This speech synthesis API then prepares the response for the user as audio.
[0257] Step 10:
[0258] The server sends the generated audio data to the terminal. This transmission also takes place via the internet connection.
[0259] Step 11:
[0260] The device plays the audio data it received from the server. Using the device's speaker or headphones, it provides the user with an audio response. The message played is: "To replace the ink cartridge, first open the printer cover, remove the old cartridge, and then insert the new cartridge."
[0261] Step 12:
[0262] The emotion engine analyzes the original audio data and evaluates the user's emotional state. For example, it can detect feelings of confusion or frustration.
[0263] Step 13:
[0264] The server adjusts the responses generated based on the emotion engine's analysis. For example, if the server detects that the user is confused, it adjusts the response to be more helpful and detailed. It might change to something like, "Don't worry, first try opening the printer cover. Next, carefully remove the old ink cartridge. Then, firmly insert the new cartridge."
[0265] Step 14:
[0266] The server sends the newly processed audio data back to the speech synthesis API, where it is converted back into audio data.
[0267] Step 15:
[0268] The server sends new audio data to the device, and the device plays this data. The user can then hear the corrected, gentler-toned response.
[0269] (Example 2)
[0270] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0271] Conventional voice dialogue systems can provide appropriate answers to user inputs, such as operational instructions and questions, but they struggle to provide flexible responses that adapt to the user's emotional state. Furthermore, they lacked the means to recognize and appropriately respond to user feelings of confusion or anxiety. As a result, the user experience was not always optimal.
[0272] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0273] In this invention, the server includes means for converting voice data into text data, means for analyzing the text data to generate a response, and means for converting the generated response back into voice data. This enables flexible response generation that takes into account the emotional state of the user based on voice input.
[0274] "Voice input means" refers to a device that includes a microphone for the user to input voice, and a means for capturing that voice as digital data.
[0275] "Means for sending audio data to a server" refers to means for sending captured audio data to a cloud server or local server via the internet.
[0276] "Means for converting audio data to text data" refers to methods for converting audio data received using a speech recognition API into text data.
[0277] "Methods for analyzing text data and generating responses" refers to methods for analyzing text data using natural language processing models and generating appropriate responses.
[0278] "Means for converting generated responses into audio data" refers to means for converting text responses generated using a speech synthesis API into audio data.
[0279] "Means for transmitting audio data to a terminal" refers to means for transmitting generated audio data to a user's terminal via the internet.
[0280] "Means for playing audio data" refers to the means by which a terminal plays back audio data it has received through speakers or headphones.
[0281] An "emotion engine" is a means of analyzing and recognizing a user's emotional state using voice data and responding accordingly.
[0282] A "natural language processing model" is an algorithm or software for analyzing text data and understanding the meaning and intention of human language.
[0283] An "audio recognition API" is an application programming interface for converting audio data into text data.
[0284] An "audio synthesis API" is an application programming interface for converting text data into audio data.
[0285] The embodiments for implementing this invention will be described in detail. This invention is a system including a voice input means, a means for transmitting voice data to a server, a means for converting voice data into text data, a means for analyzing text data to generate an answer, a means for converting the generated answer into voice data, a means for transmitting voice data to a terminal, a means for playing back voice data, and an emotion engine for recognizing the user's emotion.
[0286] System Configuration
[0287] 1. Voice Input Means
[0288] The user inputs operation instructions and questions by voice towards the terminal. For example, the user says "The printer doesn't print." This voice input means captures voice data using a microphone and stores it as digital voice data. As specific hardware, smartphones, tablet terminals, PCs with microphones, etc. can be used.
[0289] 2. Means for Transmitting Voice Data to a Server
[0290] The terminal transmits the voice data captured by the voice input means to a cloud server or a local server. This transmission is performed via an Internet connection. HTTPS is used as a specific transmission protocol.
[0291] 3. Means for converting audio data into text data
[0292] The server uses a speech recognition API (for example, Google Cloud Speech-to-Text API) to convert the received audio data into text data. The speech recognition API generates the text "The printer ink is not coming out."
[0293] 4. Means for analyzing text data to generate responses
[0294] The server uses a natural language processing model (e.g., OpenAI®'s GPT-3) to analyze text data and understand the intent behind user questions and instructions. It then refers to a database of instruction manuals and overview booklets to generate appropriate answers to the questions.
[0295] 5. Means for converting the generated response into audio data
[0296] The server converts the generated response into audio data using a speech synthesis API (e.g., Amazon Polly). This speech synthesis API then provides the response to the user as audio.
[0297] 6. Means for transmitting audio data to a terminal
[0298] The server sends the generated audio data to the terminal. This transmission also takes place via the internet connection. HTTPS is used as the transmission protocol.
[0299] 7. Means for playing audio data
[0300] The device plays back the received audio data. It provides the user with an audio response using speakers or headphones. Specific hardware options include smartphones, tablets, and speakers or headphones connected to a PC.
[0301] 8. Emotional Engine
[0302] The emotional engine analyzes and recognizes the user's emotional state using voice data. This emotional analysis is performed based on features such as the tone, speed, and intonation of the voice. As specific software, an emotional analysis engine (for example, IBM Watson (registered trademark) Tone Analyzer) can be used. Based on the analysis results, the server adjusts the response content and generates a more friendly response if necessary.
[0303] Specific Example
[0304] When the user asks "The printer is not printing ink", the series of processing steps are as follows:
[0305] 1. The user speaks to the terminal "The printer is not printing ink".
[0306] 2. The terminal captures this voice and sends the voice data to the server.
[0307] 3. The server converts the voice data into text data using the voice recognition API.
[0308] 4. The server analyzes the text data using the natural language processing model and searches for information related to the problem "The printer is not printing ink".
[0309] 5. The server generates an appropriate response to the user based on the user manual. For example, it generates text such as "To replace the ink cartridge, first open the printer cover, remove the old cartridge, and set the new cartridge."
[0310] 6. The server converts the text response generated into voice data using the text-to-speech API.
[0311] <00 8. The device plays audio data and provides the user with an answer.
[0313] 9. The emotion engine analyzes the user's original voice data and, if it determines that the user is confused, the server modifies the response to be more helpful and detailed. For example, it might adjust the tone and content of the response to something like, "Don't worry, first try opening the printer cover. Next, remove the old ink cartridge. Then, carefully insert the new cartridge."
[0314] Examples of prompt statements
[0315] Examples of prompt statements are as follows:
[0316] User: What should I do if my printer ink isn't coming out?
[0317] System: To replace an ink cartridge, first open the printer cover, remove the old cartridge, and then install the new cartridge.
[0318] By inputting this prompt into the AI model, the system generates an appropriate answer to the user's question. This mechanism allows users to use the system more comfortably and effectively.
[0319] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0320] Step 1:
[0321] Voice input
[0322] Action: The user speaks into the terminal saying, "The printer ink isn't coming out."
[0323] Input: User's voice.
[0324] Output: Digital audio data.
[0325] Details: The microphone built into the device captures the audio and saves it as digital audio data. Examples of devices that can be used include smartphones, tablets, and PCs with built-in microphones.
[0326] Step 2:
[0327] Sending audio data
[0328] Operation: The device sends audio data to the server.
[0329] Input: Digital audio data.
[0330] Output: Audio data transferred to the server.
[0331] Details: The device sends the captured audio data to a cloud server or local server via the internet. This communication uses the HTTPS protocol and is secure.
[0332] Step 3:
[0333] Convert audio data to text data
[0334] Operation: The server uses a speech recognition API to convert speech data into text data.
[0335] Input: Audio data transferred to the server.
[0336] Output: Text data. The content is "The printer ink isn't coming out."
[0337] Details: The server calls the Google Cloud Speech-to-Text API to convert the received audio data into text data.
[0338] Step 4:
[0339] Analyze text data to generate answers.
[0340] Operation: The server uses a natural language processing model to analyze text data and generate appropriate responses.
[0341] Input: Converted text data.
[0342] Output: Text data of the response.
[0343] Details: The server uses OpenAI's GPT-3 model to analyze the content of text data and understand the user's question and intent. It then refers to a database of instruction manuals and overview booklets to generate appropriate answers to the questions. A specific example answer is: "To replace the ink cartridge, first open the printer cover, remove the old cartridge, and then install the new cartridge."
[0344] Step 5:
[0345] Convert the generated response into audio data.
[0346] Operation: The server converts the generated response into audio data.
[0347] Input: Text data of the response.
[0348] Output: Audio data.
[0349] Details: The server calls the Amazon Polly API to convert the generated text responses into audio data.
[0350] Step 6:
[0351] Sending audio data
[0352] Operation: The server sends the generated audio data to the terminal.
[0353] Input: Audio data.
[0354] Output: Audio data transferred to the terminal.
[0355] Details: The server sends the generated audio data to the user's terminal via the internet. This communication also uses the HTTPS protocol, ensuring security.
[0356] Step 7:
[0357] Playback of audio data
[0358] Operation: Plays audio data received by the device.
[0359] Input: Audio data transferred to the device.
[0360] Output: Played audio.
[0361] Details: The device decodes the received audio data and plays the audio through the speaker or headphones, thereby providing the user with a response.
[0362] Step 8:
[0363] Sentiment analysis and response adjustment
[0364] Operation: The server uses an emotion engine to analyze the user's emotional state and adjusts its responses as needed.
[0365] Input: Audio data and its analysis results.
[0366] Output: Adjusted response data.
[0367] Details: The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze the tone, speed, and intonation of the user's voice data. Based on the analysis, if it is determined that the user is confused, the original response is modified to be more helpful and detailed. For example, the response's tone and content might be adjusted to something like, "Don't worry, first try opening the printer cover. Next, remove the old ink cartridge. Then, carefully insert the new cartridge."
[0368] (Application Example 2)
[0369] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0370] Current voice control systems fail to recognize the user's emotional state, resulting in problems such as being unable to respond appropriately when the user is confused or in an emergency. This is especially true in security services, where users may act inappropriately when panicked, requiring a swift and accurate response. However, existing systems cannot meet these needs.
[0371] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice input means, a means for transmitting voice data to the server, and a means for converting voice data into text data. This makes it possible to convert the user's voice input into text and analyze it. Furthermore, it includes a means for analyzing the text data to generate a response, a means for converting the generated response into voice data, an emotion engine for analyzing and recognizing the emotional state, a means for adjusting the response based on the emotion analysis results, and a means for contacting emergency contacts. This enables flexible responses based on the user's emotional state and rapid emergency response as needed.
[0372] "Voice input means" refers to devices or functions that allow users to input instructions or questions into a system via voice.
[0373] "Means of sending audio data to a server" refers to the function of sending captured audio data to a cloud server or local server via the internet or other means.
[0374] "Means for converting audio data to text data" refers to devices or functions that use speech recognition technology to convert audio data into text data.
[0375] "Means for analyzing text data and generating responses" refers to a function or technical means that analyzes input text data and generates appropriate responses.
[0376] "Means for converting generated responses into audio data" refers to a function that converts text-based responses into audio data using speech synthesis technology.
[0377] "Means for transmitting audio data to a terminal" refers to communication means for transmitting generated audio data to the user's terminal.
[0378] "Means for playing audio data" refers to the function of playing audio data received by a device through a speaker or headphones.
[0379] An "emotion engine that analyzes and recognizes the user's emotional state" is a software engine that analyzes and recognizes a user's emotions based on characteristics such as tone, speed, and intonation of their voice.
[0380] "Means for adjusting responses based on emotion analysis results" refers to a function that appropriately adjusts the tone and content of responses based on emotion analysis results obtained by the emotion engine.
[0381] The "means of contacting emergency contacts" feature is a function that sends a notification to pre-set emergency contacts when the user's emotional state is determined to be an emergency situation.
[0382] This invention is a system that combines user voice input and emotion recognition, and is particularly applicable to emergency response in security services. Specific embodiments for carrying out the invention are described below.
[0383] System Overview
[0384] This system includes means for voice input, means for transmitting voice data to a server, means for converting voice data into text data, means for analyzing text data to generate a response, means for converting the generated response into voice data, means for transmitting voice data to a terminal, means for playing back voice data, an emotion engine for analyzing and recognizing the user's emotional state, and means for adjusting the response based on the emotion analysis results. Furthermore, it also includes means for contacting emergency contacts in emergencies.
[0385] Hardware and software configuration
[0386] This system consists of the following hardware and software:
[0387] Audio input method: Use an audio capture device such as a microphone.
[0388] Server: Performs necessary data processing on cloud servers or local servers.
[0389] Speech recognition APIs such as Google Cloud Speech-to-Text are used to convert speech data into text data.
[0390] Natural language processing model: A model for understanding user questions and instructions, such as using the Google Cloud Natural Language API.
[0391] Text-to-speech APIs such as Google Cloud Text-to-Speech are used to convert text data into speech data.
[0392] Emotion Engine: Software that analyzes the tone, speed, intonation, etc. of a voice to recognize the user's emotions.
[0393] Emergency contact method: A means of communication used to send notifications to designated contacts in the event of an emergency.
[0394] Detailed operation
[0395] Voice input and data transmission
[0396] First, the user speaks a question or instruction into the smartphone's microphone. For example, a prompt might say, "There might be a suspicious person in my house, what should I do?" This audio data is captured and sent to the server.
[0397] Speech recognition and text conversion
[0398] The server converts the transmitted audio data into text data using a speech recognition API. The Google Cloud Speech-to-Text API is used, and this process generates the text "There might be a suspicious person in my house, what should I do?"
[0399] Text analysis and response generation
[0400] The text data is analyzed using a natural language processing model to generate appropriate responses. For example, in response to the above prompt, the following response is generated: "Stay calm. I will contact the police immediately. Move to a safe location."
[0401] Emotion recognition and response adjustment
[0402] Furthermore, this response is adjusted by an emotion engine according to the user's emotional state. If the tone and speed of the voice indicate that the user is in a state of panic, the tone of the response will be changed to a calmer and more reassuring one.
[0403] Speech synthesis and data transmission
[0404] The generated responses are converted into audio data using a speech synthesis API and sent from the server to the device. The Google Cloud Text-to-Speech API performs this conversion.
[0405] Audio playback
[0406] The user's device plays the received audio data. Playback occurs through speakers or headphones, and the user is provided with a response.
[0407] Emergency response
[0408] If the system determines that a user's emotional state is urgent, it sends a notification to a pre-configured emergency contact. This allows the user to receive appropriate assistance immediately.
[0409] Specific example
[0410] Let's explain the detailed operation using the previously shown prompt, "There might be a suspicious person in my house, what should I do?" as an example. When the user speaks this question into the voice input device, the voice data is captured and sent to the server. The server converts the voice data into text data and analyzes it using a natural language processing model. The generated response will be, "Please stay calm. I will contact the police immediately. Please move to a safe place." The emotion engine recognizes the user's panic state and further adjusts the response to be calmer. The response, converted into voice data using a speech synthesis API, is then sent to the device and played back. Simultaneously, a notification is sent to emergency contacts, allowing the user to receive quick assistance from the difficult situation.
[0411] The above describes a specific embodiment for carrying out this invention. By using this system, users can obtain answers to their questions and instructions safely and quickly, and can also respond in emergencies.
[0412] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0413] Step 1:
[0414] The user speaks questions or instructions using a voice input device.
[0415] As a concrete example, the user speaks into their smartphone's microphone, "There might be a suspicious person in my house, what should I do?" The input data is captured as audio data.
[0416] Step 2:
[0417] The device sends the captured audio data to the server.
[0418] In this process, the audio data is sent via the internet to a cloud server or a local server. The input is the captured audio data, and the output is the audio data sent to the server.
[0419] Step 3:
[0420] The server converts the audio data into text data using a speech recognition API.
[0421] Specifically, the Google Cloud Speech-to-Text API is used to generate text data for the sentence, "There might be a suspicious person in my house, what should I do?". The input is audio data, and the output is text data.
[0422] Step 4:
[0423] The server uses a natural language processing model to analyze text data and generate appropriate responses.
[0424] This example uses the Google Cloud Natural Language API to analyze the text "There might be a suspicious person in my house, what should I do?". Based on the analysis, it generates the response "Stay calm. I will contact the police immediately. Move to a safe location." The input is text data, and the output is the response text.
[0425] Step 5:
[0426] The server uses an emotion engine to analyze the user's emotional state.
[0427] Specifically, it evaluates the tone, speed, and intonation of the voice to determine if the user is in a state of panic. The input is voice data, and the output is the result of the emotion analysis.
[0428] Step 6:
[0429] Based on the sentiment analysis results, the server adjusts the response.
[0430] For example, if the user is in a state of panic, the tone of the response is calmed and the content is changed to something reassuring. The adjusted response would be, "Please don't worry, I'll contact the police immediately. Please move to a safe location." The input is the result of the sentiment analysis and the response text, and the output is the adjusted response text.
[0431] Step 7:
[0432] The generated responses are converted into audio data using a speech synthesis API.
[0433] Specifically, the Google Cloud Text-to-Speech API is used to convert the adjusted response text into audio data. The input is the adjusted response text, and the output is the generated audio data.
[0434] Step 8:
[0435] The server sends the audio data to the terminal.
[0436] This data transmission takes place over the internet and is sent to the user's terminal. The input is the generated audio data, and the output is the audio data sent to the terminal.
[0437] Step 9:
[0438] The device plays the audio data.
[0439] Users can receive responses from the system via audio through their smartphone's speaker or headphones. The input is the audio data sent to the device, and the output is the played audio.
[0440] Step 10:
[0441] The server will contact the emergency contact.
[0442] If the server determines that the user's emotional state is urgent, it sends a notification to a pre-configured emergency contact. The input is the result of the emotional analysis, and the output is the emergency notification that was sent.
[0443] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0444] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0445] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0446] [Second Embodiment]
[0447] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0448] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0449] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0450] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0451] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0452] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0453] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0454] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0455] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0456] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0457] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0458] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0459] The present invention will now describe embodiments for carrying out this invention. The present invention is a system in which a user can give operating instructions or ask questions using their voice and receive answers in voice. The system comprises a voice input means, a means for transmitting voice data to a server, a means for converting voice data into text data, a means for analyzing the text data to generate answers, a means for converting the generated answers into voice data, a means for transmitting voice data to a terminal, and a means for playing back the voice data.
[0460] System Configuration
[0461] 1. Voice input means
[0462] The user provides operating instructions and asks questions via voice to the terminal. This voice input method uses voice input devices such as microphones. In addition, voice recognition software is incorporated to capture the voice as digital data.
[0463] 2. Means for sending audio data to the server
[0464] The device sends the captured audio data to a cloud server or local server. This transmission takes place via an internet connection or local network.
[0465] 3. Means for converting audio data into text data
[0466] The server uses a speech recognition API (e.g., a cloud-based speech recognition service) to convert the received audio data into text data. This ensures that instructions and questions entered by voice are represented as text.
[0467] 4. Means for analyzing text data to generate responses
[0468] The server analyzes text data using a natural language processing model (e.g., pre-trained models such as BERT or GPT) to understand the intent behind the user's questions and instructions. It then refers to a specific database (e.g., instruction manuals or overview booklets) to generate appropriate answers to the questions.
[0469] 5. Means for converting the generated response into audio data
[0470] The generated responses are converted into audio data using a speech synthesis API (e.g., a cloud-based speech synthesis service). This provides the user with an audio response.
[0471] 6. Means for transmitting audio data to a terminal
[0472] The server sends the generated audio data to the terminal. This transmission takes place via an internet connection or a local network.
[0473] 7. Means for playing audio data
[0474] The device plays the received audio data. This allows the user to hear the response in audio. Audio output devices such as speakers or headphones are used.
[0475] Specific example
[0476] For example, if a user asks "The printer ink isn't coming out," the following process will be performed.
[0477] The user speaks into the voice input device and says, "The printer ink isn't coming out."
[0478] The device captures this audio data and sends it to the server.
[0479] The server converts the audio data into text data using a speech recognition API.
[0480] The text data is analyzed using a natural language processing model to search for relevant sections in the instruction manual related to the problem "the printer ink is not coming out."
[0481] The server generates the appropriate response from the instruction manual, for example, "You need to replace the ink cartridge. Open the printer cover, remove the old cartridge, and insert the new cartridge."
[0482] Convert this answer into speech data using a speech synthesis API.
[0483] Audio data is sent to the terminal, which then plays it back to the user through an audio output device.
[0484] In this way, users can ask questions and receive answers by voice without having to type. This provides an easy-to-use support system for users who have difficulty typing or who are visually impaired.
[0485] The following describes the processing flow.
[0486] Step 1:
[0487] The user inputs operating instructions or questions by voice into the terminal. For example, the user might say, "The printer ink isn't coming out."
[0488] Step 2:
[0489] The device uses its built-in microphone to capture the user's voice and temporarily stores it as digital audio data.
[0490] Step 3:
[0491] The device sends the captured audio data to the server. This transmission takes place via an internet connection.
[0492] Step 4:
[0493] The server sends the received audio data to a speech recognition API. Typically, the speech recognition API uses a cloud-based speech recognition service.
[0494] Step 5:
[0495] The server processes the text data received from the speech recognition API. For example, the text "The printer ink isn't coming out" is generated.
[0496] Step 6:
[0497] The server uses a natural language processing model to analyze text data. This model performs analysis to understand the user's intent and find relevant information.
[0498] Step 7:
[0499] The server searches a database of instruction manuals and summary booklets to obtain appropriate answers to user questions. For example, this might include questions such as "Providing instructions on how to replace ink cartridges."
[0500] Step 8:
[0501] The server generates answers in a user-friendly format based on the relevant section of the instruction manual. For example, it might generate text such as, "To replace the ink cartridge, first open the printer cover, remove the old cartridge, and then insert the new cartridge."
[0502] Step 9:
[0503] The server generates text responses, which are then sent to a speech synthesis API to be converted into audio data. This speech synthesis API also often uses a cloud-based service.
[0504] Step 10:
[0505] The server sends the generated audio data to the terminal. This transmission also takes place via the internet connection.
[0506] Step 11:
[0507] The device plays the audio data it received from the server. Using the device's speaker or a headset or other audio output device, it provides the user with an audio response. The message played is: "To replace the ink cartridge, first open the printer cover, remove the old cartridge, and then insert the new cartridge."
[0508] Through the steps described above, this system responds to user voice input with voice commands. Therefore, it is easy to use for users who have difficulty with text input or those with visual impairments.
[0509] (Example 1)
[0510] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0511] In recent years, voice-based interfaces have attracted attention, becoming a particularly important technology for users with visual impairments or those who have difficulty with text input. However, current voice response systems face several challenges. For example, accurately converting voice input data into text and then quickly analyzing it to generate appropriate responses is difficult. Furthermore, there is a need to convert the generated responses into natural and easy-to-understand speech. In addition, optimization is required to perform these processes in real time. This invention aims to solve these problems and provide quick and accurate voice responses to user questions and operational instructions.
[0512] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0513] In this invention, the server includes means for understanding the intent of questions and operational instructions by analyzing voice data into text data and the text data; means for generating appropriate answers from a database based on the analysis results; and means for converting the generated answer text into voice data. This makes it possible to quickly and accurately perform text analysis and answer generation for questions and instructions entered by the user via voice, and to provide those answers as natural and easy-to-understand voice.
[0514] "Voice input means" refers to a device that allows users to ask questions or give instructions using their voice.
[0515] "Means for sending audio data to a server" refers to the technology for sending captured audio data to a server via a network.
[0516] "Means of converting audio data to text data" refers to technology that converts captured audio data into text format.
[0517] "Methods for analyzing text data and generating responses" refers to technologies that analyze data expressed as text and generate appropriate responses to user questions or instructions.
[0518] "Means of converting generated responses into audio data" refers to technology that converts text-based responses into audio data.
[0519] "Means for transmitting audio data to a terminal" refers to the technology for transmitting generated audio data to a user's terminal via a network.
[0520] "Means of playing audio data" refers to the devices and software used to play audio data on the user's device.
[0521] "Means of converting to digital audio data in real time" refers to technology that converts user voice input into digital data in real time.
[0522] "Text data and means for analyzing that text data" refers to technologies for analyzing converted text data and understanding its intent and content.
[0523] "Means for generating appropriate answers from a database" refers to technologies that refer to a database based on analysis results and extract appropriate answers to user questions and operational instructions.
[0524] "Methods for converting response text into audio data" refers to technologies that convert generated responses into natural and easy-to-understand audio data.
[0525] This invention is a system that allows users to give operating instructions and ask questions using their voice and receive answers in voice. This system combines a voice input device, a network, a server, a natural language processing model, and speech synthesis technology.
[0526] Voice input method
[0527] The user uses a microphone to speak into the device, asking questions or giving instructions. The device is equipped with speech recognition software, such as the Google Speech-to-Text API, which converts the voice into digital audio data in real time. For example, if the user says, "The printer ink isn't coming out," this voice is captured as digital data.
[0528] Sending audio data
[0529] The device transmits the captured audio data to the server via an internet connection or local network. This transmission process uses encryption technologies such as TLS to ensure data security.
[0530] Converting audio data to text
[0531] The server receives the incoming audio data and uses the Amazon Transcribe API to convert the audio into text data. For example, the user's audio data, "The printer ink isn't coming out," is converted into the text format "The printer ink isn't coming out."
[0532] Text data analysis
[0533] The server analyzes the converted text data using natural language processing models (e.g., GPT-3, BERT) to understand the intent behind questions and instructions. For example, the server analyzes the text data "The printer ink isn't coming out" and understands its meaning. This process includes generating text tokens and interpreting intent.
[0534] Answer generation
[0535] Based on the analysis results, the server extracts and generates the appropriate answer from a database (e.g., a printer instruction manual database). For example, it might generate an answer such as, "You need to replace the ink cartridge. Open the printer cover, remove the old cartridge, and insert the new cartridge."
[0536] Voice conversion of the answer
[0537] The generated response text is converted into speech data using speech synthesis technology such as the Google Text-to-Speech API. This ensures that the user's requested response is produced as natural-sounding speech data.
[0538] Sending audio data
[0539] The server then transmits the generated audio data back to the terminal via the internet or local network. Security is ensured during this process by encrypting the data using TLS or similar protocols.
[0540] Playback of audio data
[0541] The device plays the received audio data through an audio output device such as a speaker or headphones. The user can then listen to this audio and take the necessary actions or responses. For example, the device's speaker might play an audio message saying, "You need to replace the ink cartridge. Open the printer cover, remove the old cartridge, and insert the new cartridge."
[0542] Specific example
[0543] For example, if a user asks "The printer ink isn't coming out," the following process will be performed.
[0544] The user speaks into the voice input device and says, "The printer ink isn't coming out."
[0545] The device captures this audio data and sends it to the server.
[0546] The server converts the audio data into text data using a speech recognition API.
[0547] The text data is analyzed using a natural language processing model to search for relevant sections in the instruction manual related to the problem "the printer ink is not coming out."
[0548] The server generates the appropriate response from the instruction manual, for example, "You need to replace the ink cartridge. Open the printer cover, remove the old cartridge, and insert the new cartridge."
[0549] Convert this answer into speech data using a speech synthesis API.
[0550] Audio data is sent to the terminal, which then plays it back to the user through an audio output device.
[0551] Example of a prompt
[0552] User input: "The printer ink isn't coming out."
[0553] System response: "You need to replace the ink cartridge. Open the printer cover, remove the old cartridge, and insert the new one."
[0554] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0555] Step 1: Capture voice input
[0556] The user speaks into a voice input device (e.g., a microphone) to ask a question or give an instruction. The device converts this voice into digital speech data using the Google Speech-to-Text API. The input is the user's utterance, "The printer ink isn't coming out," and the output is the text data of this utterance. Specifically, the device captures the user's voice in real time and converts it into a digital signal.
[0557] Step 2: Sending the audio data
[0558] The terminal sends the generated digital audio data to the cloud server using encryption technology such as TLS. The input is digital audio data, and the output is the transmission of encrypted data. Specifically, the terminal sends the data to the server via an internet connection or a local network.
[0559] Step 3: Converting audio data to text
[0560] The server converts received audio data into text data using the Amazon Transcribe API. The input is digital audio data, and the output is the corresponding text data. Specifically, the server receives audio data and makes an API request to convert it into text data.
[0561] Step 4: Analyzing Text Data
[0562] The server receives text data and performs analysis using a natural language processing model (e.g., BERT, GPT-3). The input is text data, and the output is semantic understanding data resulting from the analysis. Specifically, the server tokenizes the text and uses the model to analyze the intent of questions and instructions.
[0563] Step 5: Generating the answer
[0564] Based on the analysis results, the server extracts and generates appropriate answers from a database (e.g., an instruction manual database). The input is semantic comprehension data, and the output is the generated answer text. Specifically, the server searches the queried database, extracts matching answers, and assembles them.
[0565] Step 6: Voice conversion of the answer
[0566] The generated response text is converted into speech data using the Google Text-to-Speech API. The input is the response text, and the output is the corresponding speech data. Specifically, the server makes an API request to convert the text to speech.
[0567] Step 7: Sending audio data
[0568] The server re-encrypts the generated audio data using TLS or similar methods and sends it to the terminal. The input is the audio data, and the output is the transmission of the encrypted data. Specifically, the server sends the data back to the terminal via the network.
[0569] Step 8: Play back audio data
[0570] The device decodes the received audio data and plays it back through an audio output device such as a speaker or headphones. The input is audio data, and the output is the audio that the user hears. Specifically, the device decodes the encryption and outputs the audio using an audio playback device.
[0571] (Application Example 1)
[0572] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0573] Traditional food delivery systems require users to place orders via a screen, which presents challenges for users with dirty hands or those with visual impairments. Furthermore, checking order status and estimated delivery times also requires manual operation, resulting in a lack of convenience.
[0574] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0575] In this invention, the server includes a voice input means, a means for transmitting voice data to the server, a means for converting voice data into text data, a means for analyzing the text data to generate a response corresponding to the order, a means for converting the generated response into voice data, a means for transmitting voice data to a terminal, and a means for playing back the voice data. This allows users to order food and check the order status using only their voice, improving convenience for people with visual impairments and users whose hands are occupied.
[0576] A "voice input method" refers to a device or system that allows a user to give instructions by voice.
[0577] "Means for sending audio data to a server" refers to the mechanisms and protocols for sending captured audio data to a server via the internet.
[0578] "Means for converting audio data to text data" refers to technologies for converting audio data captured using speech recognition technology into text format.
[0579] "A means of analyzing text data to generate responses corresponding to orders" refers to a system that uses natural language processing technology to analyze text data and generate appropriate responses from the obtained information.
[0580] "Means for converting generated responses into audio data" refers to a mechanism or system that converts text-based responses into audio data using speech synthesis technology.
[0581] "Means for transmitting audio data to a terminal" refers to the mechanisms and protocols for transmitting generated audio data to a terminal used by the user.
[0582] "Means for playing audio data" refers to devices or systems that play transmitted audio data and allow the user to listen to it.
[0583] "Data related to food and beverage orders" refers to data including menus, prices, and restaurant information handled within the food delivery system.
[0584] A "natural language processing model" is a machine learning technique or algorithm used to analyze user text data, understand user intent, and generate appropriate responses.
[0585] "Responses regarding orders and delivery status" refers to responses that include information about the products ordered by the user and their delivery status.
[0586] The system for carrying out this invention allows users to place food delivery orders and check the order status using voice commands. The system includes the following main components:
[0587] 1. Voice input means
[0588] Users input voice commands, such as ordering food or checking delivery status, by speaking into the microphone of their device (smartphone or smart glasses). This voice input method utilizes the built-in microphone of the smartphone or smart glasses, and captures the voice as digital data using the Google Speech-to-Text API.
[0589] 2. Means for sending audio data to the server
[0590] The device sends the captured audio data to a cloud server via an internet connection. HTTPS communication is used for data transmission.
[0591] 3. Means for converting audio data into text data
[0592] The server converts the received audio data into text data using speech recognition technology. Specifically, it uses the Google Speech-to-Text API to convert the audio data into text format.
[0593] 4. A means of analyzing text data to generate responses corresponding to orders.
[0594] The server analyzes the data, which has been converted to text format, using a natural language processing model. It utilizes advanced pre-trained models such as GPT-4 to understand user intent and generate appropriate responses. In this process, it references a food delivery database, providing data such as menu information, pricing information, and estimated delivery times.
[0595] 5. Means for converting the generated response into audio data
[0596] The server converts the generated text-based responses into speech data using speech synthesis technology. This conversion utilizes the Google Text-to-Speech API.
[0597] 6. Means for transmitting audio data to a terminal
[0598] The server then sends the generated audio data back to the terminal. This is also done via HTTPS communication over the internet connection.
[0599] 7. Means for playing audio data
[0600] Finally, the device plays back the received audio data and provides the user with an answer. A smartphone, smart glasses speaker, or headphones are used as the playback device.
[0601] Adding specific examples
[0602] For example, if a user uses voice input to say "I want to order a pizza," the following process will occur.
[0603] The user says, "I want to order a pizza."
[0604] The device captures this audio and sends it to the server.
[0605] The server uses the Google Speech-to-Text API to convert the audio data into text data.
[0606] A natural language processing model (e.g., GPT-4) is used to analyze text data such as "I want to order a pizza" and provide information on appropriate restaurants and menus.
[0607] The server generates an answer such as "Which restaurant would you order pizza from?", and this answer is converted into speech data using the Google Text-to-Speech API.
[0608] The server generates audio data and sends it to the terminal, which then plays the audio.
[0609] Example of a prompt
[0610] Analyze user comments and suggest corresponding food delivery restaurants and menus: {User comments}
[0611] Follow user instructions, confirm the order, and communicate with the backend to check its status.
[0612] In this way, users can perform a series of operations, from ordering food delivery to checking the delivery status, using only their voice, without using a visual interface. This significantly improves convenience for users whose hands are full or who have visual impairments.
[0613] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0614] Step 1:
[0615] The user speaks into the device's voice input and says, "I want to order a pizza." This captures the audio data. The input is what the user said, and the output is the captured audio data.
[0616] Step 2:
[0617] The device sends the captured audio data to the cloud server via the internet connection. The input is the captured audio data, and the output is the audio data received on the server side. Specifically, the operation involves sending the audio data using HTTPS communication.
[0618] Step 3:
[0619] The server converts the received audio data into text data using the Google Speech-to-Text API. The input is the received audio data, and the output is text data. Specifically, it analyzes the audio data using speech recognition technology and converts it into the corresponding text.
[0620] Step 4:
[0621] The server analyzes the converted text data using a natural language processing model (e.g., GPT-4). The input is text data, and the output is the user's intent or question content as a result of the analysis. Specifically, it analyzes the text using a generative AI model and understands the content of the corresponding instructions or questions.
[0622] Step 5:
[0623] The server generates appropriate restaurant and menu information based on the analysis results and creates a text-based response. The input is the analysis results, and the output is the generated text-based response. Specifically, it refers to an internal restaurant database, extracts appropriate information to meet the user's request, and generates a response.
[0624] Step 6:
[0625] The server converts the generated text-based response into audio data using the Google Text-to-Speech API. The input is a text-based response, and the output is audio data. Specifically, it uses speech synthesis technology to convert text into speech.
[0626] Step 7:
[0627] The server sends the generated audio data back to the terminal. The input is the generated audio data, and the output is the audio data received by the terminal. Specifically, the operation involves sending the audio data to the terminal using HTTPS communication.
[0628] Step 8:
[0629] The device plays the received audio data and provides a response to the user. The input is the received audio data, and the output is the audio information to be heard by the user. Specifically, the device plays the audio data using its speaker or headphones.
[0630] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0631] The embodiments for carrying out this invention will be described in detail. The present invention combines an emotion recognition function with a system in which a user inputs operation instructions or questions by voice and receives answers by voice. The system comprises a voice input means, a means for transmitting voice data to a server, a means for converting voice data into text data, a means for analyzing the text data to generate answers, a means for converting the generated answers into voice data, a means for transmitting voice data to a terminal, a means for playing back the voice data, and an emotion engine that recognizes the user's emotions.
[0632] System Configuration
[0633] 1. Voice input means
[0634] The user inputs operating instructions or questions by voice into the terminal. For example, the user might say, "The printer ink isn't coming out." This voice input method uses a microphone to capture the voice data and saves it as digital audio data.
[0635] 2. Means for sending audio data to the server
[0636] The device sends audio data captured via voice input to a cloud server or local server. This transmission takes place via an internet connection.
[0637] 3. Means for converting audio data into text data
[0638] The server uses a speech recognition API to convert the received audio data into text data. The speech recognition API generates the text "The printer ink isn't coming out."
[0639] 4. Means for analyzing text data to generate responses
[0640] The server uses a natural language processing model to analyze text data and understand the intent behind user questions and instructions. It then consults a database of instruction manuals and overview booklets to generate appropriate answers to the questions.
[0641] 5. Means for converting the generated response into audio data
[0642] The server converts the generated response into audio data using a speech synthesis API. This speech synthesis API then prepares the response for the user as audio.
[0643] 6. Means for transmitting audio data to a terminal
[0644] The server sends the generated audio data to the terminal. This transmission also takes place via the internet connection.
[0645] 7. Means for playing audio data
[0646] The device plays the received audio data. It then provides the user with the answer via voice using a speaker or headphones.
[0647] 8. Emotional Engine
[0648] The emotion engine analyzes and recognizes the user's emotional state using voice data. This emotion analysis is based on features such as voice tone, speed, and intonation.
[0649] Specific example
[0650] For example, the sequence of operations when a user asks "The printer ink isn't coming out" is as follows:
[0651] The user speaks into the voice input device and says, "The printer ink isn't coming out."
[0652] The device captures this audio and sends the audio data to the server.
[0653] The server converts the audio data into text data using a speech recognition API.
[0654] The server uses a natural language processing model to analyze text data and search for information related to the problem "the printer ink isn't coming out."
[0655] The server generates appropriate responses to the user based on the instruction manual. For example, it might generate text such as, "To replace the ink cartridge, first open the printer cover, remove the old cartridge, and then insert the new cartridge."
[0656] The text response generated by the server is converted into speech data using a speech synthesis API.
[0657] The server sends the audio data to the terminal.
[0658] The device plays audio data and provides the user with an answer.
[0659] The emotion engine analyzes the user's original voice data, and if it determines that the user is confused, the server modifies the response to be more helpful and detailed. For example, it might adjust the tone and content of the response to something like, "Don't worry, first try opening the printer cover. Next, carefully remove the old ink cartridge. Then, firmly insert the new cartridge."
[0660] Through these steps, the system can provide flexible responses tailored to the user's emotional state. This allows users to use the system more comfortably and effectively.
[0661] The following describes the processing flow.
[0662] Step 1:
[0663] The user inputs operating instructions or questions by voice into the terminal. For example, the user might say, "The printer ink isn't coming out."
[0664] Step 2:
[0665] The device uses its built-in microphone to capture the user's voice and temporarily stores it as digital audio data.
[0666] Step 3:
[0667] The device sends the captured audio data to the server. This transmission takes place via an internet connection.
[0668] Step 4:
[0669] The server sends audio data to a speech recognition API, which then converts the audio data into text data. For example, a cloud-based speech recognition service could be used.
[0670] Step 5:
[0671] The server processes the text data "Printer ink is not coming out" received from the speech recognition API.
[0672] Step 6:
[0673] The server uses a natural language processing model to analyze the text data. This analysis helps the server understand what the user wants. For example, it might determine that the user wants to solve a printer problem.
[0674] Step 7:
[0675] The server searches a database of instruction manuals and summary booklets to obtain appropriate answers to user questions. For example, it might find information such as, "The ink cartridge needs to be replaced."
[0676] Step 8:
[0677] The server generates appropriate responses for the user based on the instruction manual. For example, it might generate text such as, "To replace the ink cartridge, first open the printer cover, remove the old cartridge, and then insert the new cartridge."
[0678] Step 9:
[0679] The server sends the generated text response to a speech synthesis API, which converts it into audio data. This speech synthesis API then prepares the response for the user as audio.
[0680] Step 10:
[0681] The server sends the generated audio data to the terminal. This transmission also takes place via the internet connection.
[0682] Step 11:
[0683] The device plays the audio data it received from the server. Using the device's speaker or headphones, it provides the user with an audio response. The message played is: "To replace the ink cartridge, first open the printer cover, remove the old cartridge, and then insert the new cartridge."
[0684] Step 12:
[0685] The emotion engine analyzes the original audio data and evaluates the user's emotional state. For example, it can detect feelings of confusion or frustration.
[0686] Step 13:
[0687] The server adjusts the responses generated based on the emotion engine's analysis. For example, if the server detects that the user is confused, it adjusts the response to be more helpful and detailed. It might change to something like, "Don't worry, first try opening the printer cover. Next, carefully remove the old ink cartridge. Then, firmly insert the new cartridge."
[0688] Step 14:
[0689] The server sends the newly processed audio data back to the speech synthesis API, where it is converted back into audio data.
[0690] Step 15:
[0691] The server sends new audio data to the device, and the device plays this data. The user can then hear the corrected, gentler-toned response.
[0692] (Example 2)
[0693] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0694] Conventional voice dialogue systems can provide appropriate answers to user inputs, such as operational instructions and questions, but they struggle to provide flexible responses that adapt to the user's emotional state. Furthermore, they lacked the means to recognize and appropriately respond to user feelings of confusion or anxiety. As a result, the user experience was not always optimal.
[0695] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0696] In this invention, the server includes means for converting voice data into text data, means for analyzing the text data to generate a response, and means for converting the generated response back into voice data. This enables flexible response generation that takes into account the emotional state of the user based on voice input.
[0697] "Voice input means" refers to a device that includes a microphone for the user to input voice, and a means for capturing that voice as digital data.
[0698] "Means for sending audio data to a server" refers to means for sending captured audio data to a cloud server or local server via the internet.
[0699] "Means for converting audio data to text data" refers to methods for converting audio data received using a speech recognition API into text data.
[0700] "Methods for analyzing text data and generating responses" refers to methods for analyzing text data using natural language processing models and generating appropriate responses.
[0701] "Means for converting generated responses into audio data" refers to means for converting text responses generated using a speech synthesis API into audio data.
[0702] "Means for transmitting audio data to a terminal" refers to means for transmitting generated audio data to a user's terminal via the internet.
[0703] "Means for playing audio data" refers to the means by which a terminal plays back audio data it has received through speakers or headphones.
[0704] An "emotion engine" is a means of analyzing and recognizing a user's emotional state using voice data and responding accordingly.
[0705] A "natural language processing model" is an algorithm or software that analyzes text data to understand the meaning and intent of human language.
[0706] A "speech recognition API" is an application programming interface for converting speech data into text data.
[0707] A "speech synthesis API" is an application programming interface for converting text data into speech data.
[0708] The embodiments for carrying out this invention will be described in detail. The present invention is a system including voice input means, means for transmitting voice data to a server, means for converting voice data to text data, means for analyzing text data to generate a response, means for converting the generated response to voice data, means for transmitting voice data to a terminal, means for playing back voice data, and an emotion engine for recognizing the user's emotions.
[0709] System Configuration
[0710] 1. Voice input means
[0711] The user inputs operating instructions or questions by voice into the device. For example, the user might say, "The printer ink isn't coming out." This voice input method uses a microphone to capture the voice data and saves it as digital audio data. Specific hardware that can be used include smartphones, tablets, and PCs with microphones.
[0712] 2. Means for sending audio data to the server
[0713] The device sends the audio data captured via voice input to a cloud server or local server. This transmission takes place over the internet. HTTPS is used as the specific transmission protocol.
[0714] 3. Means for converting audio data into text data
[0715] The server uses a speech recognition API (for example, Google Cloud Speech-to-Text API) to convert the received audio data into text data. The speech recognition API generates the text "The printer ink is not coming out."
[0716] 4. Means for analyzing text data to generate responses
[0717] The server uses a natural language processing model (e.g., OpenAI's GPT-3) to analyze text data and understand the intent behind user questions and instructions. It then consults a database of instruction manuals and overview booklets to generate appropriate answers to the questions.
[0718] 5. Means for converting the generated response into audio data
[0719] The server converts the generated response into audio data using a speech synthesis API (e.g., Amazon Polly). This speech synthesis API then provides the response to the user as audio.
[0720] 6. Means for transmitting audio data to a terminal
[0721] The server sends the generated audio data to the terminal. This transmission also takes place via the internet connection. HTTPS is used as the transmission protocol.
[0722] 7. Means for playing audio data
[0723] The device plays back the received audio data. It provides the user with an audio response using speakers or headphones. Specific hardware options include smartphones, tablets, and speakers or headphones connected to a PC.
[0724] 8. Emotional Engine
[0725] The emotion engine analyzes and recognizes the user's emotional state using voice data. This emotion analysis is based on features such as voice tone, speed, and intonation. Specific software such as an emotion analysis engine (e.g., IBM Watson Tone Analyzer) can be used. Based on the analysis results, the server adjusts its responses and generates more helpful responses as needed.
[0726] Specific example
[0727] The following is the sequence of events that occurs when a user asks, "The printer ink isn't coming out":
[0728] 1. The user speaks into the device and says, "The printer ink isn't coming out."
[0729] 2. The device captures this audio and sends the audio data to the server.
[0730] 3. The server converts the audio data into text data using a speech recognition API.
[0731] 4. The server uses a natural language processing model to analyze the text data and search for information related to the problem "the printer ink is not coming out."
[0732] 5. The server generates appropriate responses to the user based on the instruction manual. For example, it might generate text such as, "To replace the ink cartridge, first open the printer cover, remove the old cartridge, and then insert the new cartridge."
[0733] 6. The text response generated by the server is converted into speech data using a speech synthesis API.
[0734] 7. The server sends the audio data to the terminal.
[0735] 8. The device plays audio data and provides the user with an answer.
[0736] 9. The emotion engine analyzes the user's original voice data and, if it determines that the user is confused, the server modifies the response to be more helpful and detailed. For example, it might adjust the tone and content of the response to something like, "Don't worry, first try opening the printer cover. Next, remove the old ink cartridge. Then, carefully insert the new cartridge."
[0737] Examples of prompt statements
[0738] Examples of prompt statements are as follows:
[0739] User: What should I do if my printer ink isn't coming out?
[0740] System: To replace an ink cartridge, first open the printer cover, remove the old cartridge, and then install the new cartridge.
[0741] By inputting this prompt into the AI model, the system generates an appropriate answer to the user's question. This mechanism allows users to use the system more comfortably and effectively.
[0742] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0743] Step 1:
[0744] Voice input
[0745] Action: The user speaks into the terminal saying, "The printer ink isn't coming out."
[0746] Input: User's voice.
[0747] Output: Digital audio data.
[0748] Details: The microphone built into the device captures the audio and saves it as digital audio data. Examples of devices that can be used include smartphones, tablets, and PCs with built-in microphones.
[0749] Step 2:
[0750] Sending audio data
[0751] Operation: The device sends audio data to the server.
[0752] Input: Digital audio data.
[0753] Output: Audio data transferred to the server.
[0754] Details: The device sends the captured audio data to a cloud server or local server via the internet. This communication uses the HTTPS protocol and is secure.
[0755] Step 3:
[0756] Convert audio data to text data
[0757] Operation: The server uses a speech recognition API to convert speech data into text data.
[0758] Input: Audio data transferred to the server.
[0759] Output: Text data. The content is "The printer ink isn't coming out."
[0760] Details: The server calls the Google Cloud Speech-to-Text API to convert the received audio data into text data.
[0761] Step 4:
[0762] Analyze text data to generate answers.
[0763] Operation: The server uses a natural language processing model to analyze text data and generate appropriate responses.
[0764] Input: Converted text data.
[0765] Output: Text data of the response.
[0766] Details: The server uses OpenAI's GPT-3 model to analyze the content of text data and understand the user's question and intent. It then refers to a database of instruction manuals and overview booklets to generate appropriate answers to the questions. A specific example answer is: "To replace the ink cartridge, first open the printer cover, remove the old cartridge, and then install the new cartridge."
[0767] Step 5:
[0768] Convert the generated response into audio data.
[0769] Operation: The server converts the generated response into audio data.
[0770] Input: Text data of the response.
[0771] Output: Audio data.
[0772] Details: The server calls the Amazon Polly API to convert the generated text responses into audio data.
[0773] Step 6:
[0774] Sending audio data
[0775] Operation: The server sends the generated audio data to the terminal.
[0776] Input: Audio data.
[0777] Output: Audio data transferred to the terminal.
[0778] Details: The server sends the generated audio data to the user's terminal via the internet. This communication also uses the HTTPS protocol, ensuring security.
[0779] Step 7:
[0780] Playback of audio data
[0781] Operation: Plays audio data received by the device.
[0782] Input: Audio data transferred to the device.
[0783] Output: Played audio.
[0784] Details: The device decodes the received audio data and plays the audio through the speaker or headphones, thereby providing the user with a response.
[0785] Step 8:
[0786] Sentiment analysis and response adjustment
[0787] Operation: The server uses an emotion engine to analyze the user's emotional state and adjusts its responses as needed.
[0788] Input: Audio data and its analysis results.
[0789] Output: Adjusted response data.
[0790] Details: The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze the tone, speed, and intonation of the user's voice data. Based on the analysis, if it is determined that the user is confused, the original response is modified to be more helpful and detailed. For example, the response's tone and content might be adjusted to something like, "Don't worry, first try opening the printer cover. Next, remove the old ink cartridge. Then, carefully insert the new cartridge."
[0791] (Application Example 2)
[0792] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0793] Current voice control systems fail to recognize the user's emotional state, resulting in problems such as being unable to respond appropriately when the user is confused or in an emergency. This is especially true in security services, where users may act inappropriately when panicked, requiring a swift and accurate response. However, existing systems cannot meet these needs.
[0794] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice input means, a means for transmitting voice data to the server, and a means for converting voice data into text data. This makes it possible to convert the user's voice input into text and analyze it. Furthermore, it includes a means for analyzing the text data to generate a response, a means for converting the generated response into voice data, an emotion engine for analyzing and recognizing the emotional state, a means for adjusting the response based on the emotion analysis results, and a means for contacting emergency contacts. This enables flexible responses based on the user's emotional state and rapid emergency response as needed.
[0795] "Voice input means" refers to devices or functions that allow users to input instructions or questions into a system via voice.
[0796] "Means of sending audio data to a server" refers to the function of sending captured audio data to a cloud server or local server via the internet or other means.
[0797] "Means for converting audio data to text data" refers to devices or functions that use speech recognition technology to convert audio data into text data.
[0798] "Means for analyzing text data and generating responses" refers to a function or technical means that analyzes input text data and generates appropriate responses.
[0799] "Means for converting generated responses into audio data" refers to a function that converts text-based responses into audio data using speech synthesis technology.
[0800] "Means for transmitting audio data to a terminal" refers to communication means for transmitting generated audio data to the user's terminal.
[0801] "Means for playing audio data" refers to the function of playing audio data received by a device through a speaker or headphones.
[0802] An "emotion engine that analyzes and recognizes the user's emotional state" is a software engine that analyzes and recognizes a user's emotions based on characteristics such as tone, speed, and intonation of their voice.
[0803] "Means for adjusting responses based on emotion analysis results" refers to a function that appropriately adjusts the tone and content of responses based on emotion analysis results obtained by the emotion engine.
[0804] The "means of contacting emergency contacts" feature is a function that sends a notification to pre-set emergency contacts when the user's emotional state is determined to be an emergency situation.
[0805] This invention is a system that combines user voice input and emotion recognition, and is particularly applicable to emergency response in security services. Specific embodiments for carrying out the invention are described below.
[0806] System Overview
[0807] This system includes means for voice input, means for transmitting voice data to a server, means for converting voice data into text data, means for analyzing text data to generate a response, means for converting the generated response into voice data, means for transmitting voice data to a terminal, means for playing back voice data, an emotion engine for analyzing and recognizing the user's emotional state, and means for adjusting the response based on the emotion analysis results. Furthermore, it also includes means for contacting emergency contacts in emergencies.
[0808] Hardware and software configuration
[0809] This system consists of the following hardware and software:
[0810] Audio input method: Use an audio capture device such as a microphone.
[0811] Server: Performs necessary data processing on cloud servers or local servers.
[0812] Speech recognition APIs such as Google Cloud Speech-to-Text are used to convert speech data into text data.
[0813] Natural language processing model: A model for understanding user questions and instructions, such as using the Google Cloud Natural Language API.
[0814] Text-to-speech APIs such as Google Cloud Text-to-Speech are used to convert text data into speech data.
[0815] Emotion Engine: Software that analyzes the tone, speed, intonation, etc. of a voice to recognize the user's emotions.
[0816] Emergency contact method: A means of communication used to send notifications to designated contacts in the event of an emergency.
[0817] Detailed operation
[0818] Voice input and data transmission
[0819] First, the user speaks a question or instruction into the smartphone's microphone. For example, a prompt might say, "There might be a suspicious person in my house, what should I do?" This audio data is captured and sent to the server.
[0820] Speech recognition and text conversion
[0821] The server converts the transmitted audio data into text data using a speech recognition API. The Google Cloud Speech-to-Text API is used, and this process generates the text "There might be a suspicious person in my house, what should I do?"
[0822] Text analysis and response generation
[0823] The text data is analyzed using a natural language processing model to generate appropriate responses. For example, in response to the above prompt, the following response is generated: "Stay calm. I will contact the police immediately. Move to a safe location."
[0824] Emotion recognition and response adjustment
[0825] Furthermore, this response is adjusted by an emotion engine according to the user's emotional state. If the tone and speed of the voice indicate that the user is in a state of panic, the tone of the response will be changed to a calmer and more reassuring one.
[0826] Speech synthesis and data transmission
[0827] The generated responses are converted into audio data using a speech synthesis API and sent from the server to the device. The Google Cloud Text-to-Speech API performs this conversion.
[0828] Audio playback
[0829] The user's device plays the received audio data. Playback occurs through speakers or headphones, and the user is provided with a response.
[0830] Emergency response
[0831] If the system determines that a user's emotional state is urgent, it sends a notification to a pre-configured emergency contact. This allows the user to receive appropriate assistance immediately.
[0832] Specific example
[0833] Let's explain the detailed operation using the previously shown prompt, "There might be a suspicious person in my house, what should I do?" as an example. When the user speaks this question into the voice input device, the voice data is captured and sent to the server. The server converts the voice data into text data and analyzes it using a natural language processing model. The generated response will be, "Please stay calm. I will contact the police immediately. Please move to a safe place." The emotion engine recognizes the user's panic state and further adjusts the response to be calmer. The response, converted into voice data using a speech synthesis API, is then sent to the device and played back. Simultaneously, a notification is sent to emergency contacts, allowing the user to receive quick assistance from the difficult situation.
[0834] The above describes a specific embodiment for carrying out this invention. By using this system, users can obtain answers to their questions and instructions safely and quickly, and can also respond in emergencies.
[0835] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0836] Step 1:
[0837] The user speaks questions or instructions using a voice input device.
[0838] As a concrete example, the user speaks into their smartphone's microphone, "There might be a suspicious person in my house, what should I do?" The input data is captured as audio data.
[0839] Step 2:
[0840] The device sends the captured audio data to the server.
[0841] In this process, the audio data is sent via the internet to a cloud server or a local server. The input is the captured audio data, and the output is the audio data sent to the server.
[0842] Step 3:
[0843] The server converts the audio data into text data using a speech recognition API.
[0844] Specifically, the Google Cloud Speech-to-Text API is used to generate text data for the sentence, "There might be a suspicious person in my house, what should I do?". The input is audio data, and the output is text data.
[0845] Step 4:
[0846] The server uses a natural language processing model to analyze text data and generate appropriate responses.
[0847] This example uses the Google Cloud Natural Language API to analyze the text "There might be a suspicious person in my house, what should I do?". Based on the analysis, it generates the response "Stay calm. I will contact the police immediately. Move to a safe location." The input is text data, and the output is the response text.
[0848] Step 5:
[0849] The server uses an emotion engine to analyze the user's emotional state.
[0850] Specifically, it evaluates the tone, speed, and intonation of the voice to determine if the user is in a state of panic. The input is voice data, and the output is the result of the emotion analysis.
[0851] Step 6:
[0852] Based on the sentiment analysis results, the server adjusts the response.
[0853] For example, if the user is in a state of panic, the tone of the response is calmed and the content is changed to something reassuring. The adjusted response would be, "Please don't worry, I'll contact the police immediately. Please move to a safe location." The input is the result of the sentiment analysis and the response text, and the output is the adjusted response text.
[0854] Step 7:
[0855] The generated responses are converted into audio data using a speech synthesis API.
[0856] Specifically, the Google Cloud Text-to-Speech API is used to convert the adjusted response text into audio data. The input is the adjusted response text, and the output is the generated audio data.
[0857] Step 8:
[0858] The server sends the audio data to the terminal.
[0859] This data transmission takes place over the internet and is sent to the user's terminal. The input is the generated audio data, and the output is the audio data sent to the terminal.
[0860] Step 9:
[0861] The device plays the audio data.
[0862] Users can receive responses from the system via audio through their smartphone's speaker or headphones. The input is the audio data sent to the device, and the output is the played audio.
[0863] Step 10:
[0864] The server will contact the emergency contact.
[0865] If the server determines that the user's emotional state is urgent, it sends a notification to a pre-configured emergency contact. The input is the result of the emotional analysis, and the output is the emergency notification that was sent.
[0866] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0867] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0868] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0869] [Third Embodiment]
[0870] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0871] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0872] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0873] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0874] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0875] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0876] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0877] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0878] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0879] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0880] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0881] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0882] The present invention will now describe embodiments for carrying out this invention. The present invention is a system in which a user can give operating instructions or ask questions using their voice and receive answers in voice. The system comprises a voice input means, a means for transmitting voice data to a server, a means for converting voice data into text data, a means for analyzing the text data to generate answers, a means for converting the generated answers into voice data, a means for transmitting voice data to a terminal, and a means for playing back the voice data.
[0883] System Configuration
[0884] 1. Voice input means
[0885] The user provides operating instructions and asks questions via voice to the terminal. This voice input method uses voice input devices such as microphones. In addition, voice recognition software is incorporated to capture the voice as digital data.
[0886] 2. Means for sending audio data to the server
[0887] The device sends the captured audio data to a cloud server or local server. This transmission takes place via an internet connection or local network.
[0888] 3. Means for converting audio data into text data
[0889] The server uses a speech recognition API (e.g., a cloud-based speech recognition service) to convert the received audio data into text data. This ensures that instructions and questions entered by voice are represented as text.
[0890] 4. Means for analyzing text data to generate responses
[0891] The server analyzes text data using a natural language processing model (e.g., pre-trained models such as BERT or GPT) to understand the intent behind the user's questions and instructions. It then refers to a specific database (e.g., instruction manuals or overview booklets) to generate appropriate answers to the questions.
[0892] 5. Means for converting the generated response into audio data
[0893] The generated responses are converted into audio data using a speech synthesis API (e.g., a cloud-based speech synthesis service). This provides the user with an audio response.
[0894] 6. Means for transmitting audio data to a terminal
[0895] The server sends the generated audio data to the terminal. This transmission takes place via an internet connection or a local network.
[0896] 7. Means for playing audio data
[0897] The device plays the received audio data. This allows the user to hear the response in audio. Audio output devices such as speakers or headphones are used.
[0898] Specific example
[0899] For example, if a user asks "The printer ink isn't coming out," the following process will be performed.
[0900] The user speaks into the voice input device and says, "The printer ink isn't coming out."
[0901] The device captures this audio data and sends it to the server.
[0902] The server converts the audio data into text data using a speech recognition API.
[0903] The text data is analyzed using a natural language processing model to search for relevant sections in the instruction manual related to the problem "the printer ink is not coming out."
[0904] The server generates the appropriate response from the instruction manual, for example, "You need to replace the ink cartridge. Open the printer cover, remove the old cartridge, and insert the new cartridge."
[0905] Convert this answer into speech data using a speech synthesis API.
[0906] Audio data is sent to the terminal, which then plays it back to the user through an audio output device.
[0907] In this way, users can ask questions and receive answers by voice without having to type. This provides an easy-to-use support system for users who have difficulty typing or who are visually impaired.
[0908] The following describes the processing flow.
[0909] Step 1:
[0910] The user inputs operating instructions or questions by voice into the terminal. For example, the user might say, "The printer ink isn't coming out."
[0911] Step 2:
[0912] The device uses its built-in microphone to capture the user's voice and temporarily stores it as digital audio data.
[0913] Step 3:
[0914] The device sends the captured audio data to the server. This transmission takes place via an internet connection.
[0915] Step 4:
[0916] The server sends the received audio data to a speech recognition API. Typically, the speech recognition API uses a cloud-based speech recognition service.
[0917] Step 5:
[0918] The server processes the text data received from the speech recognition API. For example, the text "The printer ink isn't coming out" is generated.
[0919] Step 6:
[0920] The server uses a natural language processing model to analyze text data. This model performs analysis to understand the user's intent and find relevant information.
[0921] Step 7:
[0922] The server searches a database of instruction manuals and summary booklets to obtain appropriate answers to user questions. For example, this might include questions such as "Providing instructions on how to replace ink cartridges."
[0923] Step 8:
[0924] The server generates answers in a user-friendly format based on the relevant section of the instruction manual. For example, it might generate text such as, "To replace the ink cartridge, first open the printer cover, remove the old cartridge, and then insert the new cartridge."
[0925] Step 9:
[0926] The server generates text responses, which are then sent to a speech synthesis API to be converted into audio data. This speech synthesis API also often uses a cloud-based service.
[0927] Step 10:
[0928] The server sends the generated audio data to the terminal. This transmission also takes place via the internet connection.
[0929] Step 11:
[0930] The device plays the audio data it received from the server. Using the device's speaker or a headset or other audio output device, it provides the user with an audio response. The message played is: "To replace the ink cartridge, first open the printer cover, remove the old cartridge, and then insert the new cartridge."
[0931] Through the steps described above, this system responds to user voice input with voice commands. Therefore, it is easy to use for users who have difficulty with text input or those with visual impairments.
[0932] (Example 1)
[0933] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0934] In recent years, voice-based interfaces have attracted attention, becoming a particularly important technology for users with visual impairments or those who have difficulty with text input. However, current voice response systems face several challenges. For example, accurately converting voice input data into text and then quickly analyzing it to generate appropriate responses is difficult. Furthermore, there is a need to convert the generated responses into natural and easy-to-understand speech. In addition, optimization is required to perform these processes in real time. This invention aims to solve these problems and provide quick and accurate voice responses to user questions and operational instructions.
[0935] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0936] In this invention, the server includes means for understanding the intent of questions and operational instructions by analyzing voice data into text data and the text data; means for generating appropriate answers from a database based on the analysis results; and means for converting the generated answer text into voice data. This makes it possible to quickly and accurately perform text analysis and answer generation for questions and instructions entered by the user via voice, and to provide those answers as natural and easy-to-understand voice.
[0937] "Voice input means" refers to a device that allows users to ask questions or give instructions using their voice.
[0938] "Means for sending audio data to a server" refers to the technology for sending captured audio data to a server via a network.
[0939] "Means of converting audio data to text data" refers to technology that converts captured audio data into text format.
[0940] "Methods for analyzing text data and generating responses" refers to technologies that analyze data expressed as text and generate appropriate responses to user questions or instructions.
[0941] "Means of converting generated responses into audio data" refers to technology that converts text-based responses into audio data.
[0942] "Means for transmitting audio data to a terminal" refers to the technology for transmitting generated audio data to a user's terminal via a network.
[0943] "Means of playing audio data" refers to the devices and software used to play audio data on the user's device.
[0944] "Means of converting to digital audio data in real time" refers to technology that converts user voice input into digital data in real time.
[0945] "Text data and means for analyzing that text data" refers to technologies for analyzing converted text data and understanding its intent and content.
[0946] "Means for generating appropriate answers from a database" refers to technologies that refer to a database based on analysis results and extract appropriate answers to user questions and operational instructions.
[0947] "Methods for converting response text into audio data" refers to technologies that convert generated responses into natural and easy-to-understand audio data.
[0948] This invention is a system that allows users to give operating instructions and ask questions using their voice and receive answers in voice. This system combines a voice input device, a network, a server, a natural language processing model, and speech synthesis technology.
[0949] Voice input method
[0950] The user uses a microphone to speak into the device, asking questions or giving instructions. The device is equipped with speech recognition software, such as the Google Speech-to-Text API, which converts the voice into digital audio data in real time. For example, if the user says, "The printer ink isn't coming out," this voice is captured as digital data.
[0951] Sending audio data
[0952] The device transmits the captured audio data to the server via an internet connection or local network. This transmission process uses encryption technologies such as TLS to ensure data security.
[0953] Converting audio data to text
[0954] The server receives the incoming audio data and uses the Amazon Transcribe API to convert the audio into text data. For example, the user's audio data, "The printer ink isn't coming out," is converted into the text format "The printer ink isn't coming out."
[0955] Text data analysis
[0956] The server analyzes the converted text data using natural language processing models (e.g., GPT-3, BERT) to understand the intent behind questions and instructions. For example, the server analyzes the text data "The printer ink isn't coming out" and understands its meaning. This process includes generating text tokens and interpreting intent.
[0957] Answer generation
[0958] Based on the analysis results, the server extracts and generates the appropriate answer from a database (e.g., a printer instruction manual database). For example, it might generate an answer such as, "You need to replace the ink cartridge. Open the printer cover, remove the old cartridge, and insert the new cartridge."
[0959] Voice conversion of the answer
[0960] The generated response text is converted into speech data using speech synthesis technology such as the Google Text-to-Speech API. This ensures that the user's requested response is produced as natural-sounding speech data.
[0961] Sending audio data
[0962] The server then transmits the generated audio data back to the terminal via the internet or local network. Security is ensured during this process by encrypting the data using TLS or similar protocols.
[0963] Playback of audio data
[0964] The device plays the received audio data through an audio output device such as a speaker or headphones. The user can then listen to this audio and take the necessary actions or responses. For example, the device's speaker might play an audio message saying, "You need to replace the ink cartridge. Open the printer cover, remove the old cartridge, and insert the new cartridge."
[0965] Specific example
[0966] For example, if a user asks "The printer ink isn't coming out," the following process will be performed.
[0967] The user speaks into the voice input device and says, "The printer ink isn't coming out."
[0968] The device captures this audio data and sends it to the server.
[0969] The server converts the audio data into text data using a speech recognition API.
[0970] The text data is analyzed using a natural language processing model to search for relevant sections in the instruction manual related to the problem "the printer ink is not coming out."
[0971] The server generates the appropriate response from the instruction manual, for example, "You need to replace the ink cartridge. Open the printer cover, remove the old cartridge, and insert the new cartridge."
[0972] Convert this answer into speech data using a speech synthesis API.
[0973] Audio data is sent to the terminal, which then plays it back to the user through an audio output device.
[0974] Example of a prompt
[0975] User input: "The printer ink isn't coming out."
[0976] System response: "You need to replace the ink cartridge. Open the printer cover, remove the old cartridge, and insert the new one."
[0977] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0978] Step 1: Capture voice input
[0979] The user speaks into a voice input device (e.g., a microphone) to ask a question or give an instruction. The device converts this voice into digital speech data using the Google Speech-to-Text API. The input is the user's utterance, "The printer ink isn't coming out," and the output is the text data of this utterance. Specifically, the device captures the user's voice in real time and converts it into a digital signal.
[0980] Step 2: Sending the audio data
[0981] The terminal sends the generated digital audio data to the cloud server using encryption technology such as TLS. The input is digital audio data, and the output is the transmission of encrypted data. Specifically, the terminal sends the data to the server via an internet connection or a local network.
[0982] Step 3: Converting audio data to text
[0983] The server converts received audio data into text data using the Amazon Transcribe API. The input is digital audio data, and the output is the corresponding text data. Specifically, the server receives audio data and makes an API request to convert it into text data.
[0984] Step 4: Analyzing Text Data
[0985] The server receives text data and performs analysis using a natural language processing model (e.g., BERT, GPT-3). The input is text data, and the output is semantic understanding data resulting from the analysis. Specifically, the server tokenizes the text and uses the model to analyze the intent of questions and instructions.
[0986] Step 5: Generating the answer
[0987] Based on the analysis results, the server extracts and generates appropriate answers from a database (e.g., an instruction manual database). The input is semantic comprehension data, and the output is the generated answer text. Specifically, the server searches the queried database, extracts matching answers, and assembles them.
[0988] Step 6: Voice conversion of the answer
[0989] The generated response text is converted into speech data using the Google Text-to-Speech API. The input is the response text, and the output is the corresponding speech data. Specifically, the server makes an API request to convert the text to speech.
[0990] Step 7: Sending audio data
[0991] The server re-encrypts the generated audio data using TLS or similar methods and sends it to the terminal. The input is the audio data, and the output is the transmission of the encrypted data. Specifically, the server sends the data back to the terminal via the network.
[0992] Step 8: Play back audio data
[0993] The device decodes the received audio data and plays it back through an audio output device such as a speaker or headphones. The input is audio data, and the output is the audio that the user hears. Specifically, the device decodes the encryption and outputs the audio using an audio playback device.
[0994] (Application Example 1)
[0995] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0996] Traditional food delivery systems require users to place orders via a screen, which presents challenges for users with dirty hands or those with visual impairments. Furthermore, checking order status and estimated delivery times also requires manual operation, resulting in a lack of convenience.
[0997] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0998] In this invention, the server includes a voice input means, a means for transmitting voice data to the server, a means for converting voice data into text data, a means for analyzing the text data to generate a response corresponding to the order, a means for converting the generated response into voice data, a means for transmitting voice data to a terminal, and a means for playing back the voice data. This allows users to order food and check the order status using only their voice, improving convenience for people with visual impairments and users whose hands are occupied.
[0999] A "voice input method" refers to a device or system that allows a user to give instructions by voice.
[1000] "Means for sending audio data to a server" refers to the mechanisms and protocols for sending captured audio data to a server via the internet.
[1001] "Means for converting audio data to text data" refers to technologies for converting audio data captured using speech recognition technology into text format.
[1002] "A means of analyzing text data to generate responses corresponding to orders" refers to a system that uses natural language processing technology to analyze text data and generate appropriate responses from the obtained information.
[1003] "Means for converting generated responses into audio data" refers to a mechanism or system that converts text-based responses into audio data using speech synthesis technology.
[1004] "Means for transmitting audio data to a terminal" refers to the mechanisms and protocols for transmitting generated audio data to a terminal used by the user.
[1005] "Means for playing audio data" refers to devices or systems that play transmitted audio data and allow the user to listen to it.
[1006] "Data related to food and beverage orders" refers to data including menus, prices, and restaurant information handled within the food delivery system.
[1007] A "natural language processing model" is a machine learning technique or algorithm used to analyze user text data, understand user intent, and generate appropriate responses.
[1008] "Responses regarding orders and delivery status" refers to responses that include information about the products ordered by the user and their delivery status.
[1009] The system for carrying out this invention allows users to place food delivery orders and check the order status using voice commands. The system includes the following main components:
[1010] 1. Voice input means
[1011] Users input voice commands, such as ordering food or checking delivery status, by speaking into the microphone of their device (smartphone or smart glasses). This voice input method utilizes the built-in microphone of the smartphone or smart glasses, and captures the voice as digital data using the Google Speech-to-Text API.
[1012] 2. Means for sending audio data to the server
[1013] The device sends the captured audio data to a cloud server via an internet connection. HTTPS communication is used for data transmission.
[1014] 3. Means for converting audio data into text data
[1015] The server converts the received audio data into text data using speech recognition technology. Specifically, it uses the Google Speech-to-Text API to convert the audio data into text format.
[1016] 4. A means of analyzing text data to generate responses corresponding to orders.
[1017] The server analyzes the data, which has been converted to text format, using a natural language processing model. It utilizes advanced pre-trained models such as GPT-4 to understand user intent and generate appropriate responses. In this process, it references a food delivery database, providing data such as menu information, pricing information, and estimated delivery times.
[1018] 5. Means for converting the generated response into audio data
[1019] The server converts the generated text-based responses into speech data using speech synthesis technology. This conversion utilizes the Google Text-to-Speech API.
[1020] 6. Means for transmitting audio data to a terminal
[1021] The server then sends the generated audio data back to the terminal. This is also done via HTTPS communication over the internet connection.
[1022] 7. Means for playing audio data
[1023] Finally, the device plays back the received audio data and provides the user with an answer. A smartphone, smart glasses speaker, or headphones are used as the playback device.
[1024] Adding specific examples
[1025] For example, if a user uses voice input to say "I want to order a pizza," the following process will occur.
[1026] The user says, "I want to order a pizza."
[1027] The device captures this audio and sends it to the server.
[1028] The server uses the Google Speech-to-Text API to convert the audio data into text data.
[1029] A natural language processing model (e.g., GPT-4) is used to analyze text data such as "I want to order a pizza" and provide information on appropriate restaurants and menus.
[1030] The server generates an answer such as "Which restaurant would you order pizza from?", and this answer is converted into speech data using the Google Text-to-Speech API.
[1031] The server generates audio data and sends it to the terminal, which then plays the audio.
[1032] Example of a prompt
[1033] Analyze user comments and suggest corresponding food delivery restaurants and menus: {User comments}
[1034] Follow user instructions, confirm the order, and communicate with the backend to check its status.
[1035] In this way, users can perform a series of operations, from ordering food delivery to checking the delivery status, using only their voice, without using a visual interface. This significantly improves convenience for users whose hands are full or who have visual impairments.
[1036] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1037] Step 1:
[1038] The user speaks into the device's voice input and says, "I want to order a pizza." This captures the audio data. The input is what the user said, and the output is the captured audio data.
[1039] Step 2:
[1040] The device sends the captured audio data to the cloud server via the internet connection. The input is the captured audio data, and the output is the audio data received on the server side. Specifically, the operation involves sending the audio data using HTTPS communication.
[1041] Step 3:
[1042] The server converts the received audio data into text data using the Google Speech-to-Text API. The input is the received audio data, and the output is text data. Specifically, it analyzes the audio data using speech recognition technology and converts it into the corresponding text.
[1043] Step 4:
[1044] The server analyzes the converted text data using a natural language processing model (e.g., GPT-4). The input is text data, and the output is the user's intent or question content as a result of the analysis. Specifically, it analyzes the text using a generative AI model and understands the content of the corresponding instructions or questions.
[1045] Step 5:
[1046] The server generates appropriate restaurant and menu information based on the analysis results and creates a text-based response. The input is the analysis results, and the output is the generated text-based response. Specifically, it refers to an internal restaurant database, extracts appropriate information to meet the user's request, and generates a response.
[1047] Step 6:
[1048] The server converts the generated text-based response into audio data using the Google Text-to-Speech API. The input is a text-based response, and the output is audio data. Specifically, it uses speech synthesis technology to convert text into speech.
[1049] Step 7:
[1050] The server sends the generated audio data back to the terminal. The input is the generated audio data, and the output is the audio data received by the terminal. Specifically, the operation involves sending the audio data to the terminal using HTTPS communication.
[1051] Step 8:
[1052] The device plays the received audio data and provides a response to the user. The input is the received audio data, and the output is the audio information to be heard by the user. Specifically, the device plays the audio data using its speaker or headphones.
[1053] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1054] The embodiments for carrying out this invention will be described in detail. The present invention combines an emotion recognition function with a system in which a user inputs operation instructions or questions by voice and receives answers by voice. The system comprises a voice input means, a means for transmitting voice data to a server, a means for converting voice data into text data, a means for analyzing the text data to generate answers, a means for converting the generated answers into voice data, a means for transmitting voice data to a terminal, a means for playing back the voice data, and an emotion engine that recognizes the user's emotions.
[1055] System Configuration
[1056] 1. Voice input means
[1057] The user inputs operating instructions or questions by voice into the terminal. For example, the user might say, "The printer ink isn't coming out." This voice input method uses a microphone to capture the voice data and saves it as digital audio data.
[1058] 2. Means for sending audio data to the server
[1059] The device sends audio data captured via voice input to a cloud server or local server. This transmission takes place via an internet connection.
[1060] 3. Means for converting audio data into text data
[1061] The server uses a speech recognition API to convert the received audio data into text data. The speech recognition API generates the text "The printer ink isn't coming out."
[1062] 4. Means for analyzing text data to generate responses
[1063] The server uses a natural language processing model to analyze text data and understand the intent behind user questions and instructions. It then consults a database of instruction manuals and overview booklets to generate appropriate answers to the questions.
[1064] 5. Means for converting the generated response into audio data
[1065] The server converts the generated response into audio data using a speech synthesis API. This speech synthesis API then prepares the response for the user as audio.
[1066] 6. Means for transmitting audio data to a terminal
[1067] The server sends the generated audio data to the terminal. This transmission also takes place via the internet connection.
[1068] 7. Means for playing audio data
[1069] The device plays the received audio data. It then provides the user with the answer via voice using a speaker or headphones.
[1070] 8. Emotional Engine
[1071] The emotion engine analyzes and recognizes the user's emotional state using voice data. This emotion analysis is based on features such as voice tone, speed, and intonation.
[1072] Specific example
[1073] For example, the sequence of operations when a user asks "The printer ink isn't coming out" is as follows:
[1074] The user speaks into the voice input device and says, "The printer ink isn't coming out."
[1075] The device captures this audio and sends the audio data to the server.
[1076] The server converts the audio data into text data using a speech recognition API.
[1077] The server uses a natural language processing model to analyze text data and search for information related to the problem "the printer ink isn't coming out."
[1078] The server generates appropriate responses to the user based on the instruction manual. For example, it might generate text such as, "To replace the ink cartridge, first open the printer cover, remove the old cartridge, and then insert the new cartridge."
[1079] The text response generated by the server is converted into speech data using a speech synthesis API.
[1080] The server sends the audio data to the terminal.
[1081] The device plays audio data and provides the user with an answer.
[1082] The emotion engine analyzes the user's original voice data, and if it determines that the user is confused, the server modifies the response to be more helpful and detailed. For example, it might adjust the tone and content of the response to something like, "Don't worry, first try opening the printer cover. Next, carefully remove the old ink cartridge. Then, firmly insert the new cartridge."
[1083] Through these steps, the system can provide flexible responses tailored to the user's emotional state. This allows users to use the system more comfortably and effectively.
[1084] The following describes the processing flow.
[1085] Step 1:
[1086] The user inputs operating instructions or questions by voice into the terminal. For example, the user might say, "The printer ink isn't coming out."
[1087] Step 2:
[1088] The device uses its built-in microphone to capture the user's voice and temporarily stores it as digital audio data.
[1089] Step 3:
[1090] The device sends the captured audio data to the server. This transmission takes place via an internet connection.
[1091] Step 4:
[1092] The server sends audio data to a speech recognition API, which then converts the audio data into text data. For example, a cloud-based speech recognition service could be used.
[1093] Step 5:
[1094] The server processes the text data "Printer ink is not coming out" received from the speech recognition API.
[1095] Step 6:
[1096] The server uses a natural language processing model to analyze the text data. This analysis helps the server understand what the user wants. For example, it might determine that the user wants to solve a printer problem.
[1097] Step 7:
[1098] The server searches a database of instruction manuals and summary booklets to obtain appropriate answers to user questions. For example, it might find information such as, "The ink cartridge needs to be replaced."
[1099] Step 8:
[1100] The server generates appropriate responses for the user based on the instruction manual. For example, it might generate text such as, "To replace the ink cartridge, first open the printer cover, remove the old cartridge, and then insert the new cartridge."
[1101] Step 9:
[1102] The server sends the generated text response to a speech synthesis API, which converts it into audio data. This speech synthesis API then prepares the response for the user as audio.
[1103] Step 10:
[1104] The server sends the generated audio data to the terminal. This transmission also takes place via the internet connection.
[1105] Step 11:
[1106] The device plays the audio data it received from the server. Using the device's speaker or headphones, it provides the user with an audio response. The message played is: "To replace the ink cartridge, first open the printer cover, remove the old cartridge, and then insert the new cartridge."
[1107] Step 12:
[1108] The emotion engine analyzes the original audio data and evaluates the user's emotional state. For example, it can detect feelings of confusion or frustration.
[1109] Step 13:
[1110] The server adjusts the responses generated based on the emotion engine's analysis. For example, if the server detects that the user is confused, it adjusts the response to be more helpful and detailed. It might change to something like, "Don't worry, first try opening the printer cover. Next, carefully remove the old ink cartridge. Then, firmly insert the new cartridge."
[1111] Step 14:
[1112] The server sends the newly processed audio data back to the speech synthesis API, where it is converted back into audio data.
[1113] Step 15:
[1114] The server sends new audio data to the device, and the device plays this data. The user can then hear the corrected, gentler-toned response.
[1115] (Example 2)
[1116] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1117] Conventional voice dialogue systems can provide appropriate answers to user inputs, such as operational instructions and questions, but they struggle to provide flexible responses that adapt to the user's emotional state. Furthermore, they lacked the means to recognize and appropriately respond to user feelings of confusion or anxiety. As a result, the user experience was not always optimal.
[1118] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1119] In this invention, the server includes means for converting voice data into text data, means for analyzing the text data to generate a response, and means for converting the generated response back into voice data. This enables flexible response generation that takes into account the emotional state of the user based on voice input.
[1120] "Voice input means" refers to a device that includes a microphone for the user to input voice, and a means for capturing that voice as digital data.
[1121] "Means for sending audio data to a server" refers to means for sending captured audio data to a cloud server or local server via the internet.
[1122] "Means for converting audio data to text data" refers to methods for converting audio data received using a speech recognition API into text data.
[1123] "Methods for analyzing text data and generating responses" refers to methods for analyzing text data using natural language processing models and generating appropriate responses.
[1124] "Means for converting generated responses into audio data" refers to means for converting text responses generated using a speech synthesis API into audio data.
[1125] "Means for transmitting audio data to a terminal" refers to means for transmitting generated audio data to a user's terminal via the internet.
[1126] "Means for playing audio data" refers to the means by which a terminal plays back audio data it has received through speakers or headphones.
[1127] An "emotion engine" is a means of analyzing and recognizing a user's emotional state using voice data and responding accordingly.
[1128] A "natural language processing model" is an algorithm or software that analyzes text data to understand the meaning and intent of human language.
[1129] A "speech recognition API" is an application programming interface for converting speech data into text data.
[1130] A "speech synthesis API" is an application programming interface for converting text data into speech data.
[1131] The embodiments for carrying out this invention will be described in detail. The present invention is a system including voice input means, means for transmitting voice data to a server, means for converting voice data to text data, means for analyzing text data to generate a response, means for converting the generated response to voice data, means for transmitting voice data to a terminal, means for playing back voice data, and an emotion engine for recognizing the user's emotions.
[1132] System Configuration
[1133] 1. Voice input means
[1134] The user inputs operating instructions or questions by voice into the device. For example, the user might say, "The printer ink isn't coming out." This voice input method uses a microphone to capture the voice data and saves it as digital audio data. Specific hardware that can be used include smartphones, tablets, and PCs with microphones.
[1135] 2. Means for sending audio data to the server
[1136] The device sends the audio data captured via voice input to a cloud server or local server. This transmission takes place over the internet. HTTPS is used as the specific transmission protocol.
[1137] 3. Means for converting audio data into text data
[1138] The server uses a speech recognition API (for example, Google Cloud Speech-to-Text API) to convert the received audio data into text data. The speech recognition API generates the text "The printer ink is not coming out."
[1139] 4. Means for analyzing text data to generate responses
[1140] The server uses a natural language processing model (e.g., OpenAI's GPT-3) to analyze text data and understand the intent behind user questions and instructions. It then consults a database of instruction manuals and overview booklets to generate appropriate answers to the questions.
[1141] 5. Means for converting the generated response into audio data
[1142] The server converts the generated response into audio data using a speech synthesis API (e.g., Amazon Polly). This speech synthesis API then provides the response to the user as audio.
[1143] 6. Means for transmitting audio data to a terminal
[1144] The server sends the generated audio data to the terminal. This transmission also takes place via the internet connection. HTTPS is used as the transmission protocol.
[1145] 7. Means for playing audio data
[1146] The device plays back the received audio data. It provides the user with an audio response using speakers or headphones. Specific hardware options include smartphones, tablets, and speakers or headphones connected to a PC.
[1147] 8. Emotional Engine
[1148] The emotion engine analyzes and recognizes the user's emotional state using voice data. This emotion analysis is based on features such as voice tone, speed, and intonation. Specific software such as an emotion analysis engine (e.g., IBM Watson Tone Analyzer) can be used. Based on the analysis results, the server adjusts its responses and generates more helpful responses as needed.
[1149] Specific example
[1150] The following is the sequence of events that occurs when a user asks, "The printer ink isn't coming out":
[1151] 1. The user speaks into the device and says, "The printer ink isn't coming out."
[1152] 2. The device captures this audio and sends the audio data to the server.
[1153] 3. The server converts the audio data into text data using a speech recognition API.
[1154] 4. The server uses a natural language processing model to analyze the text data and search for information related to the problem "the printer ink is not coming out."
[1155] 5. The server generates appropriate responses to the user based on the instruction manual. For example, it might generate text such as, "To replace the ink cartridge, first open the printer cover, remove the old cartridge, and then insert the new cartridge."
[1156] 6. The text response generated by the server is converted into speech data using a speech synthesis API.
[1157] 7. The server sends the audio data to the terminal.
[1158] 8. The device plays audio data and provides the user with an answer.
[1159] 9. The emotion engine analyzes the user's original voice data and, if it determines that the user is confused, the server modifies the response to be more helpful and detailed. For example, it might adjust the tone and content of the response to something like, "Don't worry, first try opening the printer cover. Next, remove the old ink cartridge. Then, carefully insert the new cartridge."
[1160] Examples of prompt statements
[1161] Examples of prompt statements are as follows:
[1162] User: What should I do if my printer ink isn't coming out?
[1163] System: To replace an ink cartridge, first open the printer cover, remove the old cartridge, and then install the new cartridge.
[1164] By inputting this prompt into the AI model, the system generates an appropriate answer to the user's question. This mechanism allows users to use the system more comfortably and effectively.
[1165] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1166] Step 1:
[1167] Voice input
[1168] Action: The user speaks into the terminal saying, "The printer ink isn't coming out."
[1169] Input: User's voice.
[1170] Output: Digital audio data.
[1171] Details: The microphone built into the device captures the audio and saves it as digital audio data. Examples of devices that can be used include smartphones, tablets, and PCs with built-in microphones.
[1172] Step 2:
[1173] Sending audio data
[1174] Operation: The device sends audio data to the server.
[1175] Input: Digital audio data.
[1176] Output: Audio data transferred to the server.
[1177] Details: The device sends the captured audio data to a cloud server or local server via the internet. This communication uses the HTTPS protocol and is secure.
[1178] Step 3:
[1179] Convert audio data to text data
[1180] Operation: The server uses a speech recognition API to convert speech data into text data.
[1181] Input: Audio data transferred to the server.
[1182] Output: Text data. The content is "The printer ink isn't coming out."
[1183] Details: The server calls the Google Cloud Speech-to-Text API to convert the received audio data into text data.
[1184] Step 4:
[1185] Analyze text data to generate answers.
[1186] Operation: The server uses a natural language processing model to analyze text data and generate appropriate responses.
[1187] Input: Converted text data.
[1188] Output: Text data of the response.
[1189] Details: The server uses OpenAI's GPT-3 model to analyze the content of text data and understand the user's question and intent. It then refers to a database of instruction manuals and overview booklets to generate appropriate answers to the questions. A specific example answer is: "To replace the ink cartridge, first open the printer cover, remove the old cartridge, and then install the new cartridge."
[1190] Step 5:
[1191] Convert the generated response into audio data.
[1192] Operation: The server converts the generated response into audio data.
[1193] Input: Text data of the response.
[1194] Output: Audio data.
[1195] Details: The server calls the Amazon Polly API to convert the generated text responses into audio data.
[1196] Step 6:
[1197] Sending audio data
[1198] Operation: The server sends the generated audio data to the terminal.
[1199] Input: Audio data.
[1200] Output: Audio data transferred to the terminal.
[1201] Details: The server sends the generated audio data to the user's terminal via the internet. This communication also uses the HTTPS protocol, ensuring security.
[1202] Step 7:
[1203] Playback of audio data
[1204] Operation: Plays audio data received by the device.
[1205] Input: Audio data transferred to the device.
[1206] Output: Played audio.
[1207] Details: The device decodes the received audio data and plays the audio through the speaker or headphones, thereby providing the user with a response.
[1208] Step 8:
[1209] Sentiment analysis and response adjustment
[1210] Operation: The server uses an emotion engine to analyze the user's emotional state and adjusts its responses as needed.
[1211] Input: Audio data and its analysis results.
[1212] Output: Adjusted response data.
[1213] Details: The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze the tone, speed, and intonation of the user's voice data. Based on the analysis, if it is determined that the user is confused, the original response is modified to be more helpful and detailed. For example, the response's tone and content might be adjusted to something like, "Don't worry, first try opening the printer cover. Next, remove the old ink cartridge. Then, carefully insert the new cartridge."
[1214] (Application Example 2)
[1215] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1216] Current voice control systems fail to recognize the user's emotional state, resulting in problems such as being unable to respond appropriately when the user is confused or in an emergency. This is especially true in security services, where users may act inappropriately when panicked, requiring a swift and accurate response. However, existing systems cannot meet these needs.
[1217] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice input means, a means for transmitting voice data to the server, and a means for converting voice data into text data. This makes it possible to convert the user's voice input into text and analyze it. Furthermore, it includes a means for analyzing the text data to generate a response, a means for converting the generated response into voice data, an emotion engine for analyzing and recognizing the emotional state, a means for adjusting the response based on the emotion analysis results, and a means for contacting emergency contacts. This enables flexible responses based on the user's emotional state and rapid emergency response as needed.
[1218] "Voice input means" refers to devices or functions that allow users to input instructions or questions into a system via voice.
[1219] "Means of sending audio data to a server" refers to the function of sending captured audio data to a cloud server or local server via the internet or other means.
[1220] "Means for converting audio data to text data" refers to devices or functions that use speech recognition technology to convert audio data into text data.
[1221] "Means for analyzing text data and generating responses" refers to a function or technical means that analyzes input text data and generates appropriate responses.
[1222] "Means for converting generated responses into audio data" refers to a function that converts text-based responses into audio data using speech synthesis technology.
[1223] "Means for transmitting audio data to a terminal" refers to communication means for transmitting generated audio data to the user's terminal.
[1224] "Means for playing audio data" refers to the function of playing audio data received by a device through a speaker or headphones.
[1225] An "emotion engine that analyzes and recognizes the user's emotional state" is a software engine that analyzes and recognizes a user's emotions based on characteristics such as tone, speed, and intonation of their voice.
[1226] "Means for adjusting responses based on emotion analysis results" refers to a function that appropriately adjusts the tone and content of responses based on emotion analysis results obtained by the emotion engine.
[1227] The "means of contacting emergency contacts" feature is a function that sends a notification to pre-set emergency contacts when the user's emotional state is determined to be an emergency situation.
[1228] This invention is a system that combines user voice input and emotion recognition, and is particularly applicable to emergency response in security services. Specific embodiments for carrying out the invention are described below.
[1229] System Overview
[1230] This system includes means for voice input, means for transmitting voice data to a server, means for converting voice data into text data, means for analyzing text data to generate a response, means for converting the generated response into voice data, means for transmitting voice data to a terminal, means for playing back voice data, an emotion engine for analyzing and recognizing the user's emotional state, and means for adjusting the response based on the emotion analysis results. Furthermore, it also includes means for contacting emergency contacts in emergencies.
[1231] Hardware and software configuration
[1232] This system consists of the following hardware and software:
[1233] Audio input method: Use an audio capture device such as a microphone.
[1234] Server: Performs necessary data processing on cloud servers or local servers.
[1235] Speech recognition APIs such as Google Cloud Speech-to-Text are used to convert speech data into text data.
[1236] Natural language processing model: A model for understanding user questions and instructions, such as using the Google Cloud Natural Language API.
[1237] Text-to-speech APIs such as Google Cloud Text-to-Speech are used to convert text data into speech data.
[1238] Emotion Engine: Software that analyzes the tone, speed, intonation, etc. of a voice to recognize the user's emotions.
[1239] Emergency contact method: A means of communication used to send notifications to designated contacts in the event of an emergency.
[1240] Detailed operation
[1241] Voice input and data transmission
[1242] First, the user speaks a question or instruction into the smartphone's microphone. For example, a prompt might say, "There might be a suspicious person in my house, what should I do?" This audio data is captured and sent to the server.
[1243] Speech recognition and text conversion
[1244] The server converts the transmitted audio data into text data using a speech recognition API. The Google Cloud Speech-to-Text API is used, and this process generates the text "There might be a suspicious person in my house, what should I do?"
[1245] Text analysis and response generation
[1246] The text data is analyzed using a natural language processing model to generate appropriate responses. For example, in response to the above prompt, the following response is generated: "Stay calm. I will contact the police immediately. Move to a safe location."
[1247] Emotion recognition and response adjustment
[1248] Furthermore, this response is adjusted by an emotion engine according to the user's emotional state. If the tone and speed of the voice indicate that the user is in a state of panic, the tone of the response will be changed to a calmer and more reassuring one.
[1249] Speech synthesis and data transmission
[1250] The generated responses are converted into audio data using a speech synthesis API and sent from the server to the device. The Google Cloud Text-to-Speech API performs this conversion.
[1251] Audio playback
[1252] The user's device plays the received audio data. Playback occurs through speakers or headphones, and the user is provided with a response.
[1253] Emergency response
[1254] If the system determines that a user's emotional state is urgent, it sends a notification to a pre-configured emergency contact. This allows the user to receive appropriate assistance immediately.
[1255] Specific example
[1256] Let's explain the detailed operation using the previously shown prompt, "There might be a suspicious person in my house, what should I do?" as an example. When the user speaks this question into the voice input device, the voice data is captured and sent to the server. The server converts the voice data into text data and analyzes it using a natural language processing model. The generated response will be, "Please stay calm. I will contact the police immediately. Please move to a safe place." The emotion engine recognizes the user's panic state and further adjusts the response to be calmer. The response, converted into voice data using a speech synthesis API, is then sent to the device and played back. Simultaneously, a notification is sent to emergency contacts, allowing the user to receive quick assistance from the difficult situation.
[1257] The above describes a specific embodiment for carrying out this invention. By using this system, users can obtain answers to their questions and instructions safely and quickly, and can also respond in emergencies.
[1258] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1259] Step 1:
[1260] The user speaks questions or instructions using a voice input device.
[1261] As a concrete example, the user speaks into their smartphone's microphone, "There might be a suspicious person in my house, what should I do?" The input data is captured as audio data.
[1262] Step 2:
[1263] The device sends the captured audio data to the server.
[1264] In this process, the audio data is sent via the internet to a cloud server or a local server. The input is the captured audio data, and the output is the audio data sent to the server.
[1265] Step 3:
[1266] The server converts the audio data into text data using a speech recognition API.
[1267] Specifically, the Google Cloud Speech-to-Text API is used to generate text data for the sentence, "There might be a suspicious person in my house, what should I do?". The input is audio data, and the output is text data.
[1268] Step 4:
[1269] The server uses a natural language processing model to analyze text data and generate appropriate responses.
[1270] This example uses the Google Cloud Natural Language API to analyze the text "There might be a suspicious person in my house, what should I do?". Based on the analysis, it generates the response "Stay calm. I will contact the police immediately. Move to a safe location." The input is text data, and the output is the response text.
[1271] Step 5:
[1272] The server uses an emotion engine to analyze the user's emotional state.
[1273] Specifically, it evaluates the tone, speed, and intonation of the voice to determine if the user is in a state of panic. The input is voice data, and the output is the result of the emotion analysis.
[1274] Step 6:
[1275] Based on the sentiment analysis results, the server adjusts the response.
[1276] For example, if the user is in a state of panic, the tone of the response is calmed and the content is changed to something reassuring. The adjusted response would be, "Please don't worry, I'll contact the police immediately. Please move to a safe location." The input is the result of the sentiment analysis and the response text, and the output is the adjusted response text.
[1277] Step 7:
[1278] The generated responses are converted into audio data using a speech synthesis API.
[1279] Specifically, the Google Cloud Text-to-Speech API is used to convert the adjusted response text into audio data. The input is the adjusted response text, and the output is the generated audio data.
[1280] Step 8:
[1281] The server sends the audio data to the terminal.
[1282] This data transmission takes place over the internet and is sent to the user's terminal. The input is the generated audio data, and the output is the audio data sent to the terminal.
[1283] Step 9:
[1284] The device plays the audio data.
[1285] Users can receive responses from the system via audio through their smartphone's speaker or headphones. The input is the audio data sent to the device, and the output is the played audio.
[1286] Step 10:
[1287] The server will contact the emergency contact.
[1288] If the server determines that the user's emotional state is urgent, it sends a notification to a pre-configured emergency contact. The input is the result of the emotional analysis, and the output is the emergency notification that was sent.
[1289] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1290] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1291] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1292] [Fourth Embodiment]
[1293] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1294] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1295] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1296] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1297] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1298] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1299] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1300] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1301] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1302] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1303] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1304] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1305] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1306] The present invention will now describe embodiments for carrying out this invention. The present invention is a system in which a user can give operating instructions or ask questions using their voice and receive answers in voice. The system comprises a voice input means, a means for transmitting voice data to a server, a means for converting voice data into text data, a means for analyzing the text data to generate answers, a means for converting the generated answers into voice data, a means for transmitting voice data to a terminal, and a means for playing back the voice data.
[1307] System Configuration
[1308] 1. Voice input means
[1309] The user provides operating instructions and asks questions via voice to the terminal. This voice input method uses voice input devices such as microphones. In addition, voice recognition software is incorporated to capture the voice as digital data.
[1310] 2. Means for sending audio data to the server
[1311] The device sends the captured audio data to a cloud server or local server. This transmission takes place via an internet connection or local network.
[1312] 3. Means for converting audio data into text data
[1313] The server uses a speech recognition API (e.g., a cloud-based speech recognition service) to convert the received audio data into text data. This ensures that instructions and questions entered by voice are represented as text.
[1314] 4. Means for analyzing text data to generate responses
[1315] The server analyzes text data using a natural language processing model (e.g., pre-trained models such as BERT or GPT) to understand the intent behind the user's questions and instructions. It then refers to a specific database (e.g., instruction manuals or overview booklets) to generate appropriate answers to the questions.
[1316] 5. Means for converting the generated response into audio data
[1317] The generated responses are converted into audio data using a speech synthesis API (e.g., a cloud-based speech synthesis service). This provides the user with an audio response.
[1318] 6. Means for transmitting audio data to a terminal
[1319] The server sends the generated audio data to the terminal. This transmission takes place via an internet connection or a local network.
[1320] 7. Means for playing audio data
[1321] The device plays the received audio data. This allows the user to hear the response in audio. Audio output devices such as speakers or headphones are used.
[1322] Specific example
[1323] For example, if a user asks "The printer ink isn't coming out," the following process will be performed.
[1324] The user speaks into the voice input device and says, "The printer ink isn't coming out."
[1325] The device captures this audio data and sends it to the server.
[1326] The server converts the audio data into text data using a speech recognition API.
[1327] The text data is analyzed using a natural language processing model to search for relevant sections in the instruction manual related to the problem "the printer ink is not coming out."
[1328] The server generates the appropriate response from the instruction manual, for example, "You need to replace the ink cartridge. Open the printer cover, remove the old cartridge, and insert the new cartridge."
[1329] Convert this answer into speech data using a speech synthesis API.
[1330] Audio data is sent to the terminal, which then plays it back to the user through an audio output device.
[1331] In this way, users can ask questions and receive answers by voice without having to type. This provides an easy-to-use support system for users who have difficulty typing or who are visually impaired.
[1332] The following describes the processing flow.
[1333] Step 1:
[1334] The user inputs operating instructions or questions by voice into the terminal. For example, the user might say, "The printer ink isn't coming out."
[1335] Step 2:
[1336] The device uses its built-in microphone to capture the user's voice and temporarily stores it as digital audio data.
[1337] Step 3:
[1338] The device sends the captured audio data to the server. This transmission takes place via an internet connection.
[1339] Step 4:
[1340] The server sends the received audio data to a speech recognition API. Typically, the speech recognition API uses a cloud-based speech recognition service.
[1341] Step 5:
[1342] The server processes the text data received from the speech recognition API. For example, the text "The printer ink isn't coming out" is generated.
[1343] Step 6:
[1344] The server uses a natural language processing model to analyze text data. This model performs analysis to understand the user's intent and find relevant information.
[1345] Step 7:
[1346] The server searches a database of instruction manuals and summary booklets to obtain appropriate answers to user questions. For example, this might include questions such as "Providing instructions on how to replace ink cartridges."
[1347] Step 8:
[1348] The server generates answers in a user-friendly format based on the relevant section of the instruction manual. For example, it might generate text such as, "To replace the ink cartridge, first open the printer cover, remove the old cartridge, and then insert the new cartridge."
[1349] Step 9:
[1350] The server generates text responses, which are then sent to a speech synthesis API to be converted into audio data. This speech synthesis API also often uses a cloud-based service.
[1351] Step 10:
[1352] The server sends the generated audio data to the terminal. This transmission also takes place via the internet connection.
[1353] Step 11:
[1354] The device plays the audio data it received from the server. Using the device's speaker or a headset or other audio output device, it provides the user with an audio response. The message played is: "To replace the ink cartridge, first open the printer cover, remove the old cartridge, and then insert the new cartridge."
[1355] Through the steps described above, this system responds to user voice input with voice commands. Therefore, it is easy to use for users who have difficulty with text input or those with visual impairments.
[1356] (Example 1)
[1357] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1358] In recent years, voice-based interfaces have attracted attention, becoming a particularly important technology for users with visual impairments or those who have difficulty with text input. However, current voice response systems face several challenges. For example, accurately converting voice input data into text and then quickly analyzing it to generate appropriate responses is difficult. Furthermore, there is a need to convert the generated responses into natural and easy-to-understand speech. In addition, optimization is required to perform these processes in real time. This invention aims to solve these problems and provide quick and accurate voice responses to user questions and operational instructions.
[1359] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1360] In this invention, the server includes means for understanding the intent of questions and operational instructions by analyzing voice data into text data and the text data; means for generating appropriate answers from a database based on the analysis results; and means for converting the generated answer text into voice data. This makes it possible to quickly and accurately perform text analysis and answer generation for questions and instructions entered by the user via voice, and to provide those answers as natural and easy-to-understand voice.
[1361] "Voice input means" refers to a device that allows users to ask questions or give instructions using their voice.
[1362] "Means for sending audio data to a server" refers to the technology for sending captured audio data to a server via a network.
[1363] "Means of converting audio data to text data" refers to technology that converts captured audio data into text format.
[1364] "Methods for analyzing text data and generating responses" refers to technologies that analyze data expressed as text and generate appropriate responses to user questions or instructions.
[1365] "Means of converting generated responses into audio data" refers to technology that converts text-based responses into audio data.
[1366] "Means for transmitting audio data to a terminal" refers to the technology for transmitting generated audio data to a user's terminal via a network.
[1367] "Means of playing audio data" refers to the devices and software used to play audio data on the user's device.
[1368] "Means of converting to digital audio data in real time" refers to technology that converts user voice input into digital data in real time.
[1369] "Text data and means for analyzing that text data" refers to technologies for analyzing converted text data and understanding its intent and content.
[1370] "Means for generating appropriate answers from a database" refers to technologies that refer to a database based on analysis results and extract appropriate answers to user questions and operational instructions.
[1371] "Methods for converting response text into audio data" refers to technologies that convert generated responses into natural and easy-to-understand audio data.
[1372] This invention is a system that allows users to give operating instructions and ask questions using their voice and receive answers in voice. This system combines a voice input device, a network, a server, a natural language processing model, and speech synthesis technology.
[1373] Voice input method
[1374] The user uses a microphone to speak into the device, asking questions or giving instructions. The device is equipped with speech recognition software, such as the Google Speech-to-Text API, which converts the voice into digital audio data in real time. For example, if the user says, "The printer ink isn't coming out," this voice is captured as digital data.
[1375] Sending audio data
[1376] The device transmits the captured audio data to the server via an internet connection or local network. This transmission process uses encryption technologies such as TLS to ensure data security.
[1377] Converting audio data to text
[1378] The server receives the incoming audio data and uses the Amazon Transcribe API to convert the audio into text data. For example, the user's audio data, "The printer ink isn't coming out," is converted into the text format "The printer ink isn't coming out."
[1379] Text data analysis
[1380] The server analyzes the converted text data using natural language processing models (e.g., GPT-3, BERT) to understand the intent behind questions and instructions. For example, the server analyzes the text data "The printer ink isn't coming out" and understands its meaning. This process includes generating text tokens and interpreting intent.
[1381] Answer generation
[1382] Based on the analysis results, the server extracts and generates the appropriate answer from a database (e.g., a printer instruction manual database). For example, it might generate an answer such as, "You need to replace the ink cartridge. Open the printer cover, remove the old cartridge, and insert the new cartridge."
[1383] Voice conversion of the answer
[1384] The generated response text is converted into speech data using speech synthesis technology such as the Google Text-to-Speech API. This ensures that the user's requested response is produced as natural-sounding speech data.
[1385] Sending audio data
[1386] The server then transmits the generated audio data back to the terminal via the internet or local network. Security is ensured during this process by encrypting the data using TLS or similar protocols.
[1387] Playback of audio data
[1388] The device plays the received audio data through an audio output device such as a speaker or headphones. The user can then listen to this audio and take the necessary actions or responses. For example, the device's speaker might play an audio message saying, "You need to replace the ink cartridge. Open the printer cover, remove the old cartridge, and insert the new cartridge."
[1389] Specific example
[1390] For example, if a user asks "The printer ink isn't coming out," the following process will be performed.
[1391] The user speaks into the voice input device and says, "The printer ink isn't coming out."
[1392] The device captures this audio data and sends it to the server.
[1393] The server converts the audio data into text data using a speech recognition API.
[1394] The text data is analyzed using a natural language processing model to search for relevant sections in the instruction manual related to the problem "the printer ink is not coming out."
[1395] The server generates the appropriate response from the instruction manual, for example, "You need to replace the ink cartridge. Open the printer cover, remove the old cartridge, and insert the new cartridge."
[1396] Convert this answer into speech data using a speech synthesis API.
[1397] Audio data is sent to the terminal, which then plays it back to the user through an audio output device.
[1398] Example of a prompt
[1399] User input: "The printer ink isn't coming out."
[1400] System response: "You need to replace the ink cartridge. Open the printer cover, remove the old cartridge, and insert the new one."
[1401] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1402] Step 1: Capture voice input
[1403] The user speaks into a voice input device (e.g., a microphone) to ask a question or give an instruction. The device converts this voice into digital speech data using the Google Speech-to-Text API. The input is the user's utterance, "The printer ink isn't coming out," and the output is the text data of this utterance. Specifically, the device captures the user's voice in real time and converts it into a digital signal.
[1404] Step 2: Sending the audio data
[1405] The terminal sends the generated digital audio data to the cloud server using encryption technology such as TLS. The input is digital audio data, and the output is the transmission of encrypted data. Specifically, the terminal sends the data to the server via an internet connection or a local network.
[1406] Step 3: Converting audio data to text
[1407] The server converts received audio data into text data using the Amazon Transcribe API. The input is digital audio data, and the output is the corresponding text data. Specifically, the server receives audio data and makes an API request to convert it into text data.
[1408] Step 4: Analyzing Text Data
[1409] The server receives text data and performs analysis using a natural language processing model (e.g., BERT, GPT-3). The input is text data, and the output is semantic understanding data resulting from the analysis. Specifically, the server tokenizes the text and uses the model to analyze the intent of questions and instructions.
[1410] Step 5: Generating the answer
[1411] Based on the analysis results, the server extracts and generates appropriate answers from a database (e.g., an instruction manual database). The input is semantic comprehension data, and the output is the generated answer text. Specifically, the server searches the queried database, extracts matching answers, and assembles them.
[1412] Step 6: Voice conversion of the answer
[1413] The generated response text is converted into speech data using the Google Text-to-Speech API. The input is the response text, and the output is the corresponding speech data. Specifically, the server makes an API request to convert the text to speech.
[1414] Step 7: Sending audio data
[1415] The server re-encrypts the generated audio data using TLS or similar methods and sends it to the terminal. The input is the audio data, and the output is the transmission of the encrypted data. Specifically, the server sends the data back to the terminal via the network.
[1416] Step 8: Play back audio data
[1417] The device decodes the received audio data and plays it back through an audio output device such as a speaker or headphones. The input is audio data, and the output is the audio that the user hears. Specifically, the device decodes the encryption and outputs the audio using an audio playback device.
[1418] (Application Example 1)
[1419] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1420] Traditional food delivery systems require users to place orders via a screen, which presents challenges for users with dirty hands or those with visual impairments. Furthermore, checking order status and estimated delivery times also requires manual operation, resulting in a lack of convenience.
[1421] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1422] In this invention, the server includes a voice input means, a means for transmitting voice data to the server, a means for converting voice data into text data, a means for analyzing the text data to generate a response corresponding to the order, a means for converting the generated response into voice data, a means for transmitting voice data to a terminal, and a means for playing back the voice data. This allows users to order food and check the order status using only their voice, improving convenience for people with visual impairments and users whose hands are occupied.
[1423] A "voice input method" refers to a device or system that allows a user to give instructions by voice.
[1424] "Means for sending audio data to a server" refers to the mechanisms and protocols for sending captured audio data to a server via the internet.
[1425] "Means for converting audio data to text data" refers to technologies for converting audio data captured using speech recognition technology into text format.
[1426] "A means of analyzing text data to generate responses corresponding to orders" refers to a system that uses natural language processing technology to analyze text data and generate appropriate responses from the obtained information.
[1427] "Means for converting generated responses into audio data" refers to a mechanism or system that converts text-based responses into audio data using speech synthesis technology.
[1428] "Means for transmitting audio data to a terminal" refers to the mechanisms and protocols for transmitting generated audio data to a terminal used by the user.
[1429] "Means for playing audio data" refers to devices or systems that play transmitted audio data and allow the user to listen to it.
[1430] "Data related to food and beverage orders" refers to data including menus, prices, and restaurant information handled within the food delivery system.
[1431] A "natural language processing model" is a machine learning technique or algorithm used to analyze user text data, understand user intent, and generate appropriate responses.
[1432] "Responses regarding orders and delivery status" refers to responses that include information about the products ordered by the user and their delivery status.
[1433] The system for carrying out this invention allows users to place food delivery orders and check the order status using voice commands. The system includes the following main components:
[1434] 1. Voice input means
[1435] Users input voice commands, such as ordering food or checking delivery status, by speaking into the microphone of their device (smartphone or smart glasses). This voice input method utilizes the built-in microphone of the smartphone or smart glasses, and captures the voice as digital data using the Google Speech-to-Text API.
[1436] 2. Means for sending audio data to the server
[1437] The device sends the captured audio data to a cloud server via an internet connection. HTTPS communication is used for data transmission.
[1438] 3. Means for converting audio data into text data
[1439] The server converts the received audio data into text data using speech recognition technology. Specifically, it uses the Google Speech-to-Text API to convert the audio data into text format.
[1440] 4. A means of analyzing text data to generate responses corresponding to orders.
[1441] The server analyzes the data, which has been converted to text format, using a natural language processing model. It utilizes advanced pre-trained models such as GPT-4 to understand user intent and generate appropriate responses. In this process, it references a food delivery database, providing data such as menu information, pricing information, and estimated delivery times.
[1442] 5. Means for converting the generated response into audio data
[1443] The server converts the generated text-based responses into speech data using speech synthesis technology. This conversion utilizes the Google Text-to-Speech API.
[1444] 6. Means for transmitting audio data to a terminal
[1445] The server then sends the generated audio data back to the terminal. This is also done via HTTPS communication over the internet connection.
[1446] 7. Means for playing audio data
[1447] Finally, the device plays back the received audio data and provides the user with an answer. A smartphone, smart glasses speaker, or headphones are used as the playback device.
[1448] Adding specific examples
[1449] For example, if a user uses voice input to say "I want to order a pizza," the following process will occur.
[1450] The user says, "I want to order a pizza."
[1451] The device captures this audio and sends it to the server.
[1452] The server uses the Google Speech-to-Text API to convert the audio data into text data.
[1453] A natural language processing model (e.g., GPT-4) is used to analyze text data such as "I want to order a pizza" and provide information on appropriate restaurants and menus.
[1454] The server generates an answer such as "Which restaurant would you order pizza from?", and this answer is converted into speech data using the Google Text-to-Speech API.
[1455] The server generates audio data and sends it to the terminal, which then plays the audio.
[1456] Example of a prompt
[1457] Analyze user comments and suggest corresponding food delivery restaurants and menus: {User comments}
[1458] Follow user instructions, confirm the order, and communicate with the backend to check its status.
[1459] In this way, users can perform a series of operations, from ordering food delivery to checking the delivery status, using only their voice, without using a visual interface. This significantly improves convenience for users whose hands are full or who have visual impairments.
[1460] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1461] Step 1:
[1462] The user speaks into the device's voice input and says, "I want to order a pizza." This captures the audio data. The input is what the user said, and the output is the captured audio data.
[1463] Step 2:
[1464] The device sends the captured audio data to the cloud server via the internet connection. The input is the captured audio data, and the output is the audio data received on the server side. Specifically, the operation involves sending the audio data using HTTPS communication.
[1465] Step 3:
[1466] The server converts the received audio data into text data using the Google Speech-to-Text API. The input is the received audio data, and the output is text data. Specifically, it analyzes the audio data using speech recognition technology and converts it into the corresponding text.
[1467] Step 4:
[1468] The server analyzes the converted text data using a natural language processing model (e.g., GPT-4). The input is text data, and the output is the user's intent or question content as a result of the analysis. Specifically, it analyzes the text using a generative AI model and understands the content of the corresponding instructions or questions.
[1469] Step 5:
[1470] The server generates appropriate restaurant and menu information based on the analysis results and creates a text-based response. The input is the analysis results, and the output is the generated text-based response. Specifically, it refers to an internal restaurant database, extracts appropriate information to meet the user's request, and generates a response.
[1471] Step 6:
[1472] The server converts the generated text-based response into audio data using the Google Text-to-Speech API. The input is a text-based response, and the output is audio data. Specifically, it uses speech synthesis technology to convert text into speech.
[1473] Step 7:
[1474] The server sends the generated audio data back to the terminal. The input is the generated audio data, and the output is the audio data received by the terminal. Specifically, the operation involves sending the audio data to the terminal using HTTPS communication.
[1475] Step 8:
[1476] The device plays the received audio data and provides a response to the user. The input is the received audio data, and the output is the audio information to be heard by the user. Specifically, the device plays the audio data using its speaker or headphones.
[1477] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1478] The embodiments for carrying out this invention will be described in detail. The present invention combines an emotion recognition function with a system in which a user inputs operation instructions or questions by voice and receives answers by voice. The system comprises a voice input means, a means for transmitting voice data to a server, a means for converting voice data into text data, a means for analyzing the text data to generate answers, a means for converting the generated answers into voice data, a means for transmitting voice data to a terminal, a means for playing back the voice data, and an emotion engine that recognizes the user's emotions.
[1479] System Configuration
[1480] 1. Voice input means
[1481] The user inputs operating instructions or questions by voice into the terminal. For example, the user might say, "The printer ink isn't coming out." This voice input method uses a microphone to capture the voice data and saves it as digital audio data.
[1482] 2. Means for sending audio data to the server
[1483] The device sends audio data captured via voice input to a cloud server or local server. This transmission takes place via an internet connection.
[1484] 3. Means for converting audio data into text data
[1485] The server uses a speech recognition API to convert the received audio data into text data. The speech recognition API generates the text "The printer ink isn't coming out."
[1486] 4. Means for analyzing text data to generate responses
[1487] The server uses a natural language processing model to analyze text data and understand the intent behind user questions and instructions. It then consults a database of instruction manuals and overview booklets to generate appropriate answers to the questions.
[1488] 5. Means for converting the generated response into audio data
[1489] The server converts the generated response into audio data using a speech synthesis API. This speech synthesis API then prepares the response for the user as audio.
[1490] 6. Means for transmitting audio data to a terminal
[1491] The server sends the generated audio data to the terminal. This transmission also takes place via the internet connection.
[1492] 7. Means for playing audio data
[1493] The device plays the received audio data. It then provides the user with the answer via voice using a speaker or headphones.
[1494] 8. Emotional Engine
[1495] The emotion engine analyzes and recognizes the user's emotional state using voice data. This emotion analysis is based on features such as voice tone, speed, and intonation.
[1496] Specific example
[1497] For example, the sequence of operations when a user asks "The printer ink isn't coming out" is as follows:
[1498] The user speaks into the voice input device and says, "The printer ink isn't coming out."
[1499] The device captures this audio and sends the audio data to the server.
[1500] The server converts the audio data into text data using a speech recognition API.
[1501] The server uses a natural language processing model to analyze text data and search for information related to the problem "the printer ink isn't coming out."
[1502] The server generates appropriate responses to the user based on the instruction manual. For example, it might generate text such as, "To replace the ink cartridge, first open the printer cover, remove the old cartridge, and then insert the new cartridge."
[1503] The text response generated by the server is converted into speech data using a speech synthesis API.
[1504] The server sends the audio data to the terminal.
[1505] The device plays audio data and provides the user with an answer.
[1506] The emotion engine analyzes the user's original voice data, and if it determines that the user is confused, the server modifies the response to be more helpful and detailed. For example, it might adjust the tone and content of the response to something like, "Don't worry, first try opening the printer cover. Next, carefully remove the old ink cartridge. Then, firmly insert the new cartridge."
[1507] Through these steps, the system can provide flexible responses tailored to the user's emotional state. This allows users to use the system more comfortably and effectively.
[1508] The following describes the processing flow.
[1509] Step 1:
[1510] The user inputs operating instructions or questions by voice into the terminal. For example, the user might say, "The printer ink isn't coming out."
[1511] Step 2:
[1512] The device uses its built-in microphone to capture the user's voice and temporarily stores it as digital audio data.
[1513] Step 3:
[1514] The device sends the captured audio data to the server. This transmission takes place via an internet connection.
[1515] Step 4:
[1516] The server sends audio data to a speech recognition API, which then converts the audio data into text data. For example, a cloud-based speech recognition service could be used.
[1517] Step 5:
[1518] The server processes the text data "Printer ink is not coming out" received from the speech recognition API.
[1519] Step 6:
[1520] The server uses a natural language processing model to analyze the text data. This analysis helps the server understand what the user wants. For example, it might determine that the user wants to solve a printer problem.
[1521] Step 7:
[1522] The server searches a database of instruction manuals and summary booklets to obtain appropriate answers to user questions. For example, it might find information such as, "The ink cartridge needs to be replaced."
[1523] Step 8:
[1524] The server generates appropriate responses for the user based on the instruction manual. For example, it might generate text such as, "To replace the ink cartridge, first open the printer cover, remove the old cartridge, and then insert the new cartridge."
[1525] Step 9:
[1526] The server sends the generated text response to a speech synthesis API, which converts it into audio data. This speech synthesis API then prepares the response for the user as audio.
[1527] Step 10:
[1528] The server sends the generated audio data to the terminal. This transmission also takes place via the internet connection.
[1529] Step 11:
[1530] The device plays the audio data it received from the server. Using the device's speaker or headphones, it provides the user with an audio response. The message played is: "To replace the ink cartridge, first open the printer cover, remove the old cartridge, and then insert the new cartridge."
[1531] Step 12:
[1532] The emotion engine analyzes the original audio data and evaluates the user's emotional state. For example, it can detect feelings of confusion or frustration.
[1533] Step 13:
[1534] The server adjusts the responses generated based on the emotion engine's analysis. For example, if the server detects that the user is confused, it adjusts the response to be more helpful and detailed. It might change to something like, "Don't worry, first try opening the printer cover. Next, carefully remove the old ink cartridge. Then, firmly insert the new cartridge."
[1535] Step 14:
[1536] The server sends the newly processed audio data back to the speech synthesis API, where it is converted back into audio data.
[1537] Step 15:
[1538] The server sends new audio data to the device, and the device plays this data. The user can then hear the corrected, gentler-toned response.
[1539] (Example 2)
[1540] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1541] Conventional voice dialogue systems can provide appropriate answers to user inputs, such as operational instructions and questions, but they struggle to provide flexible responses that adapt to the user's emotional state. Furthermore, they lacked the means to recognize and appropriately respond to user feelings of confusion or anxiety. As a result, the user experience was not always optimal.
[1542] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1543] In this invention, the server includes means for converting voice data into text data, means for analyzing the text data to generate a response, and means for converting the generated response back into voice data. This enables flexible response generation that takes into account the emotional state of the user based on voice input.
[1544] "Voice input means" refers to a device that includes a microphone for the user to input voice, and a means for capturing that voice as digital data.
[1545] "Means for sending audio data to a server" refers to means for sending captured audio data to a cloud server or local server via the internet.
[1546] "Means for converting audio data to text data" refers to methods for converting audio data received using a speech recognition API into text data.
[1547] "Methods for analyzing text data and generating responses" refers to methods for analyzing text data using natural language processing models and generating appropriate responses.
[1548] "Means for converting generated responses into audio data" refers to means for converting text responses generated using a speech synthesis API into audio data.
[1549] "Means for transmitting audio data to a terminal" refers to means for transmitting generated audio data to a user's terminal via the internet.
[1550] "Means for playing audio data" refers to the means by which a terminal plays back audio data it has received through speakers or headphones.
[1551] An "emotion engine" is a means of analyzing and recognizing a user's emotional state using voice data and responding accordingly.
[1552] A "natural language processing model" is an algorithm or software that analyzes text data to understand the meaning and intent of human language.
[1553] A "speech recognition API" is an application programming interface for converting speech data into text data.
[1554] A "speech synthesis API" is an application programming interface for converting text data into speech data.
[1555] The embodiments for carrying out this invention will be described in detail. The present invention is a system including voice input means, means for transmitting voice data to a server, means for converting voice data to text data, means for analyzing text data to generate a response, means for converting the generated response to voice data, means for transmitting voice data to a terminal, means for playing back voice data, and an emotion engine for recognizing the user's emotions.
[1556] System Configuration
[1557] 1. Voice input means
[1558] The user inputs operating instructions or questions by voice into the device. For example, the user might say, "The printer ink isn't coming out." This voice input method uses a microphone to capture the voice data and saves it as digital audio data. Specific hardware that can be used include smartphones, tablets, and PCs with microphones.
[1559] 2. Means for sending audio data to the server
[1560] The device sends the audio data captured via voice input to a cloud server or local server. This transmission takes place over the internet. HTTPS is used as the specific transmission protocol.
[1561] 3. Means for converting audio data into text data
[1562] The server uses a speech recognition API (for example, Google Cloud Speech-to-Text API) to convert the received audio data into text data. The speech recognition API generates the text "The printer ink is not coming out."
[1563] 4. Means for analyzing text data to generate responses
[1564] The server uses a natural language processing model (e.g., OpenAI's GPT-3) to analyze text data and understand the intent behind user questions and instructions. It then consults a database of instruction manuals and overview booklets to generate appropriate answers to the questions.
[1565] 5. Means for converting the generated response into audio data
[1566] The server converts the generated response into audio data using a speech synthesis API (e.g., Amazon Polly). This speech synthesis API then provides the response to the user as audio.
[1567] 6. Means for transmitting audio data to a terminal
[1568] The server sends the generated audio data to the terminal. This transmission also takes place via the internet connection. HTTPS is used as the transmission protocol.
[1569] 7. Means for playing audio data
[1570] The device plays back the received audio data. It provides the user with an audio response using speakers or headphones. Specific hardware options include smartphones, tablets, and speakers or headphones connected to a PC.
[1571] 8. Emotional Engine
[1572] The emotion engine analyzes and recognizes the user's emotional state using voice data. This emotion analysis is based on features such as voice tone, speed, and intonation. Specific software such as an emotion analysis engine (e.g., IBM Watson Tone Analyzer) can be used. Based on the analysis results, the server adjusts its responses and generates more helpful responses as needed.
[1573] Specific example
[1574] The following is the sequence of events that occurs when a user asks, "The printer ink isn't coming out":
[1575] 1. The user speaks into the device and says, "The printer ink isn't coming out."
[1576] 2. The device captures this audio and sends the audio data to the server.
[1577] 3. The server converts the audio data into text data using a speech recognition API.
[1578] 4. The server uses a natural language processing model to analyze the text data and search for information related to the problem "the printer ink is not coming out."
[1579] 5. The server generates appropriate responses to the user based on the instruction manual. For example, it might generate text such as, "To replace the ink cartridge, first open the printer cover, remove the old cartridge, and then insert the new cartridge."
[1580] 6. The text response generated by the server is converted into speech data using a speech synthesis API.
[1581] 7. The server sends the audio data to the terminal.
[1582] 8. The device plays audio data and provides the user with an answer.
[1583] 9. The emotion engine analyzes the user's original voice data and, if it determines that the user is confused, the server modifies the response to be more helpful and detailed. For example, it might adjust the tone and content of the response to something like, "Don't worry, first try opening the printer cover. Next, remove the old ink cartridge. Then, carefully insert the new cartridge."
[1584] Examples of prompt statements
[1585] Examples of prompt statements are as follows:
[1586] User: What should I do if my printer ink isn't coming out?
[1587] System: To replace an ink cartridge, first open the printer cover, remove the old cartridge, and then install the new cartridge.
[1588] By inputting this prompt into the AI model, the system generates an appropriate answer to the user's question. This mechanism allows users to use the system more comfortably and effectively.
[1589] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1590] Step 1:
[1591] Voice input
[1592] Action: The user speaks into the terminal saying, "The printer ink isn't coming out."
[1593] Input: User's voice.
[1594] Output: Digital audio data.
[1595] Details: The microphone built into the device captures the audio and saves it as digital audio data. Examples of devices that can be used include smartphones, tablets, and PCs with built-in microphones.
[1596] Step 2:
[1597] Sending audio data
[1598] Operation: The device sends audio data to the server.
[1599] Input: Digital audio data.
[1600] Output: Audio data transferred to the server.
[1601] Details: The device sends the captured audio data to a cloud server or local server via the internet. This communication uses the HTTPS protocol and is secure.
[1602] Step 3:
[1603] Convert audio data to text data
[1604] Operation: The server uses a speech recognition API to convert speech data into text data.
[1605] Input: Audio data transferred to the server.
[1606] Output: Text data. The content is "The printer ink isn't coming out."
[1607] Details: The server calls the Google Cloud Speech-to-Text API to convert the received audio data into text data.
[1608] Step 4:
[1609] Analyze text data to generate answers.
[1610] Operation: The server uses a natural language processing model to analyze text data and generate appropriate responses.
[1611] Input: Converted text data.
[1612] Output: Text data of the response.
[1613] Details: The server uses OpenAI's GPT-3 model to analyze the content of text data and understand the user's question and intent. It then refers to a database of instruction manuals and overview booklets to generate appropriate answers to the questions. A specific example answer is: "To replace the ink cartridge, first open the printer cover, remove the old cartridge, and then install the new cartridge."
[1614] Step 5:
[1615] Convert the generated response into audio data.
[1616] Operation: The server converts the generated response into audio data.
[1617] Input: Text data of the response.
[1618] Output: Audio data.
[1619] Details: The server calls the Amazon Polly API to convert the generated text responses into audio data.
[1620] Step 6:
[1621] Sending audio data
[1622] Operation: The server sends the generated audio data to the terminal.
[1623] Input: Audio data.
[1624] Output: Audio data transferred to the terminal.
[1625] Details: The server sends the generated audio data to the user's terminal via the internet. This communication also uses the HTTPS protocol, ensuring security.
[1626] Step 7:
[1627] Playback of audio data
[1628] Operation: Plays audio data received by the device.
[1629] Input: Audio data transferred to the device.
[1630] Output: Played audio.
[1631] Details: The device decodes the received audio data and plays the audio through the speaker or headphones, thereby providing the user with a response.
[1632] Step 8:
[1633] Sentiment analysis and response adjustment
[1634] Operation: The server uses an emotion engine to analyze the user's emotional state and adjusts its responses as needed.
[1635] Input: Audio data and its analysis results.
[1636] Output: Adjusted response data.
[1637] Details: The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze the tone, speed, and intonation of the user's voice data. Based on the analysis, if it is determined that the user is confused, the original response is modified to be more helpful and detailed. For example, the response's tone and content might be adjusted to something like, "Don't worry, first try opening the printer cover. Next, remove the old ink cartridge. Then, carefully insert the new cartridge."
[1638] (Application Example 2)
[1639] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1640] Current voice control systems fail to recognize the user's emotional state, resulting in problems such as being unable to respond appropriately when the user is confused or in an emergency. This is especially true in security services, where users may act inappropriately when panicked, requiring a swift and accurate response. However, existing systems cannot meet these needs.
[1641] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice input means, a means for transmitting voice data to the server, and a means for converting voice data into text data. This makes it possible to convert the user's voice input into text and analyze it. Furthermore, it includes a means for analyzing the text data to generate a response, a means for converting the generated response into voice data, an emotion engine for analyzing and recognizing the emotional state, a means for adjusting the response based on the emotion analysis results, and a means for contacting emergency contacts. This enables flexible responses based on the user's emotional state and rapid emergency response as needed.
[1642] "Voice input means" refers to devices or functions that allow users to input instructions or questions into a system via voice.
[1643] "Means of sending audio data to a server" refers to the function of sending captured audio data to a cloud server or local server via the internet or other means.
[1644] "Means for converting audio data to text data" refers to devices or functions that use speech recognition technology to convert audio data into text data.
[1645] "Means for analyzing text data and generating responses" refers to a function or technical means that analyzes input text data and generates appropriate responses.
[1646] "Means for converting generated responses into audio data" refers to a function that converts text-based responses into audio data using speech synthesis technology.
[1647] "Means for transmitting audio data to a terminal" refers to communication means for transmitting generated audio data to the user's terminal.
[1648] "Means for playing audio data" refers to the function of playing audio data received by a device through a speaker or headphones.
[1649] An "emotion engine that analyzes and recognizes the user's emotional state" is a software engine that analyzes and recognizes a user's emotions based on characteristics such as tone, speed, and intonation of their voice.
[1650] "Means for adjusting responses based on emotion analysis results" refers to a function that appropriately adjusts the tone and content of responses based on emotion analysis results obtained by the emotion engine.
[1651] The "means of contacting emergency contacts" feature is a function that sends a notification to pre-set emergency contacts when the user's emotional state is determined to be an emergency situation.
[1652] This invention is a system that combines user voice input and emotion recognition, and is particularly applicable to emergency response in security services. Specific embodiments for carrying out the invention are described below.
[1653] System Overview
[1654] This system includes means for voice input, means for transmitting voice data to a server, means for converting voice data into text data, means for analyzing text data to generate a response, means for converting the generated response into voice data, means for transmitting voice data to a terminal, means for playing back voice data, an emotion engine for analyzing and recognizing the user's emotional state, and means for adjusting the response based on the emotion analysis results. Furthermore, it also includes means for contacting emergency contacts in emergencies.
[1655] Hardware and software configuration
[1656] This system consists of the following hardware and software:
[1657] Audio input method: Use an audio capture device such as a microphone.
[1658] Server: Performs necessary data processing on cloud servers or local servers.
[1659] Speech recognition APIs such as Google Cloud Speech-to-Text are used to convert speech data into text data.
[1660] Natural language processing model: A model for understanding user questions and instructions, such as using the Google Cloud Natural Language API.
[1661] Text-to-speech APIs such as Google Cloud Text-to-Speech are used to convert text data into speech data.
[1662] Emotion Engine: Software that analyzes the tone, speed, intonation, etc. of a voice to recognize the user's emotions.
[1663] Emergency contact method: A means of communication used to send notifications to designated contacts in the event of an emergency.
[1664] Detailed operation
[1665] Voice input and data transmission
[1666] First, the user speaks a question or instruction into the smartphone's microphone. For example, a prompt might say, "There might be a suspicious person in my house, what should I do?" This audio data is captured and sent to the server.
[1667] Speech recognition and text conversion
[1668] The server converts the transmitted audio data into text data using a speech recognition API. The Google Cloud Speech-to-Text API is used, and this process generates the text "There might be a suspicious person in my house, what should I do?"
[1669] Text analysis and response generation
[1670] The text data is analyzed using a natural language processing model to generate appropriate responses. For example, in response to the above prompt, the following response is generated: "Stay calm. I will contact the police immediately. Move to a safe location."
[1671] Emotion recognition and response adjustment
[1672] Furthermore, this response is adjusted by an emotion engine according to the user's emotional state. If the tone and speed of the voice indicate that the user is in a state of panic, the tone of the response will be changed to a calmer and more reassuring one.
[1673] Speech synthesis and data transmission
[1674] The generated responses are converted into audio data using a speech synthesis API and sent from the server to the device. The Google Cloud Text-to-Speech API performs this conversion.
[1675] Audio playback
[1676] The user's device plays the received audio data. Playback occurs through speakers or headphones, and the user is provided with a response.
[1677] Emergency response
[1678] If the system determines that a user's emotional state is urgent, it sends a notification to a pre-configured emergency contact. This allows the user to receive appropriate assistance immediately.
[1679] Specific example
[1680] Let's explain the detailed operation using the previously shown prompt, "There might be a suspicious person in my house, what should I do?" as an example. When the user speaks this question into the voice input device, the voice data is captured and sent to the server. The server converts the voice data into text data and analyzes it using a natural language processing model. The generated response will be, "Please stay calm. I will contact the police immediately. Please move to a safe place." The emotion engine recognizes the user's panic state and further adjusts the response to be calmer. The response, converted into voice data using a speech synthesis API, is then sent to the device and played back. Simultaneously, a notification is sent to emergency contacts, allowing the user to receive quick assistance from the difficult situation.
[1681] The above describes a specific embodiment for carrying out this invention. By using this system, users can obtain answers to their questions and instructions safely and quickly, and can also respond in emergencies.
[1682] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1683] Step 1:
[1684] The user speaks questions or instructions using a voice input device.
[1685] As a concrete example, the user speaks into their smartphone's microphone, "There might be a suspicious person in my house, what should I do?" The input data is captured as audio data.
[1686] Step 2:
[1687] The device sends the captured audio data to the server.
[1688] In this process, the audio data is sent via the internet to a cloud server or a local server. The input is the captured audio data, and the output is the audio data sent to the server.
[1689] Step 3:
[1690] The server converts the audio data into text data using a speech recognition API.
[1691] Specifically, the Google Cloud Speech-to-Text API is used to generate text data for the sentence, "There might be a suspicious person in my house, what should I do?". The input is audio data, and the output is text data.
[1692] Step 4:
[1693] The server uses a natural language processing model to analyze text data and generate appropriate responses.
[1694] This example uses the Google Cloud Natural Language API to analyze the text "There might be a suspicious person in my house, what should I do?". Based on the analysis, it generates the response "Stay calm. I will contact the police immediately. Move to a safe location." The input is text data, and the output is the response text.
[1695] Step 5:
[1696] The server uses an emotion engine to analyze the user's emotional state.
[1697] Specifically, it evaluates the tone, speed, and intonation of the voice to determine if the user is in a state of panic. The input is voice data, and the output is the result of the emotion analysis.
[1698] Step 6:
[1699] Based on the sentiment analysis results, the server adjusts the response.
[1700] For example, if the user is in a state of panic, the tone of the response is calmed and the content is changed to something reassuring. The adjusted response would be, "Please don't worry, I'll contact the police immediately. Please move to a safe location." The input is the result of the sentiment analysis and the response text, and the output is the adjusted response text.
[1701] Step 7:
[1702] The generated responses are converted into audio data using a speech synthesis API.
[1703] Specifically, the Google Cloud Text-to-Speech API is used to convert the adjusted response text into audio data. The input is the adjusted response text, and the output is the generated audio data.
[1704] Step 8:
[1705] The server sends the audio data to the terminal.
[1706] This data transmission takes place over the internet and is sent to the user's terminal. The input is the generated audio data, and the output is the audio data sent to the terminal.
[1707] Step 9:
[1708] The device plays the audio data.
[1709] Users can receive responses from the system via audio through their smartphone's speaker or headphones. The input is the audio data sent to the device, and the output is the played audio.
[1710] Step 10:
[1711] The server will contact the emergency contact.
[1712] If the server determines that the user's emotional state is urgent, it sends a notification to a pre-configured emergency contact. The input is the result of the emotional analysis, and the output is the emergency notification that was sent.
[1713] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1714] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1715] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1716] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1717] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1718] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1719] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1720] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1721] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1722] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1723] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1724] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1725] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1726] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1727] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1728] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1729] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1730] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1731] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1732] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[1733] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[1734] The following is further disclosed regarding the embodiments described above.
[1735] (Claim 1)
[1736] Voice input method,
[1737] A means of sending audio data to a server,
[1738] A means of converting audio data into text data,
[1739] A means of analyzing text data to generate answers,
[1740] A means of converting the generated response into audio data,
[1741] A means of transmitting audio data to a terminal,
[1742] A means of playing audio data,
[1743] A system that includes this.
[1744] (Claim 2)
[1745] The system according to claim 1, further comprising means for analyzing text data using data from instruction manuals and overview booklets to generate a response.
[1746] (Claim 3)
[1747] The system according to claim 1, further comprising means for analyzing text data using a natural language processing model.
[1748] "Example 1"
[1749] (Claim 1)
[1750] Voice input method,
[1751] A means of sending audio data to a server,
[1752] A means of converting audio data into text data,
[1753] A means of analyzing text data to generate answers,
[1754] A means of converting the generated response into audio data,
[1755] A means of transmitting audio data to a terminal,
[1756] A means of playing audio data,
[1757] A means of converting voice input data from a user's questions or operational instructions into digital voice data in real time,
[1758] By analyzing audio data into text data and that text data, we can understand the intent behind questions and instructions.
[1759] A means for generating an appropriate answer from a database based on the analysis results,
[1760] A means of converting the generated response text into audio data,
[1761] A system that includes this.
[1762] (Claim 2)
[1763] The system according to claim 1, further comprising means for analyzing text data using data from instruction manuals and overview booklets to generate a response.
[1764] (Claim 3)
[1765] The system according to claim 1, further comprising means for analyzing text data using a natural language processing model to understand the intent of questions or instructions.
[1766] "Application Example 1"
[1767] (Claim 1)
[1768] Voice input method,
[1769] A means of sending audio data to a server,
[1770] A means of converting audio data into text data,
[1771] A means of analyzing text data to generate responses corresponding to orders,
[1772] A means of converting the generated response into audio data,
[1773] A means of transmitting audio data to a terminal,
[1774] A means of playing audio data,
[1775] A system that includes this.
[1776] (Claim 2)
[1777] The system according to claim 1, further comprising means for analyzing text data using data relating to food and beverage orders and generating a response.
[1778] (Claim 3)
[1779] The system according to claim 1, further comprising means for analyzing text data using a natural language processing model and generating responses regarding order and delivery status.
[1780] "Example 2 of combining an emotion engine"
[1781] (Claim 1)
[1782] Voice input method,
[1783] A means of sending audio data to a server,
[1784] A means of converting audio data into text data,
[1785] A means of analyzing text data to generate answers,
[1786] A means of converting the generated response into audio data,
[1787] A means of transmitting audio data to a terminal,
[1788] A means of playing audio data,
[1789] A system that includes an emotion engine to analyze the user's emotional state.
[1790] (Claim 2)
[1791] The system according to claim 1, further comprising means for analyzing text data using data from instruction manuals and overview booklets to generate a response.
[1792] (Claim 3)
[1793] The system according to claim 1, further comprising means for analyzing text data using a natural language processing model.
[1794] (Claim 4)
[1795] The system according to claim 1, further comprising means of using a speech synthesis API when converting the generated response into speech data.
[1796] (Claim 5)
[1797] The system according to claim 1, further comprising means for using a speech recognition API when converting captured audio data into text data.
[1798] (Claim 6)
[1799] The system according to claim 1, further comprising means for adjusting the response based on the emotional state analyzed by the emotion engine.
[1800] "Application example 2 when combining with an emotional engine"
[1801] (Claim 1)
[1802] Voice input method,
[1803] A means of sending audio data to a server,
[1804] A means of converting audio data into text data,
[1805] A means of analyzing text data to generate answers,
[1806] A means of converting the generated response into audio data,
[1807] A means of transmitting audio data to a terminal,
[1808] A means of playing audio data,
[1809] An emotion engine that analyzes and recognizes the user's emotional state,
[1810] A means of adjusting responses based on emotion analysis results,
[1811] A system that includes this.
[1812] (Claim 2)
[1813] The system according to claim 1, further comprising means for analyzing text data using data from instruction manuals and overview booklets to generate a response.
[1814] (Claim 3)
[1815] The system according to claim 1, further comprising means for analyzing text data and contacting emergency contacts. [Explanation of symbols]
[1816] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. Voice input method, A means of sending audio data to a server, A means of converting audio data into text data, A means of analyzing text data to generate answers, A means of converting the generated response into audio data, A means of transmitting audio data to a terminal, A means of playing audio data, A system that includes this.
2. The system according to claim 1, further comprising means for analyzing text data using data from instruction manuals and overview booklets to generate a response.
3. The system according to claim 1, further comprising means for analyzing text data using a natural language processing model.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A