system
A system using generative AI to convert voice input into text, analyze intent, retrieve database information, and generate voice responses addresses the challenge of providing accurate and emotionally tailored answers in customer service operations.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-10-02
- Publication Date
- 2026-04-14
AI Technical Summary
Conventional customer service operations face challenges in providing quick and accurate answers to technical questions due to crew members' lack of knowledge, leading to deteriorated service quality and customer confusion regarding technical terms and concepts.
A system utilizing generative artificial intelligence to acquire user voice input, convert it into text data, analyze intent, retrieve relevant information from a database, generate a natural language response, and convert it back into voice data, incorporating speech recognition and synthesis engines.
Enables rapid and accurate provision of information, enhancing user experience by providing quick and emotionally responsive answers to technical questions.
Smart Images

Figure 2026064653000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In conventional customer service operations, it has been difficult to provide quick and accurate answers to technical questions such as mobile device specifications and Internet connection methods due to the lack of knowledge among crew members. In particular, when crew members who are not good at technical knowledge handle such questions, the quality of service provided to customers may deteriorate. In addition, customers often cannot understand technical terms and concepts, resulting in difficulties in making appropriate product selections and solving problems. To solve such problems, an automatic response system using generative artificial intelligence is required.
Means for Solving the Problems
[0005] To solve the above problems, the present invention provides the following means: a system including means for acquiring user voice input, means for converting voice input into text data, means for analyzing the intent of the text data, means for acquiring relevant information from a database based on the analysis results, means for generating a natural language response from the acquired information, means for converting the generated response into voice data, and means for providing the voice data to the user. In particular, the means for generating a natural language response based on the acquired information uses a generative artificial intelligence model, and the means for converting voice input into text data uses a speech recognition engine. This system compensates for the lack of knowledge among crew members and enables the rapid and accurate provision of information to customers.
[0006] A "user" is a person who inputs questions or requests to the robot.
[0007] "Voice input" refers to the voice data generated when a user speaks to a robot.
[0008] "Text data" refers to data resulting from the conversion of voice input into text information using speech recognition.
[0009] A "speech recognition engine" is software or hardware that analyzes speech data and converts it into text data.
[0010] Natural Language Processing (NLP) is a technology for analyzing and understanding the meaning and intent of text data.
[0011] A "database" is a data storage system used to store data such as mobile phone model information and internet connection methods.
[0012] A "generative artificial intelligence model" is an artificial intelligence algorithm that generates output in natural language based on input data.
[0013] A "speech synthesis engine" is software or hardware used to convert text data into speech data.
[0014] A "system" is a combination of devices and software that perform a series of processes including user voice input, generation of text data, and output as voice data. [Brief explanation of the drawing]
[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13]It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.
Modes for Carrying Out the Invention
[0016] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0019] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0020] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0023] [First Embodiment]
[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0036] This invention is a system for quickly and accurately answering technical questions, such as mobile phone model descriptions and internet connection methods, based on user voice input. Specific embodiments of this system are described below.
[0037] This system performs a series of processes: it acquires the user's voice input, converts it into text data, analyzes the intent of the text, retrieves relevant information from a database, generates a natural language response, converts it back into voice data, and provides it to the user.
[0038] System Configuration
[0039] Get user input
[0040] The user asks the robot a question using voice. This voice data is acquired through the microphone built into the robot.
[0041] Speech recognition
[0042] The acquired audio data is sent to a server and converted into text data by a speech recognition engine. For example, a general-purpose speech recognition engine can be used.
[0043] Natural Language Processing
[0044] The converted text data is analyzed for intent by a natural language processing (NLP) engine on the server. The NLP engine understands the user's inquiry and performs analysis to process it appropriately.
[0045] Database query
[0046] Based on the analysis results, the server queries the database to retrieve the necessary information. The database stores information such as mobile phone model and internet connection details.
[0047] Answer generation using generative artificial intelligence models
[0048] The acquired information is generated into natural language responses by the server's generative artificial intelligence model. The generative AI model generates text that provides the information in a way that is easy for the user to understand.
[0049] Speech synthesis
[0050] The generated text responses are converted into audio data by the server's speech synthesis engine. For example, a general-purpose speech synthesis engine can be used.
[0051] Response to the user
[0052] The generated audio data is provided to the user through the robot's speaker.
[0053] Specific example
[0054] User Questions
[0055] For example, a user might ask the robot, "What are the latest smartphone models?"
[0056] Server-side processing
[0057] The server receives the voice data and uses a speech recognition engine to convert it into text data. The converted text will say, "Please tell me the latest smartphone model." The NLP engine analyzes this content and recognizes that the request is for information about the latest smartphones.
[0058] Based on these analysis results, the server queries the database to retrieve information about the latest smartphones. For example, it might retrieve data indicating that "the latest smartphone is the Model X."
[0059] Based on the acquired information, a generative artificial intelligence model generates a natural language response: "The latest smartphone is the Model X." This text response is then converted into speech data by a speech synthesis engine.
[0060] Robot response
[0061] Finally, the generated audio data is delivered to the user through the robot's speaker, and the robot says, "The latest smartphone is the Model X."
[0062] In this way, the system of the present invention can provide users with quick and accurate answers to their technical questions.
[0063] The following describes the processing flow.
[0064] Step 1:
[0065] The user inputs a question to the robot using voice. For example, they might say, "Please tell me the latest smartphone model."
[0066] Step 2:
[0067] The robot acquires the user's voice input and sends the voice data to the server.
[0068] Step 3:
[0069] The server passes the received audio data to the speech recognition engine, which converts the audio into text data. For example, the text data "Please tell me the latest smartphone model" is generated.
[0070] Step 4:
[0071] The server passes the converted text data to a natural language processing (NLP) engine to analyze the text's intent. This engine recognizes that the user is seeking information about the latest smartphones.
[0072] Step 5:
[0073] The server uses the analysis results obtained from the NLP engine to query the database and retrieve relevant information. For example, the database might return information such as "The latest smartphone is the Model X."
[0074] Step 6:
[0075] The server passes the acquired information to a generative artificial intelligence model to generate a response in natural language form. The generative AI generates the text response "The latest smartphone is the Model X."
[0076] Step 7:
[0077] The server passes the generated text response to the speech synthesis engine, which converts it into speech data. For example, it might generate speech data saying, "The latest smartphone is the Model X."
[0078] Step 8:
[0079] The server sends the generated audio data to the robot. The robot plays the audio data to the user through its speaker. The user hears the voice response, "The latest smartphone is the Model X."
[0080] This will enable a system where users can obtain immediate and accurate answers to their questions.
[0081] (Example 1)
[0082] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0083] Conventional systems for answering technical questions had limited ability to provide quick and accurate answers based on user voice input. Furthermore, they struggled to properly analyze voice input, understand its intent, and provide relevant information, and lacked sufficient means to efficiently search for desired information from large databases. As a result, it was difficult for users to quickly and accurately access the information they needed.
[0084] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0085] In this invention, the server includes means for acquiring user voice input, means for converting the voice input into text data, means for analyzing the intent of the text data, means for generating a natural language response from the acquired information, means for converting the generated response into voice data, and means for providing the voice data to the user. This makes it possible to provide quick and accurate answers to the user's technical questions.
[0086] "User voice input" refers to questions or instructions that a user makes to the system using their voice.
[0087] "Means for acquiring voice input" refers to devices or technologies for capturing the user's voice, such as a microphone.
[0088] "Means of converting voice input into text data" refers to technologies for converting acquired speech into text information, such as speech recognition engines.
[0089] "Means for analyzing the intent of text data" refers to technologies for understanding the content of converted text data and analyzing the user's intent, such as natural language processing engines.
[0090] "Means of obtaining relevant information from a database based on analysis results" refers to technologies and systems that execute queries on a database according to the analyzed intent and obtain the necessary information.
[0091] "Means for generating natural language responses from acquired information" refers to technologies that create human-readable text based on information obtained from a database, such as generative artificial intelligence models.
[0092] "Means of converting generated responses into audio data" refers to technologies that convert natural language text into speech, such as speech synthesis engines.
[0093] "Means of providing audio data to the user" refers to devices and technologies for transmitting generated audio to the user, such as speakers.
[0094] This invention is a system for quickly and accurately answering technical questions based on user voice input. The system of this invention performs a series of processes, including acquiring user voice input, converting it into text data, analyzing the intent of the text, retrieving relevant information from a database, generating a natural language response, and converting it back into voice data to provide to the user. Specific hardware and software are used for each processing step.
[0095] First, the user asks a question to the robot using voice. This voice data is acquired through a microphone built into the robot (for example, an omnidirectional microphone). The acquired voice data is then transmitted to a server via the internet. This process utilizes a built-in microphone and a network module.
[0096] Next, the server converts the received audio data into text data using a speech recognition engine (for example, a general-purpose speech recognition engine). The speech recognition engine analyzes the waveform data of the audio and converts it into corresponding character information. This process converts the audio into text data.
[0097] The converted text data is analyzed for intent by a natural language processing (NLP) engine on the server (for example, a general-purpose NLP engine). The NLP engine analyzes keywords and context within the text to understand the user's inquiry. This analysis enables queries to be made to the appropriate database.
[0098] The server then queries the database based on the analysis results to retrieve the necessary information. This database may contain information such as mobile phone model details or internet connection information. The server creates an SQL query to retrieve the relevant information from the database.
[0099] Based on the acquired information, the server's generative artificial intelligence model (for example, a general generative AI model) generates a natural language response. The generative AI model creates natural-sounding sentences and provides the response in a way that is easy for the user to understand.
[0100] Next, the generated text responses are converted into audio data by the server's speech synthesis engine (for example, a general-purpose speech synthesis engine). The speech synthesis engine converts the text data into audio waveform data and generates a natural-sounding human voice.
[0101] Ultimately, the audio data is delivered to the user through the robot's speaker (for example, a standard speaker). The generated audio data is played back through the speaker and provided in a format that the user can hear.
[0102] As a concrete example, consider a scenario where a user asks a robot, "What is the latest smartphone model?" The user's voice data is acquired through a microphone and sent to a server. The server uses a speech recognition engine to convert the voice data into text data, and an NLP engine analyzes the intent of the text data. Based on the analysis results, the server queries a database to obtain the latest smartphone information. A generative artificial intelligence model generates a natural language response, "The latest smartphone is the Model X," which is then converted into voice data by a speech synthesis engine. Finally, the voice message "The latest smartphone is the Model X" is delivered to the user from the robot's speaker. In this way, the system of the present invention can provide quick and accurate answers to users' technical questions.
[0103] To realize the invention, it is crucial to properly configure and coordinate these hardware and software components.
[0104] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0105] Step 1:
[0106] The user asks the robot a question using voice. This question becomes the system's input. For example, the user might say, "Please tell me the latest smartphone model." The user's voice input is captured by the device's microphone.
[0107] Step 2:
[0108] The terminal (robot) sends the acquired voice data to the server. Specifically, it sends the acquired voice data to the server via the internet. At this time, the input voice data is transmitted directly to the server, and the server receives the voice data.
[0109] Step 3:
[0110] The server converts the received audio data into text data using a general speech recognition engine. Specifically, it analyzes the audio data and outputs text data as character information. For example, the audio data "Please tell me the latest smartphone model" is converted into the corresponding string.
[0111] Step 4:
[0112] The server uses a natural language processing (NLP) engine to analyze the intent of the converted text data. This process analyzes the context and keywords of the text and outputs the intent of the user's question. For example, from the text data "Please tell me the latest smartphone models," the server recognizes that the user is seeking information about the latest smartphones.
[0113] Step 5:
[0114] The server queries the database based on the analysis results to retrieve the necessary information. In this step, the analysis results are used as input to generate an SQL query, which is then sent to the database to output information about the latest smartphones. For example, the server retrieves the information that "the latest smartphone is the Model X" from the database.
[0115] Step 6:
[0116] The server uses a generative artificial intelligence model (for example, a general generative AI model) to generate natural language responses based on the acquired information. It uses the acquired information as input to generate text in a user-friendly format. For example, it might output the text response, "The latest smartphone is the Model X."
[0117] Step 7:
[0118] The server converts the generated text responses into speech data using a common speech synthesis engine. It uses the input text data to output speech waveform data. For example, the text "The latest smartphone is the Model X" is converted into natural-sounding speech.
[0119] Step 8:
[0120] The terminal (robot) provides the user with audio data sent from the server via its speaker. By receiving the audio data output from the server and playing it back through the speaker, it provides the user with an audio response such as, "The latest smartphone is the Model X."
[0121] (Application Example 1)
[0122] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0123] In traditional virtual stores, users lacked sufficient means to obtain appropriate and timely answers when asking questions about specific products or popular items. Furthermore, answers were often provided only in text format, limiting the user experience. Additionally, traditional systems struggled to efficiently analyze user voice input, accurately understand their intent, and generate responses.
[0124] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0125] In this invention, the server includes means for acquiring user voice input, means for converting the voice input into text data, means for analyzing the intent of the text data, means for acquiring relevant information from a database based on the analysis results, means for generating a natural language response from the acquired information, means for converting the generated response into voice data, means for providing the voice data to the user, and means for responding to questions regarding product information in the virtual store. This makes it possible to provide quick and accurate voice responses within the virtual store to questions asked by the user by voice.
[0126] "Means for acquiring user voice input" refers to a device or method used to acquire the voice spoken by the user.
[0127] "Means for converting voice input into text data" refers to a device or method for converting acquired voice into text information.
[0128] "Means for analyzing the intent of text data" refers to a device or method for understanding the user's intent or the content of a question from text.
[0129] "Means for obtaining relevant information from a database based on analysis results" refers to a device or method for searching and obtaining necessary data from a database according to the analyzed information.
[0130] "Means for generating natural language responses from acquired information" refers to a device or method for generating responses in natural language that humans can understand, based on acquired data.
[0131] "Means for converting generated responses into audio data" refers to a device or method for converting text-based responses into audio format.
[0132] "Means for providing audio data to a user" refers to a device or method for delivering generated audio data to a user.
[0133] "Means for responding to product information questions in a virtual store" refers to a device or method for users to ask questions about products within a virtual environment and for providing answers to those questions.
[0134] The system of this invention is designed to allow users in a virtual store to ask questions about product information using voice, and to provide quick and accurate answers to those questions. Specific embodiments of this system are described below.
[0135] 1. System Overview
[0136] Users ask voice questions within a virtual store via smartphone, smart glasses, or head-mounted display. The system receives the voice input, processes it on a server, and then provides a voice response. The main components are as follows:
[0137] User devices: Smartphones, smart glasses, head-mounted displays, etc.
[0138] server:
[0139] Speech recognition engine: Uses Google® Cloud Speech-to-Text API.
[0140] Natural Language Processing (NLP) engine: Uses Amazon Comprehend from AWS®.
[0141] Database: AWS RDS (Relational Database Service) is used.
[0142] Generative AI model: OpenAI® GPT-3® is used.
[0143] Text-to-speech engine: Uses Google Cloud Text-to-Speech.
[0144] 2. Processing Flow
[0145] 1. Acquisition of voice input:
[0146] The user terminal acquires voice input and sends it to the server. The terminal is equipped with a built-in microphone.
[0147] 2. Speech recognition:
[0148] The server converts the received audio data into text data using the Google Cloud Speech-to-Text API.
[0149] 3. Natural Language Processing:
[0150] The converted text data is analyzed using Amazon Comprehend on AWS to understand the user's intent.
[0151] 4. Database query:
[0152] Based on the analysis results, queries are sent to AWS RDS to retrieve relevant information.
[0153] 5. Generating answers using generative AI models:
[0154] Based on the acquired information, OpenAI GPT-3 is used to generate a natural language response.
[0155] 6. Speech synthesis:
[0156] The generated text responses are converted into audio data using Google Cloud Text-to-Speech.
[0157] 7. Responding to the user:
[0158] The generated audio data is sent to the user's device, and the response is played back through the built-in speaker.
[0159] 3. Specific examples
[0160] When a user asks a question via voice, such as "What is the most popular item in this store?", the system's server performs the following actions.
[0161] The speech recognition engine converts the speech into text data that reads, "What is the most popular item in this store?"
[0162] A natural language processing engine analyzes this text and understands that the user is asking about popular products.
[0163] The database retrieves the information that "The most popular product is 'Product A'."
[0164] The generative AI model generates a natural language response that says, "The most popular product is 'Product A'."
[0165] The speech synthesis engine converts the generated text response into audio data.
[0166] Finally, the user's device plays an audio message saying, "The most popular item is 'Product A'."
[0167] Example of a prompt
[0168] When a user asks a question, it will look like this:
[0169] User voice input: What is the most popular item in this store?
[0170] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0171] Step 1:
[0172] The user terminal acquires voice input. For example, if the user says, "What is the most popular item in this store?", the terminal's microphone picks up the voice and sends it to the server as audio data. The input is the user's voice data, and the output is the audio data sent to the server.
[0173] Step 2:
[0174] The server uses a speech recognition engine to convert the audio data into text data. The Google Cloud Speech-to-Text API is used to transcribe the audio data. The input is the user's voice data, and the output is the text data "What is the most popular item in this store?".
[0175] Step 3:
[0176] The server uses a natural language processing engine to analyze the intent of text data. Using Amazon Comprehend from AWS, it understands that the text data is a question about popular products. The input is text data, and the output is intent data indicating "I am requesting information about popular products."
[0177] Step 4:
[0178] The server sends a query to the database based on the data analysis results and retrieves relevant information. It executes a query to AWS RDS to "get popular products" and retrieves the data "The popular product is 'Product A'." The input is intent data, and the output is informational data "The popular product is 'Product A'."
[0179] Step 5:
[0180] The server uses a generation AI model to generate natural language responses based on the acquired information. Using OpenAI GPT-3, it creates a natural-sounding sentence such as "The most popular product is 'Product A'." The input is informational data, and the output is the natural language response "The most popular product is 'Product A'."
[0181] Step 6:
[0182] The server converts natural language responses generated using a speech synthesis engine into audio data. The Google Cloud Text-to-Speech API is used to convert text data into audio data. The input is text data, and the output is audio data.
[0183] Step 7:
[0184] The server sends the generated audio data to the user's terminal, which then outputs it through the terminal's speaker. The user's terminal's built-in speaker plays the audio saying, "The most popular product is 'Product A'." The input is audio data, and the output is the audio response provided to the user.
[0185] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0186] The present invention is a system for quickly and accurately answering technical questions such as mobile phone model descriptions and internet connection methods based on the user's voice input, and furthermore, it is a system that provides more human-like interaction by recognizing the user's emotions and adjusting the response accordingly. Specific embodiments of this system are described below.
[0187] This system performs a series of processes: it acquires the user's voice input, converts it into text data, analyzes the intent of the text, retrieves relevant information from a database, generates a natural language response, converts it back into voice data, and provides it to the user. It also recognizes the user's emotions from their voice and adjusts the tone and content of the response accordingly.
[0188] System Configuration
[0189] Get user input
[0190] The user asks the robot a question using voice. This voice data is acquired through the microphone built into the robot.
[0191] Speech recognition
[0192] The acquired audio data is sent to a server and converted into text data by a speech recognition engine. For example, a general-purpose speech recognition engine can be used.
[0193] Natural Language Processing
[0194] The converted text data is analyzed for intent by a natural language processing (NLP) engine on the server. The NLP engine understands the user's inquiry and performs analysis to process it appropriately.
[0195] emotion recognition
[0196] Simultaneously, the acquired audio data is analyzed by an emotion engine. The emotion engine recognizes emotions based on the tone and pitch of the user's voice. For example, it determines emotions such as anger, sadness, and happiness.
[0197] Database query
[0198] Based on the analysis results, the server queries the database to retrieve the necessary information. The database stores information such as mobile phone model and internet connection details.
[0199] Answer generation using generative artificial intelligence models
[0200] The acquired information is used by the server's generative artificial intelligence model to generate a natural language response. The generative AI model considers emotional information from the emotion engine and adjusts the tone and content of the response.
[0201] Speech synthesis
[0202] The generated text responses are converted into audio data by the server's speech synthesis engine. For example, a general-purpose speech synthesis engine can be used.
[0203] Response to the user
[0204] The generated voice data is delivered to the user through the robot's speaker. The robot responds in a tone that matches the user's emotions.
[0205] Specific example
[0206] User Questions
[0207] The user asks the robot, "Please tell me the latest smartphone model."
[0208] Server-side processing
[0209] The server receives the voice data and uses a speech recognition engine to convert it into text data. The converted text will say, "Please tell me the latest smartphone model." The NLP engine analyzes this content and recognizes that the request is for information about the latest smartphones.
[0210] At the same time, the emotion engine determines the user's emotions from their voice. For example, it might determine that the user is excited.
[0211] Based on the analysis results, the server queries the database and retrieves the information that "the latest smartphone is the Model X."
[0212] Based on the acquired information, a generative AI model generates a natural language response such as "The latest smartphone is the Model X." Taking into account the information from the emotion engine, the generative AI model generates a response in an appropriate tone to the user's emotions, such as "Sorry to keep you waiting! The latest smartphone is the Model X."
[0213] Robot response
[0214] Finally, the generated audio data is delivered to the user through the robot's speaker, and the robot says, "Sorry to keep you waiting! The latest smartphone is the Model X."
[0215] In this way, the system of the present invention can improve the user experience by providing quick and accurate answers to the user's technical questions, as well as by providing responses that take the user's feelings into consideration.
[0216] The following describes the processing flow.
[0217] Step 1:
[0218] The user inputs a question to the robot using voice. For example, they might say, "Please tell me the latest smartphone model."
[0219] Step 2:
[0220] The robot acquires the user's voice input and sends the voice data to the server.
[0221] Step 3:
[0222] The server passes the received audio data to the speech recognition engine, which converts the audio into text data. For example, the text data "Please tell me the latest smartphone model" is generated.
[0223] Step 4:
[0224] The server passes the converted text data to a natural language processing (NLP) engine to analyze the text's intent. This engine recognizes that the user is seeking information about the latest smartphones.
[0225] Step 5:
[0226] The server simultaneously passes the voice data to the emotion engine to analyze the user's emotions. The emotion engine analyzes the tone, pitch, and speed of the voice to determine the user's emotional state, such as being excited, angry, or sad.
[0227] Step 6:
[0228] Based on the analysis results of the NLP engine, the server queries the database to retrieve relevant information. For example, the database might return information such as "The latest smartphone is the Model X."
[0229] Step 7:
[0230] The server passes the acquired information to a generative artificial intelligence model to generate a response in natural language form. This model also takes emotional information from the emotion engine into consideration, and generates responses such as, "Sorry to keep you waiting! The latest smartphone is the Model X."
[0231] Step 8:
[0232] The server passes the generated text response to the speech synthesis engine, which converts it into speech data. For example, it might generate speech data such as, "Sorry to keep you waiting! The latest smartphone is the Model X."
[0233] Step 9:
[0234] The server sends the generated audio data to the robot. The robot plays the audio data to the user through its speaker. The user hears the voice response, "Thank you for waiting! The latest smartphone is the Model X."
[0235] This will create a system where users can not only get quick and accurate answers to their questions, but also receive responses that are tailored to their emotions.
[0236] (Example 2)
[0237] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0238] Conventional technical question answering systems have a problem of degrading the user experience because they cannot respond in a way that takes into account the user's emotional state. Furthermore, there are limitations in processing speed and accuracy for reliably acquiring information from voice input and providing appropriate answers. As a result, while users desire accurate and rapid answers, as well as emotionally responsive interaction, it has been difficult to effectively provide these.
[0239] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0240] In this invention, the server includes means for acquiring user voice input, means for converting the voice input into text data, means for analyzing the intent of the text data, means for recognizing emotions based on the acquired voice data, means for acquiring relevant information from an information organizing device based on the analysis results, means for converting the generated response into a natural language response, means for adjusting the tone and content of the response based on the emotion recognition results, means for converting the generated response into voice data, and means for providing the voice data to the user. This enables the user to obtain quick and accurate technical answers and to experience more human-like interactions through emotionally responsive interactions.
[0241] "Means for obtaining user voice input" refers to a device or part of a device used to collect voice data, and includes, for example, a microphone.
[0242] "Means for converting voice input into text data" refers to a process or device that converts voice data into a string of characters, and a speech recognition engine is used as an example.
[0243] "Means for analyzing the intent of text data" refers to processes or devices for understanding the content of text data and identifying its purpose or requirements, and includes, for example, natural language processing engines.
[0244] "Means of recognizing emotions based on acquired audio data" refers to a process or device that analyzes the tone and pitch of speech to determine the speaker's emotional state, and includes, for example, an emotion recognition engine.
[0245] "Means for obtaining relevant information from an information organization device based on analysis results" refers to a process or device for searching for and obtaining necessary information based on analysis results, and includes database queries as an example.
[0246] "Means for converting generated responses into natural language responses" refers to a process or device that converts acquired information into a natural language format that humans can understand, and generative AI models are used as an example.
[0247] "Means for adjusting the tone and content of responses based on the results of emotion recognition" refers to a process or device that appropriately adjusts the generated response content and expression method while taking into account the user's emotional state.
[0248] "Means for converting generated responses into audio data" refers to a process or device that converts text responses back into audio format, and includes, for example, a speech synthesis engine.
[0249] "Means of providing audio data to the user" refers to a process or device for making the generated audio data listen to the user, and a speaker is used as an example.
[0250] Modes for carrying out the invention
[0251] The present invention is a system that answers technical questions quickly and accurately based on the user's voice input, and further provides a more human-like interaction by recognizing the user's emotions and adjusting the response accordingly. Specific embodiments of this system are described below.
[0252] First, the user asks the robot a question using their voice. For example, they might say, "What are the latest smartphone models?" This voice data is captured through the microphone built into the robot.
[0253] The acquired voice data is sent from the robot to the server. The server uses a speech recognition engine (for example, Google Speech-to-Text API) to convert the voice data into text data. As a result, the voice data is converted into text data that says, "Please tell me the latest smartphone model."
[0254] Next, this text data is passed to a natural language processing (NLP) engine on the server (for example, Google Cloud Natural Language API) where the intent is analyzed. In this case, the NLP engine analyzes that the user is looking for information about the latest smartphones.
[0255] Simultaneously, the acquired audio data is analyzed by an emotion recognition engine on the server (for example, IBM Watson® Tone Analyzer). Based on the tone and pitch of the voice, the emotion recognition engine determines whether the user is excited, angry, happy, or otherwise experiencing other emotions.
[0256] The server retrieves relevant information from the database based on the analysis results of the NLP engine. The database contains technical information and connection methods for a wide variety of devices. In this case, let's assume we retrieve the information that "the latest smartphone is the Model X."
[0257] Based on the acquired information, the server uses a generative artificial intelligence model (e.g., OpenAI GPT-3) to generate a natural language response. This model takes into account the emotional information obtained from the emotion recognition engine and adjusts the tone and content of the response. As a result, it generates a response such as, "Sorry to keep you waiting! The latest smartphone is the Model X."
[0258] Next, the generated text response is passed to the server's speech synthesis engine (e.g., Amazon Polly) and converted into audio data. This audio data is then delivered to the user through the robot's speaker. The robot then says, "Sorry to keep you waiting! The latest smartphone is the Model X," in a tone appropriate to the user's mood.
[0259] In this way, the system of this invention can provide a better user experience by offering quick and accurate answers to users' technical questions, as well as responses that take into account the user's feelings.
[0260] Examples of prompt statements
[0261] "Tell me about the latest iPhone (registered trademark) models."
[0262] "What should I do if I can't connect to Wi-Fi?"
[0263] "How do you perceive your current emotions?"
[0264] As a result, when providing information to users, it becomes possible to accurately consider the user's emotional state and achieve natural dialogue.
[0265] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0266] Step 1: The user asks the robot a question using voice.
[0267] The user asks, "Please tell me the latest smartphone model." The robot's built-in microphone then captures the voice data.
[0268] Input: User's voice
[0269] Output: Acquired audio data
[0270] Step 2: The device sends the acquired audio data to the server.
[0271] The robot transmits the acquired audio data to the server using wireless communication (e.g., Wi-Fi).
[0272] Input: Acquired audio data
[0273] Output: Audio data sent to the server
[0274] Step 3: The server uses a speech recognition engine to convert the speech data into text data.
[0275] The server calls the speech recognition engine (e.g., Google Speech-to-Text API), inputs the voice data of "Please tell me the latest smartphone models", and outputs the text data.
[0276] Input: Voice data sent to the server
[0277] Output: Converted text data
[0278] Step 4: The server analyzes the intention of the text data using the natural language processing (NLP) engine.
[0279] The converted text data is input into the NLP engine (e.g., Google Cloud Natural Language API) to analyze that the user is seeking information about the latest smartphones. The output is the analysis result.
[0280] Input: Converted text data
[0281] Output: Analysis result
[0282] Step 5: The server recognizes the emotion based on the acquired voice data.
[0283] The server inputs the voice data into the emotion recognition engine (e.g., IBM Watson Tone Analyzer) and outputs the result of emotion recognition. For example, the result that the user is excited can be obtained.
[0284] Input: Voice data sent to the server
[0285] Output: Emotion recognition result
[0286] Step 6: The server obtains relevant information from the information sorting device based on the analysis result.
[0287] The server executes a query against the database, asking "What is the latest smartphone?" and retrieving relevant information. For example, it might retrieve information such as "The latest smartphone is the Model X."
[0288] Input: Analysis results
[0289] Output: Related information obtained
[0290] Step 7: The server generates a natural language response based on the information it has obtained.
[0291] The server inputs the prompt "What is the latest smartphone?" into a generated AI model (e.g., OpenAI GPT-3) and generates the answer "The latest smartphone is the Model X." Taking sentiment into account, it generates the answer in an appropriate tone, such as "Sorry to keep you waiting! The latest smartphone is the Model X."
[0292] Input: Acquired related information, emotion recognition results
[0293] Output: Generated text answer
[0294] Step 8: The server converts the generated text response into speech data using a speech synthesis engine.
[0295] The generated text responses are passed to a speech synthesis engine (e.g., Amazon Polly) and converted into audio data.
[0296] Input: Generated text answer
[0297] Output: Audio data
[0298] Step 9: The device provides the generated audio data to the user.
[0299] The robot's speaker uses the converted voice data to say, "Sorry to keep you waiting! The latest smartphone is the Model X."
[0300] Input: Audio data
[0301] Output: Voice response by robot
[0302] (Application Example 2)
[0303] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0304] In modern brick-and-mortar stores, there is a demand for quick and accurate responses to customer questions and inquiries. However, conventional methods often fail to provide flexible responses that respond to customer emotions, which can lead to decreased customer satisfaction. This invention aims to solve the problem of improving customer satisfaction by not only answering customers' technical questions but also recognizing their emotions and adjusting the content and tone of responses accordingly, thereby providing a more humane interaction.
[0305] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring user voice input, means for converting voice input into text data, means for analyzing the intent of the text data, means for acquiring relevant information from a database based on the analysis results, means for generating a natural language response from the acquired information, means for converting the generated response into voice data, means for providing the voice data to the user, and means for recognizing emotions from the user's voice and adjusting the response. This makes it possible to answer customer questions quickly and accurately in a physical store while providing appropriate responses that correspond to the customer's emotions.
[0306] "Means for acquiring voice input" refers to devices or interfaces that allow users to input questions or instructions to a system using their voice.
[0307] The "means for converting voice input into text data" refers to software or algorithms used to convert the acquired voice data into character information.
[0308] The "means for analyzing the intention of text data" refers to natural language processing technology for understanding the content of a user's request or question from character information.
[0309] The "means for retrieving relevant information from a database based on the analysis result" refers to a mechanism for searching and retrieving relevant answers and data from a database based on the analyzed intention.
[0310] The "means for generating a natural language answer from the acquired information" refers to technologies such as generative artificial intelligence models for generating an answer in a form understandable by humans based on the acquired data.
[0311] The "means for converting the generated answer into voice data" refers to voice synthesis technology for converting a text-form answer into voice format.
[0312] The "means for providing voice data to the user" refers to speakers or other output devices for playing the generated voice data to the user.
[0313] The "means for recognizing the emotion from the user's voice and adjusting the response" refers to algorithms or engines for detecting emotion from the user's voice tone, pitch, etc., and adjusting the response in an appropriate tone and content.
[0314] As a form for implementing the present invention, a system of a customer service robot in a physical store is exemplified. Specifically, it is a system in which a customer asks a question to the robot by voice, and the robot quickly and accurately answers according to the content of the question, and also recognizes the emotion of the customer and adjusts the content and tone of the response.
[0315] Configuration of the System
[0316] This system includes the following components.
[0317] Get user input
[0318] The user asks the robot a question using voice. For example, "Where can I find this product?" This voice data is captured through a microphone built into the robot. Common microphones can be used as hardware (e.g., Behringer ECM8000, Shure MV88, etc.).
[0319] Speech recognition
[0320] The acquired audio data is sent to a cloud service (e.g., Amazon Transcribe, Google Speech-to-Text) and converted into text data by a speech recognition engine. In this step, the audio data is converted into the text "Where can I find this product?".
[0321] Natural Language Processing
[0322] The converted text data is analyzed by a natural language processing (NLP) engine on the server (e.g., Google Cloud Natural Language API, Microsoft® Azure® Text Analytics API, etc.). At this stage, the intent of the text is understood to be "checking product location."
[0323] emotion recognition
[0324] Simultaneously, the acquired audio data is analyzed by an emotion engine (e.g., IBM Watson Tone Analyzer, Amazon Comprehend). Based on the voice tone and pitch, the emotion engine determines, for example, that the customer is "irritated."
[0325] Database query
[0326] Based on the analysis results, the server queries a database (e.g., MySQL®, Firebase, etc.) to retrieve the necessary information. For example, it might retrieve information such as, "Product A is located on the central shelf of the third street."
[0327] Answer generation using generative artificial intelligence models
[0328] Based on the acquired information, a generative artificial intelligence model (e.g., OpenAI GPT-4®, Hugging Face Transformers, etc.) generates a natural language response. Taking into account the results of emotion recognition, a response such as "Excuse me, I apologize for the inconvenience. Item A is on the central shelf of the third lane" is generated.
[0329] Speech synthesis
[0330] The generated text responses are converted into speech data by a speech synthesis engine (e.g., Google Text-to-Speech, Amazon Polly, etc.).
[0331] Response to the user
[0332] The generated voice data is delivered to the user through the robot's speaker (e.g., JBL Charge 4, Sonos Move, etc.). The robot says, "Excuse me, I'm sorry to trouble you. Item A is on the central shelf on the third street."
[0333] Specific example
[0334] User Questions
[0335] The user asks the robot, "Where can I find this product?"
[0336] Server-side processing
[0337] The server receives the voice data and uses a speech recognition engine to convert it into text data. The converted text will say, "Where is this product?" The NLP engine analyzes this and recognizes that the user is requesting the location information of a product. At the same time, the emotion engine determines the emotion from the user's voice. For example, it may determine that the user is irritated. Based on the analysis results, the server queries the database and obtains the information that "Product A is on the central shelf of the third street." Based on the obtained information, a generative artificial intelligence model generates a natural language response: "Excuse me, we apologize for the inconvenience. Product A is on the central shelf of the third street."
[0338] Robot response
[0339] Finally, the generated audio data is delivered to the user through the robot's speaker, and the robot says, "Excuse me, I'm sorry to trouble you. Item A is on the central shelf on the third street."
[0340] Example of a prompt
[0341] User question: "Where can I find this product?"
[0342] NLP analysis results: The product's location information is being sought.
[0343] Sentiment analysis result: The customer is irritated.
[0344] Database query result: "Product A is located on the central shelf of the third aisle."
[0345] Output from the generated AI model: "Excuse me, we apologize for the inconvenience. Product A is located on the central shelf of the third aisle."
[0346] In this way, the system of the present invention can improve the customer experience by providing quick and accurate answers to customers' technical questions in physical stores and by responding in a way that takes customer emotions into consideration.
[0347] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0348] Step 1:
[0349] The user asks the robot questions using voice. For example, a voice input such as "Where can I find this product?" is acquired through the robot's microphone (e.g., Behringer ECM8000, Shure MV88, etc.). The input is the user's voice data, and the output is that voice data itself.
[0350] Step 2:
[0351] The audio data is sent to a server and converted into text data by a speech recognition engine (e.g., Amazon Transcribe, Google Speech-to-Text). In this step, the audio data is converted into the text "Where can I find this product?". The input is the audio data, and the output is the text data that is a conversion of that audio data.
[0352] Step 3:
[0353] The server analyzes the converted text data using a natural language processing (NLP) engine (e.g., Google Cloud Natural Language API, Microsoft Azure Text Analytics API, etc.) to understand the user's intent. In this step, the intent "check product location" is analyzed. The input is text data, and the output is the analyzed intent information.
[0354] Step 4:
[0355] Simultaneously, the voice data is analyzed by an emotion engine (e.g., IBM Watson Tone Analyzer, Amazon Comprehend) to recognize the customer's emotions. For example, it might determine that the customer is "irritated." The input is voice data, and the output is information that expresses emotion.
[0356] Step 5:
[0357] Based on the analysis results of the NLP engine, the server queries a database (e.g., MySQL, Firebase, etc.) to retrieve relevant information. In this step, the information "Product A is on the central shelf of the third street" is retrieved. The input is the analyzed intent information, and the output is the relevant information retrieved from the database.
[0358] Step 6:
[0359] The server uses a generative artificial intelligence model (e.g., OpenAI GPT-4, Hugging Face Transformers, etc.) to generate a natural language response based on the acquired information. Sentiment recognition results are also considered during this process. For example, a response like, "Excuse me, I apologize for the inconvenience. Product A is on the central shelf of the third street," might be generated. The input consists of relevant and sentiment information, and the output is the generated natural language response.
[0360] Step 7:
[0361] The generated text responses are converted into audio data by a text-to-speech engine (e.g., Google Text-to-Speech, Amazon Polly, etc.). In this step, the text-formatted responses are converted into audio format. The input is the text-formatted responses, and the output is audio data.
[0362] Step 8:
[0363] Finally, the generated audio data is delivered to the user through the robot's speaker (e.g., JBL Charge 4, Sonos Move, etc.). The robot says, "Excuse me, I'm sorry to trouble you. Item A is on the central shelf on the third street." The input is audio data, and the output is the audio the user hears.
[0364] In this way, the system of the present invention can provide quick and accurate answers to users' technical questions, as well as respond in a way that takes the user's feelings into consideration. This series of processes is expected to improve the customer experience in physical stores.
[0365] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0366] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0367] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0368] [Second Embodiment]
[0369] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0370] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0371] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0372] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0373] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0374] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0375] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0376] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0377] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0378] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0379] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0380] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0381] This invention is a system for quickly and accurately answering technical questions, such as mobile phone model descriptions and internet connection methods, based on user voice input. Specific embodiments of this system are described below.
[0382] This system performs a series of processes: it acquires the user's voice input, converts it into text data, analyzes the intent of the text, retrieves relevant information from a database, generates a natural language response, converts it back into voice data, and provides it to the user.
[0383] System Configuration
[0384] Get user input
[0385] The user asks the robot a question using voice. This voice data is acquired through the microphone built into the robot.
[0386] Speech recognition
[0387] The acquired audio data is sent to a server and converted into text data by a speech recognition engine. For example, a general-purpose speech recognition engine can be used.
[0388] Natural Language Processing
[0389] The converted text data is analyzed for intent by a natural language processing (NLP) engine on the server. The NLP engine understands the user's inquiry and performs analysis to process it appropriately.
[0390] Database query
[0391] Based on the analysis results, the server queries the database to retrieve the necessary information. The database stores information such as mobile phone model and internet connection details.
[0392] Answer generation using generative artificial intelligence models
[0393] The acquired information is generated into natural language responses by the server's generative artificial intelligence model. The generative AI model generates text that provides the information in a way that is easy for the user to understand.
[0394] Speech synthesis
[0395] The generated text responses are converted into audio data by the server's speech synthesis engine. For example, a general-purpose speech synthesis engine can be used.
[0396] Response to the user
[0397] The generated audio data is provided to the user through the robot's speaker.
[0398] Specific example
[0399] User Questions
[0400] For example, a user might ask the robot, "What are the latest smartphone models?"
[0401] Server-side processing
[0402] The server receives the voice data and uses a speech recognition engine to convert it into text data. The converted text will say, "Please tell me the latest smartphone model." The NLP engine analyzes this content and recognizes that the request is for information about the latest smartphones.
[0403] Based on these analysis results, the server queries the database to retrieve information about the latest smartphones. For example, it might retrieve data indicating that "the latest smartphone is the Model X."
[0404] Based on the acquired information, a generative artificial intelligence model generates a natural language response: "The latest smartphone is the Model X." This text response is then converted into speech data by a speech synthesis engine.
[0405] Robot response
[0406] Finally, the generated audio data is delivered to the user through the robot's speaker, and the robot says, "The latest smartphone is the Model X."
[0407] In this way, the system of the present invention can provide users with quick and accurate answers to their technical questions.
[0408] The following describes the processing flow.
[0409] Step 1:
[0410] The user inputs a question to the robot using voice. For example, they might say, "Please tell me the latest smartphone model."
[0411] Step 2:
[0412] The robot acquires the user's voice input and sends the voice data to the server.
[0413] Step 3:
[0414] The server passes the received audio data to the speech recognition engine, which converts the audio into text data. For example, the text data "Please tell me the latest smartphone model" is generated.
[0415] Step 4:
[0416] The server passes the converted text data to a natural language processing (NLP) engine to analyze the text's intent. This engine recognizes that the user is seeking information about the latest smartphones.
[0417] Step 5:
[0418] The server uses the analysis results obtained from the NLP engine to query the database and retrieve relevant information. For example, the database might return information such as "The latest smartphone is the Model X."
[0419] Step 6:
[0420] The server passes the acquired information to a generative artificial intelligence model to generate a response in natural language form. The generative AI generates the text response "The latest smartphone is the Model X."
[0421] Step 7:
[0422] The server passes the generated text response to the speech synthesis engine, which converts it into speech data. For example, it might generate speech data saying, "The latest smartphone is the Model X."
[0423] Step 8:
[0424] The server sends the generated audio data to the robot. The robot plays the audio data to the user through its speaker. The user hears the voice response, "The latest smartphone is the Model X."
[0425] This will enable a system where users can obtain immediate and accurate answers to their questions.
[0426] (Example 1)
[0427] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0428] Conventional systems for answering technical questions had limited ability to provide quick and accurate answers based on user voice input. Furthermore, they struggled to properly analyze voice input, understand its intent, and provide relevant information, and lacked sufficient means to efficiently search for desired information from large databases. As a result, it was difficult for users to quickly and accurately access the information they needed.
[0429] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0430] In this invention, the server includes means for acquiring user voice input, means for converting the voice input into text data, means for analyzing the intent of the text data, means for generating a natural language response from the acquired information, means for converting the generated response into voice data, and means for providing the voice data to the user. This makes it possible to provide quick and accurate answers to the user's technical questions.
[0431] "User voice input" refers to questions or instructions that a user makes to the system using their voice.
[0432] "Means for acquiring voice input" refers to devices or technologies for capturing the user's voice, such as a microphone.
[0433] "Means of converting voice input into text data" refers to technologies for converting acquired speech into text information, such as speech recognition engines.
[0434] "Means for analyzing the intent of text data" refers to technologies for understanding the content of converted text data and analyzing the user's intent, such as natural language processing engines.
[0435] "Means of obtaining relevant information from a database based on analysis results" refers to technologies and systems that execute queries on a database according to the analyzed intent and obtain the necessary information.
[0436] "Means for generating natural language responses from acquired information" refers to technologies that create human-readable text based on information obtained from a database, such as generative artificial intelligence models.
[0437] "Means of converting generated responses into audio data" refers to technologies that convert natural language text into speech, such as speech synthesis engines.
[0438] "Means of providing audio data to the user" refers to devices and technologies for transmitting generated audio to the user, such as speakers.
[0439] This invention is a system for quickly and accurately answering technical questions based on user voice input. The system of this invention performs a series of processes, including acquiring user voice input, converting it into text data, analyzing the intent of the text, retrieving relevant information from a database, generating a natural language response, and converting it back into voice data to provide to the user. Specific hardware and software are used for each processing step.
[0440] First, the user asks a question to the robot using voice. This voice data is acquired through a microphone built into the robot (for example, an omnidirectional microphone). The acquired voice data is then transmitted to a server via the internet. This process utilizes a built-in microphone and a network module.
[0441] Next, the server converts the received audio data into text data using a speech recognition engine (for example, a general-purpose speech recognition engine). The speech recognition engine analyzes the waveform data of the audio and converts it into corresponding character information. This process converts the audio into text data.
[0442] The converted text data is analyzed for intent by a natural language processing (NLP) engine on the server (for example, a general-purpose NLP engine). The NLP engine analyzes keywords and context within the text to understand the user's inquiry. This analysis enables queries to be made to the appropriate database.
[0443] The server then queries the database based on the analysis results to retrieve the necessary information. This database may contain information such as mobile phone model details or internet connection information. The server creates an SQL query to retrieve the relevant information from the database.
[0444] Based on the acquired information, the server's generative artificial intelligence model (for example, a general generative AI model) generates a natural language response. The generative AI model creates natural-sounding sentences and provides the response in a way that is easy for the user to understand.
[0445] Next, the generated text responses are converted into audio data by the server's speech synthesis engine (for example, a general-purpose speech synthesis engine). The speech synthesis engine converts the text data into audio waveform data and generates a natural-sounding human voice.
[0446] Ultimately, the audio data is delivered to the user through the robot's speaker (for example, a standard speaker). The generated audio data is played back through the speaker and provided in a format that the user can hear.
[0447] As a concrete example, consider a scenario where a user asks a robot, "What is the latest smartphone model?" The user's voice data is acquired through a microphone and sent to a server. The server uses a speech recognition engine to convert the voice data into text data, and an NLP engine analyzes the intent of the text data. Based on the analysis results, the server queries a database to obtain the latest smartphone information. A generative artificial intelligence model generates a natural language response, "The latest smartphone is the Model X," which is then converted into voice data by a speech synthesis engine. Finally, the voice message "The latest smartphone is the Model X" is delivered to the user from the robot's speaker. In this way, the system of the present invention can provide quick and accurate answers to users' technical questions.
[0448] To realize the invention, it is crucial to properly configure and coordinate these hardware and software components.
[0449] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0450] Step 1:
[0451] The user asks the robot a question using voice. This question becomes the system's input. For example, the user might say, "Please tell me the latest smartphone model." The user's voice input is captured by the device's microphone.
[0452] Step 2:
[0453] The terminal (robot) sends the acquired voice data to the server. Specifically, it sends the acquired voice data to the server via the internet. At this time, the input voice data is transmitted directly to the server, and the server receives the voice data.
[0454] Step 3:
[0455] The server converts the received audio data into text data using a general speech recognition engine. Specifically, it analyzes the audio data and outputs text data as character information. For example, the audio data "Please tell me the latest smartphone model" is converted into the corresponding string.
[0456] Step 4:
[0457] The server uses a natural language processing (NLP) engine to analyze the intent of the converted text data. This process analyzes the context and keywords of the text and outputs the intent of the user's question. For example, from the text data "Please tell me the latest smartphone models," the server recognizes that the user is seeking information about the latest smartphones.
[0458] Step 5:
[0459] The server queries the database based on the analysis results to retrieve the necessary information. In this step, the analysis results are used as input to generate an SQL query, which is then sent to the database to output information about the latest smartphones. For example, the server retrieves the information that "the latest smartphone is the Model X" from the database.
[0460] Step 6:
[0461] The server uses a generative artificial intelligence model (for example, a general generative AI model) to generate natural language responses based on the acquired information. It uses the acquired information as input to generate text in a user-friendly format. For example, it might output the text response, "The latest smartphone is the Model X."
[0462] Step 7:
[0463] The server converts the generated text responses into speech data using a common speech synthesis engine. It uses the input text data to output speech waveform data. For example, the text "The latest smartphone is the Model X" is converted into natural-sounding speech.
[0464] Step 8:
[0465] The terminal (robot) provides the user with audio data sent from the server via its speaker. By receiving the audio data output from the server and playing it back through the speaker, it provides the user with an audio response such as, "The latest smartphone is the Model X."
[0466] (Application Example 1)
[0467] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0468] In traditional virtual stores, users lacked sufficient means to obtain appropriate and timely answers when asking questions about specific products or popular items. Furthermore, answers were often provided only in text format, limiting the user experience. Additionally, traditional systems struggled to efficiently analyze user voice input, accurately understand their intent, and generate responses.
[0469] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0470] In this invention, the server includes means for acquiring user voice input, means for converting the voice input into text data, means for analyzing the intent of the text data, means for acquiring relevant information from a database based on the analysis results, means for generating a natural language response from the acquired information, means for converting the generated response into voice data, means for providing the voice data to the user, and means for responding to questions regarding product information in the virtual store. This makes it possible to provide quick and accurate voice responses within the virtual store to questions asked by the user by voice.
[0471] "Means for acquiring user voice input" refers to a device or method used to acquire the voice spoken by the user.
[0472] "Means for converting voice input into text data" refers to a device or method for converting acquired voice into text information.
[0473] "Means for analyzing the intent of text data" refers to a device or method for understanding the user's intent or the content of a question from text.
[0474] "Means for obtaining relevant information from a database based on analysis results" refers to a device or method for searching and obtaining necessary data from a database according to the analyzed information.
[0475] "Means for generating natural language responses from acquired information" refers to a device or method for generating responses in natural language that humans can understand, based on acquired data.
[0476] "Means for converting generated responses into audio data" refers to a device or method for converting text-based responses into audio format.
[0477] "Means for providing audio data to a user" refers to a device or method for delivering generated audio data to a user.
[0478] "Means for responding to product information questions in a virtual store" refers to a device or method for users to ask questions about products within a virtual environment and for providing answers to those questions.
[0479] The system of this invention is designed to allow users in a virtual store to ask questions about product information using voice, and to provide quick and accurate answers to those questions. Specific embodiments of this system are described below.
[0480] 1. System Overview
[0481] Users ask voice questions within a virtual store via smartphone, smart glasses, or head-mounted display. The system receives the voice input, processes it on a server, and then provides a voice response. The main components are as follows:
[0482] User devices: Smartphones, smart glasses, head-mounted displays, etc.
[0483] server:
[0484] Speech recognition engine: Uses the Google Cloud Speech-to-Text API.
[0485] Natural Language Processing (NLP) engine: We use Amazon Comprehend from AWS.
[0486] Database: AWS RDS (Relational Database Service) is used.
[0487] Generative AI model: OpenAI GPT-3 is used.
[0488] Text-to-speech engine: Uses Google Cloud Text-to-Speech.
[0489] 2. Processing Flow
[0490] 1. Acquisition of voice input:
[0491] The user terminal acquires voice input and sends it to the server. The terminal is equipped with a built-in microphone.
[0492] 2. Speech recognition:
[0493] The server converts the received audio data into text data using the Google Cloud Speech-to-Text API.
[0494] 3. Natural Language Processing:
[0495] The converted text data is analyzed using Amazon Comprehend on AWS to understand the user's intent.
[0496] 4. Database query:
[0497] Based on the analysis results, queries are sent to AWS RDS to retrieve relevant information.
[0498] 5. Generating answers using generative AI models:
[0499] Based on the acquired information, OpenAI GPT-3 is used to generate a natural language response.
[0500] 6. Speech synthesis:
[0501] The generated text responses are converted into audio data using Google Cloud Text-to-Speech.
[0502] 7. Responding to the user:
[0503] The generated audio data is sent to the user's device, and the response is played back through the built-in speaker.
[0504] 3. Specific examples
[0505] When a user asks a question via voice, such as "What is the most popular item in this store?", the system's server performs the following actions.
[0506] The speech recognition engine converts the speech into text data that reads, "What is the most popular item in this store?"
[0507] A natural language processing engine analyzes this text and understands that the user is asking about popular products.
[0508] The database retrieves the information that "The most popular product is 'Product A'."
[0509] The generative AI model generates a natural language response that says, "The most popular product is 'Product A'."
[0510] The speech synthesis engine converts the generated text response into audio data.
[0511] Finally, the user's device plays an audio message saying, "The most popular item is 'Product A'."
[0512] Example of a prompt
[0513] When a user asks a question, it will look like this:
[0514] User voice input: What is the most popular item in this store?
[0515] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0516] Step 1:
[0517] The user terminal acquires voice input. For example, if the user says, "What is the most popular item in this store?", the terminal's microphone picks up the voice and sends it to the server as audio data. The input is the user's voice data, and the output is the audio data sent to the server.
[0518] Step 2:
[0519] The server uses a speech recognition engine to convert the audio data into text data. The Google Cloud Speech-to-Text API is used to transcribe the audio data. The input is the user's voice data, and the output is the text data "What is the most popular item in this store?".
[0520] Step 3:
[0521] The server uses a natural language processing engine to analyze the intent of text data. Using Amazon Comprehend from AWS, it understands that the text data is a question about popular products. The input is text data, and the output is intent data indicating "I am requesting information about popular products."
[0522] Step 4:
[0523] The server sends a query to the database based on the data analysis results and retrieves relevant information. It executes a query to AWS RDS to "get popular products" and retrieves the data "The popular product is 'Product A'." The input is intent data, and the output is informational data "The popular product is 'Product A'."
[0524] Step 5:
[0525] The server uses a generation AI model to generate natural language responses based on the acquired information. Using OpenAI GPT-3, it creates a natural-sounding sentence such as "The most popular product is 'Product A'." The input is informational data, and the output is the natural language response "The most popular product is 'Product A'."
[0526] Step 6:
[0527] The server converts natural language responses generated using a speech synthesis engine into audio data. The Google Cloud Text-to-Speech API is used to convert text data into audio data. The input is text data, and the output is audio data.
[0528] Step 7:
[0529] The server sends the generated audio data to the user's terminal, which then outputs it through the terminal's speaker. The user's terminal's built-in speaker plays the audio saying, "The most popular product is 'Product A'." The input is audio data, and the output is the audio response provided to the user.
[0530] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0531] The present invention is a system for quickly and accurately answering technical questions such as mobile phone model descriptions and internet connection methods based on the user's voice input, and furthermore, it is a system that provides more human-like interaction by recognizing the user's emotions and adjusting the response accordingly. Specific embodiments of this system are described below.
[0532] This system performs a series of processes: it acquires the user's voice input, converts it into text data, analyzes the intent of the text, retrieves relevant information from a database, generates a natural language response, converts it back into voice data, and provides it to the user. It also recognizes the user's emotions from their voice and adjusts the tone and content of the response accordingly.
[0533] System Configuration
[0534] Get user input
[0535] The user asks the robot a question using voice. This voice data is acquired through the microphone built into the robot.
[0536] Speech recognition
[0537] The acquired audio data is sent to a server and converted into text data by a speech recognition engine. For example, a general-purpose speech recognition engine can be used.
[0538] Natural Language Processing
[0539] The converted text data is analyzed for intent by a natural language processing (NLP) engine on the server. The NLP engine understands the user's inquiry and performs analysis to process it appropriately.
[0540] emotion recognition
[0541] Simultaneously, the acquired audio data is analyzed by an emotion engine. The emotion engine recognizes emotions based on the tone and pitch of the user's voice. For example, it determines emotions such as anger, sadness, and happiness.
[0542] Database query
[0543] Based on the analysis results, the server queries the database to retrieve the necessary information. The database stores information such as mobile phone model and internet connection details.
[0544] Answer generation using generative artificial intelligence models
[0545] The acquired information is used by the server's generative artificial intelligence model to generate a natural language response. The generative AI model considers emotional information from the emotion engine and adjusts the tone and content of the response.
[0546] Speech synthesis
[0547] The generated text responses are converted into audio data by the server's speech synthesis engine. For example, a general-purpose speech synthesis engine can be used.
[0548] Response to the user
[0549] The generated voice data is delivered to the user through the robot's speaker. The robot responds in a tone that matches the user's emotions.
[0550] Specific example
[0551] User Questions
[0552] The user asks the robot, "Please tell me the latest smartphone model."
[0553] Server-side processing
[0554] The server receives the voice data and uses a speech recognition engine to convert it into text data. The converted text will say, "Please tell me the latest smartphone model." The NLP engine analyzes this content and recognizes that the request is for information about the latest smartphones.
[0555] At the same time, the emotion engine determines the user's emotions from their voice. For example, it might determine that the user is excited.
[0556] Based on the analysis results, the server queries the database and retrieves the information that "the latest smartphone is the Model X."
[0557] Based on the acquired information, a generative AI model generates a natural language response such as "The latest smartphone is the Model X." Taking into account the information from the emotion engine, the generative AI model generates a response in an appropriate tone to the user's emotions, such as "Sorry to keep you waiting! The latest smartphone is the Model X."
[0558] Robot response
[0559] Finally, the generated audio data is delivered to the user through the robot's speaker, and the robot says, "Sorry to keep you waiting! The latest smartphone is the Model X."
[0560] In this way, the system of the present invention can improve the user experience by providing quick and accurate answers to the user's technical questions, as well as by providing responses that take the user's feelings into consideration.
[0561] The following describes the processing flow.
[0562] Step 1:
[0563] The user inputs a question to the robot using voice. For example, they might say, "Please tell me the latest smartphone model."
[0564] Step 2:
[0565] The robot acquires the user's voice input and sends the voice data to the server.
[0566] Step 3:
[0567] The server passes the received audio data to the speech recognition engine, which converts the audio into text data. For example, the text data "Please tell me the latest smartphone model" is generated.
[0568] Step 4:
[0569] The server passes the converted text data to a natural language processing (NLP) engine to analyze the text's intent. This engine recognizes that the user is seeking information about the latest smartphones.
[0570] Step 5:
[0571] The server simultaneously passes the voice data to the emotion engine to analyze the user's emotions. The emotion engine analyzes the tone, pitch, and speed of the voice to determine the user's emotional state, such as being excited, angry, or sad.
[0572] Step 6:
[0573] Based on the analysis results of the NLP engine, the server queries the database to retrieve relevant information. For example, the database might return information such as "The latest smartphone is the Model X."
[0574] Step 7:
[0575] The server passes the acquired information to a generative artificial intelligence model to generate a response in natural language form. This model also takes emotional information from the emotion engine into consideration, and generates responses such as, "Sorry to keep you waiting! The latest smartphone is the Model X."
[0576] Step 8:
[0577] The server passes the generated text response to the speech synthesis engine, which converts it into speech data. For example, it might generate speech data such as, "Sorry to keep you waiting! The latest smartphone is the Model X."
[0578] Step 9:
[0579] The server sends the generated audio data to the robot. The robot plays the audio data to the user through its speaker. The user hears the voice response, "Thank you for waiting! The latest smartphone is the Model X."
[0580] This will create a system where users can not only get quick and accurate answers to their questions, but also receive responses that are tailored to their emotions.
[0581] (Example 2)
[0582] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0583] Conventional technical question answering systems have a problem of degrading the user experience because they cannot respond in a way that takes into account the user's emotional state. Furthermore, there are limitations in processing speed and accuracy for reliably acquiring information from voice input and providing appropriate answers. As a result, while users desire accurate and rapid answers, as well as emotionally responsive interaction, it has been difficult to effectively provide these.
[0584] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0585] In this invention, the server includes means for acquiring user voice input, means for converting the voice input into text data, means for analyzing the intent of the text data, means for recognizing emotions based on the acquired voice data, means for acquiring relevant information from an information organizing device based on the analysis results, means for converting the generated response into a natural language response, means for adjusting the tone and content of the response based on the emotion recognition results, means for converting the generated response into voice data, and means for providing the voice data to the user. This enables the user to obtain quick and accurate technical answers and to experience more human-like interactions through emotionally responsive interactions.
[0586] "Means for obtaining user voice input" refers to a device or part of a device used to collect voice data, and includes, for example, a microphone.
[0587] "Means for converting voice input into text data" refers to a process or device that converts voice data into a string of characters, and a speech recognition engine is used as an example.
[0588] "Means for analyzing the intent of text data" refers to processes or devices for understanding the content of text data and identifying its purpose or requirements, and includes, for example, natural language processing engines.
[0589] "Means of recognizing emotions based on acquired audio data" refers to a process or device that analyzes the tone and pitch of speech to determine the speaker's emotional state, and includes, for example, an emotion recognition engine.
[0590] "Means for obtaining relevant information from an information organization device based on analysis results" refers to a process or device for searching for and obtaining necessary information based on analysis results, and includes database queries as an example.
[0591] "Means for converting generated responses into natural language responses" refers to a process or device that converts acquired information into a natural language format that humans can understand, and generative AI models are used as an example.
[0592] "Means for adjusting the tone and content of responses based on the results of emotion recognition" refers to a process or device that appropriately adjusts the generated response content and expression method while taking into account the user's emotional state.
[0593] "Means for converting generated responses into audio data" refers to a process or device that converts text responses back into audio format, and includes, for example, a speech synthesis engine.
[0594] "Means of providing audio data to the user" refers to a process or device for making the generated audio data listen to the user, and a speaker is used as an example.
[0595] Modes for carrying out the invention
[0596] The present invention is a system that answers technical questions quickly and accurately based on the user's voice input, and further provides a more human-like interaction by recognizing the user's emotions and adjusting the response accordingly. Specific embodiments of this system are described below.
[0597] First, the user asks the robot a question using their voice. For example, they might say, "What are the latest smartphone models?" This voice data is captured through the microphone built into the robot.
[0598] The acquired voice data is sent from the robot to the server. The server uses a speech recognition engine (for example, Google Speech-to-Text API) to convert the voice data into text data. As a result, the voice data is converted into text data that says, "Please tell me the latest smartphone model."
[0599] Next, this text data is passed to a natural language processing (NLP) engine on the server (for example, Google Cloud Natural Language API) where the intent is analyzed. In this case, the NLP engine analyzes that the user is looking for information about the latest smartphones.
[0600] Simultaneously, the acquired audio data is analyzed by an emotion recognition engine on the server (for example, IBM Watson Tone Analyzer). Based on the tone and pitch of the voice, the emotion recognition engine determines whether the user is excited, angry, happy, or otherwise experiencing other emotions.
[0601] The server retrieves relevant information from the database based on the analysis results of the NLP engine. The database contains technical information and connection methods for a wide variety of devices. In this case, let's assume we retrieve the information that "the latest smartphone is the Model X."
[0602] Based on the acquired information, the server uses a generative artificial intelligence model (e.g., OpenAI GPT-3) to generate a natural language response. This model takes into account the emotional information obtained from the emotion recognition engine and adjusts the tone and content of the response. As a result, it generates a response such as, "Sorry to keep you waiting! The latest smartphone is the Model X."
[0603] Next, the generated text response is passed to the server's speech synthesis engine (e.g., Amazon Polly) and converted into audio data. This audio data is then delivered to the user through the robot's speaker. The robot then says, "Sorry to keep you waiting! The latest smartphone is the Model X," in a tone appropriate to the user's mood.
[0604] In this way, the system of this invention can provide a better user experience by offering quick and accurate answers to users' technical questions, as well as responses that take into account the user's feelings.
[0605] Examples of prompt statements
[0606] "Tell me about the latest iPhone models."
[0607] "What should I do if I can't connect to Wi-Fi?"
[0608] "How do you perceive your current emotions?"
[0609] As a result, when providing information to users, it becomes possible to accurately consider the user's emotional state and achieve natural dialogue.
[0610] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0611] Step 1: The user asks the robot a question using voice.
[0612] The user asks, "Please tell me the latest smartphone model." The robot's built-in microphone then captures the voice data.
[0613] Input: User's voice
[0614] Output: Acquired audio data
[0615] Step 2: The device sends the acquired audio data to the server.
[0616] The robot transmits the acquired audio data to the server using wireless communication (e.g., Wi-Fi).
[0617] Input: Acquired audio data
[0618] Output: Audio data sent to the server
[0619] Step 3: The server uses a speech recognition engine to convert the speech data into text data.
[0620] The server calls a speech recognition engine (e.g., Google Speech-to-Text API), takes the voice data "Please tell me the latest smartphone model" as input, and outputs text data.
[0621] Input: Audio data sent to the server
[0622] Output: Converted text data
[0623] Step 4: The server uses a natural language processing (NLP) engine to analyze the intent of the text data.
[0624] The converted text data is input into an NLP engine (e.g., Google Cloud Natural Language API) to analyze that the user is seeking information about the latest smartphones. The output is the analysis result.
[0625] Input: Converted text data
[0626] Output: Analysis results
[0627] Step 5: The server recognizes emotions based on the acquired audio data.
[0628] The server inputs voice data into an emotion recognition engine (e.g., IBM Watson Tone Analyzer) and outputs the emotion recognition result. For example, it might produce a result indicating that the user is excited.
[0629] Input: Audio data sent to the server
[0630] Output: Emotion recognition result
[0631] Step 6: The server retrieves relevant information from the information organization device based on the analysis results.
[0632] The server executes a query against the database, asking "What is the latest smartphone?" and retrieving relevant information. For example, it might retrieve information such as "The latest smartphone is the Model X."
[0633] Input: Analysis results
[0634] Output: Related information obtained
[0635] Step 7: The server generates a natural language response based on the information it has obtained.
[0636] The server inputs the prompt "What is the latest smartphone?" into a generated AI model (e.g., OpenAI GPT-3) and generates the answer "The latest smartphone is the Model X." Taking sentiment into account, it generates the answer in an appropriate tone, such as "Sorry to keep you waiting! The latest smartphone is the Model X."
[0637] Input: Acquired related information, emotion recognition results
[0638] Output: Generated text answer
[0639] Step 8: The server converts the generated text response into speech data using a speech synthesis engine.
[0640] The generated text responses are passed to a speech synthesis engine (e.g., Amazon Polly) and converted into audio data.
[0641] Input: Generated text answer
[0642] Output: Audio data
[0643] Step 9: The device provides the generated audio data to the user.
[0644] The robot's speaker uses the converted voice data to say, "Sorry to keep you waiting! The latest smartphone is the Model X."
[0645] Input: Audio data
[0646] Output: Voice response by robot
[0647] (Application Example 2)
[0648] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0649] In modern brick-and-mortar stores, there is a demand for quick and accurate responses to customer questions and inquiries. However, conventional methods often fail to provide flexible responses that respond to customer emotions, which can lead to decreased customer satisfaction. This invention aims to solve the problem of improving customer satisfaction by not only answering customers' technical questions but also recognizing their emotions and adjusting the content and tone of responses accordingly, thereby providing a more humane interaction.
[0650] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring user voice input, means for converting voice input into text data, means for analyzing the intent of the text data, means for acquiring relevant information from a database based on the analysis results, means for generating a natural language response from the acquired information, means for converting the generated response into voice data, means for providing the voice data to the user, and means for recognizing emotions from the user's voice and adjusting the response. This makes it possible to answer customer questions quickly and accurately in a physical store while providing appropriate responses that correspond to the customer's emotions.
[0651] "Means for acquiring voice input" refers to devices or interfaces that allow users to input questions or instructions to a system using their voice.
[0652] "Means for converting voice input into text data" refers to software or algorithms used to convert acquired voice data into text information.
[0653] "Means for analyzing the intent of text data" refers to natural language processing techniques used to understand the content of user requests and questions from textual information.
[0654] "Means for obtaining relevant information from a database based on analysis results" refers to a mechanism for searching and retrieving relevant answers and data from a database based on the analyzed intent.
[0655] "Means for generating natural language responses from acquired information" refers to technologies such as generative artificial intelligence models that generate responses in a human-understandable format based on acquired data.
[0656] "Means for converting generated responses into audio data" refers to speech synthesis technology for converting text-based responses into audio format.
[0657] "Means of providing audio data to the user" refers to speakers or other output devices that allow the user to listen to the generated audio data.
[0658] "Means for recognizing emotions from a user's voice and adjusting responses" refers to algorithms or engines that detect emotions from the user's voice tone and pitch, and adjust responses to an appropriate tone and content.
[0659] An example of an embodiment of the present invention is a customer service robot system for use in physical stores. Specifically, this system allows customers to ask questions to the robot by voice, and the robot provides quick and accurate answers according to the content of the questions, while also recognizing the customer's emotions and adjusting the content and tone of the response.
[0660] System Configuration
[0661] This system includes the following components.
[0662] Get user input
[0663] The user asks the robot a question using voice. For example, "Where can I find this product?" This voice data is captured through a microphone built into the robot. Common microphones can be used as hardware (e.g., Behringer ECM8000, Shure MV88, etc.).
[0664] Speech recognition
[0665] The acquired audio data is sent to a cloud service (e.g., Amazon Transcribe, Google Speech-to-Text) and converted into text data by a speech recognition engine. In this step, the audio data is converted into the text "Where can I find this product?".
[0666] Natural Language Processing
[0667] The converted text data is analyzed by a natural language processing (NLP) engine on the server (e.g., Google Cloud Natural Language API, Microsoft Azure Text Analytics API, etc.). At this stage, the intent of the text is understood to be "checking product location."
[0668] emotion recognition
[0669] Simultaneously, the acquired audio data is analyzed by an emotion engine (e.g., IBM Watson Tone Analyzer, Amazon Comprehend). Based on the voice tone and pitch, the emotion engine determines, for example, that the customer is "irritated."
[0670] Database query
[0671] Based on the analysis results, the server queries a database (e.g., MySQL, Firebase, etc.) to retrieve the necessary information. For example, it might retrieve information such as, "Product A is located on the central shelf of the third street."
[0672] Answer generation using generative artificial intelligence models
[0673] Based on the acquired information, a generative artificial intelligence model (e.g., OpenAI GPT-4, Hugging Face Transformers, etc.) generates a natural language response. Taking into account the results of emotion recognition, a response such as "Excuse me, I apologize for the inconvenience. Item A is on the central shelf of the third lane" is generated.
[0674] Speech synthesis
[0675] The generated text responses are converted into speech data by a speech synthesis engine (e.g., Google Text-to-Speech, Amazon Polly, etc.).
[0676] Response to the user
[0677] The generated voice data is delivered to the user through the robot's speaker (e.g., JBL Charge 4, Sonos Move, etc.). The robot says, "Excuse me, I'm sorry to trouble you. Item A is on the central shelf on the third street."
[0678] Specific example
[0679] User Questions
[0680] The user asks the robot, "Where can I find this product?"
[0681] Server-side processing
[0682] The server receives the voice data and uses a speech recognition engine to convert it into text data. The converted text will say, "Where is this product?" The NLP engine analyzes this and recognizes that the user is requesting the location information of a product. At the same time, the emotion engine determines the emotion from the user's voice. For example, it may determine that the user is irritated. Based on the analysis results, the server queries the database and obtains the information that "Product A is on the central shelf of the third street." Based on the obtained information, a generative artificial intelligence model generates a natural language response: "Excuse me, we apologize for the inconvenience. Product A is on the central shelf of the third street."
[0683] Robot response
[0684] Finally, the generated audio data is delivered to the user through the robot's speaker, and the robot says, "Excuse me, I'm sorry to trouble you. Item A is on the central shelf on the third street."
[0685] Example of a prompt
[0686] User question: "Where can I find this product?"
[0687] NLP analysis results: The product's location information is being sought.
[0688] Sentiment analysis result: The customer is irritated.
[0689] Database query result: "Product A is located on the central shelf of the third aisle."
[0690] Output from the generated AI model: "Excuse me, we apologize for the inconvenience. Product A is located on the central shelf of the third aisle."
[0691] In this way, the system of the present invention can improve the customer experience by providing quick and accurate answers to customers' technical questions in physical stores and by responding in a way that takes customer emotions into consideration.
[0692] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0693] Step 1:
[0694] The user asks the robot questions using voice. For example, a voice input such as "Where can I find this product?" is acquired through the robot's microphone (e.g., Behringer ECM8000, Shure MV88, etc.). The input is the user's voice data, and the output is that voice data itself.
[0695] Step 2:
[0696] The audio data is sent to a server and converted into text data by a speech recognition engine (e.g., Amazon Transcribe, Google Speech-to-Text). In this step, the audio data is converted into the text "Where can I find this product?". The input is the audio data, and the output is the text data that is a conversion of that audio data.
[0697] Step 3:
[0698] The server analyzes the converted text data using a natural language processing (NLP) engine (e.g., Google Cloud Natural Language API, Microsoft Azure Text Analytics API, etc.) to understand the user's intent. In this step, the intent "check product location" is analyzed. The input is text data, and the output is the analyzed intent information.
[0699] Step 4:
[0700] Simultaneously, the voice data is analyzed by an emotion engine (e.g., IBM Watson Tone Analyzer, Amazon Comprehend) to recognize the customer's emotions. For example, it might determine that the customer is "irritated." The input is voice data, and the output is information that expresses emotion.
[0701] Step 5:
[0702] Based on the analysis results of the NLP engine, the server queries a database (e.g., MySQL, Firebase, etc.) to retrieve relevant information. In this step, the information "Product A is on the central shelf of the third street" is retrieved. The input is the analyzed intent information, and the output is the relevant information retrieved from the database.
[0703] Step 6:
[0704] The server uses a generative artificial intelligence model (e.g., OpenAI GPT-4, Hugging Face Transformers, etc.) to generate a natural language response based on the acquired information. Sentiment recognition results are also considered during this process. For example, a response like, "Excuse me, I apologize for the inconvenience. Product A is on the central shelf of the third street," might be generated. The input consists of relevant and sentiment information, and the output is the generated natural language response.
[0705] Step 7:
[0706] The generated text responses are converted into audio data by a text-to-speech engine (e.g., Google Text-to-Speech, Amazon Polly, etc.). In this step, the text-formatted responses are converted into audio format. The input is the text-formatted responses, and the output is audio data.
[0707] Step 8:
[0708] Finally, the generated audio data is delivered to the user through the robot's speaker (e.g., JBL Charge 4, Sonos Move, etc.). The robot says, "Excuse me, I'm sorry to trouble you. Item A is on the central shelf on the third street." The input is audio data, and the output is the audio the user hears.
[0709] In this way, the system of the present invention can provide quick and accurate answers to users' technical questions, as well as respond in a way that takes the user's feelings into consideration. This series of processes is expected to improve the customer experience in physical stores.
[0710] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0711] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0712] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0713] [Third Embodiment]
[0714] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0715] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0716] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0717] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0718] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0719] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0720] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0721] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0722] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0723] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0724] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0725] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0726] This invention is a system for quickly and accurately answering technical questions, such as mobile phone model descriptions and internet connection methods, based on user voice input. Specific embodiments of this system are described below.
[0727] This system performs a series of processes: it acquires the user's voice input, converts it into text data, analyzes the intent of the text, retrieves relevant information from a database, generates a natural language response, converts it back into voice data, and provides it to the user.
[0728] System Configuration
[0729] Get user input
[0730] The user asks the robot a question using voice. This voice data is acquired through the microphone built into the robot.
[0731] Speech recognition
[0732] The acquired audio data is sent to a server and converted into text data by a speech recognition engine. For example, a general-purpose speech recognition engine can be used.
[0733] Natural Language Processing
[0734] The converted text data is analyzed for intent by a natural language processing (NLP) engine on the server. The NLP engine understands the user's inquiry and performs analysis to process it appropriately.
[0735] Database query
[0736] Based on the analysis results, the server queries the database to retrieve the necessary information. The database stores information such as mobile phone model and internet connection details.
[0737] Answer generation using generative artificial intelligence models
[0738] The acquired information is generated into natural language responses by the server's generative artificial intelligence model. The generative AI model generates text that provides the information in a way that is easy for the user to understand.
[0739] Speech synthesis
[0740] The generated text responses are converted into audio data by the server's speech synthesis engine. For example, a general-purpose speech synthesis engine can be used.
[0741] Response to the user
[0742] The generated audio data is provided to the user through the robot's speaker.
[0743] Specific example
[0744] User Questions
[0745] For example, a user might ask the robot, "What are the latest smartphone models?"
[0746] Server-side processing
[0747] The server receives the voice data and uses a speech recognition engine to convert it into text data. The converted text will say, "Please tell me the latest smartphone model." The NLP engine analyzes this content and recognizes that the request is for information about the latest smartphones.
[0748] Based on these analysis results, the server queries the database to retrieve information about the latest smartphones. For example, it might retrieve data indicating that "the latest smartphone is the Model X."
[0749] Based on the acquired information, a generative artificial intelligence model generates a natural language response: "The latest smartphone is the Model X." This text response is then converted into speech data by a speech synthesis engine.
[0750] Robot response
[0751] Finally, the generated audio data is delivered to the user through the robot's speaker, and the robot says, "The latest smartphone is the Model X."
[0752] In this way, the system of the present invention can provide users with quick and accurate answers to their technical questions.
[0753] The following describes the processing flow.
[0754] Step 1:
[0755] The user inputs a question to the robot using voice. For example, they might say, "Please tell me the latest smartphone model."
[0756] Step 2:
[0757] The robot acquires the user's voice input and sends the voice data to the server.
[0758] Step 3:
[0759] The server passes the received audio data to the speech recognition engine, which converts the audio into text data. For example, the text data "Please tell me the latest smartphone model" is generated.
[0760] Step 4:
[0761] The server passes the converted text data to a natural language processing (NLP) engine to analyze the text's intent. This engine recognizes that the user is seeking information about the latest smartphones.
[0762] Step 5:
[0763] The server uses the analysis results obtained from the NLP engine to query the database and retrieve relevant information. For example, the database might return information such as "The latest smartphone is the Model X."
[0764] Step 6:
[0765] The server passes the acquired information to a generative artificial intelligence model to generate a response in natural language form. The generative AI generates the text response "The latest smartphone is the Model X."
[0766] Step 7:
[0767] The server passes the generated text response to the speech synthesis engine, which converts it into speech data. For example, it might generate speech data saying, "The latest smartphone is the Model X."
[0768] Step 8:
[0769] The server sends the generated audio data to the robot. The robot plays the audio data to the user through its speaker. The user hears the voice response, "The latest smartphone is the Model X."
[0770] This will enable a system where users can obtain immediate and accurate answers to their questions.
[0771] (Example 1)
[0772] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0773] Conventional systems for answering technical questions had limited ability to provide quick and accurate answers based on user voice input. Furthermore, they struggled to properly analyze voice input, understand its intent, and provide relevant information, and lacked sufficient means to efficiently search for desired information from large databases. As a result, it was difficult for users to quickly and accurately access the information they needed.
[0774] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0775] In this invention, the server includes means for acquiring user voice input, means for converting the voice input into text data, means for analyzing the intent of the text data, means for generating a natural language response from the acquired information, means for converting the generated response into voice data, and means for providing the voice data to the user. This makes it possible to provide quick and accurate answers to the user's technical questions.
[0776] "User voice input" refers to questions or instructions that a user makes to the system using their voice.
[0777] "Means for acquiring voice input" refers to devices or technologies for capturing the user's voice, such as a microphone.
[0778] "Means of converting voice input into text data" refers to technologies for converting acquired speech into text information, such as speech recognition engines.
[0779] "Means for analyzing the intent of text data" refers to technologies for understanding the content of converted text data and analyzing the user's intent, such as natural language processing engines.
[0780] "Means of obtaining relevant information from a database based on analysis results" refers to technologies and systems that execute queries on a database according to the analyzed intent and obtain the necessary information.
[0781] "Means for generating natural language responses from acquired information" refers to technologies that create human-readable text based on information obtained from a database, such as generative artificial intelligence models.
[0782] "Means of converting generated responses into audio data" refers to technologies that convert natural language text into speech, such as speech synthesis engines.
[0783] "Means of providing audio data to the user" refers to devices and technologies for transmitting generated audio to the user, such as speakers.
[0784] This invention is a system for quickly and accurately answering technical questions based on user voice input. The system of this invention performs a series of processes, including acquiring user voice input, converting it into text data, analyzing the intent of the text, retrieving relevant information from a database, generating a natural language response, and converting it back into voice data to provide to the user. Specific hardware and software are used for each processing step.
[0785] First, the user asks a question to the robot using voice. This voice data is acquired through a microphone built into the robot (for example, an omnidirectional microphone). The acquired voice data is then transmitted to a server via the internet. This process utilizes a built-in microphone and a network module.
[0786] Next, the server converts the received audio data into text data using a speech recognition engine (for example, a general-purpose speech recognition engine). The speech recognition engine analyzes the waveform data of the audio and converts it into corresponding character information. This process converts the audio into text data.
[0787] The converted text data is analyzed for intent by a natural language processing (NLP) engine on the server (for example, a general-purpose NLP engine). The NLP engine analyzes keywords and context within the text to understand the user's inquiry. This analysis enables queries to be made to the appropriate database.
[0788] The server then queries the database based on the analysis results to retrieve the necessary information. This database may contain information such as mobile phone model details or internet connection information. The server creates an SQL query to retrieve the relevant information from the database.
[0789] Based on the acquired information, the server's generative artificial intelligence model (for example, a general generative AI model) generates a natural language response. The generative AI model creates natural-sounding sentences and provides the response in a way that is easy for the user to understand.
[0790] Next, the generated text responses are converted into audio data by the server's speech synthesis engine (for example, a general-purpose speech synthesis engine). The speech synthesis engine converts the text data into audio waveform data and generates a natural-sounding human voice.
[0791] Ultimately, the audio data is delivered to the user through the robot's speaker (for example, a standard speaker). The generated audio data is played back through the speaker and provided in a format that the user can hear.
[0792] As a concrete example, consider a scenario where a user asks a robot, "What is the latest smartphone model?" The user's voice data is acquired through a microphone and sent to a server. The server uses a speech recognition engine to convert the voice data into text data, and an NLP engine analyzes the intent of the text data. Based on the analysis results, the server queries a database to obtain the latest smartphone information. A generative artificial intelligence model generates a natural language response, "The latest smartphone is the Model X," which is then converted into voice data by a speech synthesis engine. Finally, the voice message "The latest smartphone is the Model X" is delivered to the user from the robot's speaker. In this way, the system of the present invention can provide quick and accurate answers to users' technical questions.
[0793] To realize the invention, it is crucial to properly configure and coordinate these hardware and software components.
[0794] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0795] Step 1:
[0796] The user asks the robot a question using voice. This question becomes the system's input. For example, the user might say, "Please tell me the latest smartphone model." The user's voice input is captured by the device's microphone.
[0797] Step 2:
[0798] The terminal (robot) sends the acquired voice data to the server. Specifically, it sends the acquired voice data to the server via the internet. At this time, the input voice data is transmitted directly to the server, and the server receives the voice data.
[0799] Step 3:
[0800] The server converts the received audio data into text data using a general speech recognition engine. Specifically, it analyzes the audio data and outputs text data as character information. For example, the audio data "Please tell me the latest smartphone model" is converted into the corresponding string.
[0801] Step 4:
[0802] The server uses a natural language processing (NLP) engine to analyze the intent of the converted text data. This process analyzes the context and keywords of the text and outputs the intent of the user's question. For example, from the text data "Please tell me the latest smartphone models," the server recognizes that the user is seeking information about the latest smartphones.
[0803] Step 5:
[0804] The server queries the database based on the analysis results to retrieve the necessary information. In this step, the analysis results are used as input to generate an SQL query, which is then sent to the database to output information about the latest smartphones. For example, the server retrieves the information that "the latest smartphone is the Model X" from the database.
[0805] Step 6:
[0806] The server uses a generative artificial intelligence model (for example, a general generative AI model) to generate natural language responses based on the acquired information. It uses the acquired information as input to generate text in a user-friendly format. For example, it might output the text response, "The latest smartphone is the Model X."
[0807] Step 7:
[0808] The server converts the generated text responses into speech data using a common speech synthesis engine. It uses the input text data to output speech waveform data. For example, the text "The latest smartphone is the Model X" is converted into natural-sounding speech.
[0809] Step 8:
[0810] The terminal (robot) provides the user with audio data sent from the server via its speaker. By receiving the audio data output from the server and playing it back through the speaker, it provides the user with an audio response such as, "The latest smartphone is the Model X."
[0811] (Application Example 1)
[0812] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0813] In traditional virtual stores, users lacked sufficient means to obtain appropriate and timely answers when asking questions about specific products or popular items. Furthermore, answers were often provided only in text format, limiting the user experience. Additionally, traditional systems struggled to efficiently analyze user voice input, accurately understand their intent, and generate responses.
[0814] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0815] In this invention, the server includes means for acquiring user voice input, means for converting the voice input into text data, means for analyzing the intent of the text data, means for acquiring relevant information from a database based on the analysis results, means for generating a natural language response from the acquired information, means for converting the generated response into voice data, means for providing the voice data to the user, and means for responding to questions regarding product information in the virtual store. This makes it possible to provide quick and accurate voice responses within the virtual store to questions asked by the user by voice.
[0816] "Means for acquiring user voice input" refers to a device or method used to acquire the voice spoken by the user.
[0817] "Means for converting voice input into text data" refers to a device or method for converting acquired voice into text information.
[0818] "Means for analyzing the intent of text data" refers to a device or method for understanding the user's intent or the content of a question from text.
[0819] "Means for obtaining relevant information from a database based on analysis results" refers to a device or method for searching and obtaining necessary data from a database according to the analyzed information.
[0820] "Means for generating natural language responses from acquired information" refers to a device or method for generating responses in natural language that humans can understand, based on acquired data.
[0821] "Means for converting generated responses into audio data" refers to a device or method for converting text-based responses into audio format.
[0822] "Means for providing audio data to a user" refers to a device or method for delivering generated audio data to a user.
[0823] "Means for responding to product information questions in a virtual store" refers to a device or method for users to ask questions about products within a virtual environment and for providing answers to those questions.
[0824] The system of this invention is designed to allow users in a virtual store to ask questions about product information using voice, and to provide quick and accurate answers to those questions. Specific embodiments of this system are described below.
[0825] 1. System Overview
[0826] Users ask voice questions within a virtual store via smartphone, smart glasses, or head-mounted display. The system receives the voice input, processes it on a server, and then provides a voice response. The main components are as follows:
[0827] User devices: Smartphones, smart glasses, head-mounted displays, etc.
[0828] server:
[0829] Speech recognition engine: Uses the Google Cloud Speech-to-Text API.
[0830] Natural Language Processing (NLP) engine: We use Amazon Comprehend from AWS.
[0831] Database: AWS RDS (Relational Database Service) is used.
[0832] Generative AI model: OpenAI GPT-3 is used.
[0833] Text-to-speech engine: Uses Google Cloud Text-to-Speech.
[0834] 2. Processing Flow
[0835] 1. Acquisition of voice input:
[0836] The user terminal acquires voice input and sends it to the server. The terminal is equipped with a built-in microphone.
[0837] 2. Speech recognition:
[0838] The server converts the received audio data into text data using the Google Cloud Speech-to-Text API.
[0839] 3. Natural Language Processing:
[0840] The converted text data is analyzed using Amazon Comprehend on AWS to understand the user's intent.
[0841] 4. Database query:
[0842] Based on the analysis results, queries are sent to AWS RDS to retrieve relevant information.
[0843] 5. Generating answers using generative AI models:
[0844] Based on the acquired information, OpenAI GPT-3 is used to generate a natural language response.
[0845] 6. Speech synthesis:
[0846] The generated text responses are converted into audio data using Google Cloud Text-to-Speech.
[0847] 7. Responding to the user:
[0848] The generated audio data is sent to the user's device, and the response is played back through the built-in speaker.
[0849] 3. Specific examples
[0850] When a user asks a question via voice, such as "What is the most popular item in this store?", the system's server performs the following actions.
[0851] The speech recognition engine converts the speech into text data that reads, "What is the most popular item in this store?"
[0852] A natural language processing engine analyzes this text and understands that the user is asking about popular products.
[0853] The database retrieves the information that "The most popular product is 'Product A'."
[0854] The generative AI model generates a natural language response that says, "The most popular product is 'Product A'."
[0855] The speech synthesis engine converts the generated text response into audio data.
[0856] Finally, the user's device plays an audio message saying, "The most popular item is 'Product A'."
[0857] Example of a prompt
[0858] When a user asks a question, it will look like this:
[0859] User voice input: What is the most popular item in this store?
[0860] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0861] Step 1:
[0862] The user terminal acquires voice input. For example, if the user says, "What is the most popular item in this store?", the terminal's microphone picks up the voice and sends it to the server as audio data. The input is the user's voice data, and the output is the audio data sent to the server.
[0863] Step 2:
[0864] The server uses a speech recognition engine to convert the audio data into text data. The Google Cloud Speech-to-Text API is used to transcribe the audio data. The input is the user's voice data, and the output is the text data "What is the most popular item in this store?".
[0865] Step 3:
[0866] The server uses a natural language processing engine to analyze the intent of text data. Using Amazon Comprehend from AWS, it understands that the text data is a question about popular products. The input is text data, and the output is intent data indicating "I am requesting information about popular products."
[0867] Step 4:
[0868] The server sends a query to the database based on the data analysis results and retrieves relevant information. It executes a query to AWS RDS to "get popular products" and retrieves the data "The popular product is 'Product A'." The input is intent data, and the output is informational data "The popular product is 'Product A'."
[0869] Step 5:
[0870] The server uses a generation AI model to generate natural language responses based on the acquired information. Using OpenAI GPT-3, it creates a natural-sounding sentence such as "The most popular product is 'Product A'." The input is informational data, and the output is the natural language response "The most popular product is 'Product A'."
[0871] Step 6:
[0872] The server converts natural language responses generated using a speech synthesis engine into audio data. The Google Cloud Text-to-Speech API is used to convert text data into audio data. The input is text data, and the output is audio data.
[0873] Step 7:
[0874] The server sends the generated audio data to the user's terminal, which then outputs it through the terminal's speaker. The user's terminal's built-in speaker plays the audio saying, "The most popular product is 'Product A'." The input is audio data, and the output is the audio response provided to the user.
[0875] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0876] The present invention is a system for quickly and accurately answering technical questions such as mobile phone model descriptions and internet connection methods based on the user's voice input, and furthermore, it is a system that provides more human-like interaction by recognizing the user's emotions and adjusting the response accordingly. Specific embodiments of this system are described below.
[0877] This system performs a series of processes: it acquires the user's voice input, converts it into text data, analyzes the intent of the text, retrieves relevant information from a database, generates a natural language response, converts it back into voice data, and provides it to the user. It also recognizes the user's emotions from their voice and adjusts the tone and content of the response accordingly.
[0878] System Configuration
[0879] Get user input
[0880] The user asks the robot a question using voice. This voice data is acquired through the microphone built into the robot.
[0881] Speech recognition
[0882] The acquired audio data is sent to a server and converted into text data by a speech recognition engine. For example, a general-purpose speech recognition engine can be used.
[0883] Natural Language Processing
[0884] The converted text data is analyzed for intent by a natural language processing (NLP) engine on the server. The NLP engine understands the user's inquiry and performs analysis to process it appropriately.
[0885] emotion recognition
[0886] Simultaneously, the acquired audio data is analyzed by an emotion engine. The emotion engine recognizes emotions based on the tone and pitch of the user's voice. For example, it determines emotions such as anger, sadness, and happiness.
[0887] Database query
[0888] Based on the analysis results, the server queries the database to retrieve the necessary information. The database stores information such as mobile phone model and internet connection details.
[0889] Answer generation using generative artificial intelligence models
[0890] The acquired information is used by the server's generative artificial intelligence model to generate a natural language response. The generative AI model considers emotional information from the emotion engine and adjusts the tone and content of the response.
[0891] Speech synthesis
[0892] The generated text responses are converted into audio data by the server's speech synthesis engine. For example, a general-purpose speech synthesis engine can be used.
[0893] Response to the user
[0894] The generated voice data is delivered to the user through the robot's speaker. The robot responds in a tone that matches the user's emotions.
[0895] Specific example
[0896] User Questions
[0897] The user asks the robot, "Please tell me the latest smartphone model."
[0898] Server-side processing
[0899] The server receives the voice data and uses a speech recognition engine to convert it into text data. The converted text will say, "Please tell me the latest smartphone model." The NLP engine analyzes this content and recognizes that the request is for information about the latest smartphones.
[0900] At the same time, the emotion engine determines the user's emotions from their voice. For example, it might determine that the user is excited.
[0901] Based on the analysis results, the server queries the database and retrieves the information that "the latest smartphone is the Model X."
[0902] Based on the acquired information, a generative AI model generates a natural language response such as "The latest smartphone is the Model X." Taking into account the information from the emotion engine, the generative AI model generates a response in an appropriate tone to the user's emotions, such as "Sorry to keep you waiting! The latest smartphone is the Model X."
[0903] Robot response
[0904] Finally, the generated audio data is delivered to the user through the robot's speaker, and the robot says, "Sorry to keep you waiting! The latest smartphone is the Model X."
[0905] In this way, the system of the present invention can improve the user experience by providing quick and accurate answers to the user's technical questions, as well as by providing responses that take the user's feelings into consideration.
[0906] The following describes the processing flow.
[0907] Step 1:
[0908] The user inputs a question to the robot using voice. For example, they might say, "Please tell me the latest smartphone model."
[0909] Step 2:
[0910] The robot acquires the user's voice input and sends the voice data to the server.
[0911] Step 3:
[0912] The server passes the received audio data to the speech recognition engine, which converts the audio into text data. For example, the text data "Please tell me the latest smartphone model" is generated.
[0913] Step 4:
[0914] The server passes the converted text data to a natural language processing (NLP) engine to analyze the text's intent. This engine recognizes that the user is seeking information about the latest smartphones.
[0915] Step 5:
[0916] The server simultaneously passes the voice data to the emotion engine to analyze the user's emotions. The emotion engine analyzes the tone, pitch, and speed of the voice to determine the user's emotional state, such as being excited, angry, or sad.
[0917] Step 6:
[0918] Based on the analysis results of the NLP engine, the server queries the database to retrieve relevant information. For example, the database might return information such as "The latest smartphone is the Model X."
[0919] Step 7:
[0920] The server passes the acquired information to a generative artificial intelligence model to generate a response in natural language form. This model also takes emotional information from the emotion engine into consideration, and generates responses such as, "Sorry to keep you waiting! The latest smartphone is the Model X."
[0921] Step 8:
[0922] The server passes the generated text response to the speech synthesis engine, which converts it into speech data. For example, it might generate speech data such as, "Sorry to keep you waiting! The latest smartphone is the Model X."
[0923] Step 9:
[0924] The server sends the generated audio data to the robot. The robot plays the audio data to the user through its speaker. The user hears the voice response, "Thank you for waiting! The latest smartphone is the Model X."
[0925] This will create a system where users can not only get quick and accurate answers to their questions, but also receive responses that are tailored to their emotions.
[0926] (Example 2)
[0927] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0928] Conventional technical question answering systems have a problem of degrading the user experience because they cannot respond in a way that takes into account the user's emotional state. Furthermore, there are limitations in processing speed and accuracy for reliably acquiring information from voice input and providing appropriate answers. As a result, while users desire accurate and rapid answers, as well as emotionally responsive interaction, it has been difficult to effectively provide these.
[0929] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0930] In this invention, the server includes means for acquiring user voice input, means for converting the voice input into text data, means for analyzing the intent of the text data, means for recognizing emotions based on the acquired voice data, means for acquiring relevant information from an information organizing device based on the analysis results, means for converting the generated response into a natural language response, means for adjusting the tone and content of the response based on the emotion recognition results, means for converting the generated response into voice data, and means for providing the voice data to the user. This enables the user to obtain quick and accurate technical answers and to experience more human-like interactions through emotionally responsive interactions.
[0931] "Means for obtaining user voice input" refers to a device or part of a device used to collect voice data, and includes, for example, a microphone.
[0932] "Means for converting voice input into text data" refers to a process or device that converts voice data into a string of characters, and a speech recognition engine is used as an example.
[0933] "Means for analyzing the intent of text data" refers to processes or devices for understanding the content of text data and identifying its purpose or requirements, and includes, for example, natural language processing engines.
[0934] "Means of recognizing emotions based on acquired audio data" refers to a process or device that analyzes the tone and pitch of speech to determine the speaker's emotional state, and includes, for example, an emotion recognition engine.
[0935] "Means for obtaining relevant information from an information organization device based on analysis results" refers to a process or device for searching for and obtaining necessary information based on analysis results, and includes database queries as an example.
[0936] "Means for converting generated responses into natural language responses" refers to a process or device that converts acquired information into a natural language format that humans can understand, and generative AI models are used as an example.
[0937] "Means for adjusting the tone and content of responses based on the results of emotion recognition" refers to a process or device that appropriately adjusts the generated response content and expression method while taking into account the user's emotional state.
[0938] "Means for converting generated responses into audio data" refers to a process or device that converts text responses back into audio format, and includes, for example, a speech synthesis engine.
[0939] "Means of providing audio data to the user" refers to a process or device for making the generated audio data listen to the user, and a speaker is used as an example.
[0940] Modes for carrying out the invention
[0941] The present invention is a system that answers technical questions quickly and accurately based on the user's voice input, and further provides a more human-like interaction by recognizing the user's emotions and adjusting the response accordingly. Specific embodiments of this system are described below.
[0942] First, the user asks the robot a question using their voice. For example, they might say, "What are the latest smartphone models?" This voice data is captured through the microphone built into the robot.
[0943] The acquired voice data is sent from the robot to the server. The server uses a speech recognition engine (for example, Google Speech-to-Text API) to convert the voice data into text data. As a result, the voice data is converted into text data that says, "Please tell me the latest smartphone model."
[0944] Next, this text data is passed to a natural language processing (NLP) engine on the server (for example, Google Cloud Natural Language API) where the intent is analyzed. In this case, the NLP engine analyzes that the user is looking for information about the latest smartphones.
[0945] Simultaneously, the acquired audio data is analyzed by an emotion recognition engine on the server (for example, IBM Watson Tone Analyzer). Based on the tone and pitch of the voice, the emotion recognition engine determines whether the user is excited, angry, happy, or otherwise experiencing other emotions.
[0946] The server retrieves relevant information from the database based on the analysis results of the NLP engine. The database contains technical information and connection methods for a wide variety of devices. In this case, let's assume we retrieve the information that "the latest smartphone is the Model X."
[0947] Based on the acquired information, the server uses a generative artificial intelligence model (e.g., OpenAI GPT-3) to generate a natural language response. This model takes into account the emotional information obtained from the emotion recognition engine and adjusts the tone and content of the response. As a result, it generates a response such as, "Sorry to keep you waiting! The latest smartphone is the Model X."
[0948] Next, the generated text response is passed to the server's speech synthesis engine (e.g., Amazon Polly) and converted into audio data. This audio data is then delivered to the user through the robot's speaker. The robot then says, "Sorry to keep you waiting! The latest smartphone is the Model X," in a tone appropriate to the user's mood.
[0949] In this way, the system of this invention can provide a better user experience by offering quick and accurate answers to users' technical questions, as well as responses that take into account the user's feelings.
[0950] Examples of prompt statements
[0951] "Tell me about the latest iPhone models."
[0952] "What should I do if I can't connect to Wi-Fi?"
[0953] "How do you perceive your current emotions?"
[0954] As a result, when providing information to users, it becomes possible to accurately consider the user's emotional state and achieve natural dialogue.
[0955] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0956] Step 1: The user asks the robot a question using voice.
[0957] The user asks, "Please tell me the latest smartphone model." The robot's built-in microphone then captures the voice data.
[0958] Input: User's voice
[0959] Output: Acquired audio data
[0960] Step 2: The device sends the acquired audio data to the server.
[0961] The robot transmits the acquired audio data to the server using wireless communication (e.g., Wi-Fi).
[0962] Input: Acquired audio data
[0963] Output: Audio data sent to the server
[0964] Step 3: The server uses a speech recognition engine to convert the speech data into text data.
[0965] The server calls a speech recognition engine (e.g., Google Speech-to-Text API), takes the voice data "Please tell me the latest smartphone model" as input, and outputs text data.
[0966] Input: Audio data sent to the server
[0967] Output: Converted text data
[0968] Step 4: The server uses a natural language processing (NLP) engine to analyze the intent of the text data.
[0969] The converted text data is input into an NLP engine (e.g., Google Cloud Natural Language API) to analyze that the user is seeking information about the latest smartphones. The output is the analysis result.
[0970] Input: Converted text data
[0971] Output: Analysis results
[0972] Step 5: The server recognizes emotions based on the acquired audio data.
[0973] The server inputs voice data into an emotion recognition engine (e.g., IBM Watson Tone Analyzer) and outputs the emotion recognition result. For example, it might produce a result indicating that the user is excited.
[0974] Input: Audio data sent to the server
[0975] Output: Emotion recognition result
[0976] Step 6: The server retrieves relevant information from the information organization device based on the analysis results.
[0977] The server executes a query against the database, asking "What is the latest smartphone?" and retrieving relevant information. For example, it might retrieve information such as "The latest smartphone is the Model X."
[0978] Input: Analysis results
[0979] Output: Related information obtained
[0980] Step 7: The server generates a natural language response based on the information it has obtained.
[0981] The server inputs the prompt "What is the latest smartphone?" into a generated AI model (e.g., OpenAI GPT-3) and generates the answer "The latest smartphone is the Model X." Taking sentiment into account, it generates the answer in an appropriate tone, such as "Sorry to keep you waiting! The latest smartphone is the Model X."
[0982] Input: Acquired related information, emotion recognition results
[0983] Output: Generated text answer
[0984] Step 8: The server converts the generated text response into speech data using a speech synthesis engine.
[0985] The generated text responses are passed to a speech synthesis engine (e.g., Amazon Polly) and converted into audio data.
[0986] Input: Generated text answer
[0987] Output: Audio data
[0988] Step 9: The device provides the generated audio data to the user.
[0989] The robot's speaker uses the converted voice data to say, "Sorry to keep you waiting! The latest smartphone is the Model X."
[0990] Input: Audio data
[0991] Output: Voice response by robot
[0992] (Application Example 2)
[0993] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0994] In modern brick-and-mortar stores, there is a demand for quick and accurate responses to customer questions and inquiries. However, conventional methods often fail to provide flexible responses that respond to customer emotions, which can lead to decreased customer satisfaction. This invention aims to solve the problem of improving customer satisfaction by not only answering customers' technical questions but also recognizing their emotions and adjusting the content and tone of responses accordingly, thereby providing a more humane interaction.
[0995] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring user voice input, means for converting voice input into text data, means for analyzing the intent of the text data, means for acquiring relevant information from a database based on the analysis results, means for generating a natural language response from the acquired information, means for converting the generated response into voice data, means for providing the voice data to the user, and means for recognizing emotions from the user's voice and adjusting the response. This makes it possible to answer customer questions quickly and accurately in a physical store while providing appropriate responses that correspond to the customer's emotions.
[0996] "Means for acquiring voice input" refers to devices or interfaces that allow users to input questions or instructions to a system using their voice.
[0997] "Means for converting voice input into text data" refers to software or algorithms used to convert acquired voice data into text information.
[0998] "Means for analyzing the intent of text data" refers to natural language processing techniques used to understand the content of user requests and questions from textual information.
[0999] "Means for obtaining relevant information from a database based on analysis results" refers to a mechanism for searching and retrieving relevant answers and data from a database based on the analyzed intent.
[1000] "Means for generating natural language responses from acquired information" refers to technologies such as generative artificial intelligence models that generate responses in a human-understandable format based on acquired data.
[1001] "Means for converting generated responses into audio data" refers to speech synthesis technology for converting text-based responses into audio format.
[1002] "Means of providing audio data to the user" refers to speakers or other output devices that allow the user to listen to the generated audio data.
[1003] "Means for recognizing emotions from a user's voice and adjusting responses" refers to algorithms or engines that detect emotions from the user's voice tone and pitch, and adjust responses to an appropriate tone and content.
[1004] An example of an embodiment of the present invention is a customer service robot system for use in physical stores. Specifically, this system allows customers to ask questions to the robot by voice, and the robot provides quick and accurate answers according to the content of the questions, while also recognizing the customer's emotions and adjusting the content and tone of the response.
[1005] System Configuration
[1006] This system includes the following components.
[1007] Get user input
[1008] The user asks the robot a question using voice. For example, "Where can I find this product?" This voice data is captured through a microphone built into the robot. Common microphones can be used as hardware (e.g., Behringer ECM8000, Shure MV88, etc.).
[1009] Speech recognition
[1010] The acquired audio data is sent to a cloud service (e.g., Amazon Transcribe, Google Speech-to-Text) and converted into text data by a speech recognition engine. In this step, the audio data is converted into the text "Where can I find this product?".
[1011] Natural Language Processing
[1012] The converted text data is analyzed by a natural language processing (NLP) engine on the server (e.g., Google Cloud Natural Language API, Microsoft Azure Text Analytics API, etc.). At this stage, the intent of the text is understood to be "checking product location."
[1013] emotion recognition
[1014] Simultaneously, the acquired audio data is analyzed by an emotion engine (e.g., IBM Watson Tone Analyzer, Amazon Comprehend). Based on the voice tone and pitch, the emotion engine determines, for example, that the customer is "irritated."
[1015] Database query
[1016] Based on the analysis results, the server queries a database (e.g., MySQL, Firebase, etc.) to retrieve the necessary information. For example, it might retrieve information such as, "Product A is located on the central shelf of the third street."
[1017] Answer generation using generative artificial intelligence models
[1018] Based on the acquired information, a generative artificial intelligence model (e.g., OpenAI GPT-4, Hugging Face Transformers, etc.) generates a natural language response. Taking into account the results of emotion recognition, a response such as "Excuse me, I apologize for the inconvenience. Item A is on the central shelf of the third lane" is generated.
[1019] Speech synthesis
[1020] The generated text responses are converted into speech data by a speech synthesis engine (e.g., Google Text-to-Speech, Amazon Polly, etc.).
[1021] Response to the user
[1022] The generated voice data is delivered to the user through the robot's speaker (e.g., JBL Charge 4, Sonos Move, etc.). The robot says, "Excuse me, I'm sorry to trouble you. Item A is on the central shelf on the third street."
[1023] Specific example
[1024] User Questions
[1025] The user asks the robot, "Where can I find this product?"
[1026] Server-side processing
[1027] The server receives the voice data and uses a speech recognition engine to convert it into text data. The converted text will say, "Where is this product?" The NLP engine analyzes this and recognizes that the user is requesting the location information of a product. At the same time, the emotion engine determines the emotion from the user's voice. For example, it may determine that the user is irritated. Based on the analysis results, the server queries the database and obtains the information that "Product A is on the central shelf of the third street." Based on the obtained information, a generative artificial intelligence model generates a natural language response: "Excuse me, we apologize for the inconvenience. Product A is on the central shelf of the third street."
[1028] Robot response
[1029] Finally, the generated audio data is delivered to the user through the robot's speaker, and the robot says, "Excuse me, I'm sorry to trouble you. Item A is on the central shelf on the third street."
[1030] Example of a prompt
[1031] User question: "Where can I find this product?"
[1032] NLP analysis results: The product's location information is being sought.
[1033] Sentiment analysis result: The customer is irritated.
[1034] Database query result: "Product A is located on the central shelf of the third aisle."
[1035] Output from the generated AI model: "Excuse me, we apologize for the inconvenience. Product A is located on the central shelf of the third aisle."
[1036] In this way, the system of the present invention can improve the customer experience by providing quick and accurate answers to customers' technical questions in physical stores and by responding in a way that takes customer emotions into consideration.
[1037] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1038] Step 1:
[1039] The user asks the robot questions using voice. For example, a voice input such as "Where can I find this product?" is acquired through the robot's microphone (e.g., Behringer ECM8000, Shure MV88, etc.). The input is the user's voice data, and the output is that voice data itself.
[1040] Step 2:
[1041] The audio data is sent to a server and converted into text data by a speech recognition engine (e.g., Amazon Transcribe, Google Speech-to-Text). In this step, the audio data is converted into the text "Where can I find this product?". The input is the audio data, and the output is the text data that is a conversion of that audio data.
[1042] Step 3:
[1043] The server analyzes the converted text data using a natural language processing (NLP) engine (e.g., Google Cloud Natural Language API, Microsoft Azure Text Analytics API, etc.) to understand the user's intent. In this step, the intent "check product location" is analyzed. The input is text data, and the output is the analyzed intent information.
[1044] Step 4:
[1045] Simultaneously, the voice data is analyzed by an emotion engine (e.g., IBM Watson Tone Analyzer, Amazon Comprehend) to recognize the customer's emotions. For example, it might determine that the customer is "irritated." The input is voice data, and the output is information that expresses emotion.
[1046] Step 5:
[1047] Based on the analysis results of the NLP engine, the server queries a database (e.g., MySQL, Firebase, etc.) to retrieve relevant information. In this step, the information "Product A is on the central shelf of the third street" is retrieved. The input is the analyzed intent information, and the output is the relevant information retrieved from the database.
[1048] Step 6:
[1049] The server uses a generative artificial intelligence model (e.g., OpenAI GPT-4, Hugging Face Transformers, etc.) to generate a natural language response based on the acquired information. Sentiment recognition results are also considered during this process. For example, a response like, "Excuse me, I apologize for the inconvenience. Product A is on the central shelf of the third street," might be generated. The input consists of relevant and sentiment information, and the output is the generated natural language response.
[1050] Step 7:
[1051] The generated text responses are converted into audio data by a text-to-speech engine (e.g., Google Text-to-Speech, Amazon Polly, etc.). In this step, the text-formatted responses are converted into audio format. The input is the text-formatted responses, and the output is audio data.
[1052] Step 8:
[1053] Finally, the generated audio data is delivered to the user through the robot's speaker (e.g., JBL Charge 4, Sonos Move, etc.). The robot says, "Excuse me, I'm sorry to trouble you. Item A is on the central shelf on the third street." The input is audio data, and the output is the audio the user hears.
[1054] In this way, the system of the present invention can provide quick and accurate answers to users' technical questions, as well as respond in a way that takes the user's feelings into consideration. This series of processes is expected to improve the customer experience in physical stores.
[1055] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1056] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1057] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1058] [Fourth Embodiment]
[1059] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1060] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1061] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1062] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1063] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1064] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1065] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1066] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1067] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1068] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1069] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1070] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1071] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1072] This invention is a system for quickly and accurately answering technical questions, such as mobile phone model descriptions and internet connection methods, based on user voice input. Specific embodiments of this system are described below.
[1073] This system performs a series of processes: it acquires the user's voice input, converts it into text data, analyzes the intent of the text, retrieves relevant information from a database, generates a natural language response, converts it back into voice data, and provides it to the user.
[1074] System Configuration
[1075] Get user input
[1076] The user asks the robot a question using voice. This voice data is acquired through the microphone built into the robot.
[1077] Speech recognition
[1078] The acquired audio data is sent to a server and converted into text data by a speech recognition engine. For example, a general-purpose speech recognition engine can be used.
[1079] Natural Language Processing
[1080] The converted text data is analyzed for intent by a natural language processing (NLP) engine on the server. The NLP engine understands the user's inquiry and performs analysis to process it appropriately.
[1081] Database query
[1082] Based on the analysis results, the server queries the database to retrieve the necessary information. The database stores information such as mobile phone model and internet connection details.
[1083] Answer generation using generative artificial intelligence models
[1084] The acquired information is generated into natural language responses by the server's generative artificial intelligence model. The generative AI model generates text that provides the information in a way that is easy for the user to understand.
[1085] Speech synthesis
[1086] The generated text responses are converted into audio data by the server's speech synthesis engine. For example, a general-purpose speech synthesis engine can be used.
[1087] Response to the user
[1088] The generated audio data is provided to the user through the robot's speaker.
[1089] Specific example
[1090] User Questions
[1091] For example, a user might ask the robot, "What are the latest smartphone models?"
[1092] Server-side processing
[1093] The server receives the voice data and uses a speech recognition engine to convert it into text data. The converted text will say, "Please tell me the latest smartphone model." The NLP engine analyzes this content and recognizes that the request is for information about the latest smartphones.
[1094] Based on these analysis results, the server queries the database to retrieve information about the latest smartphones. For example, it might retrieve data indicating that "the latest smartphone is the Model X."
[1095] Based on the acquired information, a generative artificial intelligence model generates a natural language response: "The latest smartphone is the Model X." This text response is then converted into speech data by a speech synthesis engine.
[1096] Robot response
[1097] Finally, the generated audio data is delivered to the user through the robot's speaker, and the robot says, "The latest smartphone is the Model X."
[1098] In this way, the system of the present invention can provide users with quick and accurate answers to their technical questions.
[1099] The following describes the processing flow.
[1100] Step 1:
[1101] The user inputs a question to the robot using voice. For example, they might say, "Please tell me the latest smartphone model."
[1102] Step 2:
[1103] The robot acquires the user's voice input and sends the voice data to the server.
[1104] Step 3:
[1105] The server passes the received audio data to the speech recognition engine, which converts the audio into text data. For example, the text data "Please tell me the latest smartphone model" is generated.
[1106] Step 4:
[1107] The server passes the converted text data to a natural language processing (NLP) engine to analyze the text's intent. This engine recognizes that the user is seeking information about the latest smartphones.
[1108] Step 5:
[1109] The server uses the analysis results obtained from the NLP engine to query the database and retrieve relevant information. For example, the database might return information such as "The latest smartphone is the Model X."
[1110] Step 6:
[1111] The server passes the acquired information to a generative artificial intelligence model to generate a response in natural language form. The generative AI generates the text response "The latest smartphone is the Model X."
[1112] Step 7:
[1113] The server passes the generated text response to the speech synthesis engine, which converts it into speech data. For example, it might generate speech data saying, "The latest smartphone is the Model X."
[1114] Step 8:
[1115] The server sends the generated audio data to the robot. The robot plays the audio data to the user through its speaker. The user hears the voice response, "The latest smartphone is the Model X."
[1116] This will enable a system where users can obtain immediate and accurate answers to their questions.
[1117] (Example 1)
[1118] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1119] Conventional systems for answering technical questions had limited ability to provide quick and accurate answers based on user voice input. Furthermore, they struggled to properly analyze voice input, understand its intent, and provide relevant information, and lacked sufficient means to efficiently search for desired information from large databases. As a result, it was difficult for users to quickly and accurately access the information they needed.
[1120] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1121] In this invention, the server includes means for acquiring user voice input, means for converting the voice input into text data, means for analyzing the intent of the text data, means for generating a natural language response from the acquired information, means for converting the generated response into voice data, and means for providing the voice data to the user. This makes it possible to provide quick and accurate answers to the user's technical questions.
[1122] "User voice input" refers to questions or instructions that a user makes to the system using their voice.
[1123] "Means for acquiring voice input" refers to devices or technologies for capturing the user's voice, such as a microphone.
[1124] "Means of converting voice input into text data" refers to technologies for converting acquired speech into text information, such as speech recognition engines.
[1125] "Means for analyzing the intent of text data" refers to technologies for understanding the content of converted text data and analyzing the user's intent, such as natural language processing engines.
[1126] "Means of obtaining relevant information from a database based on analysis results" refers to technologies and systems that execute queries on a database according to the analyzed intent and obtain the necessary information.
[1127] "Means for generating natural language responses from acquired information" refers to technologies that create human-readable text based on information obtained from a database, such as generative artificial intelligence models.
[1128] "Means of converting generated responses into audio data" refers to technologies that convert natural language text into speech, such as speech synthesis engines.
[1129] "Means of providing audio data to the user" refers to devices and technologies for transmitting generated audio to the user, such as speakers.
[1130] This invention is a system for quickly and accurately answering technical questions based on user voice input. The system of this invention performs a series of processes, including acquiring user voice input, converting it into text data, analyzing the intent of the text, retrieving relevant information from a database, generating a natural language response, and converting it back into voice data to provide to the user. Specific hardware and software are used for each processing step.
[1131] First, the user asks a question to the robot using voice. This voice data is acquired through a microphone built into the robot (for example, an omnidirectional microphone). The acquired voice data is then transmitted to a server via the internet. This process utilizes a built-in microphone and a network module.
[1132] Next, the server converts the received audio data into text data using a speech recognition engine (for example, a general-purpose speech recognition engine). The speech recognition engine analyzes the waveform data of the audio and converts it into corresponding character information. This process converts the audio into text data.
[1133] The converted text data is analyzed for intent by a natural language processing (NLP) engine on the server (for example, a general-purpose NLP engine). The NLP engine analyzes keywords and context within the text to understand the user's inquiry. This analysis enables queries to be made to the appropriate database.
[1134] The server then queries the database based on the analysis results to retrieve the necessary information. This database may contain information such as mobile phone model details or internet connection information. The server creates an SQL query to retrieve the relevant information from the database.
[1135] Based on the acquired information, the server's generative artificial intelligence model (for example, a general generative AI model) generates a natural language response. The generative AI model creates natural-sounding sentences and provides the response in a way that is easy for the user to understand.
[1136] Next, the generated text responses are converted into audio data by the server's speech synthesis engine (for example, a general-purpose speech synthesis engine). The speech synthesis engine converts the text data into audio waveform data and generates a natural-sounding human voice.
[1137] Ultimately, the audio data is delivered to the user through the robot's speaker (for example, a standard speaker). The generated audio data is played back through the speaker and provided in a format that the user can hear.
[1138] As a concrete example, consider a scenario where a user asks a robot, "What is the latest smartphone model?" The user's voice data is acquired through a microphone and sent to a server. The server uses a speech recognition engine to convert the voice data into text data, and an NLP engine analyzes the intent of the text data. Based on the analysis results, the server queries a database to obtain the latest smartphone information. A generative artificial intelligence model generates a natural language response, "The latest smartphone is the Model X," which is then converted into voice data by a speech synthesis engine. Finally, the voice message "The latest smartphone is the Model X" is delivered to the user from the robot's speaker. In this way, the system of the present invention can provide quick and accurate answers to users' technical questions.
[1139] To realize the invention, it is crucial to properly configure and coordinate these hardware and software components.
[1140] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1141] Step 1:
[1142] The user asks the robot a question using voice. This question becomes the system's input. For example, the user might say, "Please tell me the latest smartphone model." The user's voice input is captured by the device's microphone.
[1143] Step 2:
[1144] The terminal (robot) sends the acquired voice data to the server. Specifically, it sends the acquired voice data to the server via the internet. At this time, the input voice data is transmitted directly to the server, and the server receives the voice data.
[1145] Step 3:
[1146] The server converts the received audio data into text data using a general speech recognition engine. Specifically, it analyzes the audio data and outputs text data as character information. For example, the audio data "Please tell me the latest smartphone model" is converted into the corresponding string.
[1147] Step 4:
[1148] The server uses a natural language processing (NLP) engine to analyze the intent of the converted text data. This process analyzes the context and keywords of the text and outputs the intent of the user's question. For example, from the text data "Please tell me the latest smartphone models," the server recognizes that the user is seeking information about the latest smartphones.
[1149] Step 5:
[1150] The server queries the database based on the analysis results to retrieve the necessary information. In this step, the analysis results are used as input to generate an SQL query, which is then sent to the database to output information about the latest smartphones. For example, the server retrieves the information that "the latest smartphone is the Model X" from the database.
[1151] Step 6:
[1152] The server uses a generative artificial intelligence model (for example, a general generative AI model) to generate natural language responses based on the acquired information. It uses the acquired information as input to generate text in a user-friendly format. For example, it might output the text response, "The latest smartphone is the Model X."
[1153] Step 7:
[1154] The server converts the generated text responses into speech data using a common speech synthesis engine. It uses the input text data to output speech waveform data. For example, the text "The latest smartphone is the Model X" is converted into natural-sounding speech.
[1155] Step 8:
[1156] The terminal (robot) provides the user with audio data sent from the server via its speaker. By receiving the audio data output from the server and playing it back through the speaker, it provides the user with an audio response such as, "The latest smartphone is the Model X."
[1157] (Application Example 1)
[1158] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1159] In traditional virtual stores, users lacked sufficient means to obtain appropriate and timely answers when asking questions about specific products or popular items. Furthermore, answers were often provided only in text format, limiting the user experience. Additionally, traditional systems struggled to efficiently analyze user voice input, accurately understand their intent, and generate responses.
[1160] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1161] In this invention, the server includes means for acquiring user voice input, means for converting the voice input into text data, means for analyzing the intent of the text data, means for acquiring relevant information from a database based on the analysis results, means for generating a natural language response from the acquired information, means for converting the generated response into voice data, means for providing the voice data to the user, and means for responding to questions regarding product information in the virtual store. This makes it possible to provide quick and accurate voice responses within the virtual store to questions asked by the user by voice.
[1162] "Means for acquiring user voice input" refers to a device or method used to acquire the voice spoken by the user.
[1163] "Means for converting voice input into text data" refers to a device or method for converting acquired voice into text information.
[1164] "Means for analyzing the intent of text data" refers to a device or method for understanding the user's intent or the content of a question from text.
[1165] "Means for obtaining relevant information from a database based on analysis results" refers to a device or method for searching and obtaining necessary data from a database according to the analyzed information.
[1166] "Means for generating natural language responses from acquired information" refers to a device or method for generating responses in natural language that humans can understand, based on acquired data.
[1167] "Means for converting generated responses into audio data" refers to a device or method for converting text-based responses into audio format.
[1168] "Means for providing audio data to a user" refers to a device or method for delivering generated audio data to a user.
[1169] "Means for responding to product information questions in a virtual store" refers to a device or method for users to ask questions about products within a virtual environment and for providing answers to those questions.
[1170] The system of this invention is designed to allow users in a virtual store to ask questions about product information using voice, and to provide quick and accurate answers to those questions. Specific embodiments of this system are described below.
[1171] 1. System Overview
[1172] Users ask voice questions within a virtual store via smartphone, smart glasses, or head-mounted display. The system receives the voice input, processes it on a server, and then provides a voice response. The main components are as follows:
[1173] User devices: Smartphones, smart glasses, head-mounted displays, etc.
[1174] server:
[1175] Speech recognition engine: Uses the Google Cloud Speech-to-Text API.
[1176] Natural Language Processing (NLP) engine: We use Amazon Comprehend from AWS.
[1177] Database: AWS RDS (Relational Database Service) is used.
[1178] Generative AI model: OpenAI GPT-3 is used.
[1179] Text-to-speech engine: Uses Google Cloud Text-to-Speech.
[1180] 2. Processing Flow
[1181] 1. Acquisition of voice input:
[1182] The user terminal acquires voice input and sends it to the server. The terminal is equipped with a built-in microphone.
[1183] 2. Speech recognition:
[1184] The server converts the received audio data into text data using the Google Cloud Speech-to-Text API.
[1185] 3. Natural Language Processing:
[1186] The converted text data is analyzed using Amazon Comprehend on AWS to understand the user's intent.
[1187] 4. Database query:
[1188] Based on the analysis results, queries are sent to AWS RDS to retrieve relevant information.
[1189] 5. Generating answers using generative AI models:
[1190] Based on the acquired information, OpenAI GPT-3 is used to generate a natural language response.
[1191] 6. Speech synthesis:
[1192] The generated text responses are converted into audio data using Google Cloud Text-to-Speech.
[1193] 7. Responding to the user:
[1194] The generated audio data is sent to the user's device, and the response is played back through the built-in speaker.
[1195] 3. Specific examples
[1196] When a user asks a question via voice, such as "What is the most popular item in this store?", the system's server performs the following actions.
[1197] The speech recognition engine converts the speech into text data that reads, "What is the most popular item in this store?"
[1198] A natural language processing engine analyzes this text and understands that the user is asking about popular products.
[1199] The database retrieves the information that "The most popular product is 'Product A'."
[1200] The generative AI model generates a natural language response that says, "The most popular product is 'Product A'."
[1201] The speech synthesis engine converts the generated text response into audio data.
[1202] Finally, the user's device plays an audio message saying, "The most popular item is 'Product A'."
[1203] Example of a prompt
[1204] When a user asks a question, it will look like this:
[1205] User voice input: What is the most popular item in this store?
[1206] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1207] Step 1:
[1208] The user terminal acquires voice input. For example, if the user says, "What is the most popular item in this store?", the terminal's microphone picks up the voice and sends it to the server as audio data. The input is the user's voice data, and the output is the audio data sent to the server.
[1209] Step 2:
[1210] The server uses a speech recognition engine to convert the audio data into text data. The Google Cloud Speech-to-Text API is used to transcribe the audio data. The input is the user's voice data, and the output is the text data "What is the most popular item in this store?".
[1211] Step 3:
[1212] The server uses a natural language processing engine to analyze the intent of text data. Using Amazon Comprehend from AWS, it understands that the text data is a question about popular products. The input is text data, and the output is intent data indicating "I am requesting information about popular products."
[1213] Step 4:
[1214] The server sends a query to the database based on the data analysis results and retrieves relevant information. It executes a query to AWS RDS to "get popular products" and retrieves the data "The popular product is 'Product A'." The input is intent data, and the output is informational data "The popular product is 'Product A'."
[1215] Step 5:
[1216] The server uses a generation AI model to generate natural language responses based on the acquired information. Using OpenAI GPT-3, it creates a natural-sounding sentence such as "The most popular product is 'Product A'." The input is informational data, and the output is the natural language response "The most popular product is 'Product A'."
[1217] Step 6:
[1218] The server converts natural language responses generated using a speech synthesis engine into audio data. The Google Cloud Text-to-Speech API is used to convert text data into audio data. The input is text data, and the output is audio data.
[1219] Step 7:
[1220] The server sends the generated audio data to the user's terminal, which then outputs it through the terminal's speaker. The user's terminal's built-in speaker plays the audio saying, "The most popular product is 'Product A'." The input is audio data, and the output is the audio response provided to the user.
[1221] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1222] The present invention is a system for quickly and accurately answering technical questions such as mobile phone model descriptions and internet connection methods based on the user's voice input, and furthermore, it is a system that provides more human-like interaction by recognizing the user's emotions and adjusting the response accordingly. Specific embodiments of this system are described below.
[1223] This system performs a series of processes: it acquires the user's voice input, converts it into text data, analyzes the intent of the text, retrieves relevant information from a database, generates a natural language response, converts it back into voice data, and provides it to the user. It also recognizes the user's emotions from their voice and adjusts the tone and content of the response accordingly.
[1224] System Configuration
[1225] Get user input
[1226] The user asks the robot a question using voice. This voice data is acquired through the microphone built into the robot.
[1227] Speech recognition
[1228] The acquired audio data is sent to a server and converted into text data by a speech recognition engine. For example, a general-purpose speech recognition engine can be used.
[1229] Natural Language Processing
[1230] The converted text data is analyzed for intent by a natural language processing (NLP) engine on the server. The NLP engine understands the user's inquiry and performs analysis to process it appropriately.
[1231] emotion recognition
[1232] Simultaneously, the acquired audio data is analyzed by an emotion engine. The emotion engine recognizes emotions based on the tone and pitch of the user's voice. For example, it determines emotions such as anger, sadness, and happiness.
[1233] Database query
[1234] Based on the analysis results, the server queries the database to retrieve the necessary information. The database stores information such as mobile phone model and internet connection details.
[1235] Answer generation using generative artificial intelligence models
[1236] The acquired information is used by the server's generative artificial intelligence model to generate a natural language response. The generative AI model considers emotional information from the emotion engine and adjusts the tone and content of the response.
[1237] Speech synthesis
[1238] The generated text responses are converted into audio data by the server's speech synthesis engine. For example, a general-purpose speech synthesis engine can be used.
[1239] Response to the user
[1240] The generated voice data is delivered to the user through the robot's speaker. The robot responds in a tone that matches the user's emotions.
[1241] Specific example
[1242] User Questions
[1243] The user asks the robot, "Please tell me the latest smartphone model."
[1244] Server-side processing
[1245] The server receives the voice data and uses a speech recognition engine to convert it into text data. The converted text will say, "Please tell me the latest smartphone model." The NLP engine analyzes this content and recognizes that the request is for information about the latest smartphones.
[1246] At the same time, the emotion engine determines the user's emotions from their voice. For example, it might determine that the user is excited.
[1247] Based on the analysis results, the server queries the database and retrieves the information that "the latest smartphone is the Model X."
[1248] Based on the acquired information, a generative AI model generates a natural language response such as "The latest smartphone is the Model X." Taking into account the information from the emotion engine, the generative AI model generates a response in an appropriate tone to the user's emotions, such as "Sorry to keep you waiting! The latest smartphone is the Model X."
[1249] Robot response
[1250] Finally, the generated audio data is delivered to the user through the robot's speaker, and the robot says, "Sorry to keep you waiting! The latest smartphone is the Model X."
[1251] In this way, the system of the present invention can improve the user experience by providing quick and accurate answers to the user's technical questions, as well as by providing responses that take the user's feelings into consideration.
[1252] The following describes the processing flow.
[1253] Step 1:
[1254] The user inputs a question to the robot using voice. For example, they might say, "Please tell me the latest smartphone model."
[1255] Step 2:
[1256] The robot acquires the user's voice input and sends the voice data to the server.
[1257] Step 3:
[1258] The server passes the received audio data to the speech recognition engine, which converts the audio into text data. For example, the text data "Please tell me the latest smartphone model" is generated.
[1259] Step 4:
[1260] The server passes the converted text data to a natural language processing (NLP) engine to analyze the text's intent. This engine recognizes that the user is seeking information about the latest smartphones.
[1261] Step 5:
[1262] The server simultaneously passes the voice data to the emotion engine to analyze the user's emotions. The emotion engine analyzes the tone, pitch, and speed of the voice to determine the user's emotional state, such as being excited, angry, or sad.
[1263] Step 6:
[1264] Based on the analysis results of the NLP engine, the server queries the database to retrieve relevant information. For example, the database might return information such as "The latest smartphone is the Model X."
[1265] Step 7:
[1266] The server passes the acquired information to a generative artificial intelligence model to generate a response in natural language form. This model also takes emotional information from the emotion engine into consideration, and generates responses such as, "Sorry to keep you waiting! The latest smartphone is the Model X."
[1267] Step 8:
[1268] The server passes the generated text response to the speech synthesis engine, which converts it into speech data. For example, it might generate speech data such as, "Sorry to keep you waiting! The latest smartphone is the Model X."
[1269] Step 9:
[1270] The server sends the generated audio data to the robot. The robot plays the audio data to the user through its speaker. The user hears the voice response, "Thank you for waiting! The latest smartphone is the Model X."
[1271] This will create a system where users can not only get quick and accurate answers to their questions, but also receive responses that are tailored to their emotions.
[1272] (Example 2)
[1273] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1274] Conventional technical question answering systems have a problem of degrading the user experience because they cannot respond in a way that takes into account the user's emotional state. Furthermore, there are limitations in processing speed and accuracy for reliably acquiring information from voice input and providing appropriate answers. As a result, while users desire accurate and rapid answers, as well as emotionally responsive interaction, it has been difficult to effectively provide these.
[1275] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1276] In this invention, the server includes means for acquiring user voice input, means for converting the voice input into text data, means for analyzing the intent of the text data, means for recognizing emotions based on the acquired voice data, means for acquiring relevant information from an information organizing device based on the analysis results, means for converting the generated response into a natural language response, means for adjusting the tone and content of the response based on the emotion recognition results, means for converting the generated response into voice data, and means for providing the voice data to the user. This enables the user to obtain quick and accurate technical answers and to experience more human-like interactions through emotionally responsive interactions.
[1277] "Means for obtaining user voice input" refers to a device or part of a device used to collect voice data, and includes, for example, a microphone.
[1278] "Means for converting voice input into text data" refers to a process or device that converts voice data into a string of characters, and a speech recognition engine is used as an example.
[1279] "Means for analyzing the intent of text data" refers to processes or devices for understanding the content of text data and identifying its purpose or requirements, and includes, for example, natural language processing engines.
[1280] "Means of recognizing emotions based on acquired audio data" refers to a process or device that analyzes the tone and pitch of speech to determine the speaker's emotional state, and includes, for example, an emotion recognition engine.
[1281] "Means for obtaining relevant information from an information organization device based on analysis results" refers to a process or device for searching for and obtaining necessary information based on analysis results, and includes database queries as an example.
[1282] "Means for converting generated responses into natural language responses" refers to a process or device that converts acquired information into a natural language format that humans can understand, and generative AI models are used as an example.
[1283] "Means for adjusting the tone and content of responses based on the results of emotion recognition" refers to a process or device that appropriately adjusts the generated response content and expression method while taking into account the user's emotional state.
[1284] "Means for converting generated responses into audio data" refers to a process or device that converts text responses back into audio format, and includes, for example, a speech synthesis engine.
[1285] "Means of providing audio data to the user" refers to a process or device for making the generated audio data listen to the user, and a speaker is used as an example.
[1286] Modes for carrying out the invention
[1287] The present invention is a system that answers technical questions quickly and accurately based on the user's voice input, and further provides a more human-like interaction by recognizing the user's emotions and adjusting the response accordingly. Specific embodiments of this system are described below.
[1288] First, the user asks the robot a question using their voice. For example, they might say, "What are the latest smartphone models?" This voice data is captured through the microphone built into the robot.
[1289] The acquired voice data is sent from the robot to the server. The server uses a speech recognition engine (for example, Google Speech-to-Text API) to convert the voice data into text data. As a result, the voice data is converted into text data that says, "Please tell me the latest smartphone model."
[1290] Next, this text data is passed to a natural language processing (NLP) engine on the server (for example, Google Cloud Natural Language API) where the intent is analyzed. In this case, the NLP engine analyzes that the user is looking for information about the latest smartphones.
[1291] Simultaneously, the acquired audio data is analyzed by an emotion recognition engine on the server (for example, IBM Watson Tone Analyzer). Based on the tone and pitch of the voice, the emotion recognition engine determines whether the user is excited, angry, happy, or otherwise experiencing other emotions.
[1292] The server retrieves relevant information from the database based on the analysis results of the NLP engine. The database contains technical information and connection methods for a wide variety of devices. In this case, let's assume we retrieve the information that "the latest smartphone is the Model X."
[1293] Based on the acquired information, the server uses a generative artificial intelligence model (e.g., OpenAI GPT-3) to generate a natural language response. This model takes into account the emotional information obtained from the emotion recognition engine and adjusts the tone and content of the response. As a result, it generates a response such as, "Sorry to keep you waiting! The latest smartphone is the Model X."
[1294] Next, the generated text response is passed to the server's speech synthesis engine (e.g., Amazon Polly) and converted into audio data. This audio data is then delivered to the user through the robot's speaker. The robot then says, "Sorry to keep you waiting! The latest smartphone is the Model X," in a tone appropriate to the user's mood.
[1295] In this way, the system of this invention can provide a better user experience by offering quick and accurate answers to users' technical questions, as well as responses that take into account the user's feelings.
[1296] Examples of prompt statements
[1297] "Tell me about the latest iPhone models."
[1298] "What should I do if I can't connect to Wi-Fi?"
[1299] "How do you perceive your current emotions?"
[1300] As a result, when providing information to users, it becomes possible to accurately consider the user's emotional state and achieve natural dialogue.
[1301] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1302] Step 1: The user asks the robot a question using voice.
[1303] The user asks, "Please tell me the latest smartphone model." The robot's built-in microphone then captures the voice data.
[1304] Input: User's voice
[1305] Output: Acquired audio data
[1306] Step 2: The device sends the acquired audio data to the server.
[1307] The robot transmits the acquired audio data to the server using wireless communication (e.g., Wi-Fi).
[1308] Input: Acquired audio data
[1309] Output: Audio data sent to the server
[1310] Step 3: The server uses a speech recognition engine to convert the speech data into text data.
[1311] The server calls a speech recognition engine (e.g., Google Speech-to-Text API), takes the voice data "Please tell me the latest smartphone model" as input, and outputs text data.
[1312] Input: Audio data sent to the server
[1313] Output: Converted text data
[1314] Step 4: The server uses a natural language processing (NLP) engine to analyze the intent of the text data.
[1315] The converted text data is input into an NLP engine (e.g., Google Cloud Natural Language API) to analyze that the user is seeking information about the latest smartphones. The output is the analysis result.
[1316] Input: Converted text data
[1317] Output: Analysis results
[1318] Step 5: The server recognizes emotions based on the acquired audio data.
[1319] The server inputs voice data into an emotion recognition engine (e.g., IBM Watson Tone Analyzer) and outputs the emotion recognition result. For example, it might produce a result indicating that the user is excited.
[1320] Input: Audio data sent to the server
[1321] Output: Emotion recognition result
[1322] Step 6: The server retrieves relevant information from the information organization device based on the analysis results.
[1323] The server executes a query against the database, asking "What is the latest smartphone?" and retrieving relevant information. For example, it might retrieve information such as "The latest smartphone is the Model X."
[1324] Input: Analysis results
[1325] Output: Related information obtained
[1326] Step 7: The server generates a natural language response based on the information it has obtained.
[1327] The server inputs the prompt "What is the latest smartphone?" into a generated AI model (e.g., OpenAI GPT-3) and generates the answer "The latest smartphone is the Model X." Taking sentiment into account, it generates the answer in an appropriate tone, such as "Sorry to keep you waiting! The latest smartphone is the Model X."
[1328] Input: Acquired related information, emotion recognition results
[1329] Output: Generated text answer
[1330] Step 8: The server converts the generated text response into speech data using a speech synthesis engine.
[1331] The generated text responses are passed to a speech synthesis engine (e.g., Amazon Polly) and converted into audio data.
[1332] Input: Generated text answer
[1333] Output: Audio data
[1334] Step 9: The device provides the generated audio data to the user.
[1335] The robot's speaker uses the converted voice data to say, "Sorry to keep you waiting! The latest smartphone is the Model X."
[1336] Input: Audio data
[1337] Output: Voice response by robot
[1338] (Application Example 2)
[1339] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1340] In modern brick-and-mortar stores, there is a demand for quick and accurate responses to customer questions and inquiries. However, conventional methods often fail to provide flexible responses that respond to customer emotions, which can lead to decreased customer satisfaction. This invention aims to solve the problem of improving customer satisfaction by not only answering customers' technical questions but also recognizing their emotions and adjusting the content and tone of responses accordingly, thereby providing a more humane interaction.
[1341] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring user voice input, means for converting voice input into text data, means for analyzing the intent of the text data, means for acquiring relevant information from a database based on the analysis results, means for generating a natural language response from the acquired information, means for converting the generated response into voice data, means for providing the voice data to the user, and means for recognizing emotions from the user's voice and adjusting the response. This makes it possible to answer customer questions quickly and accurately in a physical store while providing appropriate responses that correspond to the customer's emotions.
[1342] "Means for acquiring voice input" refers to devices or interfaces that allow users to input questions or instructions to a system using their voice.
[1343] "Means for converting voice input into text data" refers to software or algorithms used to convert acquired voice data into text information.
[1344] "Means for analyzing the intent of text data" refers to natural language processing techniques used to understand the content of user requests and questions from textual information.
[1345] "Means for obtaining relevant information from a database based on analysis results" refers to a mechanism for searching and retrieving relevant answers and data from a database based on the analyzed intent.
[1346] "Means for generating natural language responses from acquired information" refers to technologies such as generative artificial intelligence models that generate responses in a human-understandable format based on acquired data.
[1347] "Means for converting generated responses into audio data" refers to speech synthesis technology for converting text-based responses into audio format.
[1348] "Means of providing audio data to the user" refers to speakers or other output devices that allow the user to listen to the generated audio data.
[1349] "Means for recognizing emotions from a user's voice and adjusting responses" refers to algorithms or engines that detect emotions from the user's voice tone and pitch, and adjust responses to an appropriate tone and content.
[1350] An example of an embodiment of the present invention is a customer service robot system for use in physical stores. Specifically, this system allows customers to ask questions to the robot by voice, and the robot provides quick and accurate answers according to the content of the questions, while also recognizing the customer's emotions and adjusting the content and tone of the response.
[1351] System Configuration
[1352] This system includes the following components.
[1353] Get user input
[1354] The user asks the robot a question using voice. For example, "Where can I find this product?" This voice data is captured through a microphone built into the robot. Common microphones can be used as hardware (e.g., Behringer ECM8000, Shure MV88, etc.).
[1355] Speech recognition
[1356] The acquired audio data is sent to a cloud service (e.g., Amazon Transcribe, Google Speech-to-Text) and converted into text data by a speech recognition engine. In this step, the audio data is converted into the text "Where can I find this product?".
[1357] Natural Language Processing
[1358] The converted text data is analyzed by a natural language processing (NLP) engine on the server (e.g., Google Cloud Natural Language API, Microsoft Azure Text Analytics API, etc.). At this stage, the intent of the text is understood to be "checking product location."
[1359] emotion recognition
[1360] Simultaneously, the acquired audio data is analyzed by an emotion engine (e.g., IBM Watson Tone Analyzer, Amazon Comprehend). Based on the voice tone and pitch, the emotion engine determines, for example, that the customer is "irritated."
[1361] Database query
[1362] Based on the analysis results, the server queries a database (e.g., MySQL, Firebase, etc.) to retrieve the necessary information. For example, it might retrieve information such as, "Product A is located on the central shelf of the third street."
[1363] Answer generation using generative artificial intelligence models
[1364] Based on the acquired information, a generative artificial intelligence model (e.g., OpenAI GPT-4, Hugging Face Transformers, etc.) generates a natural language response. Taking into account the results of emotion recognition, a response such as "Excuse me, I apologize for the inconvenience. Item A is on the central shelf of the third lane" is generated.
[1365] Speech synthesis
[1366] The generated text responses are converted into speech data by a speech synthesis engine (e.g., Google Text-to-Speech, Amazon Polly, etc.).
[1367] Response to the user
[1368] The generated voice data is delivered to the user through the robot's speaker (e.g., JBL Charge 4, Sonos Move, etc.). The robot says, "Excuse me, I'm sorry to trouble you. Item A is on the central shelf on the third street."
[1369] Specific example
[1370] User Questions
[1371] The user asks the robot, "Where can I find this product?"
[1372] Server-side processing
[1373] The server receives the voice data and uses a speech recognition engine to convert it into text data. The converted text will say, "Where is this product?" The NLP engine analyzes this and recognizes that the user is requesting the location information of a product. At the same time, the emotion engine determines the emotion from the user's voice. For example, it may determine that the user is irritated. Based on the analysis results, the server queries the database and obtains the information that "Product A is on the central shelf of the third street." Based on the obtained information, a generative artificial intelligence model generates a natural language response: "Excuse me, we apologize for the inconvenience. Product A is on the central shelf of the third street."
[1374] Robot response
[1375] Finally, the generated audio data is delivered to the user through the robot's speaker, and the robot says, "Excuse me, I'm sorry to trouble you. Item A is on the central shelf on the third street."
[1376] Example of a prompt
[1377] User question: "Where can I find this product?"
[1378] NLP analysis results: The product's location information is being sought.
[1379] Sentiment analysis result: The customer is irritated.
[1380] Database query result: "Product A is located on the central shelf of the third aisle."
[1381] Output from the generated AI model: "Excuse me, we apologize for the inconvenience. Product A is located on the central shelf of the third aisle."
[1382] In this way, the system of the present invention can improve the customer experience by providing quick and accurate answers to customers' technical questions in physical stores and by responding in a way that takes customer emotions into consideration.
[1383] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1384] Step 1:
[1385] The user asks the robot questions using voice. For example, a voice input such as "Where can I find this product?" is acquired through the robot's microphone (e.g., Behringer ECM8000, Shure MV88, etc.). The input is the user's voice data, and the output is that voice data itself.
[1386] Step 2:
[1387] The audio data is sent to a server and converted into text data by a speech recognition engine (e.g., Amazon Transcribe, Google Speech-to-Text). In this step, the audio data is converted into the text "Where can I find this product?". The input is the audio data, and the output is the text data that is a conversion of that audio data.
[1388] Step 3:
[1389] The server analyzes the converted text data using a natural language processing (NLP) engine (e.g., Google Cloud Natural Language API, Microsoft Azure Text Analytics API, etc.) to understand the user's intent. In this step, the intent "check product location" is analyzed. The input is text data, and the output is the analyzed intent information.
[1390] Step 4:
[1391] Simultaneously, the voice data is analyzed by an emotion engine (e.g., IBM Watson Tone Analyzer, Amazon Comprehend) to recognize the customer's emotions. For example, it might determine that the customer is "irritated." The input is voice data, and the output is information that expresses emotion.
[1392] Step 5:
[1393] Based on the analysis results of the NLP engine, the server queries a database (e.g., MySQL, Firebase, etc.) to retrieve relevant information. In this step, the information "Product A is on the central shelf of the third street" is retrieved. The input is the analyzed intent information, and the output is the relevant information retrieved from the database.
[1394] Step 6:
[1395] The server uses a generative artificial intelligence model (e.g., OpenAI GPT-4, Hugging Face Transformers, etc.) to generate a natural language response based on the acquired information. Sentiment recognition results are also considered during this process. For example, a response like, "Excuse me, I apologize for the inconvenience. Product A is on the central shelf of the third street," might be generated. The input consists of relevant and sentiment information, and the output is the generated natural language response.
[1396] Step 7:
[1397] The generated text responses are converted into audio data by a text-to-speech engine (e.g., Google Text-to-Speech, Amazon Polly, etc.). In this step, the text-formatted responses are converted into audio format. The input is the text-formatted responses, and the output is audio data.
[1398] Step 8:
[1399] Finally, the generated audio data is delivered to the user through the robot's speaker (e.g., JBL Charge 4, Sonos Move, etc.). The robot says, "Excuse me, I'm sorry to trouble you. Item A is on the central shelf on the third street." The input is audio data, and the output is the audio the user hears.
[1400] In this way, the system of the present invention can provide quick and accurate answers to users' technical questions, as well as respond in a way that takes the user's feelings into consideration. This series of processes is expected to improve the customer experience in physical stores.
[1401] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1402] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1403] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1404] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1405] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1406] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1407] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1408] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1409] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1410] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1411] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1412] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1413] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1414] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1415] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1416] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1417] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1418] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1419] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1420] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[1421] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[1422] The following is further disclosed regarding the embodiments described above.
[1423] (Claim 1)
[1424] A means of obtaining user voice input,
[1425] A means of converting voice input into text data,
[1426] A means of analyzing the intent of text data,
[1427] A means of obtaining relevant information from a database based on the analysis results,
[1428] A means of generating a natural language response from acquired information,
[1429] A means of converting the generated response into audio data,
[1430] Means of providing audio data to users,
[1431] A system that includes this.
[1432] (Claim 2)
[1433] The system according to claim 1, wherein the means for generating a natural language response based on the acquired information is a generative artificial intelligence model.
[1434] (Claim 3)
[1435] The system according to claim 1, wherein the means for converting the voice input into text data is a voice recognition engine.
[1436] "Example 1"
[1437] (Claim 1)
[1438] A means of obtaining user voice input,
[1439] A means of converting voice input into text data,
[1440] A means of analyzing the intent of text data,
[1441] A means of obtaining relevant information from a database based on the analysis results,
[1442] A means of generating a natural language response from acquired information,
[1443] A means of converting the generated response into audio data,
[1444] Means of providing audio data to users,
[1445] A system that includes this.
[1446] (Claim 2)
[1447] The system according to claim 1, wherein the means for generating a natural language response based on acquired information is a generative artificial intelligence model.
[1448] (Claim 3)
[1449] The system according to claim 1, wherein the means for converting voice input into text data is a voice recognition engine.
[1450] "Application Example 1"
[1451] (Claim 1)
[1452] A means of obtaining user voice input,
[1453] A means of converting voice input into text data,
[1454] A means of analyzing the intent of text data,
[1455] A means of obtaining relevant information from a database based on the analysis results,
[1456] A means of generating a natural language response from acquired information,
[1457] A means of converting the generated response into audio data,
[1458] Means of providing audio data to users,
[1459] A means of responding to questions regarding product information in a virtual store,
[1460] A system that includes this.
[1461] (Claim 2)
[1462] The system according to claim 1, wherein the means for generating a natural language response based on the acquired information is a generative model.
[1463] (Claim 3)
[1464] The system according to claim 1, wherein the means for converting the voice input into text data is a voice recognition engine.
[1465] "Example 2 of combining an emotion engine"
[1466] (Claim 1)
[1467] A means of obtaining user voice input,
[1468] A means of converting voice input into text data,
[1469] A means of analyzing the intent of text data,
[1470] A means of recognizing emotions based on acquired audio data,
[1471] A means for obtaining relevant information from an information organization device based on the analysis results,
[1472] A means of generating a natural language response from acquired information,
[1473] A means of adjusting the tone and content of responses based on the results of emotion recognition,
[1474] A means of converting the generated response into audio data,
[1475] Means of providing audio data to users,
[1476] A system that includes this.
[1477] (Claim 2)
[1478] The system according to claim 1, wherein the means for generating a natural language response based on acquired information is a generative artificial intelligence model.
[1479] (Claim 3)
[1480] The system according to claim 1, wherein the means for converting voice input into text data is an automatic speech recognition engine.
[1481] "Application example 2 when combining with an emotional engine"
[1482] (Claim 1)
[1483] A means of obtaining user voice input,
[1484] A means of converting voice input into text data,
[1485] A means of analyzing the intent of text data,
[1486] A means of obtaining relevant information from a database based on the analysis results,
[1487] A means of generating a natural language response from acquired information,
[1488] A means of converting the generated response into audio data,
[1489] Means of providing audio data to users,
[1490] A means of recognizing emotions from the user's voice and adjusting the response,
[1491] A system that includes this.
[1492] (Claim 2)
[1493] The system according to claim 1, wherein the means for generating a natural language response based on the acquired information is a generative artificial intelligence model.
[1494] (Claim 3)
[1495] The system according to claim 1, wherein the means for converting the voice input into text data is a voice recognition engine. [Explanation of symbols]
[1496] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of obtaining user voice input, A means of converting voice input into text data, A means of analyzing the intent of text data, A means of obtaining relevant information from a database based on the analysis results, A means of generating a natural language response from acquired information, A means of converting the generated response into audio data, Means of providing audio data to users, A system that includes this.
2. The system according to claim 1, wherein the means for generating a natural language response based on the acquired information is a generative artificial intelligence model.
3. The system according to claim 1, wherein the means for converting the voice input into text data is a voice recognition engine.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A