system
The system addresses the challenge of flexible responses in call centers by converting user voice to text, analyzing questions, selecting appropriate responders, and saving interaction history, enhancing efficiency and user satisfaction.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-01
- Publication Date
- 2026-04-13
AI Technical Summary
Conventional call center systems struggle to provide flexible responses to diverse user inquiries, leading to increased staff burden and inconsistent answers, particularly when handling repetitive questions or requests for specific voice or explanation styles.
A system that includes means for recognizing incoming calls, converting user voice to text, analyzing question content, selecting an appropriate responder, generating answer data, converting it to speech, and saving interaction history, enabling efficient and flexible responses.
The system allows for quick and flexible responses to diverse user requests, reducing staff burden and improving operational efficiency by utilizing AI to handle repetitive inquiries and tailoring responses to user preferences.
Smart Images

Figure 2026063767000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including: receiving a user utterance; adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character; encoding the prompt; and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In a conventional call center system, it is difficult to make a flexible response according to various requirements of inquiries, and there is also a problem that the response to users who frequently receive the same question becomes a burden on the staff. In particular, it has been a problem that it is impossible to appropriately respond to requests for a specific voice or a simple explanation, and the fatigue of the staff when dealing with frequent incoming calls including harassment.
Means for Solving the Problems
[0005] The present invention solves the above-mentioned problems by providing a system that includes means for recognizing incoming calls from users, means for converting user voice into text data, means for analyzing the content of user questions from the text data, means for selecting a responder based on user requests, means for generating answer data for the questions, means for converting the generated answer data into voice, means for providing the voiced answers to the user, and means for saving the response history. This system makes it possible to flexibly respond to diverse user requests and efficiently answer even frequently asked questions, thereby reducing the burden on staff.
[0006] "Means for recognizing incoming calls" refers to a function that detects incoming calls from users and initiates a response accordingly.
[0007] "Means of converting speech to text data" refers to a function that converts spoken audio from a user into text format in real time or in batch processing.
[0008] "Means for analyzing question content" refers to a function that understands the user's intended question from the converted text data and processes it as structured data.
[0009] "Means for selecting a responder" refers to a function that selects the most suitable AI responder or voice assistant based on the user's requests.
[0010] "Means for generating response data" refers to functions that generate appropriate answers from artificial intelligence or databases in response to analyzed questions.
[0011] "Means for converting response data into speech" refers to a function that converts text-based response data into a speech format that can be provided to the user using speech synthesis technology.
[0012] "Means of providing users with voiced responses" refers to a function that transmits generated voice data to users via telephone.
[0013] "Means for saving interaction history" refers to a function for saving the content, date, and time of interactions with users as a database or log. [Brief explanation of the drawing]
[0014] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when combined with an emotion engine.
Embodiments for Carrying Out the Invention
[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0018] In the following embodiments, a labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0019] In the following embodiments, a labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0022] [First Embodiment]
[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0035] The system of the present invention includes various means for providing natural and flexible responses when an end user calls a call center. The system configuration is as follows:
[0036] 1. Recognition of incoming call
[0037] When a user calls the call center, the server automatically recognizes the call and initiates the conversation. The technologies used here are general telephone systems and VoIP (Voice over Internet Protocol) technology.
[0038] 2. Collection and conversion of audio data
[0039] When a user begins speaking, the server collects their voice in real time. This voice data is converted into text using an ASR (Automatic Speech Recognition) engine. For example, if a user asks, "Please tell me the opening hours of the supermarket," the voice data is converted into the text "Please tell me the opening hours of the supermarket."
[0040] 3. Analysis of the inquiry content
[0041] The server analyzes the converted text data to understand the user's question. The technology used here is NLP (Natural Language Processing), which classifies the question into specific categories. For example, categories such as "business hours" or "return policy."
[0042] 4. Selection of a support person based on user requirements
[0043] When a user makes a specific request (for example, "I want an explanation in a female voice" or "I want an explanation that even a primary school student can understand"), the server selects an appropriate AI respondent based on that request. A specific voice profile is used for voice selection.
[0044] 5. Generating response data
[0045] Based on the question, the server retrieves relevant information from the database and constructs answer data using generative AI. For example, if the question is about opening hours, the answer "The supermarket's opening hours are from 9 am to 8 pm" will be generated.
[0046] 6. Speech synthesis
[0047] The generated response data is converted into speech format using a Text-to-Speech (TTS) engine. Natural-sounding speech is generated based on the selected voice profile.
[0048] 7. Providing a response
[0049] The generated audio data is played back to the user in real time during the call. This allows the user to receive responses in a natural conversational style.
[0050] 8. Record of interaction history
[0051] The server stores the interaction history in a database. This includes the user's question, the answer provided, and the date and time of the interaction. This information is used to respond quickly and accurately to future inquiries.
[0052] Specific example
[0053] Example 1: Request for a brief explanation in a female voice.
[0054] User A makes a phone call and asks, "Could you tell me the supermarket's opening hours? Could you please explain it in a way that even a primary school student could understand, using a woman's voice?" The server recognizes this call, converts the user's voice into text, analyzes the question, and detects a "woman's voice" as the request. The server retrieves the opening hours information from its database, and a generative AI generates the answer, "The supermarket's opening hours are from 9 am to 8 pm," which is then converted into a woman's voice and provided to User A.
[0055] Example 2: Harassment prevention measures
[0056] If a regular user B asks the same question hundreds of times every month, the server recognizes the call and converts the user's voice into text. Because the past question history is stored in the database, the server can refer to the already accumulated information and quickly and accurately generate and provide the answer, "Our business hours are from 9 am to 8 pm." This significantly reduces the burden on staff.
[0057] By following these steps, the present invention provides a system that can flexibly respond to diverse user demands and improve the efficiency of call centers.
[0058] The following describes the processing flow.
[0059] Step 1:
[0060] The user makes a phone call. The user dials a phone number and is connected to the call center.
[0061] Step 2:
[0062] The server recognizes the incoming call. The server detects the incoming signal and automatically initiates the call.
[0063] Step 3:
[0064] The server collects audio data. When the user starts speaking, the server collects audio data in real time.
[0065] Step 4:
[0066] The server converts the audio data into text data. The collected audio data is converted into text format using the ASR engine.
[0067] Step 5:
[0068] The server analyzes the text data. Using NLP (Neuro-Linguistic Programming) techniques, it analyzes the user's questions from the text data and identifies categories.
[0069] Step 6:
[0070] The server detects user requests. If there are user requests (e.g., "female voice," "make it easy enough for elementary school children to understand"), the server analyzes and recognizes them.
[0071] Step 7:
[0072] The server selects a responder. Based on the detected request, it selects the appropriate AI responder and voice profile.
[0073] Step 8:
[0074] The server generates the answer data. The generative AI retrieves information related to the question from the database and constructs the answer.
[0075] Step 9:
[0076] The server converts the response data into speech. A TTS engine is used to convert the text-based responses into speech using a selected voice profile.
[0077] Step 10:
[0078] The server provides the user with an audio response. The generated audio data is played back to the user in real time.
[0079] Step 11:
[0080] The server saves the interaction history. It records information such as the user's question, the answer provided, and the date and time of the interaction in a database.
[0081] (Example 1)
[0082] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0083] Traditional call center systems struggled to respond to user inquiries in real time, requiring a significant amount of manual work. This resulted in low efficiency in handling inquiries, increased susceptibility to human error, and inconsistent response quality. Furthermore, flexibly addressing specific requests (e.g., responses in a specific gender's voice or tailored to a specific level of understanding) was difficult, potentially leading to decreased user satisfaction. Additionally, the inability to effectively utilize past interaction history meant that responses to the same question were inconsistent, hindering efficient operation.
[0084] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0085] In this invention, the server includes means for recognizing incoming calls from users, means for converting user voice into text data, means for analyzing the user's questions from the text data, means for selecting a responder based on the user's request, means for generating answer data to the questions using natural language generation technology, means for converting the generated answer data into speech, means for providing the spoken answers to the user, and means for storing the interaction history. This makes it possible to respond quickly and flexibly to diverse user requests and improve the operational efficiency of the call center.
[0086] (Definitions of important words)
[0087] "Means for recognizing incoming calls" refers to a function that detects when a call has started when receiving a call from a user and initiates the appropriate processing.
[0088] "Methods for converting speech to text data" refers to technologies that collect spoken audio from users in real time and convert it into text data. Specifically, this utilizes ASR (Automatic Speech Recognition) technology.
[0089] "Means for analyzing question content" refers to a function that analyzes the converted text data to identify and classify the content of the user's question. It primarily uses NLP (Natural Language Processing) technology.
[0090] The "means of selecting a responder" refer to a function that chooses the optimal voice response profile based on the user's requests. This includes selecting a voice profile tailored to a specific gender or level of comprehension.
[0091] "Means for generating response data" refers to a function that collects relevant information based on the user's question and creates an appropriate response using natural language generation technology.
[0092] "Means for converting response data into speech" refers to a function that converts the generated text-based response data into natural-sounding speech using TTS (text-to-speech) technology.
[0093] "Means of providing users with voiced responses" refers to a function that provides generated voice data to users in real time via telephone.
[0094] "Means for saving interaction history" refers to a function that saves historical data such as the user's questions, the answers provided, and the date and time of the call in a database, for use in handling future inquiries.
[0095] "Natural language generation technology" is a technique that generates meaningful natural language sentences based on given input data. It commonly utilizes generative AI models.
[0096] A "prompt" is text data input to a generative AI model, containing specific instructions or questions for the model.
[0097] The system of this invention is designed to provide natural and flexible responses when end users call a call center. The system configuration is as follows:
[0098] When a user calls the call center, the server automatically recognizes the call and initiates the conversation. At this stage, a standard telephone system or VoIP (Voice over Internet Protocol) technology is used.
[0099] Next, when the user begins to speak, the server collects the user's voice in real time. This is done using microphones and acoustic detection devices. This voice data is converted into text format using an ASR (Automatic Speech Recognition) engine. Specific technologies include Google® Cloud Speech-to-Text and IBM Watson® Speech to Text. For example, if the user asks, "Please tell me the opening hours of the supermarket," the voice data is converted into text data that reads, "Please tell me the opening hours of the supermarket."
[0100] The converted text data is sent to a server and analyzed using an NLP (Natural Language Processing) engine. Technologies include Google Cloud Natural Language API and Microsoft® Azure® Text Analytics. The analysis categorizes user inquiries into specific categories (e.g., "Business Hours," "Return Policy").
[0101] Next, when a user makes a specific request (for example, "I want an explanation in a woman's voice" or "Please explain it in a way that even a primary school student can understand"), the server analyzes the request and selects an appropriate voice profile. The server accesses an internal database and selects a voice responder based on the optimal voice profile for the request.
[0102] The answer data for the questions is generated by retrieving relevant information from a database and using generative AI. Generative AI technologies such as OpenAI® GPT-3® or similar natural language generation technologies are used. For example, the answer data "The supermarket's operating hours are from 9 AM to 8 PM" might be generated.
[0103] The generated response data is converted into speech format using a Text-to-Speech (TTS) engine. Amazon Polly and Google Cloud Text-to-Speech are used as TTS engines. The converted speech data provides natural-sounding speech based on the selected speech profile.
[0104] Ultimately, the generated audio data is played back to the user in real time via the call, allowing the user to receive responses in a natural conversational style.
[0105] The interaction history is stored in a database by the server. This includes the user's question, the answer provided, and the date and time of the interaction. This information is used to respond quickly and accurately to future inquiries.
[0106] Specific example
[0107] Example 1: Request for a brief explanation in a female voice.
[0108] User A makes a phone call and asks, "Could you tell me the supermarket's opening hours? Could you please explain it in a way that even a primary school student could understand, using a woman's voice?" The server recognizes this call and converts the user's voice into text data. Next, it analyzes the question and detects a "woman's voice" as the request. The server retrieves the opening hours information from the database, and a generative AI generates the answer, "The supermarket's opening hours are from 9 am to 8 pm." This is then converted into a woman's voice and provided to User A.
[0109] Example 2: Harassment prevention measures
[0110] If a regular user B asks the same question hundreds of times every month, the server recognizes the incoming call and converts the user's voice into text. Because past question history is stored in the database, the server can refer to the already accumulated information and quickly generate and provide an answer such as, "Our business hours are from 9 am to 8 pm." This significantly reduces the burden on staff.
[0111] Example of a prompt
[0112] "Could you please tell me about the supermarket's opening hours? Could you explain it in a way that even a primary school student could understand?"
[0113] "Our business hours are from 9 AM to 8 PM."
[0114] "Do you have any other questions?"
[0115] In this way, the system can respond quickly and flexibly to diverse user requests, thereby improving the operational efficiency of the call center.
[0116] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0117] Step 1:
[0118] When a user calls the call center, the server recognizes the incoming call and initiates the conversation. At this stage, the call is connected to the call center's system using VoIP technology. In other words, a call commencement signal (output) is generated in response to the user's phone call (input). The server logs the user's number and the time the call started.
[0119] Step 2:
[0120] When a user begins speaking, the server collects their voice in real time. This voice data (input) is sent to an ASR engine and converted into text format (output). Specifically, voice data is collected using microphones or acoustic detection devices and sent to an ASR engine (such as Google Cloud Speech-to-Text). The converted text data is temporarily stored for subsequent processing.
[0121] Step 3:
[0122] The server sends the converted text data to the NLP engine, which analyzes the user's question. This text data (input) is analyzed by the NLP engine, and the question is classified into a specific category (output). Technologies used here include Google Cloud Natural Language API and Microsoft Azure Text Analytics. The category information is also used to generate the response in the next step.
[0123] Step 4:
[0124] The server detects specific requests (e.g., "female voice," "explain it clearly") from the converted text data. This text data (input) contains the user's requests, and an appropriate voice profile is selected based on the analysis results (output). The server records this information in its internal database.
[0125] Step 5:
[0126] The server retrieves relevant information from the database based on the analyzed question. This relevant information (output) corresponding to the question (input) is converted into answer data using a generative AI model (e.g., OpenAI GPT-3). The generated answer data is temporarily stored because it will be converted into speech in the next step.
[0127] Step 6:
[0128] The server sends the generated response data to the TTS engine, where it is converted into speech format. This text-based response data (input) is then converted into speech data (output) by the TTS engine (such as Amazon Polly or Google Cloud Text-to-Speech). The server generates natural-sounding speech based on the selected speech profile.
[0129] Step 7:
[0130] The server sends the generated audio data to the call channel and plays it back to the user in real time. This audio data (input) is provided to the user through the call channel (output). The server confirms that the response has been provided and continues or ends the call with the user as needed.
[0131] Step 8:
[0132] The server stores call history, including call content, generated responses, and call date and time, in a database. This call history data (input) is recorded in the database and used for future inquiries (output). This helps in quickly generating answers to the same question again and improves operational efficiency.
[0133] (Application Example 1)
[0134] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0135] Current voice assistant systems in autonomous vehicles lack the ability to provide natural and flexible responses in passenger interaction. In particular, the systems are inadequate in handling situations where passengers request answers tailored to specific voice qualities or levels of understanding, or when they repeatedly ask the same questions. This leads to decreased passenger satisfaction and compromises in-vehicle comfort.
[0136] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0137] In this invention, the server includes means for recognizing a user entering the vehicle, means for converting the user's voice into text data, means for analyzing the user's question from the text data, means for selecting a responder based on the user's request, means including a generative AI model for generating answer data to the question, means for converting the generated answer data into speech, means for providing the spoken answer to the user, and means for storing the response history. This makes it possible to provide natural and flexible responses in real time to various requests from passengers in an autonomous vehicle.
[0138] "Means for recognizing a user entering the vehicle" refers to a device or system used in an autonomous vehicle to recognize when a passenger has boarded the vehicle.
[0139] "Means for converting user speech into text data" refers to technology or equipment for converting speech spoken by passengers into text data in real time.
[0140] "Means for analyzing user questions from text data" refers to a system or technology for analyzing passengers' questions based on converted text data and understanding their intent.
[0141] "Means for selecting a responder based on user requests" refers to a system or technology for selecting an appropriate voice profile or responder in accordance with the content of a passenger's request.
[0142] "Means including a generative AI model that generates response data to a question" refers to a system or technology that includes an artificial intelligence model for generating appropriate answers to passengers' questions.
[0143] "Means for converting generated response data into audio" refers to a technology or device for converting generated text-formatted response data into audio data.
[0144] "Means of providing users with voiced responses" refers to a device or system that plays voice data in real time and provides responses to passengers.
[0145] "Means for saving interaction history" refers to a system or technology that records the history of interactions with passengers and uses it to handle future inquiries.
[0146] The system of the present invention realizes a voice recognition assistant in an autonomous vehicle by, for example, following the procedure described below. By utilizing specific hardware and software, it is possible to interact with the user in a natural and flexible manner.
[0147] System configuration and program processing
[0148] 1. Means for recognizing a vehicle entering the premises from a user.
[0149] The server uses hardware such as sensors and cameras to recognize when a user has boarded the vehicle. Machine learning models may also be used for this purpose.
[0150] 2. Means for converting user speech into text data
[0151] The server analyzes the audio data collected through the microphone in real time. The audio data is converted into text data using an ASR (Automatic Speech Recognition) engine such as AWS® Transcribe or Google Cloud Speech-to-Text API.
[0152] 3. Means for analyzing user questions from text data
[0153] The server analyzes the converted text data to understand the intent behind the user's question. This process utilizes NLP (Natural Language Processing) technologies such as SpaCy and Google Cloud Natural Language API.
[0154] 4. Means for selecting a responder based on user requests
[0155] The server selects a specific voice profile and response method. For example, it might use Azure's Cognitive Services Speech API to select a voice quality and speaking style that suits the user's request.
[0156] 5. Means including a generative AI model that generates response data to the content of a question.
[0157] The server uses a generative AI model (e.g., GPT-4®) to generate response data corresponding to the query. It uses the OpenAI API to generate appropriate responses based on the prompt.
[0158] 6. Means for converting generated response data into speech.
[0159] The server converts the generated text-based response data into speech using a Text-to-Speech (TTS) engine. Google TTS and Amazon Polly are among the engines used for this process.
[0160] 7. Means of providing users with voiced responses
[0161] The server provides the converted audio data to the user in real time through the in-car speaker system.
[0162] 8. Means for saving interaction history
[0163] The server stores the user interaction history in a database. Services such as Amazon RDS and Google FI®restore are used to facilitate future inquiries.
[0164] Specific example
[0165] One afternoon, passenger A gets into a self-driving vehicle and asks, "Please tell me the next gas station." In this case, the system operates as follows:
[0166] The server recognizes passenger A's voice and converts the audio data into text.
[0167] The converted text data is analyzed to understand that it is a question asking for the location of a gas station.
[0168] A gentle male voice is selected as the appropriate voice quality, and the generating AI model responds, "The next gas station is 2 kilometers ahead on the right."
[0169] This text is converted into audio data by the TTS engine and played back in real time through the car's speakers.
[0170] Finally, save this question and answer exchange to the database.
[0171] Example of a prompt
[0172] The following prompt is used in a situation where a passenger requests, "Tell me the next gas station," and the system responds, "The next gas station is 2 kilometers ahead on the right":
[0173] User: "Please tell me the next gas station."
[0174] System: "The next gas station is 2 kilometers ahead on the right."
[0175] The system of the present invention, through these procedures, enables highly natural and flexible interaction with the user in an autonomous vehicle.
[0176] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0177] Step 1:
[0178] The server recognizes when a user boards the vehicle. Specifically, it uses data from on-board sensors and cameras to detect entry. The server analyzes this input data to confirm the presence of a passenger.
[0179] Step 2:
[0180] When a user begins speaking in the vehicle, the server collects audio data through the microphone. This audio data is then input to the server.
[0181] Step 3:
[0182] The server converts the collected audio data into text data using AWS Transcribe or the Google Cloud Speech-to-Text API. The input audio data is passed through the ASR engine, and the converted data in text format is output.
[0183] Step 4:
[0184] The server analyzes the generated text data. Specifically, it uses NLP technologies such as SpaCy and Google Cloud Natural Language API to analyze the text data and identify the user's question. It processes the input text data and generates the intent and category of the question as output.
[0185] Step 5:
[0186] The server selects the appropriate voice profile and speaker based on the user's request. This is done by using the Azure Cognitive Services Speech API to select a voice profile according to the user's request (e.g., "a calm male voice"). The user's request data is used as input, and the selected voice profile is obtained as output.
[0187] Step 6:
[0188] The server generates answer data in response to the user's questions. Using the OpenAI GPT-4 model, it generates appropriate answers based on the text data input. This generated answer data is then output.
[0189] Step 7:
[0190] The server converts the generated text-based response data into audio data. It uses a TTS engine such as Google TTS or Amazon Polly to convert the text data to audio data. Text response data is used as input, and audio data is generated as output.
[0191] Step 8:
[0192] The server plays the generated audio data in real time through the car's speakers. This allows the user to receive an audio response. The audio data is sent to the speakers, and an audio response is output.
[0193] Step 9:
[0194] The server stores the user interaction history in a database. Amazon RDS or Google Firestore is used to store question content and answer data. Interaction data is received as input and saved to the database as output.
[0195] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0196] The system of the present invention includes various means for recognizing the user's emotions and providing natural and flexible responses when an end user makes a call to a call center. The system configuration is as follows:
[0197] 1. Recognition of incoming call
[0198] When a user calls the call center, the server automatically recognizes the call and initiates the conversation. The technologies used here are general telephone systems and VoIP (Voice over Internet Protocol) technology.
[0199] 2. Collection and conversion of audio data
[0200] When a user begins speaking, the server collects their voice in real time. This voice data is converted into text using an ASR (Automatic Speech Recognition) engine. For example, if a user asks, "Please tell me the opening hours of the supermarket," the voice data is converted into the text "Please tell me the opening hours of the supermarket."
[0201] 3. Recognition of emotions
[0202] The collected audio data is analyzed through an emotion engine. The emotion engine recognizes the user's emotions (e.g., joy, anger, sadness, surprise) based on factors such as tone, speed, and pitch of the voice. This process determines the user's current emotional state.
[0203] 4. Analysis of the inquiry content
[0204] The server analyzes the converted text data to understand the user's question. The technology used here is NLP (Natural Language Processing), which classifies the question into specific categories. For example, categories such as "business hours" or "return policy."
[0205] 5. Selection of a representative based on user requirements
[0206] When a user makes a specific request (for example, "I want an explanation in a woman's voice" or "I want an explanation that even a primary school student can understand"), the server selects an appropriate AI respondent based on that request. A specific voice profile is used for voice selection. Furthermore, the appropriate response is selected based on the emotional state recognized by the emotion engine.
[0207] 6. Generating response data
[0208] Based on the question, the server retrieves relevant information from the database and constructs answer data using generative AI. For example, if the question is about opening hours, the answer "The supermarket's opening hours are from 9 am to 8 pm" will be generated.
[0209] 7. Speech synthesis
[0210] The generated response data is converted into speech format using a Text-to-Speech (TTS) engine. Natural-sounding speech is generated based on the selected voice profile. The tone and speed of the speech are also adjusted based on the emotions recognized by the emotion engine.
[0211] 8. Providing a response
[0212] The generated audio data is played back to the user in real time during the call. This allows the user to receive responses in a natural conversational style. A key feature is that the emotion engine provides voice responses that are tailored to the user's emotions.
[0213] 9. Record of interaction history
[0214] The server stores interaction history in a database. This includes the user's question, the answer provided, the date and time of the interaction, and the perceived emotional state. This information is used to respond quickly and accurately to future inquiries.
[0215] Specific example
[0216] Example 1: Request for a simple explanation in a woman's voice and emotional recognition.
[0217] User A makes a phone call and asks, "Could you tell me the supermarket's opening hours? Could you please explain it in a way that even a primary school student could understand, using a woman's voice?" The server recognizes this call, converts the user's voice into text, analyzes the question, and detects a "woman's voice" as the request. Furthermore, because the user's voice tone is calm, the emotion engine recognizes a "calm" state. The server retrieves the opening hours information from the database, and the generative AI generates the answer, "The supermarket's opening hours are from 9 am to 8 pm," which is then converted into a woman's voice and delivered to User A in a calm tone.
[0218] Example 2: Harassment prevention and emotional recognition
[0219] If a regular user B asks the same question hundreds of times every month, the server recognizes the incoming call and converts the user's voice into text. Since the past question history is stored in the database, the server refers to the already accumulated information and quickly and accurately generates and provides the answer, "Our business hours are from 9 am to 8 pm." Furthermore, if user B's voice is agitated, the emotion engine recognizes a state of "anger" and provides a response in a calm tone. This significantly reduces the burden on staff while enabling responses that are considerate of the user's emotions.
[0220] By following these steps, the present invention provides a system that can flexibly respond to the diverse needs and emotions of users and improve the efficiency of call centers.
[0221] The following describes the processing flow.
[0222] Step 1:
[0223] The user makes a phone call. The user dials a phone number to the call center and is connected.
[0224] Step 2:
[0225] The server recognizes the incoming call. The server detects the incoming signal and automatically initiates the call.
[0226] Step 3:
[0227] The server collects audio data. When the user starts speaking, the server collects this audio data in real time.
[0228] Step 4:
[0229] The server converts the audio data into text data. The collected audio data is converted into text format using the ASR engine.
[0230] Step 5:
[0231] The server analyzes the text data. Using NLP (Neuro-Linguistic Programming) techniques, it analyzes the user's questions from the text data and categorizes them into specific categories.
[0232] Step 6:
[0233] The server uses an emotion engine to recognize the user's emotions. It analyzes the tone, speed, and pitch of the voice to identify the user's emotional state.
[0234] Step 7:
[0235] The server detects user requests. For example, if there is a request to "explain in a female voice," the server recognizes that request. Emotional states are also taken into consideration.
[0236] Step 8:
[0237] The server selects a responder. Based on the detected request and emotional state, it chooses the most appropriate AI responder and voice profile.
[0238] Step 9:
[0239] The server generates the answer data. The generative AI retrieves information related to the question from the database and constructs the answer.
[0240] Step 10:
[0241] The server converts the response data into speech. Using a TTS engine, the text-based responses are converted into speech using a selected voice profile. The voice tone and speed are adjusted based on the results of the emotion engine.
[0242] Step 11:
[0243] The server provides the user with an audio response. The generated audio data is played back to the user in real time.
[0244] Step 12:
[0245] The server saves the interaction history. It records the user's question, the answer provided, the date and time of the interaction, and the perceived emotional state in a database.
[0246] Step 13:
[0247] The server analyzes the interaction history and makes improvements as needed. Based on the stored history, data analysis is performed to improve the system's response quality.
[0248] (Example 2)
[0249] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0250] Traditional call center systems were unable to analyze user voices in real time and provide flexible responses tailored to their emotional state. Furthermore, selecting the appropriate responder based on specific requests and generating quick answers using past interaction history were difficult. Therefore, there is a need to improve both user satisfaction and response efficiency. Moreover, a system capable of responding to diverse user requests and emotions is required.
[0251] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0252] In this invention, the server includes means for recognizing incoming calls from users, means for collecting the user's voice, means for converting the collected user's voice into text data, means for recognizing the user's emotional state from the text data, means for analyzing the user's question content from the text data, means for selecting a responder based on the analyzed question content and the recognized emotional state, means for generating answer data for the question content, means for converting the generated answer data into voice, means for providing the voiced answer to the user, and means for storing the interaction history. This enables flexible responses according to the user's emotional state, selection of a responder based on specific requests, and rapid and accurate response generation utilizing past interaction history.
[0253] "Incoming call recognition means" refers to a device or method for automatically detecting a call from a user and initiating a call.
[0254] "Voice collection means" refers to a device or method for acquiring the voice spoken by a user during a phone call in real time.
[0255] "Character conversion means" refers to a device or method for converting acquired audio data into character data.
[0256] "Emotion recognition means" refers to a device or method for analyzing and recognizing a user's emotional state from text data or audio data.
[0257] "Question analysis means" refers to a device or method for understanding and analyzing the content of a user's question from text data.
[0258] "Respondent selection means" refers to a device or method for selecting an appropriate respondent based on the analyzed question content and the recognized emotional state.
[0259] "Answer generation means" refers to a device or method for generating appropriate answer data in response to a user's question.
[0260] "Voice conversion means" refers to a device or method for converting generated response data into voice data.
[0261] "Answer provision means" refers to a device or method for providing answer data converted into audio format to a user.
[0262] "Response history storage means" refers to a device or method for recording and storing the content of a user's questions and response history.
[0263] Modes for carrying out the invention
[0264] The system according to this invention recognizes the user's emotions when they call a call center and provides a natural and flexible response. This system uses the following hardware and software.
[0265] Hardware and software
[0266] Telephone system or VoIP technology: Used to recognize incoming calls and initiate conversations.
[0267] ASR (Automatic Speech Recognition) engine: Collects speech data and converts it into text data.
[0268] Emotion Engine: Used to recognize the user's emotional state from text and audio data.
[0269] NLP (Natural Language Processing) technology: Used to analyze user questions from text data.
[0270] Database: Used to store user questions and response history.
[0271] Generative AI: Used to generate answer data in response to questions.
[0272] TTS (Text-to-Speech) engine: Used to convert generated response data into speech.
[0273] Specific processing of the program
[0274] 1. Recognition of incoming call
[0275] When a user calls the call center, the server automatically recognizes the incoming call and initiates the conversation. This process utilizes telephone systems and VoIP technology.
[0276] 2. Collection and conversion of audio data
[0277] When a user begins speaking, the server collects their voice in real time. The collected voice data is input into the ASR engine and converted into text format. For example, if a user asks, "Please tell me the opening hours of the supermarket," the ASR engine converts that voice into the text, "Please tell me the opening hours of the supermarket."
[0278] 3. Recognition of emotions
[0279] The server inputs the converted text data into the emotion engine and analyzes elements such as the tone, speed, and pitch of the voice. Based on this analysis, the emotion engine recognizes the user's emotional state (e.g., happiness, anger, calm).
[0280] 4. Analysis of Inquiry Content
[0281] The server analyzes the text data using NLP technology. In this process, the server classifies the text data into specific categories (e.g., "business hours", "return method", etc.) and understands the specific content of the question.
[0282] 5. Selection of Respondent Based on User Requirements
[0283] When the user makes a specific request, the server selects an appropriate AI respondent based on that request. For example, for a request like "Please explain in a female voice", the server selects an AI respondent with a female voice profile and also takes the emotional state into consideration when selecting the respondent.
[0284] 6. Generation of Answer Data
[0285] The server retrieves relevant information from the database and constructs answer data using generative AI. For example, an answer like "The business hours of the supermarket are from 9 am to 8 pm" is generated. This answer is generated to appropriately respond to the content of the user's question.
[0286] 7. Voice Synthesis
[0287] The server inputs the generated answer data into the TTS engine and converts it into audio format. Based on the selected voice profile, a natural voice is generated. At this time, the tone and speed of the voice are also adjusted based on the emotional state recognized by the emotion engine.
[0288] 8. Provision of Answer
[0289] The server provides the generated audio data to the user in real time. Specifically, the audio data is played back during the call. This allows the user to receive responses in a natural conversational style.
[0290] 9. Record of interaction history
[0291] The server saves the interaction history to a database. The saved data includes the user's question, the answer provided, the date and time of the interaction, and the perceived emotional state. This information is used to respond quickly and accurately to future inquiries.
[0292] Specific example
[0293] Example 1: Request for a simple explanation in a woman's voice and emotional recognition.
[0294] User A calls and asks, "Could you tell me the supermarket's opening hours? Could you please explain it in a way that even a primary school student could understand, in a woman's voice?" The server recognizes this call, converts the user's voice to text, analyzes the question, and detects that it is requesting an explanation in a woman's voice. Furthermore, because the user's voice tone is calm, the emotion engine recognizes that the user is in a "calm" state. The server retrieves the opening hours information from the database, and the generative AI generates the answer, "The supermarket's opening hours are from 9 am to 8 pm," which is then converted into a woman's voice and delivered to User A in a calm tone.
[0295] Example of a prompt
[0296] "A user asked, 'Could you please tell me about supermarket opening hours? Could you explain it in a way that even a primary school student could understand, using a woman's voice?' The user's tone is calm. Please generate an answer and explain it calmly in a woman's voice."
[0297] Example 2: Harassment prevention and emotional recognition
[0298] Regular user B asks loudly, "What are the business hours?" The server refers to the past question history and quickly and accurately generates an answer: "The business hours are from 9 am to 8 pm." Additionally, since user B's voice is loud, the emotion engine recognizes the "angry" state and provides an answer in a calm tone.
[0299] Example of a prompt sentence
[0300] "User B asked loudly, 'What are the business hours?' It is known that the same question has been asked multiple times based on the past question history. Please generate an answer in a calm tone."
[0301] Through this invention, it is possible to provide a system that can flexibly respond to various requirements and emotions of users and improve the efficiency of the call center.
[0302] The flow of specific processing in Example 2 will be described using FIG. 13.
[0303] Specific processing flow of the program
[0304] Step 1:
[0305] Recognition of incoming call
[0306] When the user makes a call to the call center, the server automatically recognizes the incoming call.
[0307] Input: Phone call from the user (VoIP data or telephone system data). <000097...> Processing: The server receives the call start signal by means of the telephone system or VoIP technology.
[0309] Output: Start of the call and the user's call information (phone number, call start time, etc.). <...> Step 2:
[0311] Audio data collection and conversion
[0312] When a user begins speaking, the server collects their voice in real time. The collected voice data is then converted into text format by the ASR engine.
[0313] Input: User's voice data.
[0314] Processing: The server collects the audio data and the ASR engine converts the audio to text. For example, if the user says, "Please tell me the opening hours of the supermarket."
[0315] Output: Text data (e.g., "Please tell me about the supermarket's opening hours").
[0316] Step 3:
[0317] Recognition of emotions
[0318] The server inputs the converted text data into an emotion engine, which analyzes the tone, speed, pitch, and other characteristics of the speech.
[0319] Input: Text data and audio data.
[0320] Processing: The server uses an emotion engine to analyze text and audio data and recognize the user's emotions. For example, if the tone of voice is calm, it will be recognized as "calm."
[0321] Output: Emotional state data (e.g., "Calm").
[0322] Step 4:
[0323] Analysis of inquiry content
[0324] The server analyzes the text data using NLP (Neuro-Linguistic Programming) technology to understand the user's question.
[0325] Input: Text data.
[0326] Processing: The server uses NLP (Neuro-Linguistic Programming) technology to analyze text data and classify the questions into specific categories. For example, "business hours" or "return policy."
[0327] Output: Category data of the question content (e.g., "Business Hours").
[0328] Step 5:
[0329] Selecting a responder based on user requirements
[0330] When a user makes a specific request, the server selects an appropriate AI responder based on that request and the perceived emotional state.
[0331] Input: Question category data, emotional state data, user request data (e.g., "Explain in a female voice").
[0332] Processing: The server selects an appropriate AI responder with a voice profile based on the request and emotion. For example, it might select a voice profile with a "female voice".
[0333] Output: Selected AI responder data (e.g., "female voice").
[0334] Step 6:
[0335] Generating response data
[0336] The server retrieves relevant information from the database and uses generative AI to construct response data.
[0337] Input: Category data of the question content, information in the database.
[0338] Processing: The server uses generative AI to generate accurate answer data based on the question. For example, if the question is about opening hours, it will generate the answer "The supermarket's opening hours are from 9 am to 8 pm."
[0339] Output: Response data (Example: "The supermarket's operating hours are from 9 AM to 8 PM").
[0340] Step 7:
[0341] Speech synthesis
[0342] The server inputs the generated response data into the TTS engine and converts it into audio format.
[0343] Input: Response data, selected AI responder data.
[0344] Processing: The server uses a TTS engine to convert response data into audio data. During this process, the tone and speed of the voice are also adjusted based on the emotional state.
[0345] Output: Audio data.
[0346] Step 8:
[0347] Providing a response
[0348] The server provides the generated audio data to the user in real time.
[0349] Input: Audio data.
[0350] Processing: The server plays the audio data through the call.
[0351] Output: Receipt of the user's response in audio format.
[0352] Step 9:
[0353] Record of interaction history
[0354] The server saves the interaction history to a database.
[0355] Input: User's question, provided answer, date and time of interaction, perceived emotional state, call information.
[0356] Processing: The server records and stores this data in the database.
[0357] Output: Saved call history data.
[0358] (Application Example 2)
[0359] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0360] Traditional call center and virtual store systems could only provide standard responses to user inquiries, making it difficult to respond flexibly to users' emotional states or specific requests. Furthermore, they lacked sufficient functionality to efficiently utilize past inquiry history. Therefore, there were limitations in addressing situations where improving the user experience and reducing staff workload were required.
[0361] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for recognizing incoming calls from users, means for converting user voice into text data, means for analyzing the user's question content from the text data, means for selecting a responder based on the user's request, means for generating answer data for the question content, means for converting the generated answer data into voice, means for providing the voiced answer to the user, means for recognizing the user's emotions, means for adjusting the response based on the recognized emotions, and means for saving the response history. This enables quick and efficient responses by utilizing past inquiry history while flexibly responding to the user's emotions and specific requests.
[0362] "Means for recognizing incoming calls" refers to a function that automatically detects voice inquiries from users and notifies the system.
[0363] "A means of converting user voice into text data" refers to a function that converts input voice data into text format in real time.
[0364] "Methods for analyzing user inquiries from text data" refers to a function that converts speech into text and then analyzes that text data to understand the content of the user's inquiry.
[0365] "Means for selecting a responder based on user requests" refers to a function that selects a responder who meets specific criteria (e.g., gender, level of understanding) when a user requests them.
[0366] "Means for generating response data to questions" refers to a function that generates appropriate answers based on the user's inquiry.
[0367] "Means for converting generated response data into audio" refers to a function that converts generated text-based response data into audio format.
[0368] "Means of providing users with voiced responses" refers to a function that transmits voiced responses to users in real time.
[0369] "Means of recognizing user emotions" refers to a function that determines the user's emotional state through the analysis of voice data.
[0370] "Means of adjusting responses based on recognized emotions" refers to a function that adjusts the content and tone of responses to match the user's emotional state.
[0371] "Means for saving interaction history" refers to a function that records and saves past inquiries, the answers provided, the date and time of the interaction, and the perceived emotional state.
[0372] This invention is a system that provides flexible responses to user inquiries in a virtual store, tailored to the user's emotional state. The server performs response processing when a user makes an inquiry by voice using the following means.
[0373] First, when a user makes a voice inquiry within the virtual store, the device (e.g., smart glasses or a smartphone) collects the audio in real time and sends it to a server. The server recognizes the incoming call using a general speech recognition system. Technologies used include the Google Cloud Speech-to-Text API.
[0374] Next, the server converts the collected audio data into text data. This also uses the Google Cloud Speech-to-Text API. This conversion changes the audio data into text format.
[0375] Next, the server recognizes the user's emotions based on the text data. The emotion engine uses IBM Watson Tone Analyzer or Microsoft Azure's Text Analytics API. This allows it to determine the user's emotional state (e.g., joy, anger, sadness, surprise) from the audio data.
[0376] Furthermore, the server analyzes the text data and classifies the user's questions into specific categories (e.g., product details, return procedures, etc.). This analysis uses NLP technologies such as Google Cloud Natural Language API and SpaCy. Subsequently, it selects the appropriate respondent (e.g., a voice of a specific gender, a concise explanation, etc.) based on the user's request.
[0377] Generative AI models are used to generate response data. Specifically, OpenAI GPT-4 can be used. The server retrieves relevant information from the database and constructs response data using the generative AI model. For example, if a user asks, "How much stock do you have of this product?", the generative AI model will generate a response such as, "We currently have 10 units of this product in stock."
[0378] The generated response data is converted into speech format using a TTS engine such as Amazon Polly or Google Cloud Text-to-Speech. The server also adjusts the tone and speed of the speech based on the recognized emotions.
[0379] Finally, the transcribed response data is provided to the user in real time via the device. Users can enjoy a natural conversational experience through smart glasses or earphones.
[0380] The interaction history is stored in databases such as Microsoft SQL Server and MongoDB. This enables quick and efficient responses to the same question again.
[0381] As a concrete example, consider a scenario where a customer inquires, "How much stock do you have of this item?" in a virtual store. In this case, the system recognizes the user's emotions and provides a calm tone of voice if the user is "excited," and a detailed and polite response if the user is "confused."
[0382] Examples of prompt statements include the following formats:
[0383] User inquiry: "How much stock do you have of this product?"
[0384] Perceived emotions: "excitement," "confusion"
[0385] Tone of response: "Calm tone," "Detailed and polite."
[0386] Using a generative AI model: "Currently, there are 10 of this item in stock."
[0387] By combining these methods, we can provide a system that can respond quickly and efficiently while flexibly responding to user emotions and specific requests.
[0388] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0389] Step 1:
[0390] The device recognizes the user's voice input.
[0391] Input: User voice inquiry.
[0392] Specific operation: In a virtual store, a user asks a question by voice, for example, "How much stock do you have of this item?" The device (smart glasses or smartphone) uses its built-in microphone to collect the voice and sends it to the server as digital audio data.
[0393] Step 2:
[0394] The server recognizes the incoming call from the user.
[0395] Input: Digital audio data.
[0396] Specific operation: The server uses a speech recognition system to confirm incoming calls and initiate appropriate processing. Here, a system is in place to process user inquiries in real time using a general speech recognition system.
[0397] Step 3:
[0398] The server converts the user's voice into text data.
[0399] Input: Digital audio data.
[0400] Output: Text data.
[0401] Specific operation: Use the Google Cloud Speech-to-Text API to convert audio data into text. For example, the audio "How much stock do you have of this product?" will be converted into text data.
[0402] Step 4:
[0403] The server analyzes the user's question from the text data.
[0404] Input: Text data.
[0405] Output: Category of the question.
[0406] Specific operation: Using Google Cloud Natural Language API and SpaCy, the text data is analyzed and the questions are categorized into categories such as "stock availability check".
[0407] Step 5:
[0408] The server recognizes the user's emotions.
[0409] Input: Text data.
[0410] Output: Recognized emotion.
[0411] Specific operation: Using IBM Watson Tone Analyzer and Microsoft Azure's Text Analytics API, it recognizes the user's emotions (e.g., joy, anger, sadness, surprise) from the audio. For example, an emotion such as "confusion" might be extracted.
[0412] Step 6:
[0413] The server selects a responder based on the user's request.
[0414] Input: Category of the question, perceived emotion.
[0415] Output: Selected responder.
[0416] Specific operation: Based on user requests (e.g., specific gender, level of understanding), the system selects the most appropriate responder profile (e.g., female voice, calm tone).
[0417] Step 7:
[0418] The server generates answer data for the question.
[0419] Input: Question category, selected responder profile.
[0420] Output: Generated response data.
[0421] Specific operation: Using generative AI models such as OpenAI GPT-4, it generates answer data based on the question. For example, it might generate an answer such as, "Currently, there are 10 units of this product in stock."
[0422] Step 8:
[0423] The server converts the generated response data into speech.
[0424] Input: Generated response data, selected respondent profile.
[0425] Output: Audio data.
[0426] Specific operation: Use Amazon Polly or Google Cloud Text-to-Speech to convert text data into speech data and adjust the tone and speed of the speech.
[0427] Step 9:
[0428] The server provides the user with an audio response.
[0429] Input: Audio data.
[0430] Output: Voice response to the user.
[0431] Specific operation: The generated audio data is delivered to the user in real time via a device (smart glasses or earphones). The user can experience a natural conversation.
[0432] Step 10:
[0433] The server saves the interaction history.
[0434] Input: User's question, generated answer, date and time of interaction, perceived emotion.
[0435] Output: Saved call history.
[0436] Specific actions: Use Microsoft SQL Server or MongoDB to record this information in a database, enabling quick responses to future queries.
[0437] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0438] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0439] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0440] [Second Embodiment]
[0441] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0442] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0443] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0444] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0445] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0446] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0447] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0448] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0449] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0450] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0451] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0452] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0453] The system of the present invention includes various means for providing natural and flexible responses when an end user calls a call center. The system configuration is as follows:
[0454] 1. Recognition of incoming call
[0455] When a user calls the call center, the server automatically recognizes the call and initiates the conversation. The technologies used here are general telephone systems and VoIP (Voice over Internet Protocol) technology.
[0456] 2. Collection and conversion of audio data
[0457] When a user begins speaking, the server collects their voice in real time. This voice data is converted into text using an ASR (Automatic Speech Recognition) engine. For example, if a user asks, "Please tell me the opening hours of the supermarket," the voice data is converted into the text "Please tell me the opening hours of the supermarket."
[0458] 3. Analysis of the inquiry content
[0459] The server analyzes the converted text data to understand the user's question. The technology used here is NLP (Natural Language Processing), which classifies the question into specific categories. For example, categories such as "business hours" or "return policy."
[0460] 4. Selection of a support person based on user requirements
[0461] When a user makes a specific request (for example, "I want an explanation in a female voice" or "I want an explanation that even a primary school student can understand"), the server selects an appropriate AI respondent based on that request. A specific voice profile is used for voice selection.
[0462] 5. Generating response data
[0463] Based on the question, the server retrieves relevant information from the database and constructs answer data using generative AI. For example, if the question is about opening hours, the answer "The supermarket's opening hours are from 9 am to 8 pm" will be generated.
[0464] 6. Speech synthesis
[0465] The generated response data is converted into speech format using a Text-to-Speech (TTS) engine. Natural-sounding speech is generated based on the selected voice profile.
[0466] 7. Providing a response
[0467] The generated audio data is played back to the user in real time during the call. This allows the user to receive responses in a natural conversational style.
[0468] 8. Record of interaction history
[0469] The server stores the interaction history in a database. This includes the user's question, the answer provided, and the date and time of the interaction. This information is used to respond quickly and accurately to future inquiries.
[0470] Specific example
[0471] Example 1: Request for a brief explanation in a female voice.
[0472] User A makes a phone call and asks, "Could you tell me the supermarket's opening hours? Could you please explain it in a way that even a primary school student could understand, using a woman's voice?" The server recognizes this call, converts the user's voice into text, analyzes the question, and detects a "woman's voice" as the request. The server retrieves the opening hours information from its database, and a generative AI generates the answer, "The supermarket's opening hours are from 9 am to 8 pm," which is then converted into a woman's voice and provided to User A.
[0473] Example 2: Harassment prevention measures
[0474] If a regular user B asks the same question hundreds of times every month, the server recognizes the call and converts the user's voice into text. Because the past question history is stored in the database, the server can refer to the already accumulated information and quickly and accurately generate and provide the answer, "Our business hours are from 9 am to 8 pm." This significantly reduces the burden on staff.
[0475] By following these steps, the present invention provides a system that can flexibly respond to diverse user demands and improve the efficiency of call centers.
[0476] The following describes the processing flow.
[0477] Step 1:
[0478] The user makes a phone call. The user dials a phone number and is connected to the call center.
[0479] Step 2:
[0480] The server recognizes the incoming call. The server detects the incoming signal and automatically initiates the call.
[0481] Step 3:
[0482] The server collects audio data. When the user starts speaking, the server collects audio data in real time.
[0483] Step 4:
[0484] The server converts the audio data into text data. The collected audio data is converted into text format using the ASR engine.
[0485] Step 5:
[0486] The server analyzes the text data. Using NLP (Neuro-Linguistic Programming) techniques, it analyzes the user's questions from the text data and identifies categories.
[0487] Step 6:
[0488] The server detects user requests. If there are user requests (e.g., "female voice," "make it easy enough for elementary school children to understand"), the server analyzes and recognizes them.
[0489] Step 7:
[0490] The server selects a responder. Based on the detected request, it selects the appropriate AI responder and voice profile.
[0491] Step 8:
[0492] The server generates the answer data. The generative AI retrieves information related to the question from the database and constructs the answer.
[0493] Step 9:
[0494] The server converts the response data into speech. A TTS engine is used to convert the text-based responses into speech using a selected voice profile.
[0495] Step 10:
[0496] The server provides the user with an audio response. The generated audio data is played back to the user in real time.
[0497] Step 11:
[0498] The server saves the interaction history. It records information such as the user's question, the answer provided, and the date and time of the interaction in a database.
[0499] (Example 1)
[0500] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0501] Traditional call center systems struggled to respond to user inquiries in real time, requiring a significant amount of manual work. This resulted in low efficiency in handling inquiries, increased susceptibility to human error, and inconsistent response quality. Furthermore, flexibly addressing specific requests (e.g., responses in a specific gender's voice or tailored to a specific level of understanding) was difficult, potentially leading to decreased user satisfaction. Additionally, the inability to effectively utilize past interaction history meant that responses to the same question were inconsistent, hindering efficient operation.
[0502] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0503] In this invention, the server includes means for recognizing incoming calls from users, means for converting user voice into text data, means for analyzing the user's questions from the text data, means for selecting a responder based on the user's request, means for generating answer data to the questions using natural language generation technology, means for converting the generated answer data into speech, means for providing the spoken answers to the user, and means for storing the interaction history. This makes it possible to respond quickly and flexibly to diverse user requests and improve the operational efficiency of the call center.
[0504] (Definitions of important words)
[0505] "Means for recognizing incoming calls" refers to a function that detects when a call has started when receiving a call from a user and initiates the appropriate processing.
[0506] "Methods for converting speech to text data" refers to technologies that collect spoken audio from users in real time and convert it into text data. Specifically, this utilizes ASR (Automatic Speech Recognition) technology.
[0507] "Means for analyzing question content" refers to a function that analyzes the converted text data to identify and classify the content of the user's question. It primarily uses NLP (Natural Language Processing) technology.
[0508] The "means of selecting a responder" refer to a function that chooses the optimal voice response profile based on the user's requests. This includes selecting a voice profile tailored to a specific gender or level of comprehension.
[0509] "Means for generating response data" refers to a function that collects relevant information based on the user's question and creates an appropriate response using natural language generation technology.
[0510] "Means for converting response data into speech" refers to a function that converts the generated text-based response data into natural-sounding speech using TTS (text-to-speech) technology.
[0511] "Means of providing users with voiced responses" refers to a function that provides generated voice data to users in real time via telephone.
[0512] "Means for saving interaction history" refers to a function that saves historical data such as the user's questions, the answers provided, and the date and time of the call in a database, for use in handling future inquiries.
[0513] "Natural language generation technology" is a technique that generates meaningful natural language sentences based on given input data. It commonly utilizes generative AI models.
[0514] A "prompt" is text data input to a generative AI model, containing specific instructions or questions for the model.
[0515] The system of this invention is designed to provide natural and flexible responses when end users call a call center. The system configuration is as follows:
[0516] When a user calls the call center, the server automatically recognizes the call and initiates the conversation. At this stage, a standard telephone system or VoIP (Voice over Internet Protocol) technology is used.
[0517] Next, when the user begins to speak, the server collects the user's voice in real time. This is done using microphones and acoustic detection devices. This voice data is converted into text format using an ASR (Automatic Speech Recognition) engine. Specific technologies include Google Cloud Speech-to-Text and IBM Watson Speech to Text. For example, if the user asks, "Please tell me the opening hours of the supermarket," the voice data is converted into text data that reads, "Please tell me the opening hours of the supermarket."
[0518] The converted text data is sent to a server and analyzed using an NLP (Natural Language Processing) engine. Technologies include Google Cloud Natural Language API and Microsoft Azure Text Analytics. The analysis categorizes user inquiries into specific categories (e.g., "Business Hours," "Return Policy").
[0519] Next, when a user makes a specific request (for example, "I want an explanation in a woman's voice" or "Please explain it in a way that even a primary school student can understand"), the server analyzes the request and selects an appropriate voice profile. The server accesses an internal database and selects a voice responder based on the optimal voice profile for the request.
[0520] The answer data for the questions is generated by retrieving relevant information from a database and using generative AI. OpenAI GPT-3 or similar natural language generation technologies are used for this generative AI. For example, the answer data "The supermarket's operating hours are from 9 AM to 8 PM" might be generated.
[0521] The generated response data is converted into speech format using a Text-to-Speech (TTS) engine. Amazon Polly and Google Cloud Text-to-Speech are used as TTS engines. The converted speech data provides natural-sounding speech based on the selected speech profile.
[0522] Ultimately, the generated audio data is played back to the user in real time via the call, allowing the user to receive responses in a natural conversational style.
[0523] The interaction history is stored in a database by the server. This includes the user's question, the answer provided, and the date and time of the interaction. This information is used to respond quickly and accurately to future inquiries.
[0524] Specific example
[0525] Example 1: Request for a brief explanation in a female voice.
[0526] User A makes a phone call and asks, "Could you tell me the supermarket's opening hours? Could you please explain it in a way that even a primary school student could understand, using a woman's voice?" The server recognizes this call and converts the user's voice into text data. Next, it analyzes the question and detects a "woman's voice" as the request. The server retrieves the opening hours information from the database, and a generative AI generates the answer, "The supermarket's opening hours are from 9 am to 8 pm." This is then converted into a woman's voice and provided to User A.
[0527] Example 2: Harassment prevention measures
[0528] If a regular user B asks the same question hundreds of times every month, the server recognizes the incoming call and converts the user's voice into text. Because past question history is stored in the database, the server can refer to the already accumulated information and quickly generate and provide an answer such as, "Our business hours are from 9 am to 8 pm." This significantly reduces the burden on staff.
[0529] Example of a prompt
[0530] "Could you please tell me about the supermarket's opening hours? Could you explain it in a way that even a primary school student could understand?"
[0531] "Our business hours are from 9 AM to 8 PM."
[0532] "Do you have any other questions?"
[0533] In this way, the system can respond quickly and flexibly to diverse user requests, thereby improving the operational efficiency of the call center.
[0534] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0535] Step 1:
[0536] When a user calls the call center, the server recognizes the incoming call and initiates the conversation. At this stage, the call is connected to the call center's system using VoIP technology. In other words, a call commencement signal (output) is generated in response to the user's phone call (input). The server logs the user's number and the time the call started.
[0537] Step 2:
[0538] When a user begins speaking, the server collects their voice in real time. This voice data (input) is sent to an ASR engine and converted into text format (output). Specifically, voice data is collected using microphones or acoustic detection devices and sent to an ASR engine (such as Google Cloud Speech-to-Text). The converted text data is temporarily stored for subsequent processing.
[0539] Step 3:
[0540] The server sends the converted text data to the NLP engine, which analyzes the user's question. This text data (input) is analyzed by the NLP engine, and the question is classified into a specific category (output). Technologies used here include Google Cloud Natural Language API and Microsoft Azure Text Analytics. The category information is also used to generate the response in the next step.
[0541] Step 4:
[0542] The server detects specific requests (e.g., "female voice," "explain it clearly") from the converted text data. This text data (input) contains the user's requests, and an appropriate voice profile is selected based on the analysis results (output). The server records this information in its internal database.
[0543] Step 5:
[0544] The server retrieves relevant information from the database based on the analyzed question. This relevant information (output) corresponding to the question (input) is converted into answer data using a generative AI model (e.g., OpenAI GPT-3). The generated answer data is temporarily stored because it will be converted into speech in the next step.
[0545] Step 6:
[0546] The server sends the generated response data to the TTS engine, where it is converted into speech format. This text-based response data (input) is then converted into speech data (output) by the TTS engine (such as Amazon Polly or Google Cloud Text-to-Speech). The server generates natural-sounding speech based on the selected speech profile.
[0547] Step 7:
[0548] The server sends the generated audio data to the call channel and plays it back to the user in real time. This audio data (input) is provided to the user through the call channel (output). The server confirms that the response has been provided and continues or ends the call with the user as needed.
[0549] Step 8:
[0550] The server stores call history, including call content, generated responses, and call date and time, in a database. This call history data (input) is recorded in the database and used for future inquiries (output). This helps in quickly generating answers to the same question again and improves operational efficiency.
[0551] (Application Example 1)
[0552] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0553] Current voice assistant systems in autonomous vehicles lack the ability to provide natural and flexible responses in passenger interaction. In particular, the systems are inadequate in handling situations where passengers request answers tailored to specific voice qualities or levels of understanding, or when they repeatedly ask the same questions. This leads to decreased passenger satisfaction and compromises in-vehicle comfort.
[0554] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0555] In this invention, the server includes means for recognizing a user entering the vehicle, means for converting the user's voice into text data, means for analyzing the user's question from the text data, means for selecting a responder based on the user's request, means including a generative AI model for generating answer data to the question, means for converting the generated answer data into speech, means for providing the spoken answer to the user, and means for storing the response history. This makes it possible to provide natural and flexible responses in real time to various requests from passengers in an autonomous vehicle.
[0556] "Means for recognizing a user entering the vehicle" refers to a device or system used in an autonomous vehicle to recognize when a passenger has boarded the vehicle.
[0557] "Means for converting user speech into text data" refers to technology or equipment for converting speech spoken by passengers into text data in real time.
[0558] "Means for analyzing user questions from text data" refers to a system or technology for analyzing passengers' questions based on converted text data and understanding their intent.
[0559] "Means for selecting a responder based on user requests" refers to a system or technology for selecting an appropriate voice profile or responder in accordance with the content of a passenger's request.
[0560] "Means including a generative AI model that generates response data to a question" refers to a system or technology that includes an artificial intelligence model for generating appropriate answers to passengers' questions.
[0561] "Means for converting generated response data into audio" refers to a technology or device for converting generated text-formatted response data into audio data.
[0562] "Means of providing users with voiced responses" refers to a device or system that plays voice data in real time and provides responses to passengers.
[0563] "Means for saving interaction history" refers to a system or technology that records the history of interactions with passengers and uses it to handle future inquiries.
[0564] The system of the present invention realizes a voice recognition assistant in an autonomous vehicle by, for example, following the procedure described below. By utilizing specific hardware and software, it is possible to interact with the user in a natural and flexible manner.
[0565] System configuration and program processing
[0566] 1. Means for recognizing a vehicle entering the premises from a user.
[0567] The server uses hardware such as sensors and cameras to recognize when a user has boarded the vehicle. Machine learning models may also be used for this purpose.
[0568] 2. Means for converting user speech into text data
[0569] The server analyzes the audio data collected through the microphone in real time. The audio data is converted into text data using an ASR (Automatic Speech Recognition) engine such as AWS Transcribe or Google Cloud Speech-to-Text API.
[0570] 3. Means for analyzing user questions from text data
[0571] The server analyzes the converted text data to understand the intent behind the user's question. This process utilizes NLP (Natural Language Processing) technologies such as SpaCy and Google Cloud Natural Language API.
[0572] 4. Means for selecting a responder based on user requests
[0573] The server selects a specific voice profile and response method. For example, it might use Azure's Cognitive Services Speech API to select a voice quality and speaking style that suits the user's request.
[0574] 5. Means including a generative AI model that generates response data to the content of a question.
[0575] The server uses a generative AI model (e.g., GPT-4) to generate response data corresponding to the query. It uses the OpenAI API to generate appropriate responses based on the prompt.
[0576] 6. Means for converting generated response data into speech.
[0577] The server converts the generated text-based response data into speech using a Text-to-Speech (TTS) engine. Google TTS and Amazon Polly are among the engines used for this process.
[0578] 7. Means of providing users with voiced responses
[0579] The server provides the converted audio data to the user in real time through the in-car speaker system.
[0580] 8. Means for saving interaction history
[0581] The server stores the user interaction history in a database. Services such as Amazon RDS and Google Firestore are used to facilitate future inquiries.
[0582] Specific example
[0583] One afternoon, passenger A gets into a self-driving vehicle and asks, "Please tell me the next gas station." In this case, the system operates as follows:
[0584] The server recognizes passenger A's voice and converts the audio data into text.
[0585] The converted text data is analyzed to understand that it is a question asking for the location of a gas station.
[0586] A gentle male voice is selected as the appropriate voice quality, and the generating AI model responds, "The next gas station is 2 kilometers ahead on the right."
[0587] This text is converted into audio data by the TTS engine and played back in real time through the car's speakers.
[0588] Finally, save this question and answer exchange to the database.
[0589] Example of a prompt
[0590] The following prompt is used in a situation where a passenger requests, "Tell me the next gas station," and the system responds, "The next gas station is 2 kilometers ahead on the right":
[0591] User: "Please tell me the next gas station."
[0592] System: "The next gas station is 2 kilometers ahead on the right."
[0593] The system of the present invention, through these procedures, enables highly natural and flexible interaction with the user in an autonomous vehicle.
[0594] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0595] Step 1:
[0596] The server recognizes when a user boards the vehicle. Specifically, it uses data from on-board sensors and cameras to detect entry. The server analyzes this input data to confirm the presence of a passenger.
[0597] Step 2:
[0598] When a user begins speaking in the vehicle, the server collects audio data through the microphone. This audio data is then input to the server.
[0599] Step 3:
[0600] The server converts the collected audio data into text data using AWS Transcribe or the Google Cloud Speech-to-Text API. The input audio data is passed through the ASR engine, and the converted data in text format is output.
[0601] Step 4:
[0602] The server analyzes the generated text data. Specifically, it uses NLP technologies such as SpaCy and Google Cloud Natural Language API to analyze the text data and identify the user's question. It processes the input text data and generates the intent and category of the question as output.
[0603] Step 5:
[0604] The server selects the appropriate voice profile and speaker based on the user's request. This is done by using the Azure Cognitive Services Speech API to select a voice profile according to the user's request (e.g., "a calm male voice"). The user's request data is used as input, and the selected voice profile is obtained as output.
[0605] Step 6:
[0606] The server generates answer data in response to the user's questions. Using the OpenAI GPT-4 model, it generates appropriate answers based on the text data input. This generated answer data is then output.
[0607] Step 7:
[0608] The server converts the generated text-based response data into audio data. It uses a TTS engine such as Google TTS or Amazon Polly to convert the text data to audio data. Text response data is used as input, and audio data is generated as output.
[0609] Step 8:
[0610] The server plays the generated audio data in real time through the car's speakers. This allows the user to receive an audio response. The audio data is sent to the speakers, and an audio response is output.
[0611] Step 9:
[0612] The server stores the user interaction history in a database. Amazon RDS or Google Firestore is used to store question content and answer data. Interaction data is received as input and saved to the database as output.
[0613] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0614] The system of the present invention includes various means for recognizing the user's emotions and providing natural and flexible responses when an end user makes a call to a call center. The system configuration is as follows:
[0615] 1. Recognition of incoming call
[0616] When a user calls the call center, the server automatically recognizes the call and initiates the conversation. The technologies used here are general telephone systems and VoIP (Voice over Internet Protocol) technology.
[0617] 2. Collection and conversion of audio data
[0618] When a user begins speaking, the server collects their voice in real time. This voice data is converted into text using an ASR (Automatic Speech Recognition) engine. For example, if a user asks, "Please tell me the opening hours of the supermarket," the voice data is converted into the text "Please tell me the opening hours of the supermarket."
[0619] 3. Recognition of emotions
[0620] The collected audio data is analyzed through an emotion engine. The emotion engine recognizes the user's emotions (e.g., joy, anger, sadness, surprise) based on factors such as tone, speed, and pitch of the voice. This process determines the user's current emotional state.
[0621] 4. Analysis of the inquiry content
[0622] The server analyzes the converted text data to understand the user's question. The technology used here is NLP (Natural Language Processing), which classifies the question into specific categories. For example, categories such as "business hours" or "return policy."
[0623] 5. Selection of a representative based on user requirements
[0624] When a user makes a specific request (for example, "I want an explanation in a woman's voice" or "I want an explanation that even a primary school student can understand"), the server selects an appropriate AI respondent based on that request. A specific voice profile is used for voice selection. Furthermore, the appropriate response is selected based on the emotional state recognized by the emotion engine.
[0625] 6. Generating response data
[0626] Based on the question, the server retrieves relevant information from the database and constructs answer data using generative AI. For example, if the question is about opening hours, the answer "The supermarket's opening hours are from 9 am to 8 pm" will be generated.
[0627] 7. Speech synthesis
[0628] The generated response data is converted into speech format using a Text-to-Speech (TTS) engine. Natural-sounding speech is generated based on the selected voice profile. The tone and speed of the speech are also adjusted based on the emotions recognized by the emotion engine.
[0629] 8. Providing a response
[0630] The generated audio data is played back to the user in real time during the call. This allows the user to receive responses in a natural conversational style. A key feature is that the emotion engine provides voice responses that are tailored to the user's emotions.
[0631] 9. Record of interaction history
[0632] The server stores interaction history in a database. This includes the user's question, the answer provided, the date and time of the interaction, and the perceived emotional state. This information is used to respond quickly and accurately to future inquiries.
[0633] Specific example
[0634] Example 1: Request for a simple explanation in a woman's voice and emotional recognition.
[0635] User A makes a phone call and asks, "Could you tell me the supermarket's opening hours? Could you please explain it in a way that even a primary school student could understand, using a woman's voice?" The server recognizes this call, converts the user's voice into text, analyzes the question, and detects a "woman's voice" as the request. Furthermore, because the user's voice tone is calm, the emotion engine recognizes a "calm" state. The server retrieves the opening hours information from the database, and the generative AI generates the answer, "The supermarket's opening hours are from 9 am to 8 pm," which is then converted into a woman's voice and delivered to User A in a calm tone.
[0636] Example 2: Harassment prevention and emotional recognition
[0637] If a regular user B asks the same question hundreds of times every month, the server recognizes the incoming call and converts the user's voice into text. Since the past question history is stored in the database, the server refers to the already accumulated information and quickly and accurately generates and provides the answer, "Our business hours are from 9 am to 8 pm." Furthermore, if user B's voice is agitated, the emotion engine recognizes a state of "anger" and provides a response in a calm tone. This significantly reduces the burden on staff while enabling responses that are considerate of the user's emotions.
[0638] By following these steps, the present invention provides a system that can flexibly respond to the diverse needs and emotions of users and improve the efficiency of call centers.
[0639] The following describes the processing flow.
[0640] Step 1:
[0641] The user makes a phone call. The user dials a phone number to the call center and is connected.
[0642] Step 2:
[0643] The server recognizes the incoming call. The server detects the incoming signal and automatically initiates the call.
[0644] Step 3:
[0645] The server collects audio data. When the user starts speaking, the server collects this audio data in real time.
[0646] Step 4:
[0647] The server converts the audio data into text data. The collected audio data is converted into text format using the ASR engine.
[0648] Step 5:
[0649] The server analyzes the text data. Using NLP (Neuro-Linguistic Programming) techniques, it analyzes the user's questions from the text data and categorizes them into specific categories.
[0650] Step 6:
[0651] The server uses an emotion engine to recognize the user's emotions. It analyzes the tone, speed, and pitch of the voice to identify the user's emotional state.
[0652] Step 7:
[0653] The server detects user requests. For example, if there is a request to "explain in a female voice," the server recognizes that request. Emotional states are also taken into consideration.
[0654] Step 8:
[0655] The server selects a responder. Based on the detected request and emotional state, it chooses the most appropriate AI responder and voice profile.
[0656] Step 9:
[0657] The server generates the answer data. The generative AI retrieves information related to the question from the database and constructs the answer.
[0658] Step 10:
[0659] The server converts the response data into speech. Using a TTS engine, the text-based responses are converted into speech using a selected voice profile. The voice tone and speed are adjusted based on the results of the emotion engine.
[0660] Step 11:
[0661] The server provides the user with an audio response. The generated audio data is played back to the user in real time.
[0662] Step 12:
[0663] The server saves the interaction history. It records the user's question, the answer provided, the date and time of the interaction, and the perceived emotional state in a database.
[0664] Step 13:
[0665] The server analyzes the interaction history and makes improvements as needed. Based on the stored history, data analysis is performed to improve the system's response quality.
[0666] (Example 2)
[0667] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0668] Traditional call center systems were unable to analyze user voices in real time and provide flexible responses tailored to their emotional state. Furthermore, selecting the appropriate responder based on specific requests and generating quick answers using past interaction history were difficult. Therefore, there is a need to improve both user satisfaction and response efficiency. Moreover, a system capable of responding to diverse user requests and emotions is required.
[0669] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0670] In this invention, the server includes means for recognizing incoming calls from users, means for collecting the user's voice, means for converting the collected user's voice into text data, means for recognizing the user's emotional state from the text data, means for analyzing the user's question content from the text data, means for selecting a responder based on the analyzed question content and the recognized emotional state, means for generating answer data for the question content, means for converting the generated answer data into voice, means for providing the voiced answer to the user, and means for storing the interaction history. This enables flexible responses according to the user's emotional state, selection of a responder based on specific requests, and rapid and accurate response generation utilizing past interaction history.
[0671] "Incoming call recognition means" refers to a device or method for automatically detecting a call from a user and initiating a call.
[0672] "Voice collection means" refers to a device or method for acquiring the voice spoken by a user during a phone call in real time.
[0673] "Character conversion means" refers to a device or method for converting acquired audio data into character data.
[0674] "Emotion recognition means" refers to a device or method for analyzing and recognizing a user's emotional state from text data or audio data.
[0675] "Question analysis means" refers to a device or method for understanding and analyzing the content of a user's question from text data.
[0676] "Respondent selection means" refers to a device or method for selecting an appropriate respondent based on the analyzed question content and the recognized emotional state.
[0677] "Answer generation means" refers to a device or method for generating appropriate answer data in response to a user's question.
[0678] "Voice conversion means" refers to a device or method for converting generated response data into voice data.
[0679] "Answer provision means" refers to a device or method for providing answer data converted into audio format to a user.
[0680] "Response history storage means" refers to a device or method for recording and storing the content of a user's questions and response history.
[0681] Modes for carrying out the invention
[0682] The system according to this invention recognizes the user's emotions when they call a call center and provides a natural and flexible response. This system uses the following hardware and software.
[0683] Hardware and software
[0684] Telephone system or VoIP technology: Used to recognize incoming calls and initiate conversations.
[0685] ASR (Automatic Speech Recognition) engine: Collects speech data and converts it into text data.
[0686] Emotion Engine: Used to recognize the user's emotional state from text and audio data.
[0687] NLP (Natural Language Processing) technology: Used to analyze user questions from text data.
[0688] Database: Used to store user questions and response history.
[0689] Generative AI: Used to generate answer data in response to questions.
[0690] TTS (Text-to-Speech) engine: Used to convert generated response data into speech.
[0691] Specific processing of the program
[0692] 1. Recognition of incoming call
[0693] When a user calls the call center, the server automatically recognizes the incoming call and initiates the conversation. This process utilizes telephone systems and VoIP technology.
[0694] 2. Collection and conversion of audio data
[0695] When a user begins speaking, the server collects their voice in real time. The collected voice data is input into the ASR engine and converted into text format. For example, if a user asks, "Please tell me the opening hours of the supermarket," the ASR engine converts that voice into the text, "Please tell me the opening hours of the supermarket."
[0696] 3. Recognition of emotions
[0697] The server inputs the converted text data into the emotion engine, which analyzes elements such as tone, speed, and pitch of the voice. Based on this analysis, the emotion engine recognizes the user's emotional state (e.g., happy, angry, calm).
[0698] 4. Analysis of the inquiry content
[0699] The server analyzes the text data using NLP (Neuro-Linguistic Programming) techniques. In this process, the server classifies the text data into specific categories (e.g., "business hours," "return policy," etc.) and understands the specific content of the questions.
[0700] 5. Selection of a representative based on user requirements
[0701] When a user makes a specific request, the server selects an appropriate AI respondent based on that request. For example, in response to a request to "explain in a female voice," the server selects an AI respondent with a female voice profile, taking into account the user's emotional state when making the selection.
[0702] 6. Generating response data
[0703] The server retrieves relevant information from the database and uses generative AI to construct response data. For example, it might generate a response like, "The supermarket's opening hours are from 9 AM to 8 PM." This response is generated to appropriately address the user's question.
[0704] 7. Speech synthesis
[0705] The server inputs the generated response data into the TTS engine and converts it into speech format. Based on the selected voice profile, natural-sounding speech is generated. At this time, the tone and speed of the speech are also adjusted based on the emotional state recognized by the emotion engine.
[0706] 8. Providing a response
[0707] The server provides the generated audio data to the user in real time. Specifically, the audio data is played back during the call. This allows the user to receive responses in a natural conversational style.
[0708] 9. Record of interaction history
[0709] The server saves the interaction history to a database. The saved data includes the user's question, the answer provided, the date and time of the interaction, and the perceived emotional state. This information is used to respond quickly and accurately to future inquiries.
[0710] Specific example
[0711] Example 1: Request for a simple explanation in a woman's voice and emotional recognition.
[0712] User A calls and asks, "Could you tell me the supermarket's opening hours? Could you please explain it in a way that even a primary school student could understand, in a woman's voice?" The server recognizes this call, converts the user's voice to text, analyzes the question, and detects that it is requesting an explanation in a woman's voice. Furthermore, because the user's voice tone is calm, the emotion engine recognizes that the user is in a "calm" state. The server retrieves the opening hours information from the database, and the generative AI generates the answer, "The supermarket's opening hours are from 9 am to 8 pm," which is then converted into a woman's voice and delivered to User A in a calm tone.
[0713] Example of a prompt
[0714] "A user asked, 'Could you please tell me about supermarket opening hours? Could you explain it in a way that even a primary school student could understand, using a woman's voice?' The user's tone is calm. Please generate an answer and explain it calmly in a woman's voice."
[0715] Example 2: Harassment prevention and emotional recognition
[0716] Regular user B asks in an agitated voice, "What are your opening hours?" The server refers to past question history and quickly and accurately generates the answer, "Your opening hours are from 9 am to 8 pm." Furthermore, recognizing the agitated tone of user B, the emotion engine recognizes a state of "anger" and provides a response in a calm tone.
[0717] Example of a prompt
[0718] "User B asked in an excited voice, 'What are your opening hours?' Based on past question history, we know this same question has been asked multiple times. Please generate a calm and collected response."
[0719] Through this invention, we can provide a system that can flexibly respond to the diverse needs and emotions of users and improve the efficiency of call centers.
[0720] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0721] Specific processing flow of the program
[0722] Step 1:
[0723] Recognition of incoming call
[0724] The server automatically recognizes incoming calls when a user dials the call center.
[0725] Input: Phone call from a user (VoIP data or phone system data).
[0726] Processing: The server receives the call initiation signal via the telephone system or VoIP technology.
[0727] Output: Call start time and user call information (phone number, call start time, etc.).
[0728] Step 2:
[0729] Audio data collection and conversion
[0730] When a user begins speaking, the server collects their voice in real time. The collected voice data is then converted into text format by the ASR engine.
[0731] Input: User's voice data.
[0732] Processing: The server collects the audio data and the ASR engine converts the audio to text. For example, if the user says, "Please tell me the opening hours of the supermarket."
[0733] Output: Text data (e.g., "Please tell me about the supermarket's opening hours").
[0734] Step 3:
[0735] Recognition of emotions
[0736] The server inputs the converted text data into an emotion engine, which analyzes the tone, speed, pitch, and other characteristics of the speech.
[0737] Input: Text data and audio data.
[0738] Processing: The server uses an emotion engine to analyze text and audio data and recognize the user's emotions. For example, if the tone of voice is calm, it will be recognized as "calm."
[0739] Output: Emotional state data (e.g., "Calm").
[0740] Step 4:
[0741] Analysis of inquiry content
[0742] The server analyzes the text data using NLP (Neuro-Linguistic Programming) technology to understand the user's question.
[0743] Input: Text data.
[0744] Processing: The server uses NLP (Neuro-Linguistic Programming) technology to analyze text data and classify the questions into specific categories. For example, "business hours" or "return policy."
[0745] Output: Category data of the question content (e.g., "Business Hours").
[0746] Step 5:
[0747] Selecting a responder based on user requirements
[0748] When a user makes a specific request, the server selects an appropriate AI responder based on that request and the perceived emotional state.
[0749] Input: Question category data, emotional state data, user request data (e.g., "Explain in a female voice").
[0750] Processing: The server selects an appropriate AI responder with a voice profile based on the request and emotion. For example, it might select a voice profile with a "female voice".
[0751] Output: Selected AI responder data (e.g., "female voice").
[0752] Step 6:
[0753] Generating response data
[0754] The server retrieves relevant information from the database and uses generative AI to construct response data.
[0755] Input: Category data of the question content, information in the database.
[0756] Processing: The server uses generative AI to generate accurate answer data based on the question. For example, if the question is about opening hours, it will generate the answer "The supermarket's opening hours are from 9 am to 8 pm."
[0757] Output: Response data (Example: "The supermarket's operating hours are from 9 AM to 8 PM").
[0758] Step 7:
[0759] Speech synthesis
[0760] The server inputs the generated response data into the TTS engine and converts it into audio format.
[0761] Input: Response data, selected AI responder data.
[0762] Processing: The server uses a TTS engine to convert response data into audio data. During this process, the tone and speed of the voice are also adjusted based on the emotional state.
[0763] Output: Audio data.
[0764] Step 8:
[0765] Providing a response
[0766] The server provides the generated audio data to the user in real time.
[0767] Input: Audio data.
[0768] Processing: The server plays the audio data through the call.
[0769] Output: Receipt of the user's response in audio format.
[0770] Step 9:
[0771] Record of interaction history
[0772] The server saves the interaction history to a database.
[0773] Input: User's question, provided answer, date and time of interaction, perceived emotional state, call information.
[0774] Processing: The server records and stores this data in the database.
[0775] Output: Saved call history data.
[0776] (Application Example 2)
[0777] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0778] Traditional call center and virtual store systems could only provide standard responses to user inquiries, making it difficult to respond flexibly to users' emotional states or specific requests. Furthermore, they lacked sufficient functionality to efficiently utilize past inquiry history. Therefore, there were limitations in addressing situations where improving the user experience and reducing staff workload were required.
[0779] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for recognizing incoming calls from users, means for converting user voice into text data, means for analyzing the user's question content from the text data, means for selecting a responder based on the user's request, means for generating answer data for the question content, means for converting the generated answer data into voice, means for providing the voiced answer to the user, means for recognizing the user's emotions, means for adjusting the response based on the recognized emotions, and means for saving the response history. This enables quick and efficient responses by utilizing past inquiry history while flexibly responding to the user's emotions and specific requests.
[0780] "Means for recognizing incoming calls" refers to a function that automatically detects voice inquiries from users and notifies the system.
[0781] "A means of converting user voice into text data" refers to a function that converts input voice data into text format in real time.
[0782] "Methods for analyzing user inquiries from text data" refers to a function that converts speech into text and then analyzes that text data to understand the content of the user's inquiry.
[0783] "Means for selecting a responder based on user requests" refers to a function that selects a responder who meets specific criteria (e.g., gender, level of understanding) when a user requests them.
[0784] "Means for generating response data to questions" refers to a function that generates appropriate answers based on the user's inquiry.
[0785] "Means for converting generated response data into audio" refers to a function that converts generated text-based response data into audio format.
[0786] "Means of providing users with voiced responses" refers to a function that transmits voiced responses to users in real time.
[0787] "Means of recognizing user emotions" refers to a function that determines the user's emotional state through the analysis of voice data.
[0788] "Means of adjusting responses based on recognized emotions" refers to a function that adjusts the content and tone of responses to match the user's emotional state.
[0789] "Means for saving interaction history" refers to a function that records and saves past inquiries, the answers provided, the date and time of the interaction, and the perceived emotional state.
[0790] This invention is a system that provides flexible responses to user inquiries in a virtual store, tailored to the user's emotional state. The server performs response processing when a user makes an inquiry by voice using the following means.
[0791] First, when a user makes a voice inquiry within the virtual store, the device (e.g., smart glasses or a smartphone) collects the audio in real time and sends it to a server. The server recognizes the incoming call using a general speech recognition system. Technologies used include the Google Cloud Speech-to-Text API.
[0792] Next, the server converts the collected audio data into text data. This also uses the Google Cloud Speech-to-Text API. This conversion changes the audio data into text format.
[0793] Next, the server recognizes the user's emotions based on the text data. The emotion engine uses IBM Watson Tone Analyzer or Microsoft Azure's Text Analytics API. This allows it to determine the user's emotional state (e.g., joy, anger, sadness, surprise) from the audio data.
[0794] Furthermore, the server analyzes the text data and classifies the user's questions into specific categories (e.g., product details, return procedures, etc.). This analysis uses NLP technologies such as Google Cloud Natural Language API and SpaCy. Subsequently, it selects the appropriate respondent (e.g., a voice of a specific gender, a concise explanation, etc.) based on the user's request.
[0795] Generative AI models are used to generate response data. Specifically, OpenAI GPT-4 can be used. The server retrieves relevant information from the database and constructs response data using the generative AI model. For example, if a user asks, "How much stock do you have of this product?", the generative AI model will generate a response such as, "We currently have 10 units of this product in stock."
[0796] The generated response data is converted into speech format using a TTS engine such as Amazon Polly or Google Cloud Text-to-Speech. The server also adjusts the tone and speed of the speech based on the recognized emotions.
[0797] Finally, the transcribed response data is provided to the user in real time via the device. Users can enjoy a natural conversational experience through smart glasses or earphones.
[0798] The interaction history is stored in a database such as Microsoft SQL Server or MongoDB. This allows for quick and efficient responses to the same question again.
[0799] As a concrete example, consider a scenario where a customer inquires, "How much stock do you have of this item?" in a virtual store. In this case, the system recognizes the user's emotions and provides a calm tone of voice if the user is "excited," and a detailed and polite response if the user is "confused."
[0800] Examples of prompt statements include the following formats:
[0801] User inquiry: "How much stock do you have of this product?"
[0802] Perceived emotions: "excitement," "confusion"
[0803] Tone of response: "Calm tone," "Detailed and polite."
[0804] Using a generative AI model: "Currently, there are 10 of this item in stock."
[0805] By combining these methods, we can provide a system that can respond quickly and efficiently while flexibly responding to user emotions and specific requests.
[0806] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0807] Step 1:
[0808] The device recognizes the user's voice input.
[0809] Input: User voice inquiry.
[0810] Specific operation: In a virtual store, a user asks a question by voice, for example, "How much stock do you have of this item?" The device (smart glasses or smartphone) uses its built-in microphone to collect the voice and sends it to the server as digital audio data.
[0811] Step 2:
[0812] The server recognizes the incoming call from the user.
[0813] Input: Digital audio data.
[0814] Specific operation: The server uses a speech recognition system to confirm incoming calls and initiate appropriate processing. Here, a system is in place to process user inquiries in real time using a general speech recognition system.
[0815] Step 3:
[0816] The server converts the user's voice into text data.
[0817] Input: Digital audio data.
[0818] Output: Text data.
[0819] Specific operation: Use the Google Cloud Speech-to-Text API to convert audio data into text. For example, the audio "How much stock do you have of this product?" will be converted into text data.
[0820] Step 4:
[0821] The server analyzes the user's question from the text data.
[0822] Input: Text data.
[0823] Output: Category of the question.
[0824] Specific operation: Using Google Cloud Natural Language API and SpaCy, the text data is analyzed and the questions are categorized into categories such as "stock availability check".
[0825] Step 5:
[0826] The server recognizes the user's emotions.
[0827] Input: Text data.
[0828] Output: Recognized emotion.
[0829] Specific operation: Using IBM Watson Tone Analyzer and Microsoft Azure's Text Analytics API, it recognizes the user's emotions (e.g., joy, anger, sadness, surprise) from the audio. For example, an emotion such as "confusion" might be extracted.
[0830] Step 6:
[0831] The server selects a responder based on the user's request.
[0832] Input: Category of the question, perceived emotion.
[0833] Output: Selected responder.
[0834] Specific operation: Based on user requests (e.g., specific gender, level of understanding), the system selects the most appropriate responder profile (e.g., female voice, calm tone).
[0835] Step 7:
[0836] The server generates answer data for the question.
[0837] Input: Question category, selected responder profile.
[0838] Output: Generated response data.
[0839] Specific operation: Using generative AI models such as OpenAI GPT-4, it generates answer data based on the question. For example, it might generate an answer such as, "Currently, there are 10 units of this product in stock."
[0840] Step 8:
[0841] The server converts the generated response data into speech.
[0842] Input: Generated response data, selected respondent profile.
[0843] Output: Audio data.
[0844] Specific operation: Use Amazon Polly or Google Cloud Text-to-Speech to convert text data into speech data and adjust the tone and speed of the speech.
[0845] Step 9:
[0846] The server provides the user with an audio response.
[0847] Input: Audio data.
[0848] Output: Voice response to the user.
[0849] Specific operation: The generated audio data is delivered to the user in real time via a device (smart glasses or earphones). The user can experience a natural conversation.
[0850] Step 10:
[0851] The server saves the interaction history.
[0852] Input: User's question, generated answer, date and time of interaction, perceived emotion.
[0853] Output: Saved call history.
[0854] Specific actions: Use Microsoft SQL Server or MongoDB to record this information in a database, enabling quick responses to future queries.
[0855] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0856] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0857] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0858] [Third Embodiment]
[0859] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0860] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0861] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0862] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0863] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0864] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0865] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0866] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0867] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0868] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0869] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0870] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0871] The system of the present invention includes various means for providing natural and flexible responses when an end user calls a call center. The system configuration is as follows:
[0872] 1. Recognition of incoming call
[0873] When a user calls the call center, the server automatically recognizes the call and initiates the conversation. The technologies used here are general telephone systems and VoIP (Voice over Internet Protocol) technology.
[0874] 2. Collection and conversion of audio data
[0875] When a user begins speaking, the server collects their voice in real time. This voice data is converted into text using an ASR (Automatic Speech Recognition) engine. For example, if a user asks, "Please tell me the opening hours of the supermarket," the voice data is converted into the text "Please tell me the opening hours of the supermarket."
[0876] 3. Analysis of the inquiry content
[0877] The server analyzes the converted text data to understand the user's question. The technology used here is NLP (Natural Language Processing), which classifies the question into specific categories. For example, categories such as "business hours" or "return policy."
[0878] 4. Selection of a support person based on user requirements
[0879] When a user makes a specific request (for example, "I want an explanation in a female voice" or "I want an explanation that even a primary school student can understand"), the server selects an appropriate AI respondent based on that request. A specific voice profile is used for voice selection.
[0880] 5. Generating response data
[0881] Based on the question, the server retrieves relevant information from the database and constructs answer data using generative AI. For example, if the question is about opening hours, the answer "The supermarket's opening hours are from 9 am to 8 pm" will be generated.
[0882] 6. Speech synthesis
[0883] The generated response data is converted into speech format using a Text-to-Speech (TTS) engine. Natural-sounding speech is generated based on the selected voice profile.
[0884] 7. Providing a response
[0885] The generated audio data is played back to the user in real time during the call. This allows the user to receive responses in a natural conversational style.
[0886] 8. Record of interaction history
[0887] The server stores the interaction history in a database. This includes the user's question, the answer provided, and the date and time of the interaction. This information is used to respond quickly and accurately to future inquiries.
[0888] Specific example
[0889] Example 1: Request for a brief explanation in a female voice.
[0890] User A makes a phone call and asks, "Could you tell me the supermarket's opening hours? Could you please explain it in a way that even a primary school student could understand, using a woman's voice?" The server recognizes this call, converts the user's voice into text, analyzes the question, and detects a "woman's voice" as the request. The server retrieves the opening hours information from its database, and a generative AI generates the answer, "The supermarket's opening hours are from 9 am to 8 pm," which is then converted into a woman's voice and provided to User A.
[0891] Example 2: Harassment prevention measures
[0892] If a regular user B asks the same question hundreds of times every month, the server recognizes the call and converts the user's voice into text. Because the past question history is stored in the database, the server can refer to the already accumulated information and quickly and accurately generate and provide the answer, "Our business hours are from 9 am to 8 pm." This significantly reduces the burden on staff.
[0893] By following these steps, the present invention provides a system that can flexibly respond to diverse user demands and improve the efficiency of call centers.
[0894] The following describes the processing flow.
[0895] Step 1:
[0896] The user makes a phone call. The user dials a phone number and is connected to the call center.
[0897] Step 2:
[0898] The server recognizes the incoming call. The server detects the incoming signal and automatically initiates the call.
[0899] Step 3:
[0900] The server collects audio data. When the user starts speaking, the server collects audio data in real time.
[0901] Step 4:
[0902] The server converts the audio data into text data. The collected audio data is converted into text format using the ASR engine.
[0903] Step 5:
[0904] The server analyzes the text data. Using NLP (Neuro-Linguistic Programming) techniques, it analyzes the user's questions from the text data and identifies categories.
[0905] Step 6:
[0906] The server detects user requests. If there are user requests (e.g., "female voice," "make it easy enough for elementary school children to understand"), the server analyzes and recognizes them.
[0907] Step 7:
[0908] The server selects a responder. Based on the detected request, it selects the appropriate AI responder and voice profile.
[0909] Step 8:
[0910] The server generates the answer data. The generative AI retrieves information related to the question from the database and constructs the answer.
[0911] Step 9:
[0912] The server converts the response data into speech. A TTS engine is used to convert the text-based responses into speech using a selected voice profile.
[0913] Step 10:
[0914] The server provides the user with an audio response. The generated audio data is played back to the user in real time.
[0915] Step 11:
[0916] The server saves the interaction history. It records information such as the user's question, the answer provided, and the date and time of the interaction in a database.
[0917] (Example 1)
[0918] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0919] Traditional call center systems struggled to respond to user inquiries in real time, requiring a significant amount of manual work. This resulted in low efficiency in handling inquiries, increased susceptibility to human error, and inconsistent response quality. Furthermore, flexibly addressing specific requests (e.g., responses in a specific gender's voice or tailored to a specific level of understanding) was difficult, potentially leading to decreased user satisfaction. Additionally, the inability to effectively utilize past interaction history meant that responses to the same question were inconsistent, hindering efficient operation.
[0920] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0921] In this invention, the server includes means for recognizing incoming calls from users, means for converting user voice into text data, means for analyzing the user's questions from the text data, means for selecting a responder based on the user's request, means for generating answer data to the questions using natural language generation technology, means for converting the generated answer data into speech, means for providing the spoken answers to the user, and means for storing the interaction history. This makes it possible to respond quickly and flexibly to diverse user requests and improve the operational efficiency of the call center.
[0922] (Definitions of important words)
[0923] "Means for recognizing incoming calls" refers to a function that detects when a call has started when receiving a call from a user and initiates the appropriate processing.
[0924] "Methods for converting speech to text data" refers to technologies that collect spoken audio from users in real time and convert it into text data. Specifically, this utilizes ASR (Automatic Speech Recognition) technology.
[0925] "Means for analyzing question content" refers to a function that analyzes the converted text data to identify and classify the content of the user's question. It primarily uses NLP (Natural Language Processing) technology.
[0926] The "means of selecting a responder" refer to a function that chooses the optimal voice response profile based on the user's requests. This includes selecting a voice profile tailored to a specific gender or level of comprehension.
[0927] "Means for generating response data" refers to a function that collects relevant information based on the user's question and creates an appropriate response using natural language generation technology.
[0928] "Means for converting response data into speech" refers to a function that converts the generated text-based response data into natural-sounding speech using TTS (text-to-speech) technology.
[0929] "Means of providing users with voiced responses" refers to a function that provides generated voice data to users in real time via telephone.
[0930] "Means for saving interaction history" refers to a function that saves historical data such as the user's questions, the answers provided, and the date and time of the call in a database, for use in handling future inquiries.
[0931] "Natural language generation technology" is a technique that generates meaningful natural language sentences based on given input data. It commonly utilizes generative AI models.
[0932] A "prompt" is text data input to a generative AI model, containing specific instructions or questions for the model.
[0933] The system of this invention is designed to provide natural and flexible responses when end users call a call center. The system configuration is as follows:
[0934] When a user calls the call center, the server automatically recognizes the call and initiates the conversation. At this stage, a standard telephone system or VoIP (Voice over Internet Protocol) technology is used.
[0935] Next, when the user begins to speak, the server collects the user's voice in real time. This is done using microphones and acoustic detection devices. This voice data is converted into text format using an ASR (Automatic Speech Recognition) engine. Specific technologies include Google Cloud Speech-to-Text and IBM Watson Speech to Text. For example, if the user asks, "Please tell me the opening hours of the supermarket," the voice data is converted into text data that reads, "Please tell me the opening hours of the supermarket."
[0936] The converted text data is sent to a server and analyzed using an NLP (Natural Language Processing) engine. Technologies include Google Cloud Natural Language API and Microsoft Azure Text Analytics. The analysis categorizes user inquiries into specific categories (e.g., "Business Hours," "Return Policy").
[0937] Next, when a user makes a specific request (for example, "I want an explanation in a woman's voice" or "Please explain it in a way that even a primary school student can understand"), the server analyzes the request and selects an appropriate voice profile. The server accesses an internal database and selects a voice responder based on the optimal voice profile for the request.
[0938] The answer data for the questions is generated by retrieving relevant information from a database and using generative AI. OpenAI GPT-3 or similar natural language generation technologies are used for this generative AI. For example, the answer data "The supermarket's operating hours are from 9 AM to 8 PM" might be generated.
[0939] The generated response data is converted into speech format using a Text-to-Speech (TTS) engine. Amazon Polly and Google Cloud Text-to-Speech are used as TTS engines. The converted speech data provides natural-sounding speech based on the selected speech profile.
[0940] Ultimately, the generated audio data is played back to the user in real time via the call, allowing the user to receive responses in a natural conversational style.
[0941] The interaction history is stored in a database by the server. This includes the user's question, the answer provided, and the date and time of the interaction. This information is used to respond quickly and accurately to future inquiries.
[0942] Specific example
[0943] Example 1: Request for a brief explanation in a female voice.
[0944] User A makes a phone call and asks, "Could you tell me the supermarket's opening hours? Could you please explain it in a way that even a primary school student could understand, using a woman's voice?" The server recognizes this call and converts the user's voice into text data. Next, it analyzes the question and detects a "woman's voice" as the request. The server retrieves the opening hours information from the database, and a generative AI generates the answer, "The supermarket's opening hours are from 9 am to 8 pm." This is then converted into a woman's voice and provided to User A.
[0945] Example 2: Harassment prevention measures
[0946] If a regular user B asks the same question hundreds of times every month, the server recognizes the incoming call and converts the user's voice into text. Because past question history is stored in the database, the server can refer to the already accumulated information and quickly generate and provide an answer such as, "Our business hours are from 9 am to 8 pm." This significantly reduces the burden on staff.
[0947] Example of a prompt
[0948] "Could you please tell me about the supermarket's opening hours? Could you explain it in a way that even a primary school student could understand?"
[0949] "Our business hours are from 9 AM to 8 PM."
[0950] "Do you have any other questions?"
[0951] In this way, the system can respond quickly and flexibly to diverse user requests, thereby improving the operational efficiency of the call center.
[0952] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0953] Step 1:
[0954] When a user calls the call center, the server recognizes the incoming call and initiates the conversation. At this stage, the call is connected to the call center's system using VoIP technology. In other words, a call commencement signal (output) is generated in response to the user's phone call (input). The server logs the user's number and the time the call started.
[0955] Step 2:
[0956] When a user begins speaking, the server collects their voice in real time. This voice data (input) is sent to an ASR engine and converted into text format (output). Specifically, voice data is collected using microphones or acoustic detection devices and sent to an ASR engine (such as Google Cloud Speech-to-Text). The converted text data is temporarily stored for subsequent processing.
[0957] Step 3:
[0958] The server sends the converted text data to the NLP engine, which analyzes the user's question. This text data (input) is analyzed by the NLP engine, and the question is classified into a specific category (output). Technologies used here include Google Cloud Natural Language API and Microsoft Azure Text Analytics. The category information is also used to generate the response in the next step.
[0959] Step 4:
[0960] The server detects specific requests (e.g., "female voice," "explain it clearly") from the converted text data. This text data (input) contains the user's requests, and an appropriate voice profile is selected based on the analysis results (output). The server records this information in its internal database.
[0961] Step 5:
[0962] The server retrieves relevant information from the database based on the analyzed question. This relevant information (output) corresponding to the question (input) is converted into answer data using a generative AI model (e.g., OpenAI GPT-3). The generated answer data is temporarily stored because it will be converted into speech in the next step.
[0963] Step 6:
[0964] The server sends the generated response data to the TTS engine, where it is converted into speech format. This text-based response data (input) is then converted into speech data (output) by the TTS engine (such as Amazon Polly or Google Cloud Text-to-Speech). The server generates natural-sounding speech based on the selected speech profile.
[0965] Step 7:
[0966] The server sends the generated audio data to the call channel and plays it back to the user in real time. This audio data (input) is provided to the user through the call channel (output). The server confirms that the response has been provided and continues or ends the call with the user as needed.
[0967] Step 8:
[0968] The server stores call history, including call content, generated responses, and call date and time, in a database. This call history data (input) is recorded in the database and used for future inquiries (output). This helps in quickly generating answers to the same question again and improves operational efficiency.
[0969] (Application Example 1)
[0970] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0971] Current voice assistant systems in autonomous vehicles lack the ability to provide natural and flexible responses in passenger interaction. In particular, the systems are inadequate in handling situations where passengers request answers tailored to specific voice qualities or levels of understanding, or when they repeatedly ask the same questions. This leads to decreased passenger satisfaction and compromises in-vehicle comfort.
[0972] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0973] In this invention, the server includes means for recognizing a user entering the vehicle, means for converting the user's voice into text data, means for analyzing the user's question from the text data, means for selecting a responder based on the user's request, means including a generative AI model for generating answer data to the question, means for converting the generated answer data into speech, means for providing the spoken answer to the user, and means for storing the response history. This makes it possible to provide natural and flexible responses in real time to various requests from passengers in an autonomous vehicle.
[0974] "Means for recognizing a user entering the vehicle" refers to a device or system used in an autonomous vehicle to recognize when a passenger has boarded the vehicle.
[0975] "Means for converting user speech into text data" refers to technology or equipment for converting speech spoken by passengers into text data in real time.
[0976] "Means for analyzing user questions from text data" refers to a system or technology for analyzing passengers' questions based on converted text data and understanding their intent.
[0977] "Means for selecting a responder based on user requests" refers to a system or technology for selecting an appropriate voice profile or responder in accordance with the content of a passenger's request.
[0978] "Means including a generative AI model that generates response data to a question" refers to a system or technology that includes an artificial intelligence model for generating appropriate answers to passengers' questions.
[0979] "Means for converting generated response data into audio" refers to a technology or device for converting generated text-formatted response data into audio data.
[0980] "Means of providing users with voiced responses" refers to a device or system that plays voice data in real time and provides responses to passengers.
[0981] "Means for saving interaction history" refers to a system or technology that records the history of interactions with passengers and uses it to handle future inquiries.
[0982] The system of the present invention realizes a voice recognition assistant in an autonomous vehicle by, for example, following the procedure described below. By utilizing specific hardware and software, it is possible to interact with the user in a natural and flexible manner.
[0983] System configuration and program processing
[0984] 1. Means for recognizing a vehicle entering the premises from a user.
[0985] The server uses hardware such as sensors and cameras to recognize when a user has boarded the vehicle. Machine learning models may also be used for this purpose.
[0986] 2. Means for converting user speech into text data
[0987] The server analyzes the audio data collected through the microphone in real time. The audio data is converted into text data using an ASR (Automatic Speech Recognition) engine such as AWS Transcribe or Google Cloud Speech-to-Text API.
[0988] 3. Means for analyzing user questions from text data
[0989] The server analyzes the converted text data to understand the intent behind the user's question. This process utilizes NLP (Natural Language Processing) technologies such as SpaCy and Google Cloud Natural Language API.
[0990] 4. Means for selecting a responder based on user requests
[0991] The server selects a specific voice profile and response method. For example, it might use Azure's Cognitive Services Speech API to select a voice quality and speaking style that suits the user's request.
[0992] 5. Means including a generative AI model that generates response data to the content of a question.
[0993] The server uses a generative AI model (e.g., GPT-4) to generate response data corresponding to the query. It uses the OpenAI API to generate appropriate responses based on the prompt.
[0994] 6. Means for converting generated response data into speech.
[0995] The server converts the generated text-based response data into speech using a Text-to-Speech (TTS) engine. Google TTS and Amazon Polly are among the engines used for this process.
[0996] 7. Means of providing users with voiced responses
[0997] The server provides the converted audio data to the user in real time through the in-car speaker system.
[0998] 8. Means for saving interaction history
[0999] The server stores the user interaction history in a database. Services such as Amazon RDS and Google Firestore are used to facilitate future inquiries.
[1000] Specific example
[1001] One afternoon, passenger A gets into a self-driving vehicle and asks, "Please tell me the next gas station." In this case, the system operates as follows:
[1002] The server recognizes passenger A's voice and converts the audio data into text.
[1003] The converted text data is analyzed to understand that it is a question asking for the location of a gas station.
[1004] A gentle male voice is selected as the appropriate voice quality, and the generating AI model responds, "The next gas station is 2 kilometers ahead on the right."
[1005] This text is converted into audio data by the TTS engine and played back in real time through the car's speakers.
[1006] Finally, save this question and answer exchange to the database.
[1007] Example of a prompt
[1008] The following prompt is used in a situation where a passenger requests, "Tell me the next gas station," and the system responds, "The next gas station is 2 kilometers ahead on the right":
[1009] User: "Please tell me the next gas station."
[1010] System: "The next gas station is 2 kilometers ahead on the right."
[1011] The system of the present invention, through these procedures, enables highly natural and flexible interaction with the user in an autonomous vehicle.
[1012] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1013] Step 1:
[1014] The server recognizes when a user boards the vehicle. Specifically, it uses data from on-board sensors and cameras to detect entry. The server analyzes this input data to confirm the presence of a passenger.
[1015] Step 2:
[1016] When a user begins speaking in the vehicle, the server collects audio data through the microphone. This audio data is then input to the server.
[1017] Step 3:
[1018] The server converts the collected audio data into text data using AWS Transcribe or the Google Cloud Speech-to-Text API. The input audio data is passed through the ASR engine, and the converted data in text format is output.
[1019] Step 4:
[1020] The server analyzes the generated text data. Specifically, it uses NLP technologies such as SpaCy and Google Cloud Natural Language API to analyze the text data and identify the user's question. It processes the input text data and generates the intent and category of the question as output.
[1021] Step 5:
[1022] The server selects the appropriate voice profile and speaker based on the user's request. This is done by using the Azure Cognitive Services Speech API to select a voice profile according to the user's request (e.g., "a calm male voice"). The user's request data is used as input, and the selected voice profile is obtained as output.
[1023] Step 6:
[1024] The server generates answer data in response to the user's questions. Using the OpenAI GPT-4 model, it generates appropriate answers based on the text data input. This generated answer data is then output.
[1025] Step 7:
[1026] The server converts the generated text-based response data into audio data. It uses a TTS engine such as Google TTS or Amazon Polly to convert the text data to audio data. Text response data is used as input, and audio data is generated as output.
[1027] Step 8:
[1028] The server plays the generated audio data in real time through the car's speakers. This allows the user to receive an audio response. The audio data is sent to the speakers, and an audio response is output.
[1029] Step 9:
[1030] The server stores the user interaction history in a database. Amazon RDS or Google Firestore is used to store question content and answer data. Interaction data is received as input and saved to the database as output.
[1031] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1032] The system of the present invention includes various means for recognizing the user's emotions and providing natural and flexible responses when an end user makes a call to a call center. The system configuration is as follows:
[1033] 1. Recognition of incoming call
[1034] When a user calls the call center, the server automatically recognizes the call and initiates the conversation. The technologies used here are general telephone systems and VoIP (Voice over Internet Protocol) technology.
[1035] 2. Collection and conversion of audio data
[1036] When a user begins speaking, the server collects their voice in real time. This voice data is converted into text using an ASR (Automatic Speech Recognition) engine. For example, if a user asks, "Please tell me the opening hours of the supermarket," the voice data is converted into the text "Please tell me the opening hours of the supermarket."
[1037] 3. Recognition of emotions
[1038] The collected audio data is analyzed through an emotion engine. The emotion engine recognizes the user's emotions (e.g., joy, anger, sadness, surprise) based on factors such as tone, speed, and pitch of the voice. This process determines the user's current emotional state.
[1039] 4. Analysis of the inquiry content
[1040] The server analyzes the converted text data to understand the user's question. The technology used here is NLP (Natural Language Processing), which classifies the question into specific categories. For example, categories such as "business hours" or "return policy."
[1041] 5. Selection of a representative based on user requirements
[1042] When a user makes a specific request (for example, "I want an explanation in a woman's voice" or "I want an explanation that even a primary school student can understand"), the server selects an appropriate AI respondent based on that request. A specific voice profile is used for voice selection. Furthermore, the appropriate response is selected based on the emotional state recognized by the emotion engine.
[1043] 6. Generating response data
[1044] Based on the question, the server retrieves relevant information from the database and constructs answer data using generative AI. For example, if the question is about opening hours, the answer "The supermarket's opening hours are from 9 am to 8 pm" will be generated.
[1045] 7. Speech synthesis
[1046] The generated response data is converted into speech format using a Text-to-Speech (TTS) engine. Natural-sounding speech is generated based on the selected voice profile. The tone and speed of the speech are also adjusted based on the emotions recognized by the emotion engine.
[1047] 8. Providing a response
[1048] The generated audio data is played back to the user in real time during the call. This allows the user to receive responses in a natural conversational style. A key feature is that the emotion engine provides voice responses that are tailored to the user's emotions.
[1049] 9. Record of interaction history
[1050] The server stores interaction history in a database. This includes the user's question, the answer provided, the date and time of the interaction, and the perceived emotional state. This information is used to respond quickly and accurately to future inquiries.
[1051] Specific example
[1052] Example 1: Request for a simple explanation in a woman's voice and emotional recognition.
[1053] User A makes a phone call and asks, "Could you tell me the supermarket's opening hours? Could you please explain it in a way that even a primary school student could understand, using a woman's voice?" The server recognizes this call, converts the user's voice into text, analyzes the question, and detects a "woman's voice" as the request. Furthermore, because the user's voice tone is calm, the emotion engine recognizes a "calm" state. The server retrieves the opening hours information from the database, and the generative AI generates the answer, "The supermarket's opening hours are from 9 am to 8 pm," which is then converted into a woman's voice and delivered to User A in a calm tone.
[1054] Example 2: Harassment prevention and emotional recognition
[1055] If a regular user B asks the same question hundreds of times every month, the server recognizes the incoming call and converts the user's voice into text. Since the past question history is stored in the database, the server refers to the already accumulated information and quickly and accurately generates and provides the answer, "Our business hours are from 9 am to 8 pm." Furthermore, if user B's voice is agitated, the emotion engine recognizes a state of "anger" and provides a response in a calm tone. This significantly reduces the burden on staff while enabling responses that are considerate of the user's emotions.
[1056] By following these steps, the present invention provides a system that can flexibly respond to the diverse needs and emotions of users and improve the efficiency of call centers.
[1057] The following describes the processing flow.
[1058] Step 1:
[1059] The user makes a phone call. The user dials a phone number to the call center and is connected.
[1060] Step 2:
[1061] The server recognizes the incoming call. The server detects the incoming signal and automatically initiates the call.
[1062] Step 3:
[1063] The server collects audio data. When the user starts speaking, the server collects this audio data in real time.
[1064] Step 4:
[1065] The server converts the audio data into text data. The collected audio data is converted into text format using the ASR engine.
[1066] Step 5:
[1067] The server analyzes the text data. Using NLP (Neuro-Linguistic Programming) techniques, it analyzes the user's questions from the text data and categorizes them into specific categories.
[1068] Step 6:
[1069] The server uses an emotion engine to recognize the user's emotions. It analyzes the tone, speed, and pitch of the voice to identify the user's emotional state.
[1070] Step 7:
[1071] The server detects user requests. For example, if there is a request to "explain in a female voice," the server recognizes that request. Emotional states are also taken into consideration.
[1072] Step 8:
[1073] The server selects a responder. Based on the detected request and emotional state, it chooses the most appropriate AI responder and voice profile.
[1074] Step 9:
[1075] The server generates the answer data. The generative AI retrieves information related to the question from the database and constructs the answer.
[1076] Step 10:
[1077] The server converts the response data into speech. Using a TTS engine, the text-based responses are converted into speech using a selected voice profile. The voice tone and speed are adjusted based on the results of the emotion engine.
[1078] Step 11:
[1079] The server provides the user with an audio response. The generated audio data is played back to the user in real time.
[1080] Step 12:
[1081] The server saves the interaction history. It records the user's question, the answer provided, the date and time of the interaction, and the perceived emotional state in a database.
[1082] Step 13:
[1083] The server analyzes the interaction history and makes improvements as needed. Based on the stored history, data analysis is performed to improve the system's response quality.
[1084] (Example 2)
[1085] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1086] Traditional call center systems were unable to analyze user voices in real time and provide flexible responses tailored to their emotional state. Furthermore, selecting the appropriate responder based on specific requests and generating quick answers using past interaction history were difficult. Therefore, there is a need to improve both user satisfaction and response efficiency. Moreover, a system capable of responding to diverse user requests and emotions is required.
[1087] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1088] In this invention, the server includes means for recognizing incoming calls from users, means for collecting the user's voice, means for converting the collected user's voice into text data, means for recognizing the user's emotional state from the text data, means for analyzing the user's question content from the text data, means for selecting a responder based on the analyzed question content and the recognized emotional state, means for generating answer data for the question content, means for converting the generated answer data into voice, means for providing the voiced answer to the user, and means for storing the interaction history. This enables flexible responses according to the user's emotional state, selection of a responder based on specific requests, and rapid and accurate response generation utilizing past interaction history.
[1089] "Incoming call recognition means" refers to a device or method for automatically detecting a call from a user and initiating a call.
[1090] "Voice collection means" refers to a device or method for acquiring the voice spoken by a user during a phone call in real time.
[1091] "Character conversion means" refers to a device or method for converting acquired audio data into character data.
[1092] "Emotion recognition means" refers to a device or method for analyzing and recognizing a user's emotional state from text data or audio data.
[1093] "Question analysis means" refers to a device or method for understanding and analyzing the content of a user's question from text data.
[1094] "Respondent selection means" refers to a device or method for selecting an appropriate respondent based on the analyzed question content and the recognized emotional state.
[1095] "Answer generation means" refers to a device or method for generating appropriate answer data in response to a user's question.
[1096] "Voice conversion means" refers to a device or method for converting generated response data into voice data.
[1097] "Answer provision means" refers to a device or method for providing answer data converted into audio format to a user.
[1098] "Response history storage means" refers to a device or method for recording and storing the content of a user's questions and response history.
[1099] Modes for carrying out the invention
[1100] The system according to this invention recognizes the user's emotions when they call a call center and provides a natural and flexible response. This system uses the following hardware and software.
[1101] Hardware and software
[1102] Telephone system or VoIP technology: Used to recognize incoming calls and initiate conversations.
[1103] ASR (Automatic Speech Recognition) engine: Collects speech data and converts it into text data.
[1104] Emotion Engine: Used to recognize the user's emotional state from text and audio data.
[1105] NLP (Natural Language Processing) technology: Used to analyze user questions from text data.
[1106] Database: Used to store user questions and response history.
[1107] Generative AI: Used to generate answer data in response to questions.
[1108] TTS (Text-to-Speech) engine: Used to convert generated response data into speech.
[1109] Specific processing of the program
[1110] 1. Recognition of incoming call
[1111] When a user calls the call center, the server automatically recognizes the incoming call and initiates the conversation. This process utilizes telephone systems and VoIP technology.
[1112] 2. Collection and conversion of audio data
[1113] When a user begins speaking, the server collects their voice in real time. The collected voice data is input into the ASR engine and converted into text format. For example, if a user asks, "Please tell me the opening hours of the supermarket," the ASR engine converts that voice into the text, "Please tell me the opening hours of the supermarket."
[1114] 3. Recognition of emotions
[1115] The server inputs the converted text data into the emotion engine, which analyzes elements such as tone, speed, and pitch of the voice. Based on this analysis, the emotion engine recognizes the user's emotional state (e.g., happy, angry, calm).
[1116] 4. Analysis of the inquiry content
[1117] The server analyzes the text data using NLP (Neuro-Linguistic Programming) techniques. In this process, the server classifies the text data into specific categories (e.g., "business hours," "return policy," etc.) and understands the specific content of the questions.
[1118] 5. Selection of a representative based on user requirements
[1119] When a user makes a specific request, the server selects an appropriate AI respondent based on that request. For example, in response to a request to "explain in a female voice," the server selects an AI respondent with a female voice profile, taking into account the user's emotional state when making the selection.
[1120] 6. Generating response data
[1121] The server retrieves relevant information from the database and uses generative AI to construct response data. For example, it might generate a response like, "The supermarket's opening hours are from 9 AM to 8 PM." This response is generated to appropriately address the user's question.
[1122] 7. Speech synthesis
[1123] The server inputs the generated response data into the TTS engine and converts it into speech format. Based on the selected voice profile, natural-sounding speech is generated. At this time, the tone and speed of the speech are also adjusted based on the emotional state recognized by the emotion engine.
[1124] 8. Providing a response
[1125] The server provides the generated audio data to the user in real time. Specifically, the audio data is played back during the call. This allows the user to receive responses in a natural conversational style.
[1126] 9. Record of interaction history
[1127] The server saves the interaction history to a database. The saved data includes the user's question, the answer provided, the date and time of the interaction, and the perceived emotional state. This information is used to respond quickly and accurately to future inquiries.
[1128] Specific example
[1129] Example 1: Request for a simple explanation in a woman's voice and emotional recognition.
[1130] User A calls and asks, "Could you tell me the supermarket's opening hours? Could you please explain it in a way that even a primary school student could understand, in a woman's voice?" The server recognizes this call, converts the user's voice to text, analyzes the question, and detects that it is requesting an explanation in a woman's voice. Furthermore, because the user's voice tone is calm, the emotion engine recognizes that the user is in a "calm" state. The server retrieves the opening hours information from the database, and the generative AI generates the answer, "The supermarket's opening hours are from 9 am to 8 pm," which is then converted into a woman's voice and delivered to User A in a calm tone.
[1131] Example of a prompt
[1132] "A user asked, 'Could you please tell me about supermarket opening hours? Could you explain it in a way that even a primary school student could understand, using a woman's voice?' The user's tone is calm. Please generate an answer and explain it calmly in a woman's voice."
[1133] Example 2: Harassment prevention and emotional recognition
[1134] Regular user B asks in an agitated voice, "What are your opening hours?" The server refers to past question history and quickly and accurately generates the answer, "Your opening hours are from 9 am to 8 pm." Furthermore, recognizing the agitated tone of user B, the emotion engine recognizes a state of "anger" and provides a response in a calm tone.
[1135] Example of a prompt
[1136] "User B asked in an excited voice, 'What are your opening hours?' Based on past question history, we know this same question has been asked multiple times. Please generate a calm and collected response."
[1137] Through this invention, we can provide a system that can flexibly respond to the diverse needs and emotions of users and improve the efficiency of call centers.
[1138] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1139] Specific processing flow of the program
[1140] Step 1:
[1141] Recognition of incoming call
[1142] The server automatically recognizes incoming calls when a user dials the call center.
[1143] Input: Phone call from a user (VoIP data or phone system data).
[1144] Processing: The server receives the call initiation signal via the telephone system or VoIP technology.
[1145] Output: Call start time and user call information (phone number, call start time, etc.).
[1146] Step 2:
[1147] Audio data collection and conversion
[1148] When a user begins speaking, the server collects their voice in real time. The collected voice data is then converted into text format by the ASR engine.
[1149] Input: User's voice data.
[1150] Processing: The server collects the audio data and the ASR engine converts the audio to text. For example, if the user says, "Please tell me the opening hours of the supermarket."
[1151] Output: Text data (e.g., "Please tell me about the supermarket's opening hours").
[1152] Step 3:
[1153] Recognition of emotions
[1154] The server inputs the converted text data into an emotion engine, which analyzes the tone, speed, pitch, and other characteristics of the speech.
[1155] Input: Text data and audio data.
[1156] Processing: The server uses an emotion engine to analyze text and audio data and recognize the user's emotions. For example, if the tone of voice is calm, it will be recognized as "calm."
[1157] Output: Emotional state data (e.g., "Calm").
[1158] Step 4:
[1159] Analysis of inquiry content
[1160] The server analyzes the text data using NLP (Neuro-Linguistic Programming) technology to understand the user's question.
[1161] Input: Text data.
[1162] Processing: The server uses NLP (Neuro-Linguistic Programming) technology to analyze text data and classify the questions into specific categories. For example, "business hours" or "return policy."
[1163] Output: Category data of the question content (e.g., "Business Hours").
[1164] Step 5:
[1165] Selecting a responder based on user requirements
[1166] When a user makes a specific request, the server selects an appropriate AI responder based on that request and the perceived emotional state.
[1167] Input: Question category data, emotional state data, user request data (e.g., "Explain in a female voice").
[1168] Processing: The server selects an appropriate AI responder with a voice profile based on the request and emotion. For example, it might select a voice profile with a "female voice".
[1169] Output: Selected AI responder data (e.g., "female voice").
[1170] Step 6:
[1171] Generating response data
[1172] The server retrieves relevant information from the database and uses generative AI to construct response data.
[1173] Input: Category data of the question content, information in the database.
[1174] Processing: The server uses generative AI to generate accurate answer data based on the question. For example, if the question is about opening hours, it will generate the answer "The supermarket's opening hours are from 9 am to 8 pm."
[1175] Output: Response data (Example: "The supermarket's operating hours are from 9 AM to 8 PM").
[1176] Step 7:
[1177] Speech synthesis
[1178] The server inputs the generated response data into the TTS engine and converts it into audio format.
[1179] Input: Response data, selected AI responder data.
[1180] Processing: The server uses a TTS engine to convert response data into audio data. During this process, the tone and speed of the voice are also adjusted based on the emotional state.
[1181] Output: Audio data.
[1182] Step 8:
[1183] Providing a response
[1184] The server provides the generated audio data to the user in real time.
[1185] Input: Audio data.
[1186] Processing: The server plays the audio data through the call.
[1187] Output: Receipt of the user's response in audio format.
[1188] Step 9:
[1189] Record of interaction history
[1190] The server saves the interaction history to a database.
[1191] Input: User's question, provided answer, date and time of interaction, perceived emotional state, call information.
[1192] Processing: The server records and stores this data in the database.
[1193] Output: Saved call history data.
[1194] (Application Example 2)
[1195] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1196] Traditional call center and virtual store systems could only provide standard responses to user inquiries, making it difficult to respond flexibly to users' emotional states or specific requests. Furthermore, they lacked sufficient functionality to efficiently utilize past inquiry history. Therefore, there were limitations in addressing situations where improving the user experience and reducing staff workload were required.
[1197] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for recognizing incoming calls from users, means for converting user voice into text data, means for analyzing the user's question content from the text data, means for selecting a responder based on the user's request, means for generating answer data for the question content, means for converting the generated answer data into voice, means for providing the voiced answer to the user, means for recognizing the user's emotions, means for adjusting the response based on the recognized emotions, and means for saving the response history. This enables quick and efficient responses by utilizing past inquiry history while flexibly responding to the user's emotions and specific requests.
[1198] "Means for recognizing incoming calls" refers to a function that automatically detects voice inquiries from users and notifies the system.
[1199] "A means of converting user voice into text data" refers to a function that converts input voice data into text format in real time.
[1200] "Methods for analyzing user inquiries from text data" refers to a function that converts speech into text and then analyzes that text data to understand the content of the user's inquiry.
[1201] "Means for selecting a responder based on user requests" refers to a function that selects a responder who meets specific criteria (e.g., gender, level of understanding) when a user requests them.
[1202] "Means for generating response data to questions" refers to a function that generates appropriate answers based on the user's inquiry.
[1203] "Means for converting generated response data into audio" refers to a function that converts generated text-based response data into audio format.
[1204] "Means of providing users with voiced responses" refers to a function that transmits voiced responses to users in real time.
[1205] "Means of recognizing user emotions" refers to a function that determines the user's emotional state through the analysis of voice data.
[1206] "Means of adjusting responses based on recognized emotions" refers to a function that adjusts the content and tone of responses to match the user's emotional state.
[1207] "Means for saving interaction history" refers to a function that records and saves past inquiries, the answers provided, the date and time of the interaction, and the perceived emotional state.
[1208] This invention is a system that provides flexible responses to user inquiries in a virtual store, tailored to the user's emotional state. The server performs response processing when a user makes an inquiry by voice using the following means.
[1209] First, when a user makes a voice inquiry within the virtual store, the device (e.g., smart glasses or a smartphone) collects the audio in real time and sends it to a server. The server recognizes the incoming call using a general speech recognition system. Technologies used include the Google Cloud Speech-to-Text API.
[1210] Next, the server converts the collected audio data into text data. This also uses the Google Cloud Speech-to-Text API. This conversion changes the audio data into text format.
[1211] Next, the server recognizes the user's emotions based on the text data. The emotion engine uses IBM Watson Tone Analyzer or Microsoft Azure's Text Analytics API. This allows it to determine the user's emotional state (e.g., joy, anger, sadness, surprise) from the audio data.
[1212] Furthermore, the server analyzes the text data and classifies the user's questions into specific categories (e.g., product details, return procedures, etc.). This analysis uses NLP technologies such as Google Cloud Natural Language API and SpaCy. Subsequently, it selects the appropriate respondent (e.g., a voice of a specific gender, a concise explanation, etc.) based on the user's request.
[1213] Generative AI models are used to generate response data. Specifically, OpenAI GPT-4 can be used. The server retrieves relevant information from the database and constructs response data using the generative AI model. For example, if a user asks, "How much stock do you have of this product?", the generative AI model will generate a response such as, "We currently have 10 units of this product in stock."
[1214] The generated response data is converted into speech format using a TTS engine such as Amazon Polly or Google Cloud Text-to-Speech. The server also adjusts the tone and speed of the speech based on the recognized emotions.
[1215] Finally, the transcribed response data is provided to the user in real time via the device. Users can enjoy a natural conversational experience through smart glasses or earphones.
[1216] The interaction history is stored in a database such as Microsoft SQL Server or MongoDB. This allows for quick and efficient responses to the same question again.
[1217] As a concrete example, consider a scenario where a customer inquires, "How much stock do you have of this item?" in a virtual store. In this case, the system recognizes the user's emotions and provides a calm tone of voice if the user is "excited," and a detailed and polite response if the user is "confused."
[1218] Examples of prompt statements include the following formats:
[1219] User inquiry: "How much stock do you have of this product?"
[1220] Perceived emotions: "excitement," "confusion"
[1221] Tone of response: "Calm tone," "Detailed and polite."
[1222] Using a generative AI model: "Currently, there are 10 of this item in stock."
[1223] By combining these methods, we can provide a system that can respond quickly and efficiently while flexibly responding to user emotions and specific requests.
[1224] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1225] Step 1:
[1226] The device recognizes the user's voice input.
[1227] Input: User voice inquiry.
[1228] Specific operation: In a virtual store, a user asks a question by voice, for example, "How much stock do you have of this item?" The device (smart glasses or smartphone) uses its built-in microphone to collect the voice and sends it to the server as digital audio data.
[1229] Step 2:
[1230] The server recognizes the incoming call from the user.
[1231] Input: Digital audio data.
[1232] Specific operation: The server uses a speech recognition system to confirm incoming calls and initiate appropriate processing. Here, a system is in place to process user inquiries in real time using a general speech recognition system.
[1233] Step 3:
[1234] The server converts the user's voice into text data.
[1235] Input: Digital audio data.
[1236] Output: Text data.
[1237] Specific operation: Use the Google Cloud Speech-to-Text API to convert audio data into text. For example, the audio "How much stock do you have of this product?" will be converted into text data.
[1238] Step 4:
[1239] The server analyzes the user's question from the text data.
[1240] Input: Text data.
[1241] Output: Category of the question.
[1242] Specific operation: Using Google Cloud Natural Language API and SpaCy, the text data is analyzed and the questions are categorized into categories such as "stock availability check".
[1243] Step 5:
[1244] The server recognizes the user's emotions.
[1245] Input: Text data.
[1246] Output: Recognized emotion.
[1247] Specific operation: Using IBM Watson Tone Analyzer and Microsoft Azure's Text Analytics API, it recognizes the user's emotions (e.g., joy, anger, sadness, surprise) from the audio. For example, an emotion such as "confusion" might be extracted.
[1248] Step 6:
[1249] The server selects a responder based on the user's request.
[1250] Input: Category of the question, perceived emotion.
[1251] Output: Selected responder.
[1252] Specific operation: Based on user requests (e.g., specific gender, level of understanding), the system selects the most appropriate responder profile (e.g., female voice, calm tone).
[1253] Step 7:
[1254] The server generates answer data for the question.
[1255] Input: Question category, selected responder profile.
[1256] Output: Generated response data.
[1257] Specific operation: Using generative AI models such as OpenAI GPT-4, it generates answer data based on the question. For example, it might generate an answer such as, "Currently, there are 10 units of this product in stock."
[1258] Step 8:
[1259] The server converts the generated response data into speech.
[1260] Input: Generated response data, selected respondent profile.
[1261] Output: Audio data.
[1262] Specific operation: Use Amazon Polly or Google Cloud Text-to-Speech to convert text data into speech data and adjust the tone and speed of the speech.
[1263] Step 9:
[1264] The server provides the user with an audio response.
[1265] Input: Audio data.
[1266] Output: Voice response to the user.
[1267] Specific operation: The generated audio data is delivered to the user in real time via a device (smart glasses or earphones). The user can experience a natural conversation.
[1268] Step 10:
[1269] The server saves the interaction history.
[1270] Input: User's question, generated answer, date and time of interaction, perceived emotion.
[1271] Output: Saved call history.
[1272] Specific actions: Use Microsoft SQL Server or MongoDB to record this information in a database, enabling quick responses to future queries.
[1273] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1274] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1275] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1276] [Fourth Embodiment]
[1277] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1278] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1279] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1280] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1281] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1282] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1283] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1284] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1285] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1286] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1287] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1288] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1289] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1290] The system of the present invention includes various means for providing natural and flexible responses when an end user calls a call center. The system configuration is as follows:
[1291] 1. Recognition of incoming call
[1292] When a user calls the call center, the server automatically recognizes the call and initiates the conversation. The technologies used here are general telephone systems and VoIP (Voice over Internet Protocol) technology.
[1293] 2. Collection and conversion of audio data
[1294] When a user begins speaking, the server collects their voice in real time. This voice data is converted into text using an ASR (Automatic Speech Recognition) engine. For example, if a user asks, "Please tell me the opening hours of the supermarket," the voice data is converted into the text "Please tell me the opening hours of the supermarket."
[1295] 3. Analysis of the inquiry content
[1296] The server analyzes the converted text data to understand the user's question. The technology used here is NLP (Natural Language Processing), which classifies the question into specific categories. For example, categories such as "business hours" or "return policy."
[1297] 4. Selection of a support person based on user requirements
[1298] When a user makes a specific request (for example, "I want an explanation in a female voice" or "I want an explanation that even a primary school student can understand"), the server selects an appropriate AI respondent based on that request. A specific voice profile is used for voice selection.
[1299] 5. Generating response data
[1300] Based on the question, the server retrieves relevant information from the database and constructs answer data using generative AI. For example, if the question is about opening hours, the answer "The supermarket's opening hours are from 9 am to 8 pm" will be generated.
[1301] 6. Speech synthesis
[1302] The generated response data is converted into speech format using a Text-to-Speech (TTS) engine. Natural-sounding speech is generated based on the selected voice profile.
[1303] 7. Providing a response
[1304] The generated audio data is played back to the user in real time during the call. This allows the user to receive responses in a natural conversational style.
[1305] 8. Record of interaction history
[1306] The server stores the interaction history in a database. This includes the user's question, the answer provided, and the date and time of the interaction. This information is used to respond quickly and accurately to future inquiries.
[1307] Specific example
[1308] Example 1: Request for a brief explanation in a female voice.
[1309] User A makes a phone call and asks, "Could you tell me the supermarket's opening hours? Could you please explain it in a way that even a primary school student could understand, using a woman's voice?" The server recognizes this call, converts the user's voice into text, analyzes the question, and detects a "woman's voice" as the request. The server retrieves the opening hours information from its database, and a generative AI generates the answer, "The supermarket's opening hours are from 9 am to 8 pm," which is then converted into a woman's voice and provided to User A.
[1310] Example 2: Harassment prevention measures
[1311] If a regular user B asks the same question hundreds of times every month, the server recognizes the call and converts the user's voice into text. Because the past question history is stored in the database, the server can refer to the already accumulated information and quickly and accurately generate and provide the answer, "Our business hours are from 9 am to 8 pm." This significantly reduces the burden on staff.
[1312] By following these steps, the present invention provides a system that can flexibly respond to diverse user demands and improve the efficiency of call centers.
[1313] The following describes the processing flow.
[1314] Step 1:
[1315] The user makes a phone call. The user dials a phone number and is connected to the call center.
[1316] Step 2:
[1317] The server recognizes the incoming call. The server detects the incoming signal and automatically initiates the call.
[1318] Step 3:
[1319] The server collects audio data. When the user starts speaking, the server collects audio data in real time.
[1320] Step 4:
[1321] The server converts the audio data into text data. The collected audio data is converted into text format using the ASR engine.
[1322] Step 5:
[1323] The server analyzes the text data. Using NLP (Neuro-Linguistic Programming) techniques, it analyzes the user's questions from the text data and identifies categories.
[1324] Step 6:
[1325] The server detects user requests. If there are user requests (e.g., "female voice," "make it easy enough for elementary school children to understand"), the server analyzes and recognizes them.
[1326] Step 7:
[1327] The server selects a responder. Based on the detected request, it selects the appropriate AI responder and voice profile.
[1328] Step 8:
[1329] The server generates the answer data. The generative AI retrieves information related to the question from the database and constructs the answer.
[1330] Step 9:
[1331] The server converts the response data into speech. A TTS engine is used to convert the text-based responses into speech using a selected voice profile.
[1332] Step 10:
[1333] The server provides the user with an audio response. The generated audio data is played back to the user in real time.
[1334] Step 11:
[1335] The server saves the interaction history. It records information such as the user's question, the answer provided, and the date and time of the interaction in a database.
[1336] (Example 1)
[1337] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1338] Traditional call center systems struggled to respond to user inquiries in real time, requiring a significant amount of manual work. This resulted in low efficiency in handling inquiries, increased susceptibility to human error, and inconsistent response quality. Furthermore, flexibly addressing specific requests (e.g., responses in a specific gender's voice or tailored to a specific level of understanding) was difficult, potentially leading to decreased user satisfaction. Additionally, the inability to effectively utilize past interaction history meant that responses to the same question were inconsistent, hindering efficient operation.
[1339] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1340] In this invention, the server includes means for recognizing incoming calls from users, means for converting user voice into text data, means for analyzing the user's questions from the text data, means for selecting a responder based on the user's request, means for generating answer data to the questions using natural language generation technology, means for converting the generated answer data into speech, means for providing the spoken answers to the user, and means for storing the interaction history. This makes it possible to respond quickly and flexibly to diverse user requests and improve the operational efficiency of the call center.
[1341] (Definitions of important words)
[1342] "Means for recognizing incoming calls" refers to a function that detects when a call has started when receiving a call from a user and initiates the appropriate processing.
[1343] "Methods for converting speech to text data" refers to technologies that collect spoken audio from users in real time and convert it into text data. Specifically, this utilizes ASR (Automatic Speech Recognition) technology.
[1344] "Means for analyzing question content" refers to a function that analyzes the converted text data to identify and classify the content of the user's question. It primarily uses NLP (Natural Language Processing) technology.
[1345] The "means of selecting a responder" refer to a function that chooses the optimal voice response profile based on the user's requests. This includes selecting a voice profile tailored to a specific gender or level of comprehension.
[1346] "Means for generating response data" refers to a function that collects relevant information based on the user's question and creates an appropriate response using natural language generation technology.
[1347] "Means for converting response data into speech" refers to a function that converts the generated text-based response data into natural-sounding speech using TTS (text-to-speech) technology.
[1348] "Means of providing users with voiced responses" refers to a function that provides generated voice data to users in real time via telephone.
[1349] "Means for saving interaction history" refers to a function that saves historical data such as the user's questions, the answers provided, and the date and time of the call in a database, for use in handling future inquiries.
[1350] "Natural language generation technology" is a technique that generates meaningful natural language sentences based on given input data. It commonly utilizes generative AI models.
[1351] A "prompt" is text data input to a generative AI model, containing specific instructions or questions for the model.
[1352] The system of this invention is designed to provide natural and flexible responses when end users call a call center. The system configuration is as follows:
[1353] When a user calls the call center, the server automatically recognizes the call and initiates the conversation. At this stage, a standard telephone system or VoIP (Voice over Internet Protocol) technology is used.
[1354] Next, when the user begins to speak, the server collects the user's voice in real time. This is done using microphones and acoustic detection devices. This voice data is converted into text format using an ASR (Automatic Speech Recognition) engine. Specific technologies include Google Cloud Speech-to-Text and IBM Watson Speech to Text. For example, if the user asks, "Please tell me the opening hours of the supermarket," the voice data is converted into text data that reads, "Please tell me the opening hours of the supermarket."
[1355] The converted text data is sent to a server and analyzed using an NLP (Natural Language Processing) engine. Technologies include Google Cloud Natural Language API and Microsoft Azure Text Analytics. The analysis categorizes user inquiries into specific categories (e.g., "Business Hours," "Return Policy").
[1356] Next, when a user makes a specific request (for example, "I want an explanation in a woman's voice" or "Please explain it in a way that even a primary school student can understand"), the server analyzes the request and selects an appropriate voice profile. The server accesses an internal database and selects a voice responder based on the optimal voice profile for the request.
[1357] The answer data for the questions is generated by retrieving relevant information from a database and using generative AI. OpenAI GPT-3 or similar natural language generation technologies are used for this generative AI. For example, the answer data "The supermarket's operating hours are from 9 AM to 8 PM" might be generated.
[1358] The generated response data is converted into speech format using a Text-to-Speech (TTS) engine. Amazon Polly and Google Cloud Text-to-Speech are used as TTS engines. The converted speech data provides natural-sounding speech based on the selected speech profile.
[1359] Ultimately, the generated audio data is played back to the user in real time via the call, allowing the user to receive responses in a natural conversational style.
[1360] The interaction history is stored in a database by the server. This includes the user's question, the answer provided, and the date and time of the interaction. This information is used to respond quickly and accurately to future inquiries.
[1361] Specific example
[1362] Example 1: Request for a brief explanation in a female voice.
[1363] User A makes a phone call and asks, "Could you tell me the supermarket's opening hours? Could you please explain it in a way that even a primary school student could understand, using a woman's voice?" The server recognizes this call and converts the user's voice into text data. Next, it analyzes the question and detects a "woman's voice" as the request. The server retrieves the opening hours information from the database, and a generative AI generates the answer, "The supermarket's opening hours are from 9 am to 8 pm." This is then converted into a woman's voice and provided to User A.
[1364] Example 2: Harassment prevention measures
[1365] If a regular user B asks the same question hundreds of times every month, the server recognizes the incoming call and converts the user's voice into text. Because past question history is stored in the database, the server can refer to the already accumulated information and quickly generate and provide an answer such as, "Our business hours are from 9 am to 8 pm." This significantly reduces the burden on staff.
[1366] Example of a prompt
[1367] "Could you please tell me about the supermarket's opening hours? Could you explain it in a way that even a primary school student could understand?"
[1368] "Our business hours are from 9 AM to 8 PM."
[1369] "Do you have any other questions?"
[1370] In this way, the system can respond quickly and flexibly to diverse user requests, thereby improving the operational efficiency of the call center.
[1371] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1372] Step 1:
[1373] When a user calls the call center, the server recognizes the incoming call and initiates the conversation. At this stage, the call is connected to the call center's system using VoIP technology. In other words, a call commencement signal (output) is generated in response to the user's phone call (input). The server logs the user's number and the time the call started.
[1374] Step 2:
[1375] When a user begins speaking, the server collects their voice in real time. This voice data (input) is sent to an ASR engine and converted into text format (output). Specifically, voice data is collected using microphones or acoustic detection devices and sent to an ASR engine (such as Google Cloud Speech-to-Text). The converted text data is temporarily stored for subsequent processing.
[1376] Step 3:
[1377] The server sends the converted text data to the NLP engine, which analyzes the user's question. This text data (input) is analyzed by the NLP engine, and the question is classified into a specific category (output). Technologies used here include Google Cloud Natural Language API and Microsoft Azure Text Analytics. The category information is also used to generate the response in the next step.
[1378] Step 4:
[1379] The server detects specific requests (e.g., "female voice," "explain it clearly") from the converted text data. This text data (input) contains the user's requests, and an appropriate voice profile is selected based on the analysis results (output). The server records this information in its internal database.
[1380] Step 5:
[1381] The server retrieves relevant information from the database based on the analyzed question. This relevant information (output) corresponding to the question (input) is converted into answer data using a generative AI model (e.g., OpenAI GPT-3). The generated answer data is temporarily stored because it will be converted into speech in the next step.
[1382] Step 6:
[1383] The server sends the generated response data to the TTS engine, where it is converted into speech format. This text-based response data (input) is then converted into speech data (output) by the TTS engine (such as Amazon Polly or Google Cloud Text-to-Speech). The server generates natural-sounding speech based on the selected speech profile.
[1384] Step 7:
[1385] The server sends the generated audio data to the call channel and plays it back to the user in real time. This audio data (input) is provided to the user through the call channel (output). The server confirms that the response has been provided and continues or ends the call with the user as needed.
[1386] Step 8:
[1387] The server stores call history, including call content, generated responses, and call date and time, in a database. This call history data (input) is recorded in the database and used for future inquiries (output). This helps in quickly generating answers to the same question again and improves operational efficiency.
[1388] (Application Example 1)
[1389] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1390] Current voice assistant systems in autonomous vehicles lack the ability to provide natural and flexible responses in passenger interaction. In particular, the systems are inadequate in handling situations where passengers request answers tailored to specific voice qualities or levels of understanding, or when they repeatedly ask the same questions. This leads to decreased passenger satisfaction and compromises in-vehicle comfort.
[1391] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1392] In this invention, the server includes means for recognizing a user entering the vehicle, means for converting the user's voice into text data, means for analyzing the user's question from the text data, means for selecting a responder based on the user's request, means including a generative AI model for generating answer data to the question, means for converting the generated answer data into speech, means for providing the spoken answer to the user, and means for storing the response history. This makes it possible to provide natural and flexible responses in real time to various requests from passengers in an autonomous vehicle.
[1393] "Means for recognizing a user entering the vehicle" refers to a device or system used in an autonomous vehicle to recognize when a passenger has boarded the vehicle.
[1394] "Means for converting user speech into text data" refers to technology or equipment for converting speech spoken by passengers into text data in real time.
[1395] "Means for analyzing user questions from text data" refers to a system or technology for analyzing passengers' questions based on converted text data and understanding their intent.
[1396] "Means for selecting a responder based on user requests" refers to a system or technology for selecting an appropriate voice profile or responder in accordance with the content of a passenger's request.
[1397] "Means including a generative AI model that generates response data to a question" refers to a system or technology that includes an artificial intelligence model for generating appropriate answers to passengers' questions.
[1398] "Means for converting generated response data into audio" refers to a technology or device for converting generated text-formatted response data into audio data.
[1399] "Means of providing users with voiced responses" refers to a device or system that plays voice data in real time and provides responses to passengers.
[1400] "Means for saving interaction history" refers to a system or technology that records the history of interactions with passengers and uses it to handle future inquiries.
[1401] The system of the present invention realizes a voice recognition assistant in an autonomous vehicle by, for example, following the procedure described below. By utilizing specific hardware and software, it is possible to interact with the user in a natural and flexible manner.
[1402] System configuration and program processing
[1403] 1. Means for recognizing a vehicle entering the premises from a user.
[1404] The server uses hardware such as sensors and cameras to recognize when a user has boarded the vehicle. Machine learning models may also be used for this purpose.
[1405] 2. Means for converting user speech into text data
[1406] The server analyzes the audio data collected through the microphone in real time. The audio data is converted into text data using an ASR (Automatic Speech Recognition) engine such as AWS Transcribe or Google Cloud Speech-to-Text API.
[1407] 3. Means for analyzing user questions from text data
[1408] The server analyzes the converted text data to understand the intent behind the user's question. This process utilizes NLP (Natural Language Processing) technologies such as SpaCy and Google Cloud Natural Language API.
[1409] 4. Means for selecting a responder based on user requests
[1410] The server selects a specific voice profile and response method. For example, it might use Azure's Cognitive Services Speech API to select a voice quality and speaking style that suits the user's request.
[1411] 5. Means including a generative AI model that generates response data to the content of a question.
[1412] The server uses a generative AI model (e.g., GPT-4) to generate response data corresponding to the query. It uses the OpenAI API to generate appropriate responses based on the prompt.
[1413] 6. Means for converting generated response data into speech.
[1414] The server converts the generated text-based response data into speech using a Text-to-Speech (TTS) engine. Google TTS and Amazon Polly are among the engines used for this process.
[1415] 7. Means of providing users with voiced responses
[1416] The server provides the converted audio data to the user in real time through the in-car speaker system.
[1417] 8. Means for saving interaction history
[1418] The server stores the user interaction history in a database. Services such as Amazon RDS and Google Firestore are used to facilitate future inquiries.
[1419] Specific example
[1420] One afternoon, passenger A gets into a self-driving vehicle and asks, "Please tell me the next gas station." In this case, the system operates as follows:
[1421] The server recognizes passenger A's voice and converts the audio data into text.
[1422] The converted text data is analyzed to understand that it is a question asking for the location of a gas station.
[1423] A gentle male voice is selected as the appropriate voice quality, and the generating AI model responds, "The next gas station is 2 kilometers ahead on the right."
[1424] This text is converted into audio data by the TTS engine and played back in real time through the car's speakers.
[1425] Finally, save this question and answer exchange to the database.
[1426] Example of a prompt
[1427] The following prompt is used in a situation where a passenger requests, "Tell me the next gas station," and the system responds, "The next gas station is 2 kilometers ahead on the right":
[1428] User: "Please tell me the next gas station."
[1429] System: "The next gas station is 2 kilometers ahead on the right."
[1430] The system of the present invention, through these procedures, enables highly natural and flexible interaction with the user in an autonomous vehicle.
[1431] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1432] Step 1:
[1433] The server recognizes when a user boards the vehicle. Specifically, it uses data from on-board sensors and cameras to detect entry. The server analyzes this input data to confirm the presence of a passenger.
[1434] Step 2:
[1435] When a user begins speaking in the vehicle, the server collects audio data through the microphone. This audio data is then input to the server.
[1436] Step 3:
[1437] The server converts the collected audio data into text data using AWS Transcribe or the Google Cloud Speech-to-Text API. The input audio data is passed through the ASR engine, and the converted data in text format is output.
[1438] Step 4:
[1439] The server analyzes the generated text data. Specifically, it uses NLP technologies such as SpaCy and Google Cloud Natural Language API to analyze the text data and identify the user's question. It processes the input text data and generates the intent and category of the question as output.
[1440] Step 5:
[1441] The server selects the appropriate voice profile and speaker based on the user's request. This is done by using the Azure Cognitive Services Speech API to select a voice profile according to the user's request (e.g., "a calm male voice"). The user's request data is used as input, and the selected voice profile is obtained as output.
[1442] Step 6:
[1443] The server generates answer data in response to the user's questions. Using the OpenAI GPT-4 model, it generates appropriate answers based on the text data input. This generated answer data is then output.
[1444] Step 7:
[1445] The server converts the generated text-based response data into audio data. It uses a TTS engine such as Google TTS or Amazon Polly to convert the text data to audio data. Text response data is used as input, and audio data is generated as output.
[1446] Step 8:
[1447] The server plays the generated audio data in real time through the car's speakers. This allows the user to receive an audio response. The audio data is sent to the speakers, and an audio response is output.
[1448] Step 9:
[1449] The server stores the user interaction history in a database. Amazon RDS or Google Firestore is used to store question content and answer data. Interaction data is received as input and saved to the database as output.
[1450] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1451] The system of the present invention includes various means for recognizing the user's emotions and providing natural and flexible responses when an end user makes a call to a call center. The system configuration is as follows:
[1452] 1. Recognition of incoming call
[1453] When a user calls the call center, the server automatically recognizes the call and initiates the conversation. The technologies used here are general telephone systems and VoIP (Voice over Internet Protocol) technology.
[1454] 2. Collection and conversion of audio data
[1455] When a user begins speaking, the server collects their voice in real time. This voice data is converted into text using an ASR (Automatic Speech Recognition) engine. For example, if a user asks, "Please tell me the opening hours of the supermarket," the voice data is converted into the text "Please tell me the opening hours of the supermarket."
[1456] 3. Recognition of emotions
[1457] The collected audio data is analyzed through an emotion engine. The emotion engine recognizes the user's emotions (e.g., joy, anger, sadness, surprise) based on factors such as tone, speed, and pitch of the voice. This process determines the user's current emotional state.
[1458] 4. Analysis of the inquiry content
[1459] The server analyzes the converted text data to understand the user's question. The technology used here is NLP (Natural Language Processing), which classifies the question into specific categories. For example, categories such as "business hours" or "return policy."
[1460] 5. Selection of a representative based on user requirements
[1461] When a user makes a specific request (for example, "I want an explanation in a woman's voice" or "I want an explanation that even a primary school student can understand"), the server selects an appropriate AI respondent based on that request. A specific voice profile is used for voice selection. Furthermore, the appropriate response is selected based on the emotional state recognized by the emotion engine.
[1462] 6. Generating response data
[1463] Based on the question, the server retrieves relevant information from the database and constructs answer data using generative AI. For example, if the question is about opening hours, the answer "The supermarket's opening hours are from 9 am to 8 pm" will be generated.
[1464] 7. Speech synthesis
[1465] The generated response data is converted into speech format using a Text-to-Speech (TTS) engine. Natural-sounding speech is generated based on the selected voice profile. The tone and speed of the speech are also adjusted based on the emotions recognized by the emotion engine.
[1466] 8. Providing a response
[1467] The generated audio data is played back to the user in real time during the call. This allows the user to receive responses in a natural conversational style. A key feature is that the emotion engine provides voice responses that are tailored to the user's emotions.
[1468] 9. Record of interaction history
[1469] The server stores interaction history in a database. This includes the user's question, the answer provided, the date and time of the interaction, and the perceived emotional state. This information is used to respond quickly and accurately to future inquiries.
[1470] Specific example
[1471] Example 1: Request for a simple explanation in a woman's voice and emotional recognition.
[1472] User A makes a phone call and asks, "Could you tell me the supermarket's opening hours? Could you please explain it in a way that even a primary school student could understand, using a woman's voice?" The server recognizes this call, converts the user's voice into text, analyzes the question, and detects a "woman's voice" as the request. Furthermore, because the user's voice tone is calm, the emotion engine recognizes a "calm" state. The server retrieves the opening hours information from the database, and the generative AI generates the answer, "The supermarket's opening hours are from 9 am to 8 pm," which is then converted into a woman's voice and delivered to User A in a calm tone.
[1473] Example 2: Harassment prevention and emotional recognition
[1474] If a regular user B asks the same question hundreds of times every month, the server recognizes the incoming call and converts the user's voice into text. Since the past question history is stored in the database, the server refers to the already accumulated information and quickly and accurately generates and provides the answer, "Our business hours are from 9 am to 8 pm." Furthermore, if user B's voice is agitated, the emotion engine recognizes a state of "anger" and provides a response in a calm tone. This significantly reduces the burden on staff while enabling responses that are considerate of the user's emotions.
[1475] By following these steps, the present invention provides a system that can flexibly respond to the diverse needs and emotions of users and improve the efficiency of call centers.
[1476] The following describes the processing flow.
[1477] Step 1:
[1478] The user makes a phone call. The user dials a phone number to the call center and is connected.
[1479] Step 2:
[1480] The server recognizes the incoming call. The server detects the incoming signal and automatically initiates the call.
[1481] Step 3:
[1482] The server collects audio data. When the user starts speaking, the server collects this audio data in real time.
[1483] Step 4:
[1484] The server converts the audio data into text data. The collected audio data is converted into text format using the ASR engine.
[1485] Step 5:
[1486] The server analyzes the text data. Using NLP (Neuro-Linguistic Programming) techniques, it analyzes the user's questions from the text data and categorizes them into specific categories.
[1487] Step 6:
[1488] The server uses an emotion engine to recognize the user's emotions. It analyzes the tone, speed, and pitch of the voice to identify the user's emotional state.
[1489] Step 7:
[1490] The server detects user requests. For example, if there is a request to "explain in a female voice," the server recognizes that request. Emotional states are also taken into consideration.
[1491] Step 8:
[1492] The server selects a responder. Based on the detected request and emotional state, it chooses the most appropriate AI responder and voice profile.
[1493] Step 9:
[1494] The server generates the answer data. The generative AI retrieves information related to the question from the database and constructs the answer.
[1495] Step 10:
[1496] The server converts the response data into speech. Using a TTS engine, the text-based responses are converted into speech using a selected voice profile. The voice tone and speed are adjusted based on the results of the emotion engine.
[1497] Step 11:
[1498] The server provides the user with an audio response. The generated audio data is played back to the user in real time.
[1499] Step 12:
[1500] The server saves the interaction history. It records the user's question, the answer provided, the date and time of the interaction, and the perceived emotional state in a database.
[1501] Step 13:
[1502] The server analyzes the interaction history and makes improvements as needed. Based on the stored history, data analysis is performed to improve the system's response quality.
[1503] (Example 2)
[1504] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1505] Traditional call center systems were unable to analyze user voices in real time and provide flexible responses tailored to their emotional state. Furthermore, selecting the appropriate responder based on specific requests and generating quick answers using past interaction history were difficult. Therefore, there is a need to improve both user satisfaction and response efficiency. Moreover, a system capable of responding to diverse user requests and emotions is required.
[1506] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1507] In this invention, the server includes means for recognizing incoming calls from users, means for collecting the user's voice, means for converting the collected user's voice into text data, means for recognizing the user's emotional state from the text data, means for analyzing the user's question content from the text data, means for selecting a responder based on the analyzed question content and the recognized emotional state, means for generating answer data for the question content, means for converting the generated answer data into voice, means for providing the voiced answer to the user, and means for storing the interaction history. This enables flexible responses according to the user's emotional state, selection of a responder based on specific requests, and rapid and accurate response generation utilizing past interaction history.
[1508] "Incoming call recognition means" refers to a device or method for automatically detecting a call from a user and initiating a call.
[1509] "Voice collection means" refers to a device or method for acquiring the voice spoken by a user during a phone call in real time.
[1510] "Character conversion means" refers to a device or method for converting acquired audio data into character data.
[1511] "Emotion recognition means" refers to a device or method for analyzing and recognizing a user's emotional state from text data or audio data.
[1512] "Question analysis means" refers to a device or method for understanding and analyzing the content of a user's question from text data.
[1513] "Respondent selection means" refers to a device or method for selecting an appropriate respondent based on the analyzed question content and the recognized emotional state.
[1514] "Answer generation means" refers to a device or method for generating appropriate answer data in response to a user's question.
[1515] "Voice conversion means" refers to a device or method for converting generated response data into voice data.
[1516] "Answer provision means" refers to a device or method for providing answer data converted into audio format to a user.
[1517] "Response history storage means" refers to a device or method for recording and storing the content of a user's questions and response history.
[1518] Modes for carrying out the invention
[1519] The system according to this invention recognizes the user's emotions when they call a call center and provides a natural and flexible response. This system uses the following hardware and software.
[1520] Hardware and software
[1521] Telephone system or VoIP technology: Used to recognize incoming calls and initiate conversations.
[1522] ASR (Automatic Speech Recognition) engine: Collects speech data and converts it into text data.
[1523] Emotion Engine: Used to recognize the user's emotional state from text and audio data.
[1524] NLP (Natural Language Processing) technology: Used to analyze user questions from text data.
[1525] Database: Used to store user questions and response history.
[1526] Generative AI: Used to generate answer data in response to questions.
[1527] TTS (Text-to-Speech) engine: Used to convert generated response data into speech.
[1528] Specific processing of the program
[1529] 1. Recognition of incoming call
[1530] When a user calls the call center, the server automatically recognizes the incoming call and initiates the conversation. This process utilizes telephone systems and VoIP technology.
[1531] 2. Collection and conversion of audio data
[1532] When a user begins speaking, the server collects their voice in real time. The collected voice data is input into the ASR engine and converted into text format. For example, if a user asks, "Please tell me the opening hours of the supermarket," the ASR engine converts that voice into the text, "Please tell me the opening hours of the supermarket."
[1533] 3. Recognition of emotions
[1534] The server inputs the converted text data into the emotion engine, which analyzes elements such as tone, speed, and pitch of the voice. Based on this analysis, the emotion engine recognizes the user's emotional state (e.g., happy, angry, calm).
[1535] 4. Analysis of the inquiry content
[1536] The server analyzes the text data using NLP (Neuro-Linguistic Programming) techniques. In this process, the server classifies the text data into specific categories (e.g., "business hours," "return policy," etc.) and understands the specific content of the questions.
[1537] 5. Selection of a representative based on user requirements
[1538] When a user makes a specific request, the server selects an appropriate AI respondent based on that request. For example, in response to a request to "explain in a female voice," the server selects an AI respondent with a female voice profile, taking into account the user's emotional state when making the selection.
[1539] 6. Generating response data
[1540] The server retrieves relevant information from the database and uses generative AI to construct response data. For example, it might generate a response like, "The supermarket's opening hours are from 9 AM to 8 PM." This response is generated to appropriately address the user's question.
[1541] 7. Speech synthesis
[1542] The server inputs the generated response data into the TTS engine and converts it into speech format. Based on the selected voice profile, natural-sounding speech is generated. At this time, the tone and speed of the speech are also adjusted based on the emotional state recognized by the emotion engine.
[1543] 8. Providing a response
[1544] The server provides the generated audio data to the user in real time. Specifically, the audio data is played back during the call. This allows the user to receive responses in a natural conversational style.
[1545] 9. Record of interaction history
[1546] The server saves the interaction history to a database. The saved data includes the user's question, the answer provided, the date and time of the interaction, and the perceived emotional state. This information is used to respond quickly and accurately to future inquiries.
[1547] Specific example
[1548] Example 1: Request for a simple explanation in a woman's voice and emotional recognition.
[1549] User A calls and asks, "Could you tell me the supermarket's opening hours? Could you please explain it in a way that even a primary school student could understand, in a woman's voice?" The server recognizes this call, converts the user's voice to text, analyzes the question, and detects that it is requesting an explanation in a woman's voice. Furthermore, because the user's voice tone is calm, the emotion engine recognizes that the user is in a "calm" state. The server retrieves the opening hours information from the database, and the generative AI generates the answer, "The supermarket's opening hours are from 9 am to 8 pm," which is then converted into a woman's voice and delivered to User A in a calm tone.
[1550] Example of a prompt
[1551] "A user asked, 'Could you please tell me about supermarket opening hours? Could you explain it in a way that even a primary school student could understand, using a woman's voice?' The user's tone is calm. Please generate an answer and explain it calmly in a woman's voice."
[1552] Example 2: Harassment prevention and emotional recognition
[1553] Regular user B asks in an agitated voice, "What are your opening hours?" The server refers to past question history and quickly and accurately generates the answer, "Your opening hours are from 9 am to 8 pm." Furthermore, recognizing the agitated tone of user B, the emotion engine recognizes a state of "anger" and provides a response in a calm tone.
[1554] Example of a prompt
[1555] "User B asked in an excited voice, 'What are your opening hours?' Based on past question history, we know this same question has been asked multiple times. Please generate a calm and collected response."
[1556] Through this invention, we can provide a system that can flexibly respond to the diverse needs and emotions of users and improve the efficiency of call centers.
[1557] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1558] Specific processing flow of the program
[1559] Step 1:
[1560] Recognition of incoming call
[1561] The server automatically recognizes incoming calls when a user dials the call center.
[1562] Input: Phone call from a user (VoIP data or phone system data).
[1563] Processing: The server receives the call initiation signal via the telephone system or VoIP technology.
[1564] Output: Call start time and user call information (phone number, call start time, etc.).
[1565] Step 2:
[1566] Audio data collection and conversion
[1567] When a user begins speaking, the server collects their voice in real time. The collected voice data is then converted into text format by the ASR engine.
[1568] Input: User's voice data.
[1569] Processing: The server collects the audio data and the ASR engine converts the audio to text. For example, if the user says, "Please tell me the opening hours of the supermarket."
[1570] Output: Text data (e.g., "Please tell me about the supermarket's opening hours").
[1571] Step 3:
[1572] Recognition of emotions
[1573] The server inputs the converted text data into an emotion engine, which analyzes the tone, speed, pitch, and other characteristics of the speech.
[1574] Input: Text data and audio data.
[1575] Processing: The server uses an emotion engine to analyze text and audio data and recognize the user's emotions. For example, if the tone of voice is calm, it will be recognized as "calm."
[1576] Output: Emotional state data (e.g., "Calm").
[1577] Step 4:
[1578] Analysis of inquiry content
[1579] The server analyzes the text data using NLP (Neuro-Linguistic Programming) technology to understand the user's question.
[1580] Input: Text data.
[1581] Processing: The server uses NLP (Neuro-Linguistic Programming) technology to analyze text data and classify the questions into specific categories. For example, "business hours" or "return policy."
[1582] Output: Category data of the question content (e.g., "Business Hours").
[1583] Step 5:
[1584] Selecting a responder based on user requirements
[1585] When a user makes a specific request, the server selects an appropriate AI responder based on that request and the perceived emotional state.
[1586] Input: Question category data, emotional state data, user request data (e.g., "Explain in a female voice").
[1587] Processing: The server selects an appropriate AI responder with a voice profile based on the request and emotion. For example, it might select a voice profile with a "female voice".
[1588] Output: Selected AI responder data (e.g., "female voice").
[1589] Step 6:
[1590] Generating response data
[1591] The server retrieves relevant information from the database and uses generative AI to construct response data.
[1592] Input: Category data of the question content, information in the database.
[1593] Processing: The server uses generative AI to generate accurate answer data based on the question. For example, if the question is about opening hours, it will generate the answer "The supermarket's opening hours are from 9 am to 8 pm."
[1594] Output: Response data (Example: "The supermarket's operating hours are from 9 AM to 8 PM").
[1595] Step 7:
[1596] Speech synthesis
[1597] The server inputs the generated response data into the TTS engine and converts it into audio format.
[1598] Input: Response data, selected AI responder data.
[1599] Processing: The server uses a TTS engine to convert response data into audio data. During this process, the tone and speed of the voice are also adjusted based on the emotional state.
[1600] Output: Audio data.
[1601] Step 8:
[1602] Providing a response
[1603] The server provides the generated audio data to the user in real time.
[1604] Input: Audio data.
[1605] Processing: The server plays the audio data through the call.
[1606] Output: Receipt of the user's response in audio format.
[1607] Step 9:
[1608] Record of interaction history
[1609] The server saves the interaction history to a database.
[1610] Input: User's question, provided answer, date and time of interaction, perceived emotional state, call information.
[1611] Processing: The server records and stores this data in the database.
[1612] Output: Saved call history data.
[1613] (Application Example 2)
[1614] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1615] Traditional call center and virtual store systems could only provide standard responses to user inquiries, making it difficult to respond flexibly to users' emotional states or specific requests. Furthermore, they lacked sufficient functionality to efficiently utilize past inquiry history. Therefore, there were limitations in addressing situations where improving the user experience and reducing staff workload were required.
[1616] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for recognizing incoming calls from users, means for converting user voice into text data, means for analyzing the user's question content from the text data, means for selecting a responder based on the user's request, means for generating answer data for the question content, means for converting the generated answer data into voice, means for providing the voiced answer to the user, means for recognizing the user's emotions, means for adjusting the response based on the recognized emotions, and means for saving the response history. This enables quick and efficient responses by utilizing past inquiry history while flexibly responding to the user's emotions and specific requests.
[1617] "Means for recognizing incoming calls" refers to a function that automatically detects voice inquiries from users and notifies the system.
[1618] "A means of converting user voice into text data" refers to a function that converts input voice data into text format in real time.
[1619] "Methods for analyzing user inquiries from text data" refers to a function that converts speech into text and then analyzes that text data to understand the content of the user's inquiry.
[1620] "Means for selecting a responder based on user requests" refers to a function that selects a responder who meets specific criteria (e.g., gender, level of understanding) when a user requests them.
[1621] "Means for generating response data to questions" refers to a function that generates appropriate answers based on the user's inquiry.
[1622] "Means for converting generated response data into audio" refers to a function that converts generated text-based response data into audio format.
[1623] "Means of providing users with voiced responses" refers to a function that transmits voiced responses to users in real time.
[1624] "Means of recognizing user emotions" refers to a function that determines the user's emotional state through the analysis of voice data.
[1625] "Means of adjusting responses based on recognized emotions" refers to a function that adjusts the content and tone of responses to match the user's emotional state.
[1626] "Means for saving interaction history" refers to a function that records and saves past inquiries, the answers provided, the date and time of the interaction, and the perceived emotional state.
[1627] This invention is a system that provides flexible responses to user inquiries in a virtual store, tailored to the user's emotional state. The server performs response processing when a user makes an inquiry by voice using the following means.
[1628] First, when a user makes a voice inquiry within the virtual store, the device (e.g., smart glasses or a smartphone) collects the audio in real time and sends it to a server. The server recognizes the incoming call using a general speech recognition system. Technologies used include the Google Cloud Speech-to-Text API.
[1629] Next, the server converts the collected audio data into text data. This also uses the Google Cloud Speech-to-Text API. This conversion changes the audio data into text format.
[1630] Next, the server recognizes the user's emotions based on the text data. The emotion engine uses IBM Watson Tone Analyzer or Microsoft Azure's Text Analytics API. This allows it to determine the user's emotional state (e.g., joy, anger, sadness, surprise) from the audio data.
[1631] Furthermore, the server analyzes the text data and classifies the user's questions into specific categories (e.g., product details, return procedures, etc.). This analysis uses NLP technologies such as Google Cloud Natural Language API and SpaCy. Subsequently, it selects the appropriate respondent (e.g., a voice of a specific gender, a concise explanation, etc.) based on the user's request.
[1632] Generative AI models are used to generate response data. Specifically, OpenAI GPT-4 can be used. The server retrieves relevant information from the database and constructs response data using the generative AI model. For example, if a user asks, "How much stock do you have of this product?", the generative AI model will generate a response such as, "We currently have 10 units of this product in stock."
[1633] The generated response data is converted into speech format using a TTS engine such as Amazon Polly or Google Cloud Text-to-Speech. The server also adjusts the tone and speed of the speech based on the recognized emotions.
[1634] Finally, the transcribed response data is provided to the user in real time via the device. Users can enjoy a natural conversational experience through smart glasses or earphones.
[1635] The interaction history is stored in a database such as Microsoft SQL Server or MongoDB. This allows for quick and efficient responses to the same question again.
[1636] As a concrete example, consider a scenario where a customer inquires, "How much stock do you have of this item?" in a virtual store. In this case, the system recognizes the user's emotions and provides a calm tone of voice if the user is "excited," and a detailed and polite response if the user is "confused."
[1637] Examples of prompt statements include the following formats:
[1638] User inquiry: "How much stock do you have of this product?"
[1639] Perceived emotions: "excitement," "confusion"
[1640] Tone of response: "Calm tone," "Detailed and polite."
[1641] Using a generative AI model: "Currently, there are 10 of this item in stock."
[1642] By combining these methods, we can provide a system that can respond quickly and efficiently while flexibly responding to user emotions and specific requests.
[1643] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1644] Step 1:
[1645] The device recognizes the user's voice input.
[1646] Input: User voice inquiry.
[1647] Specific operation: In a virtual store, a user asks a question by voice, for example, "How much stock do you have of this item?" The device (smart glasses or smartphone) uses its built-in microphone to collect the voice and sends it to the server as digital audio data.
[1648] Step 2:
[1649] The server recognizes the incoming call from the user.
[1650] Input: Digital audio data.
[1651] Specific operation: The server uses a speech recognition system to confirm incoming calls and initiate appropriate processing. Here, a system is in place to process user inquiries in real time using a general speech recognition system.
[1652] Step 3:
[1653] The server converts the user's voice into text data.
[1654] Input: Digital audio data.
[1655] Output: Text data.
[1656] Specific operation: Use the Google Cloud Speech-to-Text API to convert audio data into text. For example, the audio "How much stock do you have of this product?" will be converted into text data.
[1657] Step 4:
[1658] The server analyzes the user's question from the text data.
[1659] Input: Text data.
[1660] Output: Category of the question.
[1661] Specific operation: Using Google Cloud Natural Language API and SpaCy, the text data is analyzed and the questions are categorized into categories such as "stock availability check".
[1662] Step 5:
[1663] The server recognizes the user's emotions.
[1664] Input: Text data.
[1665] Output: Recognized emotion.
[1666] Specific operation: Using IBM Watson Tone Analyzer and Microsoft Azure's Text Analytics API, it recognizes the user's emotions (e.g., joy, anger, sadness, surprise) from the audio. For example, an emotion such as "confusion" might be extracted.
[1667] Step 6:
[1668] The server selects a responder based on the user's request.
[1669] Input: Category of the question, perceived emotion.
[1670] Output: Selected responder.
[1671] Specific operation: Based on user requests (e.g., specific gender, level of understanding), the system selects the most appropriate responder profile (e.g., female voice, calm tone).
[1672] Step 7:
[1673] The server generates answer data for the question.
[1674] Input: Question category, selected responder profile.
[1675] Output: Generated response data.
[1676] Specific operation: Using generative AI models such as OpenAI GPT-4, it generates answer data based on the question. For example, it might generate an answer such as, "Currently, there are 10 units of this product in stock."
[1677] Step 8:
[1678] The server converts the generated response data into speech.
[1679] Input: Generated response data, selected respondent profile.
[1680] Output: Audio data.
[1681] Specific operation: Use Amazon Polly or Google Cloud Text-to-Speech to convert text data into speech data and adjust the tone and speed of the speech.
[1682] Step 9:
[1683] The server provides the user with an audio response.
[1684] Input: Audio data.
[1685] Output: Voice response to the user.
[1686] Specific operation: The generated audio data is delivered to the user in real time via a device (smart glasses or earphones). The user can experience a natural conversation.
[1687] Step 10:
[1688] The server saves the interaction history.
[1689] Input: User's question, generated answer, date and time of interaction, perceived emotion.
[1690] Output: Saved call history.
[1691] Specific actions: Use Microsoft SQL Server or MongoDB to record this information in a database, enabling quick responses to future queries.
[1692] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1693] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1694] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1695] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1696] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1697] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1698] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1699] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1700] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1701] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1702] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1703] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1704] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1705] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1706] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1707] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1708] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1709] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1710] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1711] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[1712] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[1713] The following is further disclosed regarding the embodiments described above.
[1714] (Claim 1)
[1715] A system in which users make inquiries by voice and the system responds to those inquiries.
[1716] A means of recognizing incoming calls from users,
[1717] A means of converting user speech into text data,
[1718] A method for analyzing user questions from text data,
[1719] A means of selecting a responder based on the user's request,
[1720] A means for generating response data to the question content,
[1721] A means of converting the generated response data into speech,
[1722] A means of providing users with voiced responses,
[1723] A means of saving the interaction history,
[1724] A system that includes this.
[1725] (Claim 2)
[1726] The system according to claim 1, wherein the user's request includes "a response in a voice of a specific gender" or "a response tailored to a specific level of understanding."
[1727] (Claim 3)
[1728] The system according to claim 1, comprising means for quickly generating an answer to the same question again based on the interaction history.
[1729] "Example 1"
[1730] (Claim 1)
[1731] A system in which users make inquiries by voice and the system responds to those inquiries.
[1732] A means of recognizing incoming calls from users,
[1733] A means of converting user speech into text data,
[1734] A method for analyzing user questions from text data,
[1735] A means of selecting a responder based on the user's request,
[1736] A means for generating response data to a question using natural language generation technology,
[1737] A means of converting the generated response data into speech,
[1738] A means of providing users with voiced responses,
[1739] A means of saving the interaction history,
[1740] A system that includes this.
[1741] (Claim 2)
[1742] The system according to claim 1, wherein the user's request includes "a response in a voice of a specific gender" or "a response tailored to a specific level of understanding."
[1743] (Claim 3)
[1744] The system according to claim 1, comprising means for quickly generating an answer to the same question again based on the interaction history.
[1745] "Application Example 1"
[1746] (Claim 1)
[1747] An in-vehicle system in which a user makes an inquiry by voice and the system responds to it,
[1748] A means of recognizing a vehicle entering from a user,
[1749] A means of converting user speech into text data,
[1750] A method for analyzing user questions from text data,
[1751] A means of selecting a responder based on the user's request,
[1752] A means including a generative AI model that generates response data to the content of a question,
[1753] A means of converting the generated response data into speech,
[1754] A means of providing users with voiced responses,
[1755] A means of saving the interaction history,
[1756] In-vehicle systems including this.
[1757] (Claim 2)
[1758] The in-vehicle system according to claim 1, wherein the user's request includes "a response in a voice of a specific gender" or "a response tailored to a specific level of understanding."
[1759] (Claim 3)
[1760] The in-vehicle system according to claim 1, comprising means for quickly generating an answer to the same question again based on the response history.
[1761] "Example 2 of combining an emotion engine"
[1762] (Claim 1)
[1763] A system in which users make inquiries by voice and the system responds to those inquiries.
[1764] A means of recognizing incoming calls from users,
[1765] Means for collecting user voices,
[1766] A means of converting collected user voice data into text data,
[1767] A means of recognizing a user's emotional state from text data,
[1768] A method for analyzing user questions from text data,
[1769] A means of selecting a respondent based on the analyzed question content and recognized emotional state,
[1770] A means for generating response data to the question content,
[1771] A means of converting the generated response data into speech,
[1772] A means of providing users with voiced responses,
[1773] A means of saving the interaction history,
[1774] A system that includes this.
[1775] (Claim 2)
[1776] The system according to claim 1, wherein the user's request includes "a response in a voice of a specific gender" or "a response tailored to a specific level of understanding."
[1777] (Claim 3)
[1778] The system according to claim 1, comprising means for quickly generating an answer to the same question again based on the interaction history.
[1779] "Application example 2 when combining with an emotional engine"
[1780] (Claim 1)
[1781] A system in which users make inquiries by voice and the system responds to those inquiries.
[1782] A means of recognizing incoming calls from users,
[1783] A means of converting user speech into text data,
[1784] A method for analyzing user questions from text data,
[1785] A means of selecting a responder based on the user's request,
[1786] A means for generating response data to the question content,
[1787] A means of converting the generated response data into speech,
[1788] A means of providing users with voiced responses,
[1789] Means of recognizing user emotions,
[1790] Means of adjusting responses based on recognized emotions,
[1791] A means of saving the interaction history,
[1792] A system that includes this.
[1793] (Claim 2)
[1794] The system according to claim 1, wherein the user's request includes "a response in a voice of a specific gender" or "a response tailored to a specific level of understanding," and further includes a response adjusted by emotion recognition.
[1795] (Claim 3)
[1796] The system according to claim 1, comprising means for quickly generating an answer to the same question again based on the interaction history, and adjusting the response based on the user's emotion recognition. [Explanation of symbols]
[1797] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A system in which users make inquiries by voice and the system responds to those inquiries. A means of recognizing incoming calls from users, A means of converting user speech into text data, A method for analyzing user questions from text data, A means of selecting a responder based on the user's request, A means for generating response data to the question content, A means of converting the generated response data into speech, A means of providing users with voiced responses, A means of saving the interaction history, A system that includes this.
2. The system according to claim 1, wherein the user's request includes "a response in a voice of a specific gender" or "a response tailored to a specific level of understanding."
3. The system according to claim 1, comprising means for quickly generating an answer to the same question again based on the interaction history.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A