system
The system automates inquiry handling by converting audio to text, classifying, generating answers, and allowing human correction, addressing inefficiencies and inaccuracies in manual processes to enhance response accuracy and customer satisfaction.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-01
- Publication Date
- 2026-04-13
AI Technical Summary
Inquiry operations require prompt and accurate responses, but manual classification and creation of answers increase man-hours, leading to inefficiencies and potential inaccuracies, thereby affecting customer satisfaction and increasing staff burden.
A system that collects audio data, converts it into text, classifies the text, generates answer suggestions, allows human review and correction, and finally sends the corrected answers to customers, utilizing an audio analysis API, natural language processing model, and database.
This system reduces workload and ensures accurate and prompt responses by automating the inquiry handling process, improving customer satisfaction and reducing staff burden.
Smart Images

Figure 2026063885000001_ABST
Abstract
Description
Technical Field
[0004] , ,
[0005] , , ,
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In inquiry operations, while prompt and accurate responses to customers are required, an increase in the number of man-hours has become a problem. Manually classifying inquiry contents and creating answers takes time and also increases the risk of incorrect answers. In such a situation, the burden on the shop staff increases, which may lead to a decrease in quality and customer satisfaction. To solve these problems, a system for efficiently and accurately handling inquiries is necessary.
Means for Solving the Problems
[0006] "Audio data" refers to data that records audio information in digital format.
[0007] "Means of collection" refers to an apparatus or method for recording and storing audio data.
[0008] "Text data" refers to data that records character information in a digital format.
[0009] "Means of conversion" refers to a device or method for converting audio data into text data.
[0010] "Means of classification" refers to a device or method for dividing text data into specific categories.
[0011] "Answer sheet" refers to data that represents the response to an inquiry.
[0012] "Generating means" refers to an apparatus or method for creating answer sheets based on classified text data.
[0013] "Means for review and correction" refers to a device or method for a human to review the generated answer and make corrections as necessary.
[0014] "The means of transmission" refers to a device or method for delivering the final answer to the customer.
[0015] "The voice analysis API" is an application programming interface for analyzing voice data and converting it into text data.
[0016] "The natural language processing model" is a model that implements algorithms and techniques for enabling a computer to understand human language.
[0017] "The database" is a system for systematically managing and storing information and making it accessible as needed.
[0018] "The shop crew" refers to users who mean store staff or customer support staff.
[0019] "The customer" refers to a customer who uses a product or service.
[0020] Based on the above definitions, each element of the system according to the claims of the patent is to be understood.
Brief Description of the Drawings
[0021] [Figure 1] It is a conceptual diagram showing an example of the configuration of the data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of the data processing device and the smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of the data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of the data processing device and the smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of the data processing system according to the third embodiment. [Figure 6]It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] [[ID=二十一]]It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.
Embodiments for Carrying Out the Invention
[0022] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0023] First, the language used in the following description will be explained.
[0024] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), and APU (Accelerated Processing Unit).
[0025] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0026] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0027] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0028] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0029] [First Embodiment]
[0030] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0031] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0032] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0033] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0034] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0035] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0036] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0037] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0038] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0039] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0040] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0041] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0042] This invention is a system that collects voice data, converts it into text data, classifies it, automatically generates answer keys, allows a human to review and correct the answer keys, and finally sends them to the customer. This system uses a voice analysis API and makes full use of natural language processing models and databases. The following describes the specific operation of this system, using the server, terminal, and user as subjects.
[0043] System program processing
[0044] Audio data collection and conversion
[0045] Terminal: The shop crew member's terminal records conversations with customers. For example, they press the record button on the terminal to start a conversation, and press the button again to stop recording when the conversation is finished. This recording file is sent from the terminal to the server.
[0046] Server: Processes the received audio data and calls the speech analysis API to convert it into text data. Specifically, it sends the audio data to the speech analysis API and retrieves the corresponding text.
[0047] Classification of text data
[0048] Server: The converted text data is passed through a natural language processing model to classify the query content into specific categories. For example, it may be classified as "questions about product sizes," "order status checks," or "return / exchange procedures." This classification is performed automatically based on the model's training results.
[0049] Generating the answer
[0050] Server: Automatically generates suggested answers based on classified inquiry content. Here, the server refers to FAQ data and past answer history in the database to create an appropriate answer. For example, in response to the inquiry "What are the dimensions of the product?", it generates a suggested answer such as "The dimensions of this product are 10cm high, 5cm wide, and 1cm deep."
[0051] Human verification and correction
[0052] Terminal: The AI-generated solution is displayed on the shop crew member's terminal. The shop crew member reviews this solution and makes corrections as needed.
[0053] User (Shop Crew): Carefully review the displayed answer to ensure its accuracy. If any information is missing or incorrect, manually correct it. Once corrections are complete, press the confirmation button to proceed to the next step.
[0054] Sending a formal response
[0055] Terminal: The shop crew will review and correct the answer and finally send it to the customer. For example, by pressing the "Send" button on the terminal, the answer will be delivered to the customer via email or chat system.
[0056] Specific example
[0057] Example: Questions about products
[0058] 1. Terminal: Starts recording with Shop Crew: "Please tell me the size of product A."
[0059] 2. Server: Receives audio data and uses an audio analysis API to convert it into text data such as "Please tell me the size of product A."
[0060] 3. Server: Classified as a "product size inquiry" by the natural language processing model.
[0061] 4. Server: Retrieves size information from the product database and generates a suggested answer such as, "The dimensions of this product are 10cm high, 5cm wide, and 1cm deep."
[0062] 5. Terminal: The shop crew checks the answer and corrects it to "The size of this product is 10cm in height."
[0063] 6. Terminal: Press the "Send" button to send the corrected answer to the customer.
[0064] As described above, this system is intended to automate and streamline inquiry handling, enabling accurate and prompt responses.
[0065] The following describes the processing flow.
[0066] Step 1:
[0067] Terminal: The shop crew member's terminal records conversations with customers. Specifically, they press the record button on the terminal to start the conversation. When the conversation ends, they press the button again to stop recording and send the recording file to the server.
[0068] Step 2:
[0069] Server: Receives audio data and converts it into text data using a speech analysis API. Specifically, it sends the received audio file to the speech analysis API and retrieves the corresponding text data.
[0070] Step 3:
[0071] Server: Inputs the transcribed data into a natural language processing (NLP) model and classifies the inquiry content into specific categories. For example, it automatically classifies inquiries into categories such as "questions about product sizes," "order status inquiries," and "return procedures."
[0072] Step 4:
[0073] Server: Automatically generates answers based on categorized inquiries. Specifically, it refers to the FAQ database and past answer history to select or create appropriate answers to generate answers corresponding to categories.
[0074] Step 5:
[0075] Server: Sends the generated answer sheet to the shop crew's terminal. Specifically, it communicates with the terminal to send the text data of the answer sheet from the server.
[0076] Step 6:
[0077] Terminal: The shop crew's terminal displays the received answer on the screen.
[0078] Step 7:
[0079] User (Shop Crew): The Shop Crew will review the displayed answer and make corrections as needed. Specifically, they will check the content of the answer and manually correct any omissions or errors.
[0080] Step 8:
[0081] Terminal: After completing the corrections to the answer, press the "Submit" button for final confirmation and send it to the customer as the official answer.
[0082] Step 9:
[0083] Terminal: The official response will be delivered to the customer via email or chat system.
[0084] (Example 1)
[0085] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0086] Traditional customer service systems often involved manual processes such as collecting and analyzing voice data, converting it to text data, and generating appropriate answers. This resulted in time-consuming and inefficient responses. Furthermore, the quality of answers depended on the individual operator's judgment, potentially leading to inconsistencies. This resulted in challenges such as decreased customer satisfaction and increased operating costs.
[0087] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0088] In this invention, the server includes means for converting audio data into text data using an audio analysis API, means for classifying the text data using a natural language processing model, and means for generating answer suggestions based on the classified text data by referring to a database. This automates the entire process from collecting audio data to generating answer suggestions and sending them to customers, enabling efficient and accurate inquiry handling.
[0089] A "terminal" is a device operated by the user to collect voice data and to review and correct answer sheets.
[0090] A "server" is a computer system that performs a series of operations: receiving audio data, converting it into text data using an audio analysis API, classifying it using a natural language processing model, and generating a solution.
[0091] "Audio data" refers to digital audio information, which is recordings of human speech collected using a device.
[0092] A "Voice Analysis API" is an application programming interface for analyzing input voice data and converting it into text data.
[0093] "Text data" refers to the character information corresponding to the audio data, converted by the speech analysis API.
[0094] A "natural language processing model" is a machine learning model that analyzes text data, understands its content, and classifies it into specific categories.
[0095] A "solution" is a document that responds to a customer inquiry, generated based on classified text data.
[0096] A "database" is a storage device that stores information such as FAQ data and past answer history necessary for generating answer templates.
[0097] This invention is a system that collects audio data, converts it into text data, classifies it, automatically generates answer keys, has a human review and correct the answer keys, and finally sends them to the customer. This system uses an audio analysis API and makes full use of natural language processing models and databases.
[0098] Audio data collection and conversion
[0099] Terminal: The terminal used by the user is a device for recording conversations with customers. Specifically, the user launches the recording application and presses the record button to start the conversation. When the conversation ends, the user presses the record button again to stop recording, and this recording file is temporarily saved to the terminal's local storage. This audio file is then sent to the server.
[0100] Server: Receives audio data sent from the terminal. The received audio file is temporarily stored, and this audio data is converted into text data using a speech analysis API such as the Google® Cloud Speech-to-Text API. Specifically, the audio data is sent to the API endpoint, and the text data is retrieved from the returned response. This text data is stored in a database and used in the next step.
[0101] Classification of text data
[0102] Server: The acquired text data is input into a natural language processing model (e.g., BERT). The model analyzes the content of the text data and classifies the query into specific categories. For example, it might classify queries into "questions about product sizes," "order status checks," or "return / exchange procedures." This classification result is stored in a database and used in the next step of generating solutions.
[0103] Generating the answer
[0104] Server: Automatically generates answers based on classified category information. The server refers to FAQ data and past answer history in the database to create appropriate answers. For example, in response to the inquiry "What are the dimensions of the product?", it generates an answer such as "The dimensions of this product are 10cm high, 5cm wide, and 1cm deep." This generated answer is stored in the database and sent to the terminal in the next step.
[0105] Human verification and correction
[0106] Terminal: The answer submitted from the server is displayed on the user's terminal. The user can use a UI to review this answer and modify it as needed. Specifically, a text field is provided that allows the user to edit the displayed answer.
[0107] User: Check if the displayed answer is correct, and manually correct any missing or incorrect information. Once you have finished correcting, press the "Confirm" button to proceed to the next step.
[0108] Sending a formal response
[0109] Terminal: The user submits the completed and corrected answer to the customer by pressing the "Submit" button. This submission is carried out using email, chat API, etc. After submission is complete, the system displays a confirmation message to the customer.
[0110] Specific example
[0111] 1. Terminal: The user starts recording by saying, "Please tell me the size of product A."
[0112] 2. Server: Receives the audio data and converts it into text data, "Please tell me the size of product A," using the Google Cloud Speech-to-Text API.
[0113] 3. Server: The query is classified as "a query about product size" using the natural language processing model BERT.
[0114] 4. Server: Retrieves size information from the product database and generates an answer such as, "The dimensions of this product are 10cm high, 5cm wide, and 1cm deep."
[0115] 5. Terminal: The user reviews the answer and corrects it to "The size of this product is 10cm in height."
[0116] 6. Terminal: Press the "Send" button to send the corrected answer to the customer.
[0117] This will automate and streamline inquiry handling, enabling quick and accurate responses.
[0118] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0119] Step 1: Collect audio data
[0120] Terminal: This device records conversations between customers and shop crew. The user launches the recording application and presses the record button to begin the conversation. When the conversation ends, they press the record button again to stop recording. This audio file is temporarily stored in the terminal's local storage.
[0121] Input: Audio of a conversation between a customer and a shop crew member.
[0122] Output: Temporarily saved audio data file.
[0123] Step 2: Sending the audio data
[0124] Terminal: Sends the collected audio data to the server. Specifically, it uses the terminal's network module to upload the audio files to a specified URL on the server.
[0125] Input: Temporarily saved audio data file.
[0126] Output: Audio data sent to the server.
[0127] Step 3: Convert speech to text
[0128] Server: Converts received audio data into text data using a speech analysis API (e.g., Google Cloud Speech-to-Text API). Specifically, it sends the audio data to the API endpoint and extracts the text data from the returned response.
[0129] Input: Audio data sent to the server.
[0130] Output: Text data.
[0131] Step 4: Classification of Text Data
[0132] Server: The acquired text data is input into a natural language processing model (e.g., BERT) for classification. Specifically, the model is used to analyze the content of the text data and classify it into predefined categories. For example, it may be classified into "questions about product sizes," "order status confirmation," and "return / exchange procedures."
[0133] Input: Text data.
[0134] Output: Classification results (category labels).
[0135] Step 5: Generating the solution
[0136] Server: Based on the classified category, it searches the database for appropriate answers and generates suggested solutions. Specifically, it refers to FAQ data and past answer history to create suggested solutions that match the content of the inquiry.
[0137] Input: Classification result (category label).
[0138] Output: Generated solution.
[0139] Step 6: Submit your answer
[0140] Server: Sends the generated solution to the user's terminal. Specifically, it uploads the solution data to a specified URL on the user's terminal.
[0141] Input: Generated solution.
[0142] Output: The answer submitted to the user's terminal.
[0143] Step 7: Human review and correction
[0144] Terminal: The answer submitted from the server is displayed on the user's terminal. The user reviews this answer and makes corrections as needed.
[0145] User: Check the displayed solution, enter any necessary corrections using the text field, and save the revised version.
[0146] Input: Submitted answer.
[0147] Output: Revised solution.
[0148] Step 8: Submit your formal response
[0149] Terminal: This terminal sends the user's reviewed and corrected answer to the customer. Specifically, the user presses the "Send" button and sends the answer via email or chat API.
[0150] Input: Revised answer.
[0151] Output: The answer submitted to the customer.
[0152] Through the processing steps described above, this system automates the entire process from collecting voice data to generating answers and sending them to customers, enabling efficient and accurate handling of inquiries.
[0153] (Application Example 1)
[0154] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0155] In infotainment systems for autonomous vehicles, it is essential that passengers' voice-activated questions and requests are answered quickly and accurately. Conventional systems require human intervention, which can lead to delays and difficulties in providing accurate answers. Therefore, there is a need for systems in autonomous vehicles that can provide appropriate answers to voice commands in real time.
[0156] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0157] In this invention, the server includes means for collecting voice data, means for converting voice data into text data, and means for classifying text data. This makes it possible to generate and provide quick and accurate answers in real time to passenger voice instructions and requests in the infotainment system of an autonomous vehicle.
[0158] "Means for collecting voice data" refers to the function in the infotainment system of an autonomous vehicle that recognizes passenger voices and collects them as digital data.
[0159] "Methods for converting audio data into text data" refers to the process of analyzing collected audio data and converting it into corresponding text data using an audio analysis API or similar method.
[0160] "Means for classifying text data" refers to a function that automatically classifies converted text data into appropriate categories using a natural language processing model.
[0161] "Methods for generating answers" refers to the process of generating appropriate responses using natural language processing models, based on classified text data, referencing databases and FAQ data.
[0162] "Means for human review and correction" refers to a function where a human reviews the automatically generated answer and corrects its content as needed.
[0163] "Means for sending confirmed and revised answers" refers to the function of sending the final confirmed and revised answers to passengers via email or chat system.
[0164] A "voice recognition system installed in an autonomous vehicle" is a system installed in an autonomous vehicle that recognizes and processes passengers' voices in real time.
[0165] A "natural language processing model" is a machine learning model that understands and analyzes text data to generate appropriate responses or actions.
[0166] An "infotainment system" is a system installed in autonomous vehicles that provides information and entertainment to passengers using voice and video.
[0167] The system for implementing this invention operates as an infotainment system for an autonomous vehicle. When a passenger gives instructions or asks a question by voice, the system collects the voice, converts it into text data, and provides an appropriate answer in real time. The main hardware and software configuration is described in detail below.
[0168] Major hardware and software
[0169] Microphone: This is an input device for collecting passenger voices. It is installed inside the autonomous vehicle and is always on standby for passenger voice commands.
[0170] Speech analysis APIs are used to convert collected speech data into text data. Examples include the Google Speech Recognition API.
[0171] Natural Language Processing Model: Used to analyze the converted text data and generate appropriate answers. This model employs a generative AI model.
[0172] Database: Stores question-and-answer collections and FAQ data that are referenced when classifying text data and generating answer suggestions.
[0173] Display: This is an output device used to display the generated answer sheet to passengers. It is commonly installed on the dashboard of autonomous vehicles.
[0174] System processing flow
[0175] 1. Collection of voice data: The terminal (microphone inside the autonomous vehicle) collects the passenger's voice instructions. For example, questions such as "How long will it take to arrive at the next destination?" are collected.
[0176] 2. Audio Data Conversion: The device sends the collected audio data to the server, which uses an audio analysis API to convert it into text data. As a result, the text data "How long will it take to arrive at the next destination?" is generated.
[0177] 3. Text Data Classification: The server uses a natural language processing model to classify the converted text data into appropriate categories. For example, it might be classified as "Confirmation of Arrival Time."
[0178] 4. Solution Generation: Based on the classified category, the server retrieves appropriate information from the database and generates a solution using a generative AI model. For example, a solution such as "It will arrive in approximately 15 minutes" might be generated.
[0179] 5. Display of the solution: The generated solution is sent to the terminal (the display in the autonomous vehicle) and presented to the passengers visually.
[0180] Specific example
[0181] As a concrete example, consider the following conversation:
[0182] Passenger: "How long will it take to get to our next destination?"
[0183] System: "Arriving in approximately 15 minutes."
[0184] Based on this prompt, the system will process the information in the following order.
[0185] Example of a prompt:
[0186] Please convert the following audio data to text: "How long will it take to arrive at our next destination?"
[0187] Please categorize the following text data: "How long will it take to arrive at the next destination?"
[0188] Please generate an answer based on the following category and text data: "How long will it take to arrive at the next destination?"
[0189] These processes enable the infotainment system of an autonomous vehicle to respond quickly and accurately to voice commands from passengers.
[0190] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0191] Step 1:
[0192] Collection of audio data
[0193] Input: Passenger voice instructions
[0194] Operation: The terminal (microphone inside the autonomous vehicle) collects passenger voices in real time. For example, it recognizes voices such as, "How long will it take to arrive at the next destination?"
[0195] Output: Audio data file
[0196] Specific operation: The microphone records sound as digital data and passes the collected audio data to subsequent processing.
[0197] Step 2:
[0198] Audio data conversion
[0199] Input: Audio data file
[0200] Operation: The server uses a speech analysis API to convert speech data into text data. Google Speech Recognition is used as the API.
[0201] Output: Text data
[0202] Specific operation: Send audio data to the voice analysis API and retrieve the text "How long will it take to arrive at the next destination?". Pass this text data to the next processing step.
[0203] Step 3:
[0204] Classification of text data
[0205] Input: Text data
[0206] Operation: The server uses a natural language processing model to categorize text data into appropriate categories. For example, it might categorize it as "Confirming arrival time".
[0207] Output: Category Information
[0208] Specific operation: Text data is input into a natural language processing model, and the model outputs categories based on its learning results. This category information is then passed to the next step.
[0209] Step 4:
[0210] Generating the answer
[0211] Input: Category information and original text data
[0212] Operation: The server uses a database and a generative AI model to generate answer suggestions based on categories. For example, it might generate an answer such as "We will arrive in approximately 15 minutes."
[0213] Output: Solution
[0214] Specific operation: Relevant information (e.g., arrival time data) is retrieved from the database, and the optimal solution is generated using a generative AI model. This solution is then passed on to the next step.
[0215] Step 5:
[0216] Display of the answer
[0217] Input: Answer
[0218] Operation: The terminal (the display in the autonomous vehicle) displays the generated answer to the passengers. The display shows "We will arrive in approximately 15 minutes."
[0219] Output: Display of the answer sheet that passengers can see.
[0220] Specific operation: The answer will be displayed on a screen and presented in a format that is easy for passengers to read. This will allow passengers to see the answer to their question in real time.
[0221] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0222] This invention is a system that collects voice data, converts it into text data, classifies it, analyzes the user's emotions using an emotion engine, automatically generates a response, has a human review and revise the response, and finally sends it to the customer. This system makes full use of a voice analysis API, a natural language processing model, a database, and an emotion engine. The following describes the specific operation of this system, using the server, terminal, and user as subjects.
[0223] System program processing
[0224] Audio data collection and conversion
[0225] Terminal: The shop crew member's terminal records conversations with customers. Specifically, they press the record button on the terminal to start the conversation, and press the button again to stop recording when the conversation is finished. This recording file is sent from the terminal to the server.
[0226] Server: Processes the received audio data and calls the speech analysis API to convert the audio data into text data. Specifically, it sends the audio data to the speech analysis API and retrieves the corresponding text.
[0227] Text data classification and sentiment analysis
[0228] Server: The converted text data is input into a natural language processing (NLP) model to classify the inquiry content into specific categories. For example, it may be classified into "questions about product sizes," "inquiries about order status," "return procedures," etc.
[0229] Server: Sends classified text data to the sentiment engine to analyze the user's emotions. The sentiment engine identifies the user's emotions (e.g., joy, anger, sadness, etc.) from words and phrases in the text data.
[0230] Generating solutions and making emotional revisions.
[0231] Server: Automatically generates suggested answers based on classified inquiry content and sentiment analysis results. Specifically, it refers to the FAQ database and past answer history to create answers that take the user's feelings into consideration. For example, if the user looks dissatisfied with the inquiry "What are the dimensions of the product?", it will respond politely with something like, "Thank you for waiting. The dimensions of this product are 10cm high, 5cm wide, and 1cm deep."
[0232] Human verification and correction
[0233] Terminal: The AI-generated solution is displayed on the shop crew member's terminal. The shop crew member reviews this solution and makes corrections as needed.
[0234] User (Shop Crew): Carefully review the displayed answer to ensure its accuracy. If any information is missing or incorrect, manually correct it. Once corrections are complete, press the confirmation button to proceed to the next step.
[0235] Sending a formal response
[0236] Terminal: The shop crew will send the customer the corrected and revised answer. Specifically, they will press the "Send" button on the terminal and deliver the answer to the customer via email or chat system.
[0237] Specific example
[0238] Example: Questions about a product and sentiment analysis
[0239] 1. Terminal: Starts recording with Shop Crew: "Please tell me the size of product A."
[0240] 2. Server: Receives audio data and uses an audio analysis API to convert it into text data such as "Please tell me the size of product A."
[0241] 3. Server: Classified as a "product size inquiry" by the natural language processing model.
[0242] 4. Server: Sends text data to the emotion engine to analyze whether the user is experiencing anxiety.
[0243] 5. Server: Based on the sentiment analysis results, it generates a polite response: "Thank you for waiting. The dimensions of this product are 10cm high, 5cm wide, and 1cm deep. Do you have any further questions?"
[0244] 6. Terminal: The shop crew checks the answer and corrects it to "The size of this product is 10cm in height."
[0245] 7. Terminal: Press the "Send" button to send the corrected answer to the customer.
[0246] As described above, this system is intended to automate and streamline inquiry handling, enabling accurate and rapid responses. Furthermore, by combining it with user sentiment analysis using an emotion engine, it is possible to provide more personalized services.
[0247] The following describes the processing flow.
[0248] Step 1:
[0249] Terminal: The shop crew member's terminal records conversations with customers. Specifically, they press the record button on the terminal to start the conversation, and press the button again to stop recording when the conversation is finished. This recording file is sent from the terminal to the server.
[0250] Step 2:
[0251] Server: Processes the received audio data and calls the audio analysis API to convert the audio data into text data. Specifically, it sends the audio file to the audio analysis API and retrieves the corresponding text data.
[0252] Step 3:
[0253] Server: The converted text data is input into a natural language processing (NLP) model, which classifies the query into a specific category. For example, if the text data is "What are the product sizes?", it will be classified as "Questions about product sizes".
[0254] Step 4:
[0255] Server: After passing text data through a natural language processing model, the emotion engine analyzes the user's emotions. Specifically, it identifies emotions such as whether the user is happy or dissatisfied based on the words and phrases in the text data.
[0256] Step 5:
[0257] Server: Automatically generates suggested answers based on classified inquiry content and sentiment analysis results. Specifically, it refers to the FAQ database and past answer history to create answers that take the user's feelings into consideration. For example, if the user seems dissatisfied with the question "What are the dimensions of the product?", it will create a polite suggested answer such as, "We apologize for the wait. The dimensions of this product are 10cm high, 5cm wide, and 1cm deep."
[0258] Step 6:
[0259] Server: Sends the generated answer sheet to the shop crew's terminal. Specifically, it communicates with the terminal to send the text data of the answer sheet from the server.
[0260] Step 7:
[0261] Terminal: The shop crew's terminal displays the received answer on the screen.
[0262] Step 8:
[0263] User (Shop Crew): The Shop Crew will review the displayed solution and make corrections as needed. Specifically, they will check the content of the solution and manually correct any omissions or errors.
[0264] Step 9:
[0265] Terminal: After completing the corrections to the answer, press the "Submit" button for final confirmation and send it to the customer as the official answer.
[0266] Step 10:
[0267] Terminal: The official response will be delivered to the customer via email or chat system.
[0268] (Example 2)
[0269] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0270] In modern customer service, efficiently collecting voice data and providing quick and accurate answers is crucial. However, converting voice data into text, classifying the content, and generating appropriate answers is a time-consuming process. Furthermore, personalized responses that take into account user emotions are required, but this is difficult to achieve with current systems. Therefore, there is a need for a system that can efficiently process voice data and provide answers that respond to user emotions.
[0271] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0272] In this invention, the server includes means for collecting voice data, means for converting the voice data into text data, means for classifying the text data, means for analyzing the user's emotions based on the classified text data, means for generating answer suggestions based on the analyzed emotions and classification results, means for a human to review and correct the generated answer suggestions, and means for transmitting the reviewed and corrected answer suggestions. This enables efficient processing of voice data and the provision of personalized answers that respond to the user's emotions in customer service operations.
[0273] "Audio data" refers to information recorded in digital format using audio.
[0274] "Text data" refers to data obtained by converting audio data into written text.
[0275] "Classifying" refers to dividing data into categories based on specific criteria.
[0276] "Analyzing emotions" refers to detecting and identifying a user's emotional state from text data.
[0277] "Answer" refers to the content of the response to a user's inquiry.
[0278] "To confirm and correct" means that a human reviews the generated answer and modifies the content as necessary.
[0279] "To send" means using communication means to deliver the corrected answer to the user.
[0280] This system realizes a series of processes including collecting voice data, converting it into text data, classifying it, analyzing the user's emotions using an emotion engine, automatically generating an answer, having a human confirm and correct the answer, and finally sending it to the customer. This system makes full use of a voice analysis API, a natural language processing model, a database, and an emotion engine.
[0281] Details of Hardware and Software [[ID=Database: A database system for storing and managing various types of data (audio data, text data, analysis results, answer sheets, etc.).
[0289] Specific operation of the system
[0290] The terminal is used by shop staff to record conversations with customers. They press the record button to start the conversation and press it again to stop recording when finished. This recording file is then sent from the terminal to the server.
[0291] The server sends the received audio data to the Google Cloud Speech-to-Text API, where it converts the audio data into text data. The converted text data is then classified by a natural language processing model such as OpenAI GPT-4. For example, it may be categorized into "questions about product sizes," "inquiries about order status," and "return procedures."
[0292] Furthermore, the classified text data is sent to IBM Watson's sentiment engine, where the user's emotions are analyzed. Sentiment analysis is performed based on words and phrases within the text data to identify the user's emotions (e.g., joy, anger, sadness, etc.).
[0293] The server automatically generates suggested answers based on the classification results and sentiment analysis results. At this stage, it refers to the FAQ database and past answer history to create answers that take the user's emotions into consideration. For example, if a user asks "What are the product sizes?" and sentiment analysis detects anxiety, the prompt to the generating AI model will be "Answer when the user seems anxious: 'What are the product sizes?'" to generate a polite answer.
[0294] The generated answer is displayed on the terminal, where a shop crew member reviews and corrects it. They carefully examine whether the answer is correct and make corrections as needed. Once the review is complete, the answer is sent to the customer by pressing the "Send" button on the terminal. This entire process automates and streamlines inquiry handling, enabling appropriate responses that are tailored to the user's emotions.
[0295] Specific example
[0296] As an example, the flow of questions about a product and sentiment analysis is shown below.
[0297] 1. Terminal: The shop crew member starts recording, saying, "Please tell me the size of product A."
[0298] 2. Server: Receives the audio data and converts it into text data, "Please tell me the size of product A," using the Google Cloud Speech-to-Text API.
[0299] 3. Server: Use OpenAI GPT-4 to classify text data into "product size inquiries".
[0300] 4. Server: Sends text data to IBM Watson's sentiment engine to analyze whether the user is experiencing anxiety.
[0301] 5. Server: Based on the sentiment analysis results, it generates a polite response: "Thank you for waiting. The dimensions of this product are 10cm high, 5cm wide, and 1cm deep. Do you have any further questions?"
[0302] 6. Terminal: The shop crew checks the answer and corrects it to "The size of this product is 10cm in height."
[0303] 7. Terminal: Press the "Send" button to send the corrected answer to the customer.
[0304] In this way, the entire system smoothly executes a series of processes to improve the efficiency and quality of customer service operations.
[0305] The flow of the specific process in Example 2 will be described using FIG. 13.
[0306] Step 1:
[0307] Collection of voice data
[0308] Subject: Terminal
[0309] Specific operation: The shop staff presses the recording button on the terminal to record the conversation with the customer. After the conversation ends, the button is pressed again to stop the recording.
[0310] Input: Voice of the conversation between the shop staff and the customer
[0311] Output: Recorded voice file (saved in digital format)
[0312] Step 2:
[0313] Transmission of voice data
[0314] Subject: Terminal
[0315] Specific operation: The recorded voice file is automatically uploaded to the server. After the upload is complete, a notification is displayed.
[0316] Input: Recorded voice file
[0317] Output: Voice file transmitted to the server
[0318] Step 3:
[0319] Storage of voice data
[0320] Subject: Server
[0321] Specific action: The received audio data is saved to the server's storage.
[0322] Input: Audio file sent from the device
[0323] Output: Audio data stored on the server
[0324] Step 4:
[0325] Converting audio data to text
[0326] Subject: Server
[0327] Specific operation: The saved audio data is sent to the Google Cloud Speech-to-Text API and converted into text data.
[0328] Input: Audio data on the server
[0329] Output: Text data returned from the Google Cloud Speech-to-Text API
[0330] Step 5:
[0331] Saving text data
[0332] Subject: Server
[0333] Specific action: Save the converted text data to the database.
[0334] Input: Text data
[0335] Output: Text data stored in the database
[0336] Step 6:
[0337] Classification of text data
[0338] Subject: Server
[0339] Specific operation: Text data is input into a natural language processing model such as OpenAI GPT-4, and the query content is classified.
[0340] Input: Text data read from a database
[0341] Output: Classification results returned from the natural language processing model
[0342] Step 7:
[0343] Saving classification results
[0344] Subject: Server
[0345] Specific action: Save the classification results to the database.
[0346] Input: Classification result
[0347] Output: Classification results stored in the database
[0348] Step 8:
[0349] Emotion analysis
[0350] Subject: Server
[0351] Specific operation: Classified text data is sent to IBM Watson's sentiment engine for sentiment analysis.
[0352] Input: Classified text data
[0353] Output: Sentiment analysis results returned from IBM Watson
[0354] Step 9:
[0355] Saving emotion analysis results
[0356] Subject: Server
[0357] Specific action: Save the emotion analysis results to the database.
[0358] Input: Sentiment analysis results
[0359] Output: Sentiment analysis results stored in the database
[0360] Step 10:
[0361] Generating the answer
[0362] Subject: Server
[0363] Specific operation: Automatically generates suggested answers based on classification results and sentiment analysis results. Refers to the FAQ database and past answer history to generate answers that take the user's emotions into consideration.
[0364] Input: Classification results and sentiment analysis results
[0365] Output: Solution returned by the generative AI model
[0366] Step 11:
[0367] Display of the answer
[0368] Subject: terminal
[0369] Specific action: Display the generated solution on the shop crew's terminal.
[0370] Input: Generated solution
[0371] Output: Answer displayed on the terminal
[0372] Step 12:
[0373] Review and correct the answer sheet.
[0374] Subject: User (Shop Crew)
[0375] Specific actions: The shop crew will review the displayed answer and make any necessary corrections. Once the corrections are complete, they will press the confirmation button.
[0376] Input: Displayed answer
[0377] Output: Revised solution
[0378] Step 13:
[0379] Sending a formal response
[0380] Subject: terminal
[0381] Specific actions: Send the corrected answer to the customer. Press the "Send" button on the device to deliver the answer to the customer via email or chat system.
[0382] Input: Revised answer
[0383] Output: Official response sent to the customer
[0384] (Application Example 2)
[0385] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0386] Currently, customer support operations, including food delivery services, require prompt and appropriate responses to customer inquiries. However, manual responses can lead to delays and decreased customer satisfaction. Furthermore, responses tend to be formulaic, making it difficult to provide personalized responses that address customer emotions and circumstances. Therefore, there is a need to achieve both automated and personalized responses to inquiries.
[0387] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0388] In this invention, the server includes means for collecting voice data, means for converting voice data into text data, means for classifying the text data, means for analyzing emotions based on the text data, means for generating answer suggestions based on the classified text data and emotion analysis results, means for a human to review and correct the generated answer suggestions, and means for transmitting the reviewed and corrected answer suggestions. This enables rapid and appropriate automated responses in customer support operations and makes it possible to provide personalized answers that take into account the customer's emotions.
[0389] "Audio data" refers to data that records audio in digital format.
[0390] "Text data" refers to string information converted from audio data.
[0391] A "Voice Analysis API" is an application programming interface for converting voice data into text data.
[0392] A "natural language processing model" is a machine learning or deep learning model used to analyze, classify, and interpret text data.
[0393] A "sentiment analysis engine" is an algorithm or system that identifies the user's emotions contained in text data.
[0394] A "database" is a system for efficiently managing, searching, and updating information.
[0395] A "solution" is a written response to a customer's inquiry.
[0396] "Customer support" refers to the work of responding to inquiries and requests from customers.
[0397] A "server" is a computer system used for storing and processing data.
[0398] A "food delivery service" is a service that delivers food ordered online by customers to a specified location.
[0399] "Personalization" refers to optimizing responses according to the individual needs and circumstances of each customer.
[0400] "Classification of inquiry content" is the process of assigning text data to a specific category.
[0401] "Emotion analysis results" refer to emotional information identified by the emotion analysis engine from text data.
[0402] This invention is a system that collects voice data, converts it into text data, classifies it, analyzes the user's emotions using an emotion engine, automatically generates a response, has a human review and revise the response, and finally sends it to the customer. This system is realized by utilizing a voice analysis API, a natural language processing model, a database, and an emotion engine.
[0403] System Configuration
[0404] 1. Collection of audio data
[0405] Device: The customer support representative's device (smartphone or computer) will record the conversation with the customer. They will press the record button to start the conversation and press it again to stop recording at the end.
[0406] 2. Converting audio data
[0407] Server: Processes the received audio data and converts it into text data by calling a speech analysis API. Specifically, it sends the collected audio data to a speech analysis API, such as Google's speech recognition API, and retrieves the corresponding text.
[0408] 3. Classification and sentiment analysis of text data
[0409] Server: The converted text data is input into a natural language processing (NLP) model to classify the inquiry content into specific categories. For example, it may be classified as "order status confirmation," "complaint handling," or "product-related questions."
[0410] Server: Sends classified text data to the sentiment engine to analyze the user's emotions. The sentiment engine identifies the user's emotions (e.g., joy, anger, sadness, etc.) from words and phrases in the text data.
[0411] 4. Generating solutions and making emotional revisions
[0412] Server: Automatically generates suggested answers based on classified inquiry content and sentiment analysis results. Specifically, it refers to the FAQ database and past answer history to create answers that take the user's emotions into consideration. For example, if a user expresses anger in response to an inquiry about a delayed delivery, it will respond politely with something like, "We apologize. We will take immediate action to resolve your dissatisfaction. Could you please provide more details?"
[0413] 5. Human review and correction
[0414] Terminal: The AI-generated solution is displayed on the customer support representative's terminal. The customer support representative reviews this solution and makes corrections as needed.
[0415] User (Customer Support Representative): Review the displayed solution to ensure its accuracy. If any information is missing or incorrect, manually correct it. Once corrections are complete, press the confirmation button to proceed to the next step.
[0416] 6. Submitting the formal response
[0417] Terminal: Customer support staff will send the reviewed and corrected answers to the customer. Specifically, they will press the "Send" button on the terminal to deliver the answer to the customer via email or chat system.
[0418] Specific example
[0419] As an example, consider a case where a customer inquires that their product delivery is delayed. A customer support representative records this inquiry and sends the audio data to a server. The server converts the audio into text data and categorizes it as "order status inquiry." The emotion analysis then identifies "anger." Based on this, a response is generated such as, "We apologize. We will promptly address your issue to resolve it. Could you please provide more details?"
[0420] Example of a prompt:
[0421] "Text: 'My delivery is delayed and I'm having trouble. What's going on?' Please categorize the content. Estimate the category and analyze the sentiment."
[0422] This system enables the automation and streamlining of inquiry handling, allowing for customer service that is sensitive to customer emotions.
[0423] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0424] Step 1:
[0425] Collection of audio data
[0426] Terminal: Customer support staff record conversations with customers. They press the record button to start collecting audio data, and press the button again to stop recording when the conversation ends. This saves the audio file to the terminal. This audio file becomes the input data for subsequent processing.
[0427] Step 2:
[0428] Sending and converting audio data
[0429] Server: Receives audio files sent from the terminal. The received audio files are sent to a speech analysis API for conversion into text data. Specifically, the speech analysis API analyzes the audio signal and outputs it as string data. The audio file is the input data, and the text data is the output data.
[0430] Step 3:
[0431] Classification of text data
[0432] Server: The converted text data is input into a natural language processing (NLP) model, which classifies the inquiry content into categories. This model is a pre-trained machine learning model that automatically assigns text data to categories such as "order status confirmation," "complaint handling," and "product-related questions." Text data is the input data, and category information is the output data.
[0433] Step 4:
[0434] Emotion analysis
[0435] Server: Sends classified text data to the sentiment engine to analyze the user's emotions. The sentiment engine identifies the user's emotions (e.g., joy, anger, sadness, etc.) from words and phrases in the text data. The text data is the input data for sentiment analysis, and the emotional information is the output data.
[0436] Step 5:
[0437] Automatic generation of answer sheets
[0438] Server: Automatically generates suggested answers based on classified category information and sentiment analysis results. Specifically, it refers to the FAQ database and past answer history to obtain the corresponding standard answer. Then, considering the sentiment analysis results, it generates a response that includes a response tailored to the user's emotions. Category information and sentiment information are input data, and suggested answers are output data.
[0439] Step 6:
[0440] Review and correction of the answer sheet
[0441] Terminal: The AI-generated answer is displayed on the customer support representative's terminal. The representative reviews this answer and makes corrections as needed. Once corrections are complete, they press the confirmation button to proceed to the next step. The initial answer is the input data, and the reviewed and corrected answer is the output data.
[0442] Step 7:
[0443] Sending a formal response
[0444] Terminal: Customer support staff send the reviewed and corrected answers to the customer. Specifically, they press the "Send" button on the terminal to deliver the answer to the customer via email or chat system. The corrected answer is the input data, and the answer sent to the customer is the output data.
[0445] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0446] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0447] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0448] [Second Embodiment]
[0449] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0450] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0451] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0452] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0453] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0454] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0455] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0456] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0457] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0458] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0459] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0460] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0461] This invention is a system that collects voice data, converts it into text data, classifies it, automatically generates answer keys, allows a human to review and correct the answer keys, and finally sends them to the customer. This system uses a voice analysis API and makes full use of natural language processing models and databases. The following describes the specific operation of this system, using the server, terminal, and user as subjects.
[0462] System program processing
[0463] Audio data collection and conversion
[0464] Terminal: The shop crew member's terminal records conversations with customers. For example, they press the record button on the terminal to start a conversation, and press the button again to stop recording when the conversation is finished. This recording file is sent from the terminal to the server.
[0465] Server: Processes the received audio data and calls the speech analysis API to convert it into text data. Specifically, it sends the audio data to the speech analysis API and retrieves the corresponding text.
[0466] Classification of text data
[0467] Server: The converted text data is passed through a natural language processing model to classify the query content into specific categories. For example, it may be classified as "questions about product sizes," "order status checks," or "return / exchange procedures." This classification is performed automatically based on the model's training results.
[0468] Generating the answer
[0469] Server: Automatically generates suggested answers based on classified inquiry content. Here, the server refers to FAQ data and past answer history in the database to create an appropriate answer. For example, in response to the inquiry "What are the dimensions of the product?", it generates a suggested answer such as "The dimensions of this product are 10cm high, 5cm wide, and 1cm deep."
[0470] Human verification and correction
[0471] Terminal: The AI-generated solution is displayed on the shop crew member's terminal. The shop crew member reviews this solution and makes corrections as needed.
[0472] User (Shop Crew): Carefully review the displayed answer to ensure its accuracy. If any information is missing or incorrect, manually correct it. Once corrections are complete, press the confirmation button to proceed to the next step.
[0473] Sending a formal response
[0474] Terminal: The shop crew will review and correct the answer and finally send it to the customer. For example, by pressing the "Send" button on the terminal, the answer will be delivered to the customer via email or chat system.
[0475] Specific example
[0476] Example: Questions about products
[0477] 1. Terminal: Starts recording with Shop Crew: "Please tell me the size of product A."
[0478] 2. Server: Receives audio data and uses an audio analysis API to convert it into text data such as "Please tell me the size of product A."
[0479] 3. Server: Classified as a "product size inquiry" by the natural language processing model.
[0480] 4. Server: Retrieves size information from the product database and generates a suggested answer such as, "The dimensions of this product are 10cm high, 5cm wide, and 1cm deep."
[0481] 5. Terminal: The shop crew checks the answer and corrects it to "The size of this product is 10cm in height."
[0482] 6. Terminal: Press the "Send" button to send the corrected answer to the customer.
[0483] As described above, this system is intended to automate and streamline inquiry handling, enabling accurate and prompt responses.
[0484] The following describes the processing flow.
[0485] Step 1:
[0486] Terminal: The shop crew member's terminal records conversations with customers. Specifically, they press the record button on the terminal to start the conversation. When the conversation ends, they press the button again to stop recording and send the recording file to the server.
[0487] Step 2:
[0488] Server: Receives audio data and converts it into text data using a speech analysis API. Specifically, it sends the received audio file to the speech analysis API and retrieves the corresponding text data.
[0489] Step 3:
[0490] Server: Inputs the transcribed data into a natural language processing (NLP) model and classifies the inquiry content into specific categories. For example, it automatically classifies inquiries into categories such as "questions about product sizes," "order status inquiries," and "return procedures."
[0491] Step 4:
[0492] Server: Automatically generates answers based on categorized inquiries. Specifically, it refers to the FAQ database and past answer history to select or create appropriate answers to generate answers corresponding to categories.
[0493] Step 5:
[0494] Server: Sends the generated answer sheet to the shop crew's terminal. Specifically, it communicates with the terminal to send the text data of the answer sheet from the server.
[0495] Step 6:
[0496] Terminal: The shop crew's terminal displays the received answer on the screen.
[0497] Step 7:
[0498] User (Shop Crew): The Shop Crew will review the displayed solution and make corrections as needed. Specifically, they will check the content of the solution and manually correct any omissions or errors.
[0499] Step 8:
[0500] Terminal: After completing the corrections to the answer, press the "Submit" button for final confirmation and send it to the customer as the official answer.
[0501] Step 9:
[0502] Terminal: The official response will be delivered to the customer via email or chat system.
[0503] (Example 1)
[0504] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0505] Traditional customer service systems often involved manual processes such as collecting and analyzing voice data, converting it to text data, and generating appropriate answers. This resulted in time-consuming and inefficient responses. Furthermore, the quality of answers depended on the individual operator's judgment, potentially leading to inconsistencies. This resulted in challenges such as decreased customer satisfaction and increased operating costs.
[0506] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0507] In this invention, the server includes means for converting audio data into text data using an audio analysis API, means for classifying the text data using a natural language processing model, and means for generating answer suggestions based on the classified text data by referring to a database. This automates the entire process from collecting audio data to generating answer suggestions and sending them to customers, enabling efficient and accurate inquiry handling.
[0508] A "terminal" is a device operated by the user to collect voice data and to review and correct answer sheets.
[0509] A "server" is a computer system that performs a series of operations: receiving audio data, converting it into text data using an audio analysis API, classifying it using a natural language processing model, and generating a solution.
[0510] "Audio data" refers to digital audio information, which is recordings of human speech collected using a device.
[0511] A "Voice Analysis API" is an application programming interface for analyzing input voice data and converting it into text data.
[0512] "Text data" refers to the character information corresponding to the audio data, converted by the speech analysis API.
[0513] A "natural language processing model" is a machine learning model that analyzes text data, understands its content, and classifies it into specific categories.
[0514] A "solution" is a document that responds to a customer inquiry, generated based on classified text data.
[0515] A "database" is a storage device that stores information such as FAQ data and past answer history necessary for generating answer templates.
[0516] This invention is a system that collects audio data, converts it into text data, classifies it, automatically generates answer keys, has a human review and correct the answer keys, and finally sends them to the customer. This system uses an audio analysis API and makes full use of natural language processing models and databases.
[0517] Audio data collection and conversion
[0518] Terminal: The terminal used by the user is a device for recording conversations with customers. Specifically, the user launches the recording application and presses the record button to start the conversation. When the conversation ends, the user presses the record button again to stop recording, and this recording file is temporarily saved to the terminal's local storage. This audio file is then sent to the server.
[0519] Server: Receives audio data sent from the terminal. The received audio file is temporarily stored, and this audio data is converted into text data using a speech analysis API such as the Google Cloud Speech-to-Text API. Specifically, the audio data is sent to the API endpoint, and the text data is retrieved from the returned response. This text data is stored in a database and used in the next step.
[0520] Classification of text data
[0521] Server: The acquired text data is input into a natural language processing model (e.g., BERT). The model analyzes the content of the text data and classifies the query into specific categories. For example, it might classify queries into "questions about product sizes," "order status checks," or "return / exchange procedures." This classification result is stored in a database and used in the next step of generating solutions.
[0522] Generating the answer
[0523] Server: Automatically generates answers based on classified category information. The server refers to FAQ data and past answer history in the database to create appropriate answers. For example, in response to the inquiry "What are the dimensions of the product?", it generates an answer such as "The dimensions of this product are 10cm high, 5cm wide, and 1cm deep." This generated answer is stored in the database and sent to the terminal in the next step.
[0524] Human verification and correction
[0525] Terminal: The answer submitted from the server is displayed on the user's terminal. The user can use a UI to review this answer and modify it as needed. Specifically, a text field is provided that allows the user to edit the displayed answer.
[0526] User: Check if the displayed answer is correct, and manually correct any missing or incorrect information. Once you have finished correcting, press the "Confirm" button to proceed to the next step.
[0527] Sending a formal response
[0528] Terminal: The user submits the completed and corrected answer to the customer by pressing the "Submit" button. This submission is carried out using email, chat API, etc. After submission is complete, the system displays a confirmation message to the customer.
[0529] Specific example
[0530] 1. Terminal: The user starts recording by saying, "Please tell me the size of product A."
[0531] 2. Server: Receives the audio data and converts it into text data, "Please tell me the size of product A," using the Google Cloud Speech-to-Text API.
[0532] 3. Server: The query is classified as "a query about product size" using the natural language processing model BERT.
[0533] 4. Server: Retrieves size information from the product database and generates an answer such as, "The dimensions of this product are 10cm high, 5cm wide, and 1cm deep."
[0534] 5. Terminal: The user reviews the answer and corrects it to "The size of this product is 10cm in height."
[0535] 6. Terminal: Press the "Send" button to send the corrected answer to the customer.
[0536] This will automate and streamline inquiry handling, enabling quick and accurate responses.
[0537] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0538] Step 1: Collect audio data
[0539] Terminal: This device records conversations between customers and shop crew. The user launches the recording application and presses the record button to begin the conversation. When the conversation ends, they press the record button again to stop recording. This audio file is temporarily stored in the terminal's local storage.
[0540] Input: Audio of a conversation between a customer and a shop crew member.
[0541] Output: Temporarily saved audio data file.
[0542] Step 2: Sending the audio data
[0543] Terminal: Sends the collected audio data to the server. Specifically, it uses the terminal's network module to upload the audio files to a specified URL on the server.
[0544] Input: Temporarily saved audio data file.
[0545] Output: Audio data sent to the server.
[0546] Step 3: Convert speech to text
[0547] Server: Converts received audio data into text data using a speech analysis API (e.g., Google Cloud Speech-to-Text API). Specifically, it sends the audio data to the API endpoint and extracts the text data from the returned response.
[0548] Input: Audio data sent to the server.
[0549] Output: Text data.
[0550] Step 4: Classification of Text Data
[0551] Server: The acquired text data is input into a natural language processing model (e.g., BERT) for classification. Specifically, the model is used to analyze the content of the text data and classify it into predefined categories. For example, it may be classified into "questions about product sizes," "order status confirmation," and "return / exchange procedures."
[0552] Input: Text data.
[0553] Output: Classification results (category labels).
[0554] Step 5: Generating the solution
[0555] Server: Based on the classified category, it searches the database for appropriate answers and generates suggested solutions. Specifically, it refers to FAQ data and past answer history to create suggested solutions that match the content of the inquiry.
[0556] Input: Classification result (category label).
[0557] Output: Generated solution.
[0558] Step 6: Submit your answer
[0559] Server: Sends the generated solution to the user's terminal. Specifically, it uploads the solution data to a specified URL on the user's terminal.
[0560] Input: Generated solution.
[0561] Output: The answer submitted to the user's terminal.
[0562] Step 7: Human review and correction
[0563] Terminal: The answer submitted from the server is displayed on the user's terminal. The user reviews this answer and makes corrections as needed.
[0564] User: Check the displayed solution, enter any necessary corrections using the text field, and save the revised version.
[0565] Input: Submitted answer.
[0566] Output: Revised solution.
[0567] Step 8: Submit your formal response
[0568] Terminal: This terminal sends the user's reviewed and corrected answer to the customer. Specifically, the user presses the "Send" button and sends the answer via email or chat API.
[0569] Input: Revised answer.
[0570] Output: The answer submitted to the customer.
[0571] Through the processing steps described above, this system automates the entire process from collecting voice data to generating answers and sending them to customers, enabling efficient and accurate handling of inquiries.
[0572] (Application Example 1)
[0573] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0574] In infotainment systems for autonomous vehicles, it is essential that passengers' voice-activated questions and requests are answered quickly and accurately. Conventional systems require human intervention, which can lead to delays and difficulties in providing accurate answers. Therefore, there is a need for systems in autonomous vehicles that can provide appropriate answers to voice commands in real time.
[0575] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0576] In this invention, the server includes means for collecting voice data, means for converting voice data into text data, and means for classifying text data. This makes it possible to generate and provide quick and accurate answers in real time to passenger voice instructions and requests in the infotainment system of an autonomous vehicle.
[0577] "Means for collecting voice data" refers to the function in the infotainment system of an autonomous vehicle that recognizes passenger voices and collects them as digital data.
[0578] "Methods for converting audio data into text data" refers to the process of analyzing collected audio data and converting it into corresponding text data using an audio analysis API or similar method.
[0579] "Means for classifying text data" refers to a function that automatically classifies converted text data into appropriate categories using a natural language processing model.
[0580] "Methods for generating answers" refers to the process of generating appropriate responses using natural language processing models, based on classified text data, referencing databases and FAQ data.
[0581] "Means for human review and correction" refers to a function where a human reviews the automatically generated answer and corrects its content as needed.
[0582] "Means for sending confirmed and revised answers" refers to the function of sending the final confirmed and revised answers to passengers via email or chat system.
[0583] A "voice recognition system installed in an autonomous vehicle" is a system installed in an autonomous vehicle that recognizes and processes passengers' voices in real time.
[0584] A "natural language processing model" is a machine learning model that understands and analyzes text data to generate appropriate responses or actions.
[0585] An "infotainment system" is a system installed in autonomous vehicles that provides information and entertainment to passengers using voice and video.
[0586] The system for implementing this invention operates as an infotainment system for an autonomous vehicle. When a passenger gives instructions or asks a question by voice, the system collects the voice, converts it into text data, and provides an appropriate answer in real time. The main hardware and software configuration is described in detail below.
[0587] Major hardware and software
[0588] Microphone: This is an input device for collecting passenger voices. It is installed inside the autonomous vehicle and is always on standby for passenger voice commands.
[0589] Speech analysis APIs are used to convert collected speech data into text data. Examples include the Google Speech Recognition API.
[0590] Natural Language Processing Model: Used to analyze the converted text data and generate appropriate answers. This model employs a generative AI model.
[0591] Database: Stores question-and-answer collections and FAQ data that are referenced when classifying text data and generating answer suggestions.
[0592] Display: This is an output device used to display the generated answer sheet to passengers. It is commonly installed on the dashboard of autonomous vehicles.
[0593] System processing flow
[0594] 1. Collection of voice data: The terminal (microphone inside the autonomous vehicle) collects the passenger's voice instructions. For example, questions such as "How long will it take to arrive at the next destination?" are collected.
[0595] 2. Audio Data Conversion: The device sends the collected audio data to the server, which uses an audio analysis API to convert it into text data. As a result, the text data "How long will it take to arrive at the next destination?" is generated.
[0596] 3. Text Data Classification: The server uses a natural language processing model to classify the converted text data into appropriate categories. For example, it might be classified as "Confirmation of Arrival Time."
[0597] 4. Solution Generation: Based on the classified category, the server retrieves appropriate information from the database and generates a solution using a generative AI model. For example, a solution such as "It will arrive in approximately 15 minutes" might be generated.
[0598] 5. Display of the solution: The generated solution is sent to the terminal (the display in the autonomous vehicle) and presented to the passengers visually.
[0599] Specific example
[0600] As a concrete example, consider the following conversation:
[0601] Passenger: "How long will it take to get to our next destination?"
[0602] System: "Arriving in approximately 15 minutes."
[0603] Based on this prompt, the system will process the information in the following order.
[0604] Example of a prompt:
[0605] Please convert the following audio data to text: "How long will it take to arrive at our next destination?"
[0606] Please categorize the following text data: "How long will it take to arrive at the next destination?"
[0607] Please generate an answer based on the following category and text data: "How long will it take to arrive at the next destination?"
[0608] These processes enable the infotainment system of an autonomous vehicle to respond quickly and accurately to voice commands from passengers.
[0609] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0610] Step 1:
[0611] Collection of audio data
[0612] Input: Passenger voice instructions
[0613] Operation: The terminal (microphone inside the autonomous vehicle) collects passenger voices in real time. For example, it recognizes voices such as, "How long will it take to arrive at the next destination?"
[0614] Output: Audio data file
[0615] Specific operation: The microphone records sound as digital data and passes the collected audio data to subsequent processing.
[0616] Step 2:
[0617] Audio data conversion
[0618] Input: Audio data file
[0619] Operation: The server uses a speech analysis API to convert speech data into text data. Google Speech Recognition is used as the API.
[0620] Output: Text data
[0621] Specific operation: Send audio data to the voice analysis API and retrieve the text "How long will it take to arrive at the next destination?". Pass this text data to the next processing step.
[0622] Step 3:
[0623] Classification of text data
[0624] Input: Text data
[0625] Operation: The server uses a natural language processing model to categorize text data into appropriate categories. For example, it might categorize it as "Confirming arrival time".
[0626] Output: Category Information
[0627] Specific operation: Text data is input into a natural language processing model, and the model outputs categories based on its learning results. This category information is then passed to the next step.
[0628] Step 4:
[0629] Generating the answer
[0630] Input: Category information and original text data
[0631] Operation: The server uses a database and a generative AI model to generate answer suggestions based on categories. For example, it might generate an answer such as "We will arrive in approximately 15 minutes."
[0632] Output: Solution
[0633] Specific operation: Relevant information (e.g., arrival time data) is retrieved from the database, and the optimal solution is generated using a generative AI model. This solution is then passed on to the next step.
[0634] Step 5:
[0635] Display of the answer
[0636] Input: Answer
[0637] Operation: The terminal (the display in the autonomous vehicle) displays the generated answer to the passengers. The display shows "We will arrive in approximately 15 minutes."
[0638] Output: Display of the answer sheet that passengers can see.
[0639] Specific operation: The answer will be displayed on a screen and presented in a format that is easy for passengers to read. This will allow passengers to see the answer to their question in real time.
[0640] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0641] This invention is a system that collects voice data, converts it into text data, classifies it, analyzes the user's emotions using an emotion engine, automatically generates a response, has a human review and revise the response, and finally sends it to the customer. This system makes full use of a voice analysis API, a natural language processing model, a database, and an emotion engine. The following describes the specific operation of this system, using the server, terminal, and user as subjects.
[0642] System program processing
[0643] Audio data collection and conversion
[0644] Terminal: The shop crew member's terminal records conversations with customers. Specifically, they press the record button on the terminal to start the conversation, and press the button again to stop recording when the conversation is finished. This recording file is sent from the terminal to the server.
[0645] Server: Processes the received audio data and calls the speech analysis API to convert the audio data into text data. Specifically, it sends the audio data to the speech analysis API and retrieves the corresponding text.
[0646] Text data classification and sentiment analysis
[0647] Server: The converted text data is input into a natural language processing (NLP) model to classify the inquiry content into specific categories. For example, it may be classified into "questions about product sizes," "inquiries about order status," "return procedures," etc.
[0648] Server: Sends classified text data to the sentiment engine to analyze the user's emotions. The sentiment engine identifies the user's emotions (e.g., joy, anger, sadness, etc.) from words and phrases in the text data.
[0649] Generating solutions and making emotional revisions.
[0650] Server: Automatically generates suggested answers based on classified inquiry content and sentiment analysis results. Specifically, it refers to the FAQ database and past answer history to create answers that take the user's feelings into consideration. For example, if the user looks dissatisfied with the inquiry "What are the dimensions of the product?", it will respond politely with something like, "Thank you for waiting. The dimensions of this product are 10cm high, 5cm wide, and 1cm deep."
[0651] Human verification and correction
[0652] Terminal: The AI-generated solution is displayed on the shop crew member's terminal. The shop crew member reviews this solution and makes corrections as needed.
[0653] User (Shop Crew): Carefully review the displayed answer to ensure its accuracy. If any information is missing or incorrect, manually correct it. Once corrections are complete, press the confirmation button to proceed to the next step.
[0654] Sending a formal response
[0655] Terminal: The shop crew will send the customer the corrected and revised answer. Specifically, they will press the "Send" button on the terminal and deliver the answer to the customer via email or chat system.
[0656] Specific example
[0657] Example: Questions about a product and sentiment analysis
[0658] 1. Terminal: Starts recording with Shop Crew: "Please tell me the size of product A."
[0659] 2. Server: Receives audio data and uses an audio analysis API to convert it into text data such as "Please tell me the size of product A."
[0660] 3. Server: Classified as a "product size inquiry" by the natural language processing model.
[0661] 4. Server: Sends text data to the emotion engine to analyze whether the user is experiencing anxiety.
[0662] 5. Server: Based on the sentiment analysis results, it generates a polite response: "Thank you for waiting. The dimensions of this product are 10cm high, 5cm wide, and 1cm deep. Do you have any further questions?"
[0663] 6. Terminal: The shop crew checks the answer and corrects it to "The size of this product is 10cm in height."
[0664] 7. Terminal: Press the "Send" button to send the corrected answer to the customer.
[0665] As described above, this system is intended to automate and streamline inquiry handling, enabling accurate and rapid responses. Furthermore, by combining it with user sentiment analysis using an emotion engine, it is possible to provide more personalized services.
[0666] The following describes the processing flow.
[0667] Step 1:
[0668] Terminal: The shop crew member's terminal records conversations with customers. Specifically, they press the record button on the terminal to start the conversation, and press the button again to stop recording when the conversation is finished. This recording file is sent from the terminal to the server.
[0669] Step 2:
[0670] Server: Processes the received audio data and calls the audio analysis API to convert the audio data into text data. Specifically, it sends the audio file to the audio analysis API and retrieves the corresponding text data.
[0671] Step 3:
[0672] Server: The converted text data is input into a natural language processing (NLP) model, which classifies the query into a specific category. For example, if the text data is "What are the product sizes?", it will be classified as "Questions about product sizes".
[0673] Step 4:
[0674] Server: After passing text data through a natural language processing model, the emotion engine analyzes the user's emotions. Specifically, it identifies emotions such as whether the user is happy or dissatisfied based on the words and phrases in the text data.
[0675] Step 5:
[0676] Server: Automatically generates suggested answers based on classified inquiry content and sentiment analysis results. Specifically, it refers to the FAQ database and past answer history to create answers that take the user's feelings into consideration. For example, if the user seems dissatisfied with the question "What are the dimensions of the product?", it will create a polite suggested answer such as, "We apologize for the wait. The dimensions of this product are 10cm high, 5cm wide, and 1cm deep."
[0677] Step 6:
[0678] Server: Sends the generated answer sheet to the shop crew's terminal. Specifically, it communicates with the terminal to send the text data of the answer sheet from the server.
[0679] Step 7:
[0680] Terminal: The shop crew's terminal displays the received answer on the screen.
[0681] Step 8:
[0682] User (Shop Crew): The Shop Crew will review the displayed solution and make corrections as needed. Specifically, they will check the content of the solution and manually correct any omissions or errors.
[0683] Step 9:
[0684] Terminal: After completing the corrections to the answer, press the "Submit" button for final confirmation and send it to the customer as the official answer.
[0685] Step 10:
[0686] Terminal: The official response will be delivered to the customer via email or chat system.
[0687] (Example 2)
[0688] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0689] In modern customer service, efficiently collecting voice data and providing quick and accurate answers is crucial. However, converting voice data into text, classifying the content, and generating appropriate answers is a time-consuming process. Furthermore, personalized responses that take into account user emotions are required, but this is difficult to achieve with current systems. Therefore, there is a need for a system that can efficiently process voice data and provide answers that respond to user emotions.
[0690] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0691] In this invention, the server includes means for collecting voice data, means for converting the voice data into text data, means for classifying the text data, means for analyzing the user's emotions based on the classified text data, means for generating answer suggestions based on the analyzed emotions and classification results, means for a human to review and correct the generated answer suggestions, and means for transmitting the reviewed and corrected answer suggestions. This enables efficient processing of voice data and the provision of personalized answers that respond to the user's emotions in customer service operations.
[0692] "Audio data" refers to information recorded in digital format using audio.
[0693] "Text data" refers to data obtained by converting audio data into written text.
[0694] "Classifying" refers to dividing data into categories based on specific criteria.
[0695] "Analyzing emotions" refers to detecting and identifying a user's emotional state from text data.
[0696] "Answer" refers to the content of the response to a user's inquiry.
[0697] "Review and correct" means that a human reviews the generated answer and makes corrections as necessary.
[0698] "Sending" refers to using communication methods to deliver the revised answer to the user.
[0699] This system implements a series of processes: collecting voice data, converting it to text data, classifying it, analyzing the user's emotions using an emotion engine, automatically generating suggested answers, having humans review and revise those answers, and finally sending them to the customer. The system utilizes a voice analysis API, natural language processing models, a database, and an emotion engine.
[0700] Hardware and software details
[0701] The following main hardware and software will be used to implement the system.
[0702] Device: A smartphone or tablet used by the shop crew to record conversations.
[0703] Server: A high-performance server used for processing and storing data.
[0704] Speech analysis APIs: Software for converting speech data into text data, such as the Google Cloud Speech-to-Text API.
[0705] Natural Language Processing Models (NLP models): Software such as OpenAI GPT-4 used to analyze and classify text data.
[0706] Emotion engine: Software such as IBM Watson that analyzes emotions from text.
[0707] Database: A database system for storing and managing various types of data (audio data, text data, analysis results, answer sheets, etc.).
[0708] Specific operation of the system
[0709] The terminal is used by shop staff to record conversations with customers. They press the record button to start the conversation and press it again to stop recording when finished. This recording file is then sent from the terminal to the server.
[0710] The server sends the received audio data to the Google Cloud Speech-to-Text API, where it converts the audio data into text data. The converted text data is then classified by a natural language processing model such as OpenAI GPT-4. For example, it may be categorized into "questions about product sizes," "inquiries about order status," and "return procedures."
[0711] Furthermore, the classified text data is sent to IBM Watson's sentiment engine, where the user's emotions are analyzed. Sentiment analysis is performed based on words and phrases within the text data to identify the user's emotions (e.g., joy, anger, sadness, etc.).
[0712] The server automatically generates suggested answers based on the classification results and sentiment analysis results. At this stage, it refers to the FAQ database and past answer history to create answers that take the user's emotions into consideration. For example, if a user asks "What are the product sizes?" and sentiment analysis detects anxiety, the prompt to the generating AI model will be "Answer when the user seems anxious: 'What are the product sizes?'" to generate a polite answer.
[0713] The generated answer is displayed on the terminal, where a shop crew member reviews and corrects it. They carefully examine whether the answer is correct and make corrections as needed. Once the review is complete, the answer is sent to the customer by pressing the "Send" button on the terminal. This entire process automates and streamlines inquiry handling, enabling appropriate responses that are tailored to the user's emotions.
[0714] Specific example
[0715] As an example, the flow of questions about a product and sentiment analysis is shown below.
[0716] 1. Terminal: The shop crew member starts recording, saying, "Please tell me the size of product A."
[0717] 2. Server: Receives the audio data and converts it into text data, "Please tell me the size of product A," using the Google Cloud Speech-to-Text API.
[0718] 3. Server: Use OpenAI GPT-4 to classify text data into "product size inquiries".
[0719] 4. Server: Sends text data to IBM Watson's sentiment engine to analyze whether the user is experiencing anxiety.
[0720] 5. Server: Based on the sentiment analysis results, it generates a polite response: "Thank you for waiting. The dimensions of this product are 10cm high, 5cm wide, and 1cm deep. Do you have any further questions?"
[0721] 6. Terminal: The shop crew checks the answer and corrects it to "The size of this product is 10cm in height."
[0722] 7. Terminal: Press the "Send" button to send the corrected answer to the customer.
[0723] In this way, the entire system ensures that a series of processes run smoothly, improving the efficiency and effectiveness of customer service operations.
[0724] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0725] Step 1:
[0726] Collection of audio data
[0727] Subject: terminal
[0728] Specific actions: The shop crew member presses the record button on the terminal to record the conversation with the customer. After the conversation ends, they press the button again to stop recording.
[0729] Input: Audio of a conversation between a shop crew member and a customer.
[0730] Output: Recorded audio file (saved in digital format)
[0731] Step 2:
[0732] Sending audio data
[0733] Subject: terminal
[0734] Specific operation: The recorded audio file is automatically uploaded to the server. A notification is displayed after the upload is complete.
[0735] Input: Recorded audio file
[0736] Output: Audio file sent to the server
[0737] Step 3:
[0738] Saving audio data
[0739] Subject: Server
[0740] Specific action: The received audio data is saved to the server's storage.
[0741] Input: Audio file sent from the device
[0742] Output: Audio data stored on the server
[0743] Step 4:
[0744] Converting audio data to text
[0745] Subject: Server
[0746] Specific operation: The saved audio data is sent to the Google Cloud Speech-to-Text API and converted into text data.
[0747] Input: Audio data on the server
[0748] Output: Text data returned from the Google Cloud Speech-to-Text API
[0749] Step 5:
[0750] Saving text data
[0751] Subject: Server
[0752] Specific action: Save the converted text data to the database.
[0753] Input: Text data
[0754] Output: Text data stored in the database
[0755] Step 6:
[0756] Classification of text data
[0757] Subject: Server
[0758] Specific operation: Text data is input into a natural language processing model such as OpenAI GPT-4, and the query content is classified.
[0759] Input: Text data read from a database
[0760] Output: Classification results returned from the natural language processing model
[0761] Step 7:
[0762] Saving classification results
[0763] Subject: Server
[0764] Specific action: Save the classification results to the database.
[0765] Input: Classification result
[0766] Output: Classification results stored in the database
[0767] Step 8:
[0768] Emotion analysis
[0769] Subject: Server
[0770] Specific operation: Classified text data is sent to IBM Watson's sentiment engine for sentiment analysis.
[0771] Input: Classified text data
[0772] Output: Sentiment analysis results returned from IBM Watson
[0773] Step 9:
[0774] Saving emotion analysis results
[0775] Subject: Server
[0776] Specific action: Save the emotion analysis results to the database.
[0777] Input: Sentiment analysis results
[0778] Output: Sentiment analysis results stored in the database
[0779] Step 10:
[0780] Generating the answer
[0781] Subject: Server
[0782] Specific operation: Automatically generates suggested answers based on classification results and sentiment analysis results. Refers to the FAQ database and past answer history to generate answers that take the user's emotions into consideration.
[0783] Input: Classification results and sentiment analysis results
[0784] Output: Solution returned by the generative AI model
[0785] Step 11:
[0786] Display of the answer
[0787] Subject: terminal
[0788] Specific action: Display the generated solution on the shop crew's terminal.
[0789] Input: Generated solution
[0790] Output: Answer displayed on the terminal
[0791] Step 12:
[0792] Review and correct the answer sheet.
[0793] Subject: User (Shop Crew)
[0794] Specific actions: The shop crew will review the displayed answer and make any necessary corrections. Once the corrections are complete, they will press the confirmation button.
[0795] Input: Displayed answer
[0796] Output: Revised solution
[0797] Step 13:
[0798] Sending a formal response
[0799] Subject: terminal
[0800] Specific actions: Send the corrected answer to the customer. Press the "Send" button on the device to deliver the answer to the customer via email or chat system.
[0801] Input: Revised answer
[0802] Output: Official response sent to the customer
[0803] (Application Example 2)
[0804] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0805] Currently, customer support operations, including food delivery services, require prompt and appropriate responses to customer inquiries. However, manual responses can lead to delays and decreased customer satisfaction. Furthermore, responses tend to be formulaic, making it difficult to provide personalized responses that address customer emotions and circumstances. Therefore, there is a need to achieve both automated and personalized responses to inquiries.
[0806] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0807] In this invention, the server includes means for collecting voice data, means for converting voice data into text data, means for classifying the text data, means for analyzing emotions based on the text data, means for generating answer suggestions based on the classified text data and emotion analysis results, means for a human to review and correct the generated answer suggestions, and means for transmitting the reviewed and corrected answer suggestions. This enables rapid and appropriate automated responses in customer support operations and makes it possible to provide personalized answers that take into account the customer's emotions.
[0808] "Audio data" refers to data that records audio in digital format.
[0809] "Text data" refers to string information converted from audio data.
[0810] A "Voice Analysis API" is an application programming interface for converting voice data into text data.
[0811] A "natural language processing model" is a machine learning or deep learning model used to analyze, classify, and interpret text data.
[0812] A "sentiment analysis engine" is an algorithm or system that identifies the user's emotions contained in text data.
[0813] A "database" is a system for efficiently managing, searching, and updating information.
[0814] A "solution" is a written response to a customer's inquiry.
[0815] "Customer support" refers to the work of responding to inquiries and requests from customers.
[0816] A "server" is a computer system used for storing and processing data.
[0817] A "food delivery service" is a service that delivers food ordered online by customers to a specified location.
[0818] "Personalization" refers to optimizing responses according to the individual needs and circumstances of each customer.
[0819] "Classification of inquiry content" is the process of assigning text data to a specific category.
[0820] "Emotion analysis results" refer to emotional information identified by the emotion analysis engine from text data.
[0821] This invention is a system that collects voice data, converts it into text data, classifies it, analyzes the user's emotions using an emotion engine, automatically generates a response, has a human review and revise the response, and finally sends it to the customer. This system is realized by utilizing a voice analysis API, a natural language processing model, a database, and an emotion engine.
[0822] System Configuration
[0823] 1. Collection of audio data
[0824] Device: The customer support representative's device (smartphone or computer) will record the conversation with the customer. They will press the record button to start the conversation and press it again to stop recording at the end.
[0825] 2. Converting audio data
[0826] Server: Processes the received audio data and converts it into text data by calling a speech analysis API. Specifically, it sends the collected audio data to a speech analysis API, such as Google's speech recognition API, and retrieves the corresponding text.
[0827] 3. Classification and sentiment analysis of text data
[0828] Server: The converted text data is input into a natural language processing (NLP) model to classify the inquiry content into specific categories. For example, it may be classified as "order status confirmation," "complaint handling," or "product-related questions."
[0829] Server: Sends classified text data to the sentiment engine to analyze the user's emotions. The sentiment engine identifies the user's emotions (e.g., joy, anger, sadness, etc.) from words and phrases in the text data.
[0830] 4. Generating solutions and making emotional revisions
[0831] Server: Automatically generates suggested answers based on classified inquiry content and sentiment analysis results. Specifically, it refers to the FAQ database and past answer history to create answers that take the user's emotions into consideration. For example, if a user expresses anger in response to an inquiry about a delayed delivery, it will respond politely with something like, "We apologize. We will take immediate action to resolve your dissatisfaction. Could you please provide more details?"
[0832] 5. Human review and correction
[0833] Terminal: The AI-generated solution is displayed on the customer support representative's terminal. The customer support representative reviews this solution and makes corrections as needed.
[0834] User (Customer Support Representative): Review the displayed solution to ensure its accuracy. If any information is missing or incorrect, manually correct it. Once corrections are complete, press the confirmation button to proceed to the next step.
[0835] 6. Submitting the formal response
[0836] Terminal: Customer support staff will send the reviewed and corrected answers to the customer. Specifically, they will press the "Send" button on the terminal to deliver the answer to the customer via email or chat system.
[0837] Specific example
[0838] As an example, consider a case where a customer inquires that their product delivery is delayed. A customer support representative records this inquiry and sends the audio data to a server. The server converts the audio into text data and categorizes it as "order status inquiry." The emotion analysis then identifies "anger." Based on this, a response is generated such as, "We apologize. We will promptly address your issue to resolve it. Could you please provide more details?"
[0839] Example of a prompt:
[0840] "Text: 'My delivery is delayed and I'm having trouble. What's going on?' Please categorize the content. Estimate the category and analyze the sentiment."
[0841] This system enables the automation and streamlining of inquiry handling, allowing for customer service that is sensitive to customer emotions.
[0842] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0843] Step 1:
[0844] Collection of audio data
[0845] Terminal: Customer support staff record conversations with customers. They press the record button to start collecting audio data, and press the button again to stop recording when the conversation ends. This saves the audio file to the terminal. This audio file becomes the input data for subsequent processing.
[0846] Step 2:
[0847] Sending and converting audio data
[0848] Server: Receives audio files sent from the terminal. The received audio files are sent to a speech analysis API for conversion into text data. Specifically, the speech analysis API analyzes the audio signal and outputs it as string data. The audio file is the input data, and the text data is the output data.
[0849] Step 3:
[0850] Classification of text data
[0851] Server: The converted text data is input into a natural language processing (NLP) model, which classifies the inquiry content into categories. This model is a pre-trained machine learning model that automatically assigns text data to categories such as "order status confirmation," "complaint handling," and "product-related questions." Text data is the input data, and category information is the output data.
[0852] Step 4:
[0853] Emotion analysis
[0854] Server: Sends classified text data to the sentiment engine to analyze the user's emotions. The sentiment engine identifies the user's emotions (e.g., joy, anger, sadness, etc.) from words and phrases in the text data. The text data is the input data for sentiment analysis, and the emotional information is the output data.
[0855] Step 5:
[0856] Automatic generation of answer sheets
[0857] Server: Automatically generates suggested answers based on classified category information and sentiment analysis results. Specifically, it refers to the FAQ database and past answer history to obtain the corresponding standard answer. Then, considering the sentiment analysis results, it generates a response that includes a response tailored to the user's emotions. Category information and sentiment information are input data, and suggested answers are output data.
[0858] Step 6:
[0859] Review and correction of the answer sheet
[0860] Terminal: The AI-generated answer is displayed on the customer support representative's terminal. The representative reviews this answer and makes corrections as needed. Once corrections are complete, they press the confirmation button to proceed to the next step. The initial answer is the input data, and the reviewed and corrected answer is the output data.
[0861] Step 7:
[0862] Sending a formal response
[0863] Terminal: Customer support staff send the reviewed and corrected answers to the customer. Specifically, they press the "Send" button on the terminal to deliver the answer to the customer via email or chat system. The corrected answer is the input data, and the answer sent to the customer is the output data.
[0864] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0865] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0866] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0867] [Third Embodiment]
[0868] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0869] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0870] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0871] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0872] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0873] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0874] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0875] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0876] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0877] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0878] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0879] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0880] This invention is a system that collects voice data, converts it into text data, classifies it, automatically generates answer keys, allows a human to review and correct the answer keys, and finally sends them to the customer. This system uses a voice analysis API and makes full use of natural language processing models and databases. The following describes the specific operation of this system, using the server, terminal, and user as subjects.
[0881] System program processing
[0882] Audio data collection and conversion
[0883] Terminal: The shop crew member's terminal records conversations with customers. For example, they press the record button on the terminal to start a conversation, and press the button again to stop recording when the conversation is finished. This recording file is sent from the terminal to the server.
[0884] Server: Processes the received audio data and calls the speech analysis API to convert it into text data. Specifically, it sends the audio data to the speech analysis API and retrieves the corresponding text.
[0885] Classification of text data
[0886] Server: The converted text data is passed through a natural language processing model to classify the query content into specific categories. For example, it may be classified as "questions about product sizes," "order status checks," or "return / exchange procedures." This classification is performed automatically based on the model's training results.
[0887] Generating the answer
[0888] Server: Automatically generates suggested answers based on classified inquiry content. Here, the server refers to FAQ data and past answer history in the database to create an appropriate answer. For example, in response to the inquiry "What are the dimensions of the product?", it generates a suggested answer such as "The dimensions of this product are 10cm high, 5cm wide, and 1cm deep."
[0889] Human verification and correction
[0890] Terminal: The AI-generated solution is displayed on the shop crew member's terminal. The shop crew member reviews this solution and makes corrections as needed.
[0891] User (Shop Crew): Carefully review the displayed answer to ensure its accuracy. If any information is missing or incorrect, manually correct it. Once corrections are complete, press the confirmation button to proceed to the next step.
[0892] Sending a formal response
[0893] Terminal: The shop crew will review and correct the answer and finally send it to the customer. For example, by pressing the "Send" button on the terminal, the answer will be delivered to the customer via email or chat system.
[0894] Specific example
[0895] Example: Questions about products
[0896] 1. Terminal: Starts recording with Shop Crew: "Please tell me the size of product A."
[0897] 2. Server: Receives audio data and uses an audio analysis API to convert it into text data such as "Please tell me the size of product A."
[0898] 3. Server: Classified as a "product size inquiry" by the natural language processing model.
[0899] 4. Server: Retrieves size information from the product database and generates a suggested answer such as, "The dimensions of this product are 10cm high, 5cm wide, and 1cm deep."
[0900] 5. Terminal: The shop crew checks the answer and corrects it to "The size of this product is 10cm in height."
[0901] 6. Terminal: Press the "Send" button to send the corrected answer to the customer.
[0902] As described above, this system is intended to automate and streamline inquiry handling, enabling accurate and prompt responses.
[0903] The following describes the processing flow.
[0904] Step 1:
[0905] Terminal: The shop crew member's terminal records conversations with customers. Specifically, they press the record button on the terminal to start the conversation. When the conversation ends, they press the button again to stop recording and send the recording file to the server.
[0906] Step 2:
[0907] Server: Receives audio data and converts it into text data using a speech analysis API. Specifically, it sends the received audio file to the speech analysis API and retrieves the corresponding text data.
[0908] Step 3:
[0909] Server: Inputs the transcribed data into a natural language processing (NLP) model and classifies the inquiry content into specific categories. For example, it automatically classifies inquiries into categories such as "questions about product sizes," "order status inquiries," and "return procedures."
[0910] Step 4:
[0911] Server: Automatically generates answers based on categorized inquiries. Specifically, it refers to the FAQ database and past answer history to select or create appropriate answers to generate answers corresponding to categories.
[0912] Step 5:
[0913] Server: Sends the generated answer sheet to the shop crew's terminal. Specifically, it communicates with the terminal to send the text data of the answer sheet from the server.
[0914] Step 6:
[0915] Terminal: The shop crew's terminal displays the received answer on the screen.
[0916] Step 7:
[0917] User (Shop Crew): The Shop Crew will review the displayed solution and make corrections as needed. Specifically, they will check the content of the solution and manually correct any omissions or errors.
[0918] Step 8:
[0919] Terminal: After completing the corrections to the answer, press the "Submit" button for final confirmation and send it to the customer as the official answer.
[0920] Step 9:
[0921] Terminal: The official response will be delivered to the customer via email or chat system.
[0922] (Example 1)
[0923] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0924] Traditional customer service systems often involved manual processes such as collecting and analyzing voice data, converting it to text data, and generating appropriate answers. This resulted in time-consuming and inefficient responses. Furthermore, the quality of answers depended on the individual operator's judgment, potentially leading to inconsistencies. This resulted in challenges such as decreased customer satisfaction and increased operating costs.
[0925] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0926] In this invention, the server includes means for converting audio data into text data using an audio analysis API, means for classifying the text data using a natural language processing model, and means for generating answer suggestions based on the classified text data by referring to a database. This automates the entire process from collecting audio data to generating answer suggestions and sending them to customers, enabling efficient and accurate inquiry handling.
[0927] A "terminal" is a device operated by the user to collect voice data and to review and correct answer sheets.
[0928] A "server" is a computer system that performs a series of operations: receiving audio data, converting it into text data using an audio analysis API, classifying it using a natural language processing model, and generating a solution.
[0929] "Audio data" refers to digital audio information, which is recordings of human speech collected using a device.
[0930] A "Voice Analysis API" is an application programming interface for analyzing input voice data and converting it into text data.
[0931] "Text data" refers to the character information corresponding to the audio data, converted by the speech analysis API.
[0932] A "natural language processing model" is a machine learning model that analyzes text data, understands its content, and classifies it into specific categories.
[0933] A "solution" is a document that responds to a customer inquiry, generated based on classified text data.
[0934] A "database" is a storage device that stores information such as FAQ data and past answer history necessary for generating answer templates.
[0935] This invention is a system that collects audio data, converts it into text data, classifies it, automatically generates answer keys, has a human review and correct the answer keys, and finally sends them to the customer. This system uses an audio analysis API and makes full use of natural language processing models and databases.
[0936] Audio data collection and conversion
[0937] Terminal: The terminal used by the user is a device for recording conversations with customers. Specifically, the user launches the recording application and presses the record button to start the conversation. When the conversation ends, the user presses the record button again to stop recording, and this recording file is temporarily saved to the terminal's local storage. This audio file is then sent to the server.
[0938] Server: Receives audio data sent from the terminal. The received audio file is temporarily stored, and this audio data is converted into text data using a speech analysis API such as the Google Cloud Speech-to-Text API. Specifically, the audio data is sent to the API endpoint, and the text data is retrieved from the returned response. This text data is stored in a database and used in the next step.
[0939] Classification of text data
[0940] Server: The acquired text data is input into a natural language processing model (e.g., BERT). The model analyzes the content of the text data and classifies the query into specific categories. For example, it might classify queries into "questions about product sizes," "order status checks," or "return / exchange procedures." This classification result is stored in a database and used in the next step of generating solutions.
[0941] Generating the answer
[0942] Server: Automatically generates answers based on classified category information. The server refers to FAQ data and past answer history in the database to create appropriate answers. For example, in response to the inquiry "What are the dimensions of the product?", it generates an answer such as "The dimensions of this product are 10cm high, 5cm wide, and 1cm deep." This generated answer is stored in the database and sent to the terminal in the next step.
[0943] Human verification and correction
[0944] Terminal: The answer submitted from the server is displayed on the user's terminal. The user can use a UI to review this answer and modify it as needed. Specifically, a text field is provided that allows the user to edit the displayed answer.
[0945] User: Check if the displayed answer is correct, and manually correct any missing or incorrect information. Once you have finished correcting, press the "Confirm" button to proceed to the next step.
[0946] Sending a formal response
[0947] Terminal: The user submits the completed and corrected answer to the customer by pressing the "Submit" button. This submission is carried out using email, chat API, etc. After submission is complete, the system displays a confirmation message to the customer.
[0948] Specific example
[0949] 1. Terminal: The user starts recording by saying, "Please tell me the size of product A."
[0950] 2. Server: Receives the audio data and converts it into text data, "Please tell me the size of product A," using the Google Cloud Speech-to-Text API.
[0951] 3. Server: The query is classified as "a query about product size" using the natural language processing model BERT.
[0952] 4. Server: Retrieves size information from the product database and generates an answer such as, "The dimensions of this product are 10cm high, 5cm wide, and 1cm deep."
[0953] 5. Terminal: The user reviews the answer and corrects it to "The size of this product is 10cm in height."
[0954] 6. Terminal: Press the "Send" button to send the corrected answer to the customer.
[0955] This will automate and streamline inquiry handling, enabling quick and accurate responses.
[0956] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0957] Step 1: Collect audio data
[0958] Terminal: This device records conversations between customers and shop crew. The user launches the recording application and presses the record button to begin the conversation. When the conversation ends, they press the record button again to stop recording. This audio file is temporarily stored in the terminal's local storage.
[0959] Input: Audio of a conversation between a customer and a shop crew member.
[0960] Output: Temporarily saved audio data file.
[0961] Step 2: Sending the audio data
[0962] Terminal: Sends the collected audio data to the server. Specifically, it uses the terminal's network module to upload the audio files to a specified URL on the server.
[0963] Input: Temporarily saved audio data file.
[0964] Output: Audio data sent to the server.
[0965] Step 3: Convert speech to text
[0966] Server: Converts received audio data into text data using a speech analysis API (e.g., Google Cloud Speech-to-Text API). Specifically, it sends the audio data to the API endpoint and extracts the text data from the returned response.
[0967] Input: Audio data sent to the server.
[0968] Output: Text data.
[0969] Step 4: Classification of Text Data
[0970] Server: The acquired text data is input into a natural language processing model (e.g., BERT) for classification. Specifically, the model is used to analyze the content of the text data and classify it into predefined categories. For example, it may be classified into "questions about product sizes," "order status confirmation," and "return / exchange procedures."
[0971] Input: Text data.
[0972] Output: Classification results (category labels).
[0973] Step 5: Generating the solution
[0974] Server: Based on the classified category, it searches the database for appropriate answers and generates suggested solutions. Specifically, it refers to FAQ data and past answer history to create suggested solutions that match the content of the inquiry.
[0975] Input: Classification result (category label).
[0976] Output: Generated solution.
[0977] Step 6: Submit your answer
[0978] Server: Sends the generated solution to the user's terminal. Specifically, it uploads the solution data to a specified URL on the user's terminal.
[0979] Input: Generated solution.
[0980] Output: The answer submitted to the user's terminal.
[0981] Step 7: Human review and correction
[0982] Terminal: The answer submitted from the server is displayed on the user's terminal. The user reviews this answer and makes corrections as needed.
[0983] User: Check the displayed solution, enter any necessary corrections using the text field, and save the revised version.
[0984] Input: Submitted answer.
[0985] Output: Revised solution.
[0986] Step 8: Submit your formal response
[0987] Terminal: This terminal sends the user's reviewed and corrected answer to the customer. Specifically, the user presses the "Send" button and sends the answer via email or chat API.
[0988] Input: Revised answer.
[0989] Output: The answer submitted to the customer.
[0990] Through the processing steps described above, this system automates the entire process from collecting voice data to generating answers and sending them to customers, enabling efficient and accurate handling of inquiries.
[0991] (Application Example 1)
[0992] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0993] In infotainment systems for autonomous vehicles, it is essential that passengers' voice-activated questions and requests are answered quickly and accurately. Conventional systems require human intervention, which can lead to delays and difficulties in providing accurate answers. Therefore, there is a need for systems in autonomous vehicles that can provide appropriate answers to voice commands in real time.
[0994] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0995] In this invention, the server includes means for collecting voice data, means for converting voice data into text data, and means for classifying text data. This makes it possible to generate and provide quick and accurate answers in real time to passenger voice instructions and requests in the infotainment system of an autonomous vehicle.
[0996] "Means for collecting voice data" refers to the function in the infotainment system of an autonomous vehicle that recognizes passenger voices and collects them as digital data.
[0997] "Methods for converting audio data into text data" refers to the process of analyzing collected audio data and converting it into corresponding text data using an audio analysis API or similar method.
[0998] "Means for classifying text data" refers to a function that automatically classifies converted text data into appropriate categories using a natural language processing model.
[0999] "Methods for generating answers" refers to the process of generating appropriate responses using natural language processing models, based on classified text data, referencing databases and FAQ data.
[1000] "Means for human review and correction" refers to a function where a human reviews the automatically generated answer and corrects its content as needed.
[1001] "Means for sending confirmed and revised answers" refers to the function of sending the final confirmed and revised answers to passengers via email or chat system.
[1002] A "voice recognition system installed in an autonomous vehicle" is a system installed in an autonomous vehicle that recognizes and processes passengers' voices in real time.
[1003] A "natural language processing model" is a machine learning model that understands and analyzes text data to generate appropriate responses or actions.
[1004] An "infotainment system" is a system installed in autonomous vehicles that provides information and entertainment to passengers using voice and video.
[1005] The system for implementing this invention operates as an infotainment system for an autonomous vehicle. When a passenger gives instructions or asks a question by voice, the system collects the voice, converts it into text data, and provides an appropriate answer in real time. The main hardware and software configuration is described in detail below.
[1006] Major hardware and software
[1007] Microphone: This is an input device for collecting passenger voices. It is installed inside the autonomous vehicle and is always on standby for passenger voice commands.
[1008] Speech analysis APIs are used to convert collected speech data into text data. Examples include the Google Speech Recognition API.
[1009] Natural Language Processing Model: Used to analyze the converted text data and generate appropriate answers. This model employs a generative AI model.
[1010] Database: Stores question-and-answer collections and FAQ data that are referenced when classifying text data and generating answer suggestions.
[1011] Display: This is an output device used to display the generated answer sheet to passengers. It is commonly installed on the dashboard of autonomous vehicles.
[1012] System processing flow
[1013] 1. Collection of voice data: The terminal (microphone inside the autonomous vehicle) collects the passenger's voice instructions. For example, questions such as "How long will it take to arrive at the next destination?" are collected.
[1014] 2. Audio Data Conversion: The device sends the collected audio data to the server, which uses an audio analysis API to convert it into text data. As a result, the text data "How long will it take to arrive at the next destination?" is generated.
[1015] 3. Text Data Classification: The server uses a natural language processing model to classify the converted text data into appropriate categories. For example, it might be classified as "Confirmation of Arrival Time."
[1016] 4. Solution Generation: Based on the classified category, the server retrieves appropriate information from the database and generates a solution using a generative AI model. For example, a solution such as "It will arrive in approximately 15 minutes" might be generated.
[1017] 5. Display of the solution: The generated solution is sent to the terminal (the display in the autonomous vehicle) and presented to the passengers visually.
[1018] Specific example
[1019] As a concrete example, consider the following conversation:
[1020] Passenger: "How long will it take to get to our next destination?"
[1021] System: "Arriving in approximately 15 minutes."
[1022] Based on this prompt, the system will process the information in the following order.
[1023] Example of a prompt:
[1024] Please convert the following audio data to text: "How long will it take to arrive at our next destination?"
[1025] Please categorize the following text data: "How long will it take to arrive at the next destination?"
[1026] Please generate an answer based on the following category and text data: "How long will it take to arrive at the next destination?"
[1027] These processes enable the infotainment system of an autonomous vehicle to respond quickly and accurately to voice commands from passengers.
[1028] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1029] Step 1:
[1030] Collection of audio data
[1031] Input: Passenger voice instructions
[1032] Operation: The terminal (microphone inside the autonomous vehicle) collects passenger voices in real time. For example, it recognizes voices such as, "How long will it take to arrive at the next destination?"
[1033] Output: Audio data file
[1034] Specific operation: The microphone records sound as digital data and passes the collected audio data to subsequent processing.
[1035] Step 2:
[1036] Audio data conversion
[1037] Input: Audio data file
[1038] Operation: The server uses a speech analysis API to convert speech data into text data. Google Speech Recognition is used as the API.
[1039] Output: Text data
[1040] Specific operation: Send audio data to the voice analysis API and retrieve the text "How long will it take to arrive at the next destination?". Pass this text data to the next processing step.
[1041] Step 3:
[1042] Classification of text data
[1043] Input: Text data
[1044] Operation: The server uses a natural language processing model to categorize text data into appropriate categories. For example, it might categorize it as "Confirming arrival time".
[1045] Output: Category Information
[1046] Specific operation: Text data is input into a natural language processing model, and the model outputs categories based on its learning results. This category information is then passed to the next step.
[1047] Step 4:
[1048] Generating the answer
[1049] Input: Category information and original text data
[1050] Operation: The server uses a database and a generative AI model to generate answer suggestions based on categories. For example, it might generate an answer such as "We will arrive in approximately 15 minutes."
[1051] Output: Solution
[1052] Specific operation: Relevant information (e.g., arrival time data) is retrieved from the database, and the optimal solution is generated using a generative AI model. This solution is then passed on to the next step.
[1053] Step 5:
[1054] Display of the answer
[1055] Input: Answer
[1056] Operation: The terminal (the display in the autonomous vehicle) displays the generated answer to the passengers. The display shows "We will arrive in approximately 15 minutes."
[1057] Output: Display of the answer sheet that passengers can see.
[1058] Specific operation: The answer will be displayed on a screen and presented in a format that is easy for passengers to read. This will allow passengers to see the answer to their question in real time.
[1059] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1060] This invention is a system that collects voice data, converts it into text data, classifies it, analyzes the user's emotions using an emotion engine, automatically generates a response, has a human review and revise the response, and finally sends it to the customer. This system makes full use of a voice analysis API, a natural language processing model, a database, and an emotion engine. The following describes the specific operation of this system, using the server, terminal, and user as subjects.
[1061] System program processing
[1062] Audio data collection and conversion
[1063] Terminal: The shop crew member's terminal records conversations with customers. Specifically, they press the record button on the terminal to start the conversation, and press the button again to stop recording when the conversation is finished. This recording file is sent from the terminal to the server.
[1064] Server: Processes the received audio data and calls the speech analysis API to convert the audio data into text data. Specifically, it sends the audio data to the speech analysis API and retrieves the corresponding text.
[1065] Text data classification and sentiment analysis
[1066] Server: The converted text data is input into a natural language processing (NLP) model to classify the inquiry content into specific categories. For example, it may be classified into "questions about product sizes," "inquiries about order status," "return procedures," etc.
[1067] Server: Sends classified text data to the sentiment engine to analyze the user's emotions. The sentiment engine identifies the user's emotions (e.g., joy, anger, sadness, etc.) from words and phrases in the text data.
[1068] Generating solutions and making emotional revisions.
[1069] Server: Automatically generates suggested answers based on classified inquiry content and sentiment analysis results. Specifically, it refers to the FAQ database and past answer history to create answers that take the user's feelings into consideration. For example, if the user looks dissatisfied with the inquiry "What are the dimensions of the product?", it will respond politely with something like, "Thank you for waiting. The dimensions of this product are 10cm high, 5cm wide, and 1cm deep."
[1070] Human verification and correction
[1071] Terminal: The AI-generated solution is displayed on the shop crew member's terminal. The shop crew member reviews this solution and makes corrections as needed.
[1072] User (Shop Crew): Carefully review the displayed answer to ensure its accuracy. If any information is missing or incorrect, manually correct it. Once corrections are complete, press the confirmation button to proceed to the next step.
[1073] Sending a formal response
[1074] Terminal: The shop crew will send the customer the corrected and revised answer. Specifically, they will press the "Send" button on the terminal and deliver the answer to the customer via email or chat system.
[1075] Specific example
[1076] Example: Questions about a product and sentiment analysis
[1077] 1. Terminal: Starts recording with Shop Crew: "Please tell me the size of product A."
[1078] 2. Server: Receives audio data and uses an audio analysis API to convert it into text data such as "Please tell me the size of product A."
[1079] 3. Server: Classified as a "product size inquiry" by the natural language processing model.
[1080] 4. Server: Sends text data to the emotion engine to analyze whether the user is experiencing anxiety.
[1081] 5. Server: Based on the sentiment analysis results, it generates a polite response: "Thank you for waiting. The dimensions of this product are 10cm high, 5cm wide, and 1cm deep. Do you have any further questions?"
[1082] 6. Terminal: The shop crew checks the answer and corrects it to "The size of this product is 10cm in height."
[1083] 7. Terminal: Press the "Send" button to send the corrected answer to the customer.
[1084] As described above, this system is intended to automate and streamline inquiry handling, enabling accurate and rapid responses. Furthermore, by combining it with user sentiment analysis using an emotion engine, it is possible to provide more personalized services.
[1085] The following describes the processing flow.
[1086] Step 1:
[1087] Terminal: The shop crew member's terminal records conversations with customers. Specifically, they press the record button on the terminal to start the conversation, and press the button again to stop recording when the conversation is finished. This recording file is sent from the terminal to the server.
[1088] Step 2:
[1089] Server: Processes the received audio data and calls the audio analysis API to convert the audio data into text data. Specifically, it sends the audio file to the audio analysis API and retrieves the corresponding text data.
[1090] Step 3:
[1091] Server: The converted text data is input into a natural language processing (NLP) model, which classifies the query into a specific category. For example, if the text data is "What are the product sizes?", it will be classified as "Questions about product sizes".
[1092] Step 4:
[1093] Server: After passing text data through a natural language processing model, the emotion engine analyzes the user's emotions. Specifically, it identifies emotions such as whether the user is happy or dissatisfied based on the words and phrases in the text data.
[1094] Step 5:
[1095] Server: Automatically generates suggested answers based on classified inquiry content and sentiment analysis results. Specifically, it refers to the FAQ database and past answer history to create answers that take the user's feelings into consideration. For example, if the user seems dissatisfied with the question "What are the dimensions of the product?", it will create a polite suggested answer such as, "We apologize for the wait. The dimensions of this product are 10cm high, 5cm wide, and 1cm deep."
[1096] Step 6:
[1097] Server: Sends the generated answer sheet to the shop crew's terminal. Specifically, it communicates with the terminal to send the text data of the answer sheet from the server.
[1098] Step 7:
[1099] Terminal: The shop crew's terminal displays the received answer on the screen.
[1100] Step 8:
[1101] User (Shop Crew): The Shop Crew will review the displayed solution and make corrections as needed. Specifically, they will check the content of the solution and manually correct any omissions or errors.
[1102] Step 9:
[1103] Terminal: After completing the corrections to the answer, press the "Submit" button for final confirmation and send it to the customer as the official answer.
[1104] Step 10:
[1105] Terminal: The official response will be delivered to the customer via email or chat system.
[1106] (Example 2)
[1107] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1108] In modern customer service, efficiently collecting voice data and providing quick and accurate answers is crucial. However, converting voice data into text, classifying the content, and generating appropriate answers is a time-consuming process. Furthermore, personalized responses that take into account user emotions are required, but this is difficult to achieve with current systems. Therefore, there is a need for a system that can efficiently process voice data and provide answers that respond to user emotions.
[1109] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1110] In this invention, the server includes means for collecting voice data, means for converting the voice data into text data, means for classifying the text data, means for analyzing the user's emotions based on the classified text data, means for generating answer suggestions based on the analyzed emotions and classification results, means for a human to review and correct the generated answer suggestions, and means for transmitting the reviewed and corrected answer suggestions. This enables efficient processing of voice data and the provision of personalized answers that respond to the user's emotions in customer service operations.
[1111] "Audio data" refers to information recorded in digital format using audio.
[1112] "Text data" refers to data obtained by converting audio data into written text.
[1113] "Classifying" refers to dividing data into categories based on specific criteria.
[1114] "Analyzing emotions" refers to detecting and identifying a user's emotional state from text data.
[1115] "Answer" refers to the content of the response to a user's inquiry.
[1116] "Review and correct" means that a human reviews the generated answer and makes corrections as necessary.
[1117] "Sending" refers to using communication methods to deliver the revised answer to the user.
[1118] This system implements a series of processes: collecting voice data, converting it to text data, classifying it, analyzing the user's emotions using an emotion engine, automatically generating suggested answers, having humans review and revise those answers, and finally sending them to the customer. The system utilizes a voice analysis API, natural language processing models, a database, and an emotion engine.
[1119] Hardware and software details
[1120] The following main hardware and software will be used to implement the system.
[1121] Device: A smartphone or tablet used by the shop crew to record conversations.
[1122] Server: A high-performance server used for processing and storing data.
[1123] Speech analysis APIs: Software for converting speech data into text data, such as the Google Cloud Speech-to-Text API.
[1124] Natural Language Processing Models (NLP models): Software such as OpenAI GPT-4 used to analyze and classify text data.
[1125] Emotion engine: Software such as IBM Watson that analyzes emotions from text.
[1126] Database: A database system for storing and managing various types of data (audio data, text data, analysis results, answer sheets, etc.).
[1127] Specific operation of the system
[1128] The terminal is used by shop staff to record conversations with customers. They press the record button to start the conversation and press it again to stop recording when finished. This recording file is then sent from the terminal to the server.
[1129] The server sends the received audio data to the Google Cloud Speech-to-Text API, where it converts the audio data into text data. The converted text data is then classified by a natural language processing model such as OpenAI GPT-4. For example, it may be categorized into "questions about product sizes," "inquiries about order status," and "return procedures."
[1130] Furthermore, the classified text data is sent to IBM Watson's sentiment engine, where the user's emotions are analyzed. Sentiment analysis is performed based on words and phrases within the text data to identify the user's emotions (e.g., joy, anger, sadness, etc.).
[1131] The server automatically generates suggested answers based on the classification results and sentiment analysis results. At this stage, it refers to the FAQ database and past answer history to create answers that take the user's emotions into consideration. For example, if a user asks "What are the product sizes?" and sentiment analysis detects anxiety, the prompt to the generating AI model will be "Answer when the user seems anxious: 'What are the product sizes?'" to generate a polite answer.
[1132] The generated answer is displayed on the terminal, where a shop crew member reviews and corrects it. They carefully examine whether the answer is correct and make corrections as needed. Once the review is complete, the answer is sent to the customer by pressing the "Send" button on the terminal. This entire process automates and streamlines inquiry handling, enabling appropriate responses that are tailored to the user's emotions.
[1133] Specific example
[1134] As an example, the flow of questions about a product and sentiment analysis is shown below.
[1135] 1. Terminal: The shop crew member starts recording, saying, "Please tell me the size of product A."
[1136] 2. Server: Receives the audio data and converts it into text data, "Please tell me the size of product A," using the Google Cloud Speech-to-Text API.
[1137] 3. Server: Use OpenAI GPT-4 to classify text data into "product size inquiries".
[1138] 4. Server: Sends text data to IBM Watson's sentiment engine to analyze whether the user is experiencing anxiety.
[1139] 5. Server: Based on the sentiment analysis results, it generates a polite response: "Thank you for waiting. The dimensions of this product are 10cm high, 5cm wide, and 1cm deep. Do you have any further questions?"
[1140] 6. Terminal: The shop crew checks the answer and corrects it to "The size of this product is 10cm in height."
[1141] 7. Terminal: Press the "Send" button to send the corrected answer to the customer.
[1142] In this way, the entire system ensures that a series of processes run smoothly, improving the efficiency and effectiveness of customer service operations.
[1143] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1144] Step 1:
[1145] Collection of audio data
[1146] Subject: terminal
[1147] Specific actions: The shop crew member presses the record button on the terminal to record the conversation with the customer. After the conversation ends, they press the button again to stop recording.
[1148] Input: Audio of a conversation between a shop crew member and a customer.
[1149] Output: Recorded audio file (saved in digital format)
[1150] Step 2:
[1151] Sending audio data
[1152] Subject: terminal
[1153] Specific operation: The recorded audio file is automatically uploaded to the server. A notification is displayed after the upload is complete.
[1154] Input: Recorded audio file
[1155] Output: Audio file sent to the server
[1156] Step 3:
[1157] Saving audio data
[1158] Subject: Server
[1159] Specific action: The received audio data is saved to the server's storage.
[1160] Input: Audio file sent from the device
[1161] Output: Audio data stored on the server
[1162] Step 4:
[1163] Converting audio data to text
[1164] Subject: Server
[1165] Specific operation: The saved audio data is sent to the Google Cloud Speech-to-Text API and converted into text data.
[1166] Input: Audio data on the server
[1167] Output: Text data returned from the Google Cloud Speech-to-Text API
[1168] Step 5:
[1169] Saving text data
[1170] Subject: Server
[1171] Specific action: Save the converted text data to the database.
[1172] Input: Text data
[1173] Output: Text data stored in the database
[1174] Step 6:
[1175] Classification of text data
[1176] Subject: Server
[1177] Specific operation: Text data is input into a natural language processing model such as OpenAI GPT-4, and the query content is classified.
[1178] Input: Text data read from a database
[1179] Output: Classification results returned from the natural language processing model
[1180] Step 7:
[1181] Saving classification results
[1182] Subject: Server
[1183] Specific action: Save the classification results to the database.
[1184] Input: Classification result
[1185] Output: Classification results stored in the database
[1186] Step 8:
[1187] Emotion analysis
[1188] Subject: Server
[1189] Specific operation: Classified text data is sent to IBM Watson's sentiment engine for sentiment analysis.
[1190] Input: Classified text data
[1191] Output: Sentiment analysis results returned from IBM Watson
[1192] Step 9:
[1193] Saving emotion analysis results
[1194] Subject: Server
[1195] Specific action: Save the emotion analysis results to the database.
[1196] Input: Sentiment analysis results
[1197] Output: Sentiment analysis results stored in the database
[1198] Step 10:
[1199] Generating the answer
[1200] Subject: Server
[1201] Specific operation: Automatically generates suggested answers based on classification results and sentiment analysis results. Refers to the FAQ database and past answer history to generate answers that take the user's emotions into consideration.
[1202] Input: Classification results and sentiment analysis results
[1203] Output: Solution returned by the generative AI model
[1204] Step 11:
[1205] Display of the answer
[1206] Subject: terminal
[1207] Specific action: Display the generated solution on the shop crew's terminal.
[1208] Input: Generated solution
[1209] Output: Answer displayed on the terminal
[1210] Step 12:
[1211] Review and correct the answer sheet.
[1212] Subject: User (Shop Crew)
[1213] Specific actions: The shop crew will review the displayed answer and make any necessary corrections. Once the corrections are complete, they will press the confirmation button.
[1214] Input: Displayed answer
[1215] Output: Revised solution
[1216] Step 13:
[1217] Sending a formal response
[1218] Subject: terminal
[1219] Specific actions: Send the corrected answer to the customer. Press the "Send" button on the device to deliver the answer to the customer via email or chat system.
[1220] Input: Revised answer
[1221] Output: Official response sent to the customer
[1222] (Application Example 2)
[1223] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1224] Currently, customer support operations, including food delivery services, require prompt and appropriate responses to customer inquiries. However, manual responses can lead to delays and decreased customer satisfaction. Furthermore, responses tend to be formulaic, making it difficult to provide personalized responses that address customer emotions and circumstances. Therefore, there is a need to achieve both automated and personalized responses to inquiries.
[1225] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1226] In this invention, the server includes means for collecting voice data, means for converting voice data into text data, means for classifying the text data, means for analyzing emotions based on the text data, means for generating answer suggestions based on the classified text data and emotion analysis results, means for a human to review and correct the generated answer suggestions, and means for transmitting the reviewed and corrected answer suggestions. This enables rapid and appropriate automated responses in customer support operations and makes it possible to provide personalized answers that take into account the customer's emotions.
[1227] "Audio data" refers to data that records audio in digital format.
[1228] "Text data" refers to string information converted from audio data.
[1229] A "Voice Analysis API" is an application programming interface for converting voice data into text data.
[1230] A "natural language processing model" is a machine learning or deep learning model used to analyze, classify, and interpret text data.
[1231] A "sentiment analysis engine" is an algorithm or system that identifies the user's emotions contained in text data.
[1232] A "database" is a system for efficiently managing, searching, and updating information.
[1233] A "solution" is a written response to a customer's inquiry.
[1234] "Customer support" refers to the work of responding to inquiries and requests from customers.
[1235] A "server" is a computer system used for storing and processing data.
[1236] A "food delivery service" is a service that delivers food ordered online by customers to a specified location.
[1237] "Personalization" refers to optimizing responses according to the individual needs and circumstances of each customer.
[1238] "Classification of inquiry content" is the process of assigning text data to a specific category.
[1239] "Emotion analysis results" refer to emotional information identified by the emotion analysis engine from text data.
[1240] This invention is a system that collects voice data, converts it into text data, classifies it, analyzes the user's emotions using an emotion engine, automatically generates a response, has a human review and revise the response, and finally sends it to the customer. This system is realized by utilizing a voice analysis API, a natural language processing model, a database, and an emotion engine.
[1241] System Configuration
[1242] 1. Collection of audio data
[1243] Device: The customer support representative's device (smartphone or computer) will record the conversation with the customer. They will press the record button to start the conversation and press it again to stop recording at the end.
[1244] 2. Converting audio data
[1245] Server: Processes the received audio data and converts it into text data by calling a speech analysis API. Specifically, it sends the collected audio data to a speech analysis API, such as Google's speech recognition API, and retrieves the corresponding text.
[1246] 3. Classification and sentiment analysis of text data
[1247] Server: The converted text data is input into a natural language processing (NLP) model to classify the inquiry content into specific categories. For example, it may be classified as "order status confirmation," "complaint handling," or "product-related questions."
[1248] Server: Sends classified text data to the sentiment engine to analyze the user's emotions. The sentiment engine identifies the user's emotions (e.g., joy, anger, sadness, etc.) from words and phrases in the text data.
[1249] 4. Generating solutions and making emotional revisions
[1250] Server: Automatically generates suggested answers based on classified inquiry content and sentiment analysis results. Specifically, it refers to the FAQ database and past answer history to create answers that take the user's emotions into consideration. For example, if a user expresses anger in response to an inquiry about a delayed delivery, it will respond politely with something like, "We apologize. We will take immediate action to resolve your dissatisfaction. Could you please provide more details?"
[1251] 5. Human review and correction
[1252] Terminal: The AI-generated solution is displayed on the customer support representative's terminal. The customer support representative reviews this solution and makes corrections as needed.
[1253] User (Customer Support Representative): Review the displayed solution to ensure its accuracy. If any information is missing or incorrect, manually correct it. Once corrections are complete, press the confirmation button to proceed to the next step.
[1254] 6. Submitting the formal response
[1255] Terminal: Customer support staff will send the reviewed and corrected answers to the customer. Specifically, they will press the "Send" button on the terminal to deliver the answer to the customer via email or chat system.
[1256] Specific example
[1257] As an example, consider a case where a customer inquires that their product delivery is delayed. A customer support representative records this inquiry and sends the audio data to a server. The server converts the audio into text data and categorizes it as "order status inquiry." The emotion analysis then identifies "anger." Based on this, a response is generated such as, "We apologize. We will promptly address your issue to resolve it. Could you please provide more details?"
[1258] Example of a prompt:
[1259] "Text: 'My delivery is delayed and I'm having trouble. What's going on?' Please categorize the content. Estimate the category and analyze the sentiment."
[1260] This system enables the automation and streamlining of inquiry handling, allowing for customer service that is sensitive to customer emotions.
[1261] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1262] Step 1:
[1263] Collection of audio data
[1264] Terminal: Customer support staff record conversations with customers. They press the record button to start collecting audio data, and press the button again to stop recording when the conversation ends. This saves the audio file to the terminal. This audio file becomes the input data for subsequent processing.
[1265] Step 2:
[1266] Sending and converting audio data
[1267] Server: Receives audio files sent from the terminal. The received audio files are sent to a speech analysis API for conversion into text data. Specifically, the speech analysis API analyzes the audio signal and outputs it as string data. The audio file is the input data, and the text data is the output data.
[1268] Step 3:
[1269] Classification of text data
[1270] Server: The converted text data is input into a natural language processing (NLP) model, which classifies the inquiry content into categories. This model is a pre-trained machine learning model that automatically assigns text data to categories such as "order status confirmation," "complaint handling," and "product-related questions." Text data is the input data, and category information is the output data.
[1271] Step 4:
[1272] Emotion analysis
[1273] Server: Sends classified text data to the sentiment engine to analyze the user's emotions. The sentiment engine identifies the user's emotions (e.g., joy, anger, sadness, etc.) from words and phrases in the text data. The text data is the input data for sentiment analysis, and the emotional information is the output data.
[1274] Step 5:
[1275] Automatic generation of answer sheets
[1276] Server: Automatically generates suggested answers based on classified category information and sentiment analysis results. Specifically, it refers to the FAQ database and past answer history to obtain the corresponding standard answer. Then, considering the sentiment analysis results, it generates a response that includes a response tailored to the user's emotions. Category information and sentiment information are input data, and suggested answers are output data.
[1277] Step 6:
[1278] Review and correction of the answer sheet
[1279] Terminal: The AI-generated answer is displayed on the customer support representative's terminal. The representative reviews this answer and makes corrections as needed. Once corrections are complete, they press the confirmation button to proceed to the next step. The initial answer is the input data, and the reviewed and corrected answer is the output data.
[1280] Step 7:
[1281] Sending a formal response
[1282] Terminal: Customer support staff send the reviewed and corrected answers to the customer. Specifically, they press the "Send" button on the terminal to deliver the answer to the customer via email or chat system. The corrected answer is the input data, and the answer sent to the customer is the output data.
[1283] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1284] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1285] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1286] [Fourth Embodiment]
[1287] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1288] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1289] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1290] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1291] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1292] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1293] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1294] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1295] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1296] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1297] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1298] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1299] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1300] This invention is a system that collects voice data, converts it into text data, classifies it, automatically generates answer keys, allows a human to review and correct the answer keys, and finally sends them to the customer. This system uses a voice analysis API and makes full use of natural language processing models and databases. The following describes the specific operation of this system, using the server, terminal, and user as subjects.
[1301] System program processing
[1302] Audio data collection and conversion
[1303] Terminal: The shop crew member's terminal records conversations with customers. For example, they press the record button on the terminal to start a conversation, and press the button again to stop recording when the conversation is finished. This recording file is sent from the terminal to the server.
[1304] Server: Processes the received audio data and calls the speech analysis API to convert it into text data. Specifically, it sends the audio data to the speech analysis API and retrieves the corresponding text.
[1305] Classification of text data
[1306] Server: The converted text data is passed through a natural language processing model to classify the query content into specific categories. For example, it may be classified as "questions about product sizes," "order status checks," or "return / exchange procedures." This classification is performed automatically based on the model's training results.
[1307] Generating the answer
[1308] Server: Automatically generates suggested answers based on classified inquiry content. Here, the server refers to FAQ data and past answer history in the database to create an appropriate answer. For example, in response to the inquiry "What are the dimensions of the product?", it generates a suggested answer such as "The dimensions of this product are 10cm high, 5cm wide, and 1cm deep."
[1309] Human verification and correction
[1310] Terminal: The AI-generated solution is displayed on the shop crew member's terminal. The shop crew member reviews this solution and makes corrections as needed.
[1311] User (Shop Crew): Carefully review the displayed answer to ensure its accuracy. If any information is missing or incorrect, manually correct it. Once corrections are complete, press the confirmation button to proceed to the next step.
[1312] Sending a formal response
[1313] Terminal: The shop crew will review and correct the answer and finally send it to the customer. For example, by pressing the "Send" button on the terminal, the answer will be delivered to the customer via email or chat system.
[1314] Specific example
[1315] Example: Questions about products
[1316] 1. Terminal: Starts recording with Shop Crew: "Please tell me the size of product A."
[1317] 2. Server: Receives audio data and uses an audio analysis API to convert it into text data such as "Please tell me the size of product A."
[1318] 3. Server: Classified as a "product size inquiry" by the natural language processing model.
[1319] 4. Server: Retrieves size information from the product database and generates a suggested answer such as, "The dimensions of this product are 10cm high, 5cm wide, and 1cm deep."
[1320] 5. Terminal: The shop crew checks the answer and corrects it to "The size of this product is 10cm in height."
[1321] 6. Terminal: Press the "Send" button to send the corrected answer to the customer.
[1322] As described above, this system is intended to automate and streamline inquiry handling, enabling accurate and prompt responses.
[1323] The following describes the processing flow.
[1324] Step 1:
[1325] Terminal: The shop crew member's terminal records conversations with customers. Specifically, they press the record button on the terminal to start the conversation. When the conversation ends, they press the button again to stop recording and send the recording file to the server.
[1326] Step 2:
[1327] Server: Receives audio data and converts it into text data using a speech analysis API. Specifically, it sends the received audio file to the speech analysis API and retrieves the corresponding text data.
[1328] Step 3:
[1329] Server: Inputs the transcribed data into a natural language processing (NLP) model and classifies the inquiry content into specific categories. For example, it automatically classifies inquiries into categories such as "questions about product sizes," "order status inquiries," and "return procedures."
[1330] Step 4:
[1331] Server: Automatically generates answers based on categorized inquiries. Specifically, it refers to the FAQ database and past answer history to select or create appropriate answers to generate answers corresponding to categories.
[1332] Step 5:
[1333] Server: Sends the generated answer sheet to the shop crew's terminal. Specifically, it communicates with the terminal to send the text data of the answer sheet from the server.
[1334] Step 6:
[1335] Terminal: The shop crew's terminal displays the received answer on the screen.
[1336] Step 7:
[1337] User (Shop Crew): The Shop Crew will review the displayed solution and make corrections as needed. Specifically, they will check the content of the solution and manually correct any omissions or errors.
[1338] Step 8:
[1339] Terminal: After completing the corrections to the answer, press the "Submit" button for final confirmation and send it to the customer as the official answer.
[1340] Step 9:
[1341] Terminal: The official response will be delivered to the customer via email or chat system.
[1342] (Example 1)
[1343] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1344] Traditional customer service systems often involved manual processes such as collecting and analyzing voice data, converting it to text data, and generating appropriate answers. This resulted in time-consuming and inefficient responses. Furthermore, the quality of answers depended on the individual operator's judgment, potentially leading to inconsistencies. This resulted in challenges such as decreased customer satisfaction and increased operating costs.
[1345] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1346] In this invention, the server includes means for converting audio data into text data using an audio analysis API, means for classifying the text data using a natural language processing model, and means for generating answer suggestions based on the classified text data by referring to a database. This automates the entire process from collecting audio data to generating answer suggestions and sending them to customers, enabling efficient and accurate inquiry handling.
[1347] A "terminal" is a device operated by the user to collect voice data and to review and correct answer sheets.
[1348] A "server" is a computer system that performs a series of operations: receiving audio data, converting it into text data using an audio analysis API, classifying it using a natural language processing model, and generating a solution.
[1349] "Audio data" refers to digital audio information, which is recordings of human speech collected using a device.
[1350] A "Voice Analysis API" is an application programming interface for analyzing input voice data and converting it into text data.
[1351] "Text data" refers to the character information corresponding to the audio data, converted by the speech analysis API.
[1352] A "natural language processing model" is a machine learning model that analyzes text data, understands its content, and classifies it into specific categories.
[1353] A "solution" is a document that responds to a customer inquiry, generated based on classified text data.
[1354] A "database" is a storage device that stores information such as FAQ data and past answer history necessary for generating answer templates.
[1355] This invention is a system that collects audio data, converts it into text data, classifies it, automatically generates answer keys, has a human review and correct the answer keys, and finally sends them to the customer. This system uses an audio analysis API and makes full use of natural language processing models and databases.
[1356] Audio data collection and conversion
[1357] Terminal: The terminal used by the user is a device for recording conversations with customers. Specifically, the user launches the recording application and presses the record button to start the conversation. When the conversation ends, the user presses the record button again to stop recording, and this recording file is temporarily saved to the terminal's local storage. This audio file is then sent to the server.
[1358] Server: Receives audio data sent from the terminal. The received audio file is temporarily stored, and this audio data is converted into text data using a speech analysis API such as the Google Cloud Speech-to-Text API. Specifically, the audio data is sent to the API endpoint, and the text data is retrieved from the returned response. This text data is stored in a database and used in the next step.
[1359] Classification of text data
[1360] Server: The acquired text data is input into a natural language processing model (e.g., BERT). The model analyzes the content of the text data and classifies the query into specific categories. For example, it might classify queries into "questions about product sizes," "order status checks," or "return / exchange procedures." This classification result is stored in a database and used in the next step of generating solutions.
[1361] Generating the answer
[1362] Server: Automatically generates answers based on classified category information. The server refers to FAQ data and past answer history in the database to create appropriate answers. For example, in response to the inquiry "What are the dimensions of the product?", it generates an answer such as "The dimensions of this product are 10cm high, 5cm wide, and 1cm deep." This generated answer is stored in the database and sent to the terminal in the next step.
[1363] Human verification and correction
[1364] Terminal: The answer submitted from the server is displayed on the user's terminal. The user can use a UI to review this answer and modify it as needed. Specifically, a text field is provided that allows the user to edit the displayed answer.
[1365] User: Check if the displayed answer is correct, and manually correct any missing or incorrect information. Once you have finished correcting, press the "Confirm" button to proceed to the next step.
[1366] Sending a formal response
[1367] Terminal: The user submits the completed and corrected answer to the customer by pressing the "Submit" button. This submission is carried out using email, chat API, etc. After submission is complete, the system displays a confirmation message to the customer.
[1368] Specific example
[1369] 1. Terminal: The user starts recording by saying, "Please tell me the size of product A."
[1370] 2. Server: Receives the audio data and converts it into text data, "Please tell me the size of product A," using the Google Cloud Speech-to-Text API.
[1371] 3. Server: The query is classified as "a query about product size" using the natural language processing model BERT.
[1372] 4. Server: Retrieves size information from the product database and generates an answer such as, "The dimensions of this product are 10cm high, 5cm wide, and 1cm deep."
[1373] 5. Terminal: The user reviews the answer and corrects it to "The size of this product is 10cm in height."
[1374] 6. Terminal: Press the "Send" button to send the corrected answer to the customer.
[1375] This will automate and streamline inquiry handling, enabling quick and accurate responses.
[1376] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1377] Step 1: Collect audio data
[1378] Terminal: This device records conversations between customers and shop crew. The user launches the recording application and presses the record button to begin the conversation. When the conversation ends, they press the record button again to stop recording. This audio file is temporarily stored in the terminal's local storage.
[1379] Input: Audio of a conversation between a customer and a shop crew member.
[1380] Output: Temporarily saved audio data file.
[1381] Step 2: Sending the audio data
[1382] Terminal: Sends the collected audio data to the server. Specifically, it uses the terminal's network module to upload the audio files to a specified URL on the server.
[1383] Input: Temporarily saved audio data file.
[1384] Output: Audio data sent to the server.
[1385] Step 3: Convert speech to text
[1386] Server: Converts received audio data into text data using a speech analysis API (e.g., Google Cloud Speech-to-Text API). Specifically, it sends the audio data to the API endpoint and extracts the text data from the returned response.
[1387] Input: Audio data sent to the server.
[1388] Output: Text data.
[1389] Step 4: Classification of Text Data
[1390] Server: The acquired text data is input into a natural language processing model (e.g., BERT) for classification. Specifically, the model is used to analyze the content of the text data and classify it into predefined categories. For example, it may be classified into "questions about product sizes," "order status confirmation," and "return / exchange procedures."
[1391] Input: Text data.
[1392] Output: Classification results (category labels).
[1393] Step 5: Generating the solution
[1394] Server: Based on the classified category, it searches the database for appropriate answers and generates suggested solutions. Specifically, it refers to FAQ data and past answer history to create suggested solutions that match the content of the inquiry.
[1395] Input: Classification result (category label).
[1396] Output: Generated solution.
[1397] Step 6: Submit your answer
[1398] Server: Sends the generated solution to the user's terminal. Specifically, it uploads the solution data to a specified URL on the user's terminal.
[1399] Input: Generated solution.
[1400] Output: The answer submitted to the user's terminal.
[1401] Step 7: Human review and correction
[1402] Terminal: The answer submitted from the server is displayed on the user's terminal. The user reviews this answer and makes corrections as needed.
[1403] User: Check the displayed solution, enter any necessary corrections using the text field, and save the revised version.
[1404] Input: Submitted answer.
[1405] Output: Revised solution.
[1406] Step 8: Submit your formal response
[1407] Terminal: This terminal sends the user's reviewed and corrected answer to the customer. Specifically, the user presses the "Send" button and sends the answer via email or chat API.
[1408] Input: Revised answer.
[1409] Output: The answer submitted to the customer.
[1410] Through the processing steps described above, this system automates the entire process from collecting voice data to generating answers and sending them to customers, enabling efficient and accurate handling of inquiries.
[1411] (Application Example 1)
[1412] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1413] In infotainment systems for autonomous vehicles, it is essential that passengers' voice-activated questions and requests are answered quickly and accurately. Conventional systems require human intervention, which can lead to delays and difficulties in providing accurate answers. Therefore, there is a need for systems in autonomous vehicles that can provide appropriate answers to voice commands in real time.
[1414] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1415] In this invention, the server includes means for collecting voice data, means for converting voice data into text data, and means for classifying text data. This makes it possible to generate and provide quick and accurate answers in real time to passenger voice instructions and requests in the infotainment system of an autonomous vehicle.
[1416] "Means for collecting voice data" refers to the function in the infotainment system of an autonomous vehicle that recognizes passenger voices and collects them as digital data.
[1417] "Methods for converting audio data into text data" refers to the process of analyzing collected audio data and converting it into corresponding text data using an audio analysis API or similar method.
[1418] "Means for classifying text data" refers to a function that automatically classifies converted text data into appropriate categories using a natural language processing model.
[1419] "Methods for generating answers" refers to the process of generating appropriate responses using natural language processing models, based on classified text data, referencing databases and FAQ data.
[1420] "Means for human review and correction" refers to a function where a human reviews the automatically generated answer and corrects its content as needed.
[1421] "Means for sending confirmed and revised answers" refers to the function of sending the final confirmed and revised answers to passengers via email or chat system.
[1422] A "voice recognition system installed in an autonomous vehicle" is a system installed in an autonomous vehicle that recognizes and processes passengers' voices in real time.
[1423] A "natural language processing model" is a machine learning model that understands and analyzes text data to generate appropriate responses or actions.
[1424] An "infotainment system" is a system installed in autonomous vehicles that provides information and entertainment to passengers using voice and video.
[1425] The system for implementing this invention operates as an infotainment system for an autonomous vehicle. When a passenger gives instructions or asks a question by voice, the system collects the voice, converts it into text data, and provides an appropriate answer in real time. The main hardware and software configuration is described in detail below.
[1426] Major hardware and software
[1427] Microphone: This is an input device for collecting passenger voices. It is installed inside the autonomous vehicle and is always on standby for passenger voice commands.
[1428] Speech analysis APIs are used to convert collected speech data into text data. Examples include the Google Speech Recognition API.
[1429] Natural Language Processing Model: Used to analyze the converted text data and generate appropriate answers. This model employs a generative AI model.
[1430] Database: Stores question-and-answer collections and FAQ data that are referenced when classifying text data and generating answer suggestions.
[1431] Display: This is an output device used to display the generated answer sheet to passengers. It is commonly installed on the dashboard of autonomous vehicles.
[1432] System processing flow
[1433] 1. Collection of voice data: The terminal (microphone inside the autonomous vehicle) collects the passenger's voice instructions. For example, questions such as "How long will it take to arrive at the next destination?" are collected.
[1434] 2. Audio Data Conversion: The device sends the collected audio data to the server, which uses an audio analysis API to convert it into text data. As a result, the text data "How long will it take to arrive at the next destination?" is generated.
[1435] 3. Text Data Classification: The server uses a natural language processing model to classify the converted text data into appropriate categories. For example, it might be classified as "Confirmation of Arrival Time."
[1436] 4. Solution Generation: Based on the classified category, the server retrieves appropriate information from the database and generates a solution using a generative AI model. For example, a solution such as "It will arrive in approximately 15 minutes" might be generated.
[1437] 5. Display of the solution: The generated solution is sent to the terminal (the display in the autonomous vehicle) and presented to the passengers visually.
[1438] Specific example
[1439] As a concrete example, consider the following conversation:
[1440] Passenger: "How long will it take to get to our next destination?"
[1441] System: "Arriving in approximately 15 minutes."
[1442] Based on this prompt, the system will process the information in the following order.
[1443] Example of a prompt:
[1444] Please convert the following audio data to text: "How long will it take to arrive at our next destination?"
[1445] Please categorize the following text data: "How long will it take to arrive at the next destination?"
[1446] Please generate an answer based on the following category and text data: "How long will it take to arrive at the next destination?"
[1447] These processes enable the infotainment system of an autonomous vehicle to respond quickly and accurately to voice commands from passengers.
[1448] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1449] Step 1:
[1450] Collection of audio data
[1451] Input: Passenger voice instructions
[1452] Operation: The terminal (microphone inside the autonomous vehicle) collects passenger voices in real time. For example, it recognizes voices such as, "How long will it take to arrive at the next destination?"
[1453] Output: Audio data file
[1454] Specific operation: The microphone records sound as digital data and passes the collected audio data to subsequent processing.
[1455] Step 2:
[1456] Audio data conversion
[1457] Input: Audio data file
[1458] Operation: The server uses a speech analysis API to convert speech data into text data. Google Speech Recognition is used as the API.
[1459] Output: Text data
[1460] Specific operation: Send audio data to the voice analysis API and retrieve the text "How long will it take to arrive at the next destination?". Pass this text data to the next processing step.
[1461] Step 3:
[1462] Classification of text data
[1463] Input: Text data
[1464] Operation: The server uses a natural language processing model to categorize text data into appropriate categories. For example, it might categorize it as "Confirming arrival time".
[1465] Output: Category Information
[1466] Specific operation: Text data is input into a natural language processing model, and the model outputs categories based on its learning results. This category information is then passed to the next step.
[1467] Step 4:
[1468] Generating the answer
[1469] Input: Category information and original text data
[1470] Operation: The server uses a database and a generative AI model to generate answer suggestions based on categories. For example, it might generate an answer such as "We will arrive in approximately 15 minutes."
[1471] Output: Solution
[1472] Specific operation: Relevant information (e.g., arrival time data) is retrieved from the database, and the optimal solution is generated using a generative AI model. This solution is then passed on to the next step.
[1473] Step 5:
[1474] Display of the answer
[1475] Input: Answer
[1476] Operation: The terminal (the display in the autonomous vehicle) displays the generated answer to the passengers. The display shows "We will arrive in approximately 15 minutes."
[1477] Output: Display of the answer sheet that passengers can see.
[1478] Specific operation: The answer will be displayed on a screen and presented in a format that is easy for passengers to read. This will allow passengers to see the answer to their question in real time.
[1479] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1480] This invention is a system that collects voice data, converts it into text data, classifies it, analyzes the user's emotions using an emotion engine, automatically generates a response, has a human review and revise the response, and finally sends it to the customer. This system makes full use of a voice analysis API, a natural language processing model, a database, and an emotion engine. The following describes the specific operation of this system, using the server, terminal, and user as subjects.
[1481] System program processing
[1482] Audio data collection and conversion
[1483] Terminal: The shop crew member's terminal records conversations with customers. Specifically, they press the record button on the terminal to start the conversation, and press the button again to stop recording when the conversation is finished. This recording file is sent from the terminal to the server.
[1484] Server: Processes the received audio data and calls the speech analysis API to convert the audio data into text data. Specifically, it sends the audio data to the speech analysis API and retrieves the corresponding text.
[1485] Text data classification and sentiment analysis
[1486] Server: The converted text data is input into a natural language processing (NLP) model to classify the inquiry content into specific categories. For example, it may be classified into "questions about product sizes," "inquiries about order status," "return procedures," etc.
[1487] Server: Sends classified text data to the sentiment engine to analyze the user's emotions. The sentiment engine identifies the user's emotions (e.g., joy, anger, sadness, etc.) from words and phrases in the text data.
[1488] Generating solutions and making emotional revisions.
[1489] Server: Automatically generates suggested answers based on classified inquiry content and sentiment analysis results. Specifically, it refers to the FAQ database and past answer history to create answers that take the user's feelings into consideration. For example, if the user looks dissatisfied with the inquiry "What are the dimensions of the product?", it will respond politely with something like, "Thank you for waiting. The dimensions of this product are 10cm high, 5cm wide, and 1cm deep."
[1490] Human verification and correction
[1491] Terminal: The AI-generated solution is displayed on the shop crew member's terminal. The shop crew member reviews this solution and makes corrections as needed.
[1492] User (Shop Crew): Carefully review the displayed answer to ensure its accuracy. If any information is missing or incorrect, manually correct it. Once corrections are complete, press the confirmation button to proceed to the next step.
[1493] Sending a formal response
[1494] Terminal: The shop crew will send the customer the corrected and revised answer. Specifically, they will press the "Send" button on the terminal and deliver the answer to the customer via email or chat system.
[1495] Specific example
[1496] Example: Questions about a product and sentiment analysis
[1497] 1. Terminal: Starts recording with Shop Crew: "Please tell me the size of product A."
[1498] 2. Server: Receives audio data and uses an audio analysis API to convert it into text data such as "Please tell me the size of product A."
[1499] 3. Server: Classified as a "product size inquiry" by the natural language processing model.
[1500] 4. Server: Sends text data to the emotion engine to analyze whether the user is experiencing anxiety.
[1501] 5. Server: Based on the sentiment analysis results, it generates a polite response: "Thank you for waiting. The dimensions of this product are 10cm high, 5cm wide, and 1cm deep. Do you have any further questions?"
[1502] 6. Terminal: The shop crew checks the answer and corrects it to "The size of this product is 10cm in height."
[1503] 7. Terminal: Press the "Send" button to send the corrected answer to the customer.
[1504] As described above, this system is intended to automate and streamline inquiry handling, enabling accurate and rapid responses. Furthermore, by combining it with user sentiment analysis using an emotion engine, it is possible to provide more personalized services.
[1505] The following describes the processing flow.
[1506] Step 1:
[1507] Terminal: The shop crew member's terminal records conversations with customers. Specifically, they press the record button on the terminal to start the conversation, and press the button again to stop recording when the conversation is finished. This recording file is sent from the terminal to the server.
[1508] Step 2:
[1509] Server: Processes the received audio data and calls the audio analysis API to convert the audio data into text data. Specifically, it sends the audio file to the audio analysis API and retrieves the corresponding text data.
[1510] Step 3:
[1511] Server: The converted text data is input into a natural language processing (NLP) model, which classifies the query into a specific category. For example, if the text data is "What are the product sizes?", it will be classified as "Questions about product sizes".
[1512] Step 4:
[1513] Server: After passing text data through a natural language processing model, the emotion engine analyzes the user's emotions. Specifically, it identifies emotions such as whether the user is happy or dissatisfied based on the words and phrases in the text data.
[1514] Step 5:
[1515] Server: Automatically generates suggested answers based on classified inquiry content and sentiment analysis results. Specifically, it refers to the FAQ database and past answer history to create answers that take the user's feelings into consideration. For example, if the user seems dissatisfied with the question "What are the dimensions of the product?", it will create a polite suggested answer such as, "We apologize for the wait. The dimensions of this product are 10cm high, 5cm wide, and 1cm deep."
[1516] Step 6:
[1517] Server: Sends the generated answer sheet to the shop crew's terminal. Specifically, it communicates with the terminal to send the text data of the answer sheet from the server.
[1518] Step 7:
[1519] Terminal: The shop crew's terminal displays the received answer on the screen.
[1520] Step 8:
[1521] User (Shop Crew): The Shop Crew will review the displayed solution and make corrections as needed. Specifically, they will check the content of the solution and manually correct any omissions or errors.
[1522] Step 9:
[1523] Terminal: After completing the corrections to the answer, press the "Submit" button for final confirmation and send it to the customer as the official answer.
[1524] Step 10:
[1525] Terminal: The official response will be delivered to the customer via email or chat system.
[1526] (Example 2)
[1527] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1528] In modern customer service, efficiently collecting voice data and providing quick and accurate answers is crucial. However, converting voice data into text, classifying the content, and generating appropriate answers is a time-consuming process. Furthermore, personalized responses that take into account user emotions are required, but this is difficult to achieve with current systems. Therefore, there is a need for a system that can efficiently process voice data and provide answers that respond to user emotions.
[1529] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1530] In this invention, the server includes means for collecting voice data, means for converting the voice data into text data, means for classifying the text data, means for analyzing the user's emotions based on the classified text data, means for generating answer suggestions based on the analyzed emotions and classification results, means for a human to review and correct the generated answer suggestions, and means for transmitting the reviewed and corrected answer suggestions. This enables efficient processing of voice data and the provision of personalized answers that respond to the user's emotions in customer service operations.
[1531] "Audio data" refers to information recorded in digital format using audio.
[1532] "Text data" refers to data obtained by converting audio data into written text.
[1533] "Classifying" refers to dividing data into categories based on specific criteria.
[1534] "Analyzing emotions" refers to detecting and identifying a user's emotional state from text data.
[1535] "Answer" refers to the content of the response to a user's inquiry.
[1536] "Review and correct" means that a human reviews the generated answer and makes corrections as necessary.
[1537] "Sending" refers to using communication methods to deliver the revised answer to the user.
[1538] This system implements a series of processes: collecting voice data, converting it to text data, classifying it, analyzing the user's emotions using an emotion engine, automatically generating suggested answers, having humans review and revise those answers, and finally sending them to the customer. The system utilizes a voice analysis API, natural language processing models, a database, and an emotion engine.
[1539] Hardware and software details
[1540] The following main hardware and software will be used to implement the system.
[1541] Device: A smartphone or tablet used by the shop crew to record conversations.
[1542] Server: A high-performance server used for processing and storing data.
[1543] Speech analysis APIs: Software for converting speech data into text data, such as the Google Cloud Speech-to-Text API.
[1544] Natural Language Processing Models (NLP models): Software such as OpenAI GPT-4 used to analyze and classify text data.
[1545] Emotion engine: Software such as IBM Watson that analyzes emotions from text.
[1546] Database: A database system for storing and managing various types of data (audio data, text data, analysis results, answer sheets, etc.).
[1547] Specific operation of the system
[1548] The terminal is used by shop staff to record conversations with customers. They press the record button to start the conversation and press it again to stop recording when finished. This recording file is then sent from the terminal to the server.
[1549] The server sends the received audio data to the Google Cloud Speech-to-Text API, where it converts the audio data into text data. The converted text data is then classified by a natural language processing model such as OpenAI GPT-4. For example, it may be categorized into "questions about product sizes," "inquiries about order status," and "return procedures."
[1550] Furthermore, the classified text data is sent to IBM Watson's sentiment engine, where the user's emotions are analyzed. Sentiment analysis is performed based on words and phrases within the text data to identify the user's emotions (e.g., joy, anger, sadness, etc.).
[1551] The server automatically generates suggested answers based on the classification results and sentiment analysis results. At this stage, it refers to the FAQ database and past answer history to create answers that take the user's emotions into consideration. For example, if a user asks "What are the product sizes?" and sentiment analysis detects anxiety, the prompt to the generating AI model will be "Answer when the user seems anxious: 'What are the product sizes?'" to generate a polite answer.
[1552] The generated answer is displayed on the terminal, where a shop crew member reviews and corrects it. They carefully examine whether the answer is correct and make corrections as needed. Once the review is complete, the answer is sent to the customer by pressing the "Send" button on the terminal. This entire process automates and streamlines inquiry handling, enabling appropriate responses that are tailored to the user's emotions.
[1553] Specific example
[1554] As an example, the flow of questions about a product and sentiment analysis is shown below.
[1555] 1. Terminal: The shop crew member starts recording, saying, "Please tell me the size of product A."
[1556] 2. Server: Receives the audio data and converts it into text data, "Please tell me the size of product A," using the Google Cloud Speech-to-Text API.
[1557] 3. Server: Use OpenAI GPT-4 to classify text data into "product size inquiries".
[1558] 4. Server: Sends text data to IBM Watson's sentiment engine to analyze whether the user is experiencing anxiety.
[1559] 5. Server: Based on the sentiment analysis results, it generates a polite response: "Thank you for waiting. The dimensions of this product are 10cm high, 5cm wide, and 1cm deep. Do you have any further questions?"
[1560] 6. Terminal: The shop crew checks the answer and corrects it to "The size of this product is 10cm in height."
[1561] 7. Terminal: Press the "Send" button to send the corrected answer to the customer.
[1562] In this way, the entire system ensures that a series of processes run smoothly, improving the efficiency and effectiveness of customer service operations.
[1563] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1564] Step 1:
[1565] Collection of audio data
[1566] Subject: terminal
[1567] Specific actions: The shop crew member presses the record button on the terminal to record the conversation with the customer. After the conversation ends, they press the button again to stop recording.
[1568] Input: Audio of a conversation between a shop crew member and a customer.
[1569] Output: Recorded audio file (saved in digital format)
[1570] Step 2:
[1571] Sending audio data
[1572] Subject: terminal
[1573] Specific operation: The recorded audio file is automatically uploaded to the server. A notification is displayed after the upload is complete.
[1574] Input: Recorded audio file
[1575] Output: Audio file sent to the server
[1576] Step 3:
[1577] Saving audio data
[1578] Subject: Server
[1579] Specific action: The received audio data is saved to the server's storage.
[1580] Input: Audio file sent from the device
[1581] Output: Audio data stored on the server
[1582] Step 4:
[1583] Converting audio data to text
[1584] Subject: Server
[1585] Specific operation: The saved audio data is sent to the Google Cloud Speech-to-Text API and converted into text data.
[1586] Input: Audio data on the server
[1587] Output: Text data returned from the Google Cloud Speech-to-Text API
[1588] Step 5:
[1589] Saving text data
[1590] Subject: Server
[1591] Specific action: Save the converted text data to the database.
[1592] Input: Text data
[1593] Output: Text data stored in the database
[1594] Step 6:
[1595] Classification of text data
[1596] Subject: Server
[1597] Specific operation: Text data is input into a natural language processing model such as OpenAI GPT-4, and the query content is classified.
[1598] Input: Text data read from a database
[1599] Output: Classification results returned from the natural language processing model
[1600] Step 7:
[1601] Saving classification results
[1602] Subject: Server
[1603] Specific action: Save the classification results to the database.
[1604] Input: Classification result
[1605] Output: Classification results stored in the database
[1606] Step 8:
[1607] Emotion analysis
[1608] Subject: Server
[1609] Specific operation: Classified text data is sent to IBM Watson's sentiment engine for sentiment analysis.
[1610] Input: Classified text data
[1611] Output: Sentiment analysis results returned from IBM Watson
[1612] Step 9:
[1613] Saving emotion analysis results
[1614] Subject: Server
[1615] Specific action: Save the emotion analysis results to the database.
[1616] Input: Sentiment analysis results
[1617] Output: Sentiment analysis results stored in the database
[1618] Step 10:
[1619] Generating the answer
[1620] Subject: Server
[1621] Specific operation: Automatically generates suggested answers based on classification results and sentiment analysis results. Refers to the FAQ database and past answer history to generate answers that take the user's emotions into consideration.
[1622] Input: Classification results and sentiment analysis results
[1623] Output: Solution returned by the generative AI model
[1624] Step 11:
[1625] Display of the answer
[1626] Subject: terminal
[1627] Specific action: Display the generated solution on the shop crew's terminal.
[1628] Input: Generated solution
[1629] Output: Answer displayed on the terminal
[1630] Step 12:
[1631] Review and correct the answer sheet.
[1632] Subject: User (Shop Crew)
[1633] Specific actions: The shop crew will review the displayed answer and make any necessary corrections. Once the corrections are complete, they will press the confirmation button.
[1634] Input: Displayed answer
[1635] Output: Revised solution
[1636] Step 13:
[1637] Sending a formal response
[1638] Subject: terminal
[1639] Specific actions: Send the corrected answer to the customer. Press the "Send" button on the device to deliver the answer to the customer via email or chat system.
[1640] Input: Revised answer
[1641] Output: Official response sent to the customer
[1642] (Application Example 2)
[1643] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1644] Currently, customer support operations, including food delivery services, require prompt and appropriate responses to customer inquiries. However, manual responses can lead to delays and decreased customer satisfaction. Furthermore, responses tend to be formulaic, making it difficult to provide personalized responses that address customer emotions and circumstances. Therefore, there is a need to achieve both automated and personalized responses to inquiries.
[1645] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1646] In this invention, the server includes means for collecting voice data, means for converting voice data into text data, means for classifying the text data, means for analyzing emotions based on the text data, means for generating answer suggestions based on the classified text data and emotion analysis results, means for a human to review and correct the generated answer suggestions, and means for transmitting the reviewed and corrected answer suggestions. This enables rapid and appropriate automated responses in customer support operations and makes it possible to provide personalized answers that take into account the customer's emotions.
[1647] "Audio data" refers to data that records audio in digital format.
[1648] "Text data" refers to string information converted from audio data.
[1649] A "Voice Analysis API" is an application programming interface for converting voice data into text data.
[1650] A "natural language processing model" is a machine learning or deep learning model used to analyze, classify, and interpret text data.
[1651] A "sentiment analysis engine" is an algorithm or system that identifies the user's emotions contained in text data.
[1652] A "database" is a system for efficiently managing, searching, and updating information.
[1653] A "solution" is a written response to a customer's inquiry.
[1654] "Customer support" refers to the work of responding to inquiries and requests from customers.
[1655] A "server" is a computer system used for storing and processing data.
[1656] A "food delivery service" is a service that delivers food ordered online by customers to a specified location.
[1657] "Personalization" refers to optimizing responses according to the individual needs and circumstances of each customer.
[1658] "Classification of inquiry content" is the process of assigning text data to a specific category.
[1659] "Emotion analysis results" refer to emotional information identified by the emotion analysis engine from text data.
[1660] This invention is a system that collects voice data, converts it into text data, classifies it, analyzes the user's emotions using an emotion engine, automatically generates a response, has a human review and revise the response, and finally sends it to the customer. This system is realized by utilizing a voice analysis API, a natural language processing model, a database, and an emotion engine.
[1661] System Configuration
[1662] 1. Collection of audio data
[1663] Device: The customer support representative's device (smartphone or computer) will record the conversation with the customer. They will press the record button to start the conversation and press it again to stop recording at the end.
[1664] 2. Converting audio data
[1665] Server: Processes the received audio data and converts it into text data by calling a speech analysis API. Specifically, it sends the collected audio data to a speech analysis API, such as Google's speech recognition API, and retrieves the corresponding text.
[1666] 3. Classification and sentiment analysis of text data
[1667] Server: The converted text data is input into a natural language processing (NLP) model to classify the inquiry content into specific categories. For example, it may be classified as "order status confirmation," "complaint handling," or "product-related questions."
[1668] Server: Sends classified text data to the sentiment engine to analyze the user's emotions. The sentiment engine identifies the user's emotions (e.g., joy, anger, sadness, etc.) from words and phrases in the text data.
[1669] 4. Generating solutions and making emotional revisions
[1670] Server: Automatically generates suggested answers based on classified inquiry content and sentiment analysis results. Specifically, it refers to the FAQ database and past answer history to create answers that take the user's emotions into consideration. For example, if a user expresses anger in response to an inquiry about a delayed delivery, it will respond politely with something like, "We apologize. We will take immediate action to resolve your dissatisfaction. Could you please provide more details?"
[1671] 5. Human review and correction
[1672] Terminal: The AI-generated solution is displayed on the customer support representative's terminal. The customer support representative reviews this solution and makes corrections as needed.
[1673] User (Customer Support Representative): Review the displayed solution to ensure its accuracy. If any information is missing or incorrect, manually correct it. Once corrections are complete, press the confirmation button to proceed to the next step.
[1674] 6. Submitting the formal response
[1675] Terminal: Customer support staff will send the reviewed and corrected answers to the customer. Specifically, they will press the "Send" button on the terminal to deliver the answer to the customer via email or chat system.
[1676] Specific example
[1677] As an example, consider a case where a customer inquires that their product delivery is delayed. A customer support representative records this inquiry and sends the audio data to a server. The server converts the audio into text data and categorizes it as "order status inquiry." The emotion analysis then identifies "anger." Based on this, a response is generated such as, "We apologize. We will promptly address your issue to resolve it. Could you please provide more details?"
[1678] Example of a prompt:
[1679] "Text: 'My delivery is delayed and I'm having trouble. What's going on?' Please categorize the content. Estimate the category and analyze the sentiment."
[1680] This system enables the automation and streamlining of inquiry handling, allowing for customer service that is sensitive to customer emotions.
[1681] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1682] Step 1:
[1683] Collection of audio data
[1684] Terminal: Customer support staff record conversations with customers. They press the record button to start collecting audio data, and press the button again to stop recording when the conversation ends. This saves the audio file to the terminal. This audio file becomes the input data for subsequent processing.
[1685] Step 2:
[1686] Sending and converting audio data
[1687] Server: Receives audio files sent from the terminal. The received audio files are sent to a speech analysis API for conversion into text data. Specifically, the speech analysis API analyzes the audio signal and outputs it as string data. The audio file is the input data, and the text data is the output data.
[1688] Step 3:
[1689] Classification of text data
[1690] Server: The converted text data is input into a natural language processing (NLP) model, which classifies the inquiry content into categories. This model is a pre-trained machine learning model that automatically assigns text data to categories such as "order status confirmation," "complaint handling," and "product-related questions." Text data is the input data, and category information is the output data.
[1691] Step 4:
[1692] Emotion analysis
[1693] Server: Sends classified text data to the sentiment engine to analyze the user's emotions. The sentiment engine identifies the user's emotions (e.g., joy, anger, sadness, etc.) from words and phrases in the text data. The text data is the input data for sentiment analysis, and the emotional information is the output data.
[1694] Step 5:
[1695] Automatic generation of answer sheets
[1696] Server: Automatically generates suggested answers based on classified category information and sentiment analysis results. Specifically, it refers to the FAQ database and past answer history to obtain the corresponding standard answer. Then, considering the sentiment analysis results, it generates a response that includes a response tailored to the user's emotions. Category information and sentiment information are input data, and suggested answers are output data.
[1697] Step 6:
[1698] Review and correction of the answer sheet
[1699] Terminal: The AI-generated answer is displayed on the customer support representative's terminal. The representative reviews this answer and makes corrections as needed. Once corrections are complete, they press the confirmation button to proceed to the next step. The initial answer is the input data, and the reviewed and corrected answer is the output data.
[1700] Step 7:
[1701] Sending a formal response
[1702] Terminal: Customer support staff send the reviewed and corrected answers to the customer. Specifically, they press the "Send" button on the terminal to deliver the answer to the customer via email or chat system. The corrected answer is the input data, and the answer sent to the customer is the output data.
[1703] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1704] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1705] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1706] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1707] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1708] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1709] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1710] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1711] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1712] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1713] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1714] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1715] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1716] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1717] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1718] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1719] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1720] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1721] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1722] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[1723] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[1724] The following is further disclosed regarding the embodiments described above.
[1725] (Claim 1)
[1726] Means of collecting audio data,
[1727] Means for converting the aforementioned audio data into text data,
[1728] Means for classifying the aforementioned text data,
[1729] means for generating answer sheets based on the classified text data,
[1730] The aforementioned generated answer sheet includes means for a human to review and correct it,
[1731] Means for transmitting the confirmed and corrected answer sheet,
[1732] A system that includes this.
[1733] (Claim 2)
[1734] The system according to claim 1, characterized in that it uses a speech analysis API as a means for converting the aforementioned speech data into text data.
[1735] (Claim 3)
[1736] The system according to claim 1, characterized in that it uses a natural language processing model and a database as means for generating the aforementioned answer.
[1737] "Example 1"
[1738] (Claim 1)
[1739] The means by which the device collects voice data,
[1740] The server provides means for converting the aforementioned audio data into text data using an audio analysis API,
[1741] The server provides means for classifying the text data using a natural language processing model,
[1742] The server provides means for generating answer keys by referring to a database based on the classified text data,
[1743] The terminal provides means for a human to review and correct the generated answer sheet,
[1744] The terminal is a means for transmitting the confirmed and corrected answer sheet,
[1745] A system that includes this.
[1746] (Claim 2)
[1747] The system according to claim 1, characterized in that it uses a speech analysis API as a means for converting the aforementioned speech data into text data.
[1748] (Claim 3)
[1749] The system according to claim 1, characterized in that a natural language processing model is used as a means for classifying the text data.
[1750] "Application Example 1"
[1751] (Claim 1)
[1752] Means of collecting audio data,
[1753] Means for converting the aforementioned audio data into text data,
[1754] Means for classifying the aforementioned text data,
[1755] means for generating answer sheets based on the classified text data,
[1756] The aforementioned generated answer sheet includes means for a human to review and correct it,
[1757] Means for transmitting the confirmed and corrected answer sheet,
[1758] Voice recognition systems installed in autonomous vehicles,
[1759] The speech recognition system includes a natural language processing model that generates answer suggestions based on the collected speech data,
[1760] A means for displaying the generated solution on the infotainment system of an autonomous vehicle,
[1761] A system that includes this.
[1762] (Claim 2)
[1763] The system according to claim 1, characterized in that it uses a speech analysis API as a means of converting speech data into text data.
[1764] (Claim 3)
[1765] The system according to claim 1, characterized in that it uses a natural language processing model and a database as means for generating answer suggestions.
[1766] "Example 2 of combining an emotion engine"
[1767] (Claim 1)
[1768] Means of collecting audio data,
[1769] Means for converting the aforementioned audio data into text data,
[1770] Means for classifying the aforementioned text data,
[1771] A means for analyzing the user's emotions based on the classified text data,
[1772] A means for generating an answer based on the analyzed emotions and classification results,
[1773] The aforementioned generated answer sheet includes means for a human to review and correct it,
[1774] Means for transmitting the confirmed and corrected answer sheet,
[1775] A system that includes this.
[1776] (Claim 2)
[1777] The system according to claim 1, characterized in that a speech analysis interface is used as a means for converting the aforementioned speech data into text data.
[1778] (Claim 3)
[1779] The system according to claim 1, characterized in that it uses a natural language processing model and a knowledge base as means for generating the aforementioned answer.
[1780] "Application example 2 when combining with an emotional engine"
[1781] (Claim 1)
[1782] Means of collecting audio data,
[1783] Means for converting the aforementioned audio data into text data,
[1784] Means for classifying the aforementioned text data,
[1785] A means for analyzing emotions based on the aforementioned text data,
[1786] A means for generating answer suggestions based on the classified text data and sentiment analysis results,
[1787] The aforementioned generated answer sheet includes means for a human to review and correct it,
[1788] Means for transmitting the confirmed and corrected answer sheet,
[1789] A system that includes this.
[1790] (Claim 2)
[1791] The system according to claim 1, characterized in that it uses a speech analysis API as a means for converting the aforementioned speech data into text data.
[1792] (Claim 3)
[1793] The system according to claim 1, characterized in that it uses a natural language processing model, a database, and an emotion analysis engine as means for generating the aforementioned answer. [Explanation of symbols]
[1794] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. Means of collecting audio data, Means for converting the aforementioned audio data into text data, Means for classifying the aforementioned text data, means for generating answer sheets based on the classified text data, The aforementioned generated answer sheet includes means for a human to review and correct it, Means for transmitting the confirmed and corrected answer sheet, A system that includes this.
2. The system according to claim 1, characterized in that it uses a speech analysis API as a means for converting the aforementioned speech data into text data.
3. The system according to claim 1, characterized in that it uses a natural language processing model and a database as means for generating the aforementioned answer.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A