system

The system addresses the challenge of inconsistent customer service by processing voice inputs to identify user intent and emotions, providing personalized responses, thus enhancing satisfaction and quality.

JP2026060619APending Publication Date: 2026-04-08SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-27
Publication Date
2026-04-08

AI Technical Summary

Technical Problem

Conventional customer service methods struggle to accurately understand customer intentions and nuances, leading to inconsistent service quality and a decline in customer satisfaction, with a lack of individualized responses.

Method used

A system that processes user voice input through preprocessing, speech recognition, natural language processing, and intent identification, utilizing AI models to generate personalized responses, accounting for nuances and emotions.

Benefits of technology

Enables precise and individualized customer service by accurately understanding user intent and emotions, improving satisfaction and service quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026060619000001_ABST
    Figure 2026060619000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means for inputting the user's voice, A means for sending input audio data to a server, A means for preprocessing audio data, A means for converting pre-processed audio data into text data, A means of identifying customer intent from text data, A means of generating the optimal response based on identified intent, A means of notifying the user of the generated response, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In modern customer service, it is required to accurately and quickly grasp various customer demands and provide appropriate responses. However, with conventional customer service methods, it is difficult to accurately understand the intentions and nuances of customers, resulting in a decline in customer satisfaction and variations in service quality. In addition, it is difficult to provide individualized responses, and improving customer loyalty has also been an issue. The purpose of this invention is to solve these problems and provide a system for improving customer satisfaction and service quality.

Means for Solving the Problems

[0005] To solve the above-mentioned problems, this invention provides the following means: a system including means for inputting the user's voice, means for transmitting the input voice data to a server, means for preprocessing the voice data, means for converting the preprocessed voice data into text data, means for identifying the customer's intent from the text data, means for generating an optimal response based on the identified intent, and means for notifying the user of the generated response. Furthermore, by providing means for analyzing nuances and intonation when identifying the customer's intent, and means for notifying the user of the generated response as text data and voice data, it becomes possible to respond to customer requests in a more precise and individualized manner. In this way, it is possible to understand the customer's intent with high accuracy and provide high-quality service.

[0006] A "user" is an individual or group that uses the system.

[0007] "Audio data" refers to data that records the voice spoken by a user in digital format.

[0008] A "server" is a computer system that processes, stores, and analyzes audio data over a network.

[0009] "Preprocessing" refers to the process of performing actions on audio data to improve its quality, such as noise reduction and volume normalization.

[0010] "Speech recognition" is the process of converting input speech data into text data.

[0011] "Text data" refers to character information generated by speech recognition.

[0012] Natural Language Processing (NLP) is a technology that analyzes text data to understand the meaning and intent of human language.

[0013] "Nuance" refers to the subtle intonation and emotional changes contained in audio data.

[0014] "Intention" refers to the content or purpose that the user wishes to express through voice.

[0015] "Answer generation" is a process of providing an appropriate answer based on the specified intention.

[0016] "Notification means" refers to the method of conveying the generated answer to the user, including text display and voice playback.

Brief Explanation of Drawings

[0017] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12]It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.

Mode for Carrying Out the Invention

[0018] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be explained.

[0020] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.

[0021] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0022] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0023] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0025] [First Embodiment]

[0026] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0027] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0030] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0033] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0037] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0038] System Overview

[0039] This invention relates to a customer service system that utilizes speech recognition and natural language processing technologies. It takes user voice input, analyzes the voice data in real time to identify the customer's intent, and provides the optimal response. The system consists of a user, a terminal, and a server.

[0040] System Configuration

[0041] User

[0042] Users input questions and requests via voice through the system. When a user speaks into the microphone, the voice data is transmitted to the device.

[0043] terminal

[0044] The terminal records the user's voice in digital format and sends it to the server. Some pre-processing is performed as needed before the received audio data is sent to the server.

[0045] server

[0046] The server is the core of this system and performs the following main functions:

[0047] Audio data preprocessing

[0048] Speech recognition

[0049] Natural Language Processing

[0050] Intent Extraction

[0051] Answer generation

[0052] Submit your response

[0053] Processing flow

[0054] Voice input and transmission

[0055] 1. User: Speak your question or request into the microphone.

[0056] 2. Terminal: Records the user's voice as digital data and sends that data to the server.

[0057] Audio data preprocessing

[0058] 3. Server: The server performs preprocessing on the received audio data, such as noise reduction and volume normalization, and converts it into a format suitable for analysis.

[0059] Speech recognition

[0060] 4. Server: Sends the pre-processed audio data to the speech recognition engine and converts the audio data into text data.

[0061] Natural language processing and intent extraction

[0062] 5. Server: The text data generated by speech recognition is fed into a natural language processing engine to analyze the user's nuances and emotions. Next, the user's intent is identified from the text data.

[0063] Response generation and submission

[0064] 6. Server: Utilizes an AI model to generate the optimal response based on the identified intent. The generated response is then formatted as text and audio data.

[0065] 7. Server: Sends the formatted response to the terminal. The terminal displays or plays the received data in audio format.

[0066] Specific example

[0067] Scenario 1: Inquiry about pricing plans

[0068] 1. User: "What is the recommended pricing plan?"

[0069] 2. Terminal: Records audio and sends it to the server.

[0070] 3. Server: Performs noise reduction and volume normalization.

[0071] 4. Server: Converts audio data to text. "What is the recommended pricing plan?"

[0072] 5. Server: Analyzes text data to identify user intent. "Questions about pricing plans"

[0073] 6. Server: Based on information about pricing plans, the AI ​​generates the optimal response. "Currently, there are three available pricing plans: A, B, and C. Their features are..."

[0074] 7. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[0075] Scenario 2: Buying a new smartphone

[0076] 1. User: "I want to buy a new smartphone."

[0077] 2. Terminal: Records audio and sends it to the server.

[0078] 3. Server: Performs noise reduction and volume normalization.

[0079] 4. Server: Converts audio data to text. "I want to buy a new smartphone."

[0080] 5. Server: Analyzes text data to identify user intent. "Question regarding the purchase of a new smartphone"

[0081] 6. Server: Based on information about the smartphone, the AI ​​generates the optimal answer. "The current new models are X, Y, and Z, and their respective features are..."

[0082] 7. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[0083] In this way, the system can accurately understand the user's intent from their voice input and provide the optimal response. This can lead to improved customer satisfaction and consistent service quality.

[0084] The following describes the processing flow.

[0085] Step 1:

[0086] User: Speak your questions or requests into the microphone.

[0087] Step 2:

[0088] Terminal: Records the user's voice as digital data.

[0089] Step 3:

[0090] Terminal: Sends the recorded audio data to the server via an HTTP request.

[0091] Step 4:

[0092] Server: Analyzes the received audio data and applies a noise reduction filter.

[0093] Step 5:

[0094] Server: Normalizes the volume of the data to a certain level.

[0095] Step 6:

[0096] Server: Passes the pre-processed audio data to the speech recognition engine.

[0097] Step 7:

[0098] Server: The speech recognition engine converts the speech data into text data.

[0099] Step 8:

[0100] Server: Transfers the generated text data to the natural language processing (NLP) engine.

[0101] Step 9:

[0102] Server: The NLP engine performs sentiment analysis on text data, analyzing the customer's nuances and intonation.

[0103] Step 10:

[0104] Server: Uses text intent analysis algorithms to identify customer questions and requests from text data.

[0105] Step 11:

[0106] Server: Based on customer intent, the response generation module uses AI to generate the optimal response.

[0107] Step 12:

[0108] Server: Formats the generated responses into text and audio data.

[0109] Step 13:

[0110] Server: Sends the formatted response to the terminal as an HTTP response.

[0111] Step 14:

[0112] Terminal: Displays text data received from the server on the screen.

[0113] Step 15:

[0114] Terminal: Plays received audio data through the speaker.

[0115] Step 16:

[0116] User: Review the provided answers and re-enter any questions or requests as needed.

[0117] (Example 1)

[0118] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0119] Conventional speech recognition systems have limited processes for generating appropriate responses from voice input, making it difficult to accurately identify user intent. Furthermore, the generated responses are often uniform and may not adequately address customer questions or requests. There is a need for solutions to these problems and improve the efficiency and quality of customer service.

[0120] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0121] In this invention, the server includes means for converting pre-processed voice data into text data, means for identifying the customer's intent from the text data, and means for generating the optimal response using a generative AI model. This enables highly accurate intent extraction and flexible response generation from voice input.

[0122] "User voice" refers to the audio signals that a user emits to the system.

[0123] "Input method" refers to any hardware or software component used to collect the user's voice. Examples include microphones and voice acquisition software.

[0124] A "terminal" is an electronic device that receives voice data from a user and transmits it to a server. Specifically, this includes smartphones and computers.

[0125] A "server" is a computer system that performs key processes such as preprocessing of audio data, speech recognition, natural language processing, intent extraction, and response generation.

[0126] "Preprocessing methods" refer to software or algorithms used to remove noise and normalize the volume of audio data, and to convert it into a format suitable for analysis. Examples include "SoX" and "Librosa."

[0127] "Means of converting to text data" refers to speech recognition engines that convert audio data into text data. A specific example is a speech recognition API.

[0128] "Means of identifying customer intent" refers to techniques that use natural language processing engines or pre-trained AI models to extract user intent from text data.

[0129] A "generative AI model" is artificial intelligence that generates the optimal response based on the user's intent. Examples include generative AIs such as GPT-3 (registered trademark).

[0130] "Means for generating answers" refers to software or algorithms that use generative AI models to generate answers based on the user's intent.

[0131] "Means of notifying the user of the answer" refers to a system for notifying the user of the generated answer. Specifically, this includes means of sending the answer to the terminal as text data or audio data, and displaying or playing it back as audio.

[0132] "Nuance and emotion" refers to elements that represent the customer's nuances and emotional expressions, extracted from text and audio data.

[0133] "Natural language processing" is a technique for analyzing text data and extracting meaning and intent. Examples include tools and algorithms such as "spaCy" and "BERT."

[0134] Modes for carrying out the invention

[0135] This invention relates to a customer service system that utilizes speech recognition and natural language processing technologies. It takes user voice input, analyzes the voice data in real time to identify the customer's intent, and provides the optimal response. The system consists of a user, a terminal, and a server.

[0136] System Configuration

[0137] User

[0138] Users input questions and requests via voice through the system. When a user speaks into the microphone, the voice data is transmitted to the device.

[0139] terminal

[0140] The terminal records the user's voice in digital format and sends it to the server. Some pre-processing is performed as needed before the received audio data is sent to the server. This pre-processing includes noise reduction and volume normalization. Specifically, "SoX (Sound eXchange)" or "Librosa" is used as the audio collection software on the terminal.

[0141] server

[0142] The server is the core of this system and performs the following main functions:

[0143] Audio data preprocessing (noise reduction and volume normalization)

[0144] Speech recognition (converting speech data to text data)

[0145] Natural language processing (analyzing user nuances and emotions from text data)

[0146] Intent extraction (identifying the user's purpose from text data)

[0147] Answer generation (generates the optimal answer using a generation AI model based on identified intent)

[0148] Submit the response (notify the user of the generated response)

[0149] Hardware and software to be used

[0150] Microphone (hardware): A device for voice input.

[0151] Audio acquisition software for the device: Anvil Studio or Audacity

[0152] Server audio processing software: SoX, Librosa

[0153] Speech recognition engine: Google® Cloud Speech-to-Text API

[0154] Natural language processing engines: spaCy, BERT

[0155] Generative AI model: OpenAI(registered trademark) GPT-3

[0156] Text-to-speech conversion: Google Cloud Text-to-Speech API

[0157] Specific example

[0158] Scenario 1: Inquiry about pricing plans

[0159] 1. User: "What is the recommended pricing plan?"

[0160] 2. Terminal: Records audio and sends it to the server.

[0161] 3. Server: Performs noise reduction and volume normalization.

[0162] 4. Server: Converts audio data to text. "What is the recommended pricing plan?"

[0163] 5. Server: Analyzes text data to identify user intent. "Questions about pricing plans"

[0164] 6. Server: Based on information about pricing plans, the AI ​​generates the optimal response. "Currently, there are three available pricing plans: A, B, and C. Their features are..."

[0165] 7. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[0166] Scenario 2: Buying a new smartphone

[0167] 1. User: "I want to buy a new smartphone."

[0168] 2. Terminal: Records audio and sends it to the server.

[0169] 3. Server: Performs noise reduction and volume normalization.

[0170] 4. Server: Converts audio data to text. "I want to buy a new smartphone."

[0171] 5. Server: Analyzes text data to identify user intent. "Question regarding the purchase of a new smartphone"

[0172] 6. Server: Based on information about the smartphone, the AI ​​generates the optimal answer. "The current new models are X, Y, and Z, and their respective features are..."

[0173] 7. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[0174] Examples of prompt messages include "What is your recommended pricing plan?" and "I'd like to buy a new smartphone."

[0175] In this way, the system can accurately understand the user's intent from their voice input and provide the optimal response. This can lead to improved customer satisfaction and consistent service quality.

[0176] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0177] Step 1:

[0178] User: Speaks a question or request into the microphone. For example, "What is the recommended pricing plan?" This generates audio data.

[0179] Step 2:

[0180] Terminal: Records the user's voice in digital format (e.g., WAV file format). Using voice collection software on the terminal (e.g., Anvil Studio, Audacity), the voice data is saved as digital data. Once recording is complete, the voice data is sent to the server using an HTTP request. To ensure data security, the TLS (Transport Layer Security) encryption protocol is used.

[0181] Step 3:

[0182] Server: Receives audio data and performs preprocessing such as noise reduction and volume normalization. Specifically, it uses libraries such as "SoX (Sound eXchange)" and Python's "Librosa". This process converts the audio data into a format suitable for analysis. For example, noise reduction reduces background noise and improves speech clarity.

[0183] Step 4:

[0184] Server: The server sends the pre-processed audio data to the speech recognition engine and converts the audio data into text data. The Google Cloud Speech-to-Text API is used here. The API is used to generate text data from the audio data. For example, the audio data "What is the best pricing plan?" is converted to the text data "What is the best pricing plan?".

[0185] Step 5:

[0186] Server: Text data generated by speech recognition is processed by a natural language processing engine (e.g., spaCy) to analyze the user's nuances and emotions. Simultaneously, a pre-trained AI model (e.g., BERT) is used to identify the user's intent from the text data. This process extracts meaning and intent from the text data. For example, "a question about pricing plans" might be extracted as the user's intent.

[0187] Step 6:

[0188] Server: Uses a generative AI model (e.g., OpenAI GPT-3) to generate the best answer based on the identified intent. The generative AI model is given a prompt and generates an appropriate answer. For example, based on the intent "Question about pricing plans," it generates an answer such as "Currently available pricing plans are A, B, and C. Their features are..."

[0189] Step 7:

[0190] Server: Formats the generated responses as text and audio data and sends them to the device. The Google Cloud Text-to-Speech API is used to generate the audio data. The device displays or plays the received data audibly to the user, allowing the user to visually or audibly confirm the responses.

[0191] For example, the generated response, "Currently available pricing plans are A, B, and C. Their features are...", is sent to the device, and the device plays the response aloud, allowing the user to hear the content. In this way, the system accurately analyzes the user's intent through concrete actions and provides the most appropriate response.

[0192] (Application Example 1)

[0193] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0194] Traditional customer service systems had a problem in that they made it difficult for users to get quick and accurate answers when interacting directly with staff in physical stores. This could lead to decreased customer satisfaction and increased workload for store staff, resulting in inconsistent service quality.

[0195] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0196] In this invention, the server includes means for transmitting the user's voice data to a smart device, means for displaying or playing back the answer to the user's question via the smart device, and means for analyzing nuances and intonation. This makes it possible to respond quickly and accurately to user questions in physical stores.

[0197] "Means for inputting user voice" refers to a mechanism for recording and acquiring user voice information using a digital device.

[0198] "Means for sending input audio data to a server" refers to communication means for sending audio data captured by a device to a server via a network.

[0199] "Methods for pre-processing audio data" refer to methods for processing captured audio data, such as noise reduction and volume normalization, to prepare it for easier analysis.

[0200] "Means for converting pre-processed audio data into text data" refers to a mechanism that utilizes speech recognition technology to convert pre-processed audio data into text information.

[0201] "Methods for identifying customer intent from text data" refer to methods that use natural language processing technology to analyze text data generated by speech recognition and identify the user's intent and purpose.

[0202] "Means for generating optimal responses based on identified intentions" refers to methods that use AI models to create optimal responses based on relevant information, taking into account the identified user intentions.

[0203] "Means of notifying the user of the generated response" refers to the means by which the server communicates the response it has created to the user in text or audio format.

[0204] "Means for transmitting user voice data to a smart device" refers to communication means for transmitting recorded user voice data to devices such as smartphones and smart glasses.

[0205] "Means for displaying or playing back the user's question via a smart device" refers to means such as a display device for visualizing the received answer on the smart device or a speaker for playing it back as sound.

[0206] System Overview

[0207] This invention is a customer service system that utilizes speech recognition and natural language processing technologies. The system takes voice input from the user, analyzes the voice data in real time to identify the customer's intent, and provides the optimal response. The system consists of a user, a terminal, and a server.

[0208] System Configuration

[0209] User

[0210] Users input questions and requests via voice through the system. When a user speaks into the microphone, the voice data is transmitted to their smart device.

[0211] terminal

[0212] The terminal consists of smart devices such as smartphones and smart glasses, which digitally record the user's voice and send the data to a server. Some pre-processing is performed on the received voice data before it is sent to the server.

[0213] server

[0214] The server is the core of this system and performs the following main functions:

[0215] Audio data preprocessing

[0216] Speech recognition

[0217] Natural Language Processing

[0218] Intent Extraction

[0219] Answer generation

[0220] Submit your response

[0221] Hardware and software to be used

[0222] Hardware:

[0223] Smartphones and smart glasses: Record user voice and communicate with the server.

[0224] Server: Performs data preprocessing, speech recognition, natural language processing, intent extraction, and response generation.

[0225] software:

[0226] Speech recognition engine (e.g., Google Cloud Speech-to-Text)

[0227] Natural language processing engines (e.g., SpaCy, BERT)

[0228] AI models (e.g., GPT-4(registered trademark))

[0229] Data processing and data calculation

[0230] Audio data preprocessing: The server performs preprocessing on the received audio data, such as noise reduction and volume normalization, and converts it into a format suitable for analysis.

[0231] Speech recognition: Pre-processed audio data is sent to the speech recognition engine, which converts the audio data into text data.

[0232] Natural Language Processing and Intent Extraction: Text data generated by speech recognition is processed by a natural language processing engine to analyze the user's nuances and intonation. Next, the user's intent is identified from the text data.

[0233] Response Generation: An AI model is used to generate the optimal response based on the identified intent. The generated response is formatted as text and audio data and sent to the device.

[0234] Specific example

[0235] Check product availability

[0236] 1. User: "Do you have this shirt in stock?"

[0237] 2. Terminal: Records audio and sends it to the server.

[0238] 3. Server: Performs noise reduction and volume normalization, and converts the audio data into text.

[0239] 4. Server: Identify user intent from text data.

[0240] 5. Server: Based on product inventory information, the AI ​​generates the optimal response. "This shirt is currently only available in size M. Other sizes are expected to arrive within a week."

[0241] 6. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[0242] Prompt example

[0243] "Do you have this shirt in stock?"

[0244] → AI model: Generates answers based on inventory information.

[0245] Example output: "This shirt is currently only available in size M. Other sizes will be available within a week."

[0246] In this way, customer service support tools utilizing speech recognition and natural language processing technologies can be used in physical stores. This allows store staff to respond to customer inquiries quickly and accurately, which is expected to improve customer satisfaction.

[0247] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0248] Step 1:

[0249] User voice input

[0250] The user speaks questions or requests into the microphone. The input at this time is audio data.

[0251] Step 2:

[0252] Recording and transmitting audio data

[0253] The terminal records the user's voice in digital format and sends that data to the server. The input here is the recorded voice data, and the output is the digital voice data sent to the server.

[0254] Step 3:

[0255] Audio data preprocessing

[0256] The server performs preprocessing on the received audio data, such as noise reduction and volume normalization, and converts it into a format suitable for analysis. The input is digital audio data transmitted from the terminal, and the output is the preprocessed audio data.

[0257] Step 4:

[0258] Speech recognition

[0259] The server sends the pre-processed audio data to a speech recognition engine (e.g., Google Cloud Speech-to-Text) to convert the audio data into text data. The input is pre-processed audio data, and the output is text data.

[0260] Step 5:

[0261] Natural language processing and intent extraction

[0262] The server processes the text data generated by speech recognition into a natural language processing engine (e.g., SpaCy, BERT) to analyze the user's nuances and intonation. Next, it identifies the user's intent from the analyzed text data. The input is text data, and the output is the identified user intent data.

[0263] Step 6:

[0264] Answer generation

[0265] The server utilizes an AI model (e.g., GPT-4) to generate the optimal response based on the identified intent. The generated response is then formatted as text and audio data. The input is the user's intent data, and the output is the generated response data.

[0266] Step 7:

[0267] Submit your response

[0268] The server sends the generated response data to the terminal. The input is the generated response data, and the output is the response data sent to the terminal.

[0269] Step 8:

[0270] Display or play the answer

[0271] The terminal displays or plays back the received response data to the user in audio format. The input is the response data sent from the server, and the output is the display or audio playback to the user.

[0272] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0273] System Overview

[0274] This invention relates to a customer service system that combines speech recognition, natural language processing (NLP), and sentiment analysis. It takes user voice input, analyzes the voice data in real time to identify the customer's intentions and emotions, and provides the most appropriate response. The system consists of a user, a terminal, and a server.

[0275] System Configuration

[0276] User

[0277] Users input questions and requests via voice through the system. When a user speaks into the microphone, the voice data is transmitted to the device.

[0278] terminal

[0279] The terminal records the user's voice in digital format and sends it to the server. Some pre-processing is performed as needed before the received audio data is sent to the server.

[0280] Server

[0281] The server is the core of this system and performs the following main functions.

[0282] Means for preprocessing voice data

[0283] Means for performing speech recognition

[0284] Means for performing natural language processing

[0285] Means for performing sentiment analysis by a sentiment engine

[0286] Means for extracting intentions

[0287] Means for generating responses

[0288] Means for transmitting responses

[0289] Processing flow

[0290] Voice input and transmission

[0291] 1. User: Speak questions or requests towards the microphone.

[0292] 2. Terminal: Record the user's voice as digital data and transmit the data to the server.

[0293] Preprocessing of voice data

[0294] 3. Server: Perform preprocessing such as noise removal and volume normalization on the received voice data and convert it into a format suitable for analysis. <00,00930>

[0295] Speech recognition

[0296] 4. Server: Send the preprocessed voice data to a speech recognition engine and convert the voice data into text data.

[0297] Natural Language Processing and Intent Extraction

[0298] 5. Server: Apply the text data generated by speech recognition to a natural language processing engine to analyze the nuances and emotions of the user. Next, identify the user's intent from the text data.

[0299] Sentiment Analysis

[0300] 6. Server: Use a sentiment engine to analyze the emotions from the user's voice data and text data. The analysis results are used for identifying the customer's intent.

[0301] Answer Generation and Transmission

[0302] 7. Server: Utilize an AI model to generate an optimal answer based on the identified intent and sentiment analysis results. Format the generated answer as text data and voice data.

[0303] 8. Server: Transmit the formatted answer to the terminal. The terminal displays or plays back the received data in a voice format.

[0304] Specific Example

[0305] Scenario 1: Inquiry about the Pricing Plan

[0306] 1. User: Says, "What is the recommended pricing plan?"

[0307] 2. Terminal: Records the voice and transmits it to the server.

[0308] 3. Server: Perform noise removal and volume normalization.

[0309] 4. Server: Convert the voice data into text. "What is the recommended pricing plan?"

[0310] 5. Server: Analyzes text data to identify user intent. "Questions about pricing plans"

[0311] 6. Server: The emotion engine analyzes the user's emotional state and determines that they are "interested."

[0312] 7. Server: Based on information about pricing plans, the AI ​​generates the optimal response. "Currently, there are three available pricing plans: A, B, and C. Their features are..."

[0313] 8. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[0314] Scenario 2: Buying a new smartphone

[0315] 1. User: "I want to buy a new smartphone."

[0316] 2. Terminal: Records audio and sends it to the server.

[0317] 3. Server: Performs noise reduction and volume normalization.

[0318] 4. Server: Converts audio data to text. "I want to buy a new smartphone."

[0319] 5. Server: Analyzes text data to identify user intent. "Question regarding the purchase of a new smartphone"

[0320] 6. Server: The emotion engine analyzes the user's emotional state and determines that they have expectations.

[0321] 7. Server: Based on information about smartphones, the AI ​​generates the optimal answer. "The current new models are X, Y, and Z, and their respective features are..."

[0322] 8. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[0323] This system can accurately grasp the user's intent and emotions from their voice input and provide the most appropriate response. This not only improves customer satisfaction but also ensures consistent service quality.

[0324] The following describes the processing flow.

[0325] Step 1:

[0326] User: Speak your questions or requests into the microphone.

[0327] Step 2:

[0328] Terminal: Records the user's voice as digital data.

[0329] Step 3:

[0330] Terminal: Sends the recorded audio data to the server via an HTTP request.

[0331] Step 4:

[0332] Server: Analyzes the received audio data and applies a noise reduction filter.

[0333] Step 5:

[0334] Server: Normalizes the volume of the data to a certain level.

[0335] Step 6:

[0336] Server: Passes the pre-processed audio data to the speech recognition engine.

[0337] Step 7:

[0338] Server: The speech recognition engine converts the speech data into text data.

[0339] Step 8:

[0340] Server: Transfers the generated text data to the natural language processing (NLP) engine.

[0341] Step 9:

[0342] Server: The NLP engine performs sentiment analysis on text data, analyzing the customer's nuances and intonation.

[0343] Step 10:

[0344] Server: Uses an emotion engine to analyze the user's emotional state from text and audio data.

[0345] Step 11:

[0346] Server: Uses text intent analysis algorithms to identify customer questions and requests from text data.

[0347] Step 12:

[0348] Server: Based on sentiment analysis and intent analysis results, it utilizes AI models to generate appropriate answers to questions and requests.

[0349] Step 13:

[0350] Server: Adjusts the generated responses according to the user's emotional state and formats them as text and audio data.

[0351] Step 14:

[0352] Server: Sends the formatted response to the terminal as an HTTP response.

[0353] Step 15:

[0354] Terminal: Displays text data received from the server on the screen.

[0355] Step 16:

[0356] Terminal: Plays received audio data through the speaker.

[0357] Step 17:

[0358] User: Review the provided answers and re-enter any questions or requests as needed.

[0359] (Example 2)

[0360] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0361] Conventional speech recognition systems struggled to accurately analyze user intent and emotional state and provide optimal responses in real time. Furthermore, inconsistent customer service quality among different staff members led to decreased customer satisfaction. This compromised the user experience and hindered the provision of efficient customer service.

[0362] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for pre-processing received audio data, means for sending the pre-processed audio data to a speech recognition engine and converting it into text data, means for extracting the user's intent using a natural language processing engine based on the text data, means for analyzing the user's emotional state from the text data using an emotion analysis engine, means for generating an optimal response based on the intent and emotional state identified by a generation AI model, and means for sending the generated response to a terminal and notifying the user. This makes it possible to accurately grasp the intent and emotion from the user's voice input, improving customer satisfaction and standardizing the quality of service.

[0363] A "user" is the person who performs voice input.

[0364] A "terminal" is a device or system for recording a user's voice as digital data and transmitting it to a server.

[0365] A "server" is a core computer system that performs preprocessing of received audio data, speech recognition, natural language processing, sentiment analysis, and response generation.

[0366] "Audio data" refers to data that digitally represents what a user says to their device.

[0367] "Preprocessing" is the process of applying noise reduction and volume normalization to audio data and converting it into a format suitable for analysis.

[0368] A "speech recognition engine" is software or an algorithm used to convert speech data into text data.

[0369] "Text data" refers to character data converted by a speech recognition engine.

[0370] A "natural language processing engine" is software or an algorithm that analyzes text data to identify the user's intent and nuances.

[0371] An "emotion analysis engine" is software or an algorithm used to analyze a user's emotional state from text data.

[0372] A "generative AI model" is an artificial intelligence model that generates optimal answers based on the results of natural language processing and sentiment analysis.

[0373] "Answer" refers to the appropriate response to a user's inquiry, generated by a generative AI model.

[0374] "Notification" is the act of informing the user of the generated response.

[0375] The "system" is a mechanism that integrates all of these means and executes a series of processes to provide the optimal response from the user's voice input.

[0376] System Overview

[0377] This invention relates to a customer service system that combines speech recognition, natural language processing, and sentiment analysis to analyze the user's intent and emotions in real time from their voice input and provide the optimal response. The system consists of a user, a terminal, and a server.

[0378] User

[0379] Users input questions and requests via voice through the system. For example, if a user says "I want to buy a new smartphone" into the microphone, that voice is sent to the device.

[0380] terminal

[0381] The terminal records the user's voice in digital format and sends it to the server. Upon receiving the audio data, the terminal uses the "pyaudio" library to capture the audio and save it as digital data (for example, with the filename "input.wav"). The recorded digital audio data is transferred to the server in real time. Basic pre-processing such as noise reduction and volume normalization can also be performed.

[0382] server

[0383] The server, as the core of the system, performs the following main functions:

[0384] Audio data preprocessing: The server uses an audio analysis library such as "librosa" to remove noise and normalize the volume of the received audio data, and then converts it into a format suitable for analysis.

[0385] Speech Recognition: Pre-processed audio data is sent to a speech recognition engine (e.g., "Google Cloud Speech-to-Text API") to convert the audio into text data.

[0386] Natural Language Processing and Intent Extraction: Text data obtained through speech recognition is processed by natural language processing engines such as "SpaCy" and "BERT" to analyze the user's intent and nuances. For example, from the text "I want to buy a new smartphone," the intent "a question about purchasing a new smartphone" is identified.

[0387] Emotion Analysis: Using an emotion analysis engine (e.g., "IBM Watson® Tone Analyzer" or "DeepMoji"), the emotional state of the user is analyzed from their text data. The analysis results reveal emotions such as "the user has expectations."

[0388] Response Generation: Based on the identified intent and sentiment analysis results, the system generates the optimal response using a generative AI model (e.g., OpenAI's GPT-3). For example, it might generate a response such as, "The current new models are X, Y, and Z, and their respective features are..."

[0389] Notification: The generated response is sent to the device to notify the user. The device displays or plays the received response in audio format. Audio output tools such as "gTTS" or "Amazon Polly" are used for this purpose.

[0390] Specific example

[0391] Scenario 1: Inquiry about pricing plans

[0392] 1. User: "What is the recommended pricing plan?"

[0393] 2. Terminal: Records audio and sends it to the server.

[0394] 3. Server: Performs noise reduction and volume normalization.

[0395] 4. Server: Converts audio data to text. "What is the recommended pricing plan?"

[0396] 5. Server: Analyzes text data to identify user intent. "Questions about pricing plans"

[0397] 6. Server: The emotion engine analyzes the user's emotional state and determines that they are "interested."

[0398] 7. Server: Based on information about pricing plans, the AI ​​generates the optimal response. "Currently, there are three available pricing plans: A, B, and C. Their features are..."

[0399] 8. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[0400] Scenario 2: Buying a new smartphone

[0401] 1. User: "I want to buy a new smartphone."

[0402] 2. Terminal: Records audio and sends it to the server.

[0403] 3. Server: Performs noise reduction and volume normalization.

[0404] 4. Server: Converts audio data to text. "I want to buy a new smartphone."

[0405] 5. Server: Analyzes text data to identify user intent. "Question regarding the purchase of a new smartphone"

[0406] 6. Server: The emotion engine analyzes the user's emotional state and determines that they have expectations.

[0407] 7. Server: Based on information about smartphones, the AI ​​generates the optimal answer. "The current new models are X, Y, and Z, and their respective features are..."

[0408] 8. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[0409] Example of a prompt

[0410] Let's assume a user inputs "I want to buy a new smartphone" via voice. The system converts the voice to text, identifies information about new smartphones through natural language processing and sentiment analysis, and then generates the best response. Please explain the detailed processing steps for each step of this system.

[0411] This system can accurately grasp the user's intent and emotions from their voice input and provide the most appropriate response. This not only improves customer satisfaction but also ensures consistent service quality.

[0412] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0413] Step 1:

[0414] User: Speaks a question or request into the microphone. Voice input is performed, for example, saying "I want to buy a new smartphone." This voice is the user's input.

[0415] Step 2:

[0416] Terminal: Records the user's voice as digital data and sends it to the server. Specifically, it uses the "pyaudio" library to capture the audio and save it in digital format. The recorded audio data (input.wav) becomes the terminal's output. This digital data is then sent directly to the server.

[0417] Step 3:

[0418] Server: Preprocesses the received audio data. Specifically, it uses the "librosa" library to perform noise reduction and volume normalization. The input for this preprocessing is the recorded audio data (input.wav), and the output is the preprocessed audio data. For example, noise is reduced using librosa.effects.preemphasis(input_signal).

[0419] Step 4:

[0420] Server: Sends pre-processed audio data to the speech recognition engine and converts the audio data into text data. The recognize method of the "Google Cloud Speech-to-Text API" is used. The input to this process is pre-processed audio data, and the output is the text data "I want to buy a new smartphone."

[0421] Step 5:

[0422] Server: The server processes the text data obtained from speech recognition into a natural language processing engine to extract the user's intent. The "SpaCy" library is used for this process. Specifically, the text data is analyzed using NLP (text) to identify the intent as "a question about purchasing a new smartphone." The input to this process is the text data obtained from speech recognition, and the output is user intent data.

[0423] Step 6:

[0424] Server: The server processes the text data obtained through natural language processing into an emotion analysis engine to analyze the user's emotional state. Specifically, it uses "IBM Watson Tone Analyzer" to execute tone_analyzer.tone(tone_input, content_type="application / json").get_result(). The input to this process is text data that identifies intent, and the output is analyzed emotion data (for example, "has expectations").

[0425] Step 7:

[0426] Server: Based on the identified intent and sentiment analysis results, it generates the optimal response using a generative AI model. For example, "GPT-3" is used as the generative AI model. The input to this process is user intent data and sentiment data, and the output is the generated response "The current new model is X, Y, Z, and its respective features are..."

[0427] Step 8:

[0428] Server: The server sends the generated response to the device and notifies the user. HTTP requests are often used as the notification method. For example, requests.post("http: / / device.endpoint", data=response_data) is used. The device displays or plays the received response data in audio format. When playing in audio format, "gTTS" is used, and the command tts = gTTS(response_text, lang='ja') is used. The input to this process is the generated response data, and the output is the information notified to the user.

[0429] (Application Example 2)

[0430] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0431] In brick-and-mortar stores, it is crucial to respond quickly and accurately to customer questions and requests in order to improve the quality and efficiency of customer service. However, conventional customer service systems have difficulty immediately grasping customer intentions and emotions and providing appropriate answers. Furthermore, it is difficult for store staff to respond in real time, which can lead to a decrease in customer satisfaction. This invention aims to solve these problems and provide a means for efficient and effective customer service in brick-and-mortar stores.

[0432] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0433] In this invention, the server includes means for preprocessing audio data, means for converting the preprocessed audio data into text data, means for identifying the customer's intent from the text data, and a device for visually displaying the generated response. This makes it possible to analyze the customer's voice in real time, instantly grasp their intent and emotions, and provide the optimal response.

[0434] "Means of inputting user voice" refers to devices or programs that acquire voice data spoken by the user.

[0435] "Means for transmitting input audio data to a server" refers to devices or programs that transfer audio data acquired from a user to a server via a communication network.

[0436] "Means for preprocessing audio data" refers to devices or programs that perform processing on audio data, such as noise reduction and volume normalization, and convert it into a format that can be analyzed.

[0437] "Means for converting pre-processed audio data into text data" refers to devices or programs that analyze audio data and convert it into corresponding string data.

[0438] "Means for identifying customer intent from text data" refers to devices or programs that analyze text data and identify customer objectives and requirements from it.

[0439] "Means for generating optimal responses based on identified intentions" refers to devices or programs that produce appropriate responses in accordance with the customer's intentions.

[0440] "Means for notifying the user of the generated response" refers to devices or programs that communicate and provide the generated response to the user.

[0441] A "device for visually displaying generated answers" refers to a device that visualizes generated answers and presents them through a display, glasses-type device, or the like.

[0442] System Overview

[0443] This system is a customer service system for brick-and-mortar stores that takes user voice input, analyzes it to identify customer intent and emotions, and provides the most appropriate response. Its main components are the user, a terminal, and a server, which work together to function.

[0444] User

[0445] The user is a store employee wearing smart glasses. When a customer speaks to the employee, the employee receives the voice through the microphone in the smart glasses.

[0446] terminal

[0447] The device is a pair of smart glasses and has the following features:

[0448] Microphone for capturing customer voices

[0449] A display that provides visual information to the store staff.

[0450] Network function for communicating with the server

[0451] server

[0452] The server is the core of this system and performs the following main functions:

[0453] Means for preprocessing audio data

[0454] Means for performing speech recognition

[0455] Means for performing natural language processing

[0456] Means for performing emotion analysis

[0457] Means of identifying intent

[0458] Means for generating answers

[0459] Means of sending the answer to the device

[0460] Hardware and software to be used

[0461] Hardware: Smart glasses (e.g., Google Glass®, Microsoft HoloLens®)

[0462] software:

[0463] Speech recognition engine (e.g., Google Cloud Speech-to-Text API)

[0464] Natural language processing engine (e.g., Microsoft Azure Cognitive Services)

[0465] Sentiment analysis engine (e.g., IBM Watson Tone Analyzer)

[0466] Answer generation engine (e.g., OpenAI GPT-4)

[0467] Data processing flow and specific examples

[0468] 1. Voice acquisition:

[0469] A store employee wearing smart glasses uses a microphone to capture the customer's voice. For example, the customer might ask, "What products do you recommend for this season?"

[0470] 2. Sending audio data:

[0471] The acquired audio data is sent from the terminal to the server.

[0472] 3. Pre-treatment:

[0473] The server preprocesses the received audio data, performing noise reduction and volume normalization.

[0474] 4. Speech recognition:

[0475] This process converts pre-processed audio data into text data. In the previous example, the resulting text would be "What products do you recommend for this season?"

[0476] 5. Natural Language Processing:

[0477] The server analyzes the text data to identify the customer's intent. In this case, it is identified as "asking for product recommendations."

[0478] 6. Emotion analysis:

[0479] Customer emotions are analyzed from text and audio data. For example, the emotion of "interest" can be identified.

[0480] 7. Answer generation:

[0481] The server generates the optimal response based on the customer's intent and emotions. For example, it might say, "Our recommended products for this season are A, B, and C. Their features are..."

[0482] 8. Submit and display your response:

[0483] The generated response is displayed on the smart glasses' screen. The store clerk can then refer to the response and answer the customer directly.

[0484] Example of a prompt

[0485] Product: Customer support application

[0486] Model: Smart Glasses

[0487] Objective: To support customer service in physical stores.

[0488] Features: Speech recognition, natural language processing, sentiment analysis, response generation

[0489] Scenario: A customer asks a question about a product, and a store employee uses smart glasses to check the best answer and provide it instantly.

[0490] In this way, by accurately understanding the customer's intent and emotions from their voice input and providing the optimal response, the quality and efficiency of customer service in physical stores can be improved.

[0491] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0492] Step 1:

[0493] The user wears smart glasses and inputs the customer's voice. The smart glasses' microphone picks up the customer's speech and captures it as audio data. The input might be a customer question such as, "What products do you recommend for this season?"

[0494] Step 2:

[0495] The device (smart glasses) preprocesses the captured audio data. This preprocessing involves data manipulation such as noise reduction and volume normalization. As a result, audio data in a format that is easy for the server to analyze is output.

[0496] Step 3:

[0497] Pre-processed audio data is sent from the terminal to the server. In this process, the audio data is transferred to the server in real time via the internet. The input is the pre-processed audio data, and the output is the digital audio data that has reached the server.

[0498] Step 4:

[0499] The server passes the received audio data to a speech recognition engine (e.g., Google Cloud Speech-to-Text API) and converts it into text data. For example, the text "What products do you recommend this season?" is generated.

[0500] Step 5:

[0501] The server processes the generated text data using a natural language processing engine (e.g., Microsoft Azure Cognitive Services) to analyze the customer's intent. This process identifies the intent as "asking for product recommendations." The input is text data, and the output is the identified intent.

[0502] Step 6:

[0503] The server simultaneously uses an emotion analysis engine (e.g., IBM Watson Tone Analyzer) to analyze emotions from text and audio data. For example, the emotion "interested" might be identified. The input is text and audio data, and the output is the identified emotion.

[0504] Step 7:

[0505] The server uses a generative AI model (e.g., OpenAI GPT-4) to generate the most appropriate response based on identified intentions and emotions. For example, it might generate a response like, "Products A, B, and C are recommended for this season. Their characteristics are..." The input is intention and emotion, and the output is the generated response text.

[0506] Step 8:

[0507] The server sends the generated response to the device (smart glasses). In this process, the generated text data is transferred to the device via the internet. The input is the generated response text, and the output is the response data received by the device.

[0508] Step 9:

[0509] The terminal visually displays the generated response on the smart glasses' screen. The store clerk can then explain the product to the customer while looking at the displayed response. For example, the clerk might start by saying, "These are our recommended products for this season." The input is the response data received by the terminal, and the output is the visual information displayed on the screen.

[0510] In this way, a series of processes are realized in which the user provides voice input, the server analyzes it to generate the optimal response, and the terminal visualizes it. This system enables quick and accurate customer service in physical stores.

[0511] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0512] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0513] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0514] [Second Embodiment]

[0515] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0516] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0517] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0518] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0519] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0520] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0521] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0522] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0523] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0524] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0525] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0526] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0527] System Overview

[0528] This invention relates to a customer service system that utilizes speech recognition and natural language processing technologies. It takes user voice input, analyzes the voice data in real time to identify the customer's intent, and provides the optimal response. The system consists of a user, a terminal, and a server.

[0529] System Configuration

[0530] User

[0531] Users input questions and requests via voice through the system. When a user speaks into the microphone, the voice data is transmitted to the device.

[0532] terminal

[0533] The terminal records the user's voice in digital format and sends it to the server. Some pre-processing is performed as needed before the received audio data is sent to the server.

[0534] server

[0535] The server is the core of this system and performs the following main functions:

[0536] Audio data preprocessing

[0537] Speech recognition

[0538] Natural Language Processing

[0539] Intent Extraction

[0540] Answer generation

[0541] Submit your response

[0542] Processing flow

[0543] Voice input and transmission

[0544] 1. User: Speak your question or request into the microphone.

[0545] 2. Terminal: Records the user's voice as digital data and sends that data to the server.

[0546] Audio data preprocessing

[0547] 3. Server: The server performs preprocessing on the received audio data, such as noise reduction and volume normalization, and converts it into a format suitable for analysis.

[0548] Speech recognition

[0549] 4. Server: Sends the pre-processed audio data to the speech recognition engine and converts the audio data into text data.

[0550] Natural language processing and intent extraction

[0551] 5. Server: The text data generated by speech recognition is fed into a natural language processing engine to analyze the user's nuances and emotions. Next, the user's intent is identified from the text data.

[0552] Response generation and submission

[0553] 6. Server: Utilizes an AI model to generate the optimal response based on the identified intent. The generated response is then formatted as text and audio data.

[0554] 7. Server: Sends the formatted response to the terminal. The terminal displays or plays the received data in audio format.

[0555] Specific example

[0556] Scenario 1: Inquiry about pricing plans

[0557] 1. User: "What is the recommended pricing plan?"

[0558] 2. Terminal: Records audio and sends it to the server.

[0559] 3. Server: Performs noise reduction and volume normalization.

[0560] 4. Server: Converts audio data to text. "What is the recommended pricing plan?"

[0561] 5. Server: Analyzes text data to identify user intent. "Questions about pricing plans"

[0562] 6. Server: Based on information about pricing plans, the AI ​​generates the optimal response. "Currently, there are three available pricing plans: A, B, and C. Their features are..."

[0563] 7. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[0564] Scenario 2: Buying a new smartphone

[0565] 1. User: "I want to buy a new smartphone."

[0566] 2. Terminal: Records audio and sends it to the server.

[0567] 3. Server: Performs noise reduction and volume normalization.

[0568] 4. Server: Converts audio data to text. "I want to buy a new smartphone."

[0569] 5. Server: Analyzes text data to identify user intent. "Question regarding the purchase of a new smartphone"

[0570] 6. Server: Based on information about the smartphone, the AI ​​generates the optimal answer. "The current new models are X, Y, and Z, and their respective features are..."

[0571] 7. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[0572] In this way, the system can accurately understand the user's intent from their voice input and provide the optimal response. This can lead to improved customer satisfaction and consistent service quality.

[0573] The following describes the processing flow.

[0574] Step 1:

[0575] User: Speak your questions or requests into the microphone.

[0576] Step 2:

[0577] Terminal: Records the user's voice as digital data.

[0578] Step 3:

[0579] Terminal: Sends the recorded audio data to the server via an HTTP request.

[0580] Step 4:

[0581] Server: Analyzes the received audio data and applies a noise reduction filter.

[0582] Step 5:

[0583] Server: Normalizes the volume of the data to a certain level.

[0584] Step 6:

[0585] Server: Passes the pre-processed audio data to the speech recognition engine.

[0586] Step 7:

[0587] Server: The speech recognition engine converts the speech data into text data.

[0588] Step 8:

[0589] Server: Transfers the generated text data to the natural language processing (NLP) engine.

[0590] Step 9:

[0591] Server: The NLP engine performs sentiment analysis on text data, analyzing the customer's nuances and intonation.

[0592] Step 10:

[0593] Server: Uses text intent analysis algorithms to identify customer questions and requests from text data.

[0594] Step 11:

[0595] Server: Based on customer intent, the response generation module uses AI to generate the optimal response.

[0596] Step 12:

[0597] Server: Formats the generated responses into text and audio data.

[0598] Step 13:

[0599] Server: Sends the formatted response to the terminal as an HTTP response.

[0600] Step 14:

[0601] Terminal: Displays text data received from the server on the screen.

[0602] Step 15:

[0603] Terminal: Plays received audio data through the speaker.

[0604] Step 16:

[0605] User: Review the provided answers and re-enter any questions or requests as needed.

[0606] (Example 1)

[0607] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0608] Conventional speech recognition systems have limited processes for generating appropriate responses from voice input, making it difficult to accurately identify user intent. Furthermore, the generated responses are often uniform and may not adequately address customer questions or requests. There is a need for solutions to these problems and improve the efficiency and quality of customer service.

[0609] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0610] In this invention, the server includes means for converting pre-processed voice data into text data, means for identifying the customer's intent from the text data, and means for generating the optimal response using a generative AI model. This enables highly accurate intent extraction and flexible response generation from voice input.

[0611] "User voice" refers to the audio signals that a user emits to the system.

[0612] "Input method" refers to any hardware or software component used to collect the user's voice. Examples include microphones and voice acquisition software.

[0613] A "terminal" is an electronic device that receives voice data from a user and transmits it to a server. Specifically, this includes smartphones and computers.

[0614] A "server" is a computer system that performs key processes such as preprocessing of audio data, speech recognition, natural language processing, intent extraction, and response generation.

[0615] "Preprocessing methods" refer to software or algorithms used to remove noise and normalize the volume of audio data, and to convert it into a format suitable for analysis. Examples include "SoX" and "Librosa."

[0616] "Means of converting to text data" refers to speech recognition engines that convert audio data into text data. A specific example is a speech recognition API.

[0617] "Means of identifying customer intent" refers to techniques that use natural language processing engines or pre-trained AI models to extract user intent from text data.

[0618] A "generative AI model" is artificial intelligence that generates the optimal response based on the user's intent. Examples include generative AIs such as GPT-3.

[0619] "Means for generating answers" refers to software or algorithms that use generative AI models to generate answers based on the user's intent.

[0620] "Means of notifying the user of the answer" refers to a system for notifying the user of the generated answer. Specifically, this includes means of sending the answer to the terminal as text data or audio data, and displaying or playing it back as audio.

[0621] "Nuance and emotion" refers to elements that represent the customer's nuances and emotional expressions, extracted from text and audio data.

[0622] "Natural language processing" is a technique for analyzing text data and extracting meaning and intent. Examples include tools and algorithms such as "spaCy" and "BERT."

[0623] Modes for carrying out the invention

[0624] This invention relates to a customer service system that utilizes speech recognition and natural language processing technologies. It takes user voice input, analyzes the voice data in real time to identify the customer's intent, and provides the optimal response. The system consists of a user, a terminal, and a server.

[0625] System Configuration

[0626] User

[0627] Users input questions and requests via voice through the system. When a user speaks into the microphone, the voice data is transmitted to the device.

[0628] terminal

[0629] The terminal records the user's voice in digital format and sends it to the server. Some pre-processing is performed as needed before the received audio data is sent to the server. This pre-processing includes noise reduction and volume normalization. Specifically, "SoX (Sound eXchange)" or "Librosa" is used as the audio collection software on the terminal.

[0630] server

[0631] The server is the core of this system and performs the following main functions:

[0632] Audio data preprocessing (noise reduction and volume normalization)

[0633] Speech recognition (converting speech data to text data)

[0634] Natural language processing (analyzing user nuances and emotions from text data)

[0635] Intent extraction (identifying the user's purpose from text data)

[0636] Answer generation (generates the optimal answer using a generation AI model based on identified intent)

[0637] Submit the response (notify the user of the generated response)

[0638] Hardware and software to be used

[0639] Microphone (hardware): A device for voice input.

[0640] Audio acquisition software for the device: Anvil Studio or Audacity

[0641] Server audio processing software: SoX, Librosa

[0642] Speech recognition engine: Google Cloud Speech-to-Text API

[0643] Natural language processing engines: spaCy, BERT

[0644] Generative AI model: OpenAI GPT-3

[0645] Text-to-speech conversion: Google Cloud Text-to-Speech API

[0646] Specific example

[0647] Scenario 1: Inquiry about pricing plans

[0648] 1. User: "What is the recommended pricing plan?"

[0649] 2. Terminal: Records audio and sends it to the server.

[0650] 3. Server: Performs noise reduction and volume normalization.

[0651] 4. Server: Converts audio data to text. "What is the recommended pricing plan?"

[0652] 5. Server: Analyzes text data to identify user intent. "Questions about pricing plans"

[0653] 6. Server: Based on information about pricing plans, the AI ​​generates the optimal response. "Currently, there are three available pricing plans: A, B, and C. Their features are..."

[0654] 7. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[0655] Scenario 2: Buying a new smartphone

[0656] 1. User: "I want to buy a new smartphone."

[0657] 2. Terminal: Records audio and sends it to the server.

[0658] 3. Server: Performs noise reduction and volume normalization.

[0659] 4. Server: Converts audio data to text. "I want to buy a new smartphone."

[0660] 5. Server: Analyzes text data to identify user intent. "Question regarding the purchase of a new smartphone"

[0661] 6. Server: Based on information about the smartphone, the AI ​​generates the optimal answer. "The current new models are X, Y, and Z, and their respective features are..."

[0662] 7. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[0663] Examples of prompt messages include "What is your recommended pricing plan?" and "I'd like to buy a new smartphone."

[0664] In this way, the system can accurately understand the user's intent from their voice input and provide the optimal response. This can lead to improved customer satisfaction and consistent service quality.

[0665] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0666] Step 1:

[0667] User: Speaks a question or request into the microphone. For example, "What is the recommended pricing plan?" This generates audio data.

[0668] Step 2:

[0669] Terminal: Records the user's voice in digital format (e.g., WAV file format). Using voice collection software on the terminal (e.g., Anvil Studio, Audacity), the voice data is saved as digital data. Once recording is complete, the voice data is sent to the server using an HTTP request. To ensure data security, the TLS (Transport Layer Security) encryption protocol is used.

[0670] Step 3:

[0671] Server: Receives audio data and performs preprocessing such as noise reduction and volume normalization. Specifically, it uses libraries such as "SoX (Sound eXchange)" and Python's "Librosa". This process converts the audio data into a format suitable for analysis. For example, noise reduction reduces background noise and improves speech clarity.

[0672] Step 4:

[0673] Server: The server sends the pre-processed audio data to the speech recognition engine and converts the audio data into text data. The Google Cloud Speech-to-Text API is used here. The API is used to generate text data from the audio data. For example, the audio data "What is the best pricing plan?" is converted to the text data "What is the best pricing plan?".

[0674] Step 5:

[0675] Server: Text data generated by speech recognition is processed by a natural language processing engine (e.g., spaCy) to analyze the user's nuances and emotions. Simultaneously, a pre-trained AI model (e.g., BERT) is used to identify the user's intent from the text data. This process extracts meaning and intent from the text data. For example, "a question about pricing plans" might be extracted as the user's intent.

[0676] Step 6:

[0677] Server: Uses a generative AI model (e.g., OpenAI GPT-3) to generate the best answer based on the identified intent. The generative AI model is given a prompt and generates an appropriate answer. For example, based on the intent "Question about pricing plans," it generates an answer such as "Currently available pricing plans are A, B, and C. Their features are..."

[0678] Step 7:

[0679] Server: Formats the generated responses as text and audio data and sends them to the device. The Google Cloud Text-to-Speech API is used to generate the audio data. The device displays or plays the received data audibly to the user, allowing the user to visually or audibly confirm the responses.

[0680] For example, the generated response, "Currently available pricing plans are A, B, and C. Their features are...", is sent to the device, and the device plays the response aloud, allowing the user to hear the content. In this way, the system accurately analyzes the user's intent through concrete actions and provides the most appropriate response.

[0681] (Application Example 1)

[0682] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0683] Traditional customer service systems had a problem in that they made it difficult for users to get quick and accurate answers when interacting directly with staff in physical stores. This could lead to decreased customer satisfaction and increased workload for store staff, resulting in inconsistent service quality.

[0684] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0685] In this invention, the server includes means for transmitting the user's voice data to a smart device, means for displaying or playing back the answer to the user's question via the smart device, and means for analyzing nuances and intonation. This makes it possible to respond quickly and accurately to user questions in physical stores.

[0686] "Means for inputting user voice" refers to a mechanism for recording and acquiring user voice information using a digital device.

[0687] "Means for sending input audio data to a server" refers to communication means for sending audio data captured by a device to a server via a network.

[0688] "Methods for pre-processing audio data" refer to methods for processing captured audio data, such as noise reduction and volume normalization, to prepare it for easier analysis.

[0689] "Means for converting pre-processed audio data into text data" refers to a mechanism that utilizes speech recognition technology to convert pre-processed audio data into text information.

[0690] "Methods for identifying customer intent from text data" refer to methods that use natural language processing technology to analyze text data generated by speech recognition and identify the user's intent and purpose.

[0691] "Means for generating optimal responses based on identified intentions" refers to methods that use AI models to create optimal responses based on relevant information, taking into account the identified user intentions.

[0692] "Means of notifying the user of the generated response" refers to the means by which the server communicates the response it has created to the user in text or audio format.

[0693] "Means for transmitting user voice data to a smart device" refers to communication means for transmitting recorded user voice data to devices such as smartphones and smart glasses.

[0694] "Means for displaying or playing back the user's question via a smart device" refers to means such as a display device for visualizing the received answer on the smart device or a speaker for playing it back as sound.

[0695] System Overview

[0696] This invention is a customer service system that utilizes speech recognition and natural language processing technologies. The system takes voice input from the user, analyzes the voice data in real time to identify the customer's intent, and provides the optimal response. The system consists of a user, a terminal, and a server.

[0697] System Configuration

[0698] User

[0699] Users input questions and requests via voice through the system. When a user speaks into the microphone, the voice data is transmitted to their smart device.

[0700] terminal

[0701] The terminal consists of smart devices such as smartphones and smart glasses, which digitally record the user's voice and send the data to a server. Some pre-processing is performed on the received voice data before it is sent to the server.

[0702] server

[0703] The server is the core of this system and performs the following main functions:

[0704] Audio data preprocessing

[0705] Speech recognition

[0706] Natural Language Processing

[0707] Intent Extraction

[0708] Answer generation

[0709] Submit your response

[0710] Hardware and software to be used

[0711] Hardware:

[0712] Smartphones and smart glasses: Record user voice and communicate with the server.

[0713] Server: Performs data preprocessing, speech recognition, natural language processing, intent extraction, and response generation.

[0714] software:

[0715] Speech recognition engine (e.g., Google Cloud Speech-to-Text)

[0716] Natural language processing engines (e.g., SpaCy, BERT)

[0717] AI models (e.g., GPT-4)

[0718] Data processing and data calculation

[0719] Audio data preprocessing: The server performs preprocessing on the received audio data, such as noise reduction and volume normalization, and converts it into a format suitable for analysis.

[0720] Speech recognition: Pre-processed audio data is sent to the speech recognition engine, which converts the audio data into text data.

[0721] Natural Language Processing and Intent Extraction: Text data generated by speech recognition is processed by a natural language processing engine to analyze the user's nuances and intonation. Next, the user's intent is identified from the text data.

[0722] Response Generation: An AI model is used to generate the optimal response based on the identified intent. The generated response is formatted as text and audio data and sent to the device.

[0723] Specific example

[0724] Check product availability

[0725] 1. User: "Do you have this shirt in stock?"

[0726] 2. Terminal: Records audio and sends it to the server.

[0727] 3. Server: Performs noise reduction and volume normalization, and converts the audio data into text.

[0728] 4. Server: Identify user intent from text data.

[0729] 5. Server: Based on product inventory information, the AI ​​generates the optimal response. "This shirt is currently only available in size M. Other sizes are expected to arrive within a week."

[0730] 6. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[0731] Prompt example

[0732] "Do you have this shirt in stock?"

[0733] → AI model: Generates answers based on inventory information.

[0734] Example output: "This shirt is currently only available in size M. Other sizes will be available within a week."

[0735] In this way, customer service support tools utilizing speech recognition and natural language processing technologies can be used in physical stores. This allows store staff to respond to customer inquiries quickly and accurately, which is expected to improve customer satisfaction.

[0736] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0737] Step 1:

[0738] User voice input

[0739] The user speaks questions or requests into the microphone. The input at this time is audio data.

[0740] Step 2:

[0741] Recording and transmitting audio data

[0742] The terminal records the user's voice in digital format and sends that data to the server. The input here is the recorded voice data, and the output is the digital voice data sent to the server.

[0743] Step 3:

[0744] Audio data preprocessing

[0745] The server performs preprocessing on the received audio data, such as noise reduction and volume normalization, and converts it into a format suitable for analysis. The input is digital audio data transmitted from the terminal, and the output is the preprocessed audio data.

[0746] Step 4:

[0747] Speech recognition

[0748] The server sends the pre-processed audio data to a speech recognition engine (e.g., Google Cloud Speech-to-Text) to convert the audio data into text data. The input is pre-processed audio data, and the output is text data.

[0749] Step 5:

[0750] Natural language processing and intent extraction

[0751] The server processes the text data generated by speech recognition into a natural language processing engine (e.g., SpaCy, BERT) to analyze the user's nuances and intonation. Next, it identifies the user's intent from the analyzed text data. The input is text data, and the output is the identified user intent data.

[0752] Step 6:

[0753] Answer generation

[0754] The server utilizes an AI model (e.g., GPT-4) to generate the optimal response based on the identified intent. The generated response is then formatted as text and audio data. The input is the user's intent data, and the output is the generated response data.

[0755] Step 7:

[0756] Submit your response

[0757] The server sends the generated response data to the terminal. The input is the generated response data, and the output is the response data sent to the terminal.

[0758] Step 8:

[0759] Display or play the answer

[0760] The terminal displays or plays back the received response data to the user in audio format. The input is the response data sent from the server, and the output is the display or audio playback to the user.

[0761] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0762] System Overview

[0763] This invention relates to a customer service system that combines speech recognition, natural language processing (NLP), and sentiment analysis. It takes user voice input, analyzes the voice data in real time to identify the customer's intentions and emotions, and provides the most appropriate response. The system consists of a user, a terminal, and a server.

[0764] System Configuration

[0765] User

[0766] Users input questions and requests via voice through the system. When a user speaks into the microphone, the voice data is transmitted to the device.

[0767] terminal

[0768] The terminal records the user's voice in digital format and sends it to the server. Some pre-processing is performed as needed before the received audio data is sent to the server.

[0769] server

[0770] The server is the core of this system and performs the following main functions:

[0771] Means for preprocessing audio data

[0772] Means for performing speech recognition

[0773] Means of natural language processing

[0774] A method for performing emotion analysis using an emotion engine.

[0775] Means for extracting intent

[0776] Means for generating answers

[0777] Method for submitting a response

[0778] Processing flow

[0779] Voice input and transmission

[0780] 1. User: Speak your question or request into the microphone.

[0781] 2. Terminal: Records the user's voice as digital data and sends that data to the server.

[0782] Audio data preprocessing

[0783] 3. Server: The server performs preprocessing on the received audio data, such as noise reduction and volume normalization, and converts it into a format suitable for analysis.

[0784] Speech recognition

[0785] 4. Server: Sends the pre-processed audio data to the speech recognition engine and converts the audio data into text data.

[0786] Natural language processing and intent extraction

[0787] 5. Server: The text data generated by speech recognition is fed into a natural language processing engine to analyze the user's nuances and emotions. Next, the user's intent is identified from the text data.

[0788] Emotion analysis

[0789] 6. Server: Uses an emotion engine to analyze emotions from user voice and text data. The analysis results are used to identify customer intentions.

[0790] Response generation and submission

[0791] 7. Server: Utilizes an AI model to generate the optimal response based on the identified intent and sentiment analysis results. The generated response is then formatted as text and audio data.

[0792] 8. Server: Sends the formatted response to the terminal. The terminal displays or plays the received data in audio format.

[0793] Specific example

[0794] Scenario 1: Inquiry about pricing plans

[0795] 1. User: "What is the recommended pricing plan?"

[0796] 2. Terminal: Records audio and sends it to the server.

[0797] 3. Server: Performs noise reduction and volume normalization.

[0798] 4. Server: Converts audio data to text. "What is the recommended pricing plan?"

[0799] 5. Server: Analyzes text data to identify user intent. "Questions about pricing plans"

[0800] 6. Server: The emotion engine analyzes the user's emotional state and determines that they are "interested."

[0801] 7. Server: Based on information about pricing plans, the AI ​​generates the optimal response. "Currently, there are three available pricing plans: A, B, and C. Their features are…"

[0802] 8. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[0803] Scenario 2: Buying a new smartphone

[0804] 1. User: "I want to buy a new smartphone."

[0805] 2. Terminal: Records audio and sends it to the server.

[0806] 3. Server: Performs noise reduction and volume normalization.

[0807] 4. Server: Converts audio data to text. "I want to buy a new smartphone."

[0808] 5. Server: Analyzes text data to identify user intent. "Question regarding the purchase of a new smartphone"

[0809] 6. Server: The emotion engine analyzes the user's emotional state and determines that they have expectations.

[0810] 7. Server: Based on information about smartphones, the AI ​​generates the optimal answer. "The current new models are X, Y, and Z, and their respective features are..."

[0811] 8. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[0812] This system can accurately grasp the user's intent and emotions from their voice input and provide the most appropriate response. This not only improves customer satisfaction but also ensures consistent service quality.

[0813] The following describes the processing flow.

[0814] Step 1:

[0815] User: Speak your questions or requests into the microphone.

[0816] Step 2:

[0817] Terminal: Records the user's voice as digital data.

[0818] Step 3:

[0819] Terminal: Sends the recorded audio data to the server via an HTTP request.

[0820] Step 4:

[0821] Server: Analyzes the received audio data and applies a noise reduction filter.

[0822] Step 5:

[0823] Server: Normalizes the volume of the data to a certain level.

[0824] Step 6:

[0825] Server: Passes the pre-processed audio data to the speech recognition engine.

[0826] Step 7:

[0827] Server: The speech recognition engine converts the speech data into text data.

[0828] Step 8:

[0829] Server: Transfers the generated text data to the natural language processing (NLP) engine.

[0830] Step 9:

[0831] Server: The NLP engine performs sentiment analysis on text data, analyzing the customer's nuances and intonation.

[0832] Step 10:

[0833] Server: Uses an emotion engine to analyze the user's emotional state from text and audio data.

[0834] Step 11:

[0835] Server: Uses text intent analysis algorithms to identify customer questions and requests from text data.

[0836] Step 12:

[0837] Server: Based on sentiment analysis and intent analysis results, it utilizes AI models to generate appropriate answers to questions and requests.

[0838] Step 13:

[0839] Server: Adjusts the generated responses according to the user's emotional state and formats them as text and audio data.

[0840] Step 14:

[0841] Server: Sends the formatted response to the terminal as an HTTP response.

[0842] Step 15:

[0843] Terminal: Displays text data received from the server on the screen.

[0844] Step 16:

[0845] Terminal: Plays received audio data through the speaker.

[0846] Step 17:

[0847] User: Review the provided answers and re-enter any questions or requests as needed.

[0848] (Example 2)

[0849] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0850] Conventional speech recognition systems struggled to accurately analyze user intent and emotional state and provide optimal responses in real time. Furthermore, inconsistent customer service quality among different staff members led to decreased customer satisfaction. This compromised the user experience and hindered the provision of efficient customer service.

[0851] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for pre-processing received audio data, means for sending the pre-processed audio data to a speech recognition engine and converting it into text data, means for extracting the user's intent using a natural language processing engine based on the text data, means for analyzing the user's emotional state from the text data using an emotion analysis engine, means for generating an optimal response based on the intent and emotional state identified by a generation AI model, and means for sending the generated response to a terminal and notifying the user. This makes it possible to accurately grasp the intent and emotion from the user's voice input, improving customer satisfaction and standardizing the quality of service.

[0852] A "user" is the person who performs voice input.

[0853] A "terminal" is a device or system for recording a user's voice as digital data and transmitting it to a server.

[0854] A "server" is a core computer system that performs preprocessing of received audio data, speech recognition, natural language processing, sentiment analysis, and response generation.

[0855] "Audio data" refers to data that digitally represents what a user says to their device.

[0856] "Preprocessing" is the process of applying noise reduction and volume normalization to audio data and converting it into a format suitable for analysis.

[0857] A "speech recognition engine" is software or an algorithm used to convert speech data into text data.

[0858] "Text data" refers to character data converted by a speech recognition engine.

[0859] A "natural language processing engine" is software or an algorithm that analyzes text data to identify the user's intent and nuances.

[0860] An "emotion analysis engine" is software or an algorithm used to analyze a user's emotional state from text data.

[0861] A "generative AI model" is an artificial intelligence model that generates optimal answers based on the results of natural language processing and sentiment analysis.

[0862] "Answer" refers to the appropriate response to a user's inquiry, generated by a generative AI model.

[0863] "Notification" is the act of informing the user of the generated response.

[0864] The "system" is a mechanism that integrates all of these means and executes a series of processes to provide the optimal response from the user's voice input.

[0865] System Overview

[0866] This invention relates to a customer service system that combines speech recognition, natural language processing, and sentiment analysis to analyze the user's intent and emotions in real time from their voice input and provide the optimal response. The system consists of a user, a terminal, and a server.

[0867] User

[0868] Users input questions and requests via voice through the system. For example, if a user says "I want to buy a new smartphone" into the microphone, that voice is sent to the device.

[0869] terminal

[0870] The terminal records the user's voice in digital format and sends it to the server. Upon receiving the audio data, the terminal uses the "pyaudio" library to capture the audio and save it as digital data (for example, with the filename "input.wav"). The recorded digital audio data is transferred to the server in real time. Basic pre-processing such as noise reduction and volume normalization can also be performed.

[0871] server

[0872] The server, as the core of the system, performs the following main functions:

[0873] Audio data preprocessing: The server uses an audio analysis library such as "librosa" to remove noise and normalize the volume of the received audio data, and then converts it into a format suitable for analysis.

[0874] Speech Recognition: Pre-processed audio data is sent to a speech recognition engine (e.g., "Google Cloud Speech-to-Text API") to convert the audio into text data.

[0875] Natural Language Processing and Intent Extraction: Text data obtained through speech recognition is processed by natural language processing engines such as "SpaCy" and "BERT" to analyze the user's intent and nuances. For example, from the text "I want to buy a new smartphone," the intent "a question about purchasing a new smartphone" is identified.

[0876] Sentiment Analysis: Using an emotion analysis engine (e.g., "IBM Watson Tone Analyzer" or "DeepMoji"), the emotional state of the user is analyzed from their text data. The analysis results reveal emotions such as "the user has expectations."

[0877] Response Generation: Based on the identified intent and sentiment analysis results, the system generates the optimal response using a generative AI model (e.g., OpenAI's GPT-3). For example, it might generate a response such as, "The current new models are X, Y, and Z, and their respective features are..."

[0878] Notification: The generated response is sent to the device to notify the user. The device displays or plays the received response in audio format. Audio output tools such as "gTTS" or "Amazon Polly" are used for this purpose.

[0879] Specific example

[0880] Scenario 1: Inquiry about pricing plans

[0881] 1. User: "What is the recommended pricing plan?"

[0882] 2. Terminal: Records audio and sends it to the server.

[0883] 3. Server: Performs noise reduction and volume normalization.

[0884] 4. Server: Converts audio data to text. "What is the recommended pricing plan?"

[0885] 5. Server: Analyzes text data to identify user intent. "Questions about pricing plans"

[0886] 6. Server: The emotion engine analyzes the user's emotional state and determines that they are "interested."

[0887] 7. Server: Based on information about pricing plans, the AI ​​generates the optimal response. "Currently, there are three available pricing plans: A, B, and C. Their features are…"

[0888] 8. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[0889] Scenario 2: Buying a new smartphone

[0890] 1. User: "I want to buy a new smartphone."

[0891] 2. Terminal: Records audio and sends it to the server.

[0892] 3. Server: Performs noise reduction and volume normalization.

[0893] 4. Server: Converts audio data to text. "I want to buy a new smartphone."

[0894] 5. Server: Analyzes text data to identify user intent. "Question regarding the purchase of a new smartphone"

[0895] 6. Server: The emotion engine analyzes the user's emotional state and determines that they have expectations.

[0896] 7. Server: Based on information about smartphones, the AI ​​generates the optimal answer. "The current new models are X, Y, and Z, and their respective features are..."

[0897] 8. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[0898] Example of a prompt

[0899] Let's assume a user inputs "I want to buy a new smartphone" via voice. The system converts the voice to text, identifies information about new smartphones through natural language processing and sentiment analysis, and then generates the best response. Please explain the detailed processing steps for each step of this system.

[0900] This system can accurately grasp the user's intent and emotions from their voice input and provide the most appropriate response. This not only improves customer satisfaction but also ensures consistent service quality.

[0901] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0902] Step 1:

[0903] User: Speaks a question or request into the microphone. Voice input is performed, for example, saying "I want to buy a new smartphone." This voice is the user's input.

[0904] Step 2:

[0905] Terminal: Records the user's voice as digital data and sends it to the server. Specifically, it uses the "pyaudio" library to capture the audio and save it in digital format. The recorded audio data (input.wav) becomes the terminal's output. This digital data is then sent directly to the server.

[0906] Step 3:

[0907] Server: Preprocesses the received audio data. Specifically, it uses the "librosa" library to perform noise reduction and volume normalization. The input for this preprocessing is the recorded audio data (input.wav), and the output is the preprocessed audio data. For example, noise is reduced using librosa.effects.preemphasis(input_signal).

[0908] Step 4:

[0909] Server: Sends pre-processed audio data to the speech recognition engine and converts the audio data into text data. The recognize method of the "Google Cloud Speech-to-Text API" is used. The input to this process is pre-processed audio data, and the output is the text data "I want to buy a new smartphone."

[0910] Step 5:

[0911] Server: The server processes the text data obtained from speech recognition into a natural language processing engine to extract the user's intent. The "SpaCy" library is used for this process. Specifically, the text data is analyzed using NLP (text) to identify the intent as "a question about purchasing a new smartphone." The input to this process is the text data obtained from speech recognition, and the output is user intent data.

[0912] Step 6:

[0913] Server: The server processes the text data obtained through natural language processing into an emotion analysis engine to analyze the user's emotional state. Specifically, it uses "IBM Watson Tone Analyzer" to execute tone_analyzer.tone(tone_input, content_type="application / json").get_result(). The input to this process is text data that identifies intent, and the output is analyzed emotion data (for example, "has expectations").

[0914] Step 7:

[0915] Server: Based on the identified intent and sentiment analysis results, it generates the optimal response using a generative AI model. For example, "GPT-3" is used as the generative AI model. The input to this process is user intent data and sentiment data, and the output is the generated response "The current new model is X, Y, Z, and its respective features are..."

[0916] Step 8:

[0917] Server: The server sends the generated response to the device and notifies the user. HTTP requests are often used as the notification method. For example, requests.post("http: / / device.endpoint", data=response_data) is used. The device displays or plays the received response data in audio format. When playing in audio format, "gTTS" is used, and the command tts = gTTS(response_text, lang='ja') is used. The input to this process is the generated response data, and the output is the information notified to the user.

[0918] (Application Example 2)

[0919] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0920] In brick-and-mortar stores, it is crucial to respond quickly and accurately to customer questions and requests in order to improve the quality and efficiency of customer service. However, conventional customer service systems have difficulty immediately grasping customer intentions and emotions and providing appropriate answers. Furthermore, it is difficult for store staff to respond in real time, which can lead to a decrease in customer satisfaction. This invention aims to solve these problems and provide a means for efficient and effective customer service in brick-and-mortar stores.

[0921] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0922] In this invention, the server includes means for preprocessing audio data, means for converting the preprocessed audio data into text data, means for identifying the customer's intent from the text data, and a device for visually displaying the generated response. This makes it possible to analyze the customer's voice in real time, instantly grasp their intent and emotions, and provide the optimal response.

[0923] "Means of inputting user voice" refers to devices or programs that acquire voice data spoken by the user.

[0924] "Means for transmitting input audio data to a server" refers to devices or programs that transfer audio data acquired from a user to a server via a communication network.

[0925] "Means for preprocessing audio data" refers to devices or programs that perform processing on audio data, such as noise reduction and volume normalization, and convert it into a format that can be analyzed.

[0926] "Means for converting pre-processed audio data into text data" refers to devices or programs that analyze audio data and convert it into corresponding string data.

[0927] "Means for identifying customer intent from text data" refers to devices or programs that analyze text data and identify customer objectives and requirements from it.

[0928] "Means for generating optimal responses based on identified intentions" refers to devices or programs that produce appropriate responses in accordance with the customer's intentions.

[0929] "Means for notifying the user of the generated response" refers to devices or programs that communicate and provide the generated response to the user.

[0930] A "device for visually displaying generated answers" refers to a device that visualizes generated answers and presents them through a display, glasses-type device, or the like.

[0931] System Overview

[0932] This system is a customer service system for brick-and-mortar stores that takes user voice input, analyzes it to identify customer intent and emotions, and provides the most appropriate response. Its main components are the user, a terminal, and a server, which work together to function.

[0933] User

[0934] The user is a store employee wearing smart glasses. When a customer speaks to the employee, the employee receives the voice through the microphone in the smart glasses.

[0935] terminal

[0936] The device is a pair of smart glasses and has the following features:

[0937] Microphone for capturing customer voices

[0938] A display that provides visual information to the store staff.

[0939] Network function for communicating with the server

[0940] server

[0941] The server is the core of this system and performs the following main functions:

[0942] Means for preprocessing audio data

[0943] Means for performing speech recognition

[0944] Means of natural language processing

[0945] Means for performing emotion analysis

[0946] Means of identifying intent

[0947] Means for generating answers

[0948] Means of sending the answer to the device

[0949] Hardware and software to be used

[0950] Hardware: Smart glasses (e.g., Google Glass, Microsoft HoloLens)

[0951] software:

[0952] Speech recognition engine (e.g., Google Cloud Speech-to-Text API)

[0953] Natural language processing engine (e.g., Microsoft Azure Cognitive Services)

[0954] Sentiment analysis engine (e.g., IBM Watson Tone Analyzer)

[0955] Answer generation engine (e.g., OpenAI GPT-4)

[0956] Data processing flow and specific examples

[0957] 1. Voice acquisition:

[0958] A store employee wearing smart glasses uses a microphone to capture the customer's voice. For example, the customer might ask, "What products do you recommend for this season?"

[0959] 2. Sending audio data:

[0960] The acquired audio data is sent from the terminal to the server.

[0961] 3. Pre-treatment:

[0962] The server preprocesses the received audio data, performing noise reduction and volume normalization.

[0963] 4. Speech recognition:

[0964] This process converts pre-processed audio data into text data. In the previous example, the resulting text would be "What products do you recommend for this season?"

[0965] 5. Natural Language Processing:

[0966] The server analyzes the text data to identify the customer's intent. In this case, it is identified as "asking for product recommendations."

[0967] 6. Emotion analysis:

[0968] Customer emotions are analyzed from text and audio data. For example, the emotion of "interest" can be identified.

[0969] 7. Answer generation:

[0970] The server generates the optimal response based on the customer's intent and emotions. For example, it might say, "Our recommended products for this season are A, B, and C. Their features are..."

[0971] 8. Submit and display your response:

[0972] The generated response is displayed on the smart glasses' screen. The store clerk can then refer to the response and answer the customer directly.

[0973] Example of a prompt

[0974] Product: Customer support application

[0975] Model: Smart Glasses

[0976] Objective: To support customer service in physical stores.

[0977] Features: Speech recognition, natural language processing, sentiment analysis, response generation

[0978] Scenario: A customer asks a question about a product, and a store employee uses smart glasses to check the best answer and provide it instantly.

[0979] In this way, by accurately understanding the customer's intent and emotions from their voice input and providing the optimal response, the quality and efficiency of customer service in physical stores can be improved.

[0980] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0981] Step 1:

[0982] The user wears smart glasses and inputs the customer's voice. The smart glasses' microphone picks up the customer's speech and captures it as audio data. The input might be a customer question such as, "What products do you recommend for this season?"

[0983] Step 2:

[0984] The device (smart glasses) preprocesses the captured audio data. This preprocessing involves data manipulation such as noise reduction and volume normalization. As a result, audio data in a format that is easy for the server to analyze is output.

[0985] Step 3:

[0986] Pre-processed audio data is sent from the terminal to the server. In this process, the audio data is transferred to the server in real time via the internet. The input is the pre-processed audio data, and the output is the digital audio data that has reached the server.

[0987] Step 4:

[0988] The server passes the received audio data to a speech recognition engine (e.g., Google Cloud Speech-to-Text API) and converts it into text data. For example, the text "What products do you recommend this season?" is generated.

[0989] Step 5:

[0990] The server processes the generated text data using a natural language processing engine (e.g., Microsoft Azure Cognitive Services) to analyze the customer's intent. This process identifies the intent as "asking for product recommendations." The input is text data, and the output is the identified intent.

[0991] Step 6:

[0992] The server simultaneously uses an emotion analysis engine (e.g., IBM Watson Tone Analyzer) to analyze emotions from text and audio data. For example, the emotion "interested" might be identified. The input is text and audio data, and the output is the identified emotion.

[0993] Step 7:

[0994] The server uses a generative AI model (e.g., OpenAI GPT-4) to generate the most appropriate response based on identified intentions and emotions. For example, it might generate a response like, "Products A, B, and C are recommended for this season. Their characteristics are..." The input is intention and emotion, and the output is the generated response text.

[0995] Step 8:

[0996] The server sends the generated response to the device (smart glasses). In this process, the generated text data is transferred to the device via the internet. The input is the generated response text, and the output is the response data received by the device.

[0997] Step 9:

[0998] The terminal visually displays the generated response on the smart glasses' screen. The store clerk can then explain the product to the customer while looking at the displayed response. For example, the clerk might start by saying, "These are our recommended products for this season." The input is the response data received by the terminal, and the output is the visual information displayed on the screen.

[0999] In this way, a series of processes are realized in which the user receives voice input, the server analyzes it to generate the optimal response, and the terminal visualizes it. This system enables quick and accurate customer service in physical stores.

[1000] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1001] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1002] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[1003] [Third Embodiment]

[1004] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[1005] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1006] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1007] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[1008] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1009] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1010] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1011] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1012] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1013] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1014] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1015] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[1016] System Overview

[1017] This invention relates to a customer service system that utilizes speech recognition and natural language processing technologies. It takes user voice input, analyzes the voice data in real time to identify the customer's intent, and provides the optimal response. The system consists of a user, a terminal, and a server.

[1018] System Configuration

[1019] User

[1020] Users input questions and requests via voice through the system. When a user speaks into the microphone, the voice data is transmitted to the device.

[1021] terminal

[1022] The terminal records the user's voice in digital format and sends it to the server. Some pre-processing is performed as needed before the received audio data is sent to the server.

[1023] server

[1024] The server is the core of this system and performs the following main functions:

[1025] Audio data preprocessing

[1026] Speech recognition

[1027] Natural Language Processing

[1028] Intent Extraction

[1029] Answer generation

[1030] Submit your response

[1031] Processing flow

[1032] Voice input and transmission

[1033] 1. User: Speak your question or request into the microphone.

[1034] 2. Terminal: Records the user's voice as digital data and sends that data to the server.

[1035] Audio data preprocessing

[1036] 3. Server: The server performs preprocessing on the received audio data, such as noise reduction and volume normalization, and converts it into a format suitable for analysis.

[1037] Speech recognition

[1038] 4. Server: Sends the pre-processed audio data to the speech recognition engine and converts the audio data into text data.

[1039] Natural language processing and intent extraction

[1040] 5. Server: The text data generated by speech recognition is fed into a natural language processing engine to analyze the user's nuances and emotions. Next, the user's intent is identified from the text data.

[1041] Response generation and submission

[1042] 6. Server: Utilizes an AI model to generate the optimal response based on the identified intent. The generated response is then formatted as text and audio data.

[1043] 7. Server: Sends the formatted response to the terminal. The terminal displays or plays the received data in audio format.

[1044] Specific example

[1045] Scenario 1: Inquiry about pricing plans

[1046] 1. User: "What is the recommended pricing plan?"

[1047] 2. Terminal: Records audio and sends it to the server.

[1048] 3. Server: Performs noise reduction and volume normalization.

[1049] 4. Server: Converts audio data to text. "What is the recommended pricing plan?"

[1050] 5. Server: Analyzes text data to identify user intent. "Questions about pricing plans"

[1051] 6. Server: Based on information about pricing plans, the AI ​​generates the optimal response. "Currently, there are three available pricing plans: A, B, and C. Their features are..."

[1052] 7. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[1053] Scenario 2: Buying a new smartphone

[1054] 1. User: "I want to buy a new smartphone."

[1055] 2. Terminal: Records audio and sends it to the server.

[1056] 3. Server: Performs noise reduction and volume normalization.

[1057] 4. Server: Converts audio data to text. "I want to buy a new smartphone."

[1058] 5. Server: Analyzes text data to identify user intent. "Question regarding the purchase of a new smartphone"

[1059] 6. Server: Based on information about the smartphone, the AI ​​generates the optimal answer. "The current new models are X, Y, and Z, and their respective features are..."

[1060] 7. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[1061] In this way, the system can accurately understand the user's intent from their voice input and provide the optimal response. This can lead to improved customer satisfaction and consistent service quality.

[1062] The following describes the processing flow.

[1063] Step 1:

[1064] User: Speak your questions or requests into the microphone.

[1065] Step 2:

[1066] Terminal: Records the user's voice as digital data.

[1067] Step 3:

[1068] Terminal: Sends the recorded audio data to the server via an HTTP request.

[1069] Step 4:

[1070] Server: Analyzes the received audio data and applies a noise reduction filter.

[1071] Step 5:

[1072] Server: Normalizes the volume of the data to a certain level.

[1073] Step 6:

[1074] Server: Passes the pre-processed audio data to the speech recognition engine.

[1075] Step 7:

[1076] Server: The speech recognition engine converts the speech data into text data.

[1077] Step 8:

[1078] Server: Transfers the generated text data to the natural language processing (NLP) engine.

[1079] Step 9:

[1080] Server: The NLP engine performs sentiment analysis on text data, analyzing the customer's nuances and intonation.

[1081] Step 10:

[1082] Server: Uses text intent analysis algorithms to identify customer questions and requests from text data.

[1083] Step 11:

[1084] Server: Based on customer intent, the response generation module uses AI to generate the optimal response.

[1085] Step 12:

[1086] Server: Formats the generated responses into text and audio data.

[1087] Step 13:

[1088] Server: Sends the formatted response to the terminal as an HTTP response.

[1089] Step 14:

[1090] Terminal: Displays text data received from the server on the screen.

[1091] Step 15:

[1092] Terminal: Plays received audio data through the speaker.

[1093] Step 16:

[1094] User: Review the provided answers and re-enter any questions or requests as needed.

[1095] (Example 1)

[1096] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1097] Conventional speech recognition systems have limited processes for generating appropriate responses from voice input, making it difficult to accurately identify user intent. Furthermore, the generated responses are often uniform and may not adequately address customer questions or requests. There is a need for solutions to these problems and improve the efficiency and quality of customer service.

[1098] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1099] In this invention, the server includes means for converting pre-processed voice data into text data, means for identifying the customer's intent from the text data, and means for generating the optimal response using a generative AI model. This enables highly accurate intent extraction and flexible response generation from voice input.

[1100] "User voice" refers to the audio signals that a user emits to the system.

[1101] "Input method" refers to any hardware or software component used to collect the user's voice. Examples include microphones and voice acquisition software.

[1102] A "terminal" is an electronic device that receives voice data from a user and transmits it to a server. Specifically, this includes smartphones and computers.

[1103] A "server" is a computer system that performs key processes such as preprocessing of audio data, speech recognition, natural language processing, intent extraction, and response generation.

[1104] "Preprocessing methods" refer to software or algorithms used to remove noise and normalize the volume of audio data, and to convert it into a format suitable for analysis. Examples include "SoX" and "Librosa."

[1105] "Means of converting to text data" refers to speech recognition engines that convert audio data into text data. A specific example is a speech recognition API.

[1106] "Means of identifying customer intent" refers to techniques that use natural language processing engines or pre-trained AI models to extract user intent from text data.

[1107] A "generative AI model" is artificial intelligence that generates the optimal response based on the user's intent. Examples include generative AIs such as GPT-3.

[1108] "Means for generating answers" refers to software or algorithms that use generative AI models to generate answers based on the user's intent.

[1109] "Means of notifying the user of the answer" refers to a system for notifying the user of the generated answer. Specifically, this includes means of sending the answer to the terminal as text data or audio data, and displaying or playing it back as audio.

[1110] "Nuance and emotion" refers to elements that represent the customer's nuances and emotional expressions, extracted from text and audio data.

[1111] "Natural language processing" is a technique for analyzing text data and extracting meaning and intent. Examples include tools and algorithms such as "spaCy" and "BERT."

[1112] Modes for carrying out the invention

[1113] This invention relates to a customer service system that utilizes speech recognition and natural language processing technologies. It takes user voice input, analyzes the voice data in real time to identify the customer's intent, and provides the optimal response. The system consists of a user, a terminal, and a server.

[1114] System Configuration

[1115] User

[1116] Users input questions and requests via voice through the system. When a user speaks into the microphone, the voice data is transmitted to the device.

[1117] terminal

[1118] The terminal records the user's voice in digital format and sends it to the server. Some pre-processing is performed as needed before the received audio data is sent to the server. This pre-processing includes noise reduction and volume normalization. Specifically, "SoX (Sound eXchange)" or "Librosa" is used as the audio collection software on the terminal.

[1119] server

[1120] The server is the core of this system and performs the following main functions:

[1121] Audio data preprocessing (noise reduction and volume normalization)

[1122] Speech recognition (converting speech data to text data)

[1123] Natural language processing (analyzing user nuances and emotions from text data)

[1124] Intent extraction (identifying the user's purpose from text data)

[1125] Answer generation (generates the optimal answer using a generation AI model based on identified intent)

[1126] Submit the response (notify the user of the generated response)

[1127] Hardware and software to be used

[1128] Microphone (hardware): A device for voice input.

[1129] Audio acquisition software for the device: Anvil Studio or Audacity

[1130] Server audio processing software: SoX, Librosa

[1131] Speech recognition engine: Google Cloud Speech-to-Text API

[1132] Natural language processing engines: spaCy, BERT

[1133] Generative AI model: OpenAI GPT-3

[1134] Text-to-speech conversion: Google Cloud Text-to-Speech API

[1135] Specific example

[1136] Scenario 1: Inquiry about pricing plans

[1137] 1. User: "What is the recommended pricing plan?"

[1138] 2. Terminal: Records audio and sends it to the server.

[1139] 3. Server: Performs noise reduction and volume normalization.

[1140] 4. Server: Converts audio data to text. "What is the recommended pricing plan?"

[1141] 5. Server: Analyzes text data to identify user intent. "Questions about pricing plans"

[1142] 6. Server: Based on information about pricing plans, the AI ​​generates the optimal response. "Currently, there are three available pricing plans: A, B, and C. Their features are..."

[1143] 7. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[1144] Scenario 2: Buying a new smartphone

[1145] 1. User: "I want to buy a new smartphone."

[1146] 2. Terminal: Records audio and sends it to the server.

[1147] 3. Server: Performs noise reduction and volume normalization.

[1148] 4. Server: Converts audio data to text. "I want to buy a new smartphone."

[1149] 5. Server: Analyzes text data to identify user intent. "Question regarding the purchase of a new smartphone"

[1150] 6. Server: Based on information about the smartphone, the AI ​​generates the optimal answer. "The current new models are X, Y, and Z, and their respective features are..."

[1151] 7. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[1152] Examples of prompt messages include "What is your recommended pricing plan?" and "I'd like to buy a new smartphone."

[1153] In this way, the system can accurately understand the user's intent from their voice input and provide the optimal response. This can lead to improved customer satisfaction and consistent service quality.

[1154] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1155] Step 1:

[1156] User: Speaks a question or request into the microphone. For example, "What is the recommended pricing plan?" This generates audio data.

[1157] Step 2:

[1158] Terminal: Records the user's voice in digital format (e.g., WAV file format). Using voice collection software on the terminal (e.g., Anvil Studio, Audacity), the voice data is saved as digital data. Once recording is complete, the voice data is sent to the server using an HTTP request. To ensure data security, the TLS (Transport Layer Security) encryption protocol is used.

[1159] Step 3:

[1160] Server: Receives audio data and performs preprocessing such as noise reduction and volume normalization. Specifically, it uses libraries such as "SoX (Sound eXchange)" and Python's "Librosa". This process converts the audio data into a format suitable for analysis. For example, noise reduction reduces background noise and improves speech clarity.

[1161] Step 4:

[1162] Server: The server sends the pre-processed audio data to the speech recognition engine and converts the audio data into text data. The Google Cloud Speech-to-Text API is used here. The API is used to generate text data from the audio data. For example, the audio data "What is the best pricing plan?" is converted to the text data "What is the best pricing plan?".

[1163] Step 5:

[1164] Server: Text data generated by speech recognition is processed by a natural language processing engine (e.g., spaCy) to analyze the user's nuances and emotions. Simultaneously, a pre-trained AI model (e.g., BERT) is used to identify the user's intent from the text data. This process extracts meaning and intent from the text data. For example, "a question about pricing plans" might be extracted as the user's intent.

[1165] Step 6:

[1166] Server: Uses a generative AI model (e.g., OpenAI GPT-3) to generate the best answer based on the identified intent. The generative AI model is given a prompt and generates an appropriate answer. For example, based on the intent "Question about pricing plans," it generates an answer such as "Currently available pricing plans are A, B, and C. Their features are..."

[1167] Step 7:

[1168] Server: Formats the generated responses as text and audio data and sends them to the device. The Google Cloud Text-to-Speech API is used to generate the audio data. The device displays or plays the received data audibly to the user, allowing the user to visually or audibly confirm the responses.

[1169] For example, the generated response, "Currently available pricing plans are A, B, and C. Their features are...", is sent to the device, and the device plays the response aloud, allowing the user to hear the content. In this way, the system accurately analyzes the user's intent through concrete actions and provides the most appropriate response.

[1170] (Application Example 1)

[1171] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1172] Traditional customer service systems had a problem in that they made it difficult for users to get quick and accurate answers when interacting directly with staff in physical stores. This could lead to decreased customer satisfaction and increased workload for store staff, resulting in inconsistent service quality.

[1173] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1174] In this invention, the server includes means for transmitting the user's voice data to a smart device, means for displaying or playing back the answer to the user's question via the smart device, and means for analyzing nuances and intonation. This makes it possible to respond quickly and accurately to user questions in physical stores.

[1175] "Means for inputting user voice" refers to a mechanism for recording and acquiring user voice information using a digital device.

[1176] "Means for sending input audio data to a server" refers to communication means for sending audio data captured by a device to a server via a network.

[1177] "Methods for pre-processing audio data" refer to methods for processing captured audio data, such as noise reduction and volume normalization, to prepare it for easier analysis.

[1178] "Means for converting pre-processed audio data into text data" refers to a mechanism that utilizes speech recognition technology to convert pre-processed audio data into text information.

[1179] "Methods for identifying customer intent from text data" refer to methods that use natural language processing technology to analyze text data generated by speech recognition and identify the user's intent and purpose.

[1180] "Means for generating optimal responses based on identified intentions" refers to methods that use AI models to create optimal responses based on relevant information, taking into account the identified user intentions.

[1181] "Means of notifying the user of the generated response" refers to the means by which the server communicates the response it has created to the user in text or audio format.

[1182] "Means for transmitting user voice data to a smart device" refers to communication means for transmitting recorded user voice data to devices such as smartphones and smart glasses.

[1183] "Means for displaying or playing back the user's question via a smart device" refers to means such as a display device for visualizing the received answer on the smart device or a speaker for playing it back as sound.

[1184] System Overview

[1185] This invention is a customer service system that utilizes speech recognition and natural language processing technologies. The system takes voice input from the user, analyzes the voice data in real time to identify the customer's intent, and provides the optimal response. The system consists of a user, a terminal, and a server.

[1186] System Configuration

[1187] User

[1188] Users input questions and requests via voice through the system. When a user speaks into the microphone, the voice data is transmitted to their smart device.

[1189] terminal

[1190] The terminal consists of smart devices such as smartphones and smart glasses, which digitally record the user's voice and send the data to a server. Some pre-processing is performed on the received voice data before it is sent to the server.

[1191] server

[1192] The server is the core of this system and performs the following main functions:

[1193] Audio data preprocessing

[1194] Speech recognition

[1195] Natural Language Processing

[1196] Intent Extraction

[1197] Answer generation

[1198] Submit your response

[1199] Hardware and software to be used

[1200] Hardware:

[1201] Smartphones and smart glasses: Record user voice and communicate with the server.

[1202] Server: Performs data preprocessing, speech recognition, natural language processing, intent extraction, and response generation.

[1203] software:

[1204] Speech recognition engine (e.g., Google Cloud Speech-to-Text)

[1205] Natural language processing engines (e.g., SpaCy, BERT)

[1206] AI models (e.g., GPT-4)

[1207] Data processing and data calculation

[1208] Audio data preprocessing: The server performs preprocessing on the received audio data, such as noise reduction and volume normalization, and converts it into a format suitable for analysis.

[1209] Speech recognition: Pre-processed audio data is sent to the speech recognition engine, which converts the audio data into text data.

[1210] Natural Language Processing and Intent Extraction: Text data generated by speech recognition is processed by a natural language processing engine to analyze the user's nuances and intonation. Next, the user's intent is identified from the text data.

[1211] Response Generation: An AI model is used to generate the optimal response based on the identified intent. The generated response is formatted as text and audio data and sent to the device.

[1212] Specific example

[1213] Check product availability

[1214] 1. User: "Do you have this shirt in stock?"

[1215] 2. Terminal: Records audio and sends it to the server.

[1216] 3. Server: Performs noise reduction and volume normalization, and converts the audio data into text.

[1217] 4. Server: Identify user intent from text data.

[1218] 5. Server: Based on product inventory information, the AI ​​generates the optimal response. "This shirt is currently only available in size M. Other sizes are expected to arrive within a week."

[1219] 6. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[1220] Prompt example

[1221] "Do you have this shirt in stock?"

[1222] → AI model: Generates answers based on inventory information.

[1223] Example output: "This shirt is currently only available in size M. Other sizes will be available within a week."

[1224] In this way, customer service support tools utilizing speech recognition and natural language processing technologies can be used in physical stores. This allows store staff to respond to customer inquiries quickly and accurately, which is expected to improve customer satisfaction.

[1225] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1226] Step 1:

[1227] User voice input

[1228] The user speaks questions or requests into the microphone. The input at this time is audio data.

[1229] Step 2:

[1230] Recording and transmitting audio data

[1231] The terminal records the user's voice in digital format and sends that data to the server. The input here is the recorded voice data, and the output is the digital voice data sent to the server.

[1232] Step 3:

[1233] Audio data preprocessing

[1234] The server performs preprocessing on the received audio data, such as noise reduction and volume normalization, and converts it into a format suitable for analysis. The input is digital audio data transmitted from the terminal, and the output is the preprocessed audio data.

[1235] Step 4:

[1236] Speech recognition

[1237] The server sends the pre-processed audio data to a speech recognition engine (e.g., Google Cloud Speech-to-Text) to convert the audio data into text data. The input is pre-processed audio data, and the output is text data.

[1238] Step 5:

[1239] Natural language processing and intent extraction

[1240] The server processes the text data generated by speech recognition into a natural language processing engine (e.g., SpaCy, BERT) to analyze the user's nuances and intonation. Next, it identifies the user's intent from the analyzed text data. The input is text data, and the output is the identified user intent data.

[1241] Step 6:

[1242] Answer generation

[1243] The server utilizes an AI model (e.g., GPT-4) to generate the optimal response based on the identified intent. The generated response is then formatted as text and audio data. The input is the user's intent data, and the output is the generated response data.

[1244] Step 7:

[1245] Submit your response

[1246] The server sends the generated response data to the terminal. The input is the generated response data, and the output is the response data sent to the terminal.

[1247] Step 8:

[1248] Display or play the answer

[1249] The terminal displays or plays back the received response data to the user in audio format. The input is the response data sent from the server, and the output is the display or audio playback to the user.

[1250] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1251] System Overview

[1252] This invention relates to a customer service system that combines speech recognition, natural language processing (NLP), and sentiment analysis. It takes user voice input, analyzes the voice data in real time to identify the customer's intentions and emotions, and provides the most appropriate response. The system consists of a user, a terminal, and a server.

[1253] System Configuration

[1254] User

[1255] Users input questions and requests via voice through the system. When a user speaks into the microphone, the voice data is transmitted to the device.

[1256] terminal

[1257] The terminal records the user's voice in digital format and sends it to the server. Some pre-processing is performed as needed before the received audio data is sent to the server.

[1258] server

[1259] The server is the core of this system and performs the following main functions:

[1260] Means for preprocessing audio data

[1261] Means for performing speech recognition

[1262] Means of natural language processing

[1263] A method for performing emotion analysis using an emotion engine.

[1264] Means for extracting intent

[1265] Means for generating answers

[1266] Method for submitting a response

[1267] Processing flow

[1268] Voice input and transmission

[1269] 1. User: Speak your question or request into the microphone.

[1270] 2. Terminal: Records the user's voice as digital data and sends that data to the server.

[1271] Audio data preprocessing

[1272] 3. Server: The server performs preprocessing on the received audio data, such as noise reduction and volume normalization, and converts it into a format suitable for analysis.

[1273] Speech recognition

[1274] 4. Server: Sends the pre-processed audio data to the speech recognition engine and converts the audio data into text data.

[1275] Natural language processing and intent extraction

[1276] 5. Server: The text data generated by speech recognition is fed into a natural language processing engine to analyze the user's nuances and emotions. Next, the user's intent is identified from the text data.

[1277] Emotion analysis

[1278] 6. Server: Uses an emotion engine to analyze emotions from user voice and text data. The analysis results are used to identify customer intentions.

[1279] Response generation and submission

[1280] 7. Server: Utilizes an AI model to generate the optimal response based on the identified intent and sentiment analysis results. The generated response is then formatted as text and audio data.

[1281] 8. Server: Sends the formatted response to the terminal. The terminal displays or plays the received data in audio format.

[1282] Specific example

[1283] Scenario 1: Inquiry about pricing plans

[1284] 1. User: "What is the recommended pricing plan?"

[1285] 2. Terminal: Records audio and sends it to the server.

[1286] 3. Server: Performs noise reduction and volume normalization.

[1287] 4. Server: Converts audio data to text. "What is the recommended pricing plan?"

[1288] 5. Server: Analyzes text data to identify user intent. "Questions about pricing plans"

[1289] 6. Server: The emotion engine analyzes the user's emotional state and determines that they are "interested."

[1290] 7. Server: Based on information about pricing plans, the AI ​​generates the optimal response. "Currently, there are three available pricing plans: A, B, and C. Their features are…"

[1291] 8. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[1292] Scenario 2: Buying a new smartphone

[1293] 1. User: "I want to buy a new smartphone."

[1294] 2. Terminal: Records audio and sends it to the server.

[1295] 3. Server: Performs noise reduction and volume normalization.

[1296] 4. Server: Converts audio data to text. "I want to buy a new smartphone."

[1297] 5. Server: Analyzes text data to identify user intent. "Question regarding the purchase of a new smartphone"

[1298] 6. Server: The emotion engine analyzes the user's emotional state and determines that they have expectations.

[1299] 7. Server: Based on information about smartphones, the AI ​​generates the optimal answer. "The current new models are X, Y, and Z, and their respective features are..."

[1300] 8. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[1301] This system can accurately grasp the user's intent and emotions from their voice input and provide the most appropriate response. This not only improves customer satisfaction but also ensures consistent service quality.

[1302] The following describes the processing flow.

[1303] Step 1:

[1304] User: Speak your questions or requests into the microphone.

[1305] Step 2:

[1306] Terminal: Records the user's voice as digital data.

[1307] Step 3:

[1308] Terminal: Sends the recorded audio data to the server via an HTTP request.

[1309] Step 4:

[1310] Server: Analyzes the received audio data and applies a noise reduction filter.

[1311] Step 5:

[1312] Server: Normalizes the volume of the data to a certain level.

[1313] Step 6:

[1314] Server: Passes the pre-processed audio data to the speech recognition engine.

[1315] Step 7:

[1316] Server: The speech recognition engine converts the speech data into text data.

[1317] Step 8:

[1318] Server: Transfers the generated text data to the natural language processing (NLP) engine.

[1319] Step 9:

[1320] Server: The NLP engine performs sentiment analysis on text data, analyzing the customer's nuances and intonation.

[1321] Step 10:

[1322] Server: Uses an emotion engine to analyze the user's emotional state from text and audio data.

[1323] Step 11:

[1324] Server: Uses text intent analysis algorithms to identify customer questions and requests from text data.

[1325] Step 12:

[1326] Server: Based on sentiment analysis and intent analysis results, it utilizes AI models to generate appropriate answers to questions and requests.

[1327] Step 13:

[1328] Server: Adjusts the generated responses according to the user's emotional state and formats them as text and audio data.

[1329] Step 14:

[1330] Server: Sends the formatted response to the terminal as an HTTP response.

[1331] Step 15:

[1332] Terminal: Displays text data received from the server on the screen.

[1333] Step 16:

[1334] Terminal: Plays received audio data through the speaker.

[1335] Step 17:

[1336] User: Review the provided answers and re-enter any questions or requests as needed.

[1337] (Example 2)

[1338] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1339] Conventional speech recognition systems struggled to accurately analyze user intent and emotional state and provide optimal responses in real time. Furthermore, inconsistent customer service quality among different staff members led to decreased customer satisfaction. This compromised the user experience and hindered the provision of efficient customer service.

[1340] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for pre-processing received audio data, means for sending the pre-processed audio data to a speech recognition engine and converting it into text data, means for extracting the user's intent using a natural language processing engine based on the text data, means for analyzing the user's emotional state from the text data using an emotion analysis engine, means for generating an optimal response based on the intent and emotional state identified by a generation AI model, and means for sending the generated response to a terminal and notifying the user. This makes it possible to accurately grasp the intent and emotion from the user's voice input, improving customer satisfaction and standardizing the quality of service.

[1341] A "user" is the person who performs voice input.

[1342] A "terminal" is a device or system for recording a user's voice as digital data and transmitting it to a server.

[1343] A "server" is a core computer system that performs preprocessing of received audio data, speech recognition, natural language processing, sentiment analysis, and response generation.

[1344] "Audio data" refers to data that digitally represents what a user says to their device.

[1345] "Preprocessing" is the process of applying noise reduction and volume normalization to audio data and converting it into a format suitable for analysis.

[1346] A "speech recognition engine" is software or an algorithm used to convert speech data into text data.

[1347] "Text data" refers to character data converted by a speech recognition engine.

[1348] A "natural language processing engine" is software or an algorithm that analyzes text data to identify the user's intent and nuances.

[1349] An "emotion analysis engine" is software or an algorithm used to analyze a user's emotional state from text data.

[1350] A "generative AI model" is an artificial intelligence model that generates optimal answers based on the results of natural language processing and sentiment analysis.

[1351] "Answer" refers to the appropriate response to a user's inquiry, generated by a generative AI model.

[1352] "Notification" is the act of informing the user of the generated response.

[1353] The "system" is a mechanism that integrates all of these means and executes a series of processes to provide the optimal response from the user's voice input.

[1354] System Overview

[1355] This invention relates to a customer service system that combines speech recognition, natural language processing, and sentiment analysis to analyze the user's intent and emotions in real time from their voice input and provide the optimal response. The system consists of a user, a terminal, and a server.

[1356] User

[1357] Users input questions and requests via voice through the system. For example, if a user says "I want to buy a new smartphone" into the microphone, that voice is sent to the device.

[1358] terminal

[1359] The terminal records the user's voice in digital format and sends it to the server. Upon receiving the audio data, the terminal uses the "pyaudio" library to capture the audio and save it as digital data (for example, with the filename "input.wav"). The recorded digital audio data is transferred to the server in real time. Basic pre-processing such as noise reduction and volume normalization can also be performed.

[1360] server

[1361] The server, as the core of the system, performs the following main functions:

[1362] Audio data preprocessing: The server uses an audio analysis library such as "librosa" to remove noise and normalize the volume of the received audio data, and then converts it into a format suitable for analysis.

[1363] Speech Recognition: Pre-processed audio data is sent to a speech recognition engine (e.g., "Google Cloud Speech-to-Text API") to convert the audio into text data.

[1364] Natural Language Processing and Intent Extraction: Text data obtained through speech recognition is processed by natural language processing engines such as "SpaCy" and "BERT" to analyze the user's intent and nuances. For example, from the text "I want to buy a new smartphone," the intent "a question about purchasing a new smartphone" is identified.

[1365] Sentiment Analysis: Using an emotion analysis engine (e.g., "IBM Watson Tone Analyzer" or "DeepMoji"), the emotional state of the user is analyzed from their text data. The analysis results reveal emotions such as "the user has expectations."

[1366] Response Generation: Based on the identified intent and sentiment analysis results, the system generates the optimal response using a generative AI model (e.g., OpenAI's GPT-3). For example, it might generate a response such as, "The current new models are X, Y, and Z, and their respective features are..."

[1367] Notification: The generated response is sent to the device to notify the user. The device displays or plays the received response in audio format. Audio output tools such as "gTTS" or "Amazon Polly" are used for this purpose.

[1368] Specific example

[1369] Scenario 1: Inquiry about pricing plans

[1370] 1. User: "What is the recommended pricing plan?"

[1371] 2. Terminal: Records audio and sends it to the server.

[1372] 3. Server: Performs noise reduction and volume normalization.

[1373] 4. Server: Converts audio data to text. "What is the recommended pricing plan?"

[1374] 5. Server: Analyzes text data to identify user intent. "Questions about pricing plans"

[1375] 6. Server: The emotion engine analyzes the user's emotional state and determines that they are "interested."

[1376] 7. Server: Based on information about pricing plans, the AI ​​generates the optimal response. "Currently, there are three available pricing plans: A, B, and C. Their features are…"

[1377] 8. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[1378] Scenario 2: Buying a new smartphone

[1379] 1. User: "I want to buy a new smartphone."

[1380] 2. Terminal: Records audio and sends it to the server.

[1381] 3. Server: Performs noise reduction and volume normalization.

[1382] 4. Server: Converts audio data to text. "I want to buy a new smartphone."

[1383] 5. Server: Analyzes text data to identify user intent. "Question regarding the purchase of a new smartphone"

[1384] 6. Server: The emotion engine analyzes the user's emotional state and determines that they have expectations.

[1385] 7. Server: Based on information about smartphones, the AI ​​generates the optimal answer. "The current new models are X, Y, and Z, and their respective features are..."

[1386] 8. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[1387] Example of a prompt

[1388] Let's assume a user inputs "I want to buy a new smartphone" via voice. The system converts the voice to text, identifies information about new smartphones through natural language processing and sentiment analysis, and then generates the best response. Please explain the detailed processing steps for each step of this system.

[1389] This system can accurately grasp the user's intent and emotions from their voice input and provide the most appropriate response. This not only improves customer satisfaction but also ensures consistent service quality.

[1390] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1391] Step 1:

[1392] User: Speaks a question or request into the microphone. Voice input is performed, for example, saying "I want to buy a new smartphone." This voice is the user's input.

[1393] Step 2:

[1394] Terminal: Records the user's voice as digital data and sends it to the server. Specifically, it uses the "pyaudio" library to capture the audio and save it in digital format. The recorded audio data (input.wav) becomes the terminal's output. This digital data is then sent directly to the server.

[1395] Step 3:

[1396] Server: Preprocesses the received audio data. Specifically, it uses the "librosa" library to perform noise reduction and volume normalization. The input for this preprocessing is the recorded audio data (input.wav), and the output is the preprocessed audio data. For example, noise is reduced using librosa.effects.preemphasis(input_signal).

[1397] Step 4:

[1398] Server: Sends pre-processed audio data to the speech recognition engine and converts the audio data into text data. The recognize method of the "Google Cloud Speech-to-Text API" is used. The input to this process is pre-processed audio data, and the output is the text data "I want to buy a new smartphone."

[1399] Step 5:

[1400] Server: The server processes the text data obtained from speech recognition into a natural language processing engine to extract the user's intent. The "SpaCy" library is used for this process. Specifically, the text data is analyzed using NLP (text) to identify the intent as "a question about purchasing a new smartphone." The input to this process is the text data obtained from speech recognition, and the output is user intent data.

[1401] Step 6:

[1402] Server: The server processes the text data obtained through natural language processing into an emotion analysis engine to analyze the user's emotional state. Specifically, it uses "IBM Watson Tone Analyzer" to execute tone_analyzer.tone(tone_input, content_type="application / json").get_result(). The input to this process is text data that identifies intent, and the output is analyzed emotion data (for example, "has expectations").

[1403] Step 7:

[1404] Server: Based on the identified intent and sentiment analysis results, it generates the optimal response using a generative AI model. For example, "GPT-3" is used as the generative AI model. The input to this process is user intent data and sentiment data, and the output is the generated response "The current new model is X, Y, Z, and its respective features are..."

[1405] Step 8:

[1406] Server: The server sends the generated response to the device and notifies the user. HTTP requests are often used as the notification method. For example, requests.post("http: / / device.endpoint", data=response_data) is used. The device displays or plays the received response data in audio format. When playing in audio format, "gTTS" is used, and the command tts = gTTS(response_text, lang='ja') is used. The input to this process is the generated response data, and the output is the information notified to the user.

[1407] (Application Example 2)

[1408] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1409] In brick-and-mortar stores, it is crucial to respond quickly and accurately to customer questions and requests in order to improve the quality and efficiency of customer service. However, conventional customer service systems have difficulty immediately grasping customer intentions and emotions and providing appropriate answers. Furthermore, it is difficult for store staff to respond in real time, which can lead to a decrease in customer satisfaction. This invention aims to solve these problems and provide a means for efficient and effective customer service in brick-and-mortar stores.

[1410] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[1411] In this invention, the server includes means for preprocessing audio data, means for converting the preprocessed audio data into text data, means for identifying the customer's intent from the text data, and a device for visually displaying the generated response. This makes it possible to analyze the customer's voice in real time, instantly grasp their intent and emotions, and provide the optimal response.

[1412] "Means of inputting user voice" refers to devices or programs that acquire voice data spoken by the user.

[1413] "Means for transmitting input audio data to a server" refers to devices or programs that transfer audio data acquired from a user to a server via a communication network.

[1414] "Means for preprocessing audio data" refers to devices or programs that perform processing on audio data, such as noise reduction and volume normalization, and convert it into a format that can be analyzed.

[1415] "Means for converting pre-processed audio data into text data" refers to devices or programs that analyze audio data and convert it into corresponding string data.

[1416] "Means for identifying customer intent from text data" refers to devices or programs that analyze text data and identify customer objectives and requirements from it.

[1417] "Means for generating optimal responses based on identified intentions" refers to devices or programs that produce appropriate responses in accordance with the customer's intentions.

[1418] "Means for notifying the user of the generated response" refers to devices or programs that communicate and provide the generated response to the user.

[1419] A "device for visually displaying generated answers" refers to a device that visualizes generated answers and presents them through a display, glasses-type device, or the like.

[1420] System Overview

[1421] This system is a customer service system for brick-and-mortar stores that takes user voice input, analyzes it to identify customer intent and emotions, and provides the most appropriate response. Its main components are the user, a terminal, and a server, which work together to function.

[1422] User

[1423] The user is a store employee wearing smart glasses. When a customer speaks to the employee, the employee receives the voice through the microphone in the smart glasses.

[1424] terminal

[1425] The device is a pair of smart glasses and has the following features:

[1426] Microphone for capturing customer voices

[1427] A display that provides visual information to the store staff.

[1428] Network function for communicating with the server

[1429] server

[1430] The server is the core of this system and performs the following main functions:

[1431] Means for preprocessing audio data

[1432] Means for performing speech recognition

[1433] Means of natural language processing

[1434] Means for performing emotion analysis

[1435] Means of identifying intent

[1436] Means for generating answers

[1437] Means of sending the answer to the device

[1438] Hardware and software to be used

[1439] Hardware: Smart glasses (e.g., Google Glass, Microsoft HoloLens)

[1440] software:

[1441] Speech recognition engine (e.g., Google Cloud Speech-to-Text API)

[1442] Natural language processing engine (e.g., Microsoft Azure Cognitive Services)

[1443] Sentiment analysis engine (e.g., IBM Watson Tone Analyzer)

[1444] Answer generation engine (e.g., OpenAI GPT-4)

[1445] Data processing flow and specific examples

[1446] 1. Voice acquisition:

[1447] A store employee wearing smart glasses uses a microphone to capture the customer's voice. For example, the customer might ask, "What products do you recommend for this season?"

[1448] 2. Sending audio data:

[1449] The acquired audio data is sent from the terminal to the server.

[1450] 3. Pre-treatment:

[1451] The server preprocesses the received audio data, performing noise reduction and volume normalization.

[1452] 4. Speech recognition:

[1453] This process converts pre-processed audio data into text data. In the previous example, the resulting text would be "What products do you recommend for this season?"

[1454] 5. Natural Language Processing:

[1455] The server analyzes the text data to identify the customer's intent. In this case, it is identified as "asking for product recommendations."

[1456] 6. Emotion analysis:

[1457] Customer emotions are analyzed from text and audio data. For example, the emotion of "interest" can be identified.

[1458] 7. Answer generation:

[1459] The server generates the optimal response based on the customer's intent and emotions. For example, it might say, "Our recommended products for this season are A, B, and C. Their features are..."

[1460] 8. Submit and display your response:

[1461] The generated response is displayed on the smart glasses' screen. The store clerk can then refer to the response and answer the customer directly.

[1462] Example of a prompt

[1463] Product: Customer support application

[1464] Model: Smart Glasses

[1465] Objective: To support customer service in physical stores.

[1466] Features: Speech recognition, natural language processing, sentiment analysis, response generation

[1467] Scenario: A customer asks a question about a product, and a store employee uses smart glasses to check the best answer and provide it instantly.

[1468] In this way, by accurately understanding the customer's intent and emotions from their voice input and providing the optimal response, the quality and efficiency of customer service in physical stores can be improved.

[1469] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1470] Step 1:

[1471] The user wears smart glasses and inputs the customer's voice. The smart glasses' microphone picks up the customer's speech and captures it as audio data. The input might be a customer question such as, "What products do you recommend for this season?"

[1472] Step 2:

[1473] The device (smart glasses) preprocesses the captured audio data. This preprocessing involves data manipulation such as noise reduction and volume normalization. As a result, audio data in a format that is easy for the server to analyze is output.

[1474] Step 3:

[1475] Pre-processed audio data is sent from the terminal to the server. In this process, the audio data is transferred to the server in real time via the internet. The input is the pre-processed audio data, and the output is the digital audio data that has reached the server.

[1476] Step 4:

[1477] The server passes the received audio data to a speech recognition engine (e.g., Google Cloud Speech-to-Text API) and converts it into text data. For example, the text "What products do you recommend this season?" is generated.

[1478] Step 5:

[1479] The server processes the generated text data using a natural language processing engine (e.g., Microsoft Azure Cognitive Services) to analyze the customer's intent. This process identifies the intent as "asking for product recommendations." The input is text data, and the output is the identified intent.

[1480] Step 6:

[1481] The server simultaneously uses an emotion analysis engine (e.g., IBM Watson Tone Analyzer) to analyze emotions from text and audio data. For example, the emotion "interested" might be identified. The input is text and audio data, and the output is the identified emotion.

[1482] Step 7:

[1483] The server uses a generative AI model (e.g., OpenAI GPT-4) to generate the most appropriate response based on identified intentions and emotions. For example, it might generate a response like, "Products A, B, and C are recommended for this season. Their characteristics are..." The input is intention and emotion, and the output is the generated response text.

[1484] Step 8:

[1485] The server sends the generated response to the device (smart glasses). In this process, the generated text data is transferred to the device via the internet. The input is the generated response text, and the output is the response data received by the device.

[1486] Step 9:

[1487] The terminal visually displays the generated response on the smart glasses' screen. The store clerk can then explain the product to the customer while looking at the displayed response. For example, the clerk might start by saying, "These are our recommended products for this season." The input is the response data received by the terminal, and the output is the visual information displayed on the screen.

[1488] In this way, a series of processes are realized in which the user receives voice input, the server analyzes it to generate the optimal response, and the terminal visualizes it. This system enables quick and accurate customer service in physical stores.

[1489] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1490] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1491] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1492] [Fourth Embodiment]

[1493] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1494] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1495] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1496] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1497] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1498] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1499] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1500] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1501] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1502] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1503] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1504] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1505] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1506] System Overview

[1507] This invention relates to a customer service system that utilizes speech recognition and natural language processing technologies. It takes user voice input, analyzes the voice data in real time to identify the customer's intent, and provides the optimal response. The system consists of a user, a terminal, and a server.

[1508] System Configuration

[1509] User

[1510] Users input questions and requests via voice through the system. When a user speaks into the microphone, the voice data is transmitted to the device.

[1511] terminal

[1512] The terminal records the user's voice in digital format and sends it to the server. Some pre-processing is performed as needed before the received audio data is sent to the server.

[1513] server

[1514] The server is the core of this system and performs the following main functions:

[1515] Audio data preprocessing

[1516] Speech recognition

[1517] Natural Language Processing

[1518] Intent Extraction

[1519] Answer generation

[1520] Submit your response

[1521] Processing flow

[1522] Voice input and transmission

[1523] 1. User: Speak your question or request into the microphone.

[1524] 2. Terminal: Records the user's voice as digital data and sends that data to the server.

[1525] Audio data preprocessing

[1526] 3. Server: The server performs preprocessing on the received audio data, such as noise reduction and volume normalization, and converts it into a format suitable for analysis.

[1527] Speech recognition

[1528] 4. Server: Sends the pre-processed audio data to the speech recognition engine and converts the audio data into text data.

[1529] Natural language processing and intent extraction

[1530] 5. Server: The text data generated by speech recognition is fed into a natural language processing engine to analyze the user's nuances and emotions. Next, the user's intent is identified from the text data.

[1531] Response generation and submission

[1532] 6. Server: Utilizes an AI model to generate the optimal response based on the identified intent. The generated response is then formatted as text and audio data.

[1533] 7. Server: Sends the formatted response to the terminal. The terminal displays or plays the received data in audio format.

[1534] Specific example

[1535] Scenario 1: Inquiry about pricing plans

[1536] 1. User: "What is the recommended pricing plan?"

[1537] 2. Terminal: Records audio and sends it to the server.

[1538] 3. Server: Performs noise reduction and volume normalization.

[1539] 4. Server: Converts audio data to text. "What is the recommended pricing plan?"

[1540] 5. Server: Analyzes text data to identify user intent. "Questions about pricing plans"

[1541] 6. Server: Based on information about pricing plans, the AI ​​generates the optimal response. "Currently, there are three available pricing plans: A, B, and C. Their features are..."

[1542] 7. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[1543] Scenario 2: Buying a new smartphone

[1544] 1. User: "I want to buy a new smartphone."

[1545] 2. Terminal: Records audio and sends it to the server.

[1546] 3. Server: Performs noise reduction and volume normalization.

[1547] 4. Server: Converts audio data to text. "I want to buy a new smartphone."

[1548] 5. Server: Analyzes text data to identify user intent. "Question regarding the purchase of a new smartphone"

[1549] 6. Server: Based on information about the smartphone, the AI ​​generates the optimal answer. "The current new models are X, Y, and Z, and their respective features are..."

[1550] 7. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[1551] In this way, the system can accurately understand the user's intent from their voice input and provide the optimal response. This can lead to improved customer satisfaction and consistent service quality.

[1552] The following describes the processing flow.

[1553] Step 1:

[1554] User: Speak your questions or requests into the microphone.

[1555] Step 2:

[1556] Terminal: Records the user's voice as digital data.

[1557] Step 3:

[1558] Terminal: Sends the recorded audio data to the server via an HTTP request.

[1559] Step 4:

[1560] Server: Analyzes the received audio data and applies a noise reduction filter.

[1561] Step 5:

[1562] Server: Normalizes the volume of the data to a certain level.

[1563] Step 6:

[1564] Server: Passes the pre-processed audio data to the speech recognition engine.

[1565] Step 7:

[1566] Server: The speech recognition engine converts the speech data into text data.

[1567] Step 8:

[1568] Server: Transfers the generated text data to the natural language processing (NLP) engine.

[1569] Step 9:

[1570] Server: The NLP engine performs sentiment analysis on text data, analyzing the customer's nuances and intonation.

[1571] Step 10:

[1572] Server: Uses text intent analysis algorithms to identify customer questions and requests from text data.

[1573] Step 11:

[1574] Server: Based on customer intent, the response generation module uses AI to generate the optimal response.

[1575] Step 12:

[1576] Server: Formats the generated responses into text and audio data.

[1577] Step 13:

[1578] Server: Sends the formatted response to the terminal as an HTTP response.

[1579] Step 14:

[1580] Terminal: Displays text data received from the server on the screen.

[1581] Step 15:

[1582] Terminal: Plays received audio data through the speaker.

[1583] Step 16:

[1584] User: Review the provided answers and re-enter any questions or requests as needed.

[1585] (Example 1)

[1586] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1587] Conventional speech recognition systems have limited processes for generating appropriate responses from voice input, making it difficult to accurately identify user intent. Furthermore, the generated responses are often uniform and may not adequately address customer questions or requests. There is a need for solutions to these problems and improve the efficiency and quality of customer service.

[1588] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1589] In this invention, the server includes means for converting pre-processed voice data into text data, means for identifying the customer's intent from the text data, and means for generating the optimal response using a generative AI model. This enables highly accurate intent extraction and flexible response generation from voice input.

[1590] "User voice" refers to the audio signals that a user emits to the system.

[1591] "Input method" refers to any hardware or software component used to collect the user's voice. Examples include microphones and voice acquisition software.

[1592] A "terminal" is an electronic device that receives voice data from a user and transmits it to a server. Specifically, this includes smartphones and computers.

[1593] A "server" is a computer system that performs key processes such as preprocessing of audio data, speech recognition, natural language processing, intent extraction, and response generation.

[1594] "Preprocessing methods" refer to software or algorithms used to remove noise and normalize the volume of audio data, and to convert it into a format suitable for analysis. Examples include "SoX" and "Librosa."

[1595] "Means of converting to text data" refers to speech recognition engines that convert audio data into text data. A specific example is a speech recognition API.

[1596] "Means of identifying customer intent" refers to techniques that use natural language processing engines or pre-trained AI models to extract user intent from text data.

[1597] A "generative AI model" is artificial intelligence that generates the optimal response based on the user's intent. Examples include generative AIs such as GPT-3.

[1598] "Means for generating answers" refers to software or algorithms that use generative AI models to generate answers based on the user's intent.

[1599] "Means of notifying the user of the answer" refers to a system for notifying the user of the generated answer. Specifically, this includes means of sending the answer to the terminal as text data or audio data, and displaying or playing it back as audio.

[1600] "Nuance and emotion" refers to elements that represent the customer's nuances and emotional expressions, extracted from text and audio data.

[1601] "Natural language processing" is a technique for analyzing text data and extracting meaning and intent. Examples include tools and algorithms such as "spaCy" and "BERT."

[1602] Modes for carrying out the invention

[1603] This invention relates to a customer service system that utilizes speech recognition and natural language processing technologies. It takes user voice input, analyzes the voice data in real time to identify the customer's intent, and provides the optimal response. The system consists of a user, a terminal, and a server.

[1604] System Configuration

[1605] User

[1606] Users input questions and requests via voice through the system. When a user speaks into the microphone, the voice data is transmitted to the device.

[1607] terminal

[1608] The terminal records the user's voice in digital format and sends it to the server. Some pre-processing is performed as needed before the received audio data is sent to the server. This pre-processing includes noise reduction and volume normalization. Specifically, "SoX (Sound eXchange)" or "Librosa" is used as the audio collection software on the terminal.

[1609] server

[1610] The server is the core of this system and performs the following main functions:

[1611] Audio data preprocessing (noise reduction and volume normalization)

[1612] Speech recognition (converting speech data to text data)

[1613] Natural language processing (analyzing user nuances and emotions from text data)

[1614] Intent extraction (identifying the user's purpose from text data)

[1615] Answer generation (generates the optimal answer using a generation AI model based on identified intent)

[1616] Submit the response (notify the user of the generated response)

[1617] Hardware and software to be used

[1618] Microphone (hardware): A device for voice input.

[1619] Audio acquisition software for the device: Anvil Studio or Audacity

[1620] Server audio processing software: SoX, Librosa

[1621] Speech recognition engine: Google Cloud Speech-to-Text API

[1622] Natural language processing engines: spaCy, BERT

[1623] Generative AI model: OpenAI GPT-3

[1624] Text-to-speech conversion: Google Cloud Text-to-Speech API

[1625] Specific example

[1626] Scenario 1: Inquiry about pricing plans

[1627] 1. User: "What is the recommended pricing plan?"

[1628] 2. Terminal: Records audio and sends it to the server.

[1629] 3. Server: Performs noise reduction and volume normalization.

[1630] 4. Server: Converts audio data to text. "What is the recommended pricing plan?"

[1631] 5. Server: Analyzes text data to identify user intent. "Questions about pricing plans"

[1632] 6. Server: Based on information about pricing plans, the AI ​​generates the optimal response. "Currently, there are three available pricing plans: A, B, and C. Their features are..."

[1633] 7. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[1634] Scenario 2: Buying a new smartphone

[1635] 1. User: "I want to buy a new smartphone."

[1636] 2. Terminal: Records audio and sends it to the server.

[1637] 3. Server: Performs noise reduction and volume normalization.

[1638] 4. Server: Converts audio data to text. "I want to buy a new smartphone."

[1639] 5. Server: Analyzes text data to identify user intent. "Question regarding the purchase of a new smartphone"

[1640] 6. Server: Based on information about the smartphone, the AI ​​generates the optimal answer. "The current new models are X, Y, and Z, and their respective features are..."

[1641] 7. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[1642] Examples of prompt messages include "What is your recommended pricing plan?" and "I'd like to buy a new smartphone."

[1643] In this way, the system can accurately understand the user's intent from their voice input and provide the optimal response. This can lead to improved customer satisfaction and consistent service quality.

[1644] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1645] Step 1:

[1646] User: Speaks a question or request into the microphone. For example, "What is the recommended pricing plan?" This generates audio data.

[1647] Step 2:

[1648] Terminal: Records the user's voice in digital format (e.g., WAV file format). Using voice collection software on the terminal (e.g., Anvil Studio, Audacity), the voice data is saved as digital data. Once recording is complete, the voice data is sent to the server using an HTTP request. To ensure data security, the TLS (Transport Layer Security) encryption protocol is used.

[1649] Step 3:

[1650] Server: Receives audio data and performs preprocessing such as noise reduction and volume normalization. Specifically, it uses libraries such as "SoX (Sound eXchange)" and Python's "Librosa". This process converts the audio data into a format suitable for analysis. For example, noise reduction reduces background noise and improves speech clarity.

[1651] Step 4:

[1652] Server: The server sends the pre-processed audio data to the speech recognition engine and converts the audio data into text data. The Google Cloud Speech-to-Text API is used here. The API is used to generate text data from the audio data. For example, the audio data "What is the best pricing plan?" is converted to the text data "What is the best pricing plan?".

[1653] Step 5:

[1654] Server: Text data generated by speech recognition is processed by a natural language processing engine (e.g., spaCy) to analyze the user's nuances and emotions. Simultaneously, a pre-trained AI model (e.g., BERT) is used to identify the user's intent from the text data. This process extracts meaning and intent from the text data. For example, "a question about pricing plans" might be extracted as the user's intent.

[1655] Step 6:

[1656] Server: Uses a generative AI model (e.g., OpenAI GPT-3) to generate the best answer based on the identified intent. The generative AI model is given a prompt and generates an appropriate answer. For example, based on the intent "Question about pricing plans," it generates an answer such as "Currently available pricing plans are A, B, and C. Their features are..."

[1657] Step 7:

[1658] Server: Formats the generated responses as text and audio data and sends them to the device. The Google Cloud Text-to-Speech API is used to generate the audio data. The device displays or plays the received data audibly to the user, allowing the user to visually or audibly confirm the responses.

[1659] For example, the generated response, "Currently available pricing plans are A, B, and C. Their features are...", is sent to the device, and the device plays the response aloud, allowing the user to hear the content. In this way, the system accurately analyzes the user's intent through concrete actions and provides the most appropriate response.

[1660] (Application Example 1)

[1661] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1662] Traditional customer service systems had a problem in that they made it difficult for users to get quick and accurate answers when interacting directly with staff in physical stores. This could lead to decreased customer satisfaction and increased workload for store staff, resulting in inconsistent service quality.

[1663] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1664] In this invention, the server includes means for transmitting the user's voice data to a smart device, means for displaying or playing back the answer to the user's question via the smart device, and means for analyzing nuances and intonation. This makes it possible to respond quickly and accurately to user questions in physical stores.

[1665] "Means for inputting user voice" refers to a mechanism for recording and acquiring user voice information using a digital device.

[1666] "Means for sending input audio data to a server" refers to communication means for sending audio data captured by a device to a server via a network.

[1667] "Methods for pre-processing audio data" refer to methods for processing captured audio data, such as noise reduction and volume normalization, to prepare it for easier analysis.

[1668] "Means for converting pre-processed audio data into text data" refers to a mechanism that utilizes speech recognition technology to convert pre-processed audio data into text information.

[1669] "Methods for identifying customer intent from text data" refer to methods that use natural language processing technology to analyze text data generated by speech recognition and identify the user's intent and purpose.

[1670] "Means for generating optimal responses based on identified intentions" refers to methods that use AI models to create optimal responses based on relevant information, taking into account the identified user intentions.

[1671] "Means of notifying the user of the generated response" refers to the means by which the server communicates the response it has created to the user in text or audio format.

[1672] "Means for transmitting user voice data to a smart device" refers to communication means for transmitting recorded user voice data to devices such as smartphones and smart glasses.

[1673] "Means for displaying or playing back the user's question via a smart device" refers to means such as a display device for visualizing the received answer on the smart device or a speaker for playing it back as sound.

[1674] System Overview

[1675] This invention is a customer service system that utilizes speech recognition and natural language processing technologies. The system takes voice input from the user, analyzes the voice data in real time to identify the customer's intent, and provides the optimal response. The system consists of a user, a terminal, and a server.

[1676] System Configuration

[1677] User

[1678] Users input questions and requests via voice through the system. When a user speaks into the microphone, the voice data is transmitted to their smart device.

[1679] terminal

[1680] The terminal consists of smart devices such as smartphones and smart glasses, which digitally record the user's voice and send the data to a server. Some pre-processing is performed on the received voice data before it is sent to the server.

[1681] server

[1682] The server is the core of this system and performs the following main functions:

[1683] Audio data preprocessing

[1684] Speech recognition

[1685] Natural Language Processing

[1686] Intent Extraction

[1687] Answer generation

[1688] Submit your response

[1689] Hardware and software to be used

[1690] Hardware:

[1691] Smartphones and smart glasses: Record user voice and communicate with the server.

[1692] Server: Performs data preprocessing, speech recognition, natural language processing, intent extraction, and response generation.

[1693] software:

[1694] Speech recognition engine (e.g., Google Cloud Speech-to-Text)

[1695] Natural language processing engines (e.g., SpaCy, BERT)

[1696] AI models (e.g., GPT-4)

[1697] Data processing and data calculation

[1698] Audio data preprocessing: The server performs preprocessing on the received audio data, such as noise reduction and volume normalization, and converts it into a format suitable for analysis.

[1699] Speech recognition: Pre-processed audio data is sent to the speech recognition engine, which converts the audio data into text data.

[1700] Natural Language Processing and Intent Extraction: Text data generated by speech recognition is processed by a natural language processing engine to analyze the user's nuances and intonation. Next, the user's intent is identified from the text data.

[1701] Response Generation: An AI model is used to generate the optimal response based on the identified intent. The generated response is formatted as text and audio data and sent to the device.

[1702] Specific example

[1703] Check product availability

[1704] 1. User: "Do you have this shirt in stock?"

[1705] 2. Terminal: Records audio and sends it to the server.

[1706] 3. Server: Performs noise reduction and volume normalization, and converts the audio data into text.

[1707] 4. Server: Identify user intent from text data.

[1708] 5. Server: Based on product inventory information, the AI ​​generates the optimal response. "This shirt is currently only available in size M. Other sizes are expected to arrive within a week."

[1709] 6. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[1710] Prompt example

[1711] "Do you have this shirt in stock?"

[1712] → AI model: Generates answers based on inventory information.

[1713] Example output: "This shirt is currently only available in size M. Other sizes will be available within a week."

[1714] In this way, customer service support tools utilizing speech recognition and natural language processing technologies can be used in physical stores. This allows store staff to respond to customer inquiries quickly and accurately, which is expected to improve customer satisfaction.

[1715] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1716] Step 1:

[1717] User voice input

[1718] The user speaks questions or requests into the microphone. The input at this time is audio data.

[1719] Step 2:

[1720] Recording and transmitting audio data

[1721] The terminal records the user's voice in digital format and sends that data to the server. The input here is the recorded voice data, and the output is the digital voice data sent to the server.

[1722] Step 3:

[1723] Audio data preprocessing

[1724] The server performs preprocessing on the received audio data, such as noise reduction and volume normalization, and converts it into a format suitable for analysis. The input is digital audio data transmitted from the terminal, and the output is the preprocessed audio data.

[1725] Step 4:

[1726] Speech recognition

[1727] The server sends the pre-processed audio data to a speech recognition engine (e.g., Google Cloud Speech-to-Text) to convert the audio data into text data. The input is pre-processed audio data, and the output is text data.

[1728] Step 5:

[1729] Natural language processing and intent extraction

[1730] The server processes the text data generated by speech recognition into a natural language processing engine (e.g., SpaCy, BERT) to analyze the user's nuances and intonation. Next, it identifies the user's intent from the analyzed text data. The input is text data, and the output is the identified user intent data.

[1731] Step 6:

[1732] Answer generation

[1733] The server utilizes an AI model (e.g., GPT-4) to generate the optimal response based on the identified intent. The generated response is then formatted as text and audio data. The input is the user's intent data, and the output is the generated response data.

[1734] Step 7:

[1735] Submit your response

[1736] The server sends the generated response data to the terminal. The input is the generated response data, and the output is the response data sent to the terminal.

[1737] Step 8:

[1738] Display or play the answer

[1739] The terminal displays or plays back the received response data to the user in audio format. The input is the response data sent from the server, and the output is the display or audio playback to the user.

[1740] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1741] System Overview

[1742] This invention relates to a customer service system that combines speech recognition, natural language processing (NLP), and sentiment analysis. It takes user voice input, analyzes the voice data in real time to identify the customer's intentions and emotions, and provides the most appropriate response. The system consists of a user, a terminal, and a server.

[1743] System Configuration

[1744] User

[1745] Users input questions and requests via voice through the system. When a user speaks into the microphone, the voice data is transmitted to the device.

[1746] terminal

[1747] The terminal records the user's voice in digital format and sends it to the server. Some pre-processing is performed as needed before the received audio data is sent to the server.

[1748] server

[1749] The server is the core of this system and performs the following main functions:

[1750] Means for preprocessing audio data

[1751] Means for performing speech recognition

[1752] Means of natural language processing

[1753] A method for performing emotion analysis using an emotion engine.

[1754] Means for extracting intent

[1755] Means for generating answers

[1756] Method for submitting a response

[1757] Processing flow

[1758] Voice input and transmission

[1759] 1. User: Speak your question or request into the microphone.

[1760] 2. Terminal: Records the user's voice as digital data and sends that data to the server.

[1761] Audio data preprocessing

[1762] 3. Server: The server performs preprocessing on the received audio data, such as noise reduction and volume normalization, and converts it into a format suitable for analysis.

[1763] Speech recognition

[1764] 4. Server: Sends the pre-processed audio data to the speech recognition engine and converts the audio data into text data.

[1765] Natural language processing and intent extraction

[1766] 5. Server: The text data generated by speech recognition is fed into a natural language processing engine to analyze the user's nuances and emotions. Next, the user's intent is identified from the text data.

[1767] Emotion analysis

[1768] 6. Server: Uses an emotion engine to analyze emotions from user voice and text data. The analysis results are used to identify customer intentions.

[1769] Response generation and submission

[1770] 7. Server: Utilizes an AI model to generate the optimal response based on the identified intent and sentiment analysis results. The generated response is then formatted as text and audio data.

[1771] 8. Server: Sends the formatted response to the terminal. The terminal displays or plays the received data in audio format.

[1772] Specific example

[1773] Scenario 1: Inquiry about pricing plans

[1774] 1. User: "What is the recommended pricing plan?"

[1775] 2. Terminal: Records audio and sends it to the server.

[1776] 3. Server: Performs noise reduction and volume normalization.

[1777] 4. Server: Converts audio data to text. "What is the recommended pricing plan?"

[1778] 5. Server: Analyzes text data to identify user intent. "Questions about pricing plans"

[1779] 6. Server: The emotion engine analyzes the user's emotional state and determines that they are "interested."

[1780] 7. Server: Based on information about pricing plans, the AI ​​generates the optimal response. "Currently, there are three available pricing plans: A, B, and C. Their features are…"

[1781] 8. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[1782] Scenario 2: Buying a new smartphone

[1783] 1. User: "I want to buy a new smartphone."

[1784] 2. Terminal: Records audio and sends it to the server.

[1785] 3. Server: Performs noise reduction and volume normalization.

[1786] 4. Server: Converts audio data to text. "I want to buy a new smartphone."

[1787] 5. Server: Analyzes text data to identify user intent. "Question regarding the purchase of a new smartphone"

[1788] 6. Server: The emotion engine analyzes the user's emotional state and determines that they have expectations.

[1789] 7. Server: Based on information about smartphones, the AI ​​generates the optimal answer. "The current new models are X, Y, and Z, and their respective features are..."

[1790] 8. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[1791] This system can accurately grasp the user's intent and emotions from their voice input and provide the most appropriate response. This not only improves customer satisfaction but also ensures consistent service quality.

[1792] The following describes the processing flow.

[1793] Step 1:

[1794] User: Speak your questions or requests into the microphone.

[1795] Step 2:

[1796] Terminal: Records the user's voice as digital data.

[1797] Step 3:

[1798] Terminal: Sends the recorded audio data to the server via an HTTP request.

[1799] Step 4:

[1800] Server: Analyzes the received audio data and applies a noise reduction filter.

[1801] Step 5:

[1802] Server: Normalizes the volume of the data to a certain level.

[1803] Step 6:

[1804] Server: Passes the pre-processed audio data to the speech recognition engine.

[1805] Step 7:

[1806] Server: The speech recognition engine converts the speech data into text data.

[1807] Step 8:

[1808] Server: Transfers the generated text data to the natural language processing (NLP) engine.

[1809] Step 9:

[1810] Server: The NLP engine performs sentiment analysis on text data, analyzing the customer's nuances and intonation.

[1811] Step 10:

[1812] Server: Uses an emotion engine to analyze the user's emotional state from text and audio data.

[1813] Step 11:

[1814] Server: Uses text intent analysis algorithms to identify customer questions and requests from text data.

[1815] Step 12:

[1816] Server: Based on sentiment analysis and intent analysis results, it utilizes AI models to generate appropriate answers to questions and requests.

[1817] Step 13:

[1818] Server: Adjusts the generated responses according to the user's emotional state and formats them as text and audio data.

[1819] Step 14:

[1820] Server: Sends the formatted response to the terminal as an HTTP response.

[1821] Step 15:

[1822] Terminal: Displays text data received from the server on the screen.

[1823] Step 16:

[1824] Terminal: Plays received audio data through the speaker.

[1825] Step 17:

[1826] User: Review the provided answers and re-enter any questions or requests as needed.

[1827] (Example 2)

[1828] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1829] Conventional speech recognition systems struggled to accurately analyze user intent and emotional state and provide optimal responses in real time. Furthermore, inconsistent customer service quality among different staff members led to decreased customer satisfaction. This compromised the user experience and hindered the provision of efficient customer service.

[1830] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for pre-processing received audio data, means for sending the pre-processed audio data to a speech recognition engine and converting it into text data, means for extracting the user's intent using a natural language processing engine based on the text data, means for analyzing the user's emotional state from the text data using an emotion analysis engine, means for generating an optimal response based on the intent and emotional state identified by a generation AI model, and means for sending the generated response to a terminal and notifying the user. This makes it possible to accurately grasp the intent and emotion from the user's voice input, improving customer satisfaction and standardizing the quality of service.

[1831] A "user" is the person who performs voice input.

[1832] A "terminal" is a device or system for recording a user's voice as digital data and transmitting it to a server.

[1833] A "server" is a core computer system that performs preprocessing of received audio data, speech recognition, natural language processing, sentiment analysis, and response generation.

[1834] "Audio data" refers to data that digitally represents what a user says to their device.

[1835] "Preprocessing" is the process of applying noise reduction and volume normalization to audio data and converting it into a format suitable for analysis.

[1836] A "speech recognition engine" is software or an algorithm used to convert speech data into text data.

[1837] "Text data" refers to character data converted by a speech recognition engine.

[1838] A "natural language processing engine" is software or an algorithm that analyzes text data to identify the user's intent and nuances.

[1839] An "emotion analysis engine" is software or an algorithm used to analyze a user's emotional state from text data.

[1840] A "generative AI model" is an artificial intelligence model that generates optimal answers based on the results of natural language processing and sentiment analysis.

[1841] "Answer" refers to the appropriate response to a user's inquiry, generated by a generative AI model.

[1842] "Notification" is the act of informing the user of the generated response.

[1843] The "system" is a mechanism that integrates all of these means and executes a series of processes to provide the optimal response from the user's voice input.

[1844] System Overview

[1845] This invention relates to a customer service system that combines speech recognition, natural language processing, and sentiment analysis to analyze the user's intent and emotions in real time from their voice input and provide the optimal response. The system consists of a user, a terminal, and a server.

[1846] User

[1847] Users input questions and requests via voice through the system. For example, if a user says "I want to buy a new smartphone" into the microphone, that voice is sent to the device.

[1848] terminal

[1849] The terminal records the user's voice in digital format and sends it to the server. Upon receiving the audio data, the terminal uses the "pyaudio" library to capture the audio and save it as digital data (for example, with the filename "input.wav"). The recorded digital audio data is transferred to the server in real time. Basic pre-processing such as noise reduction and volume normalization can also be performed.

[1850] server

[1851] The server, as the core of the system, performs the following main functions:

[1852] Audio data preprocessing: The server uses an audio analysis library such as "librosa" to remove noise and normalize the volume of the received audio data, and then converts it into a format suitable for analysis.

[1853] Speech Recognition: Pre-processed audio data is sent to a speech recognition engine (e.g., "Google Cloud Speech-to-Text API") to convert the audio into text data.

[1854] Natural Language Processing and Intent Extraction: Text data obtained through speech recognition is processed by natural language processing engines such as "SpaCy" and "BERT" to analyze the user's intent and nuances. For example, from the text "I want to buy a new smartphone," the intent "a question about purchasing a new smartphone" is identified.

[1855] Sentiment Analysis: Using an emotion analysis engine (e.g., "IBM Watson Tone Analyzer" or "DeepMoji"), the emotional state of the user is analyzed from their text data. The analysis results reveal emotions such as "the user has expectations."

[1856] Response Generation: Based on the identified intent and sentiment analysis results, the system generates the optimal response using a generative AI model (e.g., OpenAI's GPT-3). For example, it might generate a response such as, "The current new models are X, Y, and Z, and their respective features are..."

[1857] Notification: The generated response is sent to the device to notify the user. The device displays or plays the received response in audio format. Audio output tools such as "gTTS" or "Amazon Polly" are used for this purpose.

[1858] Specific example

[1859] Scenario 1: Inquiry about pricing plans

[1860] 1. User: "What is the recommended pricing plan?"

[1861] 2. Terminal: Records audio and sends it to the server.

[1862] 3. Server: Performs noise reduction and volume normalization.

[1863] 4. Server: Converts audio data to text. "What is the recommended pricing plan?"

[1864] 5. Server: Analyzes text data to identify user intent. "Questions about pricing plans"

[1865] 6. Server: The emotion engine analyzes the user's emotional state and determines that they are "interested."

[1866] 7. Server: Based on information about pricing plans, the AI ​​generates the optimal response. "Currently, there are three available pricing plans: A, B, and C. Their features are…"

[1867] 8. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[1868] Scenario 2: Buying a new smartphone

[1869] 1. User: "I want to buy a new smartphone."

[1870] 2. Terminal: Records audio and sends it to the server.

[1871] 3. Server: Performs noise reduction and volume normalization.

[1872] 4. Server: Converts audio data to text. "I want to buy a new smartphone."

[1873] 5. Server: Analyzes text data to identify user intent. "Question regarding the purchase of a new smartphone"

[1874] 6. Server: The emotion engine analyzes the user's emotional state and determines that they have expectations.

[1875] 7. Server: Based on information about smartphones, the AI ​​generates the optimal answer. "The current new models are X, Y, and Z, and their respective features are..."

[1876] 8. Terminal: Receives responses from the server and displays or plays them aloud for the user.

[1877] Example of a prompt

[1878] Let's assume a user inputs "I want to buy a new smartphone" via voice. The system converts the voice to text, identifies information about new smartphones through natural language processing and sentiment analysis, and then generates the best response. Please explain the detailed processing steps for each step of this system.

[1879] This system can accurately grasp the user's intent and emotions from their voice input and provide the most appropriate response. This not only improves customer satisfaction but also ensures consistent service quality.

[1880] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1881] Step 1:

[1882] User: Speaks a question or request into the microphone. Voice input is performed, for example, saying "I want to buy a new smartphone." This voice is the user's input.

[1883] Step 2:

[1884] Terminal: Records the user's voice as digital data and sends it to the server. Specifically, it uses the "pyaudio" library to capture the audio and save it in digital format. The recorded audio data (input.wav) becomes the terminal's output. This digital data is then sent directly to the server.

[1885] Step 3:

[1886] Server: Preprocesses the received audio data. Specifically, it uses the "librosa" library to perform noise reduction and volume normalization. The input for this preprocessing is the recorded audio data (input.wav), and the output is the preprocessed audio data. For example, noise is reduced using librosa.effects.preemphasis(input_signal).

[1887] Step 4:

[1888] Server: Sends pre-processed audio data to the speech recognition engine and converts the audio data into text data. The recognize method of the "Google Cloud Speech-to-Text API" is used. The input to this process is pre-processed audio data, and the output is the text data "I want to buy a new smartphone."

[1889] Step 5:

[1890] Server: The server processes the text data obtained from speech recognition into a natural language processing engine to extract the user's intent. The "SpaCy" library is used for this process. Specifically, the text data is analyzed using NLP (text) to identify the intent as "a question about purchasing a new smartphone." The input to this process is the text data obtained from speech recognition, and the output is user intent data.

[1891] Step 6:

[1892] Server: The server processes the text data obtained through natural language processing into an emotion analysis engine to analyze the user's emotional state. Specifically, it uses "IBM Watson Tone Analyzer" to execute tone_analyzer.tone(tone_input, content_type="application / json").get_result(). The input to this process is text data that identifies intent, and the output is analyzed emotion data (for example, "has expectations").

[1893] Step 7:

[1894] Server: Based on the identified intent and sentiment analysis results, it generates the optimal response using a generative AI model. For example, "GPT-3" is used as the generative AI model. The input to this process is user intent data and sentiment data, and the output is the generated response "The current new model is X, Y, Z, and its respective features are..."

[1895] Step 8:

[1896] Server: The server sends the generated response to the device and notifies the user. HTTP requests are often used as the notification method. For example, requests.post("http: / / device.endpoint", data=response_data) is used. The device displays or plays the received response data in audio format. When playing in audio format, "gTTS" is used, and the command tts = gTTS(response_text, lang='ja') is used. The input to this process is the generated response data, and the output is the information notified to the user.

[1897] (Application Example 2)

[1898] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1899] In brick-and-mortar stores, it is crucial to respond quickly and accurately to customer questions and requests in order to improve the quality and efficiency of customer service. However, conventional customer service systems have difficulty immediately grasping customer intentions and emotions and providing appropriate answers. Furthermore, it is difficult for store staff to respond in real time, which can lead to a decrease in customer satisfaction. This invention aims to solve these problems and provide a means for efficient and effective customer service in brick-and-mortar stores.

[1900] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[1901] In this invention, the server includes means for preprocessing audio data, means for converting the preprocessed audio data into text data, means for identifying the customer's intent from the text data, and a device for visually displaying the generated response. This makes it possible to analyze the customer's voice in real time, instantly grasp their intent and emotions, and provide the optimal response.

[1902] "Means of inputting user voice" refers to devices or programs that acquire voice data spoken by the user.

[1903] "Means for transmitting input audio data to a server" refers to devices or programs that transfer audio data acquired from a user to a server via a communication network.

[1904] "Means for preprocessing audio data" refers to devices or programs that perform processing on audio data, such as noise reduction and volume normalization, and convert it into a format that can be analyzed.

[1905] "Means for converting pre-processed audio data into text data" refers to devices or programs that analyze audio data and convert it into corresponding string data.

[1906] "Means for identifying customer intent from text data" refers to devices or programs that analyze text data and identify customer objectives and requirements from it.

[1907] "Means for generating optimal responses based on identified intentions" refers to devices or programs that produce appropriate responses in accordance with the customer's intentions.

[1908] "Means for notifying the user of the generated response" refers to devices or programs that communicate and provide the generated response to the user.

[1909] A "device for visually displaying generated answers" refers to a device that visualizes generated answers and presents them through a display, glasses-type device, or the like.

[1910] System Overview

[1911] This system is a customer service system for brick-and-mortar stores that takes user voice input, analyzes it to identify customer intent and emotions, and provides the most appropriate response. Its main components are the user, a terminal, and a server, which work together to function.

[1912] User

[1913] The user is a store employee wearing smart glasses. When a customer speaks to the employee, the employee receives the voice through the microphone in the smart glasses.

[1914] terminal

[1915] The device is a pair of smart glasses and has the following features:

[1916] Microphone for capturing customer voices

[1917] A display that provides visual information to the store staff.

[1918] Network function for communicating with the server

[1919] server

[1920] The server is the core of this system and performs the following main functions:

[1921] Means for preprocessing audio data

[1922] Means for performing speech recognition

[1923] Means of natural language processing

[1924] Means for performing emotion analysis

[1925] Means of identifying intent

[1926] Means for generating answers

[1927] Means of sending the answer to the device

[1928] Hardware and software to be used

[1929] Hardware: Smart glasses (e.g., Google Glass, Microsoft HoloLens)

[1930] software:

[1931] Speech recognition engine (e.g., Google Cloud Speech-to-Text API)

[1932] Natural language processing engine (e.g., Microsoft Azure Cognitive Services)

[1933] Sentiment analysis engine (e.g., IBM Watson Tone Analyzer)

[1934] Answer generation engine (e.g., OpenAI GPT-4)

[1935] Data processing flow and specific examples

[1936] 1. Voice acquisition:

[1937] A store employee wearing smart glasses uses a microphone to capture the customer's voice. For example, the customer might ask, "What products do you recommend for this season?"

[1938] 2. Sending audio data:

[1939] The acquired audio data is sent from the terminal to the server.

[1940] 3. Pre-treatment:

[1941] The server preprocesses the received audio data, performing noise reduction and volume normalization.

[1942] 4. Speech recognition:

[1943] This process converts pre-processed audio data into text data. In the previous example, the resulting text would be "What products do you recommend for this season?"

[1944] 5. Natural Language Processing:

[1945] The server analyzes the text data to identify the customer's intent. In this case, it is identified as "asking for product recommendations."

[1946] 6. Emotion analysis:

[1947] Customer emotions are analyzed from text and audio data. For example, the emotion of "interest" can be identified.

[1948] 7. Answer generation:

[1949] The server generates the optimal response based on the customer's intent and emotions. For example, it might say, "Our recommended products for this season are A, B, and C. Their features are..."

[1950] 8. Submit and display your response:

[1951] The generated response is displayed on the smart glasses' screen. The store clerk can then refer to the response and answer the customer directly.

[1952] Example of a prompt

[1953] Product: Customer support application

[1954] Model: Smart Glasses

[1955] Objective: To support customer service in physical stores.

[1956] Features: Speech recognition, natural language processing, sentiment analysis, response generation

[1957] Scenario: A customer asks a question about a product, and a store employee uses smart glasses to check the best answer and provide it instantly.

[1958] In this way, by accurately understanding the customer's intent and emotions from their voice input and providing the optimal response, the quality and efficiency of customer service in physical stores can be improved.

[1959] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1960] Step 1:

[1961] The user wears smart glasses and inputs the customer's voice. The smart glasses' microphone picks up the customer's speech and captures it as audio data. The input might be a customer question such as, "What products do you recommend for this season?"

[1962] Step 2:

[1963] The device (smart glasses) preprocesses the captured audio data. This preprocessing involves data manipulation such as noise reduction and volume normalization. As a result, audio data in a format that is easy for the server to analyze is output.

[1964] Step 3:

[1965] Pre-processed audio data is sent from the terminal to the server. In this process, the audio data is transferred to the server in real time via the internet. The input is the pre-processed audio data, and the output is the digital audio data that has reached the server.

[1966] Step 4:

[1967] The server passes the received audio data to a speech recognition engine (e.g., Google Cloud Speech-to-Text API) and converts it into text data. For example, the text "What products do you recommend this season?" is generated.

[1968] Step 5:

[1969] The server processes the generated text data using a natural language processing engine (e.g., Microsoft Azure Cognitive Services) to analyze the customer's intent. This process identifies the intent as "asking for product recommendations." The input is text data, and the output is the identified intent.

[1970] Step 6:

[1971] The server simultaneously uses an emotion analysis engine (e.g., IBM Watson Tone Analyzer) to analyze emotions from text and audio data. For example, the emotion "interested" might be identified. The input is text and audio data, and the output is the identified emotion.

[1972] Step 7:

[1973] The server uses a generative AI model (e.g., OpenAI GPT-4) to generate the most appropriate response based on identified intentions and emotions. For example, it might generate a response like, "Products A, B, and C are recommended for this season. Their characteristics are..." The input is intention and emotion, and the output is the generated response text.

[1974] Step 8:

[1975] The server sends the generated response to the device (smart glasses). In this process, the generated text data is transferred to the device via the internet. The input is the generated response text, and the output is the response data received by the device.

[1976] Step 9:

[1977] The terminal visually displays the generated response on the smart glasses' screen. The store clerk can then explain the product to the customer while looking at the displayed response. For example, the clerk might start by saying, "These are our recommended products for this season." The input is the response data received by the terminal, and the output is the visual information displayed on the screen.

[1978] In this way, a series of processes are realized in which the user receives voice input, the server analyzes it to generate the optimal response, and the terminal visualizes it. This system enables quick and accurate customer service in physical stores.

[1979] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1980] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1981] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[1982] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1983] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[1984] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[1985] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[1986] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[1987] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[1988] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[1989] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[1990] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[1991] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[1992] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1993] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[1994] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[1995] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[1996] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[1997] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[1998] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[1999] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[2000] The following is further disclosed regarding the embodiments described above.

[2001] (Claim 1)

[2002] A means of inputting the user's voice,

[2003] A means for sending input audio data to a server,

[2004] A means for preprocessing audio data,

[2005] A means for converting pre-processed audio data into text data,

[2006] A means of identifying customer intent from text data,

[2007] A means of generating the optimal response based on identified intent,

[2008] A means of notifying the user of the generated response,

[2009] A system that includes this.

[2010] (Claim 2)

[2011] The system according to claim 1, further comprising means for analyzing nuances and intonation when identifying customer intent.

[2012] (Claim 3)

[2013] The system according to claim 1, further comprising means for notifying the user of the generated response as text data and audio data.

[2014] "Example 1"

[2015] (Claim 1)

[2016] A means of inputting the user's voice,

[2017] A means by which the terminal records the input audio data and sends it to the server,

[2018] A means for preprocessing audio data,

[2019] A means for converting pre-processed audio data into text data,

[2020] A means of identifying customer intent from text data,

[2021] A means of generating the optimal response using a generative AI model based on identified intent,

[2022] A means of notifying the user of the generated response,

[2023] A system that includes this.

[2024] (Claim 2)

[2025] The system according to claim 1, further comprising means for analyzing nuances and emotions using natural language processing when identifying customer intent.

[2026] (Claim 3)

[2027] The system according to claim 1, further comprising means for notifying the user of the generated response as text data and audio data.

[2028] "Application Example 1"

[2029] (Claim 1)

[2030] A means of inputting the user's voice,

[2031] A means for sending input audio data to a server,

[2032] A means for preprocessing audio data,

[2033] A means for converting pre-processed audio data into text data,

[2034] A means of identifying customer intent from text data,

[2035] A means of generating the optimal response based on identified intent,

[2036] A means of notifying the user of the generated response,

[2037] A means of transmitting user voice data to a smart device,

[2038] A means of displaying or playing aloud answers to user questions using a smart device,

[2039] A system that includes this.

[2040] (Claim 2)

[2041] The system according to claim 1, further comprising means for analyzing nuances and intonation when identifying customer intent.

[2042] (Claim 3)

[2043] The system according to claim 1, further comprising means for notifying the user of the generated response as text data and audio data.

[2044] "Example 2 of combining an emotion engine"

[2045] (Claim 1)

[2046] A means of inputting the user's voice,

[2047] A means of transmitting input audio data to a terminal and recording it as digital data,

[2048] A means of sending audio data to a server,

[2049] A means for preprocessing the received audio data,

[2050] A means for sending pre-processed audio data to a speech recognition engine and converting it into text data,

[2051] A means of extracting user intent using a natural language processing engine based on text data,

[2052] A means of analyzing a user's emotional state from text data using an emotion analysis engine,

[2053] A means for generating the optimal response based on the intent and emotional state identified by a generative AI model,

[2054] A means of sending the generated response to the terminal and notifying the user,

[2055] A system that includes this.

[2056] (Claim 2)

[2057] The system according to claim 1, further comprising an emotion engine for analyzing nuances and emotions.

[2058] (Claim 3)

[2059] The system according to claim 1, further comprising means for notifying the user of the generated response as text data and audio data.

[2060] "Application example 2 when combining with an emotional engine"

[2061] (Claim 1)

[2062] A means of inputting the user's voice,

[2063] A means for sending input audio data to a server,

[2064] A means for preprocessing audio data,

[2065] A means for converting pre-processed audio data into text data,

[2066] A means of identifying customer intent from text data,

[2067] A means of generating the optimal response based on identified intent,

[2068] A means of notifying the user of the generated response,

[2069] A device that visually displays the generated answers,

[2070] A system that includes this.

[2071] (Claim 2)

[2072] The system according to claim 1, further comprising means for analyzing nuances and intonation when identifying customer intent.

[2073] (Claim 3)

[2074] The system according to claim 1, further comprising means for notifying the user of the generated response as text data and audio data. [Explanation of Symbols]

[2075] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of inputting the user's voice, A means for sending input audio data to a server, A means for preprocessing audio data, A means for converting pre-processed audio data into text data, A means of identifying customer intent from text data, A means of generating the optimal response based on identified intent, A means of notifying the user of the generated response, A system that includes this.

2. The system according to claim 1, further comprising means for analyzing nuances and intonation when identifying customer intent.

3. The system according to claim 1, further comprising means for notifying the user of the generated response as text data and audio data.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A