System

The system addresses call center inefficiencies by converting voice data to text, using generative AI for optimal responses, and storing data for analysis, enhancing user satisfaction and service quality.

JP2026028763APending Publication Date: 2026-02-20SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024131379
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-07
Publication Date
2026-02-20

AI Technical Summary

Technical Problem

Conventional call centers face long wait times, inefficient responses, and inadequate complaint handling, with insufficient data accumulation and analysis leading to lower user satisfaction, and insufficient data analysis for improving the handling of user data for improving the handling of user data, resulting in lower user satisfaction and lower user satisfaction, and lack of effective data analysis for improving services.

Method used

A system that includes voice data conversion to text, analysis by generative AI for optimal response generation, voice synthesis, and data storage and analysis for improving user satisfaction and service quality, with the system including a server, operator terminal, and data processing devices for efficient complaint handling and data accumulation.

Benefits of technology

The system provides prompt and accurate responses, improves user satisfaction by generating optimal responses, and enables data-driven improvements in service quality through data analysis and storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026028763000001_ABST
    Figure 2026028763000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for receiving audio and converting the audio to text; means for sending the text to a generation AI, wherein the text is parsed by the generation AI to generate optimal response information; and means for receiving the generated response information, converting the response information to audio, and providing the audio to a user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In conventional call centers, users often experience long wait times and perceived inefficiency in response. Furthermore, complaint handling can sometimes lack appropriate wording and follow-up, creating a risk of lower user satisfaction. Furthermore, there has been insufficient data accumulation and analysis to effectively analyze the content of user inquiries and use the data to improve future responses. The present invention aims to solve these problems and improve call center response efficiency and user satisfaction. [Means for solving the problem]

[0005] The present invention is a system including the following means. First, it provides a means for receiving voice data and converting the voice data into text data. Next, it provides a means for transmitting the text data to a generation AI, which analyzes the text data and generates optimal response information. Then, it includes a means for receiving the generated response information, converting the response information into voice data, and providing it to the user. It further includes a means for transmitting the optimal response information to an operator terminal and displaying multiple answer candidates on the operator terminal. It also includes a means for, if the user's request is a complaint, using the generation AI to propose an optimal response and displaying the proposal on the operator terminal. It also includes a means for generating a follow-up message after handling the complaint and a means for providing the follow-up message to the user. Finally, the system includes a means for storing user requests and inquiry content and generated response information, a means for categorizing the stored data, analyzing the number of cases and trends, and a means for reporting the analysis results to the planning department or the quality department.

[0006] "Audio data" refers to data in which the user's voice or sound is recorded in digital format.

[0007] "Character data" refers to voice data converted into text format using voice recognition technology.

[0008] "Generative AI" is an artificial intelligence system that analyzes input text data and generates optimal response information based on its content.

[0009] "Response information" is information generated by the generation AI that includes appropriate answers and countermeasures to users' inquiries and requests.

[0010] A "voice recognition engine" is a software or hardware system that receives voice data and converts the voice data into text data in real time.

[0011] A "voice synthesis engine" is a technology that converts text data and response information into voice data and plays it back to the user.

[0012] An "operator terminal" is a computer or mobile device used by an operator that displays the answer candidates and suggestions provided by the generation AI.

[0013] A "complaint" is a complaint or problem that a user expresses.

[0014] A "follow-up message" is a message of additional information or gratitude provided to improve customer satisfaction after a complaint has been handled.

[0015] "Categorization" is the process of classifying and organizing accumulated data based on attributes and characteristics.

[0016] "Volume and trend analysis" is the activity of analyzing inquiry frequencies and patterns using categorized data.

[0017] A "planning or quality department" is a department within an organization that is responsible for planning, developing, and improving new products and services. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be explained.

[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0026] [First embodiment]

[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0039] The present invention is embodied in the following manner.

[0040] First, when a user calls the call center, the server detects the incoming call and activates the IVR (Interactive Voice Response) system. In the top menu of this IVR system, the user is prompted to "select an AI operator." When the user presses a button or responds by saying "AI operator," the server activates a speech recognition engine and converts the user's voice data into text data in real time.

[0041] The server then sends the converted text data to the generation AI. The generation AI analyzes the received text data and understands the user's inquiry. Based on this, the generation AI generates optimal response information. This response information is sent back to the server as text data, and the server uses a speech synthesis engine to convert the response information into voice data. This voice data is then sent back to the user.

[0042] Additionally, if an operator needs to intervene, the server sends the generated response information to the operator terminal. The operator terminal displays the information on its screen, allowing the operator to select an appropriate response. In particular, if the user files a complaint, the generation AI proposes the optimal response and displays it on the operator terminal. The operator responds according to the generation AI's proposal, and if necessary, the generation AI generates a follow-up message and sends it to the user via the server.

[0043] In addition, the server accumulates and stores data such as user inquiries and response information. This accumulated data is later categorized and used for analyzing the number of cases and trends. This allows the planning or quality department to refer to the analysis results and use them to improve services and products.

[0044] Specific examples

[0045] For example, if a user says, "My internet connection is slow," the server receives this voice data and converts it into text data using a speech recognition engine. This text data is then sent to a generation AI, which then generates a response that provides the optimal solution to the "slow internet connection," such as "Try restarting your router." The server then converts this response information into voice data and sends it back to the user.

[0046] Additionally, if a complaint occurs, for example, if the content is "The product was not delivered," the generation AI will suggest the optimal response, such as "We are sorry, we will check and contact you as soon as possible." Based on this suggestion displayed on the operator's terminal, the operator will respond quickly, and the generation AI will also generate a follow-up message that will be sent to the user.

[0047] In this way, the system of the present invention can respond to user inquiries quickly and accurately, thereby improving user satisfaction.

[0048] The processing flow will be explained below.

[0049] Step 1:

[0050] A user calls the call center. The server detects the call and activates the IVR system. The IVR system's top menu displays a prompt to "Select an AI operator."

[0051] Step 2:

[0052] The user presses a button or responds by saying "AI operator." The server activates a speech recognition engine, receives the user's voice data, and converts it into text data in real time.

[0053] Step 3:

[0054] The server sends the converted text data to the generation AI, which analyzes the text data and understands the user's inquiry.

[0055] Step 4:

[0056] The AI ​​generates the optimal response information for the user's inquiry, and the generated response information is sent to the server as text data.

[0057] Step 5:

[0058] The server converts the received response information into voice data using a voice synthesis engine, and plays the converted voice data back to the user to provide an appropriate response.

[0059] Step 6:

[0060] If necessary, the server sends the generated response information to the operator terminal, which displays answer candidates on the operator's screen, allowing the operator to select the most appropriate response.

[0061] Step 7:

[0062] When a user files a complaint, the server converts the voice data into text data and sends it to the AI ​​generator, which analyzes the content of the complaint and suggests the best response and wording.

[0063] Step 8:

[0064] The server sends the generated proposal to the operator terminal, which displays it on the screen and allows the operator to respond appropriately to the user.

[0065] Step 9:

[0066] After the complaint is handled, the AI ​​generates a follow-up message, which the server converts into audio data and sends to the user.

[0067] Step 10:

[0068] The server accumulates and stores user inquiries and response information. The accumulated data is categorized and analyzed for number of cases and trends. The analysis results are reported to the planning or quality department and used to improve services.

[0069] The above are the specific processing steps of the system. By taking appropriate measures at each step, user satisfaction can be improved.

[0070] Example 1

[0071] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0072] Conventional call center systems often provided slow or insufficient responses to user inquiries, resulting in low user satisfaction. They also faced the challenge of finding it difficult to quickly provide appropriate solutions for complaints. Furthermore, the accumulation and analysis of inquiry data was insufficient, making it difficult to utilize the data to improve services and products.

[0073] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0074] In this invention, the server includes means for receiving voice data and converting the voice data into text data, means for transmitting the text data to a generation AI, which analyzes the text data and generates optimal response information, means for receiving the generated response information, converting the response information into voice data, and providing it to the user, and means for accumulating, saving, and later analyzing data such as user inquiries and response information. This allows users to receive prompt and accurate responses, and also improves the quality of customer complaints by having the generation AI propose optimal responses. Furthermore, the accumulated data can be used to improve services and products.

[0075] "Audio data" refers to digitally converted data of an audio signal acquired via an audio input device such as a telephone or microphone.

[0076] "Text data" refers to data obtained by analyzing voice data using a voice recognition engine or the like and converting temporal voice signals into text format.

[0077] "Generative AI" is an artificial intelligence model that analyzes received text data and generates optimal response information based on the user's intentions and requests.

[0078] "Response information" refers to information about appropriate responses or solutions generated by the generation AI based on the user's inquiries or requests.

[0079] A "speech synthesis engine" is software that generates natural-sounding speech based on text data or response information.

[0080] "Operator terminal" means a computer system or device used by an operator to respond to user inquiries.

[0081] A "call center system" refers to the entire system for responding to user inquiries and complaints and providing appropriate support and services.

[0082] "Storage and preservation" means continuously recording the data obtained and retaining it for a certain period of time for future reference and analysis.

[0083] A "complaint" is a complaint made by a user reporting dissatisfaction or problems with a product or service, and requesting improvements or responses.

[0084] "Analysis" refers to quantitatively and qualitatively evaluating accumulated data using statistical methods and machine learning to identify areas for improvement and trends in the service.

[0085] This invention relates to a call center system that responds quickly and accurately to inquiries from users. This system includes technology for converting voice data into text data, analyzing it, and generating appropriate responses. Specific embodiments of this system are described below.

[0086] Hardware and Software Configuration

[0087] 1. Hardware

[0088] Server: Processes queries and transforms and stores data.

[0089] Operator terminal: An operator responds to user inquiries.

[0090] User terminal: A device used by a user to make an inquiry, such as a telephone or smartphone.

[0091] 2. Software

[0092] IVR system: Provides voice guidance in response to user inquiries and prompts users to select the appropriate menu.

[0093] Speech recognition engine: Converts voice data into text data in real time (e.g., Google Cloud Speech-to-Text).

[0094] Generative AI model: Analyzes text data and generates appropriate response information (e.g., OpenAI GPT-4).

[0095] Speech synthesis engine: Converts the generated response information into voice data (e.g., Amazon Polly).

[0096] Database: Query and response data is accumulated and stored for later analysis.

[0097] Data processing and calculation

[0098] 1. Converting audio data to text data

[0099] The server uses a speech recognition engine to convert the user's voice input into text data in real time.

[0100] 2. Analysis of character data and generation of response information

[0101] The server sends the converted text data to the generative AI model, which analyzes the data and generates optimal response information.

[0102] 3. Converting response information into voice data

[0103] The server sends the response information received from the generative AI model to the speech synthesis engine, converts it into voice data, and responds to the user.

[0104] 4. Sending information to the operator terminal

[0105] If necessary, the server transmits the generated response information to the operator terminal, and the operator responds appropriately.

[0106] 5. Data accumulation and analysis

[0107] The server accumulates the inquiry and response data and stores it in a database for later analysis, which can lead to improvements in services and products.

[0108] Specific examples

[0109] Slow Internet Connection Inquiries

[0110] When a user calls to complain about a slow internet connection, the server receives this voice data and converts it into text using a speech recognition engine. The converted text data is sent to a generative AI model, which generates a response such as "Please try restarting your router" as an analysis result. The server then converts this response data into voice data using a speech synthesis engine and sends it back to the user.

[0111] Prompt Sentence Examples

[0112] For example, below is an example of a prompt sentence for a generative AI model in response to the query "My internet connection is slow."

[0113] A user has made the following inquiry: "My internet connection is slow." Please suggest the best solution for this problem.

[0114] This allows users to receive prompt and accurate responses, and the quality of responses to complaints can be improved by the generative AI model proposing optimal responses. Furthermore, analyzing the accumulated data can lead to improvements in services and products.

[0115] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0116] Step 1:

[0117] A user calls a call center.

[0118] The user dials the call center number using a landline or smartphone, which initiates the user's inquiry as voice data.

[0119] Step 2:

[0120] The server detects the incoming call and activates the IVR system.

[0121] When the server detects an incoming call, it launches the IVR system, which provides a voice prompt to the user, saying, "Please select an AI operator." The input is the user's phone number, and the output is the launch of the IVR system.

[0122] Step 3:

[0123] The user selects "AI Operator."

[0124] The user responds by saying "AI operator" or selects "AI operator" from the menu, which inputs the voice data.

[0125] Step 4:

[0126] The server uses a speech recognition engine to convert the user's voice into text data.

[0127] The server runs a speech recognition engine and converts the user's voice data into text data in real time. The input is voice data and the output is text data. Specifically, it uses a speech recognition service such as Google Cloud Speech-to-Text.

[0128] Step 5:

[0129] The server sends the character data to the generation AI.

[0130] The server sends the converted text data to the generative AI model. The input is text data, and the output is data sent to the generative AI model. Specifically, the text data is sent through the API.

[0131] Step 6:

[0132] Generative AI analyzes the text data and generates the optimal response.

[0133] The generation AI analyzes the received text data and generates the optimal response information based on the user's inquiry. The input is text data and the output is response information. Specific operations include using OpenAI GPT-4 and other technologies to generate a response based on the prompt text.

[0134] Step 7:

[0135] The server converts the generated response into voice data and sends it back to the user.

[0136] The server sends the response information to a speech synthesis engine and converts it into voice data. The input is the response information and the output is voice data. Specific operations use a speech synthesis service such as Amazon Polly.

[0137] Step 8:

[0138] If operator intervention is required, the server sends response information to the operator terminal.

[0139] If the generating AI determines that operator intervention is necessary, the server sends response information to the operator terminal. The input is the response information, and the output is data transmission to the operator terminal. The specific operation is that the response is displayed on the operator's console display.

[0140] Step 9:

[0141] The server accumulates and stores the query data and response data.

[0142] The server accumulates query data and response data from users and stores them in a database for later analysis. The input is query data and response data, and the output is updating the database. Specifically, data is stored using a database management system.

[0143] keyword

[0144] Generative AI model, prompt sentence

[0145] (Application example 1)

[0146] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0147] User support at call centers is prone to delayed responses, human error, and increased operator workloads. Another issue is the lack of systems and methods for users to receive prompt and appropriate support on the internet and online shopping sites. In particular, it can be difficult to receive prompt and appropriate responses when users submit queries, which can lead to a decline in user satisfaction due to delays in handling complaints. To solve these problems, an efficient support system utilizing speech recognition, generative AI, and speech synthesis is needed.

[0148] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0149] In this invention, the server includes means for receiving voice data and converting the voice data into text data, means for transmitting the text data to a generation AI, which analyzes the text data and generates optimal response information, means for receiving the generated response information and converting the response information into voice data to provide to the user, means for the user to issue a query and generate an optimal answer regarding products on the mail-order site as a response to the query, and means for converting the optimal answer into voice data using a voice synthesis engine and returning it to the user. This allows users to receive quick and appropriate responses in real time, and allows for efficient handling of complaints and the like, thereby improving user satisfaction.

[0150] "Voice data" refers to information transmitted by a user through voice recorded in digital format.

[0151] "Character data" refers to data in text format of the audio content converted using voice recognition technology.

[0152] "Generative AI" refers to artificial intelligence that generates optimal responses and information based on input information.

[0153] "Response information" is text data of answers and suggestions generated by the generation AI based on the analysis results.

[0154] A "speech synthesis engine" is a technology or device that analyzes text data and converts it into speech data.

[0155] A "query" is information sent by a user as an inquiry or question.

[0156] A "user" is a person who sends an inquiry or question to the system.

[0157] An "operator terminal" is a computer or device used by an operator to display the generative AI's suggestions.

[0158] An "online shopping site" is an online sales platform where users can search for product information and make purchases.

[0159] The "optimal answer" is the most appropriate and useful response that the generative AI derives based on the user's query.

[0160] A "complaint" is a complaint or problem report from a user.

[0161] This invention is applicable to cases where a user accesses an online shopping site and sends product-related information or a query. When the user sends a query using a smartphone app, the following system functions.

[0162] First, the server receives voice data from the smartphone app. This voice data is the content of a user's inquiry. Next, the server converts the voice data into text data using a speech recognition engine (speech_recognition library). The converted text data is sent to a generative AI model (e.g., a GPT-based model).

[0163] The generative AI model analyzes the received text data and generates optimal response information. This response information is sent in text format to the server. The server then converts this text-to-speech response information into voice data using a text-to-speech engine (gTTS), and sends the voice data back to the user's smartphone app.

[0164] The specific hardware required is a microphone to capture voice data and a speaker to provide responses to the user.The software requires the speech_recognition library for speech recognition, the transformers library for implementing generative AI, and the gTTS library for speech synthesis.

[0165] For example, if a user asks, "What are the best deals on the latest smartphones?":

[0166] 1. The server receives this voice data and uses a voice recognition engine to convert it into text data such as "Please tell me about discount information on the latest smartphones."

[0167] 2. The generative AI model receives this text data and generates a response such as, "Currently, the latest smartphone models are 10% off."

[0168] 3. The server passes this response information to a speech synthesis engine, converts it into voice data, and responds to the user.

[0169] An example prompt is:

[0170] "Can you tell me about discounts on the latest smartphones?"

[0171] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0172] Step 1:

[0173] A user sends a query using a smartphone app. The user taps the microphone icon in the app and says, "Please tell me about the latest smartphone discount information." The input is the user's voice data, and the output is the voice data sent from the smartphone app to the server.

[0174] Step 2:

[0175] The server receives the voice data sent by the user. Then, the server starts a speech recognition engine (speech_recognition library) and converts this voice data into text data. The input is the user's voice data, and the output is the text data "Please tell me about discount information on the latest smartphones."

[0176] Step 3:

[0177] The server sends the converted text data to the generative AI model, which analyzes the text data and generates the optimal response information. The input is the text data (query), and the output is text data of the response information, such as "Currently, the latest model smartphones are 10% off."

[0178] Step 4:

[0179] The server receives the generated response information and passes this text data to the speech synthesis engine (gTTS). The speech synthesis engine analyzes the text data and converts it into voice data. The input is the text data of the response information, and the output is voice data.

[0180] Step 5:

[0181] The server sends the converted voice data to a smartphone app, and the user receives a response via the smartphone app. The input is voice data, and the output is the voice that the user hears.

[0182] Each processing step is performed sequentially, allowing the user to receive a prompt and appropriate response.

[0183] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0184] The present invention relates to a call center system that analyzes user voice data and incorporates an emotion engine that recognizes emotions. Specific embodiments of the present invention will be described below.

[0185] First, when a user calls the call center, the server detects the call and activates the IVR system. The IVR system's top menu displays a message saying, "Please select an AI operator." When the user presses a button or responds by saying "AI operator," the server activates a speech recognition engine and converts the user's voice data into text data in real time.

[0186] The server then sends the converted text data to the generation AI, which analyzes the text data and understands the user's inquiry. During this process, the emotion engine recognizes emotions from the user's voice data. The recognized emotion data is sent to the generation AI, which then takes this emotional information into account to generate optimal response information.

[0187] The generated response information is sent back to the server as text data. The server then converts this response information into voice data using a speech synthesis engine and provides it to the user. If necessary, the server also sends the generated response information to an operator terminal. The operator terminal displays this information on its screen, allowing the operator to select an appropriate response.

[0188] When a complaint occurs, the server converts the complaint's voice data into text data and sends it to the generation AI. The generation AI analyzes the content of the complaint and suggests the optimal response and appropriate wording based on the emotional data recognized by the emotion engine. These suggestions are displayed on the operator's terminal, and the operator responds accordingly. After the complaint is handled, the generation AI creates a follow-up message and sends it to the user via the server.

[0189] The server also accumulates and stores user inquiries, response information, and recognized emotion data. The accumulated data is later categorized and used to analyze the number of cases and trends. This allows the planning or quality department to refer to the analysis results and use them to improve services and products.

[0190] Specific examples

[0191] For example, if a user says, "My internet connection is slow," the server receives this voice data and converts it into text data using a speech recognition engine. This text data is then sent to the generation AI, which then generates a response that optimally addresses the "slow internet connection," such as "Please try restarting your router." At the same time, if the emotion engine detects "dissatisfaction" in the user's voice, the generation AI will also include an additional apology, such as "We apologize for the inconvenience."

[0192] Additionally, if a complaint is made that the product was not delivered, the generation AI will suggest the optimal response, such as "We are sorry, we will check and contact you as soon as possible." At the same time, if the emotion engine detects "anger," the generation AI will add a response that is more in tune with the emotion, such as "We are very sorry for the inconvenience." Based on the suggestions displayed on the operator terminal, the operator will respond quickly, and a generated follow-up message will be sent to the user later.

[0193] In this way, the system of the present invention can recognize the user's emotions and respond optimally to them, thereby increasing user satisfaction.

[0194] The processing flow will be explained below.

[0195] Step 1:

[0196] A user calls the call center. The server detects the call and activates the IVR system. The IVR system's top menu displays a message saying, "Please select an AI operator."

[0197] Step 2:

[0198] The user presses a button or responds by saying "AI operator." The server activates a speech recognition engine, receives the user's voice data, and converts it into text data in real time.

[0199] Step 3:

[0200] The server sends the converted text data to the generation AI, which analyzes the text data and understands the user's inquiry.

[0201] Step 4:

[0202] The emotion engine recognizes emotions from the user's voice data and sends the recognized emotion data to the generation AI.

[0203] Step 5:

[0204] The AI ​​generates the optimal response information based on the user's inquiry and emotional data. The generated response information is sent to the server as text data.

[0205] Step 6:

[0206] The server converts the received response information into voice data using a voice synthesis engine, and plays the converted voice data back to the user to provide an appropriate response.

[0207] Step 7:

[0208] If necessary, the server sends the generated response information to the operator terminal, which displays answer candidates and emotion data on the operator's screen, allowing the operator to select the optimal response.

[0209] Step 8:

[0210] When a user files a complaint, the server converts the voice data into text data and sends it to the generation AI, which analyzes the content of the complaint and suggests the best response and appropriate wording based on the emotional data recognized by the emotion engine.

[0211] Step 9:

[0212] The server sends the generated proposal to the operator terminal, which displays the proposal and emotion data on the screen, allowing the operator to respond appropriately to the user.

[0213] Step 10:

[0214] After the complaint is handled, the AI ​​generates a follow-up message, which the server converts into audio data and sends to the user.

[0215] Step 11:

[0216] The server accumulates and stores user inquiries, response information, emotional data, etc. The accumulated data is categorized and used to analyze the number of cases and trends. The results of the analysis are reported to the planning or quality department and used to improve services.

[0217] The above are the specific processing steps of this system. By taking appropriate measures at each step, user satisfaction can be improved.

[0218] Example 2

[0219] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0220] Conventional call center systems have difficulty accurately understanding users' emotions and requests and providing prompt and optimal responses. Furthermore, when handling complaints, it is difficult for operators to immediately find the optimal response, which can lead to lower user satisfaction. To solve these issues, an advanced system that can analyze users' emotional data and generate optimal responses is needed.

[0221] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0222] In this invention, the server includes means for receiving voice data and converting the voice data into character data, means for transmitting the character data to an artificial intelligence (AI) to be analyzed by the AI ​​and generating optimal response information, means for recognizing emotion data from the voice data by an emotion analysis means and transmitting the emotion data to the AI, and means for receiving the generated response information, converting the response information into voice data, and providing it to the user, thereby enabling the generation and provision of optimal responses taking the user's emotions into consideration.

[0223] "Voice data" is data that digitally represents the voice that the user utters to the system.

[0224] "Character data" is data in text format that has been converted from voice data into character information.

[0225] "Artificial intelligence" is a system that uses machine learning and data analysis to automatically perform specific tasks.

[0226] "Emotion analysis means" is a technology that recognizes the user's emotions from voice data and extracts emotion data.

[0227] "Response information" is information in response to a user's inquiry that is generated as a result of analysis by artificial intelligence.

[0228] "Operation device" refers to a terminal or computer used by an operator, and is a device that displays information necessary for responding to users.

[0229] The "optimal response" is the most appropriate response method that is generated taking into consideration the user's requests and feelings.

[0230] The present invention relates to a call center system that analyzes user voice data and incorporates an emotion engine that recognizes emotions. Specific embodiments of the present invention will be described below.

[0231] First, when a user calls the call center, the server detects the call. The server then activates the IVR (Interactive Voice Response) system and plays a message saying, "Please select an AI operator." This message is generated using a text-to-speech engine (e.g., Amazon Polly).

[0232] Next, when the user presses a button or responds by saying "AI operator," the server recognizes this and activates a speech recognition engine (for example, Google Cloud Speech-to-Text API). The server receives the voice data and converts it into text data in real time using the speech recognition engine.

[0233] The converted text data is sent from the server to a generation AI (e.g., OpenAI GPT-3). The generation AI analyzes the text data and understands the user's inquiry. At the same time, an emotion analysis means (e.g., IBM Watson Tone Analyzer) analyzes the user's voice data and recognizes emotional data. This emotional data is also sent to the generation AI.

[0234] The AI ​​generates optimal responses based on the user's query and emotional data. The server receives the responses and converts them into voice data using a speech synthesis engine (e.g., Amazon Polly) to provide them to the user.

[0235] The server also sends the generated response information to the operator terminal. The operator terminal displays this response information on the screen, allowing the operator to select a response. In particular, if the user's request is a complaint, the generation AI proposes the optimal response and displays it on the operator terminal. The operator then takes appropriate action based on the proposal.

[0236] For example, consider the case where a user inquires about a slow internet connection. The server converts the received voice data into text data using a speech recognition engine, and then sends that text data to the generation AI. The generation AI generates a response that optimally addresses the slow internet connection issue, such as "Please try restarting your router." At the same time, if the emotion engine detects "dissatisfaction" in the user's voice, the generation AI also adds an apology, such as "We apologize for the inconvenience."

[0237] As another example, consider the case of a complaint that "the product was not delivered." In this case, the generation AI will similarly generate the optimal response, such as "We are sorry, we will check and contact you as soon as possible." Furthermore, if the emotion engine detects "anger," the generation AI will also add a response that is more in tune with the emotion, such as "We are very sorry for the inconvenience." The operator will respond promptly based on the suggestions displayed on the operator terminal.

[0238] In this way, the system of the present invention can increase user satisfaction by recognizing the user's emotions and providing the optimal response. The server also accumulates and saves the inquiry content, response information, and recognized emotion data in a database for later analysis and service improvement.

[0239] An example of a prompt sentence is "Please enter your query." For example, by entering "My Internet connection is slow," the system will generate an appropriate answer.

[0240] This invention makes it possible to realize effective call center response that takes into account the user's emotions.

[0241] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0242] Step 1:

[0243] When a user calls the call center, the server detects the call. Specifically, the server uses the VoIP system to trigger an incoming call event for a specific phone number. Based on the incoming call data (input), the server activates the IVR system (output).

[0244] Step 2:

[0245] The server starts the IVR system and plays a message to the user saying, "Please select an AI operator." This message is generated by the server using a text-to-speech engine (e.g., Amazon Polly), which converts text data (input) into voice data (output).

[0246] Step 3:

[0247] When the user presses a button or responds by saying "AI operator," the server recognizes this and starts a speech recognition engine (for example, Google Cloud Speech-to-Text API), which prepares the system to convert the voice data (input) into text data (output).

[0248] Step 4:

[0249] The server receives the user's voice data and converts it into text data in real time using a speech recognition engine. Specifically, the server sends the voice data to the Google Cloud Speech-to-Text API and receives the converted text data. The voice data (input) is converted into text data (output).

[0250] Step 5:

[0251] The server sends the acquired character data to the generation AI (e.g., OpenAI GPT-3). The server sends the character data (input) to the generation AI's API and receives the analysis result, which is the response information (output).

[0252] Step 6:

[0253] The generation AI analyzes the text data and understands the user's inquiry. For example, in response to the text data (input) "My internet connection is slow," it generates the appropriate solution, "Try restarting your router."

[0254] Step 7:

[0255] An emotion analysis means (e.g., IBM Watson Tone Analyzer) analyzes the user's voice data and recognizes the emotion data. The server sends the voice data (input) to the emotion analysis API and sends the recognized emotion data (output) to the generation AI.

[0256] Step 8:

[0257] The generation AI generates optimal response information by taking into account the content of the user's inquiry and emotional data. If the emotional data indicates "dissatisfaction," the generation AI adds an apology such as "We apologize for the inconvenience." Based on the text data and emotional data (input), the AI ​​generates comprehensive response information (output).

[0258] Step 9:

[0259] The server receives the generated response information and converts it into voice data using a speech synthesis engine (e.g., Amazon Polly). Text data (input) is converted into voice data (output).

[0260] Step 10:

[0261] The server provides the converted voice data to the user, specifically by playing the voice data (input) to the user's telephone terminal (output).

[0262] Step 11:

[0263] If necessary, the server sends the generated response information to the operator terminal. Character data (input) is sent to the operation device and displayed (output).

[0264] Step 12:

[0265] When a complaint occurs, the server converts the complaint voice data into text data and sends it to the generation AI. The generation AI analyzes the content of the complaint and proposes the optimal response. This is displayed on the operator terminal, allowing the operator to select a response. The voice data (input) is converted into text data and emotional data (output), and response information is generated.

[0266] Step 13:

[0267] After handling the complaint, the generation AI generates a follow-up message and sends it to the user via the server. It sends text data (input) and provides the user with follow-up message data (output).

[0268] Step 14:

[0269] The server accumulates and stores the user's inquiry, response information, and recognized emotion data in a database. This helps to increase user satisfaction and improve services. Text and emotion data (input) are stored in the database (output).

[0270] (Application example 2)

[0271] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0272] Conventional driver assistance systems for autonomous vehicles rely primarily on visual analysis, such as the user's facial expressions and body movements, and lack the ability to analyze emotions, including vocalizations, in real time. This makes it difficult to properly detect and quickly respond to user stress and dissatisfaction. The present invention aims to solve these problems and provide a comfortable riding experience in autonomous vehicles.

[0273] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving voice data and converting the voice data into character data, means for sending the character data to a generative AI model, which analyzes the character data and generates optimal response information, means for analyzing the user's emotions and sending the emotion data to the generative AI model, means for the generative AI model to generate optimal response information in consideration of the emotion data, and means for receiving the generated response information, converting the response information into voice data, and providing it to the user. This makes it possible to analyze emotions from the user's voice in real time and quickly provide an adaptive response.

[0274] "Audio data" means a digital representation of sound waves captured using a microphone or other input device.

[0275] "Character data" is data that indicates a string of characters that is recognized as a unit of language by analyzing voice data.

[0276] A "generative AI model" is an artificial intelligence system designed to analyze input text data and automatically generate appropriate responses and processing.

[0277] "Emotion data" is data that indicates the emotional state recognized from the user's voice data.

[0278] "Operator terminal" refers to equipment used to operate the system, and in particular to the computer and display device used by the operator.

[0279] "Response information" is data that indicates an appropriate answer or suggestion to a user's inquiry or request.

[0280] "Analysis" is the process of examining and evaluating the content and characteristics of input data in detail.

[0281] "Conversion" is the process of replacing data of one format with data of another format.

[0282] A "server" is a central management device for operating the entire system and processing information, and is a computer that processes data and performs communication.

[0283] This invention relates to an emotion recognition driver assistance system for autonomous vehicles. This system has the function of analyzing user voice data and providing optimal responses based on the user's emotions. Specific embodiments are described below.

[0284] First, the vehicle's server receives voice data through a microphone, which is then converted into text data using a voice recognition engine installed in the server.

[0285] The server then sends the converted text data to the generative AI model, which analyzes the user's speech. At the same time, the emotion engine analyzes the voice data and generates the user's emotional data, which is also sent to the generative AI model.

[0286] The generative AI model takes into account the text and emotional data to generate the optimal response information. This response information is sent back to the server and converted into voice data using a speech synthesis engine. This voice data is then provided to the user through the car's speakers.

[0287] The server also sends the generated response information to the operator terminal, which displays multiple answer candidates, allowing the operator to refer to them and make the most appropriate response.

[0288] For example, if a user says, "I'm tired, today was stressful," this voice data is converted into text data and sent to the generative AI model. At the same time, the emotion engine recognizes "stress" and sends it to the generative AI model as emotion data.

[0289] The generative AI model generates the optimal response based on the following prompt:

[0290] User says: I'm tired, today was stressful

[0291] User Emotion: Stress

[0292] Generate the appropriate response for your system:

[0293] The generative AI model generates a response such as, "Thank you for your hard work. I'll play some relaxing music to help you relax." The server converts this response into audio data and provides it to the user through a speaker.

[0294] The hardware used includes a microphone to collect voice data and a speaker to play the audio back to the user, while the software uses a speech recognition engine, an emotion engine, and a generative AI model, which allows for real-time recognition of user emotions and provides adaptive services.

[0295] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0296] Step 1:

[0297] The server receives voice data through the microphone. The input is the user's voice, and the output is digital voice data. Specifically, the microphone inside the car picks up the user's speech, converts it into a digital signal, and sends it to the server.

[0298] Step 2:

[0299] The server uses a speech recognition engine to convert the received voice data into text data. The input is voice data and the output is text data. Specifically, the speech recognition engine analyzes the sound waves and generates the corresponding text in real time.

[0300] Step 3:

[0301] The server sends the converted text data to the generative AI model, which then analyzes the user's speech. The input is text data, and the output is the analysis result. Specifically, the generative AI model understands the text, analyzes the context and intent, and prepares to generate an appropriate response.

[0302] Step 4:

[0303] The server uses an emotion engine to analyze emotion data from voice data. The input is voice data, and the output is emotion data. Specifically, the emotion engine analyzes voice parameters such as tone, tempo, and pitch to estimate the user's emotional state.

[0304] Step 5:

[0305] The server sends the emotion data to the generative AI model, which then generates optimal response information taking the emotion data into consideration. The input is emotion data and text data, and the output is response information. Specifically, the generative AI model generates a prompt sentence based on the text and emotion information, and then creates an appropriate response based on that.

[0306] Step 6:

[0307] The server receives the generated response information and converts it into voice data using a voice synthesis engine. The input is the response information and the output is voice data. Specifically, the voice synthesis engine converts the text into a voice format and generates data that can be played back as a human voice.

[0308] Step 7:

[0309] The server provides the converted voice data to the user through the car's speaker. The input is voice data, and the output is the voice that the user can hear. Specifically, the server sends the voice data to the speaker, and the speaker plays it back.

[0310] Step 8:

[0311] The server sends the generated response information to the operator terminal, which displays multiple answer candidates. The input is the response information, and the output is the displayed answer candidates. The specific operation is that the operator terminal displays the received information on the screen so that the operator can confirm the options.

[0312] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0313] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0314] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0315] [Second embodiment]

[0316] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0317] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0318] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0319] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0320] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0321] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0322] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0323] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0324] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0325] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0326] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0327] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0328] The present invention is embodied in the following manner.

[0329] First, when a user calls the call center, the server detects the incoming call and activates the IVR (Interactive Voice Response) system. In the top menu of this IVR system, the user is prompted to "select an AI operator." When the user presses a button or responds by saying "AI operator," the server activates a speech recognition engine and converts the user's voice data into text data in real time.

[0330] The server then sends the converted text data to the generation AI. The generation AI analyzes the received text data and understands the user's inquiry. Based on this, the generation AI generates optimal response information. This response information is sent back to the server as text data, and the server uses a speech synthesis engine to convert the response information into voice data. This voice data is then sent back to the user.

[0331] Additionally, if an operator needs to intervene, the server sends the generated response information to the operator terminal. The operator terminal displays the information on its screen, allowing the operator to select an appropriate response. In particular, if the user files a complaint, the generation AI proposes the optimal response and displays it on the operator terminal. The operator responds according to the generation AI's proposal, and if necessary, the generation AI generates a follow-up message and sends it to the user via the server.

[0332] In addition, the server accumulates and stores data such as user inquiries and response information. This accumulated data is later categorized and used for analyzing the number of cases and trends. This allows the planning or quality department to refer to the analysis results and use them to improve services and products.

[0333] Specific examples

[0334] For example, if a user says, "My internet connection is slow," the server receives this voice data and converts it into text data using a speech recognition engine. This text data is then sent to a generation AI, which then generates a response that provides the optimal solution to the "slow internet connection," such as "Try restarting your router." The server then converts this response information into voice data and sends it back to the user.

[0335] Additionally, if a complaint occurs, for example, if the content is "The product was not delivered," the generation AI will suggest the optimal response, such as "We are sorry, we will check and contact you as soon as possible." Based on this suggestion displayed on the operator's terminal, the operator will respond quickly, and the generation AI will also generate a follow-up message that will be sent to the user.

[0336] In this way, the system of the present invention can respond to user inquiries quickly and accurately, thereby improving user satisfaction.

[0337] The processing flow will be explained below.

[0338] Step 1:

[0339] A user calls the call center. The server detects the call and activates the IVR system. The IVR system's top menu displays a prompt to "Select an AI operator."

[0340] Step 2:

[0341] The user presses a button or responds by saying "AI operator." The server activates a speech recognition engine, receives the user's voice data, and converts it into text data in real time.

[0342] Step 3:

[0343] The server sends the converted text data to the generation AI, which analyzes the text data and understands the user's inquiry.

[0344] Step 4:

[0345] The AI ​​generates the optimal response information for the user's inquiry, and the generated response information is sent to the server as text data.

[0346] Step 5:

[0347] The server converts the received response information into voice data using a voice synthesis engine, and plays the converted voice data back to the user to provide an appropriate response.

[0348] Step 6:

[0349] If necessary, the server sends the generated response information to the operator terminal, which displays answer candidates on the operator's screen, allowing the operator to select the most appropriate response.

[0350] Step 7:

[0351] When a user files a complaint, the server converts the voice data into text data and sends it to the AI ​​generator, which analyzes the content of the complaint and suggests the best response and wording.

[0352] Step 8:

[0353] The server sends the generated proposal to the operator terminal, which displays it on the screen and allows the operator to respond appropriately to the user.

[0354] Step 9:

[0355] After the complaint is handled, the AI ​​generates a follow-up message, which the server converts into audio data and sends to the user.

[0356] Step 10:

[0357] The server accumulates and stores user inquiries and response information. The accumulated data is categorized and analyzed for number of cases and trends. The analysis results are reported to the planning or quality department and used to improve services.

[0358] The above are the specific processing steps of the system. By taking appropriate measures at each step, user satisfaction can be improved.

[0359] Example 1

[0360] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0361] Conventional call center systems often provided slow or insufficient responses to user inquiries, resulting in low user satisfaction. They also faced the challenge of finding it difficult to quickly provide appropriate solutions for complaints. Furthermore, the accumulation and analysis of inquiry data was insufficient, making it difficult to utilize the data to improve services and products.

[0362] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0363] In this invention, the server includes means for receiving voice data and converting the voice data into text data, means for transmitting the text data to a generation AI, which analyzes the text data and generates optimal response information, means for receiving the generated response information, converting the response information into voice data, and providing it to the user, and means for accumulating, saving, and later analyzing data such as user inquiries and response information. This allows users to receive prompt and accurate responses, and also improves the quality of customer complaints by having the generation AI propose optimal responses. Furthermore, the accumulated data can be used to improve services and products.

[0364] "Audio data" refers to digitally converted data of an audio signal acquired via an audio input device such as a telephone or microphone.

[0365] "Text data" refers to data obtained by analyzing voice data using a voice recognition engine or the like and converting temporal voice signals into text format.

[0366] "Generative AI" is an artificial intelligence model that analyzes received text data and generates optimal response information based on the user's intentions and requests.

[0367] "Response information" refers to information about appropriate responses or solutions generated by the generation AI based on the user's inquiries or requests.

[0368] A "speech synthesis engine" is software that generates natural-sounding speech based on text data or response information.

[0369] "Operator terminal" means a computer system or device used by an operator to respond to user inquiries.

[0370] A "call center system" refers to the entire system for responding to user inquiries and complaints and providing appropriate support and services.

[0371] "Storage and preservation" means continuously recording the data obtained and retaining it for a certain period of time for future reference and analysis.

[0372] A "complaint" is a complaint made by a user reporting dissatisfaction or problems with a product or service, and requesting improvements or responses.

[0373] "Analysis" refers to quantitatively and qualitatively evaluating accumulated data using statistical methods and machine learning to identify areas for improvement and trends in the service.

[0374] This invention relates to a call center system that responds quickly and accurately to inquiries from users. This system includes technology for converting voice data into text data, analyzing it, and generating appropriate responses. Specific embodiments of this system are described below.

[0375] Hardware and Software Configuration

[0376] 1. Hardware

[0377] Server: Processes queries and transforms and stores data.

[0378] Operator terminal: An operator responds to user inquiries.

[0379] User terminal: A device used by a user to make an inquiry, such as a telephone or smartphone.

[0380] 2. Software

[0381] IVR system: Provides voice guidance in response to user inquiries and prompts users to select the appropriate menu.

[0382] Speech recognition engine: Converts voice data into text data in real time (e.g., Google Cloud Speech-to-Text).

[0383] Generative AI model: Analyzes text data and generates appropriate response information (e.g., OpenAI GPT-4).

[0384] Speech synthesis engine: Converts the generated response information into voice data (e.g., Amazon Polly).

[0385] Database: Query and response data is accumulated and stored for later analysis.

[0386] Data processing and calculation

[0387] 1. Converting audio data to text data

[0388] The server uses a speech recognition engine to convert the user's voice input into text data in real time.

[0389] 2. Analysis of character data and generation of response information

[0390] The server sends the converted text data to the generative AI model, which analyzes the data and generates optimal response information.

[0391] 3. Converting response information into voice data

[0392] The server sends the response information received from the generative AI model to the speech synthesis engine, converts it into voice data, and responds to the user.

[0393] 4. Sending information to the operator terminal

[0394] If necessary, the server transmits the generated response information to the operator terminal, and the operator responds appropriately.

[0395] 5. Data accumulation and analysis

[0396] The server accumulates the inquiry and response data and stores it in a database for later analysis, which can lead to improvements in services and products.

[0397] Specific examples

[0398] Slow Internet Connection Inquiries

[0399] When a user calls to complain about a slow internet connection, the server receives this voice data and converts it into text using a speech recognition engine. The converted text data is sent to a generative AI model, which generates a response such as "Please try restarting your router" as an analysis result. The server then converts this response data into voice data using a speech synthesis engine and sends it back to the user.

[0400] Prompt Sentence Examples

[0401] For example, below is an example of a prompt sentence for a generative AI model in response to the query "My internet connection is slow."

[0402] A user has made the following inquiry: "My internet connection is slow." Please suggest the best solution for this problem.

[0403] This allows users to receive prompt and accurate responses, and the quality of responses to complaints can be improved by the generative AI model proposing optimal responses. Furthermore, analyzing the accumulated data can lead to improvements in services and products.

[0404] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0405] Step 1:

[0406] A user calls a call center.

[0407] The user dials the call center number using a landline or smartphone, which initiates the user's inquiry as voice data.

[0408] Step 2:

[0409] The server detects the incoming call and activates the IVR system.

[0410] When the server detects an incoming call, it launches the IVR system, which provides a voice prompt to the user, saying, "Please select an AI operator." The input is the user's phone number, and the output is the launch of the IVR system.

[0411] Step 3:

[0412] The user selects "AI Operator."

[0413] The user responds by saying "AI operator" or selects "AI operator" from the menu, which inputs the voice data.

[0414] Step 4:

[0415] The server uses a speech recognition engine to convert the user's voice into text data.

[0416] The server runs a speech recognition engine and converts the user's voice data into text data in real time. The input is voice data and the output is text data. Specifically, it uses a speech recognition service such as Google Cloud Speech-to-Text.

[0417] Step 5:

[0418] The server sends the character data to the generation AI.

[0419] The server sends the converted text data to the generative AI model. The input is text data, and the output is data sent to the generative AI model. Specifically, the text data is sent through the API.

[0420] Step 6:

[0421] Generative AI analyzes the text data and generates the optimal response.

[0422] The generation AI analyzes the received text data and generates the optimal response information based on the user's inquiry. The input is text data and the output is response information. Specific operations include using OpenAI GPT-4 and other technologies to generate a response based on the prompt text.

[0423] Step 7:

[0424] The server converts the generated response into voice data and sends it back to the user.

[0425] The server sends the response information to a speech synthesis engine and converts it into voice data. The input is the response information and the output is voice data. Specific operations use a speech synthesis service such as Amazon Polly.

[0426] Step 8:

[0427] If operator intervention is required, the server sends response information to the operator terminal.

[0428] If the generating AI determines that operator intervention is necessary, the server sends response information to the operator terminal. The input is the response information, and the output is data transmission to the operator terminal. The specific operation is that the response is displayed on the operator's console display.

[0429] Step 9:

[0430] The server accumulates and stores the query data and response data.

[0431] The server accumulates query data and response data from users and stores them in a database for later analysis. The input is query data and response data, and the output is updating the database. Specifically, data is stored using a database management system.

[0432] keyword

[0433] Generative AI model, prompt sentence

[0434] (Application example 1)

[0435] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0436] User support at call centers is prone to delayed responses, human error, and increased operator workloads. Another issue is the lack of systems and methods for users to receive prompt and appropriate support on the internet and online shopping sites. In particular, it can be difficult to receive prompt and appropriate responses when users submit queries, which can lead to a decline in user satisfaction due to delays in handling complaints. To solve these problems, an efficient support system utilizing speech recognition, generative AI, and speech synthesis is needed.

[0437] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0438] In this invention, the server includes means for receiving voice data and converting the voice data into text data, means for transmitting the text data to a generation AI, which analyzes the text data and generates optimal response information, means for receiving the generated response information and converting the response information into voice data to provide to the user, means for the user to issue a query and generate an optimal answer regarding products on the mail-order site as a response to the query, and means for converting the optimal answer into voice data using a voice synthesis engine and returning it to the user. This allows users to receive quick and appropriate responses in real time, and allows for efficient handling of complaints and the like, thereby improving user satisfaction.

[0439] "Voice data" refers to information transmitted by a user through voice recorded in digital format.

[0440] "Character data" refers to data in text format of the audio content converted using voice recognition technology.

[0441] "Generative AI" refers to artificial intelligence that generates optimal responses and information based on input information.

[0442] "Response information" is text data of answers and suggestions generated by the generation AI based on the analysis results.

[0443] A "speech synthesis engine" is a technology or device that analyzes text data and converts it into speech data.

[0444] A "query" is information sent by a user as an inquiry or question.

[0445] A "user" is a person who sends an inquiry or question to the system.

[0446] An "operator terminal" is a computer or device used by an operator to display the generative AI's suggestions.

[0447] An "online shopping site" is an online sales platform where users can search for product information and make purchases.

[0448] The "optimal answer" is the most appropriate and useful response that the generative AI derives based on the user's query.

[0449] A "complaint" is a complaint or problem report from a user.

[0450] This invention is applicable to cases where a user accesses an online shopping site and sends product-related information or a query. When the user sends a query using a smartphone app, the following system functions.

[0451] First, the server receives voice data from the smartphone app. This voice data is the content of a user's inquiry. Next, the server converts the voice data into text data using a speech recognition engine (speech_recognition library). The converted text data is sent to a generative AI model (e.g., a GPT-based model).

[0452] The generative AI model analyzes the received text data and generates optimal response information. This response information is sent in text format to the server. The server then converts this text-to-speech response information into voice data using a text-to-speech engine (gTTS), and sends the voice data back to the user's smartphone app.

[0453] The specific hardware required is a microphone to capture voice data and a speaker to provide responses to the user.The software requires the speech_recognition library for speech recognition, the transformers library for implementing generative AI, and the gTTS library for speech synthesis.

[0454] For example, if a user asks, "What are the best deals on the latest smartphones?":

[0455] 1. The server receives this voice data and uses a voice recognition engine to convert it into text data such as "Please tell me about discount information on the latest smartphones."

[0456] 2. The generative AI model receives this text data and generates a response such as, "Currently, the latest smartphone models are 10% off."

[0457] 3. The server passes this response information to a speech synthesis engine, converts it into voice data, and responds to the user.

[0458] An example prompt is:

[0459] "Can you tell me about discounts on the latest smartphones?"

[0460] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0461] Step 1:

[0462] A user sends a query using a smartphone app. The user taps the microphone icon in the app and says, "Please tell me about the latest smartphone discount information." The input is the user's voice data, and the output is the voice data sent from the smartphone app to the server.

[0463] Step 2:

[0464] The server receives the voice data sent by the user. Then, the server starts a speech recognition engine (speech_recognition library) and converts this voice data into text data. The input is the user's voice data, and the output is the text data "Please tell me about discount information on the latest smartphones."

[0465] Step 3:

[0466] The server sends the converted text data to the generative AI model, which analyzes the text data and generates the optimal response information. The input is the text data (query), and the output is text data of the response information, such as "Currently, the latest model smartphones are 10% off."

[0467] Step 4:

[0468] The server receives the generated response information and passes this text data to the speech synthesis engine (gTTS). The speech synthesis engine analyzes the text data and converts it into voice data. The input is the text data of the response information, and the output is voice data.

[0469] Step 5:

[0470] The server sends the converted voice data to a smartphone app, and the user receives a response via the smartphone app. The input is voice data, and the output is the voice that the user hears.

[0471] Each processing step is performed sequentially, allowing the user to receive a prompt and appropriate response.

[0472] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0473] The present invention relates to a call center system that analyzes user voice data and incorporates an emotion engine that recognizes emotions. Specific embodiments of the present invention will be described below.

[0474] First, when a user calls the call center, the server detects the call and activates the IVR system. The IVR system's top menu displays a message saying, "Please select an AI operator." When the user presses a button or responds by saying "AI operator," the server activates a speech recognition engine and converts the user's voice data into text data in real time.

[0475] The server then sends the converted text data to the generation AI, which analyzes the text data and understands the user's inquiry. During this process, the emotion engine recognizes emotions from the user's voice data. The recognized emotion data is sent to the generation AI, which then takes this emotional information into account to generate optimal response information.

[0476] The generated response information is sent back to the server as text data. The server then converts this response information into voice data using a speech synthesis engine and provides it to the user. If necessary, the server also sends the generated response information to an operator terminal. The operator terminal displays this information on its screen, allowing the operator to select an appropriate response.

[0477] When a complaint occurs, the server converts the complaint's voice data into text data and sends it to the generation AI. The generation AI analyzes the content of the complaint and suggests the optimal response and appropriate wording based on the emotional data recognized by the emotion engine. These suggestions are displayed on the operator's terminal, and the operator responds accordingly. After the complaint is handled, the generation AI creates a follow-up message and sends it to the user via the server.

[0478] The server also accumulates and stores user inquiries, response information, and recognized emotion data. The accumulated data is later categorized and used to analyze the number of cases and trends. This allows the planning or quality department to refer to the analysis results and use them to improve services and products.

[0479] Specific examples

[0480] For example, if a user says, "My internet connection is slow," the server receives this voice data and converts it into text data using a speech recognition engine. This text data is then sent to the generation AI, which then generates a response that optimally addresses the "slow internet connection," such as "Please try restarting your router." At the same time, if the emotion engine detects "dissatisfaction" in the user's voice, the generation AI will also include an additional apology, such as "We apologize for the inconvenience."

[0481] Additionally, if a complaint is made that the product was not delivered, the generation AI will suggest the optimal response, such as "We are sorry, we will check and contact you as soon as possible." At the same time, if the emotion engine detects "anger," the generation AI will add a response that is more in tune with the emotion, such as "We are very sorry for the inconvenience." Based on the suggestions displayed on the operator terminal, the operator will respond quickly, and a generated follow-up message will be sent to the user later.

[0482] In this way, the system of the present invention can recognize the user's emotions and respond optimally to them, thereby increasing user satisfaction.

[0483] The processing flow will be explained below.

[0484] Step 1:

[0485] A user calls the call center. The server detects the call and activates the IVR system. The IVR system's top menu displays a message saying, "Please select an AI operator."

[0486] Step 2:

[0487] The user presses a button or responds by saying "AI operator." The server activates a speech recognition engine, receives the user's voice data, and converts it into text data in real time.

[0488] Step 3:

[0489] The server sends the converted text data to the generation AI, which analyzes the text data and understands the user's inquiry.

[0490] Step 4:

[0491] The emotion engine recognizes emotions from the user's voice data and sends the recognized emotion data to the generation AI.

[0492] Step 5:

[0493] The AI ​​generates the optimal response information based on the user's inquiry and emotional data. The generated response information is sent to the server as text data.

[0494] Step 6:

[0495] The server converts the received response information into voice data using a voice synthesis engine, and plays the converted voice data back to the user to provide an appropriate response.

[0496] Step 7:

[0497] If necessary, the server sends the generated response information to the operator terminal, which displays answer candidates and emotion data on the operator's screen, allowing the operator to select the optimal response.

[0498] Step 8:

[0499] When a user files a complaint, the server converts the voice data into text data and sends it to the generation AI, which analyzes the content of the complaint and suggests the best response and appropriate wording based on the emotional data recognized by the emotion engine.

[0500] Step 9:

[0501] The server sends the generated proposal to the operator terminal, which displays the proposal and emotion data on the screen, allowing the operator to respond appropriately to the user.

[0502] Step 10:

[0503] After the complaint is handled, the AI ​​generates a follow-up message, which the server converts into audio data and sends to the user.

[0504] Step 11:

[0505] The server accumulates and stores user inquiries, response information, emotional data, etc. The accumulated data is categorized and used to analyze the number of cases and trends. The results of the analysis are reported to the planning or quality department and used to improve services.

[0506] The above are the specific processing steps of this system. By taking appropriate measures at each step, user satisfaction can be improved.

[0507] Example 2

[0508] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0509] Conventional call center systems have difficulty accurately understanding users' emotions and requests and providing prompt and optimal responses. Furthermore, when handling complaints, it is difficult for operators to immediately find the optimal response, which can lead to lower user satisfaction. To solve these issues, an advanced system that can analyze users' emotional data and generate optimal responses is needed.

[0510] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0511] In this invention, the server includes means for receiving voice data and converting the voice data into character data, means for transmitting the character data to an artificial intelligence (AI) to be analyzed by the AI ​​and generating optimal response information, means for recognizing emotion data from the voice data by an emotion analysis means and transmitting the emotion data to the AI, and means for receiving the generated response information, converting the response information into voice data, and providing it to the user, thereby enabling the generation and provision of optimal responses taking the user's emotions into consideration.

[0512] "Voice data" is data that digitally represents the voice that the user utters to the system.

[0513] "Character data" is data in text format that has been converted from voice data into character information.

[0514] "Artificial intelligence" is a system that uses machine learning and data analysis to automatically perform specific tasks.

[0515] "Emotion analysis means" is a technology that recognizes the user's emotions from voice data and extracts emotion data.

[0516] "Response information" is information in response to a user's inquiry that is generated as a result of analysis by artificial intelligence.

[0517] "Operation device" refers to a terminal or computer used by an operator, and is a device that displays information necessary for responding to users.

[0518] The "optimal response" is the most appropriate response method that is generated taking into consideration the user's requests and feelings.

[0519] The present invention relates to a call center system that analyzes user voice data and incorporates an emotion engine that recognizes emotions. Specific embodiments of the present invention will be described below.

[0520] First, when a user calls the call center, the server detects the call. The server then activates the IVR (Interactive Voice Response) system and plays a message saying, "Please select an AI operator." This message is generated using a text-to-speech engine (e.g., Amazon Polly).

[0521] Next, when the user presses a button or responds by saying "AI operator," the server recognizes this and activates a speech recognition engine (for example, Google Cloud Speech-to-Text API). The server receives the voice data and converts it into text data in real time using the speech recognition engine.

[0522] The converted text data is sent from the server to a generation AI (e.g., OpenAI GPT-3). The generation AI analyzes the text data and understands the user's inquiry. At the same time, an emotion analysis means (e.g., IBM Watson Tone Analyzer) analyzes the user's voice data and recognizes emotional data. This emotional data is also sent to the generation AI.

[0523] The AI ​​generates optimal responses based on the user's query and emotional data. The server receives the responses and converts them into voice data using a speech synthesis engine (e.g., Amazon Polly) to provide them to the user.

[0524] The server also sends the generated response information to the operator terminal. The operator terminal displays this response information on the screen, allowing the operator to select a response. In particular, if the user's request is a complaint, the generation AI proposes the optimal response and displays it on the operator terminal. The operator then takes appropriate action based on the proposal.

[0525] For example, consider the case where a user inquires about a slow internet connection. The server converts the received voice data into text data using a speech recognition engine, and then sends that text data to the generation AI. The generation AI generates a response that optimally addresses the slow internet connection issue, such as "Please try restarting your router." At the same time, if the emotion engine detects "dissatisfaction" in the user's voice, the generation AI also adds an apology, such as "We apologize for the inconvenience."

[0526] As another example, consider the case of a complaint that "the product was not delivered." In this case, the generation AI will similarly generate the optimal response, such as "We are sorry, we will check and contact you as soon as possible." Furthermore, if the emotion engine detects "anger," the generation AI will also add a response that is more in tune with the emotion, such as "We are very sorry for the inconvenience." The operator will respond promptly based on the suggestions displayed on the operator terminal.

[0527] In this way, the system of the present invention can increase user satisfaction by recognizing the user's emotions and providing the optimal response. The server also accumulates and saves the inquiry content, response information, and recognized emotion data in a database for later analysis and service improvement.

[0528] An example of a prompt sentence is "Please enter your query." For example, by entering "My Internet connection is slow," the system will generate an appropriate answer.

[0529] This invention makes it possible to realize effective call center response that takes into account the user's emotions.

[0530] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0531] Step 1:

[0532] When a user calls the call center, the server detects the call. Specifically, the server uses the VoIP system to trigger an incoming call event for a specific phone number. Based on the incoming call data (input), the server activates the IVR system (output).

[0533] Step 2:

[0534] The server starts the IVR system and plays a message to the user saying, "Please select an AI operator." This message is generated by the server using a text-to-speech engine (e.g., Amazon Polly), which converts text data (input) into voice data (output).

[0535] Step 3:

[0536] When the user presses a button or responds by saying "AI operator," the server recognizes this and starts a speech recognition engine (for example, Google Cloud Speech-to-Text API), which prepares the system to convert the voice data (input) into text data (output).

[0537] Step 4:

[0538] The server receives the user's voice data and converts it into text data in real time using a speech recognition engine. Specifically, the server sends the voice data to the Google Cloud Speech-to-Text API and receives the converted text data. The voice data (input) is converted into text data (output).

[0539] Step 5:

[0540] The server sends the acquired character data to the generation AI (e.g., OpenAI GPT-3). The server sends the character data (input) to the generation AI's API and receives the analysis result, which is the response information (output).

[0541] Step 6:

[0542] The generation AI analyzes the text data and understands the user's inquiry. For example, in response to the text data (input) "My internet connection is slow," it generates the appropriate solution, "Try restarting your router."

[0543] Step 7:

[0544] An emotion analysis means (e.g., IBM Watson Tone Analyzer) analyzes the user's voice data and recognizes the emotion data. The server sends the voice data (input) to the emotion analysis API and sends the recognized emotion data (output) to the generation AI.

[0545] Step 8:

[0546] The generation AI generates optimal response information by taking into account the content of the user's inquiry and emotional data. If the emotional data indicates "dissatisfaction," the generation AI adds an apology such as "We apologize for the inconvenience." Based on the text data and emotional data (input), the AI ​​generates comprehensive response information (output).

[0547] Step 9:

[0548] The server receives the generated response information and converts it into voice data using a speech synthesis engine (e.g., Amazon Polly). Text data (input) is converted into voice data (output).

[0549] Step 10:

[0550] The server provides the converted voice data to the user, specifically by playing the voice data (input) to the user's telephone terminal (output).

[0551] Step 11:

[0552] If necessary, the server sends the generated response information to the operator terminal. Character data (input) is sent to the operation device and displayed (output).

[0553] Step 12:

[0554] When a complaint occurs, the server converts the complaint voice data into text data and sends it to the generation AI. The generation AI analyzes the content of the complaint and proposes the optimal response. This is displayed on the operator terminal, allowing the operator to select a response. The voice data (input) is converted into text data and emotional data (output), and response information is generated.

[0555] Step 13:

[0556] After handling the complaint, the generation AI generates a follow-up message and sends it to the user via the server. It sends text data (input) and provides the user with follow-up message data (output).

[0557] Step 14:

[0558] The server accumulates and stores the user's inquiry, response information, and recognized emotion data in a database. This helps to increase user satisfaction and improve services. Text and emotion data (input) are stored in the database (output).

[0559] (Application example 2)

[0560] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0561] Conventional driver assistance systems for autonomous vehicles rely primarily on visual analysis, such as the user's facial expressions and body movements, and lack the ability to analyze emotions, including vocalizations, in real time. This makes it difficult to properly detect and quickly respond to user stress and dissatisfaction. The present invention aims to solve these problems and provide a comfortable riding experience in autonomous vehicles.

[0562] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving voice data and converting the voice data into character data, means for sending the character data to a generative AI model, which analyzes the character data and generates optimal response information, means for analyzing the user's emotions and sending the emotion data to the generative AI model, means for the generative AI model to generate optimal response information in consideration of the emotion data, and means for receiving the generated response information, converting the response information into voice data, and providing it to the user. This makes it possible to analyze emotions from the user's voice in real time and quickly provide an adaptive response.

[0563] "Audio data" means a digital representation of sound waves captured using a microphone or other input device.

[0564] "Character data" is data that indicates a string of characters that is recognized as a unit of language by analyzing voice data.

[0565] A "generative AI model" is an artificial intelligence system designed to analyze input text data and automatically generate appropriate responses and processing.

[0566] "Emotion data" is data that indicates the emotional state recognized from the user's voice data.

[0567] "Operator terminal" refers to equipment used to operate the system, and in particular to the computer and display device used by the operator.

[0568] "Response information" is data that indicates an appropriate answer or suggestion to a user's inquiry or request.

[0569] "Analysis" is the process of examining and evaluating the content and characteristics of input data in detail.

[0570] "Conversion" is the process of replacing data of one format with data of another format.

[0571] A "server" is a central management device for operating the entire system and processing information, and is a computer that processes data and performs communication.

[0572] This invention relates to an emotion recognition driver assistance system for autonomous vehicles. This system has the function of analyzing user voice data and providing optimal responses based on the user's emotions. Specific embodiments are described below.

[0573] First, the vehicle's server receives voice data through a microphone, which is then converted into text data using a voice recognition engine installed in the server.

[0574] The server then sends the converted text data to the generative AI model, which analyzes the user's speech. At the same time, the emotion engine analyzes the voice data and generates the user's emotional data, which is also sent to the generative AI model.

[0575] The generative AI model takes into account the text and emotional data to generate the optimal response information. This response information is sent back to the server and converted into voice data using a speech synthesis engine. This voice data is then provided to the user through the car's speakers.

[0576] The server also sends the generated response information to the operator terminal, which displays multiple answer candidates, allowing the operator to refer to them and make the most appropriate response.

[0577] For example, if a user says, "I'm tired, today was stressful," this voice data is converted into text data and sent to the generative AI model. At the same time, the emotion engine recognizes "stress" and sends it to the generative AI model as emotion data.

[0578] The generative AI model generates the optimal response based on the following prompt:

[0579] User says: I'm tired, today was stressful

[0580] User Emotion: Stress

[0581] Generate the appropriate response for your system:

[0582] The generative AI model generates a response such as, "Thank you for your hard work. I'll play some relaxing music to help you relax." The server converts this response into audio data and provides it to the user through a speaker.

[0583] The hardware used includes a microphone to collect voice data and a speaker to play the audio back to the user, while the software uses a speech recognition engine, an emotion engine, and a generative AI model, which allows for real-time recognition of user emotions and provides adaptive services.

[0584] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0585] Step 1:

[0586] The server receives voice data through the microphone. The input is the user's voice, and the output is digital voice data. Specifically, the microphone inside the car picks up the user's speech, converts it into a digital signal, and sends it to the server.

[0587] Step 2:

[0588] The server uses a speech recognition engine to convert the received voice data into text data. The input is voice data and the output is text data. Specifically, the speech recognition engine analyzes the sound waves and generates the corresponding text in real time.

[0589] Step 3:

[0590] The server sends the converted text data to the generative AI model, which then analyzes the user's speech. The input is text data, and the output is the analysis result. Specifically, the generative AI model understands the text, analyzes the context and intent, and prepares to generate an appropriate response.

[0591] Step 4:

[0592] The server uses an emotion engine to analyze emotion data from voice data. The input is voice data, and the output is emotion data. Specifically, the emotion engine analyzes voice parameters such as tone, tempo, and pitch to estimate the user's emotional state.

[0593] Step 5:

[0594] The server sends the emotion data to the generative AI model, which then generates optimal response information taking the emotion data into consideration. The input is emotion data and text data, and the output is response information. Specifically, the generative AI model generates a prompt sentence based on the text and emotion information, and then creates an appropriate response based on that.

[0595] Step 6:

[0596] The server receives the generated response information and converts it into voice data using a voice synthesis engine. The input is the response information and the output is voice data. Specifically, the voice synthesis engine converts the text into a voice format and generates data that can be played back as a human voice.

[0597] Step 7:

[0598] The server provides the converted voice data to the user through the car's speaker. The input is voice data, and the output is the voice that the user can hear. Specifically, the server sends the voice data to the speaker, and the speaker plays it back.

[0599] Step 8:

[0600] The server sends the generated response information to the operator terminal, which displays multiple answer candidates. The input is the response information, and the output is the displayed answer candidates. The specific operation is that the operator terminal displays the received information on the screen so that the operator can confirm the options.

[0601] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0602] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0603] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0604] [Third embodiment]

[0605] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0606] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0607] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0608] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0609] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0610] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0611] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0612] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0613] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0614] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0615] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0616] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0617] The present invention is embodied in the following manner.

[0618] First, when a user calls the call center, the server detects the incoming call and activates the IVR (Interactive Voice Response) system. In the top menu of this IVR system, the user is prompted to "select an AI operator." When the user presses a button or responds by saying "AI operator," the server activates a speech recognition engine and converts the user's voice data into text data in real time.

[0619] The server then sends the converted text data to the generation AI. The generation AI analyzes the received text data and understands the user's inquiry. Based on this, the generation AI generates optimal response information. This response information is sent back to the server as text data, and the server uses a speech synthesis engine to convert the response information into voice data. This voice data is then sent back to the user.

[0620] Additionally, if an operator needs to intervene, the server sends the generated response information to the operator terminal. The operator terminal displays the information on its screen, allowing the operator to select an appropriate response. In particular, if the user files a complaint, the generation AI proposes the optimal response and displays it on the operator terminal. The operator responds according to the generation AI's proposal, and if necessary, the generation AI generates a follow-up message and sends it to the user via the server.

[0621] In addition, the server accumulates and stores data such as user inquiries and response information. This accumulated data is later categorized and used for analyzing the number of cases and trends. This allows the planning or quality department to refer to the analysis results and use them to improve services and products.

[0622] Specific examples

[0623] For example, if a user says, "My internet connection is slow," the server receives this voice data and converts it into text data using a speech recognition engine. This text data is then sent to a generation AI, which then generates a response that provides the optimal solution to the "slow internet connection," such as "Try restarting your router." The server then converts this response information into voice data and sends it back to the user.

[0624] Additionally, if a complaint occurs, for example, if the content is "The product was not delivered," the generation AI will suggest the optimal response, such as "We are sorry, we will check and contact you as soon as possible." Based on this suggestion displayed on the operator's terminal, the operator will respond quickly, and the generation AI will also generate a follow-up message that will be sent to the user.

[0625] In this way, the system of the present invention can respond to user inquiries quickly and accurately, thereby improving user satisfaction.

[0626] The processing flow will be explained below.

[0627] Step 1:

[0628] A user calls the call center. The server detects the call and activates the IVR system. The IVR system's top menu displays a prompt to "Select an AI operator."

[0629] Step 2:

[0630] The user presses a button or responds by saying "AI operator." The server activates a speech recognition engine, receives the user's voice data, and converts it into text data in real time.

[0631] Step 3:

[0632] The server sends the converted text data to the generation AI, which analyzes the text data and understands the user's inquiry.

[0633] Step 4:

[0634] The AI ​​generates the optimal response information for the user's inquiry, and the generated response information is sent to the server as text data.

[0635] Step 5:

[0636] The server converts the received response information into voice data using a voice synthesis engine, and plays the converted voice data back to the user to provide an appropriate response.

[0637] Step 6:

[0638] If necessary, the server sends the generated response information to the operator terminal, which displays answer candidates on the operator's screen, allowing the operator to select the most appropriate response.

[0639] Step 7:

[0640] When a user files a complaint, the server converts the voice data into text data and sends it to the AI ​​generator, which analyzes the content of the complaint and suggests the best response and wording.

[0641] Step 8:

[0642] The server sends the generated proposal to the operator terminal, which displays it on the screen and allows the operator to respond appropriately to the user.

[0643] Step 9:

[0644] After the complaint is handled, the AI ​​generates a follow-up message, which the server converts into audio data and sends to the user.

[0645] Step 10:

[0646] The server accumulates and stores user inquiries and response information. The accumulated data is categorized and analyzed for number of cases and trends. The analysis results are reported to the planning or quality department and used to improve services.

[0647] The above are the specific processing steps of the system. By taking appropriate measures at each step, user satisfaction can be improved.

[0648] Example 1

[0649] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0650] Conventional call center systems often provided slow or insufficient responses to user inquiries, resulting in low user satisfaction. They also faced the challenge of finding it difficult to quickly provide appropriate solutions for complaints. Furthermore, the accumulation and analysis of inquiry data was insufficient, making it difficult to utilize the data to improve services and products.

[0651] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0652] In this invention, the server includes means for receiving voice data and converting the voice data into text data, means for transmitting the text data to a generation AI, which analyzes the text data and generates optimal response information, means for receiving the generated response information, converting the response information into voice data, and providing it to the user, and means for accumulating, saving, and later analyzing data such as user inquiries and response information. This allows users to receive prompt and accurate responses, and also improves the quality of customer complaints by having the generation AI propose optimal responses. Furthermore, the accumulated data can be used to improve services and products.

[0653] "Audio data" refers to digitally converted data of an audio signal acquired via an audio input device such as a telephone or microphone.

[0654] "Text data" refers to data obtained by analyzing voice data using a voice recognition engine or the like and converting temporal voice signals into text format.

[0655] "Generative AI" is an artificial intelligence model that analyzes received text data and generates optimal response information based on the user's intentions and requests.

[0656] "Response information" refers to information about appropriate responses or solutions generated by the generation AI based on the user's inquiries or requests.

[0657] A "speech synthesis engine" is software that generates natural-sounding speech based on text data or response information.

[0658] "Operator terminal" means a computer system or device used by an operator to respond to user inquiries.

[0659] A "call center system" refers to the entire system for responding to user inquiries and complaints and providing appropriate support and services.

[0660] "Storage and preservation" means continuously recording the data obtained and retaining it for a certain period of time for future reference and analysis.

[0661] A "complaint" is a complaint made by a user reporting dissatisfaction or problems with a product or service, and requesting improvements or responses.

[0662] "Analysis" refers to quantitatively and qualitatively evaluating accumulated data using statistical methods and machine learning to identify areas for improvement and trends in the service.

[0663] This invention relates to a call center system that responds quickly and accurately to inquiries from users. This system includes technology for converting voice data into text data, analyzing it, and generating appropriate responses. Specific embodiments of this system are described below.

[0664] Hardware and Software Configuration

[0665] 1. Hardware

[0666] Server: Processes queries and transforms and stores data.

[0667] Operator terminal: An operator responds to user inquiries.

[0668] User terminal: A device used by a user to make an inquiry, such as a telephone or smartphone.

[0669] 2. Software

[0670] IVR system: Provides voice guidance in response to user inquiries and prompts users to select the appropriate menu.

[0671] Speech recognition engine: Converts voice data into text data in real time (e.g., Google Cloud Speech-to-Text).

[0672] Generative AI model: Analyzes text data and generates appropriate response information (e.g., OpenAI GPT-4).

[0673] Speech synthesis engine: Converts the generated response information into voice data (e.g., Amazon Polly).

[0674] Database: Query and response data is accumulated and stored for later analysis.

[0675] Data processing and calculation

[0676] 1. Converting audio data to text data

[0677] The server uses a speech recognition engine to convert the user's voice input into text data in real time.

[0678] 2. Analysis of character data and generation of response information

[0679] The server sends the converted text data to the generative AI model, which analyzes the data and generates optimal response information.

[0680] 3. Converting response information into voice data

[0681] The server sends the response information received from the generative AI model to the speech synthesis engine, converts it into voice data, and responds to the user.

[0682] 4. Sending information to the operator terminal

[0683] If necessary, the server transmits the generated response information to the operator terminal, and the operator responds appropriately.

[0684] 5. Data accumulation and analysis

[0685] The server accumulates the inquiry and response data and stores it in a database for later analysis, which can lead to improvements in services and products.

[0686] Specific examples

[0687] Slow Internet Connection Inquiries

[0688] When a user calls to complain about a slow internet connection, the server receives this voice data and converts it into text using a speech recognition engine. The converted text data is sent to a generative AI model, which generates a response such as "Please try restarting your router" as an analysis result. The server then converts this response data into voice data using a speech synthesis engine and sends it back to the user.

[0689] Prompt Sentence Examples

[0690] For example, below is an example of a prompt sentence for a generative AI model in response to the query "My internet connection is slow."

[0691] A user has made the following inquiry: "My internet connection is slow." Please suggest the best solution for this problem.

[0692] This allows users to receive prompt and accurate responses, and the quality of responses to complaints can be improved by the generative AI model proposing optimal responses. Furthermore, analyzing the accumulated data can lead to improvements in services and products.

[0693] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0694] Step 1:

[0695] A user calls a call center.

[0696] The user dials the call center number using a landline or smartphone, which initiates the user's inquiry as voice data.

[0697] Step 2:

[0698] The server detects the incoming call and activates the IVR system.

[0699] When the server detects an incoming call, it launches the IVR system, which provides a voice prompt to the user, saying, "Please select an AI operator." The input is the user's phone number, and the output is the launch of the IVR system.

[0700] Step 3:

[0701] The user selects "AI Operator."

[0702] The user responds by saying "AI operator" or selects "AI operator" from the menu, which inputs the voice data.

[0703] Step 4:

[0704] The server uses a speech recognition engine to convert the user's voice into text data.

[0705] The server runs a speech recognition engine and converts the user's voice data into text data in real time. The input is voice data and the output is text data. Specifically, it uses a speech recognition service such as Google Cloud Speech-to-Text.

[0706] Step 5:

[0707] The server sends the character data to the generation AI.

[0708] The server sends the converted text data to the generative AI model. The input is text data, and the output is data sent to the generative AI model. Specifically, the text data is sent through the API.

[0709] Step 6:

[0710] Generative AI analyzes the text data and generates the optimal response.

[0711] The generation AI analyzes the received text data and generates the optimal response information based on the user's inquiry. The input is text data and the output is response information. Specific operations include using OpenAI GPT-4 and other technologies to generate a response based on the prompt text.

[0712] Step 7:

[0713] The server converts the generated response into voice data and sends it back to the user.

[0714] The server sends the response information to a speech synthesis engine and converts it into voice data. The input is the response information and the output is voice data. Specific operations use a speech synthesis service such as Amazon Polly.

[0715] Step 8:

[0716] If operator intervention is required, the server sends response information to the operator terminal.

[0717] If the generating AI determines that operator intervention is necessary, the server sends response information to the operator terminal. The input is the response information, and the output is data transmission to the operator terminal. The specific operation is that the response is displayed on the operator's console display.

[0718] Step 9:

[0719] The server accumulates and stores the query data and response data.

[0720] The server accumulates query data and response data from users and stores them in a database for later analysis. The input is query data and response data, and the output is updating the database. Specifically, data is stored using a database management system.

[0721] keyword

[0722] Generative AI model, prompt sentence

[0723] (Application example 1)

[0724] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0725] User support at call centers is prone to delayed responses, human error, and increased operator workloads. Another issue is the lack of systems and methods for users to receive prompt and appropriate support on the internet and online shopping sites. In particular, it can be difficult to receive prompt and appropriate responses when users submit queries, which can lead to a decline in user satisfaction due to delays in handling complaints. To solve these problems, an efficient support system utilizing speech recognition, generative AI, and speech synthesis is needed.

[0726] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0727] In this invention, the server includes means for receiving voice data and converting the voice data into text data, means for transmitting the text data to a generation AI, which analyzes the text data and generates optimal response information, means for receiving the generated response information and converting the response information into voice data to provide to the user, means for the user to issue a query and generate an optimal answer regarding products on the mail-order site as a response to the query, and means for converting the optimal answer into voice data using a voice synthesis engine and returning it to the user. This allows users to receive quick and appropriate responses in real time, and allows for efficient handling of complaints and the like, thereby improving user satisfaction.

[0728] "Voice data" refers to information transmitted by a user through voice recorded in digital format.

[0729] "Character data" refers to data in text format of the audio content converted using voice recognition technology.

[0730] "Generative AI" refers to artificial intelligence that generates optimal responses and information based on input information.

[0731] "Response information" is text data of answers and suggestions generated by the generation AI based on the analysis results.

[0732] A "speech synthesis engine" is a technology or device that analyzes text data and converts it into speech data.

[0733] A "query" is information sent by a user as an inquiry or question.

[0734] A "user" is a person who sends an inquiry or question to the system.

[0735] An "operator terminal" is a computer or device used by an operator to display the generative AI's suggestions.

[0736] An "online shopping site" is an online sales platform where users can search for product information and make purchases.

[0737] The "optimal answer" is the most appropriate and useful response that the generative AI derives based on the user's query.

[0738] A "complaint" is a complaint or problem report from a user.

[0739] This invention is applicable to cases where a user accesses an online shopping site and sends product-related information or a query. When the user sends a query using a smartphone app, the following system functions.

[0740] First, the server receives voice data from the smartphone app. This voice data is the content of a user's inquiry. Next, the server converts the voice data into text data using a speech recognition engine (speech_recognition library). The converted text data is sent to a generative AI model (e.g., a GPT-based model).

[0741] The generative AI model analyzes the received text data and generates optimal response information. This response information is sent in text format to the server. The server then converts this text-to-speech response information into voice data using a text-to-speech engine (gTTS), and sends the voice data back to the user's smartphone app.

[0742] The specific hardware required is a microphone to capture voice data and a speaker to provide responses to the user.The software requires the speech_recognition library for speech recognition, the transformers library for implementing generative AI, and the gTTS library for speech synthesis.

[0743] For example, if a user asks, "What are the best deals on the latest smartphones?":

[0744] 1. The server receives this voice data and uses a voice recognition engine to convert it into text data such as "Please tell me about discount information on the latest smartphones."

[0745] 2. The generative AI model receives this text data and generates a response such as, "Currently, the latest smartphone models are 10% off."

[0746] 3. The server passes this response information to a speech synthesis engine, converts it into voice data, and responds to the user.

[0747] An example prompt is:

[0748] "Can you tell me about discounts on the latest smartphones?"

[0749] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0750] Step 1:

[0751] A user sends a query using a smartphone app. The user taps the microphone icon in the app and says, "Please tell me about the latest smartphone discount information." The input is the user's voice data, and the output is the voice data sent from the smartphone app to the server.

[0752] Step 2:

[0753] The server receives the voice data sent by the user. Then, the server starts a speech recognition engine (speech_recognition library) and converts this voice data into text data. The input is the user's voice data, and the output is the text data "Please tell me about discount information on the latest smartphones."

[0754] Step 3:

[0755] The server sends the converted text data to the generative AI model, which analyzes the text data and generates the optimal response information. The input is the text data (query), and the output is text data of the response information, such as "Currently, the latest model smartphones are 10% off."

[0756] Step 4:

[0757] The server receives the generated response information and passes this text data to the speech synthesis engine (gTTS). The speech synthesis engine analyzes the text data and converts it into voice data. The input is the text data of the response information, and the output is voice data.

[0758] Step 5:

[0759] The server sends the converted voice data to a smartphone app, and the user receives a response via the smartphone app. The input is voice data, and the output is the voice that the user hears.

[0760] Each processing step is performed sequentially, allowing the user to receive a prompt and appropriate response.

[0761] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0762] The present invention relates to a call center system that analyzes user voice data and incorporates an emotion engine that recognizes emotions. Specific embodiments of the present invention will be described below.

[0763] First, when a user calls the call center, the server detects the call and activates the IVR system. The IVR system's top menu displays a message saying, "Please select an AI operator." When the user presses a button or responds by saying "AI operator," the server activates a speech recognition engine and converts the user's voice data into text data in real time.

[0764] The server then sends the converted text data to the generation AI, which analyzes the text data and understands the user's inquiry. During this process, the emotion engine recognizes emotions from the user's voice data. The recognized emotion data is sent to the generation AI, which then takes this emotional information into account to generate optimal response information.

[0765] The generated response information is sent back to the server as text data. The server then converts this response information into voice data using a speech synthesis engine and provides it to the user. If necessary, the server also sends the generated response information to an operator terminal. The operator terminal displays this information on its screen, allowing the operator to select an appropriate response.

[0766] When a complaint occurs, the server converts the complaint's voice data into text data and sends it to the generation AI. The generation AI analyzes the content of the complaint and suggests the optimal response and appropriate wording based on the emotional data recognized by the emotion engine. These suggestions are displayed on the operator's terminal, and the operator responds accordingly. After the complaint is handled, the generation AI creates a follow-up message and sends it to the user via the server.

[0767] The server also accumulates and stores user inquiries, response information, and recognized emotion data. The accumulated data is later categorized and used to analyze the number of cases and trends. This allows the planning or quality department to refer to the analysis results and use them to improve services and products.

[0768] Specific examples

[0769] For example, if a user says, "My internet connection is slow," the server receives this voice data and converts it into text data using a speech recognition engine. This text data is then sent to the generation AI, which then generates a response that optimally addresses the "slow internet connection," such as "Please try restarting your router." At the same time, if the emotion engine detects "dissatisfaction" in the user's voice, the generation AI will also include an additional apology, such as "We apologize for the inconvenience."

[0770] Additionally, if a complaint is made that the product was not delivered, the generation AI will suggest the optimal response, such as "We are sorry, we will check and contact you as soon as possible." At the same time, if the emotion engine detects "anger," the generation AI will add a response that is more in tune with the emotion, such as "We are very sorry for the inconvenience." Based on the suggestions displayed on the operator terminal, the operator will respond quickly, and a generated follow-up message will be sent to the user later.

[0771] In this way, the system of the present invention can recognize the user's emotions and respond optimally to them, thereby increasing user satisfaction.

[0772] The processing flow will be explained below.

[0773] Step 1:

[0774] A user calls the call center. The server detects the call and activates the IVR system. The IVR system's top menu displays a message saying, "Please select an AI operator."

[0775] Step 2:

[0776] The user presses a button or responds by saying "AI operator." The server activates a speech recognition engine, receives the user's voice data, and converts it into text data in real time.

[0777] Step 3:

[0778] The server sends the converted text data to the generation AI, which analyzes the text data and understands the user's inquiry.

[0779] Step 4:

[0780] The emotion engine recognizes emotions from the user's voice data and sends the recognized emotion data to the generation AI.

[0781] Step 5:

[0782] The AI ​​generates the optimal response information based on the user's inquiry and emotional data. The generated response information is sent to the server as text data.

[0783] Step 6:

[0784] The server converts the received response information into voice data using a voice synthesis engine, and plays the converted voice data back to the user to provide an appropriate response.

[0785] Step 7:

[0786] If necessary, the server sends the generated response information to the operator terminal, which displays answer candidates and emotion data on the operator's screen, allowing the operator to select the optimal response.

[0787] Step 8:

[0788] When a user files a complaint, the server converts the voice data into text data and sends it to the generation AI, which analyzes the content of the complaint and suggests the best response and appropriate wording based on the emotional data recognized by the emotion engine.

[0789] Step 9:

[0790] The server sends the generated proposal to the operator terminal, which displays the proposal and emotion data on the screen, allowing the operator to respond appropriately to the user.

[0791] Step 10:

[0792] After the complaint is handled, the AI ​​generates a follow-up message, which the server converts into audio data and sends to the user.

[0793] Step 11:

[0794] The server accumulates and stores user inquiries, response information, emotional data, etc. The accumulated data is categorized and used to analyze the number of cases and trends. The results of the analysis are reported to the planning or quality department and used to improve services.

[0795] The above are the specific processing steps of this system. By taking appropriate measures at each step, user satisfaction can be improved.

[0796] Example 2

[0797] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0798] Conventional call center systems have difficulty accurately understanding users' emotions and requests and providing prompt and optimal responses. Furthermore, when handling complaints, it is difficult for operators to immediately find the optimal response, which can lead to lower user satisfaction. To solve these issues, an advanced system that can analyze users' emotional data and generate optimal responses is needed.

[0799] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0800] In this invention, the server includes means for receiving voice data and converting the voice data into character data, means for transmitting the character data to an artificial intelligence (AI) to be analyzed by the AI ​​and generating optimal response information, means for recognizing emotion data from the voice data by an emotion analysis means and transmitting the emotion data to the AI, and means for receiving the generated response information, converting the response information into voice data, and providing it to the user, thereby enabling the generation and provision of optimal responses taking the user's emotions into consideration.

[0801] "Voice data" is data that digitally represents the voice that the user utters to the system.

[0802] "Character data" is data in text format that has been converted from voice data into character information.

[0803] "Artificial intelligence" is a system that uses machine learning and data analysis to automatically perform specific tasks.

[0804] "Emotion analysis means" is a technology that recognizes the user's emotions from voice data and extracts emotion data.

[0805] "Response information" is information in response to a user's inquiry that is generated as a result of analysis by artificial intelligence.

[0806] "Operation device" refers to a terminal or computer used by an operator, and is a device that displays information necessary for responding to users.

[0807] The "optimal response" is the most appropriate response method that is generated taking into consideration the user's requests and feelings.

[0808] The present invention relates to a call center system that analyzes user voice data and incorporates an emotion engine that recognizes emotions. Specific embodiments of the present invention will be described below.

[0809] First, when a user calls the call center, the server detects the call. The server then activates the IVR (Interactive Voice Response) system and plays a message saying, "Please select an AI operator." This message is generated using a text-to-speech engine (e.g., Amazon Polly).

[0810] Next, when the user presses a button or responds by saying "AI operator," the server recognizes this and activates a speech recognition engine (for example, Google Cloud Speech-to-Text API). The server receives the voice data and converts it into text data in real time using the speech recognition engine.

[0811] The converted text data is sent from the server to a generation AI (e.g., OpenAI GPT-3). The generation AI analyzes the text data and understands the user's inquiry. At the same time, an emotion analysis means (e.g., IBM Watson Tone Analyzer) analyzes the user's voice data and recognizes emotional data. This emotional data is also sent to the generation AI.

[0812] The AI ​​generates optimal responses based on the user's query and emotional data. The server receives the responses and converts them into voice data using a speech synthesis engine (e.g., Amazon Polly) to provide them to the user.

[0813] The server also sends the generated response information to the operator terminal. The operator terminal displays this response information on the screen, allowing the operator to select a response. In particular, if the user's request is a complaint, the generation AI proposes the optimal response and displays it on the operator terminal. The operator then takes appropriate action based on the proposal.

[0814] For example, consider the case where a user inquires about a slow internet connection. The server converts the received voice data into text data using a speech recognition engine, and then sends that text data to the generation AI. The generation AI generates a response that optimally addresses the slow internet connection issue, such as "Please try restarting your router." At the same time, if the emotion engine detects "dissatisfaction" in the user's voice, the generation AI also adds an apology, such as "We apologize for the inconvenience."

[0815] As another example, consider the case of a complaint that "the product was not delivered." In this case, the generation AI will similarly generate the optimal response, such as "We are sorry, we will check and contact you as soon as possible." Furthermore, if the emotion engine detects "anger," the generation AI will also add a response that is more in tune with the emotion, such as "We are very sorry for the inconvenience." The operator will respond promptly based on the suggestions displayed on the operator terminal.

[0816] In this way, the system of the present invention can increase user satisfaction by recognizing the user's emotions and providing the optimal response. The server also accumulates and saves the inquiry content, response information, and recognized emotion data in a database for later analysis and service improvement.

[0817] An example of a prompt sentence is "Please enter your query." For example, by entering "My Internet connection is slow," the system will generate an appropriate answer.

[0818] This invention makes it possible to realize effective call center response that takes into account the user's emotions.

[0819] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0820] Step 1:

[0821] When a user calls the call center, the server detects the call. Specifically, the server uses the VoIP system to trigger an incoming call event for a specific phone number. Based on the incoming call data (input), the server activates the IVR system (output).

[0822] Step 2:

[0823] The server starts the IVR system and plays a message to the user saying, "Please select an AI operator." This message is generated by the server using a text-to-speech engine (e.g., Amazon Polly), which converts text data (input) into voice data (output).

[0824] Step 3:

[0825] When the user presses a button or responds by saying "AI operator," the server recognizes this and starts a speech recognition engine (for example, Google Cloud Speech-to-Text API), which prepares the system to convert the voice data (input) into text data (output).

[0826] Step 4:

[0827] The server receives the user's voice data and converts it into text data in real time using a speech recognition engine. Specifically, the server sends the voice data to the Google Cloud Speech-to-Text API and receives the converted text data. The voice data (input) is converted into text data (output).

[0828] Step 5:

[0829] The server sends the acquired character data to the generation AI (e.g., OpenAI GPT-3). The server sends the character data (input) to the generation AI's API and receives the analysis result, which is the response information (output).

[0830] Step 6:

[0831] The generation AI analyzes the text data and understands the user's inquiry. For example, in response to the text data (input) "My internet connection is slow," it generates the appropriate solution, "Try restarting your router."

[0832] Step 7:

[0833] An emotion analysis means (e.g., IBM Watson Tone Analyzer) analyzes the user's voice data and recognizes the emotion data. The server sends the voice data (input) to the emotion analysis API and sends the recognized emotion data (output) to the generation AI.

[0834] Step 8:

[0835] The generation AI generates optimal response information by taking into account the content of the user's inquiry and emotional data. If the emotional data indicates "dissatisfaction," the generation AI adds an apology such as "We apologize for the inconvenience." Based on the text data and emotional data (input), the AI ​​generates comprehensive response information (output).

[0836] Step 9:

[0837] The server receives the generated response information and converts it into voice data using a speech synthesis engine (e.g., Amazon Polly). Text data (input) is converted into voice data (output).

[0838] Step 10:

[0839] The server provides the converted voice data to the user, specifically by playing the voice data (input) to the user's telephone terminal (output).

[0840] Step 11:

[0841] If necessary, the server sends the generated response information to the operator terminal. Character data (input) is sent to the operation device and displayed (output).

[0842] Step 12:

[0843] When a complaint occurs, the server converts the complaint voice data into text data and sends it to the generation AI. The generation AI analyzes the content of the complaint and proposes the optimal response. This is displayed on the operator terminal, allowing the operator to select a response. The voice data (input) is converted into text data and emotional data (output), and response information is generated.

[0844] Step 13:

[0845] After handling the complaint, the generation AI generates a follow-up message and sends it to the user via the server. It sends text data (input) and provides the user with follow-up message data (output).

[0846] Step 14:

[0847] The server accumulates and stores the user's inquiry, response information, and recognized emotion data in a database. This helps to increase user satisfaction and improve services. Text and emotion data (input) are stored in the database (output).

[0848] (Application example 2)

[0849] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0850] Conventional driver assistance systems for autonomous vehicles rely primarily on visual analysis, such as the user's facial expressions and body movements, and lack the ability to analyze emotions, including vocalizations, in real time. This makes it difficult to properly detect and quickly respond to user stress and dissatisfaction. The present invention aims to solve these problems and provide a comfortable riding experience in autonomous vehicles.

[0851] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving voice data and converting the voice data into character data, means for sending the character data to a generative AI model, which analyzes the character data and generates optimal response information, means for analyzing the user's emotions and sending the emotion data to the generative AI model, means for the generative AI model to generate optimal response information in consideration of the emotion data, and means for receiving the generated response information, converting the response information into voice data, and providing it to the user. This makes it possible to analyze emotions from the user's voice in real time and quickly provide an adaptive response.

[0852] "Audio data" means a digital representation of sound waves captured using a microphone or other input device.

[0853] "Character data" is data that indicates a string of characters that is recognized as a unit of language by analyzing voice data.

[0854] A "generative AI model" is an artificial intelligence system designed to analyze input text data and automatically generate appropriate responses and processing.

[0855] "Emotion data" is data that indicates the emotional state recognized from the user's voice data.

[0856] "Operator terminal" refers to equipment used to operate the system, and in particular to the computer and display device used by the operator.

[0857] "Response information" is data that indicates an appropriate answer or suggestion to a user's inquiry or request.

[0858] "Analysis" is the process of examining and evaluating the content and characteristics of input data in detail.

[0859] "Conversion" is the process of replacing data of one format with data of another format.

[0860] A "server" is a central management device for operating the entire system and processing information, and is a computer that processes data and performs communication.

[0861] This invention relates to an emotion recognition driver assistance system for autonomous vehicles. This system has the function of analyzing user voice data and providing optimal responses based on the user's emotions. Specific embodiments are described below.

[0862] First, the vehicle's server receives voice data through a microphone, which is then converted into text data using a voice recognition engine installed in the server.

[0863] The server then sends the converted text data to the generative AI model, which analyzes the user's speech. At the same time, the emotion engine analyzes the voice data and generates the user's emotional data, which is also sent to the generative AI model.

[0864] The generative AI model takes into account the text and emotional data to generate the optimal response information. This response information is sent back to the server and converted into voice data using a speech synthesis engine. This voice data is then provided to the user through the car's speakers.

[0865] The server also sends the generated response information to the operator terminal, which displays multiple answer candidates, allowing the operator to refer to them and make the most appropriate response.

[0866] For example, if a user says, "I'm tired, today was stressful," this voice data is converted into text data and sent to the generative AI model. At the same time, the emotion engine recognizes "stress" and sends it to the generative AI model as emotion data.

[0867] The generative AI model generates the optimal response based on the following prompt:

[0868] User says: I'm tired, today was stressful

[0869] User Emotion: Stress

[0870] Generate the appropriate response for your system:

[0871] The generative AI model generates a response such as, "Thank you for your hard work. I'll play some relaxing music to help you relax." The server converts this response into audio data and provides it to the user through a speaker.

[0872] The hardware used includes a microphone to collect voice data and a speaker to play the audio back to the user, while the software uses a speech recognition engine, an emotion engine, and a generative AI model, which allows for real-time recognition of user emotions and provides adaptive services.

[0873] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0874] Step 1:

[0875] The server receives voice data through the microphone. The input is the user's voice, and the output is digital voice data. Specifically, the microphone inside the car picks up the user's speech, converts it into a digital signal, and sends it to the server.

[0876] Step 2:

[0877] The server uses a speech recognition engine to convert the received voice data into text data. The input is voice data and the output is text data. Specifically, the speech recognition engine analyzes the sound waves and generates the corresponding text in real time.

[0878] Step 3:

[0879] The server sends the converted text data to the generative AI model, which then analyzes the user's speech. The input is text data, and the output is the analysis result. Specifically, the generative AI model understands the text, analyzes the context and intent, and prepares to generate an appropriate response.

[0880] Step 4:

[0881] The server uses an emotion engine to analyze emotion data from voice data. The input is voice data, and the output is emotion data. Specifically, the emotion engine analyzes voice parameters such as tone, tempo, and pitch to estimate the user's emotional state.

[0882] Step 5:

[0883] The server sends the emotion data to the generative AI model, which then generates optimal response information taking the emotion data into consideration. The input is emotion data and text data, and the output is response information. Specifically, the generative AI model generates a prompt sentence based on the text and emotion information, and then creates an appropriate response based on that.

[0884] Step 6:

[0885] The server receives the generated response information and converts it into voice data using a voice synthesis engine. The input is the response information and the output is voice data. Specifically, the voice synthesis engine converts the text into a voice format and generates data that can be played back as a human voice.

[0886] Step 7:

[0887] The server provides the converted voice data to the user through the car's speaker. The input is voice data, and the output is the voice that the user can hear. Specifically, the server sends the voice data to the speaker, and the speaker plays it back.

[0888] Step 8:

[0889] The server sends the generated response information to the operator terminal, which displays multiple answer candidates. The input is the response information, and the output is the displayed answer candidates. The specific operation is that the operator terminal displays the received information on the screen so that the operator can confirm the options.

[0890] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0891] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0892] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[0893] [Fourth embodiment]

[0894] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[0895] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0896] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0897] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[0898] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0899] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0900] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0901] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[0902] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0903] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0904] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0905] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0906] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[0907] The present invention is embodied in the following manner.

[0908] First, when a user calls the call center, the server detects the incoming call and activates the IVR (Interactive Voice Response) system. In the top menu of this IVR system, the user is prompted to "select an AI operator." When the user presses a button or responds by saying "AI operator," the server activates a speech recognition engine and converts the user's voice data into text data in real time.

[0909] The server then sends the converted text data to the generation AI. The generation AI analyzes the received text data and understands the user's inquiry. Based on this, the generation AI generates optimal response information. This response information is sent back to the server as text data, and the server uses a speech synthesis engine to convert the response information into voice data. This voice data is then sent back to the user.

[0910] Additionally, if an operator needs to intervene, the server sends the generated response information to the operator terminal. The operator terminal displays the information on its screen, allowing the operator to select an appropriate response. In particular, if the user files a complaint, the generation AI proposes the optimal response and displays it on the operator terminal. The operator responds according to the generation AI's proposal, and if necessary, the generation AI generates a follow-up message and sends it to the user via the server.

[0911] In addition, the server accumulates and stores data such as user inquiries and response information. This accumulated data is later categorized and used for analyzing the number of cases and trends. This allows the planning or quality department to refer to the analysis results and use them to improve services and products.

[0912] Specific examples

[0913] For example, if a user says, "My internet connection is slow," the server receives this voice data and converts it into text data using a speech recognition engine. This text data is then sent to a generation AI, which then generates a response that provides the optimal solution to the "slow internet connection," such as "Try restarting your router." The server then converts this response information into voice data and sends it back to the user.

[0914] Additionally, if a complaint occurs, for example, if the content is "The product was not delivered," the generation AI will suggest the optimal response, such as "We are sorry, we will check and contact you as soon as possible." Based on this suggestion displayed on the operator's terminal, the operator will respond quickly, and the generation AI will also generate a follow-up message that will be sent to the user.

[0915] In this way, the system of the present invention can respond to user inquiries quickly and accurately, thereby improving user satisfaction.

[0916] The processing flow will be explained below.

[0917] Step 1:

[0918] A user calls the call center. The server detects the call and activates the IVR system. The IVR system's top menu displays a prompt to "Select an AI operator."

[0919] Step 2:

[0920] The user presses a button or responds by saying "AI operator." The server activates a speech recognition engine, receives the user's voice data, and converts it into text data in real time.

[0921] Step 3:

[0922] The server sends the converted text data to the generation AI, which analyzes the text data and understands the user's inquiry.

[0923] Step 4:

[0924] The AI ​​generates the optimal response information for the user's inquiry, and the generated response information is sent to the server as text data.

[0925] Step 5:

[0926] The server converts the received response information into voice data using a voice synthesis engine, and plays the converted voice data back to the user to provide an appropriate response.

[0927] Step 6:

[0928] If necessary, the server sends the generated response information to the operator terminal, which displays answer candidates on the operator's screen, allowing the operator to select the most appropriate response.

[0929] Step 7:

[0930] When a user files a complaint, the server converts the voice data into text data and sends it to the AI ​​generator, which analyzes the content of the complaint and suggests the best response and wording.

[0931] Step 8:

[0932] The server sends the generated proposal to the operator terminal, which displays it on the screen and allows the operator to respond appropriately to the user.

[0933] Step 9:

[0934] After the complaint is handled, the AI ​​generates a follow-up message, which the server converts into audio data and sends to the user.

[0935] Step 10:

[0936] The server accumulates and stores user inquiries and response information. The accumulated data is categorized and analyzed for number of cases and trends. The analysis results are reported to the planning or quality department and used to improve services.

[0937] The above are the specific processing steps of the system. By taking appropriate measures at each step, user satisfaction can be improved.

[0938] Example 1

[0939] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[0940] Conventional call center systems often provided slow or insufficient responses to user inquiries, resulting in low user satisfaction. They also faced the challenge of finding it difficult to quickly provide appropriate solutions for complaints. Furthermore, the accumulation and analysis of inquiry data was insufficient, making it difficult to utilize the data to improve services and products.

[0941] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0942] In this invention, the server includes means for receiving voice data and converting the voice data into text data, means for transmitting the text data to a generation AI, which analyzes the text data and generates optimal response information, means for receiving the generated response information, converting the response information into voice data, and providing it to the user, and means for accumulating, saving, and later analyzing data such as user inquiries and response information. This allows users to receive prompt and accurate responses, and also improves the quality of customer complaints by having the generation AI propose optimal responses. Furthermore, the accumulated data can be used to improve services and products.

[0943] "Audio data" refers to digitally converted data of an audio signal acquired via an audio input device such as a telephone or microphone.

[0944] "Text data" refers to data obtained by analyzing voice data using a voice recognition engine or the like and converting temporal voice signals into text format.

[0945] "Generative AI" is an artificial intelligence model that analyzes received text data and generates optimal response information based on the user's intentions and requests.

[0946] "Response information" refers to information about appropriate responses or solutions generated by the generation AI based on the user's inquiries or requests.

[0947] A "speech synthesis engine" is software that generates natural-sounding speech based on text data or response information.

[0948] "Operator terminal" means a computer system or device used by an operator to respond to user inquiries.

[0949] A "call center system" refers to the entire system for responding to user inquiries and complaints and providing appropriate support and services.

[0950] "Storage and preservation" means continuously recording the data obtained and retaining it for a certain period of time for future reference and analysis.

[0951] A "complaint" is a complaint made by a user reporting dissatisfaction or problems with a product or service, and requesting improvements or responses.

[0952] "Analysis" refers to quantitatively and qualitatively evaluating accumulated data using statistical methods and machine learning to identify areas for improvement and trends in the service.

[0953] This invention relates to a call center system that responds quickly and accurately to inquiries from users. This system includes technology for converting voice data into text data, analyzing it, and generating appropriate responses. Specific embodiments of this system are described below.

[0954] Hardware and Software Configuration

[0955] 1. Hardware

[0956] Server: Processes queries and transforms and stores data.

[0957] Operator terminal: An operator responds to user inquiries.

[0958] User terminal: A device used by a user to make an inquiry, such as a telephone or smartphone.

[0959] 2. Software

[0960] IVR system: Provides voice guidance in response to user inquiries and prompts users to select the appropriate menu.

[0961] Speech recognition engine: Converts voice data into text data in real time (e.g., Google Cloud Speech-to-Text).

[0962] Generative AI model: Analyzes text data and generates appropriate response information (e.g., OpenAI GPT-4).

[0963] Speech synthesis engine: Converts the generated response information into voice data (e.g., Amazon Polly).

[0964] Database: Query and response data is accumulated and stored for later analysis.

[0965] Data processing and calculation

[0966] 1. Converting audio data to text data

[0967] The server uses a speech recognition engine to convert the user's voice input into text data in real time.

[0968] 2. Analysis of character data and generation of response information

[0969] The server sends the converted text data to the generative AI model, which analyzes the data and generates optimal response information.

[0970] 3. Converting response information into voice data

[0971] The server sends the response information received from the generative AI model to the speech synthesis engine, converts it into voice data, and responds to the user.

[0972] 4. Sending information to the operator terminal

[0973] If necessary, the server transmits the generated response information to the operator terminal, and the operator responds appropriately.

[0974] 5. Data accumulation and analysis

[0975] The server accumulates the inquiry and response data and stores it in a database for later analysis, which can lead to improvements in services and products.

[0976] Specific examples

[0977] Slow Internet Connection Inquiries

[0978] When a user calls to complain about a slow internet connection, the server receives this voice data and converts it into text using a speech recognition engine. The converted text data is sent to a generative AI model, which generates a response such as "Please try restarting your router" as an analysis result. The server then converts this response data into voice data using a speech synthesis engine and sends it back to the user.

[0979] Prompt Sentence Examples

[0980] For example, below is an example of a prompt sentence for a generative AI model in response to the query "My internet connection is slow."

[0981] A user has made the following inquiry: "My internet connection is slow." Please suggest the best solution for this problem.

[0982] This allows users to receive prompt and accurate responses, and the quality of responses to complaints can be improved by the generative AI model proposing optimal responses. Furthermore, analyzing the accumulated data can lead to improvements in services and products.

[0983] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0984] Step 1:

[0985] A user calls a call center.

[0986] The user dials the call center number using a landline or smartphone, which initiates the user's inquiry as voice data.

[0987] Step 2:

[0988] The server detects the incoming call and activates the IVR system.

[0989] When the server detects an incoming call, it launches the IVR system, which provides a voice prompt to the user, saying, "Please select an AI operator." The input is the user's phone number, and the output is the launch of the IVR system.

[0990] Step 3:

[0991] The user selects "AI Operator."

[0992] The user responds by saying "AI operator" or selects "AI operator" from the menu, which inputs the voice data.

[0993] Step 4:

[0994] The server uses a speech recognition engine to convert the user's voice into text data.

[0995] The server runs a speech recognition engine and converts the user's voice data into text data in real time. The input is voice data and the output is text data. Specifically, it uses a speech recognition service such as Google Cloud Speech-to-Text.

[0996] Step 5:

[0997] The server sends the character data to the generation AI.

[0998] The server sends the converted text data to the generative AI model. The input is text data, and the output is data sent to the generative AI model. Specifically, the text data is sent through the API.

[0999] Step 6:

[1000] Generative AI analyzes the text data and generates the optimal response.

[1001] The generation AI analyzes the received text data and generates the optimal response information based on the user's inquiry. The input is text data and the output is response information. Specific operations include using OpenAI GPT-4 and other technologies to generate a response based on the prompt text.

[1002] Step 7:

[1003] The server converts the generated response into voice data and sends it back to the user.

[1004] The server sends the response information to a speech synthesis engine and converts it into voice data. The input is the response information and the output is voice data. Specific operations use a speech synthesis service such as Amazon Polly.

[1005] Step 8:

[1006] If operator intervention is required, the server sends response information to the operator terminal.

[1007] If the generating AI determines that operator intervention is necessary, the server sends response information to the operator terminal. The input is the response information, and the output is data transmission to the operator terminal. The specific operation is that the response is displayed on the operator's console display.

[1008] Step 9:

[1009] The server accumulates and stores the query data and response data.

[1010] The server accumulates query data and response data from users and stores them in a database for later analysis. The input is query data and response data, and the output is updating the database. Specifically, data is stored using a database management system.

[1011] keyword

[1012] Generative AI model, prompt sentence

[1013] (Application example 1)

[1014] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1015] User support at call centers is prone to delayed responses, human error, and increased operator workloads. Another issue is the lack of systems and methods for users to receive prompt and appropriate support on the internet and online shopping sites. In particular, it can be difficult to receive prompt and appropriate responses when users submit queries, which can lead to a decline in user satisfaction due to delays in handling complaints. To solve these problems, an efficient support system utilizing speech recognition, generative AI, and speech synthesis is needed.

[1016] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1017] In this invention, the server includes means for receiving voice data and converting the voice data into text data, means for transmitting the text data to a generation AI, which analyzes the text data and generates optimal response information, means for receiving the generated response information and converting the response information into voice data to provide to the user, means for the user to issue a query and generate an optimal answer regarding products on the mail-order site as a response to the query, and means for converting the optimal answer into voice data using a voice synthesis engine and returning it to the user. This allows users to receive quick and appropriate responses in real time, and allows for efficient handling of complaints and the like, thereby improving user satisfaction.

[1018] "Voice data" refers to information transmitted by a user through voice recorded in digital format.

[1019] "Character data" refers to data in text format of the audio content converted using voice recognition technology.

[1020] "Generative AI" refers to artificial intelligence that generates optimal responses and information based on input information.

[1021] "Response information" is text data of answers and suggestions generated by the generation AI based on the analysis results.

[1022] A "speech synthesis engine" is a technology or device that analyzes text data and converts it into speech data.

[1023] A "query" is information sent by a user as an inquiry or question.

[1024] A "user" is a person who sends an inquiry or question to the system.

[1025] An "operator terminal" is a computer or device used by an operator to display the generative AI's suggestions.

[1026] An "online shopping site" is an online sales platform where users can search for product information and make purchases.

[1027] The "optimal answer" is the most appropriate and useful response that the generative AI derives based on the user's query.

[1028] A "complaint" is a complaint or problem report from a user.

[1029] This invention is applicable to cases where a user accesses an online shopping site and sends product-related information or a query. When the user sends a query using a smartphone app, the following system functions.

[1030] First, the server receives voice data from the smartphone app. This voice data is the content of a user's inquiry. Next, the server converts the voice data into text data using a speech recognition engine (speech_recognition library). The converted text data is sent to a generative AI model (e.g., a GPT-based model).

[1031] The generative AI model analyzes the received text data and generates optimal response information. This response information is sent in text format to the server. The server then converts this text-to-speech response information into voice data using a text-to-speech engine (gTTS), and sends the voice data back to the user's smartphone app.

[1032] The specific hardware required is a microphone to capture voice data and a speaker to provide responses to the user.The software requires the speech_recognition library for speech recognition, the transformers library for implementing generative AI, and the gTTS library for speech synthesis.

[1033] For example, if a user asks, "What are the best deals on the latest smartphones?":

[1034] 1. The server receives this voice data and uses a voice recognition engine to convert it into text data such as "Please tell me about discount information on the latest smartphones."

[1035] 2. The generative AI model receives this text data and generates a response such as, "Currently, the latest smartphone models are 10% off."

[1036] 3. The server passes this response information to a speech synthesis engine, converts it into voice data, and responds to the user.

[1037] An example prompt is:

[1038] "Can you tell me about discounts on the latest smartphones?"

[1039] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1040] Step 1:

[1041] A user sends a query using a smartphone app. The user taps the microphone icon in the app and says, "Please tell me about the latest smartphone discount information." The input is the user's voice data, and the output is the voice data sent from the smartphone app to the server.

[1042] Step 2:

[1043] The server receives the voice data sent by the user. Then, the server starts a speech recognition engine (speech_recognition library) and converts this voice data into text data. The input is the user's voice data, and the output is the text data "Please tell me about discount information on the latest smartphones."

[1044] Step 3:

[1045] The server sends the converted text data to the generative AI model, which analyzes the text data and generates the optimal response information. The input is the text data (query), and the output is text data of the response information, such as "Currently, the latest model smartphones are 10% off."

[1046] Step 4:

[1047] The server receives the generated response information and passes this text data to the speech synthesis engine (gTTS). The speech synthesis engine analyzes the text data and converts it into voice data. The input is the text data of the response information, and the output is voice data.

[1048] Step 5:

[1049] The server sends the converted voice data to a smartphone app, and the user receives a response via the smartphone app. The input is voice data, and the output is the voice that the user hears.

[1050] Each processing step is performed sequentially, allowing the user to receive a prompt and appropriate response.

[1051] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1052] The present invention relates to a call center system that analyzes user voice data and incorporates an emotion engine that recognizes emotions. Specific embodiments of the present invention will be described below.

[1053] First, when a user calls the call center, the server detects the call and activates the IVR system. The IVR system's top menu displays a message saying, "Please select an AI operator." When the user presses a button or responds by saying "AI operator," the server activates a speech recognition engine and converts the user's voice data into text data in real time.

[1054] The server then sends the converted text data to the generation AI, which analyzes the text data and understands the user's inquiry. During this process, the emotion engine recognizes emotions from the user's voice data. The recognized emotion data is sent to the generation AI, which then takes this emotional information into account to generate optimal response information.

[1055] The generated response information is sent back to the server as text data. The server then converts this response information into voice data using a speech synthesis engine and provides it to the user. If necessary, the server also sends the generated response information to an operator terminal. The operator terminal displays this information on its screen, allowing the operator to select an appropriate response.

[1056] When a complaint occurs, the server converts the complaint's voice data into text data and sends it to the generation AI. The generation AI analyzes the content of the complaint and suggests the optimal response and appropriate wording based on the emotional data recognized by the emotion engine. These suggestions are displayed on the operator's terminal, and the operator responds accordingly. After the complaint is handled, the generation AI creates a follow-up message and sends it to the user via the server.

[1057] The server also accumulates and stores user inquiries, response information, and recognized emotion data. The accumulated data is later categorized and used to analyze the number of cases and trends. This allows the planning or quality department to refer to the analysis results and use them to improve services and products.

[1058] Specific examples

[1059] For example, if a user says, "My internet connection is slow," the server receives this voice data and converts it into text data using a speech recognition engine. This text data is then sent to the generation AI, which then generates a response that optimally addresses the "slow internet connection," such as "Please try restarting your router." At the same time, if the emotion engine detects "dissatisfaction" in the user's voice, the generation AI will also include an additional apology, such as "We apologize for the inconvenience."

[1060] Additionally, if a complaint is made that the product was not delivered, the generation AI will suggest the optimal response, such as "We are sorry, we will check and contact you as soon as possible." At the same time, if the emotion engine detects "anger," the generation AI will add a response that is more in tune with the emotion, such as "We are very sorry for the inconvenience." Based on the suggestions displayed on the operator terminal, the operator will respond quickly, and a generated follow-up message will be sent to the user later.

[1061] In this way, the system of the present invention can recognize the user's emotions and respond optimally to them, thereby increasing user satisfaction.

[1062] The processing flow will be explained below.

[1063] Step 1:

[1064] A user calls the call center. The server detects the call and activates the IVR system. The IVR system's top menu displays a message saying, "Please select an AI operator."

[1065] Step 2:

[1066] The user presses a button or responds by saying "AI operator." The server activates a speech recognition engine, receives the user's voice data, and converts it into text data in real time.

[1067] Step 3:

[1068] The server sends the converted text data to the generation AI, which analyzes the text data and understands the user's inquiry.

[1069] Step 4:

[1070] The emotion engine recognizes emotions from the user's voice data and sends the recognized emotion data to the generation AI.

[1071] Step 5:

[1072] The AI ​​generates the optimal response information based on the user's inquiry and emotional data. The generated response information is sent to the server as text data.

[1073] Step 6:

[1074] The server converts the received response information into voice data using a voice synthesis engine, and plays the converted voice data back to the user to provide an appropriate response.

[1075] Step 7:

[1076] If necessary, the server sends the generated response information to the operator terminal, which displays answer candidates and emotion data on the operator's screen, allowing the operator to select the optimal response.

[1077] Step 8:

[1078] When a user files a complaint, the server converts the voice data into text data and sends it to the generation AI, which analyzes the content of the complaint and suggests the best response and appropriate wording based on the emotional data recognized by the emotion engine.

[1079] Step 9:

[1080] The server sends the generated proposal to the operator terminal, which displays the proposal and emotion data on the screen, allowing the operator to respond appropriately to the user.

[1081] Step 10:

[1082] After the complaint is handled, the AI ​​generates a follow-up message, which the server converts into audio data and sends to the user.

[1083] Step 11:

[1084] The server accumulates and stores user inquiries, response information, emotional data, etc. The accumulated data is categorized and used to analyze the number of cases and trends. The results of the analysis are reported to the planning or quality department and used to improve services.

[1085] The above are the specific processing steps of this system. By taking appropriate measures at each step, user satisfaction can be improved.

[1086] Example 2

[1087] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1088] Conventional call center systems have difficulty accurately understanding users' emotions and requests and providing prompt and optimal responses. Furthermore, when handling complaints, it is difficult for operators to immediately find the optimal response, which can lead to lower user satisfaction. To solve these issues, an advanced system that can analyze users' emotional data and generate optimal responses is needed.

[1089] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1090] In this invention, the server includes means for receiving voice data and converting the voice data into character data, means for transmitting the character data to an artificial intelligence (AI) to be analyzed by the AI ​​and generating optimal response information, means for recognizing emotion data from the voice data by an emotion analysis means and transmitting the emotion data to the AI, and means for receiving the generated response information, converting the response information into voice data, and providing it to the user, thereby enabling the generation and provision of optimal responses taking the user's emotions into consideration.

[1091] "Voice data" is data that digitally represents the voice that the user utters to the system.

[1092] "Character data" is data in text format that has been converted from voice data into character information.

[1093] "Artificial intelligence" is a system that uses machine learning and data analysis to automatically perform specific tasks.

[1094] "Emotion analysis means" is a technology that recognizes the user's emotions from voice data and extracts emotion data.

[1095] "Response information" is information in response to a user's inquiry that is generated as a result of analysis by artificial intelligence.

[1096] "Operation device" refers to a terminal or computer used by an operator, and is a device that displays information necessary for responding to users.

[1097] The "optimal response" is the most appropriate response method that is generated taking into consideration the user's requests and feelings.

[1098] The present invention relates to a call center system that analyzes user voice data and incorporates an emotion engine that recognizes emotions. Specific embodiments of the present invention will be described below.

[1099] First, when a user calls the call center, the server detects the call. The server then activates the IVR (Interactive Voice Response) system and plays a message saying, "Please select an AI operator." This message is generated using a text-to-speech engine (e.g., Amazon Polly).

[1100] Next, when the user presses a button or responds by saying "AI operator," the server recognizes this and activates a speech recognition engine (for example, Google Cloud Speech-to-Text API). The server receives the voice data and converts it into text data in real time using the speech recognition engine.

[1101] The converted text data is sent from the server to a generation AI (e.g., OpenAI GPT-3). The generation AI analyzes the text data and understands the user's inquiry. At the same time, an emotion analysis means (e.g., IBM Watson Tone Analyzer) analyzes the user's voice data and recognizes emotional data. This emotional data is also sent to the generation AI.

[1102] The AI ​​generates optimal responses based on the user's query and emotional data. The server receives the responses and converts them into voice data using a speech synthesis engine (e.g., Amazon Polly) to provide them to the user.

[1103] The server also sends the generated response information to the operator terminal. The operator terminal displays this response information on the screen, allowing the operator to select a response. In particular, if the user's request is a complaint, the generation AI proposes the optimal response and displays it on the operator terminal. The operator then takes appropriate action based on the proposal.

[1104] For example, consider the case where a user inquires about a slow internet connection. The server converts the received voice data into text data using a speech recognition engine, and then sends that text data to the generation AI. The generation AI generates a response that optimally addresses the slow internet connection issue, such as "Please try restarting your router." At the same time, if the emotion engine detects "dissatisfaction" in the user's voice, the generation AI also adds an apology, such as "We apologize for the inconvenience."

[1105] As another example, consider the case of a complaint that "the product was not delivered." In this case, the generation AI will similarly generate the optimal response, such as "We are sorry, we will check and contact you as soon as possible." Furthermore, if the emotion engine detects "anger," the generation AI will also add a response that is more in tune with the emotion, such as "We are very sorry for the inconvenience." The operator will respond promptly based on the suggestions displayed on the operator terminal.

[1106] In this way, the system of the present invention can increase user satisfaction by recognizing the user's emotions and providing the optimal response. The server also accumulates and saves the inquiry content, response information, and recognized emotion data in a database for later analysis and service improvement.

[1107] An example of a prompt sentence is "Please enter your query." For example, by entering "My Internet connection is slow," the system will generate an appropriate answer.

[1108] This invention makes it possible to realize effective call center response that takes into account the user's emotions.

[1109] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1110] Step 1:

[1111] When a user calls the call center, the server detects the call. Specifically, the server uses the VoIP system to trigger an incoming call event for a specific phone number. Based on the incoming call data (input), the server activates the IVR system (output).

[1112] Step 2:

[1113] The server starts the IVR system and plays a message to the user saying, "Please select an AI operator." This message is generated by the server using a text-to-speech engine (e.g., Amazon Polly), which converts text data (input) into voice data (output).

[1114] Step 3:

[1115] When the user presses a button or responds by saying "AI operator," the server recognizes this and starts a speech recognition engine (for example, Google Cloud Speech-to-Text API), which prepares the system to convert the voice data (input) into text data (output).

[1116] Step 4:

[1117] The server receives the user's voice data and converts it into text data in real time using a speech recognition engine. Specifically, the server sends the voice data to the Google Cloud Speech-to-Text API and receives the converted text data. The voice data (input) is converted into text data (output).

[1118] Step 5:

[1119] The server sends the acquired character data to the generation AI (e.g., OpenAI GPT-3). The server sends the character data (input) to the generation AI's API and receives the analysis result, which is the response information (output).

[1120] Step 6:

[1121] The generation AI analyzes the text data and understands the user's inquiry. For example, in response to the text data (input) "My internet connection is slow," it generates the appropriate solution, "Try restarting your router."

[1122] Step 7:

[1123] An emotion analysis means (e.g., IBM Watson Tone Analyzer) analyzes the user's voice data and recognizes the emotion data. The server sends the voice data (input) to the emotion analysis API and sends the recognized emotion data (output) to the generation AI.

[1124] Step 8:

[1125] The generation AI generates optimal response information by taking into account the content of the user's inquiry and emotional data. If the emotional data indicates "dissatisfaction," the generation AI adds an apology such as "We apologize for the inconvenience." Based on the text data and emotional data (input), the AI ​​generates comprehensive response information (output).

[1126] Step 9:

[1127] The server receives the generated response information and converts it into voice data using a speech synthesis engine (e.g., Amazon Polly). Text data (input) is converted into voice data (output).

[1128] Step 10:

[1129] The server provides the converted voice data to the user, specifically by playing the voice data (input) to the user's telephone terminal (output).

[1130] Step 11:

[1131] If necessary, the server sends the generated response information to the operator terminal. Character data (input) is sent to the operation device and displayed (output).

[1132] Step 12:

[1133] When a complaint occurs, the server converts the complaint voice data into text data and sends it to the generation AI. The generation AI analyzes the content of the complaint and proposes the optimal response. This is displayed on the operator terminal, allowing the operator to select a response. The voice data (input) is converted into text data and emotional data (output), and response information is generated.

[1134] Step 13:

[1135] After handling the complaint, the generation AI generates a follow-up message and sends it to the user via the server. It sends text data (input) and provides the user with follow-up message data (output).

[1136] Step 14:

[1137] The server accumulates and stores the user's inquiry, response information, and recognized emotion data in a database. This helps to increase user satisfaction and improve services. Text and emotion data (input) are stored in the database (output).

[1138] (Application example 2)

[1139] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1140] Conventional driver assistance systems for autonomous vehicles rely primarily on visual analysis, such as the user's facial expressions and body movements, and lack the ability to analyze emotions, including vocalizations, in real time. This makes it difficult to properly detect and quickly respond to user stress and dissatisfaction. The present invention aims to solve these problems and provide a comfortable riding experience in autonomous vehicles.

[1141] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving voice data and converting the voice data into character data, means for sending the character data to a generative AI model, which analyzes the character data and generates optimal response information, means for analyzing the user's emotions and sending the emotion data to the generative AI model, means for the generative AI model to generate optimal response information in consideration of the emotion data, and means for receiving the generated response information, converting the response information into voice data, and providing it to the user. This makes it possible to analyze emotions from the user's voice in real time and quickly provide an adaptive response.

[1142] "Audio data" means a digital representation of sound waves captured using a microphone or other input device.

[1143] "Character data" is data that indicates a string of characters that is recognized as a unit of language by analyzing voice data.

[1144] A "generative AI model" is an artificial intelligence system designed to analyze input text data and automatically generate appropriate responses and processing.

[1145] "Emotion data" is data that indicates the emotional state recognized from the user's voice data.

[1146] "Operator terminal" refers to equipment used to operate the system, and in particular to the computer and display device used by the operator.

[1147] "Response information" is data that indicates an appropriate answer or suggestion to a user's inquiry or request.

[1148] "Analysis" is the process of examining and evaluating the content and characteristics of input data in detail.

[1149] "Conversion" is the process of replacing data of one format with data of another format.

[1150] A "server" is a central management device for operating the entire system and processing information, and is a computer that processes data and performs communication.

[1151] This invention relates to an emotion recognition driver assistance system for autonomous vehicles. This system has the function of analyzing user voice data and providing optimal responses based on the user's emotions. Specific embodiments are described below.

[1152] First, the vehicle's server receives voice data through a microphone, which is then converted into text data using a voice recognition engine installed in the server.

[1153] The server then sends the converted text data to the generative AI model, which analyzes the user's speech. At the same time, the emotion engine analyzes the voice data and generates the user's emotional data, which is also sent to the generative AI model.

[1154] The generative AI model takes into account the text and emotional data to generate the optimal response information. This response information is sent back to the server and converted into voice data using a speech synthesis engine. This voice data is then provided to the user through the car's speakers.

[1155] The server also sends the generated response information to the operator terminal, which displays multiple answer candidates, allowing the operator to refer to them and make the most appropriate response.

[1156] For example, if a user says, "I'm tired, today was stressful," this voice data is converted into text data and sent to the generative AI model. At the same time, the emotion engine recognizes "stress" and sends it to the generative AI model as emotion data.

[1157] The generative AI model generates the optimal response based on the following prompt:

[1158] User says: I'm tired, today was stressful

[1159] User Emotion: Stress

[1160] Generate the appropriate response for your system:

[1161] The generative AI model generates a response such as, "Thank you for your hard work. I'll play some relaxing music to help you relax." The server converts this response into audio data and provides it to the user through a speaker.

[1162] The hardware used includes a microphone to collect voice data and a speaker to play the audio back to the user, while the software uses a speech recognition engine, an emotion engine, and a generative AI model, which allows for real-time recognition of user emotions and provides adaptive services.

[1163] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1164] Step 1:

[1165] The server receives voice data through the microphone. The input is the user's voice, and the output is digital voice data. Specifically, the microphone inside the car picks up the user's speech, converts it into a digital signal, and sends it to the server.

[1166] Step 2:

[1167] The server uses a speech recognition engine to convert the received voice data into text data. The input is voice data and the output is text data. Specifically, the speech recognition engine analyzes the sound waves and generates the corresponding text in real time.

[1168] Step 3:

[1169] The server sends the converted text data to the generative AI model, which then analyzes the user's speech. The input is text data, and the output is the analysis result. Specifically, the generative AI model understands the text, analyzes the context and intent, and prepares to generate an appropriate response.

[1170] Step 4:

[1171] The server uses an emotion engine to analyze emotion data from voice data. The input is voice data, and the output is emotion data. Specifically, the emotion engine analyzes voice parameters such as tone, tempo, and pitch to estimate the user's emotional state.

[1172] Step 5:

[1173] The server sends the emotion data to the generative AI model, which then generates optimal response information taking the emotion data into consideration. The input is emotion data and text data, and the output is response information. Specifically, the generative AI model generates a prompt sentence based on the text and emotion information, and then creates an appropriate response based on that.

[1174] Step 6:

[1175] The server receives the generated response information and converts it into voice data using a voice synthesis engine. The input is the response information and the output is voice data. Specifically, the voice synthesis engine converts the text into a voice format and generates data that can be played back as a human voice.

[1176] Step 7:

[1177] The server provides the converted voice data to the user through the car's speaker. The input is voice data, and the output is the voice that the user can hear. Specifically, the server sends the voice data to the speaker, and the speaker plays it back.

[1178] Step 8:

[1179] The server sends the generated response information to the operator terminal, which displays multiple answer candidates. The input is the response information, and the output is the displayed answer candidates. The specific operation is that the operator terminal displays the received information on the screen so that the operator can confirm the options.

[1180] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1181] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1182] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1183] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1184] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1185] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1186] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1187] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1188] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1189] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1190] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1191] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1192] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1193] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1194] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1195] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1196] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1197] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1198] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1199] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1200] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1201] The following is further disclosed regarding the above embodiment.

[1202] (Claim 1)

[1203] means for receiving voice data and converting the voice data into text data;

[1204] A means for transmitting the character data to a generation AI, which analyzes the character data and generates optimal response information;

[1205] means for receiving the generated response information, converting the response information into voice data, and providing the voice data to a user;

[1206] A system including:

[1207] (Claim 2)

[1208] 2. The system according to claim 1, further comprising means for transmitting the optimal response information to an operator terminal and displaying a plurality of answer candidates on the operator terminal.

[1209] (Claim 3)

[1210] The system according to claim 1, further comprising a means for, when the user's request is a complaint, proposing an optimal response using a generation AI and displaying the proposed content on an operator terminal.

[1211] (Claim 4)

[1212] a means for generating a follow-up message after a complaint has been handled;

[1213] means for providing said follow-up message to a user;

[1214] The system of claim 3 further comprising:

[1215] (Claim 5)

[1216] means for storing the user's requests and inquiries and generated response information;

[1217] means for categorizing the accumulated data and analyzing the number and trends of cases;

[1218] A means for reporting the analysis results to a planning department or a quality department;

[1219] The system of claim 1 further comprising:

[1220] "Example 1"

[1221] (Claim 1)

[1222] means for receiving voice data and converting the voice data into text data;

[1223] A means for transmitting the character data to a generation AI, which analyzes the character data and generates optimal response information;

[1224] means for receiving the generated response information, converting the response information into voice data, and providing the voice data to a user;

[1225] A means to accumulate and store data such as user inquiries and response information for later analysis,

[1226] A system including:

[1227] (Claim 2)

[1228] 2. The system according to claim 1, further comprising means for transmitting the optimal response information to an operator terminal and displaying a plurality of answer candidates on the operator terminal.

[1229] (Claim 3)

[1230] The system according to claim 1, further comprising a means for, when the user's request is a complaint, proposing an optimal response using a generation AI and displaying the proposed content on an operator terminal.

[1231] "Application Example 1"

[1232] (Claim 1)

[1233] means for receiving voice data and converting the voice data into text data;

[1234] A means for transmitting the character data to a generation AI, which analyzes the character data and generates optimal response information;

[1235] means for receiving the generated response information, converting the response information into voice data, and providing the voice data to a user;

[1236] A means for allowing a user to issue a query and generating an optimal answer regarding a product on an online shopping site as a response to the query;

[1237] means for converting the optimal answer into voice data using a voice synthesis engine and replying to the user;

[1238] A system including:

[1239] (Claim 2)

[1240] 2. The system according to claim 1, further comprising means for transmitting the optimal response information to an operator terminal and displaying a plurality of answer candidates on the operator terminal.

[1241] (Claim 3)

[1242] The system according to claim 1, further comprising a means for, when the user's request is a complaint, proposing an optimal response using a generation AI and displaying the proposed content on an operator terminal.

[1243] "Example 2: Combining Emotion Engines"

[1244] (Claim 1)

[1245] means for receiving voice data and converting the voice data into text data;

[1246] a means for transmitting the character data to an artificial intelligence, and for the character data to be analyzed by the artificial intelligence and generate optimal response information;

[1247] a means for recognizing emotion data from the voice data by an emotion analysis means and transmitting the emotion data to the artificial intelligence;

[1248] means for receiving the generated response information, converting the response information into voice data, and providing the voice data to a user;

[1249] A system including:

[1250] (Claim 2)

[1251] 2. The system according to claim 1, further comprising means for transmitting the optimal response information to an operation device and displaying a plurality of answer candidates on the operation device.

[1252] (Claim 3)

[1253] 2. The system according to claim 1, further comprising means for proposing an optimal solution using artificial intelligence when the user's request is a complaint, and displaying the content of the proposal on the operation device.

[1254] "Application example 2 when combining emotion engines"

[1255] (Claim 1)

[1256] means for receiving voice data and converting the voice data into text data;

[1257] a means for transmitting the character data to a generating AI model, and for the character data to be analyzed by the generating AI model and generate optimal response information;

[1258] means for analyzing the user's emotions and transmitting the emotion data to a generative AI model;

[1259] A means for the generative AI model to generate optimal response information in consideration of emotional data;

[1260] means for receiving the generated response information, converting the response information into voice data, and providing the voice data to a user;

[1261] A system including:

[1262] (Claim 2)

[1263] 2. The system according to claim 1, further comprising means for transmitting the optimal response information to an operator terminal and displaying a plurality of answer candidates on the operator terminal.

[1264] (Claim 3)

[1265] The system according to claim 1, further comprising a means for proposing an optimal response using a generative AI model when the user's request is a complaint, and displaying the content of the proposal on an operator terminal.

[1266] (Claim 4)

[1267] 10. The system of claim 1, further comprising means for generating and providing an adaptive response based on emotions recognized from the user's voice. [Explanation of symbols]

[1268] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for receiving voice data and converting the voice data into text data; A means for transmitting the character data to a generation AI, which analyzes the character data and generates optimal response information; means for receiving the generated response information, converting the response information into voice data, and providing the voice data to a user; A system including:

2. The system according to claim 1, further comprising means for transmitting the optimal response information to an operator terminal and displaying a plurality of answer candidates on the operator terminal.

3. The system according to claim 1, further comprising means for, when the user's request is a complaint, proposing an optimal response using a generation AI and displaying the proposed content on an operator terminal.

4. a means for generating a follow-up message after a complaint has been handled; means for providing said follow-up message to a user; The system of claim 3 further comprising:

5. means for storing the user's requests and inquiries and generated response information; means for categorizing the accumulated data and analyzing the number and trends of cases; A means for reporting the analysis results to a planning department or a quality department; The system of claim 1 further comprising:

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A