System

The system addresses language and cultural barriers by converting voice input to text, analyzing sentiment, and delivering personalized solutions across multiple languages, improving customer service efficiency and satisfaction.

JP2026019116APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024120525
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-25
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Medium- to large-sized companies face challenges in providing efficient and personalized customer service due to language and cultural barriers, difficulty in understanding customer inquiries, and responding quickly to diverse emotional and linguistic needs.

Method used

A system that receives customer utterances as voice input, converts them into text data, analyzes sentiment, identifies needs, provides multilingual support, and retrains AI models to deliver personalized solutions.

Benefits of technology

Enables quick and efficient analysis of customer emotions and needs across multiple languages, providing personalized responses that enhance customer satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026019116000001_ABST
    Figure 2026019116000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for receiving a customer's speech as audio input and converting it to textual data by a speech recognition system; means for analyzing the textual data and evaluating the customer's emotions; means for identifying the customer's needs and generating a personalized solution; means for translating the solution into multiple languages; means for providing the translated solution to the customer; and means for saving the customer interaction and retraining the AI model.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Today's medium- to large-sized companies, especially those in industries that require global customer service, face limitations in customer service due to language and cultural barriers. Furthermore, the difficulty of responding quickly and appropriately to diverse inquiries makes it difficult to improve customer satisfaction. It is not easy to accurately understand what customers are saying and instantly analyze their emotions and needs. This creates a need for efficient and personalized service delivery. [Means for solving the problem]

[0005] In order to solve the above problems, the present invention provides the following means. It includes a means for receiving customer utterances as voice input and converting them into text data using a voice recognition system. It also includes a sentiment analysis means for analyzing the text data and evaluating the customer's emotions. This provides a means for identifying customer needs and generating personalized solutions. It also includes a multilingual support means for translating solutions into multiple languages ​​and providing the translated solutions to the customer. It also incorporates a means for saving customer interaction data and retraining the AI ​​model, thereby building a system that can always provide the latest responses.

[0006] "Customer Utterances" means voice or text inputs that represent customer questions, feedback, or comments regarding a particular product or service.

[0007] "Voice input" refers to the customer's voice data before it is processed by the voice recognition system.

[0008] A "voice recognition system" refers to a technology that converts voice data into text data.

[0009] "Text data" refers to a string representation of a customer's utterances generated by a speech recognition system.

[0010] "Means for analyzing text data" refers to technology that uses natural language processing technology to understand the meaning of text data and classify and evaluate it.

[0011] "Sentiment analysis tools" refers to technologies that use analytics algorithms to identify a customer's emotional state from text data.

[0012] "Means for identifying needs" refers to technology that identifies the information and solutions that customers are looking for based on what the customers say and the results of analysis.

[0013] "Personalized solutions" refer to specific responses or information optimized for a particular customer problem.

[0014] "Multilingual support" refers to the technology that translates solutions into multiple languages ​​and responds appropriately to customers who speak different languages.

[0015] "Means for providing to the customer" refers to the technology for displaying the generated solution on a terminal or playing it aloud.

[0016] "Interaction Data" means records of interactions with customers, including voice, text, and analytics results.

[0017] "Means to retrain AI models" refers to techniques that use collected interaction data and feedback to update AI algorithms and improve system performance.

[0018] "Database" refers to a system that efficiently stores and manages customer inquiry history and other related information. [Brief explanation of the drawings]

[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8]FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0021] First, the terms used in the following description will be explained.

[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0027] [First embodiment]

[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0040] The present invention provides a multilingual and multimodal customer service system aimed at improving customer service, and an embodiment thereof is shown below.

[0041] Program processing explanation

[0042] 1. Input and recognition of customer utterances

[0043] The user makes a voice inquiry about a problem with the product. The device records this voice and sends the data to the server. The server then uses a voice recognition system to convert the recorded voice data into text data.

[0044] 2. Sentiment and Text Analysis

[0045] The server analyzes the converted text data and uses NLP (Natural Language Processing) technology to understand the content of customer comments. It also uses a sentiment analysis system to evaluate the emotional state in the text data. For example, if a user says, "This product is completely unusable!", the sentiment analysis system will classify the comment as "anger," a negative emotion.

[0046] 3. Identifying needs and generating personalized solutions

[0047] The server retrieves past customer inquiries from a database and combines them with the current text data. It then identifies the customer's needs based on what they said and their sentiment assessment. Based on this, the AI ​​model generates a personalized solution. For example, if the customer is contacted about a software crash, it could provide instructions on how to change settings or the latest patch.

[0048] 4. Multilingual support

[0049] The server translates the generated solutions into multiple languages ​​as needed, for example, when translating a solution written in English into Japanese, it converts it into natural language while preserving the exact meaning.

[0050] 5. Providing solutions

[0051] The device displays the solution received from the server to the user, and also interactively asks the user additional questions or offers suggestions to increase engagement, such as displaying a message with instructions on how to change settings and offering a link to download the latest patch.

[0052] 6. Data Feedback and Learning

[0053] The user follows the system's instructions and attempts to solve the problem. The device sends the results and any follow-up questions to the server, which stores this interaction data and uses it to retrain the AI ​​model, improving the accuracy and personalization of responses to future inquiries.

[0054] Specific examples

[0055] Scenario: A user contacts support in Japanese.

[0056] 1. User: Asks by voice, "This software keeps crashing, what should I do?"

[0057] 2. Device: Records audio and sends it to the server.

[0058] 3. Server: Uses a speech recognition system to convert the speech into text data such as "This software keeps crashing, what should I do?"

[0059] 4. Server: Analyzes the content of the text and identifies the emotion "anger" using an emotion analysis system.

[0060] 5. Server: Retrieves past query history from the database, identifies needs, and generates solutions that provide configuration change instructions and the latest patches.

[0061] 6. Server: Translate the solution into Japanese on a multilingual system and generate the message "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link."

[0062] 7. Terminal: Display the translated solution to the user.

[0063] 8. User: Try the steps provided and enter additional questions if the problem persists.

[0064] 9. Terminal: Sends additional queries to the server.

[0065] 10. Server: Provides additional countermeasures and stores interaction data as training data.

[0066] In this way, the system of the present invention aims to improve customer satisfaction by analyzing customer comments in multiple languages ​​and multiple modalities and providing personalized solutions quickly and efficiently.

[0067] The processing flow will be explained below.

[0068] Step 1:

[0069] A user voice-inquires about a problem with a product, for example, "This software keeps crashing, what should I do?"

[0070] Step 2:

[0071] The terminal records the user's voice and transmits the voice data to the server.

[0072] Step 3:

[0073] The server inputs the received voice data into a voice recognition system and converts the voice data into text data, for example, "This software keeps crashing, what should I do?"

[0074] Step 4:

[0075] The server applies natural language processing algorithms to analyze the text data, extracting themes and keywords from the comments.

[0076] Step 5:

[0077] The server uses an emotion analysis system to identify the user's emotion from the text data, which in this case is classified as "anger."

[0078] Step 6:

[0079] The server retrieves the customer's past inquiry history from the database and searches for solutions to similar problems.

[0080] Step 7:

[0081] The server uses the results of text and sentiment analysis to identify customer needs, such as a need for a solution to a software crash.

[0082] Step 8:

[0083] The server uses AI models to generate personalized solutions for identified needs, such as instructions for changing settings or providing the latest patches.

[0084] Step 9:

[0085] The server utilizes a multilingual system to translate the generated solutions into multiple languages ​​as needed, for example, translating the solutions from Japanese to English.

[0086] Step 10:

[0087] The device displays the solution received from the server to the user, for example, "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link."

[0088] Step 11:

[0089] The user attempts the provided solution and types additional questions or feedback about the solution into the device.

[0090] Step 12:

[0091] The device sends the user's feedback and any follow-up questions back to the server.

[0092] Step 13:

[0093] The server analyzes the additional information and takes further action if necessary, and also stores the interaction data and stores it in a database for use in retraining the AI ​​model.

[0094] This series of steps enables the system of the present invention to respond quickly and efficiently to customer needs and provide a high level of personalized customer service.

[0095] Example 1

[0096] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0097] Modern companies strive to improve customer service, but implementing a multilingual, multimodal customer service system is a technically difficult challenge. Customers make inquiries in a variety of languages, and their emotions and needs vary widely, making responding to them time-consuming and labor-intensive. Conventional systems have difficulty accurately grasping customer emotions and needs, and responses tend to be uniform, failing to sufficiently increase customer satisfaction.

[0098] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0099] In this invention, the server includes means for receiving customer utterances as voice input and converting them into text data using a voice recognition system, means for analyzing the text data and evaluating customer sentiment, and means for identifying customer needs and generating personalized solutions. This makes it possible to analyze customer utterances in multiple languages ​​and multiple modalities and provide personalized solutions quickly and efficiently, thereby improving customer satisfaction.

[0100] "Voice input" is a method by which a user can query a system through speech.

[0101] A "voice recognition system" refers to a technology that analyzes recorded voice data and converts it into text data.

[0102] "Text data" refers to character information converted from audio data.

[0103] "Sentiment analysis means" refers to technology that evaluates and classifies emotional states in text data.

[0104] "Personalized solutions" are a method for providing the best possible solution to a specific problem based on the customer's needs.

[0105] "Multilingual means" refers to techniques for translating generated solutions into multiple languages.

[0106] "Interaction Data" refers to records of conversations between a customer and a system.

[0107] "Retraining methods" refers to techniques that use stored dialogue data to retrain AI models and improve their accuracy.

[0108] "Terminal" refers to the device used by the Customer to provide voice input.

[0109] "Database" refers to a repository of information where past inquiry history and other related information is stored.

[0110] "Natural language processing technology" refers to a series of technologies for analyzing text data and understanding its content.

[0111] "Engagement" refers to methods for eliciting involvement and interest through interaction with users.

[0112] An "AI model" refers to an algorithm that has been trained through machine learning to perform a specific task.

[0113] The present invention provides a multilingual and multimodal customer service system aimed at improving customer service, and an embodiment thereof is shown below.

[0114] System configuration

[0115] The system begins when a user voice-inquires about a problem with a product. When the user voice-inquires, the device records the voice and sends it to the server. The server then uses a voice recognition system to convert the voice data into text. The server then analyzes the text data, evaluates the customer's sentiment, and identifies their needs. It then uses an AI model to generate a personalized solution and translates it into multiple languages. Finally, the translated solution is provided to the user via the device.

[0116] Hardware and software used

[0117] Devices: Devices such as smartphones, computers, and tablets are used.

[0118] Server: A remote server is used for data processing and storage.

[0119] Speech recognition system: Google Cloud Speech-to-Text API, IBM Watson Speech to Text, etc. are used.

[0120] NLP technologies: Natural language processing libraries such as TensorFlow, spaCy, and NLTK are used.

[0121] Sentiment analysis system: TextBlob, NRC Emotion Lexicon, etc. are used.

[0122] Translation system: Google Cloud Translation API, DeepL API, etc. are used.

[0123] AI models: Generative AI models such as GPT and BERT are used.

[0124] Specific examples

[0125] The following are specific usage scenarios for this system:

[0126] 1. User: Makes a voice inquiry saying, "This software keeps crashing, what should I do?"

[0127] 2. Device: Records the user's voice and sends it to the server.

[0128] 3. Server: Using a speech recognition system, convert the speech into text data such as "This software keeps crashing, what should I do?"

[0129] 4. Server: Analyzes text data using NLP technology and identifies the emotion "anger."

[0130] 5. Server: Retrieves past inquiry history from a database, identifies needs, and uses AI models to generate solutions that provide configuration change instructions and the latest patches.

[0131] 6. Server: Translate the generated solution using a multilingual system. For example, the message "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link." is translated into Japanese.

[0132] 7. Terminal: Display the translated solution to the user.

[0133] Prompt Sentence Examples

[0134] Here are some examples of prompts for generative AI models:

[0135] Prompt sentences that convert user speech into text

[0136] Please convert the following audio data to text: [Audio data]

[0137] Prompt sentences that analyze the sentiment of the user's text utterances

[0138] Please rate the emotional state of the following text data: [Text data]

[0139] Prompts that identify user needs and generate solutions

[0140] Identify the user's needs and generate a solution based on the following text: [Text]

[0141] Prompts for translating generated text into multiple languages

[0142] Please translate the following text data into [language name]: [text data]

[0143] By using these procedures and prompts, the system of the present invention aims to analyze customer utterances in multiple languages ​​and multiple modalities and provide personalized solutions quickly and efficiently.

[0144] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0145] Step 1:

[0146] A user speaks to a product about a problem they are having. For example, a user might say, "This software keeps crashing. What should I do?" This is the input.

[0147] Step 2:

[0148] The device records the user's voice and sends the data to the server. The recording is performed using the microphone on a smartphone or computer. The recorded voice data is output.

[0149] Step 3:

[0150] The server uses a speech recognition system to convert the speech data into text data. Using the speech data as input, a speech recognition system such as the Google Cloud Speech-to-Text API converts the speech into text data such as "This software keeps crashing, what should I do?". The output is the text data converted from the speech.

[0151] Step 4:

[0152] The server analyzes the text data and uses NLP technology to understand the content of customer comments. It uses natural language processing tools such as TensorFlow and spaCy to analyze the structure and meaning of sentences, taking the text data as input. The output of this process is analyzed text data.

[0153] Step 5:

[0154] The server uses a sentiment analysis system to evaluate the emotional state in the text data. Using the analyzed text data as input, it uses tools such as TextBlob and the NRC Emotion Lexicon to identify the customer's emotion as "anger." The output is the analyzed sentiment data.

[0155] Step 6:

[0156] The server retrieves past query history from the database and combines it with the current text data. Using the past query history stored in the database as input, it extracts relevant data using SQL queries. The output is a combination of the past query history and the current text data.

[0157] Step 7:

[0158] The server identifies customer needs and generates personalized solutions using AI models. Using the combined text data as input, machine learning models (GPT and BERT) are used to identify needs and generate solutions. The output is a personalized solution.

[0159] Step 8:

[0160] The server translates the generated solution into multiple languages ​​as needed. The generated solution is used as input and translated using the Google Cloud Translation API or the DeepL API. The output is the translated solution in multiple languages.

[0161] Step 9:

[0162] The device displays the translated solution received from the server to the user. The translated solution is used as input and a solution message is displayed on the smartphone or PC screen. The output is the solution provided to the user.

[0163] Step 10:

[0164] The user follows the system's instructions to try to solve the problem. The user follows the displayed solution steps, changing software settings or downloading patches from links. This is the input, and the results of the setting changes and patch downloads are the output.

[0165] Step 11:

[0166] The terminal sends the results and any additional questions to the server. The terminal sends data to the server using the results of the solution execution and any new questions from the user as input. The output is the sent result data and any additional questions.

[0167] Step 12:

[0168] The server stores the dialogue data and retrains the AI ​​model. It uses newly collected dialogue data as input, stores the data, and retrains the AI ​​model. The output is a trained AI model, which improves the accuracy of responses to future inquiries.

[0169] (Application example 1)

[0170] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0171] Current customer service systems only support one language and lack multilingual support and sentiment analysis, preventing them from fully improving customer satisfaction. They also lack the means to quickly provide personalized solutions, making it difficult to respond immediately to follow-up customer inquiries. There is a need for a system that can provide multilingual and multimodal support in brick-and-mortar stores using smartphones and service robots.

[0172] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0173] In this invention, the server includes a means for converting voice input into text data using a voice recognition system, a sentiment analysis means for analyzing the text data to evaluate customer sentiment, a means for identifying customer needs and generating personalized solutions, a multilingual support means for translating the solutions into multiple languages, a means for providing the translated solutions to customers, a means for saving customer interaction data and retraining an AI model, a customer service system that functions on hardware including smartphones and service robots, and a means for providing customer service through voice, text, and translated solutions and for engaging in dialogue based on sentiment evaluation and needs, thereby enabling the provision of advanced customer service that is multilingual and emotionally responsive.

[0174] A "customer support system" is a system that responds quickly and appropriately to customer inquiries in physical stores and online.

[0175] "Voice input" is a means of acquiring customer utterances as voice data.

[0176] A "voice recognition system" is a technology or device for converting voice data into text data.

[0177] "Text data" is data expressed in the form of character information.

[0178] "Sentiment analysis means" is a means for evaluating and classifying customer emotions from text data.

[0179] "Needs identification means" are means for clarifying specific requirements or requests based on the content of a customer inquiry.

[0180] A "personalized solution" is a customized response to address a customer's specific needs.

[0181] "Multilingual support means" is a means for translating generated solutions into multiple languages.

[0182] "Interaction Data" is a record of the interaction between a customer and the system.

[0183] An "AI model" is a mathematical model for learning and prediction using artificial intelligence technology.

[0184] "Retraining methods" are methods used to improve the accuracy of existing AI models using new data.

[0185] A "smartphone" is a portable information terminal that can input voice and run applications.

[0186] A "service robot" is an autonomous robot designed to provide customer service and information.

[0187] "Multimodal" is the property of a system that includes multiple input and output formats, such as speech and text.

[0188] The present invention is a multilingual, multimodal customer service system for enhancing customer service in brick-and-mortar stores. This system combines the processes of speech recognition, emotion analysis, text analysis, needs identification, multilingual support, solution provision, data feedback, and learning. Specific embodiments for implementing the present invention are described below.

[0189] The main components of the system are the server, the terminal, and the user. The terminal consists of a smartphone or a service robot, and is responsible for receiving voice input from the user and sending it to the server.

[0190] Hardware used

[0191] Smartphone: A mobile information device that can accept voice input and run applications.

[0192] Service robot: An autonomous robot designed to provide customer service and information.

[0193] Software used

[0194] SpeechRecognition: A Python library for performing speech recognition.

[0195] transformers: A library that provides sentiment analysis, text analysis, and translation models.

[0196] Data Processing Steps

[0197] 1. Receiving and converting voice input

[0198] A user makes a voice inquiry to a smartphone or a service robot. For example, the user might say, "Please tell me how to use this product." The device records this voice and sends it to a server.

[0199] 2. Speech Recognition and Emotion Analysis

[0200] The server uses a speech recognition system (SpeechRecognition) to convert the recorded voice data into text data. The converted text data is then analyzed using sentiment analysis (transformers) to evaluate the customer's emotions. Through this process, the text is converted to "Please tell me how to use this product." and the emotion is classified as "confused."

[0201] 3. Needs Identification and Solution Generation

[0202] The server uses text analysis tools (transformers' text-classification) to identify customer needs based on the analyzed text data, retrieves personalized solutions from the database and applies them to the customer.

[0203] 4. Multilingual support

[0204] The generated solutions are translated into multiple languages ​​as needed. A multilingual solution (transformers' translation-en-to-ja) is used to properly translate from English to Japanese.

[0205] 5. Providing solutions and dialogue

[0206] The translated solution is sent to the device and displayed to the user. If the user asks additional questions, that data is also sent to the server for processing again.

[0207] 6. Data Feedback and Learning

[0208] The server stores the results of the conversation and any follow-up questions, and retrains the AI ​​model, improving its accuracy for the next inquiry.

[0209] Examples and prompts

[0210] Specific examples

[0211] 1. User: Ask, "How do I use this product?"

[0212] 2. Application: Convert speech to text and analyze sentiment and content.

[0213] 3. Application: Multilingual support and display of usage solutions.

[0214] 4. User: Enter a follow-up question.

[0215] 5. Application: Save the proposed responses and use them for future learning.

[0216] Prompt Sentence Examples

[0217] Customer: "Please tell me how to use this product."

[0218] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0219] Step 1:

[0220] Receiving and converting voice input

[0221] A user makes a voice inquiry to a smartphone or service robot. For example, the user might say, "Please tell me how to use this product." This inputs voice data into the terminal. The terminal records this voice and sends it to a server as digital voice data.

[0222] Input: User speech (e.g., "How do I use this product?")

[0223] Output: Digital audio data

[0224] Step 2:

[0225] Speech Recognition and Emotion Analysis

[0226] The server converts the received digital voice data into text data using a speech recognition system (SpeechRecognition library). For example, the voice is converted into text such as "Please tell me how to use this product." The text data is then analyzed using a sentiment analysis method (transformers' sentiment-analysis model) to evaluate the customer's emotions. As a result of the sentiment analysis, the text is classified as "confused."

[0227] Input: Digital audio data

[0228] Output: Text data (e.g., "How do I use this product?") and emotion ratings (e.g., "Confusion")

[0229] Step 3:

[0230] Needs Identification and Solution Generation

[0231] The server uses text analysis methods (Transformers' text-classification model) to identify customer needs based on the text data. For example, the tag "How to use the product" is identified. Then, it retrieves the corresponding solutions from the database and generates personalized solutions. For example, instructions on how to use the product are retrieved from the database.

[0232] Input: Text data and sentiment ratings

[0233] Output: Identified needs (e.g., how to use the product) and personalized solutions (e.g., product usage instructions)

[0234] Step 4:

[0235] Multilingual support

[0236] The server translates the generated solution as needed using multilingual support (transformers' translation-en-to-ja model). For example, a solution generated in English is translated into Japanese. This results in a solution in the target language.

[0237] Input: personalized solution (e.g. product usage instructions)

[0238] Output: Translated solution (e.g., product usage instructions in Japanese)

[0239] Step 5:

[0240] Providing solutions and dialogue

[0241] The translated solution is sent to the device, which displays it to the user and continues the dialogue via voice or text as needed. For example, if the user reviews the presented instructions and asks additional questions, their input is sent back to the server.

[0242] Input: translated solution

[0243] Output: The solution and any follow-up questions displayed to the user

[0244] Step 6:

[0245] Data Feedback and Learning

[0246] The results of the user's attempts and any follow-up questions are sent to the server, which stores this interaction data and uses it to retrain the AI ​​model, improving its accuracy for the next inquiry.

[0247] Input: User feedback and follow-up questions

[0248] Output: Dialogue data for retraining and improved AI models

[0249] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0250] The present invention provides a multilingual, multimodal customer support system aimed at improving customer service, and includes a form that combines an emotion engine. An embodiment of this system is shown below.

[0251] Program processing explanation

[0252] 1. Input and recognition of customer utterances

[0253] A user voice-inquires about a problem with a product. For example, "This software keeps crashing. What should I do?" The device records this voice and sends the voice data to the server. The server uses a voice recognition system to convert the recorded voice data into text data.

[0254] 2. Sentiment and Text Analysis

[0255] The server applies natural language processing algorithms to analyze the converted text data. This analysis extracts themes and keywords from the utterances. It then uses an emotion engine to identify the user's emotional state from the text data. For example, if a user utters, "This product is completely unusable!", the emotion engine classifies the utterance as "anger," a negative emotion. The emotion engine evaluates emotions in real time and adjusts the response as needed.

[0256] 3. Identifying needs and generating personalized solutions

[0257] The server retrieves the customer's past inquiry history from the database and combines it with the current text data for analysis. This allows it to identify the customer's needs and generate a personalized solution. For example, if the inquiry is about a software crash, it generates a message providing instructions on changing settings and the latest patches.

[0258] 4. Multilingual support

[0259] The server translates the generated solutions into multiple languages ​​as needed, for example, when translating a solution written in English into Japanese, it converts it into natural language while preserving the exact meaning.

[0260] 5. Providing solutions

[0261] The device displays the solution received from the server to the user, for example, "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link."

[0262] 6. Data Feedback and Learning

[0263] The user follows the system's instructions and attempts the proposed solution. They then enter additional questions or feedback about the solution into the device. The device then sends the user's feedback and additional questions back to the server. The server analyzes the additional information and takes further action if necessary. The interaction data is also saved and stored in a database for use in retraining the AI ​​model.

[0264] Specific examples

[0265] Scenario: A user contacts support in Japanese.

[0266] 1. User: Asks by voice, "This software keeps crashing, what should I do?"

[0267] 2. Device: Records audio and sends it to the server.

[0268] 3. Server: Uses a speech recognition system to convert the speech into text data such as "This software keeps crashing, what should I do?"

[0269] 4. Server: Analyzes the content of the text and identifies the emotion "anger" using an emotion engine.

[0270] 5. Server: Retrieves past query history from the database, identifies needs, and generates solutions that provide configuration change instructions and the latest patches.

[0271] 6. Server: Translate the solution into Japanese in a multilingual system and generate the message "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link."

[0272] 7. Terminal: Display the translated solution to the user.

[0273] 8. User: Try the steps provided and enter additional questions if the problem persists.

[0274] 9. Terminal: Sends additional queries to the server.

[0275] 10. Server: Provides additional countermeasures and stores interaction data as training data.

[0276] In this way, the system of the present invention aims to improve customer satisfaction by analyzing customer comments in multiple languages ​​and multiple modalities, using an emotion engine to evaluate the user's emotional state in real time, and providing personalized solutions quickly and efficiently.

[0277] The processing flow will be explained below.

[0278] Step 1:

[0279] A user voice-inquires about a problem with a product, for example, "This software keeps crashing, what should I do?"

[0280] Step 2:

[0281] The terminal records the user's voice and transmits the voice data to the server.

[0282] Step 3:

[0283] The server inputs the received voice data into a voice recognition system and converts the voice data into text data, for example, "This software keeps crashing, what should I do?"

[0284] Step 4:

[0285] The server then analyzes the converted text data using a natural language processing algorithm, extracting the topic and keywords of the comments and identifying the nature of the problem.

[0286] Step 5:

[0287] The server uses an emotion engine to identify the user's emotional state from the text data in real time, for example, in this case the emotion "anger."

[0288] Step 6:

[0289] The server retrieves the customer's past inquiry history from the database and collects relevant data for resolving the problem.

[0290] Step 7:

[0291] The server identifies specific customer needs based on the results of text analysis and sentiment analysis. For example, based on the frequent occurrence of software crashes, it determines that the customer needs a procedure for changing settings to prevent crashes.

[0292] Step 8:

[0293] The server uses AI models to generate personalized solutions based on identified needs, such as "Settings - Options - Crash Prevention" instructions or a "download link for the latest patch."

[0294] Step 9:

[0295] The server translates the generated solution into the required language using a multilingual system, for example translating the solution from Japanese to English.

[0296] Step 10:

[0297] The device displays the solution received from the server to the user, for example, "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link."

[0298] Step 11:

[0299] The user attempts the provided solution and enters additional questions or feedback about the solution into the terminal.

[0300] Step 12:

[0301] The device sends the user's feedback and follow-up questions to the server.

[0302] Step 13:

[0303] The server analyzes the additional information and takes further action if necessary, such as providing more detailed configuration instructions or an alternative solution, and also saves the interaction data and stores it in a database for use in retraining the AI ​​model.

[0304] Through this series of steps, the system of the present invention quickly and efficiently analyzes customer comments in multiple languages ​​and multiple modalities, and evaluates user sentiment in real time using an emotion engine, thereby significantly improving customer satisfaction by providing personalized solutions.

[0305] Example 2

[0306] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0307] Modern customer service requires multilingual and multimodal customer interactions, and it is becoming increasingly important to provide personalized solutions quickly based on customer emotions and needs. However, conventional systems struggle to accurately convert voice input into text, perform sentiment analysis, and provide solutions adapted to multiple languages, making it difficult to increase customer satisfaction. Furthermore, while continuous learning and improvement based on dialogue data is desirable, there is a lack of systems that can achieve this.

[0308] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0309] In this invention, the server includes means for receiving customer utterances as voice input and converting them into text data using speech recognition technology, sentiment analysis means for analyzing the text data and evaluating the customer's sentiment, means for identifying customer needs and generating personalized solutions, multilingual support means for translating the generated solutions into multiple languages, and means for saving customer interaction data and retraining a machine learning model. This makes it possible to analyze customer utterances in multiple languages ​​and evaluate the user's emotional state in real time using an emotion engine, thereby quickly and efficiently providing personalized solutions.

[0310] Below are definitions of important terms included in the claims.

[0311] "Voice recognition technology" is a technology that analyzes voice data and converts it into text data.

[0312] "Text data" is character information converted using voice recognition technology.

[0313] An "emotion analysis means" is an algorithm or system that analyzes text data and identifies the user's emotional state.

[0314] "Means for identifying needs" refers to a technique for extracting customer requests and requirements based on the customer's past inquiry history and current statements.

[0315] "Personalized solutions" are customized solutions or guidance provided based on a customer's specific situation and emotions.

[0316] "Multilingual solutions" are technologies or systems that translate solutions into multiple languages.

[0317] "Dialogue data" is historical information about questions and answers exchanged with customers.

[0318] A "machine learning model" is an algorithm that learns from past data and makes predictions and classifications.

[0319] "Retraining" is a technique for adding new data to improve the accuracy and performance of a machine learning model.

[0320] A "terminal" is an electronic device used by a customer that inputs voice and displays text.

[0321] A "database" is a system that stores and manages past inquiry history and customer information.

[0322] The present invention provides a multilingual, multimodal customer support system aimed at improving customer service, including a form that combines an emotion engine. The following describes the embodiments of the system in detail.

[0323] System configuration

[0324] The multilingual and multimodal customer support system of the present invention is composed of the following elements: a server, a terminal, and a user.

[0325] Server: A high-performance computer running speech recognition technology, natural language processing algorithms, emotion engines, machine learning models, and translation systems. Examples include software such as Google Cloud Speech-to-Text, Google Cloud Natural Language API, and DeepL API.

[0326] Terminal: An electronic device capable of voice input and text display, such as a smartphone, tablet, or PC.

[0327] User: A customer who contacts us regarding a problem with a product.

[0328] Program processing explanation

[0329] Customer utterance input and recognition

[0330] A user voice-inquires about a problem with a product. For example, "This software keeps crashing. What should I do?" The device records this voice and sends the voice data to a server. The server then uses voice recognition technology, such as Google Cloud Speech-to-Text, to convert the recorded voice data into text.

[0331] Sentiment and Text Analysis

[0332] The server then applies natural language processing algorithms, such as the Google Cloud Natural Language API, to analyze the converted text data. This analysis extracts the topic and keywords of the utterance. It then uses an emotion engine to identify the user's emotional state from the text data. For example, if a user says, "This product is completely unusable!", the emotion engine classifies the utterance as "anger," a negative emotion. The emotion engine evaluates the emotion in real time and adjusts the response as needed.

[0333] Identifying needs and generating personalized solutions

[0334] The server retrieves the customer's past inquiry history from the database and combines it with the current text data for analysis. This allows it to identify the customer's needs and generate a personalized solution. For example, if the inquiry is about a software crash, it generates a message providing instructions on changing settings and the latest patches.

[0335] Multilingual support

[0336] The server translates the generated solutions into multiple languages ​​as needed. For example, if a solution written in English needs to be translated into Japanese, it uses a translation system such as the DeepL API to convert it into natural language while preserving the exact meaning.

[0337] Providing solutions

[0338] The device displays the solution received from the server to the user, for example, "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link."

[0339] Data Feedback and Learning

[0340] The user follows the system's instructions and attempts the proposed solution. They then enter additional questions or feedback about the solution into the device. The device then sends the user's feedback and additional questions back to the server. The server analyzes the additional information and takes further action if necessary. The interaction data is also saved and stored in a database for use in retraining the AI ​​model.

[0341] Specific examples

[0342] Scenario: A user contacts support in Japanese.

[0343] 1. User: Asks by voice, "This software keeps crashing, what should I do?"

[0344] 2. Device: Records audio and sends it to the server.

[0345] 3. Server: Use Google Cloud Speech-to-Text to convert the speech into text data such as "This software keeps crashing, what should I do?"

[0346] 4. Server: The text content is analyzed using the Google Cloud Natural Language API, and the emotion engine identifies the emotion "anger."

[0347] 5. Server: Retrieves past query history from the database, identifies needs, and generates solutions that provide configuration change instructions and the latest patches.

[0348] 6. Server: Translate the solution into Japanese using the DeepL API and generate the message "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link."

[0349] 7. Terminal: Display the translated solution to the user.

[0350] 8. User: Try the steps provided and enter any follow-up questions if the problem persists.

[0351] 9. Terminal: Sends additional queries to the server.

[0352] 10. Server: Provides additional countermeasures and stores the dialogue data as training data.

[0353] In this way, the system of the present invention aims to improve customer satisfaction by analyzing customer comments in multiple languages ​​and multiple modalities, using an emotion engine to evaluate the user's emotional state in real time, and providing personalized solutions quickly and efficiently.

[0354] Example prompt sentence:

[0355] "A user speaks to you about a software crash. Design a system that uses speech recognition to convert the text and analyzes the user's emotional state to generate a personalized solution. Describe the process for translating the generated solution into various languages ​​and displaying it to the user."

[0356] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0357] Step 1:

[0358] The user makes a voice inquiry about a problem with the product, for example, "This software keeps crashing. What should I do?" The device records this voice and sends the voice data to the server.

[0359] Input: User voice input

[0360] Output: Audio data sent to the server

[0361] Specific operation: The user uses the device's microphone to voice-input the details of the problem, and the device sends the voice data to the server.

[0362] Step 2:

[0363] The server uses voice recognition technology such as Google Cloud Speech-to-Text to convert the recorded audio data into text data.

[0364] Input: Audio data

[0365] Output: Text data

[0366] Specific operation: The server calls the speech recognition API and converts the speech data into text data.

[0367] Step 3:

[0368] The server then applies natural language processing algorithms, such as Google Cloud Natural Language API, to analyze the converted text data, extracting themes and keywords from the speech.

[0369] Input: Text data

[0370] Output: Analysis data including themes and keywords

[0371] Specific operation: The server analyzes the text via a natural language processing API and extracts themes and keywords.

[0372] Step 4:

[0373] The server uses an emotion engine to identify the user's emotional state from the text data. For example, if a user says, "This product is completely unusable!", the emotion engine classifies the statement as the negative emotion "anger."

[0374] Input: Parsed text data

[0375] Output: Analysis data including emotional state

[0376] Specific operation: The server uses the emotion engine to determine the emotion of the text data and identify the emotional state.

[0377] Step 5:

[0378] The server retrieves the customer's past inquiry history from the database and combines it with the current text data for analysis. This identifies the customer's needs and generates a personalized solution. For example, if the inquiry is about a software crash, it generates a message providing instructions on changing settings and the latest patches.

[0379] Input: Customer text data, past inquiry history

[0380] Output: personalized solution

[0381] Specific operation: The server queries past inquiry data from the database, combines it with current text data, analyzes it, and generates a solution.

[0382] Step 6:

[0383] The server translates the generated solutions into multiple languages ​​as needed. For example, if a solution written in English is to be translated into Japanese, a translation system such as the DeepL API is used to convert it into natural language while preserving the exact meaning.

[0384] Input: personalized solutions

[0385] Output: Translated solution

[0386] Specific operation: The server calls the translation API to translate the solution text into the specified language.

[0387] Step 7:

[0388] The device displays the solution received from the server to the user, for example, "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link."

[0389] Input: translated solution

[0390] Output: The solution that is displayed to the user

[0391] Specific operation: The terminal receives the solution sent from the server and displays it on the display.

[0392] Step 8:

[0393] The user follows the system's instructions to try the proposed solution, then inputs additional questions or feedback about the solution into the terminal, which then sends the user's feedback or additional questions back to the server.

[0394] Input: User feedback and follow-up questions

[0395] Output: Additional questions and feedback sent to the server

[0396] Specific operation: The user enters additional questions or feedback into the device, which then sends it to the server.

[0397] Step 9:

[0398] The server analyzes the additional information and takes further action if necessary, and also stores the interaction data and stores it in a database for use in retraining the AI ​​model.

[0399] Input: Additional questions and feedback

[0400] Output: Additional countermeasures, saved interaction data

[0401] Specific actions: The server analyzes any additional questions or feedback, provides further solutions, and stores the interaction data in a database.

[0402] (Application example 2)

[0403] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0404] Autonomous vehicles require support systems that can quickly and effectively respond to various problems faced by drivers in real time. However, conventional systems have difficulty assessing the driver's emotional state in real time and providing personalized solutions. Furthermore, their lack of multilingual support makes it difficult to provide solutions that reflect the driver's language preferences.

[0405] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for receiving customer utterances as voice input and converting them into text data using a voice recognition system; sentiment analysis means for analyzing the text data and evaluating the customer's sentiment; means for identifying customer needs and generating personalized solutions; multilingual support means for translating the solutions into multiple languages; means for providing the translated solutions to the customer; means for saving dialogue data with the customer and retraining the AI ​​model; means for processing the driver's voice input in real time and displaying solutions on the vehicle's dashboard or head-mounted display; and means for generating and providing personalized solutions appropriate to the driver's emotional state based on the sentiment evaluation using the AI ​​model. This enables the driver to quickly obtain appropriate solutions appropriate to the driver's emotional state for problems that arise in real time inside an autonomous vehicle.

[0406] "Voice input" refers to inputting voice data spoken by a user into a terminal or system.

[0407] A "voice recognition system" is a technology or device for converting voice data into text data.

[0408] "Text data" is character string information of voice input converted by a voice recognition system.

[0409] "Emotion analysis means" refers to a technique or device that analyzes text data and identifies the user's emotional state.

[0410] The "means for identifying needs" is a technology or device that clarifies the user's requests and problems based on the content of the user's inquiry and past inquiry history.

[0411] A "personalized solution" is a specific solution or instruction provided to a user that is tailored to their specific situation and needs.

[0412] A "multilingual means" is a technique or device that translates generated solutions into multiple languages ​​and provides them to the user in the appropriate language.

[0413] "Dialogue data" refers to recorded data relating to all statements and inquiries exchanged between the user and the system.

[0414] An "artificial intelligence model" is a model built using machine learning algorithms to learn from large amounts of data and perform specific tasks.

[0415] A "dashboard" is a display device installed inside a vehicle that provides various information to the driver.

[0416] A "head-mounted display" is a device worn on the head that displays information within the field of vision.

[0417] "Driver emotional assessment" refers to the act of analyzing and identifying the driver's emotional state while driving.

[0418] "Real-time processing means" refers to technology or devices that instantly analyze and respond to user input on the spot.

[0419] A system for implementing this invention is configured as follows: First, the driver provides voice input to the system through a terminal (dashboard or head-mounted display). The terminal records the driver's voice and transmits the voice data to a server.

[0420] The server converts the voice data into text using a speech recognition system, which uses the "speech_recognition" library. For example, if a driver says, "I can't set up the navigation system in my car. What should I do?", the speech is converted into text.

[0421] The server then analyzes the converted text data using a natural language processing (NLP) algorithm and a sentiment analysis engine called "transformers" library to extract themes and keywords from the speech and identify the driver's emotional state. For example, if the driver's speech expresses frustration, the sentiment engine classifies the emotional state as "negative."

[0422] The server also retrieves the driver's past inquiry history from the database and analyzes it in combination with the current text data. This identifies the driver's needs and generates personalized solutions. For example, if a problem arises with the navigation system settings, the solution might be to "select an option from the settings menu and change the navigation settings."

[0423] The generated solutions are translated into the driver's preferred language using a multilingual system if necessary. For example, when translating an English solution into Japanese, the server converts it into natural language while preserving the exact meaning. The translated solution is then displayed on the device.

[0424] If the driver tries the proposed solutions and the problem is not resolved, they can again enter additional questions into the device via voice input. The device then sends the additional questions to the server, which analyzes them again and provides further solutions. All of this dialogue data is stored in a database and used to retrain the artificial intelligence model.

[0425] The system's unique features include its ability to assess the driver's emotional state in real time and provide personalized solutions quickly and efficiently. It also supports multiple languages, making it flexible enough for international use.

[0426] Specific examples and examples of prompts for generative AI models

[0427] A concrete example is a scenario in which a driver makes an inquiry about the settings of a navigation system.

[0428] Examples:

[0429] The driver speaks, "I can't set up the navigation system in this car. What should I do?"

[0430] Example prompt for a generative AI model:

[0431] If a user says, "I can't configure the navigation system in my car, what should I do?", convert the speech to text and analyze the sentiment. Generate the following solution: "Please reset your navigation system. Select the option from the settings menu and change your navigation settings."

[0432] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0433] Step 1:

[0434] The user speaks to the device in the autonomous vehicle, for example, saying, "I can't set up the navigation system in this car. What should I do?" This speech input is recorded by the device's microphone.

[0435] Step 2:

[0436] The device sends the recorded voice data to the server. The server receives the voice data as input and converts it into text data using a voice recognition system. The "speech_recognition" library is used for voice recognition. As a result of the conversion, the voice data is output as text data: "I can't set up the navigation system in this car. What should I do?"

[0437] Step 3:

[0438] The server analyzes the converted text data using natural language processing algorithms (NLP) and a sentiment analysis engine. Specifically, sentiment analysis is performed using the "transformers" library. From the input text data, themes and keywords are extracted, along with an assessment of the user's emotional state (e.g., frustration). The output of this step is the extracted keywords and the emotional state.

[0439] Step 4:

[0440] The server retrieves past inquiry history from the database and compares it with the analyzed text data. This identifies the driver's needs for the current problem and generates an optimal personalized solution. In this case, specific instructions for configuring the navigation system are generated. The input is the analyzed text data and emotional state, and the output is the generated solution.

[0441] Step 5:

[0442] The server translates the generated solution into the driver's preferred language using a multilingual system. For example, when translating an English solution into Japanese, it converts it into a natural expression while preserving the exact meaning. The input of this step is the generated solution (English), and the output is the translated solution (Japanese).

[0443] Step 6:

[0444] The server sends the translated solution to the device, which then displays it on the car's dashboard or head-mounted display, for example, "Select an option from the settings menu and change your navigation settings."

[0445] Step 7:

[0446] The user tries the proposed solution. If the problem is not resolved, they use voice input again to enter an additional question into the device. The device then sends this additional question back to the server. The server converts it back into text data, analyzes it, and provides additional solutions. This dialogue data is stored in a database and used to retrain the artificial intelligence model. The input is the additional question, and the output is the additional solution.

[0447] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0448] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0449] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0450] [Second embodiment]

[0451] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0452] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0453] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0454] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0455] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0456] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0457] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0458] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0459] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0460] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0461] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0462] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0463] The present invention provides a multilingual and multimodal customer service system aimed at improving customer service, and an embodiment thereof is shown below.

[0464] Program processing explanation

[0465] 1. Input and recognition of customer utterances

[0466] The user makes a voice inquiry about a problem with the product. The device records this voice and sends the data to the server. The server then uses a voice recognition system to convert the recorded voice data into text data.

[0467] 2. Sentiment and Text Analysis

[0468] The server analyzes the converted text data and uses NLP (Natural Language Processing) technology to understand the content of customer comments. It also uses a sentiment analysis system to evaluate the emotional state in the text data. For example, if a user says, "This product is completely unusable!", the sentiment analysis system will classify the comment as "anger," a negative emotion.

[0469] 3. Identifying needs and generating personalized solutions

[0470] The server retrieves past customer inquiries from a database and combines them with the current text data. It then identifies the customer's needs based on what they said and their sentiment assessment. Based on this, the AI ​​model generates a personalized solution. For example, if the customer is contacted about a software crash, it could provide instructions on how to change settings or the latest patch.

[0471] 4. Multilingual support

[0472] The server translates the generated solutions into multiple languages ​​as needed, for example, when translating a solution written in English into Japanese, it converts it into natural language while preserving the exact meaning.

[0473] 5. Providing solutions

[0474] The device displays the solution received from the server to the user, and also interactively asks the user additional questions or offers suggestions to increase engagement, such as displaying a message with instructions on how to change settings and offering a link to download the latest patch.

[0475] 6. Data Feedback and Learning

[0476] The user follows the system's instructions and attempts to solve the problem. The device sends the results and any follow-up questions to the server, which stores this interaction data and uses it to retrain the AI ​​model, improving the accuracy and personalization of responses to future inquiries.

[0477] Specific examples

[0478] Scenario: A user contacts support in Japanese.

[0479] 1. User: Asks by voice, "This software keeps crashing, what should I do?"

[0480] 2. Device: Records audio and sends it to the server.

[0481] 3. Server: Uses a speech recognition system to convert the speech into text data such as "This software keeps crashing, what should I do?"

[0482] 4. Server: Analyzes the content of the text and identifies the emotion "anger" using an emotion analysis system.

[0483] 5. Server: Retrieves past query history from the database, identifies needs, and generates solutions that provide configuration change instructions and the latest patches.

[0484] 6. Server: Translate the solution into Japanese in a multilingual system and generate the message "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link."

[0485] 7. Terminal: Display the translated solution to the user.

[0486] 8. User: Try the steps provided and enter additional questions if the problem persists.

[0487] 9. Terminal: Sends additional queries to the server.

[0488] 10. Server: Provides additional countermeasures and stores interaction data as training data.

[0489] In this way, the system of the present invention aims to improve customer satisfaction by analyzing customer comments in multiple languages ​​and multiple modalities and providing personalized solutions quickly and efficiently.

[0490] The processing flow will be explained below.

[0491] Step 1:

[0492] A user voice-inquires about a problem with a product, for example, "This software keeps crashing, what should I do?"

[0493] Step 2:

[0494] The terminal records the user's voice and transmits the voice data to the server.

[0495] Step 3:

[0496] The server inputs the received voice data into a voice recognition system and converts the voice data into text data, for example, "This software keeps crashing, what should I do?"

[0497] Step 4:

[0498] The server applies natural language processing algorithms to analyze the text data, extracting themes and keywords from the comments.

[0499] Step 5:

[0500] The server uses an emotion analysis system to identify the user's emotion from the text data, which in this case is classified as "anger."

[0501] Step 6:

[0502] The server retrieves the customer's past inquiry history from the database and searches for solutions to similar problems.

[0503] Step 7:

[0504] The server uses the results of text and sentiment analysis to identify customer needs, such as a need for a solution to a software crash.

[0505] Step 8:

[0506] The server uses AI models to generate personalized solutions for identified needs, such as instructions for changing settings or providing the latest patches.

[0507] Step 9:

[0508] The server utilizes a multilingual system to translate the generated solutions into multiple languages ​​as needed, for example, translating the solutions from Japanese to English.

[0509] Step 10:

[0510] The device displays the solution received from the server to the user, for example, "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link."

[0511] Step 11:

[0512] The user attempts the provided solution and types additional questions or feedback about the solution into the device.

[0513] Step 12:

[0514] The device sends the user's feedback and any follow-up questions back to the server.

[0515] Step 13:

[0516] The server analyzes the additional information and takes further action if necessary, and also stores the interaction data and stores it in a database for use in retraining the AI ​​model.

[0517] This series of steps enables the system of the present invention to respond quickly and efficiently to customer needs and provide a high level of personalized customer service.

[0518] Example 1

[0519] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0520] Modern companies strive to improve customer service, but implementing a multilingual, multimodal customer service system is a technically difficult challenge. Customers make inquiries in a variety of languages, and their emotions and needs vary widely, making responding to them time-consuming and labor-intensive. Conventional systems have difficulty accurately grasping customer emotions and needs, and responses tend to be uniform, failing to sufficiently increase customer satisfaction.

[0521] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0522] In this invention, the server includes means for receiving customer utterances as voice input and converting them into text data using a voice recognition system, means for analyzing the text data and evaluating customer sentiment, and means for identifying customer needs and generating personalized solutions. This makes it possible to analyze customer utterances in multiple languages ​​and multiple modalities and provide personalized solutions quickly and efficiently, thereby improving customer satisfaction.

[0523] "Voice input" is a method by which a user can query a system through speech.

[0524] A "voice recognition system" refers to a technology that analyzes recorded voice data and converts it into text data.

[0525] "Text data" refers to character information converted from audio data.

[0526] "Sentiment analysis means" refers to technology that evaluates and classifies emotional states in text data.

[0527] "Personalized solutions" are a method for providing the best possible solution to a specific problem based on the customer's needs.

[0528] "Multilingual means" refers to techniques for translating generated solutions into multiple languages.

[0529] "Interaction Data" refers to records of conversations between a customer and a system.

[0530] "Retraining methods" refers to techniques that use stored dialogue data to retrain AI models and improve their accuracy.

[0531] "Terminal" refers to the device used by the Customer to provide voice input.

[0532] "Database" refers to a repository of information where past inquiry history and other related information is stored.

[0533] "Natural language processing technology" refers to a series of technologies for analyzing text data and understanding its content.

[0534] "Engagement" refers to methods for eliciting involvement and interest through interaction with users.

[0535] An "AI model" refers to an algorithm that has been trained through machine learning to perform a specific task.

[0536] The present invention provides a multilingual and multimodal customer service system aimed at improving customer service, and an embodiment thereof is shown below.

[0537] System configuration

[0538] The system begins when a user voice-inquires about a problem with a product. When the user voice-inquires, the device records the voice and sends it to the server. The server then uses a voice recognition system to convert the voice data into text. The server then analyzes the text data, evaluates the customer's sentiment, and identifies their needs. It then uses an AI model to generate a personalized solution and translates it into multiple languages. Finally, the translated solution is provided to the user via the device.

[0539] Hardware and software used

[0540] Devices: Devices such as smartphones, computers, and tablets are used.

[0541] Server: A remote server is used for data processing and storage.

[0542] Speech recognition system: Google Cloud Speech-to-Text API, IBM Watson Speech to Text, etc. are used.

[0543] NLP technologies: Natural language processing libraries such as TensorFlow, spaCy, and NLTK are used.

[0544] Sentiment analysis system: TextBlob, NRC Emotion Lexicon, etc. are used.

[0545] Translation system: Google Cloud Translation API, DeepL API, etc. are used.

[0546] AI models: Generative AI models such as GPT and BERT are used.

[0547] Specific examples

[0548] The following are specific usage scenarios for this system:

[0549] 1. User: Makes a voice inquiry saying, "This software keeps crashing, what should I do?"

[0550] 2. Device: Records the user's voice and sends it to the server.

[0551] 3. Server: Using a speech recognition system, convert the speech into text data such as "This software keeps crashing, what should I do?"

[0552] 4. Server: Analyzes text data using NLP technology and identifies the emotion "anger."

[0553] 5. Server: Retrieves past inquiry history from a database, identifies needs, and uses AI models to generate solutions that provide configuration change instructions and the latest patches.

[0554] 6. Server: Translate the generated solution using a multilingual system. For example, the message "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link." is translated into Japanese.

[0555] 7. Terminal: Display the translated solution to the user.

[0556] Prompt Sentence Examples

[0557] Here are some examples of prompts for generative AI models:

[0558] Prompt sentences that convert user speech into text

[0559] Please convert the following audio data to text: [Audio data]

[0560] Prompt sentences that analyze the sentiment of the user's text utterances

[0561] Please rate the emotional state of the following text data: [Text data]

[0562] Prompts that identify user needs and generate solutions

[0563] Identify the user's needs and generate a solution based on the following text: [Text]

[0564] Prompts for translating generated text into multiple languages

[0565] Please translate the following text data into [language name]: [text data]

[0566] By using these procedures and prompts, the system of the present invention aims to analyze customer utterances in multiple languages ​​and multiple modalities and provide personalized solutions quickly and efficiently.

[0567] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0568] Step 1:

[0569] A user speaks to a product about a problem they are having. For example, a user might say, "This software keeps crashing. What should I do?" This is the input.

[0570] Step 2:

[0571] The device records the user's voice and sends the data to the server. The recording is performed using the microphone on a smartphone or computer. The recorded voice data is output.

[0572] Step 3:

[0573] The server uses a speech recognition system to convert the speech data into text data. Using the speech data as input, a speech recognition system such as the Google Cloud Speech-to-Text API converts the speech into text data such as "This software keeps crashing, what should I do?". The output is the text data converted from the speech.

[0574] Step 4:

[0575] The server analyzes the text data and uses NLP technology to understand the content of customer comments. It uses natural language processing tools such as TensorFlow and spaCy to analyze the structure and meaning of sentences, taking the text data as input. The output of this process is analyzed text data.

[0576] Step 5:

[0577] The server uses a sentiment analysis system to evaluate the emotional state in the text data. Using the analyzed text data as input, it uses tools such as TextBlob and the NRC Emotion Lexicon to identify the customer's emotion as "anger." The output is the analyzed sentiment data.

[0578] Step 6:

[0579] The server retrieves past query history from the database and combines it with the current text data. Using the past query history stored in the database as input, it extracts relevant data using SQL queries. The output is a combination of the past query history and the current text data.

[0580] Step 7:

[0581] The server identifies customer needs and generates personalized solutions using AI models. Using the combined text data as input, machine learning models (GPT and BERT) are used to identify needs and generate solutions. The output is a personalized solution.

[0582] Step 8:

[0583] The server translates the generated solution into multiple languages ​​as needed. The generated solution is used as input and translated using the Google Cloud Translation API or the DeepL API. The output is the translated solution in multiple languages.

[0584] Step 9:

[0585] The device displays the translated solution received from the server to the user. The translated solution is used as input and a solution message is displayed on the smartphone or PC screen. The output is the solution provided to the user.

[0586] Step 10:

[0587] The user follows the system's instructions to try to solve the problem. The user follows the displayed solution steps, changing software settings or downloading patches from links. This is the input, and the results of the setting changes and patch downloads are the output.

[0588] Step 11:

[0589] The terminal sends the results and any additional questions to the server. The terminal sends data to the server using the results of the solution execution and any new questions from the user as input. The output is the sent result data and any additional questions.

[0590] Step 12:

[0591] The server stores the dialogue data and retrains the AI ​​model. It uses newly collected dialogue data as input, stores the data, and retrains the AI ​​model. The output is a trained AI model, which improves the accuracy of responses to future inquiries.

[0592] (Application example 1)

[0593] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0594] Current customer service systems only support one language and lack multilingual support and sentiment analysis, preventing them from fully improving customer satisfaction. They also lack the means to quickly provide personalized solutions, making it difficult to respond immediately to follow-up customer inquiries. There is a need for a system that can provide multilingual and multimodal support in brick-and-mortar stores using smartphones and service robots.

[0595] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0596] In this invention, the server includes a means for converting voice input into text data using a voice recognition system, a sentiment analysis means for analyzing the text data to evaluate customer sentiment, a means for identifying customer needs and generating personalized solutions, a multilingual support means for translating the solutions into multiple languages, a means for providing the translated solutions to customers, a means for saving customer interaction data and retraining an AI model, a customer service system that functions on hardware including smartphones and service robots, and a means for providing customer service through voice, text, and translated solutions and for engaging in dialogue based on sentiment evaluation and needs, thereby enabling the provision of advanced customer service that is multilingual and emotionally responsive.

[0597] A "customer support system" is a system that responds quickly and appropriately to customer inquiries in physical stores and online.

[0598] "Voice input" is a means of acquiring customer utterances as voice data.

[0599] A "voice recognition system" is a technology or device for converting voice data into text data.

[0600] "Text data" is data expressed in the form of character information.

[0601] "Sentiment analysis means" is a means for evaluating and classifying customer emotions from text data.

[0602] "Needs identification means" are means for clarifying specific requirements or requests based on the content of a customer inquiry.

[0603] A "personalized solution" is a customized response to address a customer's specific needs.

[0604] "Multilingual support means" is a means for translating generated solutions into multiple languages.

[0605] "Interaction Data" is a record of the interaction between a customer and the system.

[0606] An "AI model" is a mathematical model for learning and prediction using artificial intelligence technology.

[0607] "Retraining methods" are methods used to improve the accuracy of existing AI models using new data.

[0608] A "smartphone" is a portable information terminal that can input voice and run applications.

[0609] A "service robot" is an autonomous robot designed to provide customer service and information.

[0610] "Multimodal" is the property of a system that includes multiple input and output formats, such as speech and text.

[0611] The present invention is a multilingual, multimodal customer service system for enhancing customer service in brick-and-mortar stores. This system combines the processes of speech recognition, emotion analysis, text analysis, needs identification, multilingual support, solution provision, data feedback, and learning. Specific embodiments for implementing the present invention are described below.

[0612] The main components of the system are the server, the terminal, and the user. The terminal consists of a smartphone or a service robot, and is responsible for receiving voice input from the user and sending it to the server.

[0613] Hardware used

[0614] Smartphone: A mobile information device that can accept voice input and run applications.

[0615] Service robot: An autonomous robot designed to provide customer service and information.

[0616] Software used

[0617] SpeechRecognition: A Python library for performing speech recognition.

[0618] transformers: A library that provides sentiment analysis, text analysis, and translation models.

[0619] Data Processing Steps

[0620] 1. Receiving and converting voice input

[0621] A user makes a voice inquiry to a smartphone or a service robot. For example, the user might say, "Please tell me how to use this product." The device records this voice and sends it to a server.

[0622] 2. Speech Recognition and Emotion Analysis

[0623] The server uses a speech recognition system (SpeechRecognition) to convert the recorded voice data into text data. The converted text data is then analyzed using sentiment analysis (transformers) to evaluate the customer's emotions. Through this process, the text is converted to "Please tell me how to use this product." and the emotion is classified as "confused."

[0624] 3. Needs Identification and Solution Generation

[0625] The server uses text analysis tools (transformers' text-classification) to identify customer needs based on the analyzed text data, retrieves personalized solutions from the database and applies them to the customer.

[0626] 4. Multilingual support

[0627] The generated solutions are translated into multiple languages ​​as needed. A multilingual solution (transformers' translation-en-to-ja) is used to properly translate from English to Japanese.

[0628] 5. Providing solutions and dialogue

[0629] The translated solution is sent to the device and displayed to the user. If the user asks additional questions, that data is also sent to the server for processing again.

[0630] 6. Data Feedback and Learning

[0631] The server stores the results of the conversation and any follow-up questions, and retrains the AI ​​model, improving its accuracy for the next inquiry.

[0632] Examples and prompts

[0633] Specific examples

[0634] 1. User: Ask, "How do I use this product?"

[0635] 2. Application: Convert speech to text and analyze sentiment and content.

[0636] 3. Application: Multilingual support and display of usage solutions.

[0637] 4. User: Enter a follow-up question.

[0638] 5. Application: Save the proposed responses and use them for future learning.

[0639] Prompt Sentence Examples

[0640] Customer: "Please tell me how to use this product."

[0641] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0642] Step 1:

[0643] Receiving and converting voice input

[0644] A user makes a voice inquiry to a smartphone or service robot. For example, the user might say, "Please tell me how to use this product." This inputs voice data into the terminal. The terminal records this voice and sends it to a server as digital voice data.

[0645] Input: User speech (e.g., "How do I use this product?")

[0646] Output: Digital audio data

[0647] Step 2:

[0648] Speech Recognition and Emotion Analysis

[0649] The server converts the received digital voice data into text data using a speech recognition system (SpeechRecognition library). For example, the voice is converted into text such as "Please tell me how to use this product." The text data is then analyzed using a sentiment analysis method (transformers' sentiment-analysis model) to evaluate the customer's emotions. As a result of the sentiment analysis, the text is classified as "confused."

[0650] Input: Digital audio data

[0651] Output: Text data (e.g., "How do I use this product?") and emotion ratings (e.g., "Confusion")

[0652] Step 3:

[0653] Needs Identification and Solution Generation

[0654] The server uses text analysis methods (Transformers' text-classification model) to identify customer needs based on the text data. For example, the tag "How to use the product" is identified. Then, it retrieves the corresponding solutions from the database and generates personalized solutions. For example, instructions on how to use the product are retrieved from the database.

[0655] Input: Text data and sentiment ratings

[0656] Output: Identified needs (e.g., how to use the product) and personalized solutions (e.g., product usage instructions)

[0657] Step 4:

[0658] Multilingual support

[0659] The server translates the generated solution as needed using multilingual support (transformers' translation-en-to-ja model). For example, a solution generated in English is translated into Japanese. This results in a solution in the target language.

[0660] Input: personalized solution (e.g. product usage instructions)

[0661] Output: Translated solution (e.g., product usage instructions in Japanese)

[0662] Step 5:

[0663] Providing solutions and dialogue

[0664] The translated solution is sent to the device, which displays it to the user and continues the dialogue via voice or text as needed. For example, if the user reviews the presented instructions and asks additional questions, their input is sent back to the server.

[0665] Input: translated solution

[0666] Output: The solution and any follow-up questions displayed to the user

[0667] Step 6:

[0668] Data Feedback and Learning

[0669] The results of the user's attempts and any follow-up questions are sent to the server, which stores this interaction data and uses it to retrain the AI ​​model, improving its accuracy for the next inquiry.

[0670] Input: User feedback and follow-up questions

[0671] Output: Dialogue data for retraining and improved AI models

[0672] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0673] The present invention provides a multilingual, multimodal customer support system aimed at improving customer service, and includes a form that combines an emotion engine. An embodiment of this system is shown below.

[0674] Program processing explanation

[0675] 1. Input and recognition of customer utterances

[0676] A user voice-inquires about a problem with a product. For example, "This software keeps crashing. What should I do?" The device records this voice and sends the voice data to the server. The server uses a voice recognition system to convert the recorded voice data into text data.

[0677] 2. Sentiment and Text Analysis

[0678] The server applies natural language processing algorithms to analyze the converted text data. This analysis extracts themes and keywords from the utterances. It then uses an emotion engine to identify the user's emotional state from the text data. For example, if a user utters, "This product is completely unusable!", the emotion engine classifies the utterance as "anger," a negative emotion. The emotion engine evaluates emotions in real time and adjusts the response as needed.

[0679] 3. Identifying needs and generating personalized solutions

[0680] The server retrieves the customer's past inquiry history from the database and combines it with the current text data for analysis. This allows it to identify the customer's needs and generate a personalized solution. For example, if the inquiry is about a software crash, it generates a message providing instructions on changing settings and the latest patches.

[0681] 4. Multilingual support

[0682] The server translates the generated solutions into multiple languages ​​as needed, for example, when translating a solution written in English into Japanese, it converts it into natural language while preserving the exact meaning.

[0683] 5. Providing solutions

[0684] The device displays the solution received from the server to the user, for example, "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link."

[0685] 6. Data Feedback and Learning

[0686] The user follows the system's instructions and attempts the proposed solution. They then enter additional questions or feedback about the solution into the device. The device then sends the user's feedback and additional questions back to the server. The server analyzes the additional information and takes further action if necessary. The interaction data is also saved and stored in a database for use in retraining the AI ​​model.

[0687] Specific examples

[0688] Scenario: A user contacts support in Japanese.

[0689] 1. User: Asks by voice, "This software keeps crashing, what should I do?"

[0690] 2. Device: Records audio and sends it to the server.

[0691] 3. Server: Uses a speech recognition system to convert the speech into text data such as "This software keeps crashing, what should I do?"

[0692] 4. Server: Analyzes the content of the text and identifies the emotion "anger" using an emotion engine.

[0693] 5. Server: Retrieves past query history from the database, identifies needs, and generates solutions that provide configuration change instructions and the latest patches.

[0694] 6. Server: Translate the solution into Japanese in a multilingual system and generate the message "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link."

[0695] 7. Terminal: Display the translated solution to the user.

[0696] 8. User: Try the steps provided and enter additional questions if the problem persists.

[0697] 9. Terminal: Sends additional queries to the server.

[0698] 10. Server: Provides additional countermeasures and stores interaction data as training data.

[0699] In this way, the system of the present invention aims to improve customer satisfaction by analyzing customer comments in multiple languages ​​and multiple modalities, using an emotion engine to evaluate the user's emotional state in real time, and providing personalized solutions quickly and efficiently.

[0700] The processing flow will be explained below.

[0701] Step 1:

[0702] A user voice-inquires about a problem with a product, for example, "This software keeps crashing, what should I do?"

[0703] Step 2:

[0704] The terminal records the user's voice and transmits the voice data to the server.

[0705] Step 3:

[0706] The server inputs the received voice data into a voice recognition system and converts the voice data into text data, for example, "This software keeps crashing, what should I do?"

[0707] Step 4:

[0708] The server then analyzes the converted text data using a natural language processing algorithm, extracting the topic and keywords of the comments and identifying the nature of the problem.

[0709] Step 5:

[0710] The server uses an emotion engine to identify the user's emotional state from the text data in real time, for example, in this case the emotion "anger."

[0711] Step 6:

[0712] The server retrieves the customer's past inquiry history from the database and collects relevant data for resolving the problem.

[0713] Step 7:

[0714] The server identifies specific customer needs based on the results of text analysis and sentiment analysis. For example, based on the frequent occurrence of software crashes, it determines that the customer needs a procedure for changing settings to prevent crashes.

[0715] Step 8:

[0716] The server uses AI models to generate personalized solutions based on identified needs, such as "Settings - Options - Crash Prevention" instructions or a "download link for the latest patch."

[0717] Step 9:

[0718] The server translates the generated solution into the required language using a multilingual system, for example translating the solution from Japanese to English.

[0719] Step 10:

[0720] The device displays the solution received from the server to the user, for example, "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link."

[0721] Step 11:

[0722] The user attempts the provided solution and enters additional questions or feedback about the solution into the terminal.

[0723] Step 12:

[0724] The device sends the user's feedback and follow-up questions to the server.

[0725] Step 13:

[0726] The server analyzes the additional information and takes further action if necessary, such as providing more detailed configuration instructions or an alternative solution, and also saves the interaction data and stores it in a database for use in retraining the AI ​​model.

[0727] Through this series of steps, the system of the present invention quickly and efficiently analyzes customer comments in multiple languages ​​and multiple modalities, and evaluates user sentiment in real time using an emotion engine, thereby significantly improving customer satisfaction by providing personalized solutions.

[0728] Example 2

[0729] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0730] Modern customer service requires multilingual and multimodal customer interactions, and it is becoming increasingly important to provide personalized solutions quickly based on customer emotions and needs. However, conventional systems struggle to accurately convert voice input into text, perform sentiment analysis, and provide solutions adapted to multiple languages, making it difficult to increase customer satisfaction. Furthermore, while continuous learning and improvement based on dialogue data is desirable, there is a lack of systems that can achieve this.

[0731] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0732] In this invention, the server includes means for receiving customer utterances as voice input and converting them into text data using speech recognition technology, sentiment analysis means for analyzing the text data and evaluating the customer's sentiment, means for identifying customer needs and generating personalized solutions, multilingual support means for translating the generated solutions into multiple languages, and means for saving customer interaction data and retraining a machine learning model. This makes it possible to analyze customer utterances in multiple languages ​​and multimodally, evaluate the user's emotional state in real time using an emotion engine, and provide personalized solutions quickly and efficiently.

[0733] Below are definitions of important terms included in the claims.

[0734] "Voice recognition technology" is a technology that analyzes voice data and converts it into text data.

[0735] "Text data" is character information converted using voice recognition technology.

[0736] An "emotion analysis means" is an algorithm or system that analyzes text data and identifies the user's emotional state.

[0737] "Means for identifying needs" refers to a technique for extracting customer requests and requirements based on the customer's past inquiry history and current statements.

[0738] "Personalized solutions" are customized solutions or guidance provided based on a customer's specific situation and emotions.

[0739] "Multilingual solutions" are technologies or systems that translate solutions into multiple languages.

[0740] "Dialogue data" is historical information about questions and answers exchanged with customers.

[0741] A "machine learning model" is an algorithm that learns from past data and makes predictions and classifications.

[0742] "Retraining" is a technique for adding new data to improve the accuracy and performance of a machine learning model.

[0743] A "terminal" is an electronic device used by a customer that inputs voice and displays text.

[0744] A "database" is a system that stores and manages past inquiry history and customer information.

[0745] The present invention provides a multilingual, multimodal customer support system aimed at improving customer service, including a form that combines an emotion engine. The following describes the embodiments of the system in detail.

[0746] System configuration

[0747] The multilingual and multimodal customer support system of the present invention is composed of the following elements: a server, a terminal, and a user.

[0748] Server: A high-performance computer running speech recognition technology, natural language processing algorithms, emotion engines, machine learning models, and translation systems. Examples include software such as Google Cloud Speech-to-Text, Google Cloud Natural Language API, and DeepL API.

[0749] Terminal: An electronic device capable of voice input and text display, such as a smartphone, tablet, or PC.

[0750] User: A customer who contacts us regarding a problem with a product.

[0751] Program processing explanation

[0752] Customer utterance input and recognition

[0753] A user voice-inquires about a problem with a product. For example, "This software keeps crashing. What should I do?" The device records this voice and sends the voice data to a server. The server then uses voice recognition technology, such as Google Cloud Speech-to-Text, to convert the recorded voice data into text.

[0754] Sentiment and Text Analysis

[0755] The server then applies natural language processing algorithms, such as the Google Cloud Natural Language API, to analyze the converted text data. This analysis extracts the topic and keywords of the utterance. It then uses an emotion engine to identify the user's emotional state from the text data. For example, if a user says, "This product is completely unusable!", the emotion engine classifies the utterance as "anger," a negative emotion. The emotion engine evaluates the emotion in real time and adjusts the response as needed.

[0756] Identifying needs and generating personalized solutions

[0757] The server retrieves the customer's past inquiry history from the database and combines it with the current text data for analysis. This allows it to identify the customer's needs and generate a personalized solution. For example, if the inquiry is about a software crash, it generates a message providing instructions on changing settings and the latest patches.

[0758] Multilingual support

[0759] The server translates the generated solutions into multiple languages ​​as needed. For example, if a solution written in English needs to be translated into Japanese, it uses a translation system such as the DeepL API to convert it into natural language while preserving the exact meaning.

[0760] Providing solutions

[0761] The device displays the solution received from the server to the user, for example, "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link."

[0762] Data Feedback and Learning

[0763] The user follows the system's instructions and attempts the proposed solution. They then enter additional questions or feedback about the solution into the device. The device then sends the user's feedback and additional questions back to the server. The server analyzes the additional information and takes further action if necessary. The interaction data is also saved and stored in a database for use in retraining the AI ​​model.

[0764] Specific examples

[0765] Scenario: A user contacts support in Japanese.

[0766] 1. User: Asks by voice, "This software keeps crashing, what should I do?"

[0767] 2. Device: Records audio and sends it to the server.

[0768] 3. Server: Use Google Cloud Speech-to-Text to convert the speech into text data such as "This software keeps crashing, what should I do?"

[0769] 4. Server: The text content is analyzed using the Google Cloud Natural Language API, and the emotion engine identifies the emotion "anger."

[0770] 5. Server: Retrieves past query history from the database, identifies needs, and generates solutions that provide configuration change instructions and the latest patches.

[0771] 6. Server: Translate the solution into Japanese using the DeepL API and generate the message "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link."

[0772] 7. Terminal: Display the translated solution to the user.

[0773] 8. User: Try the steps provided and enter any follow-up questions if the problem persists.

[0774] 9. Terminal: Sends additional queries to the server.

[0775] 10. Server: Provides additional countermeasures and stores the dialogue data as training data.

[0776] In this way, the system of the present invention aims to improve customer satisfaction by analyzing customer comments in multiple languages ​​and multiple modalities, using an emotion engine to evaluate the user's emotional state in real time, and providing personalized solutions quickly and efficiently.

[0777] Example prompt sentence:

[0778] "A user speaks to you about a software crash. Design a system that uses speech recognition to convert the text and analyzes the user's emotional state to generate a personalized solution. Describe the process for translating the generated solution into various languages ​​and displaying it to the user."

[0779] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0780] Step 1:

[0781] The user makes a voice inquiry about a problem with the product, for example, "This software keeps crashing. What should I do?" The device records this voice and sends the voice data to the server.

[0782] Input: User voice input

[0783] Output: Audio data sent to the server

[0784] Specific operation: The user uses the device's microphone to voice-input the details of the problem, and the device sends the voice data to the server.

[0785] Step 2:

[0786] The server uses voice recognition technology such as Google Cloud Speech-to-Text to convert the recorded audio data into text data.

[0787] Input: Audio data

[0788] Output: Text data

[0789] Specific operation: The server calls the speech recognition API and converts the speech data into text data.

[0790] Step 3:

[0791] The server then applies natural language processing algorithms, such as Google Cloud Natural Language API, to analyze the converted text data, extracting themes and keywords from the speech.

[0792] Input: Text data

[0793] Output: Analysis data including themes and keywords

[0794] Specific operation: The server analyzes the text via a natural language processing API and extracts themes and keywords.

[0795] Step 4:

[0796] The server uses an emotion engine to identify the user's emotional state from the text data. For example, if a user says, "This product is completely unusable!", the emotion engine classifies the statement as the negative emotion "anger."

[0797] Input: Parsed text data

[0798] Output: Analysis data including emotional state

[0799] Specific operation: The server uses the emotion engine to determine the emotion of the text data and identify the emotional state.

[0800] Step 5:

[0801] The server retrieves the customer's past inquiry history from the database and combines it with the current text data for analysis. This identifies the customer's needs and generates a personalized solution. For example, if the inquiry is about a software crash, it generates a message providing instructions on changing settings and the latest patches.

[0802] Input: Customer text data, past inquiry history

[0803] Output: personalized solution

[0804] Specific operation: The server queries past inquiry data from the database, combines it with current text data, analyzes it, and generates a solution.

[0805] Step 6:

[0806] The server translates the generated solutions into multiple languages ​​as needed. For example, if a solution written in English is to be translated into Japanese, a translation system such as the DeepL API is used to convert it into natural language while preserving the exact meaning.

[0807] Input: personalized solutions

[0808] Output: Translated solution

[0809] Specific operation: The server calls the translation API to translate the solution text into the specified language.

[0810] Step 7:

[0811] The device displays the solution received from the server to the user, for example, "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link."

[0812] Input: translated solution

[0813] Output: The solution that is displayed to the user

[0814] Specific operation: The terminal receives the solution sent from the server and displays it on the display.

[0815] Step 8:

[0816] The user follows the system's instructions to try the proposed solution, then inputs additional questions or feedback about the solution into the terminal, which then sends the user's feedback or additional questions back to the server.

[0817] Input: User feedback and follow-up questions

[0818] Output: Additional questions and feedback sent to the server

[0819] Specific operation: The user enters additional questions or feedback into the device, which then sends it to the server.

[0820] Step 9:

[0821] The server analyzes the additional information and takes further action if necessary, and also stores the interaction data and stores it in a database for use in retraining the AI ​​model.

[0822] Input: Additional questions and feedback

[0823] Output: Additional countermeasures, saved interaction data

[0824] Specific actions: The server analyzes any additional questions or feedback, provides further solutions, and stores the interaction data in a database.

[0825] (Application example 2)

[0826] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0827] Autonomous vehicles require support systems that can quickly and effectively respond to various problems faced by drivers in real time. However, conventional systems have difficulty assessing the driver's emotional state in real time and providing personalized solutions. Furthermore, their lack of multilingual support makes it difficult to provide solutions that reflect the driver's language preferences.

[0828] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for receiving customer utterances as voice input and converting them into text data using a voice recognition system; sentiment analysis means for analyzing the text data and evaluating the customer's sentiment; means for identifying customer needs and generating personalized solutions; multilingual support means for translating the solutions into multiple languages; means for providing the translated solutions to the customer; means for saving dialogue data with the customer and retraining the AI ​​model; means for processing the driver's voice input in real time and displaying solutions on the vehicle's dashboard or head-mounted display; and means for generating and providing personalized solutions appropriate to the driver's emotional state based on the sentiment evaluation using the AI ​​model. This enables the driver to quickly obtain appropriate solutions appropriate to the driver's emotional state for problems that arise in real time inside an autonomous vehicle.

[0829] "Voice input" refers to inputting voice data spoken by a user into a terminal or system.

[0830] A "voice recognition system" is a technology or device for converting voice data into text data.

[0831] "Text data" is character string information of voice input converted by a voice recognition system.

[0832] "Emotion analysis means" refers to a technique or device that analyzes text data and identifies the user's emotional state.

[0833] The "means for identifying needs" is a technology or device that clarifies the user's requests and problems based on the content of the user's inquiry and past inquiry history.

[0834] A "personalized solution" is a specific solution or instruction provided to a user that is tailored to their specific situation and needs.

[0835] A "multilingual means" is a technique or device that translates generated solutions into multiple languages ​​and provides them to the user in the appropriate language.

[0836] "Dialogue data" refers to recorded data relating to all statements and inquiries exchanged between the user and the system.

[0837] An "artificial intelligence model" is a model built using machine learning algorithms to learn from large amounts of data and perform specific tasks.

[0838] A "dashboard" is a display device installed inside a vehicle that provides various information to the driver.

[0839] A "head-mounted display" is a device worn on the head that displays information within the field of vision.

[0840] "Driver emotional assessment" refers to the act of analyzing and identifying the driver's emotional state while driving.

[0841] "Real-time processing means" refers to technology or devices that instantly analyze and respond to user input on the spot.

[0842] A system for implementing this invention is configured as follows: First, the driver provides voice input to the system through a terminal (dashboard or head-mounted display). The terminal records the driver's voice and transmits the voice data to a server.

[0843] The server converts the voice data into text using a speech recognition system, which uses the "speech_recognition" library. For example, if a driver says, "I can't set up the navigation system in my car. What should I do?", the speech is converted into text.

[0844] The server then analyzes the converted text data using a natural language processing (NLP) algorithm and a sentiment analysis engine called "transformers" library to extract themes and keywords from the speech and identify the driver's emotional state. For example, if the driver's speech expresses frustration, the sentiment engine classifies the emotional state as "negative."

[0845] The server also retrieves the driver's past inquiry history from the database and analyzes it in combination with the current text data. This identifies the driver's needs and generates personalized solutions. For example, if a problem arises with the navigation system settings, the solution might be to "select an option from the settings menu and change the navigation settings."

[0846] The generated solutions are translated into the driver's preferred language using a multilingual system if necessary. For example, when translating an English solution into Japanese, the server converts it into natural language while preserving the exact meaning. The translated solution is then displayed on the device.

[0847] If the driver tries the proposed solutions and the problem is not resolved, they can again enter additional questions into the device via voice input. The device then sends the additional questions to the server, which analyzes them again and provides further solutions. All of this dialogue data is stored in a database and used to retrain the artificial intelligence model.

[0848] The system's unique features include its ability to assess the driver's emotional state in real time and provide personalized solutions quickly and efficiently. It also supports multiple languages, making it flexible enough for international use.

[0849] Specific examples and examples of prompts for generative AI models

[0850] A concrete example is a scenario in which a driver makes an inquiry about the settings of a navigation system.

[0851] Examples:

[0852] The driver speaks, "I can't set up the navigation system in this car. What should I do?"

[0853] Example prompt for a generative AI model:

[0854] If a user says, "I can't configure the navigation system in my car, what should I do?", convert the speech to text and analyze the sentiment. Generate the following solution: "Please reset your navigation system. Select the option from the settings menu and change your navigation settings."

[0855] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0856] Step 1:

[0857] The user speaks to the device in the autonomous vehicle, for example, saying, "I can't set up the navigation system in this car. What should I do?" This speech input is recorded by the device's microphone.

[0858] Step 2:

[0859] The device sends the recorded voice data to the server. The server receives the voice data as input and converts it into text data using a voice recognition system. The "speech_recognition" library is used for voice recognition. As a result of the conversion, the voice data is output as text data: "I can't set up the navigation system in this car. What should I do?"

[0860] Step 3:

[0861] The server analyzes the converted text data using natural language processing algorithms (NLP) and a sentiment analysis engine. Specifically, sentiment analysis is performed using the "transformers" library. From the input text data, themes and keywords are extracted, along with an assessment of the user's emotional state (e.g., frustration). The output of this step is the extracted keywords and the emotional state.

[0862] Step 4:

[0863] The server retrieves past inquiry history from the database and compares it with the analyzed text data. This identifies the driver's needs for the current problem and generates an optimal personalized solution. In this case, specific instructions for configuring the navigation system are generated. The input is the analyzed text data and emotional state, and the output is the generated solution.

[0864] Step 5:

[0865] The server translates the generated solution into the driver's preferred language using a multilingual system. For example, when translating an English solution into Japanese, it converts it into a natural expression while preserving the exact meaning. The input of this step is the generated solution (English), and the output is the translated solution (Japanese).

[0866] Step 6:

[0867] The server sends the translated solution to the device, which then displays it on the car's dashboard or head-mounted display, for example, "Select an option from the settings menu and change your navigation settings."

[0868] Step 7:

[0869] The user tries the proposed solution. If the problem is not resolved, they use voice input again to enter an additional question into the device. The device then sends this additional question back to the server. The server converts it back into text data, analyzes it, and provides additional solutions. This dialogue data is stored in a database and used to retrain the artificial intelligence model. The input is the additional question, and the output is the additional solution.

[0870] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0871] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0872] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0873] [Third embodiment]

[0874] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0875] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0876] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0877] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0878] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0879] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0880] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0881] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0882] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0883] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0884] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0885] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0886] The present invention provides a multilingual and multimodal customer service system aimed at improving customer service, and an embodiment thereof is shown below.

[0887] Program processing explanation

[0888] 1. Input and recognition of customer utterances

[0889] The user makes a voice inquiry about a problem with the product. The device records this voice and sends the data to the server. The server then uses a voice recognition system to convert the recorded voice data into text data.

[0890] 2. Sentiment and Text Analysis

[0891] The server analyzes the converted text data and uses NLP (Natural Language Processing) technology to understand the content of customer comments. It also uses a sentiment analysis system to evaluate the emotional state in the text data. For example, if a user says, "This product is completely unusable!", the sentiment analysis system will classify the comment as "anger," a negative emotion.

[0892] 3. Identifying needs and generating personalized solutions

[0893] The server retrieves past customer inquiries from a database and combines them with the current text data. It then identifies the customer's needs based on what they said and their sentiment assessment. Based on this, the AI ​​model generates a personalized solution. For example, if the customer is contacted about a software crash, it could provide instructions on how to change settings or the latest patch.

[0894] 4. Multilingual support

[0895] The server translates the generated solutions into multiple languages ​​as needed, for example, when translating a solution written in English into Japanese, it converts it into natural language while preserving the exact meaning.

[0896] 5. Providing solutions

[0897] The device displays the solution received from the server to the user, and also interactively asks the user additional questions or offers suggestions to increase engagement, such as displaying a message with instructions on how to change settings and offering a link to download the latest patch.

[0898] 6. Data Feedback and Learning

[0899] The user follows the system's instructions and attempts to solve the problem. The device sends the results and any follow-up questions to the server, which stores this interaction data and uses it to retrain the AI ​​model, improving the accuracy and personalization of responses to future inquiries.

[0900] Specific examples

[0901] Scenario: A user contacts support in Japanese.

[0902] 1. User: Asks by voice, "This software keeps crashing, what should I do?"

[0903] 2. Device: Records audio and sends it to the server.

[0904] 3. Server: Uses a speech recognition system to convert the speech into text data such as "This software keeps crashing, what should I do?"

[0905] 4. Server: Analyzes the content of the text and identifies the emotion "anger" using an emotion analysis system.

[0906] 5. Server: Retrieves past query history from the database, identifies needs, and generates solutions that provide configuration change instructions and the latest patches.

[0907] 6. Server: Translate the solution into Japanese in a multilingual system and generate the message "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link."

[0908] 7. Terminal: Display the translated solution to the user.

[0909] 8. User: Try the steps provided and enter additional questions if the problem persists.

[0910] 9. Terminal: Sends additional queries to the server.

[0911] 10. Server: Provides additional countermeasures and stores interaction data as training data.

[0912] In this way, the system of the present invention aims to improve customer satisfaction by analyzing customer comments in multiple languages ​​and multiple modalities and providing personalized solutions quickly and efficiently.

[0913] The processing flow will be explained below.

[0914] Step 1:

[0915] A user voice-inquires about a problem with a product, for example, "This software keeps crashing, what should I do?"

[0916] Step 2:

[0917] The terminal records the user's voice and transmits the voice data to the server.

[0918] Step 3:

[0919] The server inputs the received voice data into a voice recognition system and converts the voice data into text data, for example, "This software keeps crashing, what should I do?"

[0920] Step 4:

[0921] The server applies natural language processing algorithms to analyze the text data, extracting themes and keywords from the comments.

[0922] Step 5:

[0923] The server uses an emotion analysis system to identify the user's emotion from the text data, which in this case is classified as "anger."

[0924] Step 6:

[0925] The server retrieves the customer's past inquiry history from the database and searches for solutions to similar problems.

[0926] Step 7:

[0927] The server uses the results of text and sentiment analysis to identify customer needs, such as a need for a solution to a software crash.

[0928] Step 8:

[0929] The server uses AI models to generate personalized solutions for identified needs, such as instructions for changing settings or providing the latest patches.

[0930] Step 9:

[0931] The server utilizes a multilingual system to translate the generated solutions into multiple languages ​​as needed, for example, translating the solutions from Japanese to English.

[0932] Step 10:

[0933] The device displays the solution received from the server to the user, for example, "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link."

[0934] Step 11:

[0935] The user attempts the provided solution and types additional questions or feedback about the solution into the device.

[0936] Step 12:

[0937] The device sends the user's feedback and any follow-up questions back to the server.

[0938] Step 13:

[0939] The server analyzes the additional information and takes further action if necessary, and also stores the interaction data and stores it in a database for use in retraining the AI ​​model.

[0940] This series of steps enables the system of the present invention to respond quickly and efficiently to customer needs and provide a high level of personalized customer service.

[0941] Example 1

[0942] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0943] Modern companies strive to improve customer service, but implementing a multilingual, multimodal customer service system is a technically difficult challenge. Customers make inquiries in a variety of languages, and their emotions and needs vary widely, making responding to them time-consuming and labor-intensive. Conventional systems have difficulty accurately grasping customer emotions and needs, and responses tend to be uniform, failing to sufficiently increase customer satisfaction.

[0944] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0945] In this invention, the server includes means for receiving customer utterances as voice input and converting them into text data using a voice recognition system, means for analyzing the text data and evaluating customer sentiment, and means for identifying customer needs and generating personalized solutions. This makes it possible to analyze customer utterances in multiple languages ​​and multiple modalities and provide personalized solutions quickly and efficiently, thereby improving customer satisfaction.

[0946] "Voice input" is a method by which a user can query a system through speech.

[0947] A "voice recognition system" refers to a technology that analyzes recorded voice data and converts it into text data.

[0948] "Text data" refers to character information converted from audio data.

[0949] "Sentiment analysis means" refers to technology that evaluates and classifies emotional states in text data.

[0950] "Personalized solutions" are a method for providing the best possible solution to a specific problem based on the customer's needs.

[0951] "Multilingual means" refers to techniques for translating generated solutions into multiple languages.

[0952] "Interaction Data" refers to records of conversations between a customer and a system.

[0953] "Retraining methods" refers to techniques that use stored dialogue data to retrain AI models and improve their accuracy.

[0954] "Terminal" refers to the device used by the Customer to provide voice input.

[0955] "Database" refers to a repository of information where past inquiry history and other related information is stored.

[0956] "Natural language processing technology" refers to a series of technologies for analyzing text data and understanding its content.

[0957] "Engagement" refers to methods for eliciting involvement and interest through interaction with users.

[0958] An "AI model" refers to an algorithm that has been trained through machine learning to perform a specific task.

[0959] The present invention provides a multilingual and multimodal customer service system aimed at improving customer service, and an embodiment thereof is shown below.

[0960] System configuration

[0961] The system begins when a user voice-inquires about a problem with a product. When the user voice-inquires, the device records the voice and sends it to the server. The server then uses a voice recognition system to convert the voice data into text. The server then analyzes the text data, evaluates the customer's sentiment, and identifies their needs. It then uses an AI model to generate a personalized solution and translates it into multiple languages. Finally, the translated solution is provided to the user via the device.

[0962] Hardware and software used

[0963] Devices: Devices such as smartphones, computers, and tablets are used.

[0964] Server: A remote server is used for data processing and storage.

[0965] Speech recognition system: Google Cloud Speech-to-Text API, IBM Watson Speech to Text, etc. are used.

[0966] NLP technologies: Natural language processing libraries such as TensorFlow, spaCy, and NLTK are used.

[0967] Sentiment analysis system: TextBlob, NRC Emotion Lexicon, etc. are used.

[0968] Translation system: Google Cloud Translation API, DeepL API, etc. are used.

[0969] AI models: Generative AI models such as GPT and BERT are used.

[0970] Specific examples

[0971] The following are specific usage scenarios for this system:

[0972] 1. User: Makes a voice inquiry saying, "This software keeps crashing, what should I do?"

[0973] 2. Device: Records the user's voice and sends it to the server.

[0974] 3. Server: Using a speech recognition system, convert the speech into text data such as "This software keeps crashing, what should I do?"

[0975] 4. Server: Analyzes text data using NLP technology and identifies the emotion "anger."

[0976] 5. Server: Retrieves past inquiry history from a database, identifies needs, and uses AI models to generate solutions that provide configuration change instructions and the latest patches.

[0977] 6. Server: Translate the generated solution using a multilingual system. For example, the message "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link." is translated into Japanese.

[0978] 7. Terminal: Display the translated solution to the user.

[0979] Prompt Sentence Examples

[0980] Here are some examples of prompts for generative AI models:

[0981] Prompt sentences that convert user speech into text

[0982] Please convert the following audio data to text: [Audio data]

[0983] Prompt sentences that analyze the sentiment of the user's text utterances

[0984] Please rate the emotional state of the following text data: [Text data]

[0985] Prompts that identify user needs and generate solutions

[0986] Identify the user's needs and generate a solution based on the following text: [Text]

[0987] Prompts for translating generated text into multiple languages

[0988] Please translate the following text data into [language name]: [text data]

[0989] By using these procedures and prompts, the system of the present invention aims to analyze customer utterances in multiple languages ​​and multiple modalities and provide personalized solutions quickly and efficiently.

[0990] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0991] Step 1:

[0992] A user speaks to a product about a problem they are having. For example, a user might say, "This software keeps crashing. What should I do?" This is the input.

[0993] Step 2:

[0994] The device records the user's voice and sends the data to the server. The recording is performed using the microphone on a smartphone or computer. The recorded voice data is output.

[0995] Step 3:

[0996] The server uses a speech recognition system to convert the speech data into text data. Using the speech data as input, a speech recognition system such as the Google Cloud Speech-to-Text API converts the speech into text data such as "This software keeps crashing, what should I do?". The output is the text data converted from the speech.

[0997] Step 4:

[0998] The server analyzes the text data and uses NLP technology to understand the content of customer comments. It uses natural language processing tools such as TensorFlow and spaCy to analyze the structure and meaning of sentences, taking the text data as input. The output of this process is analyzed text data.

[0999] Step 5:

[1000] The server uses a sentiment analysis system to evaluate the emotional state in the text data. Using the analyzed text data as input, it uses tools such as TextBlob and the NRC Emotion Lexicon to identify the customer's emotion as "anger." The output is the analyzed sentiment data.

[1001] Step 6:

[1002] The server retrieves past query history from the database and combines it with the current text data. Using the past query history stored in the database as input, it extracts relevant data using SQL queries. The output is a combination of the past query history and the current text data.

[1003] Step 7:

[1004] The server identifies customer needs and generates personalized solutions using AI models. Using the combined text data as input, machine learning models (GPT and BERT) are used to identify needs and generate solutions. The output is a personalized solution.

[1005] Step 8:

[1006] The server translates the generated solution into multiple languages ​​as needed. The generated solution is used as input and translated using the Google Cloud Translation API or the DeepL API. The output is the translated solution in multiple languages.

[1007] Step 9:

[1008] The device displays the translated solution received from the server to the user. The translated solution is used as input and a solution message is displayed on the smartphone or PC screen. The output is the solution provided to the user.

[1009] Step 10:

[1010] The user follows the system's instructions to try to solve the problem. The user follows the displayed solution steps, changing software settings or downloading patches from links. This is the input, and the results of the setting changes and patch downloads are the output.

[1011] Step 11:

[1012] The terminal sends the results and any additional questions to the server. The terminal sends data to the server using the results of the solution execution and any new questions from the user as input. The output is the sent result data and any additional questions.

[1013] Step 12:

[1014] The server stores the dialogue data and retrains the AI ​​model. It uses newly collected dialogue data as input, stores the data, and retrains the AI ​​model. The output is a trained AI model, which improves the accuracy of responses to future inquiries.

[1015] (Application example 1)

[1016] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1017] Current customer service systems only support one language and lack multilingual support and sentiment analysis, preventing them from fully improving customer satisfaction. They also lack the means to quickly provide personalized solutions, making it difficult to respond immediately to follow-up customer inquiries. There is a need for a system that can provide multilingual and multimodal support in brick-and-mortar stores using smartphones and service robots.

[1018] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1019] In this invention, the server includes a means for converting voice input into text data using a voice recognition system, a sentiment analysis means for analyzing the text data to evaluate customer sentiment, a means for identifying customer needs and generating personalized solutions, a multilingual support means for translating the solutions into multiple languages, a means for providing the translated solutions to customers, a means for saving customer interaction data and retraining an AI model, a customer service system that functions on hardware including smartphones and service robots, and a means for providing customer service through voice, text, and translated solutions and for engaging in dialogue based on sentiment evaluation and needs, thereby enabling the provision of advanced customer service that is multilingual and emotionally responsive.

[1020] A "customer support system" is a system that responds quickly and appropriately to customer inquiries in physical stores and online.

[1021] "Voice input" is a means of acquiring customer utterances as voice data.

[1022] A "voice recognition system" is a technology or device for converting voice data into text data.

[1023] "Text data" is data expressed in the form of character information.

[1024] "Sentiment analysis means" is a means for evaluating and classifying customer emotions from text data.

[1025] "Needs identification means" are means for clarifying specific requirements or requests based on the content of a customer inquiry.

[1026] A "personalized solution" is a customized response to address a customer's specific needs.

[1027] "Multilingual support means" is a means for translating generated solutions into multiple languages.

[1028] "Interaction Data" is a record of the interaction between a customer and the system.

[1029] An "AI model" is a mathematical model for learning and prediction using artificial intelligence technology.

[1030] "Retraining methods" are methods used to improve the accuracy of existing AI models using new data.

[1031] A "smartphone" is a portable information terminal that can input voice and run applications.

[1032] A "service robot" is an autonomous robot designed to provide customer service and information.

[1033] "Multimodal" is the property of a system that includes multiple input and output formats, such as speech and text.

[1034] The present invention is a multilingual, multimodal customer service system for enhancing customer service in brick-and-mortar stores. This system combines the processes of speech recognition, emotion analysis, text analysis, needs identification, multilingual support, solution provision, data feedback, and learning. Specific embodiments for implementing the present invention are described below.

[1035] The main components of the system are the server, the terminal, and the user. The terminal consists of a smartphone or a service robot, and is responsible for receiving voice input from the user and sending it to the server.

[1036] Hardware used

[1037] Smartphone: A mobile information device that can accept voice input and run applications.

[1038] Service robot: An autonomous robot designed to provide customer service and information.

[1039] Software used

[1040] SpeechRecognition: A Python library for performing speech recognition.

[1041] transformers: A library that provides sentiment analysis, text analysis, and translation models.

[1042] Data Processing Steps

[1043] 1. Receiving and converting voice input

[1044] A user makes a voice inquiry to a smartphone or a service robot. For example, the user might say, "Please tell me how to use this product." The device records this voice and sends it to a server.

[1045] 2. Speech Recognition and Emotion Analysis

[1046] The server uses a speech recognition system (SpeechRecognition) to convert the recorded voice data into text data. The converted text data is then analyzed using sentiment analysis (transformers) to evaluate the customer's emotions. Through this process, the text is converted to "Please tell me how to use this product." and the emotion is classified as "confused."

[1047] 3. Needs Identification and Solution Generation

[1048] The server uses text analysis tools (transformers' text-classification) to identify customer needs based on the analyzed text data, retrieves personalized solutions from the database and applies them to the customer.

[1049] 4. Multilingual support

[1050] The generated solutions are translated into multiple languages ​​as needed. A multilingual solution (transformers' translation-en-to-ja) is used to properly translate from English to Japanese.

[1051] 5. Providing solutions and dialogue

[1052] The translated solution is sent to the device and displayed to the user. If the user asks additional questions, that data is also sent to the server for processing again.

[1053] 6. Data Feedback and Learning

[1054] The server stores the results of the conversation and any follow-up questions, and retrains the AI ​​model, improving its accuracy for the next inquiry.

[1055] Examples and prompts

[1056] Specific examples

[1057] 1. User: Ask, "How do I use this product?"

[1058] 2. Application: Convert speech to text and analyze sentiment and content.

[1059] 3. Application: Multilingual support and display of usage solutions.

[1060] 4. User: Enter a follow-up question.

[1061] 5. Application: Save the proposed responses and use them for future learning.

[1062] Prompt Sentence Examples

[1063] Customer: "Please tell me how to use this product."

[1064] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1065] Step 1:

[1066] Receiving and converting voice input

[1067] A user makes a voice inquiry to a smartphone or service robot. For example, the user might say, "Please tell me how to use this product." This inputs voice data into the terminal. The terminal records this voice and sends it to a server as digital voice data.

[1068] Input: User speech (e.g., "How do I use this product?")

[1069] Output: Digital audio data

[1070] Step 2:

[1071] Speech Recognition and Emotion Analysis

[1072] The server converts the received digital voice data into text data using a speech recognition system (SpeechRecognition library). For example, the voice is converted into text such as "Please tell me how to use this product." The text data is then analyzed using a sentiment analysis method (transformers' sentiment-analysis model) to evaluate the customer's emotions. As a result of the sentiment analysis, the text is classified as "confused."

[1073] Input: Digital audio data

[1074] Output: Text data (e.g., "How do I use this product?") and emotion ratings (e.g., "Confusion")

[1075] Step 3:

[1076] Needs Identification and Solution Generation

[1077] The server uses text analysis methods (Transformers' text-classification model) to identify customer needs based on the text data. For example, the tag "How to use the product" is identified. Then, it retrieves the corresponding solutions from the database and generates personalized solutions. For example, instructions on how to use the product are retrieved from the database.

[1078] Input: Text data and sentiment ratings

[1079] Output: Identified needs (e.g., how to use the product) and personalized solutions (e.g., product usage instructions)

[1080] Step 4:

[1081] Multilingual support

[1082] The server translates the generated solution as needed using multilingual support (transformers' translation-en-to-ja model). For example, a solution generated in English is translated into Japanese. This results in a solution in the target language.

[1083] Input: personalized solution (e.g. product usage instructions)

[1084] Output: Translated solution (e.g., product usage instructions in Japanese)

[1085] Step 5:

[1086] Providing solutions and dialogue

[1087] The translated solution is sent to the device, which displays it to the user and continues the dialogue via voice or text as needed. For example, if the user reviews the presented instructions and asks additional questions, their input is sent back to the server.

[1088] Input: translated solution

[1089] Output: The solution and any follow-up questions displayed to the user

[1090] Step 6:

[1091] Data Feedback and Learning

[1092] The results of the user's attempts and any follow-up questions are sent to the server, which stores this interaction data and uses it to retrain the AI ​​model, improving its accuracy for the next inquiry.

[1093] Input: User feedback and follow-up questions

[1094] Output: Dialogue data for retraining and improved AI models

[1095] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1096] The present invention provides a multilingual, multimodal customer support system aimed at improving customer service, and includes a form that combines an emotion engine. An embodiment of this system is shown below.

[1097] Program processing explanation

[1098] 1. Input and recognition of customer utterances

[1099] A user voice-inquires about a problem with a product. For example, "This software keeps crashing. What should I do?" The device records this voice and sends the voice data to the server. The server uses a voice recognition system to convert the recorded voice data into text data.

[1100] 2. Sentiment and Text Analysis

[1101] The server applies natural language processing algorithms to analyze the converted text data. This analysis extracts themes and keywords from the utterances. It then uses an emotion engine to identify the user's emotional state from the text data. For example, if a user utters, "This product is completely unusable!", the emotion engine classifies the utterance as "anger," a negative emotion. The emotion engine evaluates emotions in real time and adjusts the response as needed.

[1102] 3. Identifying needs and generating personalized solutions

[1103] The server retrieves the customer's past inquiry history from the database and combines it with the current text data for analysis. This allows it to identify the customer's needs and generate a personalized solution. For example, if the inquiry is about a software crash, it generates a message providing instructions on changing settings and the latest patches.

[1104] 4. Multilingual support

[1105] The server translates the generated solutions into multiple languages ​​as needed, for example, when translating a solution written in English into Japanese, it converts it into natural language while preserving the exact meaning.

[1106] 5. Providing solutions

[1107] The device displays the solution received from the server to the user, for example, "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link."

[1108] 6. Data Feedback and Learning

[1109] The user follows the system's instructions and attempts the proposed solution. They then enter additional questions or feedback about the solution into the device. The device then sends the user's feedback and additional questions back to the server. The server analyzes the additional information and takes further action if necessary. The interaction data is also saved and stored in a database for use in retraining the AI ​​model.

[1110] Specific examples

[1111] Scenario: A user contacts support in Japanese.

[1112] 1. User: Asks by voice, "This software keeps crashing, what should I do?"

[1113] 2. Device: Records audio and sends it to the server.

[1114] 3. Server: Uses a speech recognition system to convert the speech into text data such as "This software keeps crashing, what should I do?"

[1115] 4. Server: Analyzes the content of the text and identifies the emotion "anger" using an emotion engine.

[1116] 5. Server: Retrieves past query history from the database, identifies needs, and generates solutions that provide configuration change instructions and the latest patches.

[1117] 6. Server: Translate the solution into Japanese in a multilingual system and generate the message "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link."

[1118] 7. Terminal: Display the translated solution to the user.

[1119] 8. User: Try the steps provided and enter additional questions if the problem persists.

[1120] 9. Terminal: Sends additional queries to the server.

[1121] 10. Server: Provides additional countermeasures and stores interaction data as training data.

[1122] In this way, the system of the present invention aims to improve customer satisfaction by analyzing customer comments in multiple languages ​​and multiple modalities, using an emotion engine to evaluate the user's emotional state in real time, and providing personalized solutions quickly and efficiently.

[1123] The processing flow will be explained below.

[1124] Step 1:

[1125] A user voice-inquires about a problem with a product, for example, "This software keeps crashing, what should I do?"

[1126] Step 2:

[1127] The terminal records the user's voice and transmits the voice data to the server.

[1128] Step 3:

[1129] The server inputs the received voice data into a voice recognition system and converts the voice data into text data, for example, "This software keeps crashing, what should I do?"

[1130] Step 4:

[1131] The server then analyzes the converted text data using a natural language processing algorithm, extracting the topic and keywords of the comments and identifying the nature of the problem.

[1132] Step 5:

[1133] The server uses an emotion engine to identify the user's emotional state from the text data in real time, for example, in this case the emotion "anger."

[1134] Step 6:

[1135] The server retrieves the customer's past inquiry history from the database and collects relevant data for resolving the problem.

[1136] Step 7:

[1137] The server identifies specific customer needs based on the results of text analysis and sentiment analysis. For example, based on the frequent occurrence of software crashes, it determines that the customer needs a procedure for changing settings to prevent crashes.

[1138] Step 8:

[1139] The server uses AI models to generate personalized solutions based on identified needs, such as "Settings - Options - Crash Prevention" instructions or a "download link for the latest patch."

[1140] Step 9:

[1141] The server translates the generated solution into the required language using a multilingual system, for example translating the solution from Japanese to English.

[1142] Step 10:

[1143] The device displays the solution received from the server to the user, for example, "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link."

[1144] Step 11:

[1145] The user attempts the provided solution and enters additional questions or feedback about the solution into the terminal.

[1146] Step 12:

[1147] The device sends the user's feedback and follow-up questions to the server.

[1148] Step 13:

[1149] The server analyzes the additional information and takes further action if necessary, such as providing more detailed configuration instructions or an alternative solution, and also saves the interaction data and stores it in a database for use in retraining the AI ​​model.

[1150] Through this series of steps, the system of the present invention quickly and efficiently analyzes customer comments in multiple languages ​​and multiple modalities, and evaluates user sentiment in real time using an emotion engine, thereby significantly improving customer satisfaction by providing personalized solutions.

[1151] Example 2

[1152] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1153] Modern customer service requires multilingual and multimodal customer interactions, and it is becoming increasingly important to provide personalized solutions quickly based on customer emotions and needs. However, conventional systems struggle to accurately convert voice input into text, perform sentiment analysis, and provide solutions adapted to multiple languages, making it difficult to increase customer satisfaction. Furthermore, while continuous learning and improvement based on dialogue data is desirable, there is a lack of systems that can achieve this.

[1154] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1155] In this invention, the server includes means for receiving customer utterances as voice input and converting them into text data using speech recognition technology, sentiment analysis means for analyzing the text data and evaluating the customer's sentiment, means for identifying customer needs and generating personalized solutions, multilingual support means for translating the generated solutions into multiple languages, and means for saving customer interaction data and retraining a machine learning model. This makes it possible to analyze customer utterances in multiple languages ​​and multimodally, evaluate the user's emotional state in real time using an emotion engine, and provide personalized solutions quickly and efficiently.

[1156] Below are definitions of important terms included in the claims.

[1157] "Voice recognition technology" is a technology that analyzes voice data and converts it into text data.

[1158] "Text data" is character information converted using voice recognition technology.

[1159] An "emotion analysis means" is an algorithm or system that analyzes text data and identifies the user's emotional state.

[1160] "Means for identifying needs" refers to a technique for extracting customer requests and requirements based on the customer's past inquiry history and current statements.

[1161] "Personalized solutions" are customized solutions or guidance provided based on a customer's specific situation and emotions.

[1162] "Multilingual solutions" are technologies or systems that translate solutions into multiple languages.

[1163] "Dialogue data" is historical information about questions and answers exchanged with customers.

[1164] A "machine learning model" is an algorithm that learns from past data and makes predictions and classifications.

[1165] "Retraining" is a technique for adding new data to improve the accuracy and performance of a machine learning model.

[1166] A "terminal" is an electronic device used by a customer that inputs voice and displays text.

[1167] A "database" is a system that stores and manages past inquiry history and customer information.

[1168] The present invention provides a multilingual, multimodal customer support system aimed at improving customer service, including a form that combines an emotion engine. The following describes the embodiments of the system in detail.

[1169] System configuration

[1170] The multilingual and multimodal customer support system of the present invention is composed of the following elements: a server, a terminal, and a user.

[1171] Server: A high-performance computer running speech recognition technology, natural language processing algorithms, emotion engines, machine learning models, and translation systems. Examples include software such as Google Cloud Speech-to-Text, Google Cloud Natural Language API, and DeepL API.

[1172] Terminal: An electronic device capable of voice input and text display, such as a smartphone, tablet, or PC.

[1173] User: A customer who contacts us regarding a problem with a product.

[1174] Program processing explanation

[1175] Customer utterance input and recognition

[1176] A user voice-inquires about a problem with a product. For example, "This software keeps crashing. What should I do?" The device records this voice and sends the voice data to a server. The server then uses voice recognition technology, such as Google Cloud Speech-to-Text, to convert the recorded voice data into text.

[1177] Sentiment and Text Analysis

[1178] The server then applies natural language processing algorithms, such as the Google Cloud Natural Language API, to analyze the converted text data. This analysis extracts the topic and keywords of the utterance. It then uses an emotion engine to identify the user's emotional state from the text data. For example, if a user says, "This product is completely unusable!", the emotion engine classifies the utterance as "anger," a negative emotion. The emotion engine evaluates the emotion in real time and adjusts the response as needed.

[1179] Identifying needs and generating personalized solutions

[1180] The server retrieves the customer's past inquiry history from the database and combines it with the current text data for analysis. This allows it to identify the customer's needs and generate a personalized solution. For example, if the inquiry is about a software crash, it generates a message providing instructions on changing settings and the latest patches.

[1181] Multilingual support

[1182] The server translates the generated solutions into multiple languages ​​as needed. For example, if a solution written in English needs to be translated into Japanese, it uses a translation system such as the DeepL API to convert it into natural language while preserving the exact meaning.

[1183] Providing solutions

[1184] The device displays the solution received from the server to the user, for example, "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link."

[1185] Data Feedback and Learning

[1186] The user follows the system's instructions and attempts the proposed solution. They then enter additional questions or feedback about the solution into the device. The device then sends the user's feedback and additional questions back to the server. The server analyzes the additional information and takes further action if necessary. The interaction data is also saved and stored in a database for use in retraining the AI ​​model.

[1187] Specific examples

[1188] Scenario: A user contacts support in Japanese.

[1189] 1. User: Asks by voice, "This software keeps crashing, what should I do?"

[1190] 2. Device: Records audio and sends it to the server.

[1191] 3. Server: Use Google Cloud Speech-to-Text to convert the speech into text data such as "This software keeps crashing, what should I do?"

[1192] 4. Server: The text content is analyzed using the Google Cloud Natural Language API, and the emotion engine identifies the emotion "anger."

[1193] 5. Server: Retrieves past query history from the database, identifies needs, and generates solutions that provide configuration change instructions and the latest patches.

[1194] 6. Server: Translate the solution into Japanese using the DeepL API and generate the message "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link."

[1195] 7. Terminal: Display the translated solution to the user.

[1196] 8. User: Try the steps provided and enter any follow-up questions if the problem persists.

[1197] 9. Terminal: Sends additional queries to the server.

[1198] 10. Server: Provides additional countermeasures and stores the dialogue data as training data.

[1199] In this way, the system of the present invention aims to improve customer satisfaction by analyzing customer comments in multiple languages ​​and multiple modalities, using an emotion engine to evaluate the user's emotional state in real time, and providing personalized solutions quickly and efficiently.

[1200] Example prompt sentence:

[1201] "A user speaks to you about a software crash. Design a system that uses speech recognition to convert the text and analyzes the user's emotional state to generate a personalized solution. Describe the process for translating the generated solution into various languages ​​and displaying it to the user."

[1202] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1203] Step 1:

[1204] The user makes a voice inquiry about a problem with the product, for example, "This software keeps crashing. What should I do?" The device records this voice and sends the voice data to the server.

[1205] Input: User voice input

[1206] Output: Audio data sent to the server

[1207] Specific operation: The user uses the device's microphone to voice-input the details of the problem, and the device sends the voice data to the server.

[1208] Step 2:

[1209] The server uses voice recognition technology such as Google Cloud Speech-to-Text to convert the recorded audio data into text data.

[1210] Input: Audio data

[1211] Output: Text data

[1212] Specific operation: The server calls the speech recognition API and converts the speech data into text data.

[1213] Step 3:

[1214] The server then applies natural language processing algorithms, such as Google Cloud Natural Language API, to analyze the converted text data, extracting themes and keywords from the speech.

[1215] Input: Text data

[1216] Output: Analysis data including themes and keywords

[1217] Specific operation: The server analyzes the text via a natural language processing API and extracts themes and keywords.

[1218] Step 4:

[1219] The server uses an emotion engine to identify the user's emotional state from the text data. For example, if a user says, "This product is completely unusable!", the emotion engine classifies the statement as the negative emotion "anger."

[1220] Input: Parsed text data

[1221] Output: Analysis data including emotional state

[1222] Specific operation: The server uses the emotion engine to determine the emotion of the text data and identify the emotional state.

[1223] Step 5:

[1224] The server retrieves the customer's past inquiry history from the database and combines it with the current text data for analysis. This identifies the customer's needs and generates a personalized solution. For example, if the inquiry is about a software crash, it generates a message providing instructions on changing settings and the latest patches.

[1225] Input: Customer text data, past inquiry history

[1226] Output: personalized solution

[1227] Specific operation: The server queries past inquiry data from the database, combines it with current text data, analyzes it, and generates a solution.

[1228] Step 6:

[1229] The server translates the generated solutions into multiple languages ​​as needed. For example, if a solution written in English is to be translated into Japanese, a translation system such as the DeepL API is used to convert it into natural language while preserving the exact meaning.

[1230] Input: personalized solutions

[1231] Output: Translated solution

[1232] Specific operation: The server calls the translation API to translate the solution text into the specified language.

[1233] Step 7:

[1234] The device displays the solution received from the server to the user, for example, "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link."

[1235] Input: translated solution

[1236] Output: The solution that is displayed to the user

[1237] Specific operation: The terminal receives the solution sent from the server and displays it on the display.

[1238] Step 8:

[1239] The user follows the system's instructions to try the proposed solution, then inputs additional questions or feedback about the solution into the terminal, which then sends the user's feedback or additional questions back to the server.

[1240] Input: User feedback and follow-up questions

[1241] Output: Additional questions and feedback sent to the server

[1242] Specific operation: The user enters additional questions or feedback into the device, which then sends it to the server.

[1243] Step 9:

[1244] The server analyzes the additional information and takes further action if necessary, and also stores the interaction data and stores it in a database for use in retraining the AI ​​model.

[1245] Input: Additional questions and feedback

[1246] Output: Additional countermeasures, saved interaction data

[1247] Specific actions: The server analyzes any additional questions or feedback, provides further solutions, and stores the interaction data in a database.

[1248] (Application example 2)

[1249] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1250] Autonomous vehicles require support systems that can quickly and effectively respond to various problems faced by drivers in real time. However, conventional systems have difficulty assessing the driver's emotional state in real time and providing personalized solutions. Furthermore, their lack of multilingual support makes it difficult to provide solutions that reflect the driver's language preferences.

[1251] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for receiving customer utterances as voice input and converting them into text data using a voice recognition system; sentiment analysis means for analyzing the text data and evaluating the customer's sentiment; means for identifying customer needs and generating personalized solutions; multilingual support means for translating the solutions into multiple languages; means for providing the translated solutions to the customer; means for saving dialogue data with the customer and retraining the AI ​​model; means for processing the driver's voice input in real time and displaying solutions on the vehicle's dashboard or head-mounted display; and means for generating and providing personalized solutions appropriate to the driver's emotional state based on the sentiment evaluation using the AI ​​model. This enables the driver to quickly obtain appropriate solutions appropriate to the driver's emotional state for problems that arise in real time inside an autonomous vehicle.

[1252] "Voice input" refers to inputting voice data spoken by a user into a terminal or system.

[1253] A "voice recognition system" is a technology or device for converting voice data into text data.

[1254] "Text data" is character string information of voice input converted by a voice recognition system.

[1255] "Emotion analysis means" refers to a technique or device that analyzes text data and identifies the user's emotional state.

[1256] The "means for identifying needs" is a technology or device that clarifies the user's requests and problems based on the content of the user's inquiry and past inquiry history.

[1257] A "personalized solution" is a specific solution or instruction provided to a user that is tailored to their specific situation and needs.

[1258] A "multilingual means" is a technique or device that translates generated solutions into multiple languages ​​and provides them to the user in the appropriate language.

[1259] "Dialogue data" refers to recorded data relating to all statements and inquiries exchanged between the user and the system.

[1260] An "artificial intelligence model" is a model built using machine learning algorithms to learn from large amounts of data and perform specific tasks.

[1261] A "dashboard" is a display device installed inside a vehicle that provides various information to the driver.

[1262] A "head-mounted display" is a device worn on the head that displays information within the field of vision.

[1263] "Driver emotional assessment" refers to the act of analyzing and identifying the driver's emotional state while driving.

[1264] "Real-time processing means" refers to technology or devices that instantly analyze and respond to user input on the spot.

[1265] A system for implementing this invention is configured as follows: First, the driver provides voice input to the system through a terminal (dashboard or head-mounted display). The terminal records the driver's voice and transmits the voice data to a server.

[1266] The server converts the voice data into text using a speech recognition system, which uses the "speech_recognition" library. For example, if a driver says, "I can't set up the navigation system in my car. What should I do?", the speech is converted into text.

[1267] The server then analyzes the converted text data using a natural language processing (NLP) algorithm and a sentiment analysis engine called "transformers" library to extract themes and keywords from the speech and identify the driver's emotional state. For example, if the driver's speech expresses frustration, the sentiment engine classifies the emotional state as "negative."

[1268] The server also retrieves the driver's past inquiry history from the database and analyzes it in combination with the current text data. This identifies the driver's needs and generates personalized solutions. For example, if a problem arises with the navigation system settings, the solution might be to "select an option from the settings menu and change the navigation settings."

[1269] The generated solutions are translated into the driver's preferred language using a multilingual system if necessary. For example, when translating an English solution into Japanese, the server converts it into natural language while preserving the exact meaning. The translated solution is then displayed on the device.

[1270] If the driver tries the proposed solutions and the problem is not resolved, they can again enter additional questions into the device via voice input. The device then sends the additional questions to the server, which analyzes them again and provides further solutions. All of this dialogue data is stored in a database and used to retrain the artificial intelligence model.

[1271] The system's unique features include its ability to assess the driver's emotional state in real time and provide personalized solutions quickly and efficiently. It also supports multiple languages, making it flexible enough for international use.

[1272] Specific examples and examples of prompts for generative AI models

[1273] A concrete example is a scenario in which a driver makes an inquiry about the settings of a navigation system.

[1274] Examples:

[1275] The driver speaks, "I can't set up the navigation system in this car. What should I do?"

[1276] Example prompt for a generative AI model:

[1277] If a user says, "I can't configure the navigation system in my car, what should I do?", convert the speech to text and analyze the sentiment. Generate a solution: "Please reset your navigation system. Select the option from the settings menu and change your navigation settings."

[1278] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1279] Step 1:

[1280] The user speaks to the device in the autonomous vehicle, for example, saying, "I can't set up the navigation system in this car. What should I do?" This speech input is recorded by the device's microphone.

[1281] Step 2:

[1282] The device sends the recorded voice data to the server. The server receives the voice data as input and converts it into text data using a voice recognition system. The "speech_recognition" library is used for voice recognition. As a result of the conversion, the voice data is output as text data: "I can't set up the navigation system in this car. What should I do?"

[1283] Step 3:

[1284] The server analyzes the converted text data using natural language processing algorithms (NLP) and a sentiment analysis engine. Specifically, sentiment analysis is performed using the "transformers" library. From the input text data, themes and keywords are extracted, along with an assessment of the user's emotional state (e.g., frustration). The output of this step is the extracted keywords and the emotional state.

[1285] Step 4:

[1286] The server retrieves past inquiry history from the database and compares it with the analyzed text data. This identifies the driver's needs for the current problem and generates an optimal personalized solution. In this case, specific instructions for configuring the navigation system are generated. The input is the analyzed text data and emotional state, and the output is the generated solution.

[1287] Step 5:

[1288] The server translates the generated solution into the driver's preferred language using a multilingual system. For example, when translating an English solution into Japanese, it converts it into a natural expression while preserving the exact meaning. The input of this step is the generated solution (English), and the output is the translated solution (Japanese).

[1289] Step 6:

[1290] The server sends the translated solution to the device, which then displays it on the car's dashboard or head-mounted display, for example, "Select an option from the settings menu and change your navigation settings."

[1291] Step 7:

[1292] The user tries the proposed solution. If the problem is not resolved, they use voice input again to enter an additional question into the device. The device then sends this additional question back to the server. The server converts it back into text data, analyzes it, and provides additional solutions. This dialogue data is stored in a database and used to retrain the artificial intelligence model. The input is the additional question, and the output is the additional solution.

[1293] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1294] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1295] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1296] [Fourth embodiment]

[1297] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1298] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1299] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1300] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1301] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1302] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1303] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1304] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1305] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1306] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1307] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1308] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1309] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1310] The present invention provides a multilingual and multimodal customer service system aimed at improving customer service, and an embodiment thereof is shown below.

[1311] Program processing explanation

[1312] 1. Input and recognition of customer utterances

[1313] The user makes a voice inquiry about a problem with the product. The device records this voice and sends the data to the server. The server then uses a voice recognition system to convert the recorded voice data into text data.

[1314] 2. Sentiment and Text Analysis

[1315] The server analyzes the converted text data and uses NLP (Natural Language Processing) technology to understand the content of customer comments. It also uses a sentiment analysis system to evaluate the emotional state in the text data. For example, if a user says, "This product is completely unusable!", the sentiment analysis system will classify the comment as "anger," a negative emotion.

[1316] 3. Identifying needs and generating personalized solutions

[1317] The server retrieves past customer inquiries from a database and combines them with the current text data. It then identifies the customer's needs based on what they said and their sentiment assessment. Based on this, the AI ​​model generates a personalized solution. For example, if the customer is contacted about a software crash, it could provide instructions on how to change settings or the latest patch.

[1318] 4. Multilingual support

[1319] The server translates the generated solutions into multiple languages ​​as needed, for example, when translating a solution written in English into Japanese, it converts it into natural language while preserving the exact meaning.

[1320] 5. Providing solutions

[1321] The device displays the solution received from the server to the user, and also interactively asks the user additional questions or offers suggestions to increase engagement, such as displaying a message with instructions on how to change settings and offering a link to download the latest patch.

[1322] 6. Data Feedback and Learning

[1323] The user follows the system's instructions and attempts to solve the problem. The device sends the results and any follow-up questions to the server, which stores this interaction data and uses it to retrain the AI ​​model, improving the accuracy and personalization of responses to future inquiries.

[1324] Specific examples

[1325] Scenario: A user contacts support in Japanese.

[1326] 1. User: Asks by voice, "This software keeps crashing, what should I do?"

[1327] 2. Device: Records audio and sends it to the server.

[1328] 3. Server: Uses a speech recognition system to convert the speech into text data such as "This software keeps crashing, what should I do?"

[1329] 4. Server: Analyzes the content of the text and identifies the emotion "anger" using an emotion analysis system.

[1330] 5. Server: Retrieves past query history from the database, identifies needs, and generates solutions that provide configuration change instructions and the latest patches.

[1331] 6. Server: Translate the solution into Japanese in a multilingual system and generate the message "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link."

[1332] 7. Terminal: Display the translated solution to the user.

[1333] 8. User: Try the steps provided and enter additional questions if the problem persists.

[1334] 9. Terminal: Sends additional queries to the server.

[1335] 10. Server: Provides additional countermeasures and stores interaction data as training data.

[1336] In this way, the system of the present invention aims to improve customer satisfaction by analyzing customer comments in multiple languages ​​and multiple modalities and providing personalized solutions quickly and efficiently.

[1337] The processing flow will be explained below.

[1338] Step 1:

[1339] A user voice-inquires about a problem with a product, for example, "This software keeps crashing, what should I do?"

[1340] Step 2:

[1341] The terminal records the user's voice and transmits the voice data to the server.

[1342] Step 3:

[1343] The server inputs the received voice data into a voice recognition system and converts the voice data into text data, for example, "This software keeps crashing, what should I do?"

[1344] Step 4:

[1345] The server applies natural language processing algorithms to analyze the text data, extracting themes and keywords from the comments.

[1346] Step 5:

[1347] The server uses an emotion analysis system to identify the user's emotion from the text data, which in this case is classified as "anger."

[1348] Step 6:

[1349] The server retrieves the customer's past inquiry history from the database and searches for solutions to similar problems.

[1350] Step 7:

[1351] The server uses the results of text and sentiment analysis to identify customer needs, such as a need for a solution to a software crash.

[1352] Step 8:

[1353] The server uses AI models to generate personalized solutions for identified needs, such as instructions for changing settings or providing the latest patches.

[1354] Step 9:

[1355] The server utilizes a multilingual system to translate the generated solutions into multiple languages ​​as needed, for example, translating the solutions from Japanese to English.

[1356] Step 10:

[1357] The device displays the solution received from the server to the user, for example, "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link."

[1358] Step 11:

[1359] The user attempts the provided solution and types additional questions or feedback about the solution into the device.

[1360] Step 12:

[1361] The device sends the user's feedback and any follow-up questions back to the server.

[1362] Step 13:

[1363] The server analyzes the additional information and takes further action if necessary, and also stores the interaction data and stores it in a database for use in retraining the AI ​​model.

[1364] This series of steps enables the system of the present invention to respond quickly and efficiently to customer needs and provide a high level of personalized customer service.

[1365] Example 1

[1366] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1367] Modern companies strive to improve customer service, but implementing a multilingual, multimodal customer service system is a technically difficult challenge. Customers make inquiries in a variety of languages, and their emotions and needs vary widely, making responding to them time-consuming and labor-intensive. Conventional systems have difficulty accurately grasping customer emotions and needs, and responses tend to be uniform, failing to sufficiently increase customer satisfaction.

[1368] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1369] In this invention, the server includes means for receiving customer utterances as voice input and converting them into text data using a voice recognition system, means for analyzing the text data and evaluating customer sentiment, and means for identifying customer needs and generating personalized solutions. This makes it possible to analyze customer utterances in multiple languages ​​and multiple modalities and provide personalized solutions quickly and efficiently, thereby improving customer satisfaction.

[1370] "Voice input" is a method by which a user can query a system through speech.

[1371] A "voice recognition system" refers to a technology that analyzes recorded voice data and converts it into text data.

[1372] "Text data" refers to character information converted from audio data.

[1373] "Sentiment analysis means" refers to technology that evaluates and classifies emotional states in text data.

[1374] "Personalized solutions" are a method for providing the best possible solution to a specific problem based on the customer's needs.

[1375] "Multilingual means" refers to techniques for translating generated solutions into multiple languages.

[1376] "Interaction Data" refers to records of conversations between a customer and a system.

[1377] "Retraining methods" refers to techniques that use stored dialogue data to retrain AI models and improve their accuracy.

[1378] "Terminal" refers to the device used by the Customer to provide voice input.

[1379] "Database" refers to a repository of information where past inquiry history and other related information is stored.

[1380] "Natural language processing technology" refers to a series of technologies for analyzing text data and understanding its content.

[1381] "Engagement" refers to methods for eliciting involvement and interest through interaction with users.

[1382] An "AI model" refers to an algorithm that has been trained through machine learning to perform a specific task.

[1383] The present invention provides a multilingual and multimodal customer service system aimed at improving customer service, and an embodiment thereof is shown below.

[1384] System configuration

[1385] The system begins when a user voice-inquires about a problem with a product. When the user voice-inquires, the device records the voice and sends it to the server. The server then uses a voice recognition system to convert the voice data into text. The server then analyzes the text data, evaluates the customer's sentiment, and identifies their needs. It then uses an AI model to generate a personalized solution and translates it into multiple languages. Finally, the translated solution is provided to the user via the device.

[1386] Hardware and software used

[1387] Devices: Devices such as smartphones, computers, and tablets are used.

[1388] Server: A remote server is used for data processing and storage.

[1389] Speech recognition system: Google Cloud Speech-to-Text API, IBM Watson Speech to Text, etc. are used.

[1390] NLP technologies: Natural language processing libraries such as TensorFlow, spaCy, and NLTK are used.

[1391] Sentiment analysis system: TextBlob, NRC Emotion Lexicon, etc. are used.

[1392] Translation system: Google Cloud Translation API, DeepL API, etc. are used.

[1393] AI models: Generative AI models such as GPT and BERT are used.

[1394] Specific examples

[1395] The following are specific usage scenarios for this system:

[1396] 1. User: Makes a voice inquiry saying, "This software keeps crashing, what should I do?"

[1397] 2. Device: Records the user's voice and sends it to the server.

[1398] 3. Server: Using a speech recognition system, convert the speech into text data such as "This software keeps crashing, what should I do?"

[1399] 4. Server: Analyzes text data using NLP technology and identifies the emotion "anger."

[1400] 5. Server: Retrieves past inquiry history from a database, identifies needs, and uses AI models to generate solutions that provide configuration change instructions and the latest patches.

[1401] 6. Server: Translate the generated solution using a multilingual system. For example, the message "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link." is translated into Japanese.

[1402] 7. Terminal: Display the translated solution to the user.

[1403] Prompt Sentence Examples

[1404] Here are some examples of prompts for generative AI models:

[1405] Prompt sentences that convert user speech into text

[1406] Please convert the following audio data to text: [Audio data]

[1407] Prompt sentences that analyze the sentiment of the user's text utterances

[1408] Please rate the emotional state of the following text data: [Text data]

[1409] Prompts that identify user needs and generate solutions

[1410] Identify the user's needs and generate a solution based on the following text: [Text]

[1411] Prompts for translating generated text into multiple languages

[1412] Please translate the following text data into [language name]: [text data]

[1413] By using these procedures and prompts, the system of the present invention aims to analyze customer utterances in multiple languages ​​and multiple modalities and provide personalized solutions quickly and efficiently.

[1414] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1415] Step 1:

[1416] A user speaks to a product about a problem they are having. For example, a user might say, "This software keeps crashing. What should I do?" This is the input.

[1417] Step 2:

[1418] The device records the user's voice and sends the data to the server. The recording is performed using the microphone on a smartphone or computer. The recorded voice data is output.

[1419] Step 3:

[1420] The server uses a speech recognition system to convert the speech data into text data. Using the speech data as input, a speech recognition system such as the Google Cloud Speech-to-Text API converts the speech into text data such as "This software keeps crashing, what should I do?". The output is the text data converted from the speech.

[1421] Step 4:

[1422] The server analyzes the text data and uses NLP technology to understand the content of customer comments. It uses natural language processing tools such as TensorFlow and spaCy to analyze the structure and meaning of sentences, taking the text data as input. The output of this process is analyzed text data.

[1423] Step 5:

[1424] The server uses a sentiment analysis system to evaluate the emotional state in the text data. Using the analyzed text data as input, it uses tools such as TextBlob and the NRC Emotion Lexicon to identify the customer's emotion as "anger." The output is the analyzed sentiment data.

[1425] Step 6:

[1426] The server retrieves past query history from the database and combines it with the current text data. Using the past query history stored in the database as input, it extracts relevant data using SQL queries. The output is a combination of the past query history and the current text data.

[1427] Step 7:

[1428] The server identifies customer needs and generates personalized solutions using AI models. Using the combined text data as input, machine learning models (GPT and BERT) are used to identify needs and generate solutions. The output is a personalized solution.

[1429] Step 8:

[1430] The server translates the generated solution into multiple languages ​​as needed. The generated solution is used as input and translated using the Google Cloud Translation API or the DeepL API. The output is the translated solution in multiple languages.

[1431] Step 9:

[1432] The device displays the translated solution received from the server to the user. The translated solution is used as input and a solution message is displayed on the smartphone or PC screen. The output is the solution provided to the user.

[1433] Step 10:

[1434] The user follows the system's instructions to try to solve the problem. The user follows the displayed solution steps, changing software settings or downloading patches from links. This is the input, and the results of the setting changes and patch downloads are the output.

[1435] Step 11:

[1436] The terminal sends the results and any additional questions to the server. The terminal sends data to the server using the results of the solution execution and any new questions from the user as input. The output is the sent result data and any additional questions.

[1437] Step 12:

[1438] The server stores the dialogue data and retrains the AI ​​model. It uses newly collected dialogue data as input, stores the data, and retrains the AI ​​model. The output is a trained AI model, which improves the accuracy of responses to future inquiries.

[1439] (Application example 1)

[1440] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1441] Current customer service systems only support one language and lack multilingual support and sentiment analysis, preventing them from fully improving customer satisfaction. They also lack the means to quickly provide personalized solutions, making it difficult to respond immediately to follow-up customer inquiries. There is a need for a system that can provide multilingual and multimodal support in brick-and-mortar stores using smartphones and service robots.

[1442] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1443] In this invention, the server includes a means for converting voice input into text data using a voice recognition system, a sentiment analysis means for analyzing the text data to evaluate customer sentiment, a means for identifying customer needs and generating personalized solutions, a multilingual support means for translating the solutions into multiple languages, a means for providing the translated solutions to customers, a means for saving customer interaction data and retraining an AI model, a customer service system that functions on hardware including smartphones and service robots, and a means for providing customer service through voice, text, and translated solutions and for engaging in dialogue based on sentiment evaluation and needs, thereby enabling the provision of advanced customer service that is multilingual and emotionally responsive.

[1444] A "customer support system" is a system that responds quickly and appropriately to customer inquiries in physical stores and online.

[1445] "Voice input" is a means of acquiring customer utterances as voice data.

[1446] A "voice recognition system" is a technology or device for converting voice data into text data.

[1447] "Text data" is data expressed in the form of character information.

[1448] "Sentiment analysis means" is a means for evaluating and classifying customer emotions from text data.

[1449] "Needs identification means" are means for clarifying specific requirements or requests based on the content of a customer inquiry.

[1450] A "personalized solution" is a customized response to address a customer's specific needs.

[1451] "Multilingual support means" is a means for translating generated solutions into multiple languages.

[1452] "Interaction Data" is a record of the interaction between a customer and the system.

[1453] An "AI model" is a mathematical model for learning and prediction using artificial intelligence technology.

[1454] "Retraining methods" are methods used to improve the accuracy of existing AI models using new data.

[1455] A "smartphone" is a portable information terminal that can input voice and run applications.

[1456] A "service robot" is an autonomous robot designed to provide customer service and information.

[1457] "Multimodal" is the property of a system that includes multiple input and output formats, such as speech and text.

[1458] The present invention is a multilingual, multimodal customer service system for enhancing customer service in brick-and-mortar stores. This system combines the processes of speech recognition, emotion analysis, text analysis, needs identification, multilingual support, solution provision, data feedback, and learning. Specific embodiments for implementing the present invention are described below.

[1459] The main components of the system are the server, the terminal, and the user. The terminal consists of a smartphone or a service robot, and is responsible for receiving voice input from the user and sending it to the server.

[1460] Hardware used

[1461] Smartphone: A mobile information device that can accept voice input and run applications.

[1462] Service robot: An autonomous robot designed to provide customer service and information.

[1463] Software used

[1464] SpeechRecognition: A Python library for performing speech recognition.

[1465] transformers: A library that provides sentiment analysis, text analysis, and translation models.

[1466] Data Processing Steps

[1467] 1. Receiving and converting voice input

[1468] A user makes a voice inquiry to a smartphone or a service robot. For example, the user might say, "Please tell me how to use this product." The device records this voice and sends it to a server.

[1469] 2. Speech Recognition and Emotion Analysis

[1470] The server uses a speech recognition system (SpeechRecognition) to convert the recorded voice data into text data. The converted text data is then analyzed using sentiment analysis (transformers) to evaluate the customer's emotions. Through this process, the text is converted to "Please tell me how to use this product." and the emotion is classified as "confused."

[1471] 3. Needs Identification and Solution Generation

[1472] The server uses text analysis tools (transformers' text-classification) to identify customer needs based on the analyzed text data, retrieves personalized solutions from the database and applies them to the customer.

[1473] 4. Multilingual support

[1474] The generated solutions are translated into multiple languages ​​as needed. A multilingual solution (transformers' translation-en-to-ja) is used to properly translate from English to Japanese.

[1475] 5. Providing solutions and dialogue

[1476] The translated solution is sent to the device and displayed to the user. If the user asks additional questions, that data is also sent to the server for processing again.

[1477] 6. Data Feedback and Learning

[1478] The server stores the results of the conversation and any follow-up questions, and retrains the AI ​​model, improving its accuracy for the next inquiry.

[1479] Examples and prompts

[1480] Specific examples

[1481] 1. User: Ask, "How do I use this product?"

[1482] 2. Application: Convert speech to text and analyze sentiment and content.

[1483] 3. Application: Multilingual support and display of usage solutions.

[1484] 4. User: Enter a follow-up question.

[1485] 5. Application: Save the proposed responses and use them for future learning.

[1486] Prompt Sentence Examples

[1487] Customer: "Please tell me how to use this product."

[1488] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1489] Step 1:

[1490] Receiving and converting voice input

[1491] A user makes a voice inquiry to a smartphone or service robot. For example, the user might say, "Please tell me how to use this product." This inputs voice data into the terminal. The terminal records this voice and sends it to a server as digital voice data.

[1492] Input: User speech (e.g., "How do I use this product?")

[1493] Output: Digital audio data

[1494] Step 2:

[1495] Speech Recognition and Emotion Analysis

[1496] The server converts the received digital voice data into text data using a speech recognition system (SpeechRecognition library). For example, the voice is converted into text such as "Please tell me how to use this product." The text data is then analyzed using a sentiment analysis method (transformers' sentiment-analysis model) to evaluate the customer's emotions. As a result of the sentiment analysis, the text is classified as "confused."

[1497] Input: Digital audio data

[1498] Output: Text data (e.g., "How do I use this product?") and emotion ratings (e.g., "Confusion")

[1499] Step 3:

[1500] Needs Identification and Solution Generation

[1501] The server uses text analysis methods (Transformers' text-classification model) to identify customer needs based on the text data. For example, the tag "How to use the product" is identified. Then, it retrieves the corresponding solutions from the database and generates personalized solutions. For example, instructions on how to use the product are retrieved from the database.

[1502] Input: Text data and sentiment ratings

[1503] Output: Identified needs (e.g., how to use the product) and personalized solutions (e.g., product usage instructions)

[1504] Step 4:

[1505] Multilingual support

[1506] The server translates the generated solution as needed using multilingual support (transformers' translation-en-to-ja model). For example, a solution generated in English is translated into Japanese. This results in a solution in the target language.

[1507] Input: personalized solution (e.g. product usage instructions)

[1508] Output: Translated solution (e.g., product usage instructions in Japanese)

[1509] Step 5:

[1510] Providing solutions and dialogue

[1511] The translated solution is sent to the device, which displays it to the user and continues the dialogue via voice or text as needed. For example, if the user reviews the presented instructions and asks additional questions, their input is sent back to the server.

[1512] Input: translated solution

[1513] Output: The solution and any follow-up questions displayed to the user

[1514] Step 6:

[1515] Data Feedback and Learning

[1516] The results of the user's attempts and any follow-up questions are sent to the server, which stores this interaction data and uses it to retrain the AI ​​model, improving its accuracy for the next inquiry.

[1517] Input: User feedback and follow-up questions

[1518] Output: Dialogue data for retraining and improved AI models

[1519] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1520] The present invention provides a multilingual, multimodal customer support system aimed at improving customer service, and includes a form that combines an emotion engine. An embodiment of this system is shown below.

[1521] Program processing explanation

[1522] 1. Input and recognition of customer utterances

[1523] A user voice-inquires about a problem with a product. For example, "This software keeps crashing. What should I do?" The device records this voice and sends the voice data to the server. The server uses a voice recognition system to convert the recorded voice data into text data.

[1524] 2. Sentiment and Text Analysis

[1525] The server applies natural language processing algorithms to analyze the converted text data. This analysis extracts themes and keywords from the utterances. It then uses an emotion engine to identify the user's emotional state from the text data. For example, if a user utters, "This product is completely unusable!", the emotion engine classifies the utterance as "anger," a negative emotion. The emotion engine evaluates emotions in real time and adjusts the response as needed.

[1526] 3. Identifying needs and generating personalized solutions

[1527] The server retrieves the customer's past inquiry history from the database and combines it with the current text data for analysis. This allows it to identify the customer's needs and generate a personalized solution. For example, if the inquiry is about a software crash, it generates a message providing instructions on changing settings and the latest patches.

[1528] 4. Multilingual support

[1529] The server translates the generated solutions into multiple languages ​​as needed, for example, when translating a solution written in English into Japanese, it converts it into natural language while preserving the exact meaning.

[1530] 5. Providing solutions

[1531] The device displays the solution received from the server to the user, for example, "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link."

[1532] 6. Data Feedback and Learning

[1533] The user follows the system's instructions and attempts the proposed solution. They then enter additional questions or feedback about the solution into the device. The device then sends the user's feedback and additional questions back to the server. The server analyzes the additional information and takes further action if necessary. The interaction data is also saved and stored in a database for use in retraining the AI ​​model.

[1534] Specific examples

[1535] Scenario: A user contacts support in Japanese.

[1536] 1. User: Asks by voice, "This software keeps crashing, what should I do?"

[1537] 2. Device: Records audio and sends it to the server.

[1538] 3. Server: Uses a speech recognition system to convert the speech into text data such as "This software keeps crashing, what should I do?"

[1539] 4. Server: Analyzes the content of the text and identifies the emotion "anger" using an emotion engine.

[1540] 5. Server: Retrieves past query history from the database, identifies needs, and generates solutions that provide configuration change instructions and the latest patches.

[1541] 6. Server: Translate the solution into Japanese in a multilingual system and generate the message "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link."

[1542] 7. Terminal: Display the translated solution to the user.

[1543] 8. User: Try the steps provided and enter additional questions if the problem persists.

[1544] 9. Terminal: Sends additional queries to the server.

[1545] 10. Server: Provides additional countermeasures and stores interaction data as training data.

[1546] In this way, the system of the present invention aims to improve customer satisfaction by analyzing customer comments in multiple languages ​​and multiple modalities, using an emotion engine to evaluate the user's emotional state in real time, and providing personalized solutions quickly and efficiently.

[1547] The processing flow will be explained below.

[1548] Step 1:

[1549] A user voice-inquires about a problem with a product, for example, "This software keeps crashing, what should I do?"

[1550] Step 2:

[1551] The terminal records the user's voice and transmits the voice data to the server.

[1552] Step 3:

[1553] The server inputs the received voice data into a voice recognition system and converts the voice data into text data, for example, "This software keeps crashing, what should I do?"

[1554] Step 4:

[1555] The server then analyzes the converted text data using a natural language processing algorithm, extracting the topic and keywords of the comments and identifying the nature of the problem.

[1556] Step 5:

[1557] The server uses an emotion engine to identify the user's emotional state from the text data in real time, for example, in this case the emotion "anger."

[1558] Step 6:

[1559] The server retrieves the customer's past inquiry history from the database and collects relevant data for resolving the problem.

[1560] Step 7:

[1561] The server identifies specific customer needs based on the results of text analysis and sentiment analysis. For example, based on the frequent occurrence of software crashes, it determines that the customer needs a procedure for changing settings to prevent crashes.

[1562] Step 8:

[1563] The server uses AI models to generate personalized solutions based on identified needs, such as "Settings - Options - Crash Prevention" instructions or a "download link for the latest patch."

[1564] Step 9:

[1565] The server translates the generated solution into the required language using a multilingual system, for example translating the solution from Japanese to English.

[1566] Step 10:

[1567] The device displays the solution received from the server to the user, for example, "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link."

[1568] Step 11:

[1569] The user attempts the provided solution and enters additional questions or feedback about the solution into the terminal.

[1570] Step 12:

[1571] The device sends the user's feedback and follow-up questions to the server.

[1572] Step 13:

[1573] The server analyzes the additional information and takes further action if necessary, such as providing more detailed configuration instructions or an alternative solution, and also saves the interaction data and stores it in a database for use in retraining the AI ​​model.

[1574] Through this series of steps, the system of the present invention quickly and efficiently analyzes customer comments in multiple languages ​​and multiple modalities, and evaluates user sentiment in real time using an emotion engine, thereby significantly improving customer satisfaction by providing personalized solutions.

[1575] Example 2

[1576] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1577] Modern customer service requires multilingual and multimodal customer interactions, and it is becoming increasingly important to provide personalized solutions quickly based on customer emotions and needs. However, conventional systems struggle to accurately convert voice input into text, perform sentiment analysis, and provide solutions adapted to multiple languages, making it difficult to increase customer satisfaction. Furthermore, while continuous learning and improvement based on dialogue data is desirable, there is a lack of systems that can achieve this.

[1578] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1579] In this invention, the server includes means for receiving customer utterances as voice input and converting them into text data using speech recognition technology, sentiment analysis means for analyzing the text data and evaluating the customer's sentiment, means for identifying customer needs and generating personalized solutions, multilingual support means for translating the generated solutions into multiple languages, and means for saving customer interaction data and retraining a machine learning model. This makes it possible to analyze customer utterances in multiple languages ​​and evaluate the user's emotional state in real time using an emotion engine, thereby quickly and efficiently providing personalized solutions.

[1580] Below are definitions of important terms included in the claims.

[1581] "Voice recognition technology" is a technology that analyzes voice data and converts it into text data.

[1582] "Text data" is character information converted using voice recognition technology.

[1583] An "emotion analysis means" is an algorithm or system that analyzes text data and identifies the user's emotional state.

[1584] "Means for identifying needs" refers to a technique for extracting customer requests and requirements based on the customer's past inquiry history and current statements.

[1585] "Personalized solutions" are customized solutions or guidance provided based on a customer's specific situation and emotions.

[1586] "Multilingual solutions" are technologies or systems that translate solutions into multiple languages.

[1587] "Dialogue data" is historical information about questions and answers exchanged with customers.

[1588] A "machine learning model" is an algorithm that learns from past data and makes predictions and classifications.

[1589] "Retraining" is a technique for adding new data to improve the accuracy and performance of a machine learning model.

[1590] A "terminal" is an electronic device used by a customer that inputs voice and displays text.

[1591] A "database" is a system that stores and manages past inquiry history and customer information.

[1592] The present invention provides a multilingual, multimodal customer support system aimed at improving customer service, including a form that combines an emotion engine. The following describes the embodiments of the system in detail.

[1593] System configuration

[1594] The multilingual and multimodal customer support system of the present invention is composed of the following elements: a server, a terminal, and a user.

[1595] Server: A high-performance computer running speech recognition technology, natural language processing algorithms, emotion engines, machine learning models, and translation systems. Examples include software such as Google Cloud Speech-to-Text, Google Cloud Natural Language API, and DeepL API.

[1596] Terminal: An electronic device capable of voice input and text display, such as a smartphone, tablet, or PC.

[1597] User: A customer who contacts us regarding a problem with a product.

[1598] Program processing explanation

[1599] Customer utterance input and recognition

[1600] A user voice-inquires about a problem with a product. For example, "This software keeps crashing. What should I do?" The device records this voice and sends the voice data to a server. The server then uses voice recognition technology, such as Google Cloud Speech-to-Text, to convert the recorded voice data into text.

[1601] Sentiment and Text Analysis

[1602] The server then applies natural language processing algorithms, such as the Google Cloud Natural Language API, to analyze the converted text data. This analysis extracts the topic and keywords of the utterance. It then uses an emotion engine to identify the user's emotional state from the text data. For example, if a user says, "This product is completely unusable!", the emotion engine classifies the utterance as "anger," a negative emotion. The emotion engine evaluates the emotion in real time and adjusts the response as needed.

[1603] Identifying needs and generating personalized solutions

[1604] The server retrieves the customer's past inquiry history from the database and combines it with the current text data for analysis. This allows it to identify the customer's needs and generate a personalized solution. For example, if the inquiry is about a software crash, it generates a message providing instructions on changing settings and the latest patches.

[1605] Multilingual support

[1606] The server translates the generated solutions into multiple languages ​​as needed. For example, if a solution written in English needs to be translated into Japanese, it uses a translation system such as the DeepL API to convert it into natural language while preserving the exact meaning.

[1607] Providing solutions

[1608] The device displays the solution received from the server to the user, for example, "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link."

[1609] Data Feedback and Learning

[1610] The user follows the system's instructions and attempts the proposed solution. They then enter additional questions or feedback about the solution into the device. The device then sends the user's feedback and additional questions back to the server. The server analyzes the additional information and takes further action if necessary. The interaction data is also saved and stored in a database for use in retraining the AI ​​model.

[1611] Specific examples

[1612] Scenario: A user contacts support in Japanese.

[1613] 1. User: Asks by voice, "This software keeps crashing, what should I do?"

[1614] 2. Device: Records audio and sends it to the server.

[1615] 3. Server: Use Google Cloud Speech-to-Text to convert the speech into text data such as "This software keeps crashing, what should I do?"

[1616] 4. Server: The text content is analyzed using the Google Cloud Natural Language API, and the emotion engine identifies the emotion "anger."

[1617] 5. Server: Retrieves past query history from the database, identifies needs, and generates solutions that provide configuration change instructions and the latest patches.

[1618] 6. Server: Translate the solution into Japanese using the DeepL API and generate the message "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link."

[1619] 7. Terminal: Display the translated solution to the user.

[1620] 8. User: Try the steps provided and enter any follow-up questions if the problem persists.

[1621] 9. Terminal: Sends additional queries to the server.

[1622] 10. Server: Provides additional countermeasures and stores the dialogue data as training data.

[1623] In this way, the system of the present invention aims to improve customer satisfaction by analyzing customer comments in multiple languages ​​and multiple modalities, using an emotion engine to evaluate the user's emotional state in real time, and providing personalized solutions quickly and efficiently.

[1624] Example prompt sentence:

[1625] "A user speaks to you about a software crash. Design a system that uses speech recognition to convert the text and analyzes the user's emotional state to generate a personalized solution. Describe the process for translating the generated solution into various languages ​​and displaying it to the user."

[1626] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1627] Step 1:

[1628] The user makes a voice inquiry about a problem with the product, for example, "This software keeps crashing. What should I do?" The device records this voice and sends the voice data to the server.

[1629] Input: User voice input

[1630] Output: Audio data sent to the server

[1631] Specific operation: The user uses the device's microphone to voice-input the details of the problem, and the device sends the voice data to the server.

[1632] Step 2:

[1633] The server uses voice recognition technology such as Google Cloud Speech-to-Text to convert the recorded audio data into text data.

[1634] Input: Audio data

[1635] Output: Text data

[1636] Specific operation: The server calls the speech recognition API and converts the speech data into text data.

[1637] Step 3:

[1638] The server then applies natural language processing algorithms, such as Google Cloud Natural Language API, to analyze the converted text data, extracting themes and keywords from the speech.

[1639] Input: Text data

[1640] Output: Analysis data including themes and keywords

[1641] Specific operation: The server analyzes the text via a natural language processing API and extracts themes and keywords.

[1642] Step 4:

[1643] The server uses an emotion engine to identify the user's emotional state from the text data. For example, if a user says, "This product is completely unusable!", the emotion engine classifies the statement as the negative emotion "anger."

[1644] Input: Parsed text data

[1645] Output: Analysis data including emotional state

[1646] Specific operation: The server uses the emotion engine to determine the emotion of the text data and identify the emotional state.

[1647] Step 5:

[1648] The server retrieves the customer's past inquiry history from the database and combines it with the current text data for analysis. This identifies the customer's needs and generates a personalized solution. For example, if the inquiry is about a software crash, it generates a message providing instructions on changing settings and the latest patches.

[1649] Input: Customer text data, past inquiry history

[1650] Output: personalized solution

[1651] Specific operation: The server queries past inquiry data from the database, combines it with current text data, analyzes it, and generates a solution.

[1652] Step 6:

[1653] The server translates the generated solutions into multiple languages ​​as needed. For example, if a solution written in English is to be translated into Japanese, a translation system such as the DeepL API is used to convert it into natural language while preserving the exact meaning.

[1654] Input: personalized solutions

[1655] Output: Translated solution

[1656] Specific operation: The server calls the translation API to translate the solution text into the specified language.

[1657] Step 7:

[1658] The device displays the solution received from the server to the user, for example, "Please try the following steps: Settings - Options - Crash Prevention. You can also download the latest patch from this link."

[1659] Input: translated solution

[1660] Output: The solution that is displayed to the user

[1661] Specific operation: The terminal receives the solution sent from the server and displays it on the display.

[1662] Step 8:

[1663] The user follows the system's instructions to try the proposed solution, then inputs additional questions or feedback about the solution into the terminal, which then sends the user's feedback or additional questions back to the server.

[1664] Input: User feedback and follow-up questions

[1665] Output: Additional questions and feedback sent to the server

[1666] Specific operation: The user enters additional questions or feedback into the device, which then sends it to the server.

[1667] Step 9:

[1668] The server analyzes the additional information and takes further action if necessary, and also stores the interaction data and stores it in a database for use in retraining the AI ​​model.

[1669] Input: Additional questions and feedback

[1670] Output: Additional countermeasures, saved interaction data

[1671] Specific actions: The server analyzes any additional questions or feedback, provides further solutions, and stores the interaction data in a database.

[1672] (Application example 2)

[1673] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1674] Autonomous vehicles require support systems that can quickly and effectively respond to various problems faced by drivers in real time. However, conventional systems have difficulty assessing the driver's emotional state in real time and providing personalized solutions. Furthermore, their lack of multilingual support makes it difficult to provide solutions that reflect the driver's language preferences.

[1675] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for receiving customer utterances as voice input and converting them into text data using a voice recognition system; sentiment analysis means for analyzing the text data and evaluating the customer's sentiment; means for identifying customer needs and generating personalized solutions; multilingual support means for translating the solutions into multiple languages; means for providing the translated solutions to the customer; means for saving dialogue data with the customer and retraining the AI ​​model; means for processing the driver's voice input in real time and displaying solutions on the vehicle's dashboard or head-mounted display; and means for generating and providing personalized solutions appropriate to the driver's emotional state based on the sentiment evaluation using the AI ​​model. This enables the driver to quickly obtain appropriate solutions appropriate to the driver's emotional state for problems that arise in real time inside an autonomous vehicle.

[1676] "Voice input" refers to inputting voice data spoken by a user into a terminal or system.

[1677] A "voice recognition system" is a technology or device for converting voice data into text data.

[1678] "Text data" is character string information of voice input converted by a voice recognition system.

[1679] "Emotion analysis means" refers to a technique or device that analyzes text data and identifies the user's emotional state.

[1680] The "means for identifying needs" is a technology or device that clarifies the user's requests and problems based on the content of the user's inquiry and past inquiry history.

[1681] A "personalized solution" is a specific solution or instruction provided to a user that is tailored to their specific situation and needs.

[1682] A "multilingual means" is a technique or device that translates generated solutions into multiple languages ​​and provides them to the user in the appropriate language.

[1683] "Dialogue data" refers to recorded data relating to all statements and inquiries exchanged between the user and the system.

[1684] An "artificial intelligence model" is a model built using machine learning algorithms to learn from large amounts of data and perform specific tasks.

[1685] A "dashboard" is a display device installed inside a vehicle that provides various information to the driver.

[1686] A "head-mounted display" is a device worn on the head that displays information within the field of vision.

[1687] "Driver emotional assessment" refers to the act of analyzing and identifying the driver's emotional state while driving.

[1688] "Real-time processing means" refers to technology or devices that instantly analyze and respond to user input on the spot.

[1689] A system for implementing this invention is configured as follows: First, the driver provides voice input to the system through a terminal (dashboard or head-mounted display). The terminal records the driver's voice and transmits the voice data to a server.

[1690] The server converts the voice data into text using a speech recognition system, which uses the "speech_recognition" library. For example, if a driver says, "I can't set up the navigation system in my car. What should I do?", the speech is converted into text.

[1691] The server then analyzes the converted text data using a natural language processing (NLP) algorithm and a sentiment analysis engine called "transformers" library to extract themes and keywords from the speech and identify the driver's emotional state. For example, if the driver's speech expresses frustration, the sentiment engine classifies the emotional state as "negative."

[1692] The server also retrieves the driver's past inquiry history from the database and analyzes it in combination with the current text data. This identifies the driver's needs and generates personalized solutions. For example, if a problem arises with the navigation system settings, the solution might be to "select an option from the settings menu and change the navigation settings."

[1693] The generated solutions are translated into the driver's preferred language using a multilingual system if necessary. For example, when translating an English solution into Japanese, the server converts it into natural language while preserving the exact meaning. The translated solution is then displayed on the device.

[1694] If the driver tries the proposed solutions and the problem is not resolved, they can again enter additional questions into the device via voice input. The device then sends the additional questions to the server, which analyzes them again and provides further solutions. All of this dialogue data is stored in a database and used to retrain the artificial intelligence model.

[1695] The system's unique features include its ability to assess the driver's emotional state in real time and provide personalized solutions quickly and efficiently. It also supports multiple languages, making it flexible enough for international use.

[1696] Specific examples and examples of prompts for generative AI models

[1697] A concrete example is a scenario in which a driver makes an inquiry about the settings of a navigation system.

[1698] Examples:

[1699] The driver speaks, "I can't set up the navigation system in this car. What should I do?"

[1700] Example prompt for a generative AI model:

[1701] If a user says, "I can't configure the navigation system in my car, what should I do?", convert the speech to text and analyze the sentiment. Generate a solution: "Please reset your navigation system. Select the option from the settings menu and change your navigation settings."

[1702] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1703] Step 1:

[1704] The user speaks to the device in the autonomous vehicle, for example, saying, "I can't set up the navigation system in this car. What should I do?" This speech input is recorded by the device's microphone.

[1705] Step 2:

[1706] The device sends the recorded voice data to the server. The server receives the voice data as input and converts it into text data using a voice recognition system. The "speech_recognition" library is used for voice recognition. As a result of the conversion, the voice data is output as text data: "I can't set up the navigation system in this car. What should I do?"

[1707] Step 3:

[1708] The server analyzes the converted text data using natural language processing algorithms (NLP) and a sentiment analysis engine. Specifically, sentiment analysis is performed using the "transformers" library. From the input text data, themes and keywords are extracted, along with an assessment of the user's emotional state (e.g., frustration). The output of this step is the extracted keywords and the emotional state.

[1709] Step 4:

[1710] The server retrieves past inquiry history from the database and compares it with the analyzed text data. This identifies the driver's needs for the current problem and generates an optimal personalized solution. In this case, specific instructions for configuring the navigation system are generated. The input is the analyzed text data and emotional state, and the output is the generated solution.

[1711] Step 5:

[1712] The server translates the generated solution into the driver's preferred language using a multilingual system. For example, when translating an English solution into Japanese, it converts it into a natural expression while preserving the exact meaning. The input of this step is the generated solution (English), and the output is the translated solution (Japanese).

[1713] Step 6:

[1714] The server sends the translated solution to the device, which then displays it on the car's dashboard or head-mounted display, for example, "Select an option from the settings menu and change your navigation settings."

[1715] Step 7:

[1716] The user tries the proposed solution. If the problem is not resolved, they use voice input again to enter an additional question into the device. The device then sends this additional question back to the server. The server converts it back into text data, analyzes it, and provides additional solutions. This dialogue data is stored in a database and used to retrain the artificial intelligence model. The input is the additional question, and the output is the additional solution.

[1717] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1718] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1719] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1720] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1721] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1722] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1723] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1724] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1725] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1726] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1727] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1728] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1729] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1730] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1731] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1732] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1733] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1734] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1735] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1736] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1737] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1738] The following is further disclosed regarding the above embodiment.

[1739] (Claim 1)

[1740] a means for receiving a customer's speech as a voice input and converting it into text data using a voice recognition system;

[1741] A sentiment analysis means for analyzing the text data and evaluating the sentiment of the customer;

[1742] A means of identifying customer needs and generating personalized solutions;

[1743] A multilingual means of translating the solution into multiple languages;

[1744] a means of providing translated solutions to customers;

[1745] A system that stores customer interaction data and includes a means to retrain AI models.

[1746] (Claim 2)

[1747] 10. The system of claim 1, further comprising means for retrieving a customer's past inquiry history from a database and generating a solution.

[1748] (Claim 3)

[1749] 10. The system of claim 1, further comprising a terminal that provides solutions for display to a customer.

[1750] "Example 1"

[1751] (Claim 1)

[1752] a means for receiving a customer's speech as a voice input and converting it into text data using a voice recognition system;

[1753] A means for analyzing text data and assessing customer sentiment;

[1754] A means of identifying customer needs and generating personalized solutions;

[1755] a means of translating the generated solutions into multiple languages;

[1756] a means of providing translated solutions to customers;

[1757] A means to store customer interaction data and retrain AI models;

[1758] means for transmitting voice data from the terminal to a server;

[1759] A means for retrieving past inquiry history from a database;

[1760] A means for analyzing text data using natural language processing technology;

[1761] A means to ask follow-up questions or make suggestions in a conversational format that increases engagement,

[1762] a means for sending the solution results and / or follow-up questions to the server;

[1763] A system including a means for improving the accuracy of responses to subsequent inquiries based on dialogue data.

[1764] (Claim 2)

[1765] 10. The system of claim 1, further comprising means for retrieving a customer's past inquiry history from a database and generating a solution.

[1766] (Claim 3)

[1767] 10. The system of claim 1, further comprising a terminal that provides solutions for display to a customer.

[1768] "Application Example 1"

[1769] (Claim 1)

[1770] a means for receiving a customer's speech as a voice input and converting it into text data using a voice recognition system;

[1771] A sentiment analysis means for analyzing the text data and evaluating the sentiment of the customer;

[1772] A means of identifying customer needs and generating personalized solutions;

[1773] A multilingual means of translating the solution into multiple languages;

[1774] a means of providing translated solutions to customers;

[1775] A means to store customer interaction data and retrain AI models;

[1776] Examples used include customer service systems that run on hardware including smartphones and service robots, and

[1777] A means to engage with customers through voice, text, and translated solutions, and to engage in conversations based on sentiment assessment and needs

[1778] A system including:

[1779] (Claim 2)

[1780] 10. The system of claim 1, further comprising means for retrieving a customer's past inquiry history from a database and generating a solution.

[1781] (Claim 3)

[1782] 10. The system of claim 1, further comprising a terminal that provides solutions for display to a customer.

[1783] "Example 2: Combining Emotion Engines"

[1784] (Claim 1)

[1785] A means for receiving customer speech as voice input and converting it into text data using voice recognition technology;

[1786] A sentiment analysis means for analyzing the text data and evaluating the sentiment of the customer;

[1787] A means of identifying customer needs and generating personalized solutions;

[1788] a multilingual means for translating the generated solution into multiple languages;

[1789] a means of providing translated solutions to customers;

[1790] A system that stores customer interaction data and includes a means to retrain machine learning models.

[1791] (Claim 2)

[1792] 10. The system of claim 1, further comprising means for retrieving a customer's past inquiry history from a database and generating a solution.

[1793] (Claim 3)

[1794] 10. The system of claim 1, further comprising a terminal that provides solutions for display to a customer.

[1795] "Application example 2 when combining emotion engines"

[1796] (Claim 1)

[1797] a means for receiving a customer's speech as a voice input and converting it into text data using a voice recognition system;

[1798] A sentiment analysis means for analyzing the text data and evaluating the sentiment of the customer;

[1799] A means of identifying customer needs and generating personalized solutions;

[1800] A multilingual means of translating the solution into multiple languages;

[1801] a means of providing translated solutions to customers;

[1802] a means of storing customer interaction data and retraining artificial intelligence models; and

[1803] A means to process the driver's voice input in real time and display a solution on the vehicle's dashboard or head-mounted display;

[1804] means for generating and providing a personalized solution suited to the driver's emotional state based on an emotion assessment using an artificial intelligence model;

[1805] A system including:

[1806] (Claim 2)

[1807] 10. The system of claim 1, further comprising means for retrieving a customer's past inquiry history from a database and generating a solution.

[1808] (Claim 3)

[1809] 10. The system of claim 1, further comprising a terminal that provides solutions for display to a customer. [Explanation of symbols]

[1810] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for receiving a customer's speech as a voice input and converting it into text data using a voice recognition system; A sentiment analysis means for analyzing the text data and evaluating the sentiment of the customer; A means of identifying customer needs and generating personalized solutions; A multilingual means of translating the solution into multiple languages; a means of providing translated solutions to customers; A system that stores customer interaction data and includes a means to retrain AI models.

2. 10. The system of claim 1, further comprising means for retrieving a customer's past inquiry history from a database and generating a solution.

3. 10. The system of claim 1, further comprising a terminal that provides solutions for display to a customer.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A