system

A system using natural language processing and emotion analysis generates personalized and adaptive customer service responses, addressing the challenges of quick and accurate response generation and emotional understanding in customer service systems.

JP2026101313APending Publication Date: 2026-06-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-12-10
Publication Date
2026-06-22

AI Technical Summary

Technical Problem

Existing customer service systems struggle to provide quick and accurate responses to diverse inquiries, especially in situations with limited staff, and lack the ability to continuously adapt to customer emotions and improve responses based on feedback.

Method used

A system that utilizes natural language processing to analyze voice and text data, searches relevant information from a database, and generates optimal responses in real-time, incorporating emotion analysis and user feedback to enhance communication and learning.

Benefits of technology

Enables quick and accurate customer service by providing personalized responses that adapt to customer emotions, improving staff efficiency and customer satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026101313000001_ABST
    Figure 2026101313000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means for receiving audio or text information and analyzing it using natural language processing technology, A means for retrieving information from a storage device that stores related past cases based on the analyzed information, A means for instantly generating the optimal response based on search results and displaying it in the user interface, A means of collecting user feedback and using it to optimize the system, A means of converting voice input into text and sending it to a cloud-based processing device, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In order to quickly and accurately respond to questions and requests from customers in a store, a lot of knowledge and experience are required, but the burden on this is large. Especially in a situation where there are few new staff and insufficient support, there is a problem that it is difficult to give appropriate responses. Also, since the information handled in the store is diverse, it is difficult to constantly keep track of information such as the latest plans and campaigns. It is required to enhance the reliability in communication with customers and improve work efficiency.

Means for Solving the Problems

[0005] This invention provides a means for receiving voice and text data and analyzing customer questions using natural language processing. Based on the analysis results, it has the technology to search for relevant information from a database of past cases and generate the optimal answer. Furthermore, by presenting the generated answer to staff in real time through a user interface, it enables quick and accurate customer service. By continuously collecting user feedback and incorporating it into system learning, it can provide even more accurate support. This improves the quality of customer service while reducing the burden on store staff.

[0006] "Audio or text data" refers to data obtained by converting user-generated audio into text, as well as directly entered text data.

[0007] "Natural language processing" is a technology that analyzes and processes natural language used by humans, and is a method for computers to understand text and audio data.

[0008] "Analysis" is the process of breaking down and understanding input information and extracting the important elements.

[0009] A "database of past cases" is an information repository where customer service cases and FAQs accumulated to date are stored.

[0010] "Searching" is the process of finding information within a database based on specific criteria.

[0011] The "optimal answer" is the information that can respond to the user's question most accurately and in the most timely manner.

[0012] A "user interface" refers to the screens and operating systems that allow a user to interact with a computer system.

[0013] "Presenting in real time" means processing information immediately and displaying it to the user on the spot.

[0014] "Feedback" refers to responses such as opinions and improvement requests received from users.

[0015] "System learning" is a learning process for making the system more accurate using the collected data.

Brief Explanation of Drawings

[0016] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13]It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.

Mode for Carrying Out the Invention

[0017] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0020] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0021] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0024] [First Embodiment]

[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0037] One embodiment of this invention involves constructing a system that processes voice and text data to support customer service. A specific embodiment is described below.

[0038] System Overview

[0039] terminal

[0040] Store staff handling customer inquiries use a terminal to input customer questions via voice or text. The terminal converts the voice data into text in real time, preparing it for natural language processing.

[0041] The entered data is sent to the server in an encrypted form, ensuring secure communication at all times.

[0042] server

[0043] The server receives data sent from the terminal and performs analysis using a natural language processing engine. This analysis identifies key keywords and intents contained in the customer's question.

[0044] Based on the identified information, the system searches past case databases and FAQ resources to extract the most relevant solutions.

[0045] User interface and feedback

[0046] The responses generated by the server are displayed on the terminal's screen in real time. The user interface is intuitive and used for direct explanations to customers.

[0047] Staff can review the displayed information, provide customer support, and offer feedback on the accuracy and usefulness of the answers.

[0048] Specific example

[0049] Example 1: Inquiry about campaign information

[0050] When a customer asks, "Tell me about the current campaign," the user inputs their inquiry into their device using voice input.

[0051] The device converts the audio to text and sends it to the server. The server searches its database for cases related to the "campaign" based on the analyzed results and extracts the current campaign information.

[0052] This information is quickly forwarded, and the terminal screen displays "We are currently running campaigns for Plan A and Plan B." The user sees this and responds to the customer immediately.

[0053] Example 2: Questions about the new plan

[0054] If another customer requests "Please tell me the details of the new plan," the user will enter that intention into the device.

[0055] The server retrieves the latest information about the new plan from the database and generates the most appropriate recommended plan information.

[0056] This information is displayed on the device and provided to the staff in the form of "New Plan X has the following features."

[0057] Thus, this system aims to improve the quality of customer service and optimize operational efficiency. By coordinating servers and terminals to provide users with intuitive and rapid advice, it reduces the burden on stores.

[0058] The following describes the processing flow.

[0059] Step 1:

[0060] The user inputs a question by voice into the device. The device uses speech recognition to convert the voice data into text data in real time.

[0061] Step 2:

[0062] The device temporarily stores the converted text data, encrypts it, and prepares it for transmission to the server. Secure protocols are used to ensure the data's safety.

[0063] Step 3:

[0064] The server analyzes the text data received from the terminal. This analysis uses a natural language processing engine to extract important keywords and customer intent from the input questions.

[0065] Step 4:

[0066] The server searches past case databases and FAQs based on the analysis results. It quickly extracts relevant cases and information and generates multiple candidate answers.

[0067] Step 5:

[0068] An algorithm is applied to select the most relevant and appropriate answer from multiple answer candidates generated by the server. The selected answer is then formatted appropriately and sent to the terminal.

[0069] Step 6:

[0070] The device displays the responses it receives on the user interface in real time. The device organizes the information and presents it visually to the user, making it easier to understand.

[0071] Step 7:

[0072] The user responds to the customer based on the answers displayed on the device. If necessary, additional questions can be asked, and that data can also be entered and processed on the device.

[0073] Step 8:

[0074] When a user provides feedback to the system, the terminal sends that feedback to the server. The server records this feedback and uses it as data to improve the system's accuracy.

[0075] (Example 1)

[0076] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0077] Current customer support systems struggle to process large amounts of data quickly and accurately, making it difficult to provide immediate answers to diverse customer inquiries. Furthermore, there is a need to ensure the security of acquired data, improve the accuracy of responses, and effectively utilize user feedback. Additionally, there is a lack of methods for personalizing responses for individual users and leveraging other information sources when analysis results are insufficient.

[0078] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0079] In this invention, the server includes means at a terminal for acquiring voice or text information, software means for converting the acquired information into structured data, encrypted communication means for securely communicating the acquired data, a natural language processing device for analyzing the received data, information retrieval means for searching a relevant information base based on the analysis results, user-facing display means for generating and displaying optimized responses, and feedback means for collecting user reviews to contribute to the system's adaptive learning. This enables integrated and secure automated customer service.

[0080] A "terminal for acquiring voice or text information" is an information input device that receives inquiries from customers in voice or text format and initially acquires the data necessary for subsequent processing.

[0081] "Software for converting to structured data" refers to a program that analyzes acquired audio or text information and converts it into a data format that can be processed using natural language processing.

[0082] "Encrypted communication methods" refer to communication methods that include technology for encrypting information to securely transmit data and decrypting it at the destination.

[0083] A "natural language processing device" is a software or hardware system that analyzes input text data to perform semantic understanding and intent extraction.

[0084] An "information retrieval tool" is an algorithm or software used to search a database for relevant information based on analyzed intents and keywords.

[0085] "User-facing display means" refers to an interface device or program for displaying answers generated based on search results in a format that is easy for the user to view.

[0086] A "feedback mechanism" is a system for collecting evaluations and opinions on answers provided by users and using them to improve the system and for learning.

[0087] This invention relates to a system that supports customer service using voice and text data. The system is designed to enable users to efficiently process customer inquiries. Specific embodiments are described below.

[0088] System Configuration

[0089] terminal

[0090] The user uses a terminal to input customer inquiries via voice or text form. This terminal uses speech recognition software (e.g., a speech recognition engine) to convert the voice data into text data. Data transfer is conducted via encrypted communication protocols (e.g., TLS / SSL) to ensure the security of the information.

[0091] server

[0092] The server analyzes the received data using natural language processing software (e.g., a natural language processing engine). This analysis identifies the customer's intent, and a database management system (e.g., an SQL database) is used to retrieve relevant information from the database. This then generates the most appropriate response.

[0093] User interface and feedback features

[0094] The server sends the generated responses to the terminal, allowing the user to explain the content to the customer in real time. The user interface is designed with ease of use in mind and is intuitive to operate. Users can provide feedback on the accuracy and usefulness of the provided responses, contributing to the continuous improvement of the system.

[0095] Specific example

[0096] Campaign information inquiry

[0097] When a customer asks, "Tell me about the current campaign," the user inputs the question by voice into the device. The device converts the voice into text and sends the data to the server. The server searches for data related to "campaign" and generates information such as, "We are currently running campaigns Plan A and Plan B." By displaying this information on the device, the user can immediately answer the customer's question.

[0098] Questions about the details of the new plan

[0099] In another scenario, if a customer asks, "Tell me the details of the new plan," the user enters the request into their device. The server retrieves the latest data related to the "new plan" in the most appropriate format and provides the user with information such as, "New Plan X has the following features."

[0100] In this example, a generative AI model is used to optimize prompt text, enabling specific and detailed questions such as "Please provide current campaign information." This system allows users to respond to customer needs quickly and accurately.

[0101] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0102] Step 1:

[0103] The user inputs a question from the customer via voice or text through the terminal. The terminal uses speech recognition software to convert the voice data into text data. Specifically, if the voice input is "I want to know the details of the new plan," the terminal converts this to text "I want to know the details of the new plan" and outputs the converted text.

[0104] Step 2:

[0105] The terminal encrypts the converted text data. For example, AES encryption technology is used. The encrypted data is then sent to the server while maintaining security. The input is the text data "I want to know more about the new plan," and the output is the encrypted text data.

[0106] Step 3:

[0107] The server receives encrypted data from the terminal and performs a decryption process. As a result, the original text data, "I want to know the details of the new plan," becomes the input data on the server side. The server analyzes this text using a natural language processing engine and extracts keywords and intent. For example, "new plan" is extracted as a keyword.

[0108] Step 4:

[0109] The server uses a database management system based on the extracted keywords to search for relevant information within the database. The search results obtained are information about the new plan. For example, "Price and features of new plan X" are obtained.

[0110] Step 5:

[0111] The server generates an answer based on the search results. A generative AI model is used here, which generates more natural sentences by using prompts. If the data input is "Pricing and features of the new plan X," the output information will be "The new plan X costs $100 per month and includes unlimited data usage."

[0112] Step 6:

[0113] The server sends the generated response to the terminal. The terminal displays the received information on its screen. The user reviews the details displayed on the screen and explains them to the customer. Based on the confirmed information, the user can quickly respond to the customer.

[0114] Step 7:

[0115] Users provide feedback on the accuracy and usefulness of the answers provided. This feedback is used by the system's adaptive learning algorithm to improve the quality of future answers. An example of feedback is "The information was very helpful." This helps the system strive to provide even more accurate information.

[0116] (Application Example 1)

[0117] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0118] In physical stores, a challenge exists in that sales staff often struggle to answer customer questions quickly and accurately. This challenge stems from the inability to instantly provide diverse product information, inventory status, and campaign information, which reduces the efficiency of customer service. Therefore, a system is needed that allows sales staff to obtain and provide necessary information to customers in real time.

[0119] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0120] In this invention, the server includes means for receiving voice or text information and analyzing it using natural language processing technology; means for searching for information from a storage device that stores relevant past cases based on the analyzed information; means for immediately generating an optimal response based on the search results and displaying it on the user interface; and means for converting voice input into text and transmitting it to a cloud processing device. This enables salespeople to quickly and accurately obtain the information requested by customers and provide it in real time.

[0121] "Audio or text information" refers to audio or text data received from the user.

[0122] "Natural language processing technology" is a technology that analyzes speech or text information to understand its meaning and extract necessary data.

[0123] A "storage device" refers to a device such as a database used to store past cases and related information.

[0124] A "user interface" refers to a screen or device that displays system output and allows the user to view the information.

[0125] "Voice input" refers to a method of acquiring the content of a user's speech as digital data.

[0126] A "cloud processing device" refers to a remote server or computer device used to process voice input and analyzed information.

[0127] One embodiment of this invention is based on a system designed to facilitate customer service for sales staff in physical stores. The system receives voice and text information and analyzes it using natural language processing technology. This helps sales staff to respond quickly and accurately to various questions from customers.

[0128] The server instantly converts speech data into text using the Google® Cloud Speech-to-Text API. The converted text is sent to the cloud server and analyzed using spaCy, a natural language processing technology. Based on the analysis results and relevant information stored in a MySQL® database, the optimal response is instantly generated and displayed on the salesperson's terminal via the user interface.

[0129] The terminal uses a smartphone as its hardware and is a device for sales staff to input voice commands. This terminal is responsible not only for receiving and displaying information but also for sending user feedback to a server.

[0130] For example, if a customer asks a store clerk, "Is there a discount on this product?", the system will process this voice question immediately and instantly display relevant campaign and discount information on the terminal.

[0131] An example of a prompt message used as an example of the use of the generating AI model is: "When a customer asks a question about a specific product, retrieve the product's features, stock status, and campaign information from the server in real time and display it on the smartphone."

[0132] This configuration allows sales staff to interact with customers more efficiently, resulting in a system that contributes to improved customer satisfaction.

[0133] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0134] Step 1:

[0135] The user uses their smartphone (device) to input customer questions by voice. This input voice is captured by the device as audio data. At this stage, the input is in the form of audio data.

[0136] Step 2:

[0137] The device converts the audio data to text using the Google Cloud Speech-to-Text API. This API analyzes the audio data and generates the corresponding text data. The output of this step is text data.

[0138] Step 3:

[0139] Text data is sent from the terminal to the cloud server. The server receives this text data and analyzes it using spaCy, a natural language processing technology. The analysis involves data processing to understand the meaning of the text and identify the customer's intent. The output of this step is the analyzed intent and keywords.

[0140] Step 4:

[0141] The server searches for relevant information from the MySQL database based on the analyzed intent and keywords. The retrieved information is extracted from relevant cases, FAQs, etc. The input is the analysis result, and the output is the best answer found.

[0142] Step 5:

[0143] The server immediately generates an appropriate response based on the search results. The generated response is then formatted into text using a generative AI model. At this stage, the answer is output in text format.

[0144] Step 6:

[0145] The server sends the generated text-formatted response back to the terminal and displays it in the user interface. The terminal receives this information and outputs it on the screen. At this stage, the salesperson can review the displayed information and communicate it to the customer.

[0146] Step 7:

[0147] Salespeople, acting as users, input customer feedback and their own evaluations into a terminal. This feedback is sent back to the server and used to optimize and improve the system. The input feedback becomes output that contributes to improving the system's future performance.

[0148] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0149] As an embodiment of this invention, a system is constructed that incorporates an emotion engine that processes voice and text data and supports customer service. A specific embodiment thereof is described below.

[0150] System Overview

[0151] terminal

[0152] Store staff assisting customers use a terminal to input customer questions via voice or text. The terminal converts the voice data into text, making it analyzable.

[0153] The device also features an emotion engine that analyzes emotions from voice and text, recognizing customer emotions in real time.

[0154] server

[0155] The server receives voice and text data sent from the terminal and analyzes the content using a natural language processing engine. This extracts the intent behind the customer's question and relevant information.

[0156] Emotional information from the emotion engine is also sent to the server, which then generates responses with appropriate tone and content. Each question is addressed individually based on the results of the emotion analysis.

[0157] User interface and feedback

[0158] The responses generated by the server are displayed on the user interface on the terminal. The displayed information is provided as the most effective means of communication for the user, based on sentiment analysis.

[0159] By following the guidelines provided by the emotion engine when providing responses to customers, users can communicate more effectively. Furthermore, users can continuously improve the quality of their responses by incorporating the feedback they receive into the system.

[0160] Specific example

[0161] Example 1: Responding to customer complaints

[0162] If a customer expresses dissatisfaction and asks for an explanation about a product defect, the user will input their voice message into the device.

[0163] The device's emotion engine recognizes customer dissatisfaction from their voice and sends that information to the server. The server extracts information related to the problem and generates a response in a reassuring tone.

[0164] An explanation such as "We apologize, regarding this issue..." will be displayed on the device, and the user will respond in good faith based on that information.

[0165] Example 2: High-interest inquiries about new products

[0166] If a customer expresses interest and asks, "I want to know more about this new product," the user will input their intention into the terminal.

[0167] The emotion engine recognizes the customer's level of interest, and the server uses that information to provide detailed information about relevant products in an enthusiastic and engaging manner.

[0168] Based on the displayed information, users can communicate the appeal of the new product in a more personalized way.

[0169] In this way, by integrating an emotion engine, this system enables communication that adapts to customer emotions and strengthens support for store staff. Staff will be able to provide highly accurate customer service, contributing to improved customer satisfaction.

[0170] The following describes the processing flow.

[0171] Step 1:

[0172] The user inputs a question by voice into the device. The device uses speech recognition to convert the voice data into text. The text data is temporarily stored.

[0173] Step 2:

[0174] The device's emotion engine analyzes the user's emotions from the voice data. The analysis results are tagged with an emotional state (e.g., dissatisfied, excited, calm).

[0175] Step 3:

[0176] The device sends text and sentiment data to the server. The data is encrypted and transferred using a secure protocol.

[0177] Step 4:

[0178] The server receives text data and analyzes its content using a natural language processing engine. Based on the keywords and phrases extracted through the analysis, it searches a database of past cases.

[0179] Step 5:

[0180] The server takes emotional data into account and adjusts the tone and content of responses accordingly. Depending on the emotion, it selects reassuring tones or encouraging words.

[0181] Step 6:

[0182] The server generates the final response and sends it to the terminal along with an emotion-based communication strategy.

[0183] Step 7:

[0184] The device displays the received response on the user interface. The displayed content is visually organized and in a format that is easy for the user to understand.

[0185] Step 8:

[0186] Users review the information displayed on their devices and use it in their interactions with customers. Feedback received during customer interactions is sent to the server via the device and stored as learning data for future interactions.

[0187] (Example 2)

[0188] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0189] In customer service, rapidly and accurately analyzing voice and text data and providing appropriate responses that reflect customer emotions is a challenge that conventional technologies have not adequately addressed. In particular, achieving this in real time while continuously learning and improving the system is difficult. Therefore, a system is needed that combines more advanced natural language processing and sentiment analysis technologies to provide effective responses to users.

[0190] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0191] In this invention, the server includes means for receiving voice or text data and analyzing it using natural language processing technology, means for analyzing emotional information using emotion analysis technology based on the analyzed information, and means for generating an optimal response in real time using a generative artificial intelligence model based on the analysis results and displaying it on an information display device. This makes it possible to provide appropriate and personalized responses that correspond to the customer's emotions and to continuously improve the system's performance.

[0192] "Voice or text data" refers to voice or text information that users input into the system.

[0193] "Natural language processing technology" refers to the technology that enables computers to understand, analyze, and generate human language.

[0194] "Emotional analysis technology" refers to the technology that extracts and analyzes emotional information from voice and text data.

[0195] A "generative artificial intelligence model" refers to a model that uses machine learning algorithms to generate data-driven results or answers.

[0196] An "information display device" refers to a device or software that visually displays information on a user interface.

[0197] "Feedback" refers to evaluations and opinions provided by users to improve the system's capabilities.

[0198] To implement this invention, a system is constructed that processes voice and text data to support customer service. This system utilizes emotion analysis technology to enable personalized communication.

[0199] The terminal is used by store staff to input customer questions via voice or text. The voice data is converted to text using speech recognition software. Specifically, common speech recognition software and services are used, such as open-source speech recognition libraries and cloud-based speech recognition services. The terminal is equipped with software that incorporates sentiment analysis technology to instantly analyze the customer's emotions and generate an emotion score.

[0200] The server receives voice or text data and sentiment scores transmitted from the terminal. The server analyzes the text data using natural language processing techniques and further evaluates the customer's emotional state based on sentiment analysis techniques. This process utilizes a generative AI model to automatically generate the optimal response from the analyzed information. The generated response is then customized with appropriate tone and content based on the sentiment analysis results.

[0201] This response is immediately displayed on the information display on the terminal. The user uses this information to take appropriate action with the customer. For example, if a customer is dissatisfied with a product, the server will generate a response that includes an apology and a solution to alleviate the dissatisfaction. As a concrete example, an example of a prompt message when a customer makes a complaint is shown below.

[0202] Example prompt: "Please provide more details about the defect in this product."

[0203] In this way, the system provides responses that are appropriate to the customer's emotions, supporting better customer service.

[0204] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0205] Step 1:

[0206] The device receives voice or text data entered by the user. In the case of voice input, speech recognition technology is used to convert the voice into text. In this process, it receives voice or text data as input and generates text data as output. Specifically, it collects voice data in real time, and the speech recognition engine converts it into text format.

[0207] Step 2:

[0208] The device sends the received text data to an emotion analysis engine to generate an emotion score. In this process, text data is provided as input to the emotion analysis model, and an emotion score is obtained as output. Specifically, it analyzes the wording and expressions within the text data to identify the emotional state.

[0209] Step 3:

[0210] The server receives text data and sentiment scores sent from the terminal. Here, the input received is text data and sentiment scores. The server uses natural language processing techniques to analyze the content of the text data and extract customer intent and related information. The output generates data including the analysis results and intent. In this step, natural language processing, including keyword extraction and intent understanding, is performed using the input data, and based on this, the customer's requests are clarified.

[0211] Step 4:

[0212] The server provides the generative AI model with analysis results and sentiment scores to generate the optimal response. The generative AI model receives analyzed intent data and sentiment scores as input and generates a response to convey to the customer as output. Specifically, it produces grammatically correct responses and adjusts the message to an emotionally sensitive tone.

[0213] Step 5:

[0214] The terminal displays the responses generated from the server on the user interface. Here, it receives response data from the server as input and generates text or audio data to be displayed on the screen as output. As a concrete example of its operation, the user interface visually displays the sentiment score while clearly arranging the response messages.

[0215] This process allows users to respond to customers based on the generated answers, leading to more effective communication.

[0216] (Application Example 2)

[0217] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0218] Traditional customer service systems have faced the challenge of not being able to adequately consider customer emotions when providing service. Store staff, when interacting with customers through voice and text data, lack the skills to properly understand the customer's emotional state and communicate appropriately according to that state, which hinders improvements in customer satisfaction.

[0219] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0220] In this invention, the server includes means for receiving voice or text data and analyzing it using natural language processing; means for searching for information from a database of relevant past cases based on the analyzed information and sentiment analysis results from an emotion engine; and means for generating a response in the optimal tone and content in real time based on the search results and sentiment analysis results, and providing it to a display device. This makes it possible to improve customer satisfaction by providing appropriate responses that are in line with the customer's emotions.

[0221] "Audio or text data" refers to audio and text information produced by customers, which is used for communication with customers.

[0222] "Natural language processing" is a technology that analyzes speech or text data to understand its meaning and intent.

[0223] An "emotion engine" is a technology that analyzes and quantifies or categorizes a customer's emotional state from voice and text data.

[0224] A "database of related past cases" is a collection of data that accumulates past inquiries and responses, providing useful information for similar situations.

[0225] "Optimal tone and content" refers to appropriate expressions and information that are tailored to the customer's emotions and situation, thereby enhancing the effectiveness of communication.

[0226] "User feedback" refers to information about the results and effectiveness of customer service, and is used to improve the system.

[0227] A "display device" is a device or system that provides generated information to a user visually.

[0228] To implement this invention, the system provides a technology that uses voice and text data to enable communication tailored to the customer's emotions. Specific embodiments are described below.

[0229] The server uses a speech recognition API to convert the audio data received from the terminal into text data. This converted text is then analyzed by a natural language processing engine (e.g., Dialogflow) to extract the customer's intent and the content of their question. Furthermore, an emotion engine analyzes the customer's emotional state based on this text data and sends the results to the server.

[0230] The terminal is responsible for interacting with customers at the store level. The terminal receives voice input, converts the voice into text appropriately, and sends it to the server. The analysis results and sentiment analysis results are processed on the server to generate responses with the optimal tone and content. This information is displayed in real time on the terminal's display device, allowing users to refer to it and respond appropriately to the customer's emotions.

[0231] Users can provide highly personalized responses based on emotional information derived from the customer's tone of voice and facial expressions. Because these responses are based on automatically generated guidelines, even average operators can provide high-quality customer service. Furthermore, user feedback is collected on the server and used for system optimization and further learning.

[0232] As a concrete example, if a customer is dissatisfied and requests a detailed explanation of a product defect, entering this into the terminal will cause the emotion engine to quantify the dissatisfaction, and the server will automatically generate a response in a reassuring tone. At this point, the server will return a tailored message such as, "We apologize, regarding this issue..." to communicate appropriately with the customer.

[0233] An example of a prompt is: "When a customer asks, 'Please tell me more about the new product,' and the emotion engine indicates a high level of excitement, respond appropriately using relevant information." In this way, it becomes possible to respond to a variety of customer emotional states.

[0234] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0235] Step 1:

[0236] The terminal receives voice input from the customer and converts that voice into text data using a speech recognition API. The input is the customer's voice, and the output is a sentence in text format. This conversion process enables subsequent natural language processing.

[0237] Step 2:

[0238] The server uses a natural language processing engine to analyze the text data received from the terminal. The input is the text data from step 1, and the output is information about the customer's intent and the content of their question. This process helps to understand what the customer is looking for.

[0239] Step 3:

[0240] The server uses the analyzed information and an emotion engine to analyze the customer's emotional state. The input is the text information from step 2, and the output is numerical or categorical data indicating the emotional state. This enables responses tailored to the customer's emotions.

[0241] Step 4:

[0242] The server searches a database of past cases for relevant information based on the sentiment analysis results and intent analysis results. The input is the analysis results from steps 2 and 3, and the output is information based on similar past cases. This allows for the acquisition of appropriate reference information.

[0243] Step 5:

[0244] The server generates a response with the optimal tone and content based on the previous output. The input is the information from step 4 and the emotional state from step 3, and the output is the response text provided to the customer. This process enables responses that are appropriate to the customer's emotions.

[0245] Step 6:

[0246] The terminal displays the generated response on a display device, and the user responds to the customer based on it. The input is the response text generated in step 5, and the output is the user's response to the customer. In this final step, the user can efficiently and effectively meet the customer's needs.

[0247] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0248] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0249] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0250] [Second Embodiment]

[0251] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0252] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0253] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0254] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0255] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0256] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0257] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0258] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0259] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0260] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0261] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0262] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0263] One embodiment of this invention involves constructing a system that processes voice and text data to support customer service. A specific embodiment is described below.

[0264] System Overview

[0265] terminal

[0266] Store staff handling customer inquiries use a terminal to input customer questions via voice or text. The terminal converts the voice data into text in real time, preparing it for natural language processing.

[0267] The entered data is sent to the server in an encrypted form, ensuring secure communication at all times.

[0268] server

[0269] The server receives data sent from the terminal and performs analysis using a natural language processing engine. This analysis identifies key keywords and intents contained in the customer's question.

[0270] Based on the identified information, the system searches past case databases and FAQ resources to extract the most relevant solutions.

[0271] User interface and feedback

[0272] The responses generated by the server are displayed on the terminal's screen in real time. The user interface is intuitive and used for direct explanations to customers.

[0273] Staff can review the displayed information, provide customer support, and offer feedback on the accuracy and usefulness of the answers.

[0274] Specific example

[0275] Example 1: Inquiry about campaign information

[0276] When a customer asks, "Tell me about the current campaign," the user inputs their inquiry into their device using voice input.

[0277] The device converts the audio to text and sends it to the server. The server searches its database for cases related to the "campaign" based on the analyzed results and extracts the current campaign information.

[0278] This information is quickly forwarded, and the terminal screen displays "We are currently running campaigns for Plan A and Plan B." The user sees this and responds to the customer immediately.

[0279] Example 2: Questions about the new plan

[0280] When another customer requests "Please tell me the details of the new plan", the user inputs their intention into the terminal.

[0281] The server then retrieves the latest information about the new plan from the database and generates the most appropriate recommended plan information.

[0282] This information is displayed on the terminal and provided to the staff in the form of "New Plan X has the following features".

[0283] In this way, this system aims to improve the quality of customer service and optimize business efficiency. By enabling the server and the terminal to cooperate and provide intuitive and prompt advice to the user, the burden on the store is reduced.

[0284] The following describes the processing flow.

[0285] Step 1:

[0286] The user inputs a question by voice towards the terminal. The terminal uses a speech recognition function to convert the voice data into text data in real time.

[0287] Step 2:

[0288] The terminal temporarily stores the converted text data and prepares to encrypt the data and send it to the server. A secure protocol is used to ensure the security of the data.

[0289] Step 3:

[0290] The server analyzes the text data received from the terminal. For this analysis, a natural language processing engine is used to extract important keywords and customer intentions from the input question.

[0291] Step 4:

[0292] The server searches past case databases and FAQs based on the analysis results. It quickly extracts relevant cases and information and generates multiple candidate answers.

[0293] Step 5:

[0294] An algorithm is applied to select the most relevant and appropriate answer from multiple answer candidates generated by the server. The selected answer is then formatted appropriately and sent to the terminal.

[0295] Step 6:

[0296] The device displays the responses it receives on the user interface in real time. The device organizes the information and presents it visually to the user, making it easier to understand.

[0297] Step 7:

[0298] The user responds to the customer based on the answers displayed on the device. If necessary, additional questions can be asked, and that data can also be entered and processed on the device.

[0299] Step 8:

[0300] When a user provides feedback to the system, the terminal sends that feedback to the server. The server records this feedback and uses it as data to improve the system's accuracy.

[0301] (Example 1)

[0302] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0303] The current customer support system has problems in that it is difficult to process a large amount of data quickly and accurately, and it cannot immediately answer various questions from customers. Also, while ensuring the security of the acquired data, there is a need to improve the accuracy of answers and effectively utilize user feedback. Furthermore, there is a lack of methods for personalizing answers for individual users and for utilizing other information sources when the analysis results are insufficient.

[0304] The specific processing by the specific processing unit 290 of the data processing apparatus 12 in Example 1 is realized by the following means.

[0305] In this invention, the server includes means in a terminal for acquiring voice or text information, software means for converting the acquired information into structured data, encrypted communication means for securely communicating the acquired data, a natural language processing apparatus for analyzing the received data, information retrieval means for searching an associated information base based on the analysis results, display means for users for generating and displaying an optimized answer, and feedback means for collecting reviews from users and contributing to the adaptive learning of the system. Thereby, integrated and secure automation of customer support becomes possible.

[0306] The "terminal for acquiring voice or text information" is an information input device that receives inquiries in voice or text form from customers and initially acquires the data necessary for subsequent processing.

[0307] The "software for converting into structured data" is a program that analyzes the acquired voice or text information and converts it into a data format capable of natural language processing.

[0308] The "encrypted communication means" is a communication method that includes a technique for encrypting information to securely communicate data and decrypting it at the destination.

[0309] The "natural language processing apparatus" is a software or hardware system for analyzing the input text data and performing semantic understanding and intention extraction.

[0310] An "information retrieval tool" is an algorithm or software used to search a database for relevant information based on analyzed intents and keywords.

[0311] "User-facing display means" refers to an interface device or program for displaying answers generated based on search results in a format that is easy for the user to view.

[0312] A "feedback mechanism" is a system for collecting evaluations and opinions on answers provided by users and using them to improve the system and for learning.

[0313] This invention relates to a system that supports customer service using voice and text data. The system is designed to enable users to efficiently process customer inquiries. Specific embodiments are described below.

[0314] System Configuration

[0315] terminal

[0316] The user uses a terminal to input customer inquiries via voice or text form. This terminal uses speech recognition software (e.g., a speech recognition engine) to convert the voice data into text data. Data transfer is conducted via encrypted communication protocols (e.g., TLS / SSL) to ensure the security of the information.

[0317] server

[0318] The server analyzes the received data using natural language processing software (e.g., a natural language processing engine). This analysis identifies the customer's intent, and a database management system (e.g., an SQL database) is used to retrieve relevant information from the database. This then generates the most appropriate response.

[0319] User interface and feedback features

[0320] The server sends the generated responses to the terminal, allowing the user to explain the content to the customer in real time. The user interface is designed with ease of use in mind and is intuitive to operate. Users can provide feedback on the accuracy and usefulness of the provided responses, contributing to the continuous improvement of the system.

[0321] Specific example

[0322] Campaign information inquiry

[0323] When a customer asks, "Tell me about the current campaign," the user inputs the question by voice into the device. The device converts the voice into text and sends the data to the server. The server searches for data related to "campaign" and generates information such as, "We are currently running campaigns Plan A and Plan B." By displaying this information on the device, the user can immediately answer the customer's question.

[0324] Questions about the details of the new plan

[0325] In another scenario, if a customer asks, "Tell me the details of the new plan," the user enters the request into their device. The server retrieves the latest data related to the "new plan" in the most appropriate format and provides the user with information such as, "New Plan X has the following features."

[0326] In this example, a generative AI model is used to optimize prompt text, enabling specific and detailed questions such as "Please provide current campaign information." This system allows users to respond to customer needs quickly and accurately.

[0327] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0328] Step 1:

[0329] The user inputs a question from the customer via voice or text through the terminal. The terminal uses speech recognition software to convert the voice data into text data. Specifically, if the voice input is "I want to know the details of the new plan," the terminal converts this to text "I want to know the details of the new plan" and outputs the converted text.

[0330] Step 2:

[0331] The terminal encrypts the converted text data. For example, AES encryption technology is used. The encrypted data is then sent to the server while maintaining security. The input is the text data "I want to know more about the new plan," and the output is the encrypted text data.

[0332] Step 3:

[0333] The server receives encrypted data from the terminal and performs a decryption process. As a result, the original text data, "I want to know the details of the new plan," becomes the input data on the server side. The server analyzes this text using a natural language processing engine and extracts keywords and intent. For example, "new plan" is extracted as a keyword.

[0334] Step 4:

[0335] The server uses a database management system based on the extracted keywords to search for relevant information within the database. The search results obtained are information about the new plan. For example, "Price and features of new plan X" are obtained.

[0336] Step 5:

[0337] The server generates an answer based on the search results. A generative AI model is used here, which generates more natural sentences by using prompts. If the data input is "Pricing and features of the new plan X," the output information will be "The new plan X costs $100 per month and includes unlimited data usage."

[0338] Step 6:

[0339] The server sends the generated response to the terminal. The terminal displays the received information on its screen. The user reviews the details displayed on the screen and explains them to the customer. Based on the confirmed information, the user can quickly respond to the customer.

[0340] Step 7:

[0341] Users provide feedback on the accuracy and usefulness of the answers provided. This feedback is used by the system's adaptive learning algorithm to improve the quality of future answers. An example of feedback is "The information was very helpful." This helps the system strive to provide even more accurate information.

[0342] (Application Example 1)

[0343] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0344] In physical stores, a challenge exists in that sales staff often struggle to answer customer questions quickly and accurately. This challenge stems from the inability to instantly provide diverse product information, inventory status, and campaign information, which reduces the efficiency of customer service. Therefore, a system is needed that allows sales staff to obtain and provide necessary information to customers in real time.

[0345] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0346] In this invention, the server includes means for receiving voice or text information and analyzing it using natural language processing technology; means for searching for information from a storage device that stores relevant past cases based on the analyzed information; means for immediately generating an optimal response based on the search results and displaying it on the user interface; and means for converting voice input into text and transmitting it to a cloud processing device. This enables salespeople to quickly and accurately obtain the information requested by customers and provide it in real time.

[0347] "Audio or text information" refers to audio or text data received from the user.

[0348] "Natural language processing technology" is a technology that analyzes speech or text information to understand its meaning and extract necessary data.

[0349] A "storage device" refers to a device such as a database used to store past cases and related information.

[0350] A "user interface" refers to a screen or device that displays system output and allows the user to view the information.

[0351] "Voice input" refers to a method of acquiring the content of a user's speech as digital data.

[0352] A "cloud processing device" refers to a remote server or computer device used to process voice input and analyzed information.

[0353] One embodiment of this invention is based on a system designed to facilitate customer service for sales staff in physical stores. The system receives voice and text information and analyzes it using natural language processing technology. This helps sales staff to respond quickly and accurately to various questions from customers.

[0354] The server instantly converts speech data into text using the Google Cloud Speech-to-Text API. The converted text is sent to the cloud server and analyzed using spaCy, a natural language processing technology. Based on the analysis results and relevant information stored in a MySQL database, the optimal response is instantly generated and displayed on the salesperson's terminal via the user interface.

[0355] The terminal uses a smartphone as its hardware and is a device for sales staff to input voice commands. This terminal is responsible not only for receiving and displaying information but also for sending user feedback to a server.

[0356] For example, if a customer asks a store clerk, "Is there a discount on this product?", the system will process this voice question immediately and instantly display relevant campaign and discount information on the terminal.

[0357] An example of a prompt message used as an example of the use of the generating AI model is: "When a customer asks a question about a specific product, retrieve the product's features, stock status, and campaign information from the server in real time and display it on the smartphone."

[0358] This configuration allows sales staff to interact with customers more efficiently, resulting in a system that contributes to improved customer satisfaction.

[0359] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0360] Step 1:

[0361] The user uses their smartphone (device) to input customer questions by voice. This input voice is captured by the device as audio data. At this stage, the input is in the form of audio data.

[0362] Step 2:

[0363] The device converts the audio data to text using the Google Cloud Speech-to-Text API. This API analyzes the audio data and generates the corresponding text data. The output of this step is text data.

[0364] Step 3:

[0365] Text data is sent from the terminal to the cloud server. The server receives this text data and analyzes it using spaCy, a natural language processing technology. The analysis involves data processing to understand the meaning of the text and identify the customer's intent. The output of this step is the analyzed intent and keywords.

[0366] Step 4:

[0367] The server searches for relevant information from the MySQL database based on the analyzed intent and keywords. The retrieved information is extracted from relevant cases, FAQs, etc. The input is the analysis result, and the output is the best answer found.

[0368] Step 5:

[0369] The server immediately generates an appropriate response based on the search results. The generated response is then formatted into text using a generative AI model. At this stage, the answer is output in text format.

[0370] Step 6:

[0371] The server sends the generated text-formatted response back to the terminal and displays it in the user interface. The terminal receives this information and outputs it on the screen. At this stage, the salesperson can review the displayed information and communicate it to the customer.

[0372] Step 7:

[0373] Salespeople, acting as users, input customer feedback and their own evaluations into a terminal. This feedback is sent back to the server and used to optimize and improve the system. The input feedback becomes output that contributes to improving the system's future performance.

[0374] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0375] As an embodiment of this invention, a system is constructed that incorporates an emotion engine that processes voice and text data and supports customer service. A specific embodiment thereof is described below.

[0376] System Overview

[0377] terminal

[0378] Store staff assisting customers use a terminal to input customer questions via voice or text. The terminal converts the voice data into text, making it analyzable.

[0379] The device also features an emotion engine that analyzes emotions from voice and text, recognizing customer emotions in real time.

[0380] server

[0381] The server receives voice and text data sent from the terminal and analyzes the content using a natural language processing engine. This extracts the intent behind the customer's question and relevant information.

[0382] Emotional information from the emotion engine is also sent to the server, which then generates responses with appropriate tone and content. Each question is addressed individually based on the results of the emotion analysis.

[0383] User interface and feedback

[0384] The responses generated by the server are displayed on the user interface on the terminal. The displayed information is provided as the most effective means of communication for the user, based on sentiment analysis.

[0385] By following the guidelines provided by the emotion engine when providing responses to customers, users can communicate more effectively. Furthermore, users can continuously improve the quality of their responses by incorporating the feedback they receive into the system.

[0386] Specific example

[0387] Example 1: Responding to customer complaints

[0388] If a customer expresses dissatisfaction and asks for an explanation about a product defect, the user will input their voice message into the device.

[0389] The device's emotion engine recognizes customer dissatisfaction from their voice and sends that information to the server. The server extracts information related to the problem and generates a response in a reassuring tone.

[0390] An explanation such as "We apologize, regarding this issue..." will be displayed on the device, and the user will respond in good faith based on that information.

[0391] Example 2: High-interest inquiries about new products

[0392] If a customer expresses interest and asks, "I want to know more about this new product," the user will input their intention into the terminal.

[0393] The emotion engine recognizes the customer's level of interest, and the server uses that information to provide detailed information about relevant products in an enthusiastic and engaging manner.

[0394] Based on the displayed information, users can communicate the appeal of the new product in a more personalized way.

[0395] In this way, by integrating an emotion engine, this system enables communication that adapts to customer emotions and strengthens support for store staff. Staff will be able to provide highly accurate customer service, contributing to improved customer satisfaction.

[0396] The following describes the processing flow.

[0397] Step 1:

[0398] The user inputs a question by voice into the device. The device uses speech recognition to convert the voice data into text. The text data is temporarily stored.

[0399] Step 2:

[0400] The device's emotion engine analyzes the user's emotions from the voice data. The analysis results are tagged with an emotional state (e.g., dissatisfied, excited, calm).

[0401] Step 3:

[0402] The device sends text and sentiment data to the server. The data is encrypted and transferred using a secure protocol.

[0403] Step 4:

[0404] The server receives text data and analyzes its content using a natural language processing engine. Based on the keywords and phrases extracted through the analysis, it searches a database of past cases.

[0405] Step 5:

[0406] The server takes emotional data into account and adjusts the tone and content of responses accordingly. Depending on the emotion, it selects reassuring tones or encouraging words.

[0407] Step 6:

[0408] The server generates the final response and sends it to the terminal along with an emotion-based communication strategy.

[0409] Step 7:

[0410] The device displays the received response on the user interface. The displayed content is visually organized and in a format that is easy for the user to understand.

[0411] Step 8:

[0412] Users review the information displayed on their devices and use it in their interactions with customers. Feedback received during customer interactions is sent to the server via the device and stored as learning data for future interactions.

[0413] (Example 2)

[0414] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0415] In customer service, rapidly and accurately analyzing voice and text data and providing appropriate responses that reflect customer emotions is a challenge that conventional technologies have not adequately addressed. In particular, achieving this in real time while continuously learning and improving the system is difficult. Therefore, a system is needed that combines more advanced natural language processing and sentiment analysis technologies to provide effective responses to users.

[0416] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0417] In this invention, the server includes means for receiving voice or text data and analyzing it using natural language processing technology, means for analyzing emotional information using emotion analysis technology based on the analyzed information, and means for generating an optimal response in real time using a generative artificial intelligence model based on the analysis results and displaying it on an information display device. This makes it possible to provide appropriate and personalized responses that correspond to the customer's emotions and to continuously improve the system's performance.

[0418] "Voice or text data" refers to voice or text information that users input into the system.

[0419] "Natural language processing technology" refers to the technology that enables computers to understand, analyze, and generate human language.

[0420] "Emotional analysis technology" refers to the technology that extracts and analyzes emotional information from voice and text data.

[0421] A "generative artificial intelligence model" refers to a model that uses machine learning algorithms to generate data-driven results or answers.

[0422] An "information display device" refers to a device or software that visually displays information on a user interface.

[0423] "Feedback" refers to evaluations and opinions provided by users to improve the system's capabilities.

[0424] To implement this invention, a system is constructed that processes voice and text data to support customer service. This system utilizes emotion analysis technology to enable personalized communication.

[0425] The terminal is used by store staff to input customer questions via voice or text. The voice data is converted to text using speech recognition software. Specifically, common speech recognition software and services are used, such as open-source speech recognition libraries and cloud-based speech recognition services. The terminal is equipped with software that incorporates sentiment analysis technology to instantly analyze the customer's emotions and generate an emotion score.

[0426] The server receives voice or text data and sentiment scores transmitted from the terminal. The server analyzes the text data using natural language processing techniques and further evaluates the customer's emotional state based on sentiment analysis techniques. This process utilizes a generative AI model to automatically generate the optimal response from the analyzed information. The generated response is then customized with appropriate tone and content based on the sentiment analysis results.

[0427] This response is immediately displayed on the information display on the terminal. The user uses this information to take appropriate action with the customer. For example, if a customer is dissatisfied with a product, the server will generate a response that includes an apology and a solution to alleviate the dissatisfaction. As a concrete example, an example of a prompt message when a customer makes a complaint is shown below.

[0428] Example prompt: "Please provide more details about the defect in this product."

[0429] In this way, the system provides responses that are appropriate to the customer's emotions, supporting better customer service.

[0430] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0431] Step 1:

[0432] The device receives voice or text data entered by the user. In the case of voice input, speech recognition technology is used to convert the voice into text. In this process, it receives voice or text data as input and generates text data as output. Specifically, it collects voice data in real time, and the speech recognition engine converts it into text format.

[0433] Step 2:

[0434] The device sends the received text data to an emotion analysis engine to generate an emotion score. In this process, text data is provided as input to the emotion analysis model, and an emotion score is obtained as output. Specifically, it analyzes the wording and expressions within the text data to identify the emotional state.

[0435] Step 3:

[0436] The server receives text data and sentiment scores sent from the terminal. Here, the input received is text data and sentiment scores. The server uses natural language processing techniques to analyze the content of the text data and extract customer intent and related information. The output generates data including the analysis results and intent. In this step, natural language processing, including keyword extraction and intent understanding, is performed using the input data, and based on this, the customer's requests are clarified.

[0437] Step 4:

[0438] The server provides the generative AI model with analysis results and sentiment scores to generate the optimal response. The generative AI model receives analyzed intent data and sentiment scores as input and generates a response to convey to the customer as output. Specifically, it produces grammatically correct responses and adjusts the message to an emotionally sensitive tone.

[0439] Step 5:

[0440] The terminal displays the responses generated from the server on the user interface. Here, it receives response data from the server as input and generates text or audio data to be displayed on the screen as output. As a concrete example of its operation, the user interface visually displays the sentiment score while clearly arranging the response messages.

[0441] This process allows users to respond to customers based on the generated answers, leading to more effective communication.

[0442] (Application Example 2)

[0443] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0444] Traditional customer service systems have faced the challenge of not being able to adequately consider customer emotions when providing service. Store staff, when interacting with customers through voice and text data, lack the skills to properly understand the customer's emotional state and communicate appropriately according to that state, which hinders improvements in customer satisfaction.

[0445] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0446] In this invention, the server includes means for receiving voice or text data and analyzing it using natural language processing; means for searching for information from a database of relevant past cases based on the analyzed information and sentiment analysis results from an emotion engine; and means for generating a response in the optimal tone and content in real time based on the search results and sentiment analysis results, and providing it to a display device. This makes it possible to improve customer satisfaction by providing appropriate responses that are in line with the customer's emotions.

[0447] "Audio or text data" refers to audio and text information produced by customers, which is used for communication with customers.

[0448] "Natural language processing" is a technology that analyzes speech or text data to understand its meaning and intent.

[0449] An "emotion engine" is a technology that analyzes and quantifies or categorizes a customer's emotional state from voice and text data.

[0450] A "database of related past cases" is a collection of data that accumulates past inquiries and responses, providing useful information for similar situations.

[0451] "Optimal tone and content" refers to appropriate expressions and information that are tailored to the customer's emotions and situation, thereby enhancing the effectiveness of communication.

[0452] "User feedback" refers to information about the results and effectiveness of customer service, and is used to improve the system.

[0453] A "display device" is a device or system that provides generated information to a user visually.

[0454] To implement this invention, the system provides a technology that uses voice and text data to enable communication tailored to the customer's emotions. Specific embodiments are described below.

[0455] The server uses a speech recognition API to convert the audio data received from the terminal into text data. This converted text is then analyzed by a natural language processing engine (e.g., Dialogflow) to extract the customer's intent and the content of their question. Furthermore, an emotion engine analyzes the customer's emotional state based on this text data and sends the results to the server.

[0456] The terminal is responsible for interacting with customers at the store level. The terminal receives voice input, converts the voice into text appropriately, and sends it to the server. The analysis results and sentiment analysis results are processed on the server to generate responses with the optimal tone and content. This information is displayed in real time on the terminal's display device, allowing users to refer to it and respond appropriately to the customer's emotions.

[0457] Users can provide highly personalized responses based on emotional information derived from the customer's tone of voice and facial expressions. Because these responses are based on automatically generated guidelines, even average operators can provide high-quality customer service. Furthermore, user feedback is collected on the server and used for system optimization and further learning.

[0458] As a concrete example, if a customer is dissatisfied and requests a detailed explanation of a product defect, entering this into the terminal will cause the emotion engine to quantify the dissatisfaction, and the server will automatically generate a response in a reassuring tone. At this point, the server will return a tailored message such as, "We apologize, regarding this issue..." to communicate appropriately with the customer.

[0459] An example of a prompt is: "When a customer asks, 'Please tell me more about the new product,' and the emotion engine indicates a high level of excitement, respond appropriately using relevant information." In this way, it becomes possible to respond to a variety of customer emotional states.

[0460] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0461] Step 1:

[0462] The terminal receives voice input from the customer and converts that voice into text data using a speech recognition API. The input is the customer's voice, and the output is a sentence in text format. This conversion process enables subsequent natural language processing.

[0463] Step 2:

[0464] The server uses a natural language processing engine to analyze the text data received from the terminal. The input is the text data from step 1, and the output is information about the customer's intent and the content of their question. This process helps to understand what the customer is looking for.

[0465] Step 3:

[0466] The server uses the analyzed information and an emotion engine to analyze the customer's emotional state. The input is the text information from step 2, and the output is numerical or categorical data indicating the emotional state. This enables responses tailored to the customer's emotions.

[0467] Step 4:

[0468] The server searches a database of past cases for relevant information based on the sentiment analysis results and intent analysis results. The input is the analysis results from steps 2 and 3, and the output is information based on similar past cases. This allows for the acquisition of appropriate reference information.

[0469] Step 5:

[0470] The server generates a response with the optimal tone and content based on the previous output. The input is the information from step 4 and the emotional state from step 3, and the output is the response text provided to the customer. This process enables responses that are appropriate to the customer's emotions.

[0471] Step 6:

[0472] The terminal displays the generated response on a display device, and the user responds to the customer based on it. The input is the response text generated in step 5, and the output is the user's response to the customer. In this final step, the user can efficiently and effectively meet the customer's needs.

[0473] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0474] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0475] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0476] [Third Embodiment]

[0477] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0478] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0479] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0480] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0481] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0482] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0483] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0484] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0485] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0486] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0487] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0488] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0489] One embodiment of this invention involves constructing a system that processes voice and text data to support customer service. A specific embodiment is described below.

[0490] System Overview

[0491] terminal

[0492] Store staff handling customer inquiries use a terminal to input customer questions via voice or text. The terminal converts the voice data into text in real time, preparing it for natural language processing.

[0493] The entered data is sent to the server in an encrypted form, ensuring secure communication at all times.

[0494] server

[0495] The server receives data sent from the terminal and performs analysis using a natural language processing engine. This analysis identifies key keywords and intents contained in the customer's question.

[0496] Based on the identified information, the system searches past case databases and FAQ resources to extract the most relevant solutions.

[0497] User interface and feedback

[0498] The responses generated by the server are displayed on the terminal's screen in real time. The user interface is intuitive and used for direct explanations to customers.

[0499] Staff can review the displayed information, provide customer support, and offer feedback on the accuracy and usefulness of the answers.

[0500] Specific example

[0501] Example 1: Inquiry about campaign information

[0502] When a customer asks, "Tell me about the current campaign," the user inputs their inquiry into their device using voice input.

[0503] The device converts the audio to text and sends it to the server. The server searches its database for cases related to the "campaign" based on the analyzed results and extracts the current campaign information.

[0504] This information is quickly forwarded, and the terminal screen displays "We are currently running campaigns for Plan A and Plan B." The user sees this and responds to the customer immediately.

[0505] Example 2: Questions about the new plan

[0506] If another customer requests "Please tell me the details of the new plan," the user will enter that intention into the device.

[0507] The server retrieves the latest information about the new plan from the database and generates the most appropriate recommended plan information.

[0508] This information is displayed on the device and provided to the staff in the form of "New Plan X has the following features."

[0509] Thus, this system aims to improve the quality of customer service and optimize operational efficiency. By coordinating servers and terminals to provide users with intuitive and rapid advice, it reduces the burden on stores.

[0510] The following describes the processing flow.

[0511] Step 1:

[0512] The user inputs a question by voice into the device. The device uses speech recognition to convert the voice data into text data in real time.

[0513] Step 2:

[0514] The device temporarily stores the converted text data, encrypts it, and prepares it for transmission to the server. Secure protocols are used to ensure the data's safety.

[0515] Step 3:

[0516] The server analyzes the text data received from the terminal. This analysis uses a natural language processing engine to extract important keywords and customer intent from the input questions.

[0517] Step 4:

[0518] The server searches past case databases and FAQs based on the analysis results. It quickly extracts relevant cases and information and generates multiple candidate answers.

[0519] Step 5:

[0520] An algorithm is applied to select the most relevant and appropriate answer from multiple answer candidates generated by the server. The selected answer is then formatted appropriately and sent to the terminal.

[0521] Step 6:

[0522] The device displays the responses it receives on the user interface in real time. The device organizes the information and presents it visually to the user, making it easier to understand.

[0523] Step 7:

[0524] The user responds to the customer based on the answers displayed on the device. If necessary, additional questions can be asked, and that data can also be entered and processed on the device.

[0525] Step 8:

[0526] When a user provides feedback to the system, the terminal sends that feedback to the server. The server records this feedback and uses it as data to improve the system's accuracy.

[0527] (Example 1)

[0528] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0529] Current customer support systems struggle to process large amounts of data quickly and accurately, making it difficult to provide immediate answers to diverse customer inquiries. Furthermore, there is a need to ensure the security of acquired data, improve the accuracy of responses, and effectively utilize user feedback. Additionally, there is a lack of methods for personalizing responses for individual users and leveraging other information sources when analysis results are insufficient.

[0530] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0531] In this invention, the server includes means at a terminal for acquiring voice or text information, software means for converting the acquired information into structured data, encrypted communication means for securely communicating the acquired data, a natural language processing device for analyzing the received data, information retrieval means for searching a relevant information base based on the analysis results, user-facing display means for generating and displaying optimized responses, and feedback means for collecting user reviews to contribute to the system's adaptive learning. This enables integrated and secure automated customer service.

[0532] A "terminal for acquiring voice or text information" is an information input device that receives inquiries from customers in voice or text format and initially acquires the data necessary for subsequent processing.

[0533] "Software for converting to structured data" refers to a program that analyzes acquired audio or text information and converts it into a data format that can be processed using natural language processing.

[0534] "Encrypted communication methods" refer to communication methods that include technology for encrypting information to securely transmit data and decrypting it at the destination.

[0535] A "natural language processing device" is a software or hardware system that analyzes input text data to perform semantic understanding and intent extraction.

[0536] An "information retrieval tool" is an algorithm or software used to search a database for relevant information based on analyzed intents and keywords.

[0537] "User-facing display means" refers to an interface device or program for displaying answers generated based on search results in a format that is easy for the user to view.

[0538] A "feedback mechanism" is a system for collecting evaluations and opinions on answers provided by users and using them to improve the system and for learning.

[0539] This invention relates to a system that supports customer service using voice and text data. The system is designed to enable users to efficiently process customer inquiries. Specific embodiments are described below.

[0540] System Configuration

[0541] terminal

[0542] The user uses a terminal to input customer inquiries via voice or text form. This terminal uses speech recognition software (e.g., a speech recognition engine) to convert the voice data into text data. Data transfer is conducted via encrypted communication protocols (e.g., TLS / SSL) to ensure the security of the information.

[0543] server

[0544] The server analyzes the received data using natural language processing software (e.g., a natural language processing engine). This analysis identifies the customer's intent, and a database management system (e.g., an SQL database) is used to retrieve relevant information from the database. This then generates the most appropriate response.

[0545] User interface and feedback features

[0546] The server sends the generated responses to the terminal, allowing the user to explain the content to the customer in real time. The user interface is designed with ease of use in mind and is intuitive to operate. Users can provide feedback on the accuracy and usefulness of the provided responses, contributing to the continuous improvement of the system.

[0547] Specific example

[0548] Campaign information inquiry

[0549] When a customer asks, "Tell me about the current campaign," the user inputs the question by voice into the device. The device converts the voice into text and sends the data to the server. The server searches for data related to "campaign" and generates information such as, "We are currently running campaigns Plan A and Plan B." By displaying this information on the device, the user can immediately answer the customer's question.

[0550] Questions about the details of the new plan

[0551] In another scenario, if a customer asks, "Tell me the details of the new plan," the user enters the request into their device. The server retrieves the latest data related to the "new plan" in the most appropriate format and provides the user with information such as, "New Plan X has the following features."

[0552] In this example, a generative AI model is used to optimize prompt text, enabling specific and detailed questions such as "Please provide current campaign information." This system allows users to respond to customer needs quickly and accurately.

[0553] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0554] Step 1:

[0555] The user inputs a question from the customer via voice or text through the terminal. The terminal uses speech recognition software to convert the voice data into text data. Specifically, if the voice input is "I want to know the details of the new plan," the terminal converts this to text "I want to know the details of the new plan" and outputs the converted text.

[0556] Step 2:

[0557] The terminal encrypts the converted text data. For example, AES encryption technology is used. The encrypted data is then sent to the server while maintaining security. The input is the text data "I want to know more about the new plan," and the output is the encrypted text data.

[0558] Step 3:

[0559] The server receives encrypted data from the terminal and performs a decryption process. As a result, the original text data, "I want to know the details of the new plan," becomes the input data on the server side. The server analyzes this text using a natural language processing engine and extracts keywords and intent. For example, "new plan" is extracted as a keyword.

[0560] Step 4:

[0561] The server uses a database management system based on the extracted keywords to search for relevant information within the database. The search results obtained are information about the new plan. For example, "Price and features of new plan X" are obtained.

[0562] Step 5:

[0563] The server generates an answer based on the search results. A generative AI model is used here, which generates more natural sentences by using prompts. If the data input is "Pricing and features of the new plan X," the output information will be "The new plan X costs $100 per month and includes unlimited data usage."

[0564] Step 6:

[0565] The server sends the generated response to the terminal. The terminal displays the received information on its screen. The user reviews the details displayed on the screen and explains them to the customer. Based on the confirmed information, the user can quickly respond to the customer.

[0566] Step 7:

[0567] Users provide feedback on the accuracy and usefulness of the answers provided. This feedback is used by the system's adaptive learning algorithm to improve the quality of future answers. An example of feedback is "The information was very helpful." This helps the system strive to provide even more accurate information.

[0568] (Application Example 1)

[0569] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0570] In physical stores, a challenge exists in that sales staff often struggle to answer customer questions quickly and accurately. This challenge stems from the inability to instantly provide diverse product information, inventory status, and campaign information, which reduces the efficiency of customer service. Therefore, a system is needed that allows sales staff to obtain and provide necessary information to customers in real time.

[0571] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0572] In this invention, the server includes means for receiving voice or text information and analyzing it using natural language processing technology; means for searching for information from a storage device that stores relevant past cases based on the analyzed information; means for immediately generating an optimal response based on the search results and displaying it on the user interface; and means for converting voice input into text and transmitting it to a cloud processing device. This enables salespeople to quickly and accurately obtain the information requested by customers and provide it in real time.

[0573] "Audio or text information" refers to audio or text data received from the user.

[0574] "Natural language processing technology" is a technology that analyzes speech or text information to understand its meaning and extract necessary data.

[0575] A "storage device" refers to a device such as a database used to store past cases and related information.

[0576] A "user interface" refers to a screen or device that displays system output and allows the user to view the information.

[0577] "Voice input" refers to a method of acquiring the content of a user's speech as digital data.

[0578] A "cloud processing device" refers to a remote server or computer device used to process voice input and analyzed information.

[0579] One embodiment of this invention is based on a system designed to facilitate customer service for sales staff in physical stores. The system receives voice and text information and analyzes it using natural language processing technology. This helps sales staff to respond quickly and accurately to various questions from customers.

[0580] The server instantly converts speech data into text using the Google Cloud Speech-to-Text API. The converted text is sent to the cloud server and analyzed using spaCy, a natural language processing technology. Based on the analysis results and relevant information stored in a MySQL database, the optimal response is instantly generated and displayed on the salesperson's terminal via the user interface.

[0581] The terminal uses a smartphone as its hardware and is a device for sales staff to input voice commands. This terminal is responsible not only for receiving and displaying information but also for sending user feedback to a server.

[0582] For example, if a customer asks a store clerk, "Is there a discount on this product?", the system will process this voice question immediately and instantly display relevant campaign and discount information on the terminal.

[0583] An example of a prompt message used as an example of the use of the generating AI model is: "When a customer asks a question about a specific product, retrieve the product's features, stock status, and campaign information from the server in real time and display it on the smartphone."

[0584] This configuration allows sales staff to interact with customers more efficiently, resulting in a system that contributes to improved customer satisfaction.

[0585] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0586] Step 1:

[0587] The user uses their smartphone (device) to input customer questions by voice. This input voice is captured by the device as audio data. At this stage, the input is in the form of audio data.

[0588] Step 2:

[0589] The device converts the audio data to text using the Google Cloud Speech-to-Text API. This API analyzes the audio data and generates the corresponding text data. The output of this step is text data.

[0590] Step 3:

[0591] Text data is sent from the terminal to the cloud server. The server receives this text data and analyzes it using spaCy, a natural language processing technology. The analysis involves data processing to understand the meaning of the text and identify the customer's intent. The output of this step is the analyzed intent and keywords.

[0592] Step 4:

[0593] The server searches for relevant information from the MySQL database based on the analyzed intent and keywords. The retrieved information is extracted from relevant cases, FAQs, etc. The input is the analysis result, and the output is the best answer found.

[0594] Step 5:

[0595] The server immediately generates an appropriate response based on the search results. The generated response is then formatted into text using a generative AI model. At this stage, the answer is output in text format.

[0596] Step 6:

[0597] The server sends the generated text-formatted response back to the terminal and displays it in the user interface. The terminal receives this information and outputs it on the screen. At this stage, the salesperson can review the displayed information and communicate it to the customer.

[0598] Step 7:

[0599] Salespeople, acting as users, input customer feedback and their own evaluations into a terminal. This feedback is sent back to the server and used to optimize and improve the system. The input feedback becomes output that contributes to improving the system's future performance.

[0600] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0601] As an embodiment of this invention, a system is constructed that incorporates an emotion engine that processes voice and text data and supports customer service. A specific embodiment thereof is described below.

[0602] System Overview

[0603] terminal

[0604] Store staff assisting customers use a terminal to input customer questions via voice or text. The terminal converts the voice data into text, making it analyzable.

[0605] The device also features an emotion engine that analyzes emotions from voice and text, recognizing customer emotions in real time.

[0606] server

[0607] The server receives voice and text data sent from the terminal and analyzes the content using a natural language processing engine. This extracts the intent behind the customer's question and relevant information.

[0608] Emotional information from the emotion engine is also sent to the server, which then generates responses with appropriate tone and content. Each question is addressed individually based on the results of the emotion analysis.

[0609] User interface and feedback

[0610] The responses generated by the server are displayed on the user interface on the terminal. The displayed information is provided as the most effective means of communication for the user, based on sentiment analysis.

[0611] By following the guidelines provided by the emotion engine when providing responses to customers, users can communicate more effectively. Furthermore, users can continuously improve the quality of their responses by incorporating the feedback they receive into the system.

[0612] Specific example

[0613] Example 1: Responding to customer complaints

[0614] If a customer expresses dissatisfaction and asks for an explanation about a product defect, the user will input their voice message into the device.

[0615] The device's emotion engine recognizes customer dissatisfaction from their voice and sends that information to the server. The server extracts information related to the problem and generates a response in a reassuring tone.

[0616] An explanation such as "We apologize, regarding this issue..." will be displayed on the device, and the user will respond in good faith based on that information.

[0617] Example 2: High-interest inquiries about new products

[0618] If a customer expresses interest and asks, "I want to know more about this new product," the user will input their intention into the terminal.

[0619] The emotion engine recognizes the customer's level of interest, and the server uses that information to provide detailed information about relevant products in an enthusiastic and engaging manner.

[0620] Based on the displayed information, users can communicate the appeal of the new product in a more personalized way.

[0621] In this way, by integrating an emotion engine, this system enables communication that adapts to customer emotions and strengthens support for store staff. Staff will be able to provide highly accurate customer service, contributing to improved customer satisfaction.

[0622] The following describes the processing flow.

[0623] Step 1:

[0624] The user inputs a question by voice into the device. The device uses speech recognition to convert the voice data into text. The text data is temporarily stored.

[0625] Step 2:

[0626] The device's emotion engine analyzes the user's emotions from the voice data. The analysis results are tagged with an emotional state (e.g., dissatisfied, excited, calm).

[0627] Step 3:

[0628] The device sends text and sentiment data to the server. The data is encrypted and transferred using a secure protocol.

[0629] Step 4:

[0630] The server receives text data and analyzes its content using a natural language processing engine. Based on the keywords and phrases extracted through the analysis, it searches a database of past cases.

[0631] Step 5:

[0632] The server takes emotional data into account and adjusts the tone and content of responses accordingly. Depending on the emotion, it selects reassuring tones or encouraging words.

[0633] Step 6:

[0634] The server generates the final response and sends it to the terminal along with an emotion-based communication strategy.

[0635] Step 7:

[0636] The device displays the received response on the user interface. The displayed content is visually organized and in a format that is easy for the user to understand.

[0637] Step 8:

[0638] Users review the information displayed on their devices and use it in their interactions with customers. Feedback received during customer interactions is sent to the server via the device and stored as learning data for future interactions.

[0639] (Example 2)

[0640] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0641] In customer service, rapidly and accurately analyzing voice and text data and providing appropriate responses that reflect customer emotions is a challenge that conventional technologies have not adequately addressed. In particular, achieving this in real time while continuously learning and improving the system is difficult. Therefore, a system is needed that combines more advanced natural language processing and sentiment analysis technologies to provide effective responses to users.

[0642] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0643] In this invention, the server includes means for receiving voice or text data and analyzing it using natural language processing technology, means for analyzing emotional information using emotion analysis technology based on the analyzed information, and means for generating an optimal response in real time using a generative artificial intelligence model based on the analysis results and displaying it on an information display device. This makes it possible to provide appropriate and personalized responses that correspond to the customer's emotions and to continuously improve the system's performance.

[0644] "Voice or text data" refers to voice or text information that users input into the system.

[0645] "Natural language processing technology" refers to the technology that enables computers to understand, analyze, and generate human language.

[0646] "Emotional analysis technology" refers to the technology that extracts and analyzes emotional information from voice and text data.

[0647] A "generative artificial intelligence model" refers to a model that uses machine learning algorithms to generate data-driven results or answers.

[0648] An "information display device" refers to a device or software that visually displays information on a user interface.

[0649] "Feedback" refers to evaluations and opinions provided by users to improve the system's capabilities.

[0650] To implement this invention, a system is constructed that processes voice and text data to support customer service. This system utilizes emotion analysis technology to enable personalized communication.

[0651] The terminal is used by store staff to input customer questions via voice or text. The voice data is converted to text using speech recognition software. Specifically, common speech recognition software and services are used, such as open-source speech recognition libraries and cloud-based speech recognition services. The terminal is equipped with software that incorporates sentiment analysis technology to instantly analyze the customer's emotions and generate an emotion score.

[0652] The server receives voice or text data and sentiment scores transmitted from the terminal. The server analyzes the text data using natural language processing techniques and further evaluates the customer's emotional state based on sentiment analysis techniques. This process utilizes a generative AI model to automatically generate the optimal response from the analyzed information. The generated response is then customized with appropriate tone and content based on the sentiment analysis results.

[0653] This response is immediately displayed on the information display on the terminal. The user uses this information to take appropriate action with the customer. For example, if a customer is dissatisfied with a product, the server will generate a response that includes an apology and a solution to alleviate the dissatisfaction. As a concrete example, an example of a prompt message when a customer makes a complaint is shown below.

[0654] Example prompt: "Please provide more details about the defect in this product."

[0655] In this way, the system provides responses that are appropriate to the customer's emotions, supporting better customer service.

[0656] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0657] Step 1:

[0658] The device receives voice or text data entered by the user. In the case of voice input, speech recognition technology is used to convert the voice into text. In this process, it receives voice or text data as input and generates text data as output. Specifically, it collects voice data in real time, and the speech recognition engine converts it into text format.

[0659] Step 2:

[0660] The device sends the received text data to an emotion analysis engine to generate an emotion score. In this process, text data is provided as input to the emotion analysis model, and an emotion score is obtained as output. Specifically, it analyzes the wording and expressions within the text data to identify the emotional state.

[0661] Step 3:

[0662] The server receives text data and sentiment scores sent from the terminal. Here, the input received is text data and sentiment scores. The server uses natural language processing techniques to analyze the content of the text data and extract customer intent and related information. The output generates data including the analysis results and intent. In this step, natural language processing, including keyword extraction and intent understanding, is performed using the input data, and based on this, the customer's requests are clarified.

[0663] Step 4:

[0664] The server provides the generative AI model with analysis results and sentiment scores to generate the optimal response. The generative AI model receives analyzed intent data and sentiment scores as input and generates a response to convey to the customer as output. Specifically, it produces grammatically correct responses and adjusts the message to an emotionally sensitive tone.

[0665] Step 5:

[0666] The terminal displays the responses generated from the server on the user interface. Here, it receives response data from the server as input and generates text or audio data to be displayed on the screen as output. As a concrete example of its operation, the user interface visually displays the sentiment score while clearly arranging the response messages.

[0667] This process allows users to respond to customers based on the generated answers, leading to more effective communication.

[0668] (Application Example 2)

[0669] Next, we will explain Application Example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0670] Traditional customer service systems have faced the challenge of not being able to adequately consider customer emotions when providing service. Store staff, when interacting with customers through voice and text data, lack the skills to properly understand the customer's emotional state and communicate appropriately according to that state, which hinders improvements in customer satisfaction.

[0671] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0672] In this invention, the server includes means for receiving voice or text data and analyzing it using natural language processing; means for searching for information from a database of relevant past cases based on the analyzed information and sentiment analysis results from an emotion engine; and means for generating a response in the optimal tone and content in real time based on the search results and sentiment analysis results, and providing it to a display device. This makes it possible to improve customer satisfaction by providing appropriate responses that are in line with the customer's emotions.

[0673] "Audio or text data" refers to audio and text information produced by customers, which is used for communication with customers.

[0674] "Natural language processing" is a technology that analyzes speech or text data to understand its meaning and intent.

[0675] An "emotion engine" is a technology that analyzes and quantifies or categorizes a customer's emotional state from voice and text data.

[0676] A "database of related past cases" is a collection of data that accumulates past inquiries and responses, providing useful information for similar situations.

[0677] "Optimal tone and content" refers to appropriate expressions and information that are tailored to the customer's emotions and situation, thereby enhancing the effectiveness of communication.

[0678] "User feedback" refers to information about the results and effectiveness of customer service, and is used to improve the system.

[0679] A "display device" is a device or system that provides generated information to a user visually.

[0680] To implement this invention, the system provides a technology that uses voice and text data to enable communication tailored to the customer's emotions. Specific embodiments are described below.

[0681] The server uses a speech recognition API to convert the audio data received from the terminal into text data. This converted text is then analyzed by a natural language processing engine (e.g., Dialogflow) to extract the customer's intent and the content of their question. Furthermore, an emotion engine analyzes the customer's emotional state based on this text data and sends the results to the server.

[0682] The terminal is responsible for interacting with customers at the store level. The terminal receives voice input, converts the voice into text appropriately, and sends it to the server. The analysis results and sentiment analysis results are processed on the server to generate responses with the optimal tone and content. This information is displayed in real time on the terminal's display device, allowing users to refer to it and respond appropriately to the customer's emotions.

[0683] Users can provide highly personalized responses based on emotional information derived from the customer's tone of voice and facial expressions. Because these responses are based on automatically generated guidelines, even average operators can provide high-quality customer service. Furthermore, user feedback is collected on the server and used for system optimization and further learning.

[0684] As a concrete example, if a customer is dissatisfied and requests a detailed explanation of a product defect, entering this into the terminal will cause the emotion engine to quantify the dissatisfaction, and the server will automatically generate a response in a reassuring tone. At this point, the server will return a tailored message such as, "We apologize, regarding this issue..." to communicate appropriately with the customer.

[0685] An example of a prompt is: "When a customer asks, 'Please tell me more about the new product,' and the emotion engine indicates a high level of excitement, respond appropriately using relevant information." In this way, it becomes possible to respond to a variety of customer emotional states.

[0686] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0687] Step 1:

[0688] The terminal receives voice input from the customer and converts that voice into text data using a speech recognition API. The input is the customer's voice, and the output is a sentence in text format. This conversion process enables subsequent natural language processing.

[0689] Step 2:

[0690] The server uses a natural language processing engine to analyze the text data received from the terminal. The input is the text data from step 1, and the output is information about the customer's intent and the content of their question. This process helps to understand what the customer is looking for.

[0691] Step 3:

[0692] The server uses the analyzed information and an emotion engine to analyze the customer's emotional state. The input is the text information from step 2, and the output is numerical or categorical data indicating the emotional state. This enables responses tailored to the customer's emotions.

[0693] Step 4:

[0694] The server searches a database of past cases for relevant information based on the sentiment analysis results and intent analysis results. The input is the analysis results from steps 2 and 3, and the output is information based on similar past cases. This allows for the acquisition of appropriate reference information.

[0695] Step 5:

[0696] The server generates a response with the optimal tone and content based on the previous output. The input is the information from step 4 and the emotional state from step 3, and the output is the response text provided to the customer. This process enables responses that are appropriate to the customer's emotions.

[0697] Step 6:

[0698] The terminal displays the generated response on a display device, and the user responds to the customer based on it. The input is the response text generated in step 5, and the output is the user's response to the customer. In this final step, the user can efficiently and effectively meet the customer's needs.

[0699] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0700] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0701] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0702] [Fourth Embodiment]

[0703] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0704] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0705] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0706] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0707] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0708] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0709] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0710] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0711] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0712] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0713] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0714] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0715] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0716] One embodiment of this invention involves constructing a system that processes voice and text data to support customer service. A specific embodiment is described below.

[0717] System Overview

[0718] terminal

[0719] Store staff handling customer inquiries use a terminal to input customer questions via voice or text. The terminal converts the voice data into text in real time, preparing it for natural language processing.

[0720] The entered data is sent to the server in an encrypted form, ensuring secure communication at all times.

[0721] server

[0722] The server receives data sent from the terminal and performs analysis using a natural language processing engine. This analysis identifies key keywords and intents contained in the customer's question.

[0723] Based on the identified information, the system searches past case databases and FAQ resources to extract the most relevant solutions.

[0724] User interface and feedback

[0725] The responses generated by the server are displayed on the terminal's screen in real time. The user interface is intuitive and used for direct explanations to customers.

[0726] Staff can review the displayed information, provide customer support, and offer feedback on the accuracy and usefulness of the answers.

[0727] Specific example

[0728] Example 1: Inquiry about campaign information

[0729] When a customer asks, "Tell me about the current campaign," the user inputs their inquiry into their device using voice input.

[0730] The device converts the audio to text and sends it to the server. The server searches its database for cases related to the "campaign" based on the analyzed results and extracts the current campaign information.

[0731] This information is quickly forwarded, and the terminal screen displays "We are currently running campaigns for Plan A and Plan B." The user sees this and responds to the customer immediately.

[0732] Example 2: Questions about the new plan

[0733] If another customer requests "Please tell me the details of the new plan," the user will enter that intention into their device.

[0734] The server retrieves the latest information about the new plan from the database and generates the most appropriate recommended plan information.

[0735] This information is displayed on the device and provided to the staff in the form of "New Plan X has the following features."

[0736] Thus, this system aims to improve the quality of customer service and optimize operational efficiency. By coordinating servers and terminals to provide users with intuitive and rapid advice, it reduces the burden on stores.

[0737] The following describes the processing flow.

[0738] Step 1:

[0739] The user inputs a question by voice into the device. The device uses speech recognition to convert the voice data into text data in real time.

[0740] Step 2:

[0741] The device temporarily stores the converted text data, encrypts it, and prepares it for transmission to the server. Secure protocols are used to ensure the data's safety.

[0742] Step 3:

[0743] The server analyzes the text data received from the terminal. This analysis uses a natural language processing engine to extract important keywords and customer intent from the input questions.

[0744] Step 4:

[0745] The server searches past case databases and FAQs based on the analysis results. It quickly extracts relevant cases and information and generates multiple candidate answers.

[0746] Step 5:

[0747] An algorithm is applied to select the most relevant and appropriate answer from multiple answer candidates generated by the server. The selected answer is then formatted appropriately and sent to the terminal.

[0748] Step 6:

[0749] The device displays the responses it receives on the user interface in real time. The device organizes the information and presents it visually to the user, making it easier to understand.

[0750] Step 7:

[0751] The user responds to the customer based on the answers displayed on the device. If necessary, additional questions can be asked, and that data can also be entered and processed on the device.

[0752] Step 8:

[0753] When a user provides feedback to the system, the terminal sends that feedback to the server. The server records this feedback and uses it as data to improve the system's accuracy.

[0754] (Example 1)

[0755] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0756] Current customer support systems struggle to process large amounts of data quickly and accurately, making it difficult to provide immediate answers to diverse customer inquiries. Furthermore, there is a need to ensure the security of acquired data, improve the accuracy of responses, and effectively utilize user feedback. Additionally, there is a lack of methods for personalizing responses for individual users and leveraging other information sources when analysis results are insufficient.

[0757] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0758] In this invention, the server includes means at a terminal for acquiring voice or text information, software means for converting the acquired information into structured data, encrypted communication means for securely communicating the acquired data, a natural language processing device for analyzing the received data, information retrieval means for searching a relevant information base based on the analysis results, user-facing display means for generating and displaying optimized responses, and feedback means for collecting user reviews to contribute to the system's adaptive learning. This enables integrated and secure automated customer service.

[0759] A "terminal for acquiring voice or text information" is an information input device that receives inquiries from customers in voice or text format and initially acquires the data necessary for subsequent processing.

[0760] "Software for converting to structured data" refers to a program that analyzes acquired audio or text information and converts it into a data format that can be processed using natural language processing.

[0761] "Encrypted communication methods" refer to communication methods that include technology for encrypting information to securely transmit data and decrypting it at the destination.

[0762] A "natural language processing device" is a software or hardware system that analyzes input text data to perform semantic understanding and intent extraction.

[0763] An "information retrieval tool" is an algorithm or software used to search a database for relevant information based on analyzed intents and keywords.

[0764] "User-facing display means" refers to an interface device or program for displaying answers generated based on search results in a format that is easy for the user to view.

[0765] A "feedback mechanism" is a system for collecting evaluations and opinions on answers provided by users and using them to improve the system and for learning.

[0766] This invention relates to a system that supports customer service using voice and text data. The system is designed to enable users to efficiently process customer inquiries. Specific embodiments are described below.

[0767] System Configuration

[0768] terminal

[0769] The user uses a terminal to input customer inquiries via voice or text form. This terminal uses speech recognition software (e.g., a speech recognition engine) to convert the voice data into text data. Data transfer is conducted via encrypted communication protocols (e.g., TLS / SSL) to ensure the security of the information.

[0770] server

[0771] The server analyzes the received data using natural language processing software (e.g., a natural language processing engine). This analysis identifies the customer's intent, and a database management system (e.g., an SQL database) is used to retrieve relevant information from the database. This then generates the most appropriate response.

[0772] User interface and feedback features

[0773] The server sends the generated responses to the terminal, allowing the user to explain the content to the customer in real time. The user interface is designed with ease of use in mind and is intuitive to operate. Users can provide feedback on the accuracy and usefulness of the provided responses, contributing to the continuous improvement of the system.

[0774] Specific example

[0775] Campaign information inquiry

[0776] When a customer asks, "Tell me about the current campaign," the user inputs the question by voice into the device. The device converts the voice into text and sends the data to the server. The server searches for data related to "campaign" and generates information such as, "We are currently running campaigns Plan A and Plan B." By displaying this information on the device, the user can immediately answer the customer's question.

[0777] Questions about the details of the new plan

[0778] In another scenario, if a customer asks, "Tell me the details of the new plan," the user enters the request into their device. The server retrieves the latest data related to the "new plan" in the most appropriate format and provides the user with information such as, "New Plan X has the following features."

[0779] In this example, a generative AI model is used to optimize prompt text, enabling specific and detailed questions such as "Please provide current campaign information." This system allows users to respond to customer needs quickly and accurately.

[0780] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0781] Step 1:

[0782] The user inputs a question from the customer via voice or text through the terminal. The terminal uses speech recognition software to convert the voice data into text data. Specifically, if the voice input is "I want to know the details of the new plan," the terminal converts this to text "I want to know the details of the new plan" and outputs the converted text.

[0783] Step 2:

[0784] The terminal encrypts the converted text data. For example, AES encryption technology is used. The encrypted data is then sent to the server while maintaining security. The input is the text data "I want to know more about the new plan," and the output is the encrypted text data.

[0785] Step 3:

[0786] The server receives encrypted data from the terminal and performs a decryption process. As a result, the original text data, "I want to know the details of the new plan," becomes the input data on the server side. The server analyzes this text using a natural language processing engine and extracts keywords and intent. For example, "new plan" is extracted as a keyword.

[0787] Step 4:

[0788] The server uses a database management system based on the extracted keywords to search for relevant information within the database. The search results obtained are information about the new plan. For example, "Price and features of new plan X" are obtained.

[0789] Step 5:

[0790] The server generates an answer based on the search results. A generative AI model is used here, which generates more natural sentences by using prompts. If the data input is "Pricing and features of the new plan X," the output information will be "The new plan X costs $100 per month and includes unlimited data usage."

[0791] Step 6:

[0792] The server sends the generated response to the terminal. The terminal displays the received information on its screen. The user reviews the details displayed on the screen and explains them to the customer. Based on the confirmed information, the user can quickly respond to the customer.

[0793] Step 7:

[0794] Users provide feedback on the accuracy and usefulness of the answers provided. This feedback is used by the system's adaptive learning algorithm to improve the quality of future answers. An example of feedback is "The information was very helpful." This helps the system strive to provide even more accurate information.

[0795] (Application Example 1)

[0796] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0797] In physical stores, a challenge exists in that sales staff often struggle to answer customer questions quickly and accurately. This challenge stems from the inability to instantly provide diverse product information, inventory status, and campaign information, which reduces the efficiency of customer service. Therefore, a system is needed that allows sales staff to obtain and provide necessary information to customers in real time.

[0798] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0799] In this invention, the server includes means for receiving voice or text information and analyzing it using natural language processing technology; means for searching for information from a storage device that stores relevant past cases based on the analyzed information; means for immediately generating an optimal response based on the search results and displaying it on the user interface; and means for converting voice input into text and transmitting it to a cloud processing device. This enables salespeople to quickly and accurately obtain the information requested by customers and provide it in real time.

[0800] "Audio or text information" refers to audio or text data received from the user.

[0801] "Natural language processing technology" is a technology that analyzes speech or text information to understand its meaning and extract necessary data.

[0802] A "storage device" refers to a device such as a database used to store past cases and related information.

[0803] A "user interface" refers to a screen or device that displays system output and allows the user to view the information.

[0804] "Voice input" refers to a method of acquiring the content of a user's speech as digital data.

[0805] A "cloud processing device" refers to a remote server or computer device used to process voice input and analyzed information.

[0806] One embodiment of this invention is based on a system designed to facilitate customer service for sales staff in physical stores. The system receives voice and text information and analyzes it using natural language processing technology. This helps sales staff to respond quickly and accurately to various questions from customers.

[0807] The server instantly converts speech data into text using the Google Cloud Speech-to-Text API. The converted text is sent to the cloud server and analyzed using spaCy, a natural language processing technology. Based on the analysis results and relevant information stored in a MySQL database, the optimal response is instantly generated and displayed on the salesperson's terminal via the user interface.

[0808] The terminal uses a smartphone as its hardware and is a device for sales staff to input voice commands. This terminal is responsible not only for receiving and displaying information but also for sending user feedback to a server.

[0809] For example, if a customer asks a store clerk, "Is there a discount on this product?", the system will process this voice question immediately and instantly display relevant campaign and discount information on the terminal.

[0810] An example of a prompt message used as an example of the use of the generating AI model is: "When a customer asks a question about a specific product, retrieve the product's features, stock status, and campaign information from the server in real time and display it on the smartphone."

[0811] This configuration allows sales staff to interact with customers more efficiently, resulting in a system that contributes to improved customer satisfaction.

[0812] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0813] Step 1:

[0814] The user uses their smartphone (device) to input customer questions by voice. This input voice is captured by the device as audio data. At this stage, the input is in the form of audio data.

[0815] Step 2:

[0816] The device converts the audio data to text using the Google Cloud Speech-to-Text API. This API analyzes the audio data and generates the corresponding text data. The output of this step is text data.

[0817] Step 3:

[0818] Text data is sent from the terminal to the cloud server. The server receives this text data and analyzes it using spaCy, a natural language processing technology. The analysis involves data processing to understand the meaning of the text and identify the customer's intent. The output of this step is the analyzed intent and keywords.

[0819] Step 4:

[0820] The server searches for relevant information from the MySQL database based on the analyzed intent and keywords. The retrieved information is extracted from relevant cases, FAQs, etc. The input is the analysis result, and the output is the best answer found.

[0821] Step 5:

[0822] The server immediately generates an appropriate response based on the search results. The generated response is then formatted into text using a generative AI model. At this stage, the answer is output in text format.

[0823] Step 6:

[0824] The server sends the generated text-formatted response back to the terminal and displays it in the user interface. The terminal receives this information and outputs it on the screen. At this stage, the salesperson can review the displayed information and communicate it to the customer.

[0825] Step 7:

[0826] Salespeople, acting as users, input customer feedback and their own evaluations into a terminal. This feedback is sent back to the server and used to optimize and improve the system. The input feedback becomes output that contributes to improving the system's future performance.

[0827] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0828] As an embodiment of this invention, a system is constructed that incorporates an emotion engine that processes voice and text data and supports customer service. A specific embodiment thereof is described below.

[0829] System Overview

[0830] terminal

[0831] Store staff assisting customers use a terminal to input customer questions via voice or text. The terminal converts the voice data into text, making it analyzable.

[0832] The device also features an emotion engine that analyzes emotions from voice and text, recognizing customer emotions in real time.

[0833] server

[0834] The server receives voice and text data sent from the terminal and analyzes the content using a natural language processing engine. This extracts the intent behind the customer's question and relevant information.

[0835] Emotional information from the emotion engine is also sent to the server, which then generates responses with appropriate tone and content. Each question is addressed individually based on the results of the emotion analysis.

[0836] User interface and feedback

[0837] The responses generated by the server are displayed on the user interface on the terminal. The displayed information is provided as the most effective means of communication for the user, based on sentiment analysis.

[0838] By following the guidelines provided by the emotion engine when providing responses to customers, users can communicate more effectively. Furthermore, users can continuously improve the quality of their responses by incorporating the feedback they receive into the system.

[0839] Specific example

[0840] Example 1: Responding to customer complaints

[0841] If a customer expresses dissatisfaction and asks for an explanation about a product defect, the user will input their voice message into the device.

[0842] The device's emotion engine recognizes customer dissatisfaction from their voice and sends that information to the server. The server extracts information related to the problem and generates a response in a reassuring tone.

[0843] An explanation such as "We apologize, regarding this issue..." will be displayed on the device, and the user will respond in good faith based on that information.

[0844] Example 2: High-interest inquiries about new products

[0845] If a customer expresses interest and asks, "I want to know more about this new product," the user will input their intention into the terminal.

[0846] The emotion engine recognizes the customer's level of interest, and the server uses that information to provide detailed information about relevant products in an enthusiastic and engaging manner.

[0847] Based on the displayed information, users can communicate the appeal of the new product in a more personalized way.

[0848] In this way, by integrating an emotion engine, this system enables communication that adapts to customer emotions and strengthens support for store staff. Staff will be able to provide highly accurate customer service, contributing to improved customer satisfaction.

[0849] The following describes the processing flow.

[0850] Step 1:

[0851] The user inputs a question by voice into the device. The device uses speech recognition to convert the voice data into text. The text data is temporarily stored.

[0852] Step 2:

[0853] The device's emotion engine analyzes the user's emotions from the voice data. The analysis results are tagged with an emotional state (e.g., dissatisfied, excited, calm).

[0854] Step 3:

[0855] The device sends text and sentiment data to the server. The data is encrypted and transferred using a secure protocol.

[0856] Step 4:

[0857] The server receives text data and analyzes its content using a natural language processing engine. Based on the keywords and phrases extracted through the analysis, it searches a database of past cases.

[0858] Step 5:

[0859] The server takes emotional data into account and adjusts the tone and content of responses accordingly. Depending on the emotion, it selects reassuring tones or encouraging words.

[0860] Step 6:

[0861] The server generates the final response and sends it to the terminal along with an emotion-based communication strategy.

[0862] Step 7:

[0863] The device displays the received response on the user interface. The displayed content is visually organized and in a format that is easy for the user to understand.

[0864] Step 8:

[0865] Users review the information displayed on their devices and use it in their interactions with customers. Feedback received during customer interactions is sent to the server via the device and stored as learning data for future interactions.

[0866] (Example 2)

[0867] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0868] In customer service, rapidly and accurately analyzing voice and text data and providing appropriate responses that reflect customer emotions is a challenge that conventional technologies have not adequately addressed. In particular, achieving this in real time while continuously learning and improving the system is difficult. Therefore, a system is needed that combines more advanced natural language processing and sentiment analysis technologies to provide effective responses to users.

[0869] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0870] In this invention, the server includes means for receiving voice or text data and analyzing it using natural language processing technology, means for analyzing emotional information using emotion analysis technology based on the analyzed information, and means for generating an optimal response in real time using a generative artificial intelligence model based on the analysis results and displaying it on an information display device. This makes it possible to provide appropriate and personalized responses that correspond to the customer's emotions and to continuously improve the system's performance.

[0871] "Voice or text data" refers to voice or text information that users input into the system.

[0872] "Natural language processing technology" refers to the technology that enables computers to understand, analyze, and generate human language.

[0873] "Emotional analysis technology" refers to the technology that extracts and analyzes emotional information from voice and text data.

[0874] A "generative artificial intelligence model" refers to a model that uses machine learning algorithms to generate data-driven results or answers.

[0875] An "information display device" refers to a device or software that visually displays information on a user interface.

[0876] "Feedback" refers to evaluations and opinions provided by users to improve the system's capabilities.

[0877] To implement this invention, a system is constructed that processes voice and text data to support customer service. This system utilizes emotion analysis technology to enable personalized communication.

[0878] The terminal is used by store staff to input customer questions via voice or text. The voice data is converted to text using speech recognition software. Specifically, common speech recognition software and services are used, such as open-source speech recognition libraries and cloud-based speech recognition services. The terminal is equipped with software that incorporates sentiment analysis technology to instantly analyze the customer's emotions and generate an emotion score.

[0879] The server receives voice or text data and sentiment scores transmitted from the terminal. The server analyzes the text data using natural language processing techniques and further evaluates the customer's emotional state based on sentiment analysis techniques. This process utilizes a generative AI model to automatically generate the optimal response from the analyzed information. The generated response is then customized with appropriate tone and content based on the sentiment analysis results.

[0880] This response is immediately displayed on the information display on the terminal. The user uses this information to take appropriate action with the customer. For example, if a customer is dissatisfied with a product, the server will generate a response that includes an apology and a solution to alleviate the dissatisfaction. As a concrete example, an example of a prompt message when a customer makes a complaint is shown below.

[0881] Example prompt: "Please provide more details about the defect in this product."

[0882] In this way, the system provides responses that are appropriate to the customer's emotions, supporting better customer service.

[0883] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0884] Step 1:

[0885] The device receives voice or text data entered by the user. In the case of voice input, speech recognition technology is used to convert the voice into text. In this process, it receives voice or text data as input and generates text data as output. Specifically, it collects voice data in real time, and the speech recognition engine converts it into text format.

[0886] Step 2:

[0887] The device sends the received text data to an emotion analysis engine to generate an emotion score. In this process, text data is provided as input to the emotion analysis model, and an emotion score is obtained as output. Specifically, it analyzes the wording and expressions within the text data to identify the emotional state.

[0888] Step 3:

[0889] The server receives text data and sentiment scores sent from the terminal. Here, the input received is text data and sentiment scores. The server uses natural language processing techniques to analyze the content of the text data and extract customer intent and related information. The output generates data including the analysis results and intent. In this step, natural language processing, including keyword extraction and intent understanding, is performed using the input data, and based on this, the customer's requests are clarified.

[0890] Step 4:

[0891] The server provides the generative AI model with analysis results and sentiment scores to generate the optimal response. The generative AI model receives analyzed intent data and sentiment scores as input and generates a response to convey to the customer as output. Specifically, it produces grammatically correct responses and adjusts the message to an emotionally sensitive tone.

[0892] Step 5:

[0893] The terminal displays the responses generated from the server on the user interface. Here, it receives response data from the server as input and generates text or audio data to be displayed on the screen as output. As a concrete example of its operation, the user interface visually displays the sentiment score while clearly arranging the response messages.

[0894] This process allows users to respond to customers based on the generated answers, leading to more effective communication.

[0895] (Application Example 2)

[0896] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0897] Traditional customer service systems have faced the challenge of not being able to adequately consider customer emotions when providing service. Store staff, when interacting with customers through voice and text data, lack the skills to properly understand the customer's emotional state and communicate appropriately according to that state, which hinders improvements in customer satisfaction.

[0898] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0899] In this invention, the server includes means for receiving voice or text data and analyzing it using natural language processing; means for searching for information from a database of relevant past cases based on the analyzed information and sentiment analysis results from an emotion engine; and means for generating a response in the optimal tone and content in real time based on the search results and sentiment analysis results, and providing it to a display device. This makes it possible to improve customer satisfaction by providing appropriate responses that are in line with the customer's emotions.

[0900] "Audio or text data" refers to audio and text information produced by customers, which is used for communication with customers.

[0901] "Natural language processing" is a technology that analyzes speech or text data to understand its meaning and intent.

[0902] An "emotion engine" is a technology that analyzes and quantifies or categorizes a customer's emotional state from voice and text data.

[0903] A "database of related past cases" is a collection of data that accumulates past inquiries and responses, providing useful information for similar situations.

[0904] "Optimal tone and content" refers to appropriate expressions and information that are tailored to the customer's emotions and situation, thereby enhancing the effectiveness of communication.

[0905] "User feedback" refers to information about the results and effectiveness of customer service, and is used to improve the system.

[0906] A "display device" is a device or system that provides generated information to a user visually.

[0907] To implement this invention, the system provides a technology that uses voice and text data to enable communication tailored to the customer's emotions. Specific embodiments are described below.

[0908] The server uses a speech recognition API to convert the audio data received from the terminal into text data. This converted text is then analyzed by a natural language processing engine (e.g., Dialogflow) to extract the customer's intent and the content of their question. Furthermore, an emotion engine analyzes the customer's emotional state based on this text data and sends the results to the server.

[0909] The terminal is responsible for interacting with customers at the store level. The terminal receives voice input, converts the voice into text appropriately, and sends it to the server. The analysis results and sentiment analysis results are processed on the server to generate responses with the optimal tone and content. This information is displayed in real time on the terminal's display device, allowing users to refer to it and respond appropriately to the customer's emotions.

[0910] Users can provide highly personalized responses based on emotional information derived from the customer's tone of voice and facial expressions. Because these responses are based on automatically generated guidelines, even average operators can provide high-quality customer service. Furthermore, user feedback is collected on the server and used for system optimization and further learning.

[0911] As a concrete example, if a customer is dissatisfied and requests a detailed explanation of a product defect, entering this into the terminal will cause the emotion engine to quantify the dissatisfaction, and the server will automatically generate a response in a reassuring tone. At this point, the server will return a tailored message such as, "We apologize, regarding this issue..." to communicate appropriately with the customer.

[0912] An example of a prompt is: "When a customer asks, 'Please tell me more about the new product,' and the emotion engine indicates a high level of excitement, respond appropriately using relevant information." In this way, it becomes possible to respond to a variety of customer emotional states.

[0913] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0914] Step 1:

[0915] The terminal receives voice input from the customer and converts that voice into text data using a speech recognition API. The input is the customer's voice, and the output is a sentence in text format. This conversion process enables subsequent natural language processing.

[0916] Step 2:

[0917] The server uses a natural language processing engine to analyze the text data received from the terminal. The input is the text data from step 1, and the output is information about the customer's intent and the content of their question. This process helps to understand what the customer is looking for.

[0918] Step 3:

[0919] The server uses the analyzed information and an emotion engine to analyze the customer's emotional state. The input is the text information from step 2, and the output is numerical or categorical data indicating the emotional state. This enables responses tailored to the customer's emotions.

[0920] Step 4:

[0921] The server searches a database of past cases for relevant information based on the sentiment analysis results and intent analysis results. The input is the analysis results from steps 2 and 3, and the output is information based on similar past cases. This allows for the acquisition of appropriate reference information.

[0922] Step 5:

[0923] The server generates a response with the optimal tone and content based on the previous output. The input is the information from step 4 and the emotional state from step 3, and the output is the response text provided to the customer. This process enables responses that are appropriate to the customer's emotions.

[0924] Step 6:

[0925] The terminal displays the generated response on a display device, and the user responds to the customer based on it. The input is the response text generated in step 5, and the output is the user's response to the customer. In this final step, the user can efficiently and effectively meet the customer's needs.

[0926] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0927] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0928] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0929] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0930] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0931] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0932] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0933] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0934] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0935] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0936] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0937] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0938] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0939] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0940] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0941] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0942] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0943] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0944] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0945] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0946] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0947] The following is further disclosed regarding the embodiments described above.

[0948] (Claim 1)

[0949] A means for receiving audio or text data and analyzing it using natural language processing,

[0950] A means of searching for information from a database of related past cases based on the analyzed information,

[0951] A means of generating the optimal answer in real time based on search results and displaying it in the user interface,

[0952] A means of collecting user feedback and using it to learn from and improve the system,

[0953] A system that includes this.

[0954] (Claim 2)

[0955] The system according to claim 1, which uses user-specific information to provide personalized responses in receiving and analyzing the aforementioned voice or text data.

[0956] (Claim 3)

[0957] The system according to claim 1, further comprising means for contacting the relevant department and providing feedback on the results if the analyzed information is insufficient.

[0958] "Example 1"

[0959] (Claim 1)

[0960] Means of acquiring voice or text information on a terminal,

[0961] Software means for converting acquired information into structured data,

[0962] An encrypted communication method for securely transmitting the converted data,

[0963] A natural language processing unit for analyzing received data,

[0964] An information retrieval means for searching a related information base based on the analysis results,

[0965] A user-facing display means for generating and displaying optimized answers,

[0966] A feedback mechanism that collects user reviews and contributes to the adaptive learning of the system,

[0967] A system that includes this.

[0968] (Claim 2)

[0969] The system according to claim 1, which uses information acquired by the terminal to output personalized information.

[0970] (Claim 3)

[0971] The system according to claim 1, which, when the aforementioned analysis results are limited, obtains data from other sources and provides feedback based on that information.

[0972] "Application Example 1"

[0973] (Claim 1)

[0974] A means for receiving audio or text information and analyzing it using natural language processing technology,

[0975] A means for retrieving information from a storage device that stores related past cases based on the analyzed information,

[0976] A means for instantly generating the optimal response based on search results and displaying it in the user interface,

[0977] A means of collecting user feedback and using it to optimize the system,

[0978] A means of converting voice input into text and sending it to a cloud-based processing device,

[0979] A system that includes this.

[0980] (Claim 2)

[0981] The system according to claim 1, which uses user-specific data to provide a personalized response in receiving and analyzing the aforementioned voice or text information.

[0982] (Claim 3)

[0983] The system according to claim 1, further comprising means for querying relevant categories and returning the results as an evaluation if the analyzed information is insufficient.

[0984] "Example 2 of combining an emotion engine"

[0985] (Claim 1)

[0986] A means for receiving audio or text data and analyzing it using natural language processing technology,

[0987] A means of analyzing emotional information using emotion analysis technology based on the analyzed information,

[0988] A means for generating the optimal answer in real time using an artificial intelligence model based on the analysis results and displaying it on an information display device,

[0989] A means of collecting user feedback and using it for learning and improving information systems,

[0990] A system that includes this.

[0991] (Claim 2)

[0992] The system according to claim 1, which, in receiving and analyzing the aforementioned voice or text data, uses user-specific information and provides personalized responses using sentiment analysis results.

[0993] (Claim 3)

[0994] The system according to claim 1, further comprising means for contacting relevant departments and providing feedback on the results if the analyzed information and sentiment analysis information are insufficient.

[0995] "Application example 2 when combining with an emotional engine"

[0996] (Claim 1)

[0997] A means for receiving audio or text data and analyzing it using natural language processing,

[0998] A means of searching for information from a database of relevant past cases based on the analyzed information and the results of sentiment analysis by the sentiment engine,

[0999] A means for generating and providing a response in the optimal tone and content in real time, based on search results and sentiment analysis results, to a display device,

[1000] A means of collecting user feedback and using it for learning and improving the information processing system,

[1001] A system that includes this.

[1002] (Claim 2)

[1003] The system according to claim 1, which, in receiving and analyzing the aforementioned voice or text data, uses user-specific information and emotional state to provide personalized responses.

[1004] (Claim 3)

[1005] The system according to claim 1, further comprising means for contacting the relevant department and providing feedback on the results if the analyzed information and sentiment analysis results are insufficient. [Explanation of Symbols]

[1006] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for receiving audio or text information and analyzing it using natural language processing technology, A means for retrieving information from a storage device that stores related past cases based on the analyzed information, A means for instantly generating the optimal response based on search results and displaying it in the user interface, A means of collecting user feedback and using it to optimize the system, A means of converting voice input into text and sending it to a cloud-based processing device, A system that includes this.

2. The system according to claim 1, which uses user-specific data to provide a personalized response in receiving and analyzing the aforementioned voice or text information.

3. The system according to claim 1, further comprising means for querying relevant categories and returning the results as an evaluation if the analyzed information is insufficient.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A