System

The system addresses inefficiencies in call centers by using voice recognition and natural language processing to automate inquiry handling, ensuring quick and consistent responses with real-time feedback.

JP2026024092APending Publication Date: 2026-02-13SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024126413
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-01
Publication Date
2026-02-13

Smart Images

  • Figure 2026024092000001_ABST
    Figure 2026024092000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for receiving a speech input; means for converting the speech input into text data; means for analyzing the text data to classify consultation contents; means for automatically determining a corresponding department based on the consultation contents; means for transmitting the consultation contents to the corresponding department; and means for displaying a summary of the consultation contents and a standard answer example to the corresponding department.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In conventional call centers, customers have to select from a complex menu when making an inquiry, which is time-consuming and laborious. It is also difficult for operators to respond quickly and appropriately to customer inquiries, resulting in long response times and inconsistent responses. It is necessary to solve these issues and improve the efficiency of both customer and call center operations. [Means for solving the problem]

[0005] In order to solve the above problems, the present invention provides the following means.

[0006] First, a means for receiving voice input is provided, which receives a user's inquiry by voice. Then, a means for converting the received voice input into text data is provided, which converts the voice into text. Next, a means for analyzing the text data and classifying the inquiry content is provided, and a means for automatically determining the department to handle the inquiry based on the inquiry content is provided. This saves the user time and enables a prompt response to the inquiry. Furthermore, by providing a means for transmitting the inquiry content to the department to handle the inquiry, and a means for displaying a summary of the inquiry content and a standard response example on the operator terminal of the department, the operator can provide a prompt and consistent response. Furthermore, by providing a means for providing feedback to the user about the inquiry content by voice output, the current processing status can be appropriately communicated to the user. This results in an overall improvement in customer satisfaction and an increase in the efficiency of call center operations.

[0007] "Voice input" refers to voice data when a user makes an inquiry to a call center.

[0008] A "means for receiving voice input" is a means for capturing voice data from a user and processing it within the system.

[0009] "Text data" refers to data that is the result of converting voice data into text using voice recognition technology.

[0010] "Means for converting speech to text data" refers to a speech recognition engine or related technology used to convert received speech data into text data.

[0011] The "contents of consultation" refer to the specific problems or questions that the user inquires about with the call center.

[0012] "Means for classifying consultation content" refers to natural language processing techniques and algorithms that analyze text data and assign consultation content to specific categories.

[0013] A "response department" is a specialized department within a call center established to appropriately respond to specific consultation content.

[0014] "Means for automatically determining the department to handle inquiries" refers to technologies or algorithms that automatically assign inquiries to the most appropriate specialized department based on the classification results of the consultation content.

[0015] The "means for transmitting the consultation content to the department" refers to a communication technique or function for transferring the details of the user's consultation content to the determined department.

[0016] The "summary" is information that briefly summarizes the content of the user's consultation.

[0017] "Standard response examples" are samples or guidelines of responses that are generally used in response to specific inquiries.

[0018] "Means for displaying a summary of the consultation content and a standard response example" refers to the technology and functions for displaying a summary of the consultation content and a standard response example on the terminal so that the operator can respond quickly and appropriately.

[0019] "Voice output" refers to the output format in which the system communicates information to the user using voice synthesis technology.

[0020] "Means for providing feedback to the user about the consultation content by audio output" refers to technology or functionality for notifying the user of information such as the processing status and forwarding destination by audio. [Brief explanation of the drawings]

[0021] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0022] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0023] First, the terms used in the following description will be explained.

[0024] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0025] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0026] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0027] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0028] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0029] [First embodiment]

[0030] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0031] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0032] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0033] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0034] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0035] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0036] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0037] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0038] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0039] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0040] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0041] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0042] The present invention relates to a system for improving the efficiency of inquiries handled at a call center by using a voice recognition function. The system can be implemented in the following forms.

[0043] Acquiring and converting voice input

[0044] When a user calls the call center, the server launches speech recognition software to capture the user's voice input in real time. The server temporarily stores this voice data and sends it to the speech recognition engine. The speech recognition engine converts the voice data into text data. For example, if a user says, "My Internet connection is slow," the speech recognition engine converts this speech into text: "My Internet connection is slow."

[0045] Analysis and classification of consultation content

[0046] The text data is sent to a natural language processing (NLP) engine in the server. The NLP engine analyzes the text data and classifies the inquiry into a specific category. For example, the text data "My Internet connection is slow" would be classified into the "Technical Support" category.

[0047] Allocation to departments

[0048] Based on the classified content, the server automatically determines the appropriate department and specialized operator. For example, if the call is classified as "technical support," the call will be automatically connected to an operator in the technical support department.

[0049] Operator notification

[0050] The server then sends a summary of the user's inquiry and a standard response to the operator's terminal in the selected department. The terminal displays this information in a pop-up format, allowing the operator to respond quickly and appropriately. For example, in response to a question about a slow internet connection, connection troubleshooting procedures and general solutions are displayed.

[0051] User Feedback

[0052] While waiting for the connection to the appropriate department, the server will notify the user of the current processing status by voice. For example, it may say, "Your call is now being transferred to the technical support department. Please wait a moment." This voice feedback allows the user to understand the processing status and reduces anxiety while waiting.

[0053] Specific examples

[0054] When a user calls a call center complaining of an incorrect credit card charge, the system works as follows:

[0055] 1. A user calls and says, "There's an incorrect charge on my credit card."

[0056] 2. The server captures the voice data and uses a speech recognition engine to convert it into text data such as "The credit card charge is incorrect."

[0057] 3. The text data is sent to the NLP engine on the server and classified as a "billing inquiry."

[0058] 4. The server automatically connects to an operator in the billing department.

[0059] 5. The server sends a summary of the user's inquiry and an example of an appropriate response to the operator terminal in the billing department, which then displays this information in a pop-up window.

[0060] 6. The server will notify the user by voice, "Your call will now be forwarded to the billing department. Please wait a moment."

[0061] 7. Operators will respond promptly and appropriately based on the summary and standard response examples.

[0062] In this way, the system of the present invention utilizes voice recognition and natural language processing to reduce the burden on users and improve the efficiency of call center operations.

[0063] The processing flow will be explained below.

[0064] Step 1:

[0065] A user calls a call center.

[0066] Step 2:

[0067] The server runs speech recognition software to capture the user's speech in real time.

[0068] Step 3:

[0069] The server temporarily stores the captured audio data as streaming.

[0070] Step 4:

[0071] The server sends the voice data to the voice recognition engine.

[0072] Step 5:

[0073] The server uses a voice recognition engine to convert the voice data into text data.

[0074] Step 6:

[0075] The server receives the converted text data.

[0076] Step 7:

[0077] The server sends the text data to a natural language processing (NLP) engine.

[0078] Step 8:

[0079] The server analyzes the text data using an NLP engine and classifies the consultation content into specific categories.

[0080] Step 9:

[0081] Based on the analysis results, the server automatically determines the appropriate department and specialized operator to handle the situation.

[0082] Step 10:

[0083] The server transmits the consultation content to the determined department.

[0084] Step 11:

[0085] The terminal (operator's terminal) displays a summary of the received consultation and a standard response example in a pop-up format.

[0086] Step 12:

[0087] The server notifies the user of the current processing status using a speech synthesis engine.

[0088] Step 13:

[0089] Users receive notifications and are informed of the progress of the process.

[0090] Step 14:

[0091] The operator will respond quickly based on the summary and standard response examples displayed on the terminal.

[0092] Step 15:

[0093] Once the operator has completed the response, the results are fed back to the server.

[0094] Specific examples

[0095] A detailed process flow when a user contacts a call center complaining about a slow internet connection:

[0096] Step 1:

[0097] A user says, "My internet connection is slow."

[0098] Step 2:

[0099] The server launches the speech recognition software and captures the audio.

[0100] Step 3:

[0101] The server temporarily stores the captured audio as streaming.

[0102] Step 4:

[0103] The server sends the voice data to the voice recognition engine.

[0104] Step 5:

[0105] The server converts the voice data into text data saying "Your internet connection is slow."

[0106] Step 6:

[0107] The server receives the text data.

[0108] Step 7:

[0109] The server sends the text data to the NLP engine.

[0110] Step 8:

[0111] The server analyzes it using an NLP engine and classifies it as "Technical Support."

[0112] Step 9:

[0113] The server decides to assign it to the "Technical Support" department.

[0114] Step 10:

[0115] The server sends the text data to an operator in the technical support department.

[0116] Step 11:

[0117] The device will pop up a summary that reads "Your internet connection is slow" along with a standard example response.

[0118] Step 12:

[0119] The server will notify the user by voice, "Your call will now be transferred to the technical support department. Please wait a moment."

[0120] Step 13:

[0121] Ensure the user is "transferred to technical support."

[0122] Step 14:

[0123] An operator will quickly guide you through the steps to resolving the problem based on a summary and standard answers.

[0124] Step 15:

[0125] Once the operator has completed the response, the results are fed back to the server.

[0126] Example 1

[0127] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0128] Traditional call centers lacked the technology to efficiently handle user inquiries, and in many cases, operators had to respond manually. This resulted in slow response times and lower user satisfaction. Furthermore, there were sometimes delays in assigning inquiries to the appropriate department, further extending the inquiry time. Furthermore, users could not check the current processing status while waiting, which left them with no way to alleviate their anxiety while waiting.

[0129] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0130] In this invention, the server includes means for receiving voice input, means for converting the voice input into text data, means for analyzing the text data and classifying the consultation content, means for automatically determining a department to handle the consultation based on the consultation content, means for transmitting the consultation content to the department, means for displaying a summary of the consultation content and a standard answer example to the department, and means for providing voice feedback on the processing status to the user. This makes it possible to efficiently distribute inquiries from users to the departments to handle them quickly. Furthermore, the feedback to the user can reduce anxiety while waiting and improve user satisfaction.

[0131] "Means for receiving audio input" refers to equipment or software that captures audio uttered by a user through an audio input device such as a telephone or microphone as digital audio data.

[0132] "Means for converting voice input into text data" refers to a voice recognition engine or software that analyzes captured voice data and converts it into corresponding text data.

[0133] "Means for analyzing text data and classifying consultation content" refers to a system or software that uses a natural language processing engine to analyze text data obtained through voice recognition and classify the consultation content into specific categories.

[0134] "Means for automatically determining the appropriate department based on the content of the consultation" refers to a system or algorithm for selecting the appropriate department or specialized operator based on the classified content of the consultation.

[0135] The "means for transmitting the consultation content to the department in charge" refers to a communication means or protocol for electronically transmitting the user's consultation content to the determined department or operator in charge.

[0136] "Means for displaying a summary of the consultation content and a standard response example to the relevant department" refers to software for summarizing the consultation content received and displaying corresponding standard response examples on the operator terminal.

[0137] The "means for providing the user with voice feedback on the processing status" refers to a voice synthesis engine or system that notifies the user of the current processing status while the user is waiting as a voice message.

[0138] The present invention relates to a system for improving the efficiency of inquiries at a call center by using a voice recognition function. This system includes means for receiving voice input, means for converting the voice input into text data, means for analyzing the text data and classifying the inquiry content, means for automatically determining the department to handle the inquiry based on the inquiry content, means for transmitting the inquiry content to the department, means for displaying a summary of the inquiry content and a standard response example to the department, and means for providing voice feedback on the processing status to the user. The following specific hardware and software are used to implement the invention.

[0139] Acquiring and converting voice input

[0140] When a user contacts the call center via telephone or microphone, the server invokes speech recognition software, using a speech recognition engine such as the Google Cloud Speech-to-Text API, to capture the user's voice input in real time and temporarily store it as audio data.

[0141] The server sends the captured voice data to a speech recognition engine, which converts the voice data into corresponding text data, for example, if the user says "My internet connection is slow," the speech recognition engine converts the voice into text data saying "My internet connection is slow."

[0142] Analysis and classification of consultation content

[0143] The converted text data is sent to a natural language processing engine on the server. An NLP engine such as Google Cloud Natural Language API is used here. The NLP engine analyzes the text data and classifies the inquiry into a specific category. For example, text data such as "My internet connection is slow" would be classified into the "technical support" category.

[0144] Allocation to departments

[0145] The server automatically determines the corresponding department and specialized operator based on the category information received from the NLP engine. For example, if the call is classified as "technical support," the server will automatically connect to an operator in the technical support department.

[0146] Operator notification

[0147] The server then sends a summary of the user's inquiry and a standard response to the operator's terminal in the selected department. The user's inquiry is summarized and a standard response is displayed on the operator's terminal. For example, if the user says "my internet connection is slow," connection troubleshooting procedures and general solutions are displayed in a pop-up format on the terminal.

[0148] User Feedback

[0149] While waiting for the call to be connected to the appropriate department, the server will notify the user of the current processing status by voice. This uses speech synthesis software such as Google Cloud Text-to-Speech API. For example, the server may say, "Your call is now being transferred to the technical support department. Please wait a moment." This voice feedback allows the user to understand the processing status and reduces anxiety while waiting.

[0150] Specific examples

[0151] When a user reports an incorrect charge on their credit card, the system works as follows:

[0152] 1. A user calls and says, "There's an incorrect charge on my credit card."

[0153] 2. The server captures the voice data and uses a speech recognition engine to convert it into text data such as "The credit card charge is incorrect."

[0154] 3. The text data is sent to the NLP engine on the server and classified as a "billing inquiry."

[0155] 4. The server automatically connects to an operator in the billing department.

[0156] 5. The server sends a summary of the user's inquiry and an example of an appropriate response to the operator terminal in the billing department, which then displays this information in a pop-up window.

[0157] 6. The server will notify the user by voice, "Your call will now be forwarded to the billing department. Please wait a moment."

[0158] 7. Operators will respond promptly and appropriately based on the summary and standard response examples.

[0159] Prompt Sentence Examples

[0160] Examples of prompts to be input to a generative AI model:

[0161] "Please explain how this system responds when a user contacts us with a credit card billing issue."

[0162] In this way, the system of the present invention utilizes voice recognition and natural language processing to reduce the burden on users and improve the efficiency of call center operations.

[0163] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0164] Step 1:

[0165] When a user contacts the call center via telephone or microphone, the server activates voice recognition software and captures the user's voice input in real time. The input is the user's voice data, and the output is the captured digital voice data. Specifically, the server obtains the voice data using the Google Cloud Speech-to-Text API or similar.

[0166] Step 2:

[0167] The server sends the captured voice data to the speech recognition engine. The input is the captured digital voice data, and the output is text data. Specifically, the server sends the voice data to the Google Cloud Speech-to-Text API via an HTTP request and receives the text data as a response from the API.

[0168] Step 3:

[0169] The converted text data is sent to the natural language processing engine on the server. The input is text data from the speech recognition engine, and the output is category information for the consultation content. Specifically, the server sends the text data to the Google Cloud Natural Language API and receives category information as a response from the API.

[0170] Step 4:

[0171] The server automatically determines the appropriate department and specialized operator based on the category information received from the NLP engine. The input is the category information of the consultation content, and the output is information on the corresponding department and operator. Specifically, the server references an internal database to map the category and the corresponding department.

[0172] Step 5:

[0173] The server then sends a summary of the user's consultation and a standard example response to the operator terminal of the selected department. The input is information about the department and operator and text data about the consultation, and the output is a summary and example response that are displayed on the operator terminal. Specifically, the server sends the consultation data to the operator terminal as an HTTP request, and the terminal displays this data in a pop-up.

[0174] Step 6:

[0175] While waiting for the corresponding department to connect, the server notifies the user of the current processing status by voice. The input is text data about the current processing status, and the output is synthesized voice data. Specifically, the server uses the Google Cloud Text-to-Speech API to convert the text data into voice data and play it back to the user.

[0176] (Application example 1)

[0177] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0178] Virtual stores require a means to quickly and appropriately respond to inquiries and questions users have about a wide variety of products. However, conventional text-based or operator-based responses can be slow, and a lack of human resources can lead to poor user experience. For this reason, there is a need for technology that uses voice input to automatically analyze inquiries and quickly provide appropriate information.

[0179] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0180] In this invention, the server includes means for receiving voice input, means for converting the voice input into text data, and means for analyzing the text data and classifying the consultation content, thereby making it possible to obtain product information from the database based on the classification results and to respond to the user with the obtained product information in voice and text.

[0181] "Voice input" refers to the device capturing voice uttered by the user.

[0182] "Converting to text data" refers to the process of converting voice input into written information.

[0183] "Analyzing text data" means understanding the content of converted text data and identifying its meaning and intent.

[0184] "Classifying the consultation content" means dividing the content into specific categories or groups based on the analyzed text data.

[0185] "Automatically determining the department to handle the inquiry" means automatically selecting the most appropriate department and person in charge based on the classified inquiry content.

[0186] "Sending the consultation content to the department in charge" means transferring the consultation content and its information to the determined department or person in charge.

[0187] "Displaying a summary of the consultation content and a standard response example" means displaying a concise summary of the consultation content and a recommended response on the operator's terminal.

[0188] "Acquiring product information from a database based on the classification result" means searching the database for product information related to the classified consultation content and acquiring it.

[0189] "Responding to the user with the acquired product information by voice and text" means returning the acquired product information to the user as voice and text information.

[0190] The present invention relates to an application for improving the efficiency of customer support within a virtual store. The system can be implemented in the following forms.

[0191] Hardware and Software Configuration

[0192] Hardware: Smartphone (including microphone, speaker, and display)

[0193] Software: Speech recognition engine (Google Speech-to-Text API), natural language processing engine (Google Cloud Natural Language API), database server (Google Cloud Firestore), application (compatible with Android / iOS)

[0194] Program processing flow

[0195] 1. Acquiring voice input

[0196] The user launches the smartphone app and makes a voice inquiry. The smartphone app captures the user's voice through the microphone and temporarily stores it as voice data.

[0197] 2. Audio data conversion

[0198] The server converts the acquired voice data into text data using the Google Speech-to-Text API. For example, if a user asks, "What size is this dress?", the voice data is converted directly into text information.

[0199] 3. Analysis and Classification of Text Data

[0200] The converted text data is analyzed using the Google Cloud Natural Language API on the server and classified according to the inquiry content, for example, "inquiry about size."

[0201] 4. Determining Corresponding Information

[0202] Based on the classified inquiry, the server retrieves related product information from a database (Google Cloud Firestore), such as "This dress comes in sizes S, M, L, and XL."

[0203] 5. User Feedback

[0204] The acquired product information is returned to the user in voice and text format. The smartphone application provides voice feedback through the speaker and displays text on the display. For example, "This dress is available in sizes S, M, L, and XL" is displayed along with a voice response.

[0205] 6. Notify Operator (if necessary)

[0206] If a complex inquiry is made, the server will escalate the request to a specialist operator who can provide more detailed information and provide an appropriate response.

[0207] Specific examples

[0208] If a user asks "What size is this dress?" in a virtual store, the system works as follows:

[0209] 1. The user makes a voice inquiry through a smartphone app asking, "What size is this dress?"

[0210] 2. The server captures the voice data and converts it into text data such as "What size is this dress?" using the Google Speech-to-Text API.

[0211] 3. The server analyzes the text data using the Google Cloud Natural Language API and classifies it as a "size query."

[0212] 4. Based on the classification results, the server retrieves related product information from Google Cloud Firestore, such as "This dress comes in sizes S, M, L, and XL."

[0213] 5. The server responds to the user with the acquired product information in voice and text format. The smartphone app displays the text on the screen and simultaneously responds with voice.

[0214] Example prompts for natural language generation AI models

[0215] markdown

[0216] Prompt: In a virtual shopping assistant, if a user asks "What size is this dress?", generate a prompt to respond with the appropriate size information to provide to the user.

[0217] Examples:

[0218] markdown

[0219] User input: "What size is this dress?"

[0220] Sizing Information: "This dress comes in sizes S, M, L, and XL."

[0221] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0222] Step 1:

[0223] Acquiring voice input

[0224] The user launches the smartphone app and inputs the inquiry by voice. The smartphone is equipped with a microphone, which captures the user's voice. The input voice is temporarily saved in the smartphone application.

[0225] Input: User voice input

[0226] Output: Audio data

[0227] Step 2:

[0228] Audio data conversion

[0229] The server receives voice data sent from the smartphone. The voice data is sent to the Google Speech-to-Text API and converted into text data in real time. For example, voice data such as "What size is this dress?" is converted into text data such as "What size is this dress?"

[0230] Input: Audio data

[0231] Output: Text data

[0232] Step 3:

[0233] Text data analysis and classification

[0234] The server then sends the converted text data to the Google Cloud Natural Language API, which analyzes the text data and classifies the inquiry into a specific category. For example, the text data "What size is this dress?" would be classified as a "size inquiry."

[0235] Input: Text data

[0236] Output: Classification results

[0237] Step 4:

[0238] Determining correspondence information

[0239] The server searches and retrieves related product information from a database (Google Cloud Firestore) based on the classified inquiry. For example, for an inquiry about size, it retrieves the size information of the specified product.

[0240] Input: Classification results

[0241] Output: Product information data

[0242] Step 5:

[0243] User Feedback

[0244] The server generates the acquired product information as a voice response and a text response. The generated voice data is fed back to the user through the smartphone's speaker, and the text data is displayed on the smartphone's display. For example, the server responds by voice, saying, "This dress comes in sizes S, M, L, and XL," and the same content is displayed on the display.

[0245] Input: Product information data

[0246] Output: Audio and text responses

[0247] Step 6:

[0248] Notify the operator (if necessary)

[0249] If the inquiry is complex and requires a detailed response, the server escalates the inquiry to an appropriate operator, who then sends a summary of the inquiry and a recommended response to the operator's terminal.

[0250] Input: Classification results and product information data

[0251] Output: Notification to operator (summary and standard response example)

[0252] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0253] The present invention relates to a system for improving the efficiency of inquiries handled at a call center by using a voice recognition function and an emotion engine. The system can be implemented in the following forms.

[0254] Acquiring and converting voice input

[0255] When a user calls the call center, the server launches speech recognition software to capture the user's voice input in real time. The server temporarily stores this voice data and sends it to the speech recognition engine. The speech recognition engine converts the voice data into text data. In this case, if the user says, "My Internet connection is slow," the speech recognition engine converts this speech into text: "My Internet connection is slow."

[0256] Emotion analysis

[0257] The converted text data is sent to the server's natural language processing (NLP) engine and also to the emotion engine, which analyzes the user's emotions from the voice data and identifies emotional states such as anger, joy, and sadness. This emotion analysis information is used for subsequent processing.

[0258] Analysis and classification of consultation content

[0259] The NLP engine analyzes the text data and classifies the inquiry into a specific category. For example, the text data "My internet connection is slow" will be classified into the "Technical Support" category.

[0260] Allocation to departments

[0261] The server automatically determines the department and specialized operator to respond to the inquiry based on the classified content and the user's emotions obtained from emotion analysis. For example, if the user expresses anger, the server will prioritize connecting the user to an operator who can respond more quickly and professionally.

[0262] Operator notification

[0263] The server then sends a summary of the user's consultation, a standard response example, and the user's emotional state to the operator terminal in the selected department. The terminal displays this information in a pop-up format, allowing the operator to respond quickly and appropriately while taking the user's emotions into consideration. For example, if the message "Internet connection is slow" and the emotional state of "anger" are displayed, the operator can prepare to respond calmly and quickly.

[0264] User Feedback

[0265] While waiting for the connection to the appropriate department, the server will notify the user of the current processing status by voice. For example, it may say, "Your call is now being transferred to the technical support department. Please wait a moment." This voice feedback allows the user to understand the processing status and reduces anxiety while waiting.

[0266] Specific examples

[0267] If a user angrily calls a call center complaining that their credit card was charged incorrectly, the system works like this:

[0268] 1. A user calls and angrily states that there is an error in the charge on their credit card.

[0269] 2. The server captures the voice data and uses a speech recognition engine to convert it into text data such as "The credit card charge is incorrect."

[0270] 3. The text and voice data are sent to the NLP engine and emotion engine on the server. The emotion engine analyzes the user's anger.

[0271] 4. The text data is classified as a "billing inquiry" and the emotional state of "anger" is detected along with the results of the emotion analysis.

[0272] 5. The server automatically connects to an operator in the billing department, and based on the results of emotion analysis, selects an operator who is best able to handle the emotion of "anger."

[0273] 6. The server sends a summary of the user's consultation, their emotional state, and examples of appropriate responses to the operator terminal in the billing department. The terminal displays this information in a pop-up.

[0274] 7. The server will notify the user by voice, "Your call will now be forwarded to the billing department. Please wait."

[0275] 8. The operator responds calmly, quickly, and sensitively based on the summary and the caller's emotional state.

[0276] In this way, the system of the present invention, which combines an emotion engine, not only performs speech recognition and natural language processing but also analyzes and takes into account the user's emotional state, thereby reducing the burden on the user and further improving the efficiency of call center operations.

[0277] The processing flow will be explained below.

[0278] Step 1:

[0279] A user calls a call center.

[0280] Step 2:

[0281] The server runs speech recognition software to capture the user's speech in real time.

[0282] Step 3:

[0283] The server temporarily stores the captured audio data as streaming.

[0284] Step 4:

[0285] The server sends the voice data to the voice recognition engine.

[0286] Step 5:

[0287] The server uses a speech recognition engine to convert the voice data into text data, for example, "Your internet connection is slow."

[0288] Step 6:

[0289] The server then sends the converted text and voice data to an emotion engine, which analyzes the user's emotions. For example, emotions such as "anger" or "dissatisfaction" are detected from voice patterns.

[0290] Step 7:

[0291] At the same time, the server sends the converted text data to a natural language processing (NLP) engine.

[0292] Step 8:

[0293] The server analyzes the text data using an NLP engine and classifies the inquiry into a specific category. For example, the text "My internet connection is slow" is classified into the "Technical Support" category.

[0294] Step 9:

[0295] The server automatically determines the department and specialist operator to respond to the inquiry based on the content of the inquiry and the results of emotion analysis. For example, if emotion analysis detects anger, it will select an operator from the technical support department who can respond quickly.

[0296] Step 10:

[0297] The server sends the consultation content and emotion analysis results to the determined department.

[0298] Step 11:

[0299] The terminal (operator's terminal) displays a summary of the received consultation, a standard response example, and analyzed emotional information in a pop-up format. For example, it displays the summary "My internet connection is slow" and the emotional state "Anger."

[0300] Step 12:

[0301] The server uses a speech synthesis engine to notify the user of the current status of the process, for example, "Your call is being transferred to technical support. Please wait."

[0302] Step 13:

[0303] Users receive notifications and are informed of the progress of the process.

[0304] Step 14:

[0305] The operator responds quickly and appropriately based on the summary and emotional information displayed on the terminal, for example, by speaking calmly and politely to alleviate the user's anger and try to resolve the technical problem.

[0306] Step 15:

[0307] Once the operator has completed the response, the results are fed back to the server.

[0308] Specific examples

[0309] A detailed process flow when a user contacts a call center because they are angry about an incorrect credit card charge:

[0310] Step 1:

[0311] A user says, "There's an incorrect charge on my credit card."

[0312] Step 2:

[0313] The server launches the speech recognition software and captures the audio.

[0314] Step 3:

[0315] The server temporarily stores the captured audio as streaming.

[0316] Step 4:

[0317] The server sends the voice data to the voice recognition engine.

[0318] Step 5:

[0319] The server converts the voice data into text data saying "The credit card billing is incorrect."

[0320] Step 6:

[0321] The server sends the text data and voice data to the emotion engine, which then analyzes the user's emotion of "anger."

[0322] Step 7:

[0323] The server sends the text data to the NLP engine.

[0324] Step 8:

[0325] The server analyzes the text data using an NLP engine and classifies it as a "billing inquiry."

[0326] Step 9:

[0327] The server determines the department to handle the inquiry based on the emotion of "invoice inquiry" and "anger." For example, it may give priority to an operator in the billing department who can respond quickly.

[0328] Step 10:

[0329] The server sends the consultation details and emotion analysis results to the billing department operator.

[0330] Step 11:

[0331] The device will pop up a summary saying "Credit card charge incorrect" and an emotional state of "Anger."

[0332] Step 12:

[0333] The server will announce, "Your call is now being forwarded to the billing department. Please wait a moment."

[0334] Step 13:

[0335] Verify that the user is "transferred to the billing department."

[0336] Step 14:

[0337] The operator responds calmly and quickly based on the summary and emotional information, for example, by first apologizing to calm the user's anger and then confirming the details of the problem.

[0338] Step 15:

[0339] Once the operator has completed the response, the results are fed back to the server.

[0340] In this way, the system of the present invention, which combines an emotion engine, not only performs speech recognition and natural language processing but also analyzes and takes into account the user's emotional state, thereby reducing the burden on the user and further improving the efficiency of call center operations.

[0341] Example 2

[0342] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0343] Handling inquiries at call centers requires listening to voice input, analyzing content, analyzing emotions, assigning calls to the appropriate department, and providing information to operators promptly. Conventional systems have difficulty performing these tasks efficiently, resulting in problems such as reduced user satisfaction and reduced operator work efficiency. The present invention aims to provide a system that solves these problems and enables efficient call center operation.

[0344] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0345] In this invention, the server includes means for receiving voice input, means for converting the voice input into text data, means for analyzing emotions from the text data and voice data, means for analyzing and classifying the consultation content, means for automatically determining a department to handle the consultation based on the consultation content and emotion analysis results, means for sending a summary of the consultation content and a standard response example to the department to handle the consultation, and means for displaying appropriate information on an operator terminal. This makes it possible to analyze the user's voice input in real time and quickly transfer the consultation to the appropriate department or operator.

[0346] A "means for receiving voice input" is a device or program that has the function of capturing voice from a user and recording it as digital voice data.

[0347] The "means for converting voice input into text data" is a device or program that has the function of analyzing captured voice data and converting it into corresponding text data.

[0348] The "means for analyzing text data and classifying consultation contents" refers to a device or program that has the function of analyzing the contents based on the text data and classifying them into specific categories.

[0349] The "means for analyzing emotions" is a device or program that has the function of identifying the emotional state of a user from voice data and text data.

[0350] The "means for automatically determining the department to handle the call" is a device or program that has the function of selecting the appropriate department or operator to handle the call based on the classified content of the call and the results of emotion analysis.

[0351] The "means for transmitting the consultation content to the department" is a device or program having a function for transmitting the consultation content of the user and related information to the automatically determined department.

[0352] The "means for displaying a summary of the consultation content and a standard response example" is a device or program that has the function of displaying a summary of the consultation content and a standard response example on the operator terminal, thereby supporting a prompt and appropriate response.

[0353] This invention is a system that uses voice recognition and emotion analysis functions to improve the efficiency of inquiries handled at call centers. This system converts user voice input into text data, analyzes the content of the inquiry and the user's emotional state, and automatically assigns the inquiry to the appropriate department and operator.

[0354] To implement this system, the following hardware and software are used.

[0355] Hardware

[0356] server

[0357] Operator terminal (e.g. Windows PC)

[0358] software

[0359] Voice recognition software (e.g., Google Cloud Speech-to-Text)

[0360] Natural Language Processing (NLP) engines (e.g., SpaCy)

[0361] Sentiment analysis engine (e.g. IBM Watson Tone Analyzer)

[0362] Program processing

[0363] Acquiring voice input

[0364] When a user calls the call center, the server captures the user's voice in real time through a telephone line interface and stores it as digital voice data.

[0365] Speech recognition and text conversion

[0366] The server sends the stored voice data to a speech recognition engine, which converts it into text data. For example, if a user says, "My internet connection is slow," the voice data will be converted into text data saying, "My internet connection is slow."

[0367] Emotion analysis

[0368] The converted text and voice data is sent to the emotion engine on the server, which analyzes the data and identifies the user's emotion (e.g., anger, joy, sadness).

[0369] Analysis and classification of consultation content

[0370] The server sends the text data to an NLP engine, which analyzes it and classifies it into a specific category, for example, "My internet connection is slow" would be classified into the "tech support" category.

[0371] Allocation to departments

[0372] The server automatically determines the appropriate department and operator based on the category of the consultation content and the results of emotion analysis. If the user expresses anger, the call will be connected preferentially to an operator who can respond quickly and professionally.

[0373] Operator notification

[0374] The server then sends a summary of the user's consultation and their emotional state to the operator terminal in the selected department. The terminal displays this information in a pop-up format, allowing the operator to respond promptly and appropriately.

[0375] User Feedback

[0376] While waiting for the connection to the appropriate department, the server will notify the user of the current processing status by voice. For example, by guiding the user to "Your call is now being transferred to the technical support department. Please wait a moment," the user can understand the processing status and reduce anxiety while waiting.

[0377] Specific examples

[0378] If a user angrily calls a call center complaining that their credit card was charged incorrectly, the system works like this:

[0379] 1. A user calls and angrily states that there is an error in the charge on their credit card.

[0380] 2. The server captures the voice data and uses a speech recognition engine to convert it into text data such as "The credit card charge is incorrect."

[0381] 3. The text and voice data are sent to the NLP engine and emotion engine on the server. The emotion engine analyzes the user's anger.

[0382] 4. The text data is classified as a "billing inquiry" and the emotional state of "anger" is detected along with the results of the emotion analysis.

[0383] 5. The server automatically connects to an operator in the billing department, and based on the results of emotion analysis, selects an operator who is best able to handle the emotion of "anger."

[0384] 6. The server sends a summary of the user's consultation, their emotional state, and examples of appropriate responses to the operator terminal in the billing department. The terminal displays this information in a pop-up.

[0385] 7. The server will notify the user by voice, "Your call will now be forwarded to the billing department. Please wait."

[0386] 8. The operator responds calmly, quickly, and sensitively based on the summary and the caller's emotional state.

[0387] Prompt Sentence Examples

[0388] "A user is expressing anger over a claim that their credit card is incorrectly charged. What's the best approach to calmly respond?"

[0389] "Some users are complaining about slow internet connections. Can you suggest a good solution for this?"

[0390] In this way, the system of the present invention, which combines an emotion engine, not only performs speech recognition and natural language processing but also analyzes and takes into account the user's emotional state, thereby reducing the burden on the user and further improving the efficiency of call center operations.

[0391] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0392] Step 1: Getting voice input

[0393] When the server detects an incoming call from the user, it immediately activates the voice recognition software and captures the user's voice input in real time. The input is an analog voice signal received over the telephone line. The server converts this voice signal into digital voice data and temporarily stores it.

[0394] Specific behavior:

[0395] The user makes a call.

[0396] The telephone line interface implemented in the server detects the incoming call.

[0397] The server receives voice input through a telephone line interface and stores it as digital voice data.

[0398] Input: Analog audio signal obtained through telephone lines.

[0399] Output: Digital audio data.

[0400] Step 2: Speech recognition and text conversion

[0401] The server then sends the stored digital voice data to a speech recognition engine (such as Google Cloud Speech-to-Text) to convert it into text data. The speech recognition engine then analyzes the digital voice data and generates corresponding text data.

[0402] Specific behavior:

[0403] The server calls the API to send the voice data to the voice recognition engine.

[0404] The speech recognition engine analyzes the speech and generates text data.

[0405] The server receives the generated text data.

[0406] Input: Digital audio data.

[0407] Output: Text data.

[0408] Step 3: Sentiment Analysis

[0409] The server sends the generated text data and saved voice data to an emotion engine (such as IBM Watson Tone Analyzer) to analyze the user's emotional state. The emotion engine analyzes the voice and text data to identify the user's emotions (e.g., anger, joy, sadness).

[0410] Specific behavior:

[0411] The server calls the API to send text and voice data to the emotion engine.

[0412] An emotion engine analyzes the data and identifies the user's emotional state.

[0413] The server receives the results of the sentiment analysis.

[0414] Input: text data, audio data.

[0415] Output: Sentiment analysis results (e.g. anger, joy, sadness).

[0416] Step 4: Analysis and classification of consultation content

[0417] The server sends the text data to a natural language processing (NLP) engine, which analyzes it and classifies the query into a specific category. The NLP engine analyzes the text and classifies it into a specific category (e.g., technical support, billing).

[0418] Specific behavior:

[0419] The server calls the API to send the text data to the NLP engine.

[0420] The NLP engine analyzes the text data and identifies categories.

[0421] The server receives the analysis results.

[0422] Input: Text data.

[0423] Output: Classification results (e.g., tech support, billing).

[0424] Step 5: Assign to departments

[0425] The server automatically determines the appropriate department and operator based on the category of the consultation content and the results of emotion analysis. If the user expresses anger, the caller will be connected preferentially to an operator who can provide a prompt and professional response.

[0426] Specific behavior:

[0427] The server refers to the emotion analysis and classification results and runs an algorithm to select the appropriate department or operator.

[0428] The server issues connection instructions to the selected department or operator.

[0429] Input: Sentiment analysis results, classification results.

[0430] Output: Information on the selected response department and operator.

[0431] Step 6: Notify the operator

[0432] The server then sends a summary of the user's consultation and their emotional state to the operator terminal in the selected department. The operator terminal displays this information in a pop-up format, helping to ensure a prompt and appropriate response.

[0433] Specific behavior:

[0434] The server sends a summary of the user's consultation and the results of emotion analysis to the operator terminal.

[0435] The information received by the operator terminal is displayed as a pop-up.

[0436] Input: Summary of consultation content, sentiment analysis results.

[0437] Output: Information displayed on the operator terminal.

[0438] Step 7: User feedback

[0439] While waiting for the call to be connected to the appropriate department, the server will notify the user of the current processing status by voice, for example, "Your call is now being transferred to the technical support department. Please wait a moment."

[0440] Specific behavior:

[0441] The server generates a feedback message for the user.

[0442] The server generates this message as voice data using voice synthesis software and transmits it to the user over a telephone line.

[0443] Input: Current processing status.

[0444] Output: Audio feedback to the user.

[0445] (Application example 2)

[0446] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0447] Customer service in modern brick-and-mortar stores requires a fast and accurate response to customer problems and complaints. However, traditional methods make it difficult to consider customer emotions, making it difficult to increase customer satisfaction. In particular, there is a lack of appropriate responses to emotional customers, which leads to a decline in the service quality of the entire store.

[0448] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0449] In this invention, the server includes a means for receiving voice input, a means for converting the voice input into text data, and a means for analyzing the user's emotional state from the text data and voice data, thereby enabling the server to analyze the customer's emotions in real time, automatically determine an appropriate response method, and notify store staff in a timely manner.

[0450] "Voice input" refers to voice information uttered by a user, and is information acquired by an input device.

[0451] "Text data" is data obtained by converting voice input into text format, and is information expressed as a character string.

[0452] "Consultation content" refers to the subject of the problem or question that the user communicates through voice input.

[0453] The "emotional state" is an emotional state analyzed from the user's voice input, and indicates, for example, anger, joy, sadness, etc.

[0454] The "department" refers to a specific department or person in charge within an organization that is set up to deal with inquiries from users.

[0455] The term "server" refers to a central computer system that receives voice input, converts it into text data, analyzes emotional states, classifies the content of inquiries, and determines which departments should respond.

[0456] "Speech recognition engine" refers to the software modules and algorithms used to convert voice input into text data.

[0457] An "NLP engine" is a software module for natural language processing that analyzes text data to extract meaning and intent.

[0458] "Operator" refers to the staff at the physical store who responds to customers based on the consultation content and emotional state sent from the server.

[0459] "Feedback" refers to the act of notifying the user of the current processing status by voice or the like.

[0460] The present invention relates to a system for improving the efficiency of customer service in a physical store by using a voice recognition function and an emotion engine. This system is implemented in the following manner.

[0461] Acquiring and converting voice input

[0462] The server receives customer voice input directly from the customer service robot in the physical store. This voice input is captured in real time using a microphone connected to the server. Then, the voice data is converted into text data using a speech recognition engine (e.g., Google Speech Recognition API).

[0463] Emotion analysis

[0464] The converted text and voice data are sent to the server's natural language processing (NLP) engine and emotion engine. The emotion engine uses, for example, the Hugging Face Transformers library to analyze the customer's emotions from the voice data into "anger," "happiness," "sadness," etc. This emotion analysis information is used in subsequent processing steps.

[0465] Analysis and classification of consultation content

[0466] The server's NLP engine analyzes the converted text data and classifies the consultation content into specific categories such as "product usage," "product defect," "inventory check," etc. Here, the classification process can be performed using a library such as TextBlob.

[0467] Deciding how to respond

[0468] The server automatically determines how to respond based on the customer's emotional state obtained from emotion analysis and the classification of the inquiry content. For example, if a customer is in a "confused" emotional state and says, "I don't know where the product is," the server will contact the appropriate staff member who can respond quickly.

[0469] Notification to store staff

[0470] The server then sends a summary of the customer's consultation and emotional state, along with examples of appropriate responses, to the terminal of the selected staff member. This information is displayed in a pop-up format on the staff member's terminal, allowing the staff member to respond quickly and appropriately while taking the customer's emotions into consideration.

[0471] User Feedback

[0472] While waiting for the customer to be connected to the staff member, the server notifies the customer of the current processing status by voice. This notification may include, for example, a message such as "We are currently contacting the staff member in charge. Please wait a moment." This allows the customer to understand the processing status and reduces anxiety while waiting.

[0473] Specific examples

[0474] For example, if a customer approaches a customer service robot in a confused manner, saying, "I don't know where my item is," the system works as follows:

[0475] 1. The server captures the voice data and uses a voice recognition engine to convert it into text data such as "I don't know where the product is."

[0476] 2. The text and voice data are sent to the NLP engine and emotion engine on the server. The emotion engine analyzes the customer's "confusion" emotion.

[0477] 3. The text data is classified as a “location inquiry” and the emotional state is detected along with the sentiment analysis results.

[0478] 4. The server automatically selects the staff member most suited to the consultation and sends a notification.

[0479] 5. The server will notify the customer by voice, "We are currently contacting the staff in charge. Please wait a moment."

[0480] 6. Staff respond quickly to distressed customers based on pop-up summary information and emotional state.

[0481] Example prompt sentence:

[0482] For example, for an input such as "I'm having trouble finding the product," sentiment analysis and content classification are performed. In this way, a system that can easily and effectively improve customer service in physical stores is realized.

[0483] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0484] Step 1:

[0485] The server receives customer voice input in real time from the customer service robot in the physical store. Based on this input, the server directly captures voice data using a high-performance microphone. The input is voice data saying, "I don't know where the product is."

[0486] Step 2:

[0487] The server sends the captured voice data to a voice recognition engine and converts it into text data. Here, the Google Speech Recognition API is used. The input is voice data, and the output is text data such as "I don't know where the product is."

[0488] Step 3:

[0489] The server sends the text and voice data to a natural language processing (NLP) engine and an emotion engine. The NLP engine uses a generative AI model (such as the Hugging Face Transformers library) to analyze the text data and classify the inquiry into a specific category. In this case, the input is the text data, and the output is the category "location inquiry."

[0490] Step 4:

[0491] The server uses an emotion engine to analyze the customer's emotional state from the voice data. The emotion engine also uses a generative AI model to analyze the emotion of "confused" from the voice data. The input is the voice data, and the output is the emotional state of "confused."

[0492] Step 5:

[0493] The server automatically determines how to respond based on the classified consultation content and the analyzed emotional state. In this case, a notification is sent to an available staff member for a customer who expresses "confusion" in a "location inquiry." The input is the consultation content and emotional information, and the output is the selection of the most suitable staff member.

[0494] Step 6:

[0495] The server sends a summary of the customer's consultation and emotional state, as well as examples of appropriate responses, to the staff member's terminal. This notification is displayed in a pop-up format on the terminal. The input is customer information and emotional analysis information, and the output is a notification message to the staff member.

[0496] Step 7:

[0497] The server provides the customer with audio feedback on the current processing status. For example, it generates a message such as "We are currently contacting the staff member in charge. Please wait a moment," and notifies the customer via an audio output device. The input is the response decision information, and the output is the audio notification message.

[0498] In this way, a system can be realized in physical stores that analyzes emotions and classifies content from customer voice input, and then quickly responds appropriately.

[0499] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0500] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0501] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0502] [Second embodiment]

[0503] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0504] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0505] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0506] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0507] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0508] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0509] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0510] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0511] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0512] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0513] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0514] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0515] The present invention relates to a system for improving the efficiency of inquiries handled at a call center by using a voice recognition function. The system can be implemented in the following forms.

[0516] Acquiring and converting voice input

[0517] When a user calls the call center, the server launches speech recognition software to capture the user's voice input in real time. The server temporarily stores this voice data and sends it to the speech recognition engine. The speech recognition engine converts the voice data into text data. For example, if a user says, "My Internet connection is slow," the speech recognition engine converts this speech into text: "My Internet connection is slow."

[0518] Analysis and classification of consultation content

[0519] The text data is sent to a natural language processing (NLP) engine in the server. The NLP engine analyzes the text data and classifies the inquiry into a specific category. For example, the text data "My Internet connection is slow" would be classified into the "Technical Support" category.

[0520] Allocation to departments

[0521] Based on the classified content, the server automatically determines the appropriate department and specialized operator. For example, if the call is classified as "technical support," the call will be automatically connected to an operator in the technical support department.

[0522] Operator notification

[0523] The server then sends a summary of the user's inquiry and a standard response to the operator's terminal in the selected department. The terminal displays this information in a pop-up format, allowing the operator to respond quickly and appropriately. For example, in response to a question about a slow internet connection, connection troubleshooting procedures and general solutions are displayed.

[0524] User Feedback

[0525] While waiting for the connection to the appropriate department, the server will notify the user of the current processing status by voice. For example, it may say, "Your call is now being transferred to the technical support department. Please wait a moment." This voice feedback allows the user to understand the processing status and reduces anxiety while waiting.

[0526] Specific examples

[0527] When a user calls a call center complaining of an incorrect credit card charge, the system works as follows:

[0528] 1. A user calls and says, "There's an incorrect charge on my credit card."

[0529] 2. The server captures the voice data and uses a speech recognition engine to convert it into text data such as "The credit card charge is incorrect."

[0530] 3. The text data is sent to the NLP engine on the server and classified as a "billing inquiry."

[0531] 4. The server automatically connects to an operator in the billing department.

[0532] 5. The server sends a summary of the user's inquiry and an example of an appropriate response to the operator terminal in the billing department, which then displays this information in a pop-up window.

[0533] 6. The server will notify the user by voice, "Your call will now be forwarded to the billing department. Please wait a moment."

[0534] 7. Operators will respond promptly and appropriately based on the summary and standard response examples.

[0535] In this way, the system of the present invention utilizes voice recognition and natural language processing to reduce the burden on users and improve the efficiency of call center operations.

[0536] The processing flow will be explained below.

[0537] Step 1:

[0538] A user calls a call center.

[0539] Step 2:

[0540] The server runs speech recognition software to capture the user's speech in real time.

[0541] Step 3:

[0542] The server temporarily stores the captured audio data as streaming.

[0543] Step 4:

[0544] The server sends the voice data to the voice recognition engine.

[0545] Step 5:

[0546] The server uses a voice recognition engine to convert the voice data into text data.

[0547] Step 6:

[0548] The server receives the converted text data.

[0549] Step 7:

[0550] The server sends the text data to a natural language processing (NLP) engine.

[0551] Step 8:

[0552] The server analyzes the text data using an NLP engine and classifies the consultation content into specific categories.

[0553] Step 9:

[0554] Based on the analysis results, the server automatically determines the appropriate department and specialized operator to handle the situation.

[0555] Step 10:

[0556] The server transmits the consultation content to the determined department.

[0557] Step 11:

[0558] The terminal (operator's terminal) displays a summary of the received consultation and a standard response example in a pop-up format.

[0559] Step 12:

[0560] The server notifies the user of the current processing status using a speech synthesis engine.

[0561] Step 13:

[0562] Users receive notifications and are informed of the progress of the process.

[0563] Step 14:

[0564] The operator will respond quickly based on the summary and standard response examples displayed on the terminal.

[0565] Step 15:

[0566] Once the operator has completed the response, the results are fed back to the server.

[0567] Specific examples

[0568] A detailed process flow when a user contacts a call center complaining about a slow internet connection:

[0569] Step 1:

[0570] A user says, "My internet connection is slow."

[0571] Step 2:

[0572] The server launches the speech recognition software and captures the audio.

[0573] Step 3:

[0574] The server temporarily stores the captured audio as streaming.

[0575] Step 4:

[0576] The server sends the voice data to the voice recognition engine.

[0577] Step 5:

[0578] The server converts the voice data into text data saying "Your internet connection is slow."

[0579] Step 6:

[0580] The server receives the text data.

[0581] Step 7:

[0582] The server sends the text data to the NLP engine.

[0583] Step 8:

[0584] The server analyzes it using an NLP engine and classifies it as "Technical Support."

[0585] Step 9:

[0586] The server decides to assign it to the "Technical Support" department.

[0587] Step 10:

[0588] The server sends the text data to an operator in the technical support department.

[0589] Step 11:

[0590] The device will pop up a summary that reads "Your internet connection is slow" along with a standard example response.

[0591] Step 12:

[0592] The server will notify the user by voice, "Your call will now be transferred to the technical support department. Please wait a moment."

[0593] Step 13:

[0594] Ensure the user is "transferred to technical support."

[0595] Step 14:

[0596] An operator will quickly guide you through the steps to resolving the problem based on a summary and standard answers.

[0597] Step 15:

[0598] Once the operator has completed the response, the results are fed back to the server.

[0599] Example 1

[0600] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0601] Traditional call centers lacked the technology to efficiently handle user inquiries, and in many cases, operators had to respond manually. This resulted in slow response times and lower user satisfaction. Furthermore, there were sometimes delays in assigning inquiries to the appropriate department, further extending the inquiry time. Furthermore, users could not check the current processing status while waiting, which left them with no way to alleviate their anxiety while waiting.

[0602] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0603] In this invention, the server includes means for receiving voice input, means for converting the voice input into text data, means for analyzing the text data and classifying the consultation content, means for automatically determining a department to handle the consultation based on the consultation content, means for transmitting the consultation content to the department, means for displaying a summary of the consultation content and a standard answer example to the department, and means for providing voice feedback on the processing status to the user. This makes it possible to efficiently distribute inquiries from users to the departments to handle them quickly. Furthermore, the feedback to the user can reduce anxiety while waiting and improve user satisfaction.

[0604] "Means for receiving audio input" refers to equipment or software that captures audio uttered by a user through an audio input device such as a telephone or microphone as digital audio data.

[0605] "Means for converting voice input into text data" refers to a voice recognition engine or software that analyzes captured voice data and converts it into corresponding text data.

[0606] "Means for analyzing text data and classifying consultation content" refers to a system or software that uses a natural language processing engine to analyze text data obtained through voice recognition and classify the consultation content into specific categories.

[0607] "Means for automatically determining the appropriate department based on the content of the consultation" refers to a system or algorithm for selecting the appropriate department or specialized operator based on the classified content of the consultation.

[0608] The "means for transmitting the consultation content to the department in charge" refers to a communication means or protocol for electronically transmitting the user's consultation content to the determined department or operator in charge.

[0609] "Means for displaying a summary of the consultation content and a standard response example to the relevant department" refers to software for summarizing the consultation content received and displaying corresponding standard response examples on the operator terminal.

[0610] The "means for providing the user with voice feedback on the processing status" refers to a voice synthesis engine or system that notifies the user of the current processing status while the user is waiting as a voice message.

[0611] The present invention relates to a system for improving the efficiency of inquiries at a call center by using a voice recognition function. This system includes means for receiving voice input, means for converting the voice input into text data, means for analyzing the text data and classifying the inquiry content, means for automatically determining the department to handle the inquiry based on the inquiry content, means for transmitting the inquiry content to the department, means for displaying a summary of the inquiry content and a standard response example to the department, and means for providing voice feedback on the processing status to the user. The following specific hardware and software are used to implement the invention.

[0612] Acquiring and converting voice input

[0613] When a user contacts the call center via telephone or microphone, the server invokes speech recognition software, using a speech recognition engine such as the Google Cloud Speech-to-Text API, to capture the user's voice input in real time and temporarily store it as audio data.

[0614] The server sends the captured voice data to a speech recognition engine, which converts the voice data into corresponding text data, for example, if the user says "My internet connection is slow," the speech recognition engine converts the voice into text data saying "My internet connection is slow."

[0615] Analysis and classification of consultation content

[0616] The converted text data is sent to a natural language processing engine on the server. An NLP engine such as Google Cloud Natural Language API is used here. The NLP engine analyzes the text data and classifies the inquiry into a specific category. For example, text data such as "My internet connection is slow" would be classified into the "technical support" category.

[0617] Allocation to departments

[0618] The server automatically determines the corresponding department and specialized operator based on the category information received from the NLP engine. For example, if the call is classified as "technical support," the server will automatically connect to an operator in the technical support department.

[0619] Operator notification

[0620] The server then sends a summary of the user's inquiry and a standard response to the operator's terminal in the selected department. The user's inquiry is summarized and a standard response is displayed on the operator's terminal. For example, if the user says "my internet connection is slow," connection troubleshooting procedures and general solutions are displayed in a pop-up format on the terminal.

[0621] User Feedback

[0622] While waiting for the call to be connected to the appropriate department, the server will notify the user of the current processing status by voice. This uses speech synthesis software such as Google Cloud Text-to-Speech API. For example, the server may say, "Your call is now being transferred to the technical support department. Please wait a moment." This voice feedback allows the user to understand the processing status and reduces anxiety while waiting.

[0623] Specific examples

[0624] When a user reports an incorrect charge on their credit card, the system works as follows:

[0625] 1. A user calls and says, "There's an incorrect charge on my credit card."

[0626] 2. The server captures the voice data and uses a speech recognition engine to convert it into text data such as "The credit card charge is incorrect."

[0627] 3. The text data is sent to the NLP engine on the server and classified as a "billing inquiry."

[0628] 4. The server automatically connects to an operator in the billing department.

[0629] 5. The server sends a summary of the user's inquiry and an example of an appropriate response to the operator terminal in the billing department, which then displays this information in a pop-up window.

[0630] 6. The server will notify the user by voice, "Your call will now be forwarded to the billing department. Please wait a moment."

[0631] 7. Operators will respond promptly and appropriately based on the summary and standard response examples.

[0632] Prompt Sentence Examples

[0633] Examples of prompts to be input to a generative AI model:

[0634] "Please explain how this system responds when a user contacts us with a credit card billing issue."

[0635] In this way, the system of the present invention utilizes voice recognition and natural language processing to reduce the burden on users and improve the efficiency of call center operations.

[0636] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0637] Step 1:

[0638] When a user contacts the call center via telephone or microphone, the server activates voice recognition software and captures the user's voice input in real time. The input is the user's voice data, and the output is the captured digital voice data. Specifically, the server obtains the voice data using the Google Cloud Speech-to-Text API or similar.

[0639] Step 2:

[0640] The server sends the captured voice data to the speech recognition engine. The input is the captured digital voice data, and the output is text data. Specifically, the server sends the voice data to the Google Cloud Speech-to-Text API via an HTTP request and receives the text data as a response from the API.

[0641] Step 3:

[0642] The converted text data is sent to the natural language processing engine on the server. The input is text data from the speech recognition engine, and the output is category information for the consultation content. Specifically, the server sends the text data to the Google Cloud Natural Language API and receives category information as a response from the API.

[0643] Step 4:

[0644] The server automatically determines the appropriate department and specialized operator based on the category information received from the NLP engine. The input is the category information of the consultation content, and the output is information on the corresponding department and operator. Specifically, the server references an internal database to map the category and the corresponding department.

[0645] Step 5:

[0646] The server then sends a summary of the user's consultation and a standard example response to the operator terminal of the selected department. The input is information about the department and operator and text data about the consultation, and the output is a summary and example response that are displayed on the operator terminal. Specifically, the server sends the consultation data to the operator terminal as an HTTP request, and the terminal displays this data in a pop-up.

[0647] Step 6:

[0648] While waiting for the corresponding department to connect, the server notifies the user of the current processing status by voice. The input is text data about the current processing status, and the output is synthesized voice data. Specifically, the server uses the Google Cloud Text-to-Speech API to convert the text data into voice data and play it back to the user.

[0649] (Application example 1)

[0650] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0651] Virtual stores require a means to quickly and appropriately respond to inquiries and questions users have about a wide variety of products. However, conventional text-based or operator-based responses can be slow, and a lack of human resources can lead to poor user experience. For this reason, there is a need for technology that uses voice input to automatically analyze inquiries and quickly provide appropriate information.

[0652] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0653] In this invention, the server includes means for receiving voice input, means for converting the voice input into text data, and means for analyzing the text data and classifying the consultation content, thereby making it possible to obtain product information from the database based on the classification results and to respond to the user with the obtained product information in voice and text.

[0654] "Voice input" refers to the device capturing voice uttered by the user.

[0655] "Converting to text data" refers to the process of converting voice input into written information.

[0656] "Analyzing text data" means understanding the content of converted text data and identifying its meaning and intent.

[0657] "Classifying the consultation content" means dividing the content into specific categories or groups based on the analyzed text data.

[0658] "Automatically determining the department to handle the inquiry" means automatically selecting the most appropriate department and person in charge based on the classified inquiry content.

[0659] "Sending the consultation content to the department in charge" means transferring the consultation content and its information to the determined department or person in charge.

[0660] "Displaying a summary of the consultation content and a standard response example" means displaying a concise summary of the consultation content and a recommended response on the operator's terminal.

[0661] "Acquiring product information from a database based on the classification result" means searching the database for product information related to the classified consultation content and acquiring it.

[0662] "Responding to the user with the acquired product information by voice and text" means returning the acquired product information to the user as voice and text information.

[0663] The present invention relates to an application for improving the efficiency of customer support within a virtual store. The system can be implemented in the following forms.

[0664] Hardware and Software Configuration

[0665] Hardware: Smartphone (including microphone, speaker, and display)

[0666] Software: Speech recognition engine (Google Speech-to-Text API), natural language processing engine (Google Cloud Natural Language API), database server (Google Cloud Firestore), application (compatible with Android / iOS)

[0667] Program processing flow

[0668] 1. Acquiring voice input

[0669] The user launches the smartphone app and makes a voice inquiry. The smartphone app captures the user's voice through the microphone and temporarily stores it as voice data.

[0670] 2. Audio data conversion

[0671] The server converts the acquired voice data into text data using the Google Speech-to-Text API. For example, if a user asks, "What size is this dress?", the voice data is converted directly into text information.

[0672] 3. Analysis and Classification of Text Data

[0673] The converted text data is analyzed using the Google Cloud Natural Language API on the server and classified according to the inquiry content, for example, "inquiry about size."

[0674] 4. Determining Corresponding Information

[0675] Based on the classified inquiry, the server retrieves related product information from a database (Google Cloud Firestore), such as "This dress comes in sizes S, M, L, and XL."

[0676] 5. User Feedback

[0677] The acquired product information is returned to the user in voice and text format. The smartphone application provides voice feedback through the speaker and displays text on the display. For example, "This dress is available in sizes S, M, L, and XL" is displayed along with a voice response.

[0678] 6. Notify Operator (if necessary)

[0679] If a complex inquiry is made, the server will escalate the request to a specialist operator who can provide more detailed information and provide an appropriate response.

[0680] Specific examples

[0681] If a user asks "What size is this dress?" in a virtual store, the system works as follows:

[0682] 1. The user makes a voice inquiry through a smartphone app asking, "What size is this dress?"

[0683] 2. The server captures the voice data and converts it into text data such as "What size is this dress?" using the Google Speech-to-Text API.

[0684] 3. The server analyzes the text data using the Google Cloud Natural Language API and classifies it as a "size query."

[0685] 4. Based on the classification results, the server retrieves related product information from Google Cloud Firestore, such as "This dress comes in sizes S, M, L, and XL."

[0686] 5. The server responds to the user with the acquired product information in voice and text format. The smartphone app displays the text on the screen and simultaneously responds with voice.

[0687] Example prompts for natural language generation AI models

[0688] markdown

[0689] Prompt: In a virtual shopping assistant, if a user asks "What size is this dress?", generate a prompt to respond with the appropriate size information to provide to the user.

[0690] Examples:

[0691] markdown

[0692] User input: "What size is this dress?"

[0693] Sizing Information: "This dress comes in sizes S, M, L, and XL."

[0694] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0695] Step 1:

[0696] Acquiring voice input

[0697] The user launches the smartphone app and inputs the inquiry by voice. The smartphone is equipped with a microphone, which captures the user's voice. The input voice is temporarily saved in the smartphone application.

[0698] Input: User voice input

[0699] Output: Audio data

[0700] Step 2:

[0701] Audio data conversion

[0702] The server receives voice data sent from the smartphone. The voice data is sent to the Google Speech-to-Text API and converted into text data in real time. For example, voice data such as "What size is this dress?" is converted into text data such as "What size is this dress?"

[0703] Input: Audio data

[0704] Output: Text data

[0705] Step 3:

[0706] Text data analysis and classification

[0707] The server then sends the converted text data to the Google Cloud Natural Language API, which analyzes the text data and classifies the inquiry into a specific category. For example, the text data "What size is this dress?" would be classified as a "size inquiry."

[0708] Input: Text data

[0709] Output: Classification results

[0710] Step 4:

[0711] Determining correspondence information

[0712] The server searches and retrieves related product information from a database (Google Cloud Firestore) based on the classified inquiry. For example, for an inquiry about size, it retrieves the size information of the specified product.

[0713] Input: Classification results

[0714] Output: Product information data

[0715] Step 5:

[0716] User Feedback

[0717] The server generates the acquired product information as a voice response and a text response. The generated voice data is fed back to the user through the smartphone's speaker, and the text data is displayed on the smartphone's display. For example, the server responds by voice, saying, "This dress comes in sizes S, M, L, and XL," and the same content is displayed on the display.

[0718] Input: Product information data

[0719] Output: Audio and text responses

[0720] Step 6:

[0721] Notify the operator (if necessary)

[0722] If the inquiry is complex and requires a detailed response, the server escalates the inquiry to an appropriate operator, who then sends a summary of the inquiry and a recommended response to the operator's terminal.

[0723] Input: Classification results and product information data

[0724] Output: Notification to operator (summary and standard response example)

[0725] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0726] The present invention relates to a system for improving the efficiency of inquiries handled at a call center by using a voice recognition function and an emotion engine. The system can be implemented in the following forms.

[0727] Acquiring and converting voice input

[0728] When a user calls the call center, the server launches speech recognition software to capture the user's voice input in real time. The server temporarily stores this voice data and sends it to the speech recognition engine. The speech recognition engine converts the voice data into text data. In this case, if the user says, "My Internet connection is slow," the speech recognition engine converts this speech into text: "My Internet connection is slow."

[0729] Emotion analysis

[0730] The converted text data is sent to the server's natural language processing (NLP) engine and also to the emotion engine, which analyzes the user's emotions from the voice data and identifies emotional states such as anger, joy, and sadness. This emotion analysis information is used for subsequent processing.

[0731] Analysis and classification of consultation content

[0732] The NLP engine analyzes the text data and classifies the inquiry into a specific category. For example, the text data "My internet connection is slow" will be classified into the "Technical Support" category.

[0733] Allocation to departments

[0734] The server automatically determines the department and specialized operator to respond to the inquiry based on the classified content and the user's emotions obtained from emotion analysis. For example, if the user expresses anger, the server will prioritize connecting the user to an operator who can respond more quickly and professionally.

[0735] Operator notification

[0736] The server then sends a summary of the user's consultation, a standard response example, and the user's emotional state to the operator terminal in the selected department. The terminal displays this information in a pop-up format, allowing the operator to respond quickly and appropriately while taking the user's emotions into consideration. For example, if the message "Internet connection is slow" and the emotional state of "anger" are displayed, the operator can prepare to respond calmly and quickly.

[0737] User Feedback

[0738] While waiting for the connection to the appropriate department, the server will notify the user of the current processing status by voice. For example, it may say, "Your call is now being transferred to the technical support department. Please wait a moment." This voice feedback allows the user to understand the processing status and reduces anxiety while waiting.

[0739] Specific examples

[0740] If a user angrily calls a call center complaining that their credit card was charged incorrectly, the system works like this:

[0741] 1. A user calls and angrily states that there is an error in the charge on their credit card.

[0742] 2. The server captures the voice data and uses a speech recognition engine to convert it into text data such as "The credit card charge is incorrect."

[0743] 3. The text and voice data are sent to the NLP engine and emotion engine on the server. The emotion engine analyzes the user's anger.

[0744] 4. The text data is classified as a "billing inquiry" and the emotional state of "anger" is detected along with the results of the emotion analysis.

[0745] 5. The server automatically connects to an operator in the billing department, and based on the results of emotion analysis, selects an operator who is best able to handle the emotion of "anger."

[0746] 6. The server sends a summary of the user's consultation, their emotional state, and examples of appropriate responses to the operator terminal in the billing department. The terminal displays this information in a pop-up.

[0747] 7. The server will notify the user by voice, "Your call will now be forwarded to the billing department. Please wait."

[0748] 8. The operator responds calmly, quickly, and sensitively based on the summary and the caller's emotional state.

[0749] In this way, the system of the present invention, which combines an emotion engine, not only performs speech recognition and natural language processing but also analyzes and takes into account the user's emotional state, thereby reducing the burden on the user and further improving the efficiency of call center operations.

[0750] The processing flow will be explained below.

[0751] Step 1:

[0752] A user calls a call center.

[0753] Step 2:

[0754] The server runs speech recognition software to capture the user's speech in real time.

[0755] Step 3:

[0756] The server temporarily stores the captured audio data as streaming.

[0757] Step 4:

[0758] The server sends the voice data to the voice recognition engine.

[0759] Step 5:

[0760] The server uses a speech recognition engine to convert the voice data into text data, for example, "Your internet connection is slow."

[0761] Step 6:

[0762] The server then sends the converted text and voice data to an emotion engine, which analyzes the user's emotions. For example, emotions such as "anger" or "dissatisfaction" are detected from voice patterns.

[0763] Step 7:

[0764] At the same time, the server sends the converted text data to a natural language processing (NLP) engine.

[0765] Step 8:

[0766] The server analyzes the text data using an NLP engine and classifies the inquiry into a specific category. For example, the text "My internet connection is slow" is classified into the "Technical Support" category.

[0767] Step 9:

[0768] The server automatically determines the department and specialist operator to respond to the inquiry based on the content of the inquiry and the results of emotion analysis. For example, if emotion analysis detects anger, it will select an operator from the technical support department who can respond quickly.

[0769] Step 10:

[0770] The server sends the consultation content and emotion analysis results to the determined department.

[0771] Step 11:

[0772] The terminal (operator's terminal) displays a summary of the received consultation, a standard response example, and analyzed emotional information in a pop-up format. For example, it displays the summary "My internet connection is slow" and the emotional state "Anger."

[0773] Step 12:

[0774] The server uses a speech synthesis engine to notify the user of the current status of the process, for example, "Your call is being transferred to technical support. Please wait."

[0775] Step 13:

[0776] Users receive notifications and are informed of the progress of the process.

[0777] Step 14:

[0778] The operator responds quickly and appropriately based on the summary and emotional information displayed on the terminal, for example, by speaking calmly and politely to alleviate the user's anger and try to resolve the technical problem.

[0779] Step 15:

[0780] Once the operator has completed the response, the results are fed back to the server.

[0781] Specific examples

[0782] A detailed process flow when a user contacts a call center because they are angry about an incorrect credit card charge:

[0783] Step 1:

[0784] A user says, "There's an incorrect charge on my credit card."

[0785] Step 2:

[0786] The server launches the speech recognition software and captures the audio.

[0787] Step 3:

[0788] The server temporarily stores the captured audio as streaming.

[0789] Step 4:

[0790] The server sends the voice data to the voice recognition engine.

[0791] Step 5:

[0792] The server converts the voice data into text data saying "The credit card billing is incorrect."

[0793] Step 6:

[0794] The server sends the text data and voice data to the emotion engine, which then analyzes the user's emotion of "anger."

[0795] Step 7:

[0796] The server sends the text data to the NLP engine.

[0797] Step 8:

[0798] The server analyzes the text data using an NLP engine and classifies it as a "billing inquiry."

[0799] Step 9:

[0800] The server determines the department to handle the inquiry based on the emotion of "invoice inquiry" and "anger." For example, it may give priority to an operator in the billing department who can respond quickly.

[0801] Step 10:

[0802] The server sends the consultation details and emotion analysis results to the billing department operator.

[0803] Step 11:

[0804] The device will pop up a summary saying "Credit card charge incorrect" and an emotional state of "Anger."

[0805] Step 12:

[0806] The server will announce, "Your call is now being forwarded to the billing department. Please wait a moment."

[0807] Step 13:

[0808] Verify that the user is "transferred to the billing department."

[0809] Step 14:

[0810] The operator responds calmly and quickly based on the summary and emotional information, for example, by first apologizing to calm the user's anger and then confirming the details of the problem.

[0811] Step 15:

[0812] Once the operator has completed the response, the results are fed back to the server.

[0813] In this way, the system of the present invention, which combines an emotion engine, not only performs speech recognition and natural language processing but also analyzes and takes into account the user's emotional state, thereby reducing the burden on the user and further improving the efficiency of call center operations.

[0814] Example 2

[0815] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0816] Handling inquiries at call centers requires listening to voice input, analyzing content, analyzing emotions, assigning calls to the appropriate department, and providing information to operators promptly. Conventional systems have difficulty performing these tasks efficiently, resulting in problems such as reduced user satisfaction and reduced operator work efficiency. The present invention aims to provide a system that solves these problems and enables efficient call center operation.

[0817] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0818] In this invention, the server includes means for receiving voice input, means for converting the voice input into text data, means for analyzing emotions from the text data and voice data, means for analyzing and classifying the consultation content, means for automatically determining a department to handle the consultation based on the consultation content and emotion analysis results, means for sending a summary of the consultation content and a standard response example to the department to handle the consultation, and means for displaying appropriate information on an operator terminal. This makes it possible to analyze the user's voice input in real time and quickly transfer the consultation to the appropriate department or operator.

[0819] A "means for receiving voice input" is a device or program that has the function of capturing voice from a user and recording it as digital voice data.

[0820] The "means for converting voice input into text data" is a device or program that has the function of analyzing captured voice data and converting it into corresponding text data.

[0821] The "means for analyzing text data and classifying consultation contents" refers to a device or program that has the function of analyzing the contents based on the text data and classifying them into specific categories.

[0822] The "means for analyzing emotions" is a device or program that has the function of identifying the emotional state of a user from voice data and text data.

[0823] The "means for automatically determining the department to handle the call" is a device or program that has the function of selecting the appropriate department or operator to handle the call based on the classified content of the call and the results of emotion analysis.

[0824] The "means for transmitting the consultation content to the department" is a device or program having a function for transmitting the consultation content of the user and related information to the automatically determined department.

[0825] The "means for displaying a summary of the consultation content and a standard response example" is a device or program that has the function of displaying a summary of the consultation content and a standard response example on the operator terminal, thereby supporting a prompt and appropriate response.

[0826] This invention is a system that uses voice recognition and emotion analysis functions to improve the efficiency of inquiries handled at call centers. This system converts user voice input into text data, analyzes the content of the inquiry and the user's emotional state, and automatically assigns the inquiry to the appropriate department and operator.

[0827] To implement this system, the following hardware and software are used.

[0828] Hardware

[0829] server

[0830] Operator terminal (e.g. Windows PC)

[0831] software

[0832] Voice recognition software (e.g., Google Cloud Speech-to-Text)

[0833] Natural Language Processing (NLP) engines (e.g., SpaCy)

[0834] Sentiment analysis engine (e.g. IBM Watson Tone Analyzer)

[0835] Program processing

[0836] Acquiring voice input

[0837] When a user calls the call center, the server captures the user's voice in real time through a telephone line interface and stores it as digital voice data.

[0838] Speech recognition and text conversion

[0839] The server sends the stored voice data to a speech recognition engine, which converts it into text data. For example, if a user says, "My internet connection is slow," the voice data will be converted into text data saying, "My internet connection is slow."

[0840] Emotion analysis

[0841] The converted text and voice data is sent to the emotion engine on the server, which analyzes the data and identifies the user's emotion (e.g., anger, joy, sadness).

[0842] Analysis and classification of consultation content

[0843] The server sends the text data to an NLP engine, which analyzes it and classifies it into a specific category, for example, "My internet connection is slow" would be classified into the "tech support" category.

[0844] Allocation to departments

[0845] The server automatically determines the appropriate department and operator based on the category of the consultation content and the results of emotion analysis. If the user expresses anger, the call will be connected preferentially to an operator who can respond quickly and professionally.

[0846] Operator notification

[0847] The server then sends a summary of the user's consultation and their emotional state to the operator terminal in the selected department. The terminal displays this information in a pop-up format, allowing the operator to respond promptly and appropriately.

[0848] User Feedback

[0849] While waiting for the connection to the appropriate department, the server will notify the user of the current processing status by voice. For example, by guiding the user to "Your call is now being transferred to the technical support department. Please wait a moment," the user can understand the processing status and reduce anxiety while waiting.

[0850] Specific examples

[0851] If a user angrily calls a call center complaining that their credit card was charged incorrectly, the system works like this:

[0852] 1. A user calls and angrily states that there is an error in the charge on their credit card.

[0853] 2. The server captures the voice data and uses a speech recognition engine to convert it into text data such as "The credit card charge is incorrect."

[0854] 3. The text and voice data are sent to the NLP engine and emotion engine on the server. The emotion engine analyzes the user's anger.

[0855] 4. The text data is classified as a "billing inquiry" and the emotional state of "anger" is detected along with the results of the emotion analysis.

[0856] 5. The server automatically connects to an operator in the billing department, and based on the results of emotion analysis, selects an operator who is best able to handle the emotion of "anger."

[0857] 6. The server sends a summary of the user's consultation, their emotional state, and examples of appropriate responses to the operator terminal in the billing department. The terminal displays this information in a pop-up.

[0858] 7. The server will notify the user by voice, "Your call will now be forwarded to the billing department. Please wait."

[0859] 8. The operator responds calmly, quickly, and sensitively based on the summary and the caller's emotional state.

[0860] Prompt Sentence Examples

[0861] "A user is expressing anger over a claim that their credit card is incorrectly charged. What's the best approach to calmly respond?"

[0862] "Some users are complaining about slow internet connections. Can you suggest a good solution for this?"

[0863] In this way, the system of the present invention, which combines an emotion engine, not only performs speech recognition and natural language processing but also analyzes and takes into account the user's emotional state, thereby reducing the burden on the user and further improving the efficiency of call center operations.

[0864] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0865] Step 1: Getting voice input

[0866] When the server detects an incoming call from the user, it immediately activates the voice recognition software and captures the user's voice input in real time. The input is an analog voice signal received over the telephone line. The server converts this voice signal into digital voice data and temporarily stores it.

[0867] Specific behavior:

[0868] The user makes a call.

[0869] The telephone line interface implemented in the server detects the incoming call.

[0870] The server receives voice input through a telephone line interface and stores it as digital voice data.

[0871] Input: Analog audio signal obtained through telephone lines.

[0872] Output: Digital audio data.

[0873] Step 2: Speech recognition and text conversion

[0874] The server then sends the stored digital voice data to a speech recognition engine (such as Google Cloud Speech-to-Text) to convert it into text data. The speech recognition engine then analyzes the digital voice data and generates corresponding text data.

[0875] Specific behavior:

[0876] The server calls the API to send the voice data to the voice recognition engine.

[0877] The speech recognition engine analyzes the speech and generates text data.

[0878] The server receives the generated text data.

[0879] Input: Digital audio data.

[0880] Output: Text data.

[0881] Step 3: Sentiment Analysis

[0882] The server sends the generated text data and saved voice data to an emotion engine (such as IBM Watson Tone Analyzer) to analyze the user's emotional state. The emotion engine analyzes the voice and text data to identify the user's emotions (e.g., anger, joy, sadness).

[0883] Specific behavior:

[0884] The server calls the API to send text and voice data to the emotion engine.

[0885] An emotion engine analyzes the data and identifies the user's emotional state.

[0886] The server receives the results of the sentiment analysis.

[0887] Input: text data, audio data.

[0888] Output: Sentiment analysis results (e.g. anger, joy, sadness).

[0889] Step 4: Analysis and classification of consultation content

[0890] The server sends the text data to a natural language processing (NLP) engine, which analyzes it and classifies the query into a specific category. The NLP engine analyzes the text and classifies it into a specific category (e.g., technical support, billing).

[0891] Specific behavior:

[0892] The server calls the API to send the text data to the NLP engine.

[0893] The NLP engine analyzes the text data and identifies categories.

[0894] The server receives the analysis results.

[0895] Input: Text data.

[0896] Output: Classification results (e.g., tech support, billing).

[0897] Step 5: Assign to departments

[0898] The server automatically determines the appropriate department and operator based on the category of the consultation content and the results of emotion analysis. If the user expresses anger, the caller will be connected preferentially to an operator who can provide a prompt and professional response.

[0899] Specific behavior:

[0900] The server refers to the emotion analysis and classification results and runs an algorithm to select the appropriate department or operator.

[0901] The server issues connection instructions to the selected department or operator.

[0902] Input: Sentiment analysis results, classification results.

[0903] Output: Information on the selected response department and operator.

[0904] Step 6: Notify the operator

[0905] The server then sends a summary of the user's consultation and their emotional state to the operator terminal in the selected department. The operator terminal displays this information in a pop-up format, helping to ensure a prompt and appropriate response.

[0906] Specific behavior:

[0907] The server sends a summary of the user's consultation and the results of emotion analysis to the operator terminal.

[0908] The information received by the operator terminal is displayed as a pop-up.

[0909] Input: Summary of consultation content, sentiment analysis results.

[0910] Output: Information displayed on the operator terminal.

[0911] Step 7: User feedback

[0912] While waiting for the call to be connected to the appropriate department, the server will notify the user of the current processing status by voice, for example, "Your call is now being transferred to the technical support department. Please wait a moment."

[0913] Specific behavior:

[0914] The server generates a feedback message for the user.

[0915] The server generates this message as voice data using voice synthesis software and transmits it to the user over a telephone line.

[0916] Input: Current processing status.

[0917] Output: Audio feedback to the user.

[0918] (Application example 2)

[0919] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0920] Customer service in modern brick-and-mortar stores requires a fast and accurate response to customer problems and complaints. However, traditional methods make it difficult to consider customer emotions, making it difficult to increase customer satisfaction. In particular, there is a lack of appropriate responses to emotional customers, which leads to a decline in the service quality of the entire store.

[0921] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0922] In this invention, the server includes a means for receiving voice input, a means for converting the voice input into text data, and a means for analyzing the user's emotional state from the text data and voice data, thereby enabling the server to analyze the customer's emotions in real time, automatically determine an appropriate response method, and notify store staff in a timely manner.

[0923] "Voice input" refers to voice information uttered by a user, and is information acquired by an input device.

[0924] "Text data" is data obtained by converting voice input into text format, and is information expressed as a character string.

[0925] "Consultation content" refers to the subject of the problem or question that the user communicates through voice input.

[0926] The "emotional state" is an emotional state analyzed from the user's voice input, and indicates, for example, anger, joy, sadness, etc.

[0927] The "department" refers to a specific department or person in charge within an organization that is set up to deal with inquiries from users.

[0928] The term "server" refers to a central computer system that receives voice input, converts it into text data, analyzes emotional states, classifies the content of inquiries, and determines which departments should respond.

[0929] "Speech recognition engine" refers to the software modules and algorithms used to convert voice input into text data.

[0930] An "NLP engine" is a software module for natural language processing that analyzes text data to extract meaning and intent.

[0931] "Operator" refers to the staff at the physical store who responds to customers based on the consultation content and emotional state sent from the server.

[0932] "Feedback" refers to the act of notifying the user of the current processing status by voice or the like.

[0933] The present invention relates to a system for improving the efficiency of customer service in a physical store by using a voice recognition function and an emotion engine. This system is implemented in the following manner.

[0934] Acquiring and converting voice input

[0935] The server receives customer voice input directly from the customer service robot in the physical store. This voice input is captured in real time using a microphone connected to the server. Then, the voice data is converted into text data using a speech recognition engine (e.g., Google Speech Recognition API).

[0936] Emotion analysis

[0937] The converted text and voice data are sent to the server's natural language processing (NLP) engine and emotion engine. The emotion engine uses, for example, the Hugging Face Transformers library to analyze the customer's emotions from the voice data into "anger," "happiness," "sadness," etc. This emotion analysis information is used in subsequent processing steps.

[0938] Analysis and classification of consultation content

[0939] The server's NLP engine analyzes the converted text data and classifies the consultation content into specific categories such as "product usage," "product defect," "inventory check," etc. Here, the classification process can be performed using a library such as TextBlob.

[0940] Deciding how to respond

[0941] The server automatically determines how to respond based on the customer's emotional state obtained from emotion analysis and the classification of the inquiry content. For example, if a customer is in a "confused" emotional state and says, "I don't know where the product is," the server will contact the appropriate staff member who can respond quickly.

[0942] Notification to store staff

[0943] The server then sends a summary of the customer's consultation and emotional state, along with examples of appropriate responses, to the terminal of the selected staff member. This information is displayed in a pop-up format on the staff member's terminal, allowing the staff member to respond quickly and appropriately while taking the customer's emotions into consideration.

[0944] User Feedback

[0945] While waiting for the customer to be connected to the staff member, the server notifies the customer of the current processing status by voice. This notification may include, for example, a message such as "We are currently contacting the staff member in charge. Please wait a moment." This allows the customer to understand the processing status and reduces anxiety while waiting.

[0946] Specific examples

[0947] For example, if a customer approaches a customer service robot in a confused manner, saying, "I don't know where my item is," the system works as follows:

[0948] 1. The server captures the voice data and uses a voice recognition engine to convert it into text data such as "I don't know where the product is."

[0949] 2. The text and voice data are sent to the NLP engine and emotion engine on the server. The emotion engine analyzes the customer's "confusion" emotion.

[0950] 3. The text data is classified as a “location inquiry” and the emotional state is detected along with the sentiment analysis results.

[0951] 4. The server automatically selects the staff member most suited to the consultation and sends a notification.

[0952] 5. The server will notify the customer by voice, "We are currently contacting the staff in charge. Please wait a moment."

[0953] 6. Staff respond quickly to distressed customers based on pop-up summary information and emotional state.

[0954] Example prompt sentence:

[0955] For example, for an input such as "I'm having trouble finding the product," sentiment analysis and content classification are performed. In this way, a system that can easily and effectively improve customer service in physical stores is realized.

[0956] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0957] Step 1:

[0958] The server receives customer voice input in real time from the customer service robot in the physical store. Based on this input, the server directly captures voice data using a high-performance microphone. The input is voice data saying, "I don't know where the product is."

[0959] Step 2:

[0960] The server sends the captured voice data to a voice recognition engine and converts it into text data. Here, the Google Speech Recognition API is used. The input is voice data, and the output is text data such as "I don't know where the product is."

[0961] Step 3:

[0962] The server sends the text and voice data to a natural language processing (NLP) engine and an emotion engine. The NLP engine uses a generative AI model (such as the Hugging Face Transformers library) to analyze the text data and classify the inquiry into a specific category. In this case, the input is the text data, and the output is the category "location inquiry."

[0963] Step 4:

[0964] The server uses an emotion engine to analyze the customer's emotional state from the voice data. The emotion engine also uses a generative AI model to analyze the emotion of "confused" from the voice data. The input is the voice data, and the output is the emotional state of "confused."

[0965] Step 5:

[0966] The server automatically determines how to respond based on the classified consultation content and the analyzed emotional state. In this case, a notification is sent to an available staff member for a customer who expresses "confusion" in a "location inquiry." The input is the consultation content and emotional information, and the output is the selection of the most suitable staff member.

[0967] Step 6:

[0968] The server sends a summary of the customer's consultation and emotional state, as well as examples of appropriate responses, to the staff member's terminal. This notification is displayed in a pop-up format on the terminal. The input is customer information and emotional analysis information, and the output is a notification message to the staff member.

[0969] Step 7:

[0970] The server provides the customer with audio feedback on the current processing status. For example, it generates a message such as "We are currently contacting the staff member in charge. Please wait a moment," and notifies the customer via an audio output device. The input is the response decision information, and the output is the audio notification message.

[0971] In this way, a system can be realized in physical stores that analyzes emotions and classifies content from customer voice input, and then quickly responds appropriately.

[0972] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0973] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0974] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0975] [Third embodiment]

[0976] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0977] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0978] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0979] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0980] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0981] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0982] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0983] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0984] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0985] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0986] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0987] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0988] The present invention relates to a system for improving the efficiency of inquiries handled at a call center by using a voice recognition function. The system can be implemented in the following forms.

[0989] Acquiring and converting voice input

[0990] When a user calls the call center, the server launches speech recognition software to capture the user's voice input in real time. The server temporarily stores this voice data and sends it to the speech recognition engine. The speech recognition engine converts the voice data into text data. For example, if a user says, "My Internet connection is slow," the speech recognition engine converts this speech into text: "My Internet connection is slow."

[0991] Analysis and classification of consultation content

[0992] The text data is sent to a natural language processing (NLP) engine in the server. The NLP engine analyzes the text data and classifies the inquiry into a specific category. For example, the text data "My Internet connection is slow" would be classified into the "Technical Support" category.

[0993] Allocation to departments

[0994] Based on the classified content, the server automatically determines the appropriate department and specialized operator. For example, if the call is classified as "technical support," the call will be automatically connected to an operator in the technical support department.

[0995] Operator notification

[0996] The server then sends a summary of the user's inquiry and a standard response to the operator's terminal in the selected department. The terminal displays this information in a pop-up format, allowing the operator to respond quickly and appropriately. For example, in response to a question about a slow internet connection, connection troubleshooting procedures and general solutions are displayed.

[0997] User Feedback

[0998] While waiting for the connection to the appropriate department, the server will notify the user of the current processing status by voice. For example, it may say, "Your call is now being transferred to the technical support department. Please wait a moment." This voice feedback allows the user to understand the processing status and reduces anxiety while waiting.

[0999] Specific examples

[1000] When a user calls a call center complaining of an incorrect credit card charge, the system works as follows:

[1001] 1. A user calls and says, "There's an incorrect charge on my credit card."

[1002] 2. The server captures the voice data and uses a speech recognition engine to convert it into text data such as "The credit card charge is incorrect."

[1003] 3. The text data is sent to the NLP engine on the server and classified as a "billing inquiry."

[1004] 4. The server automatically connects to an operator in the billing department.

[1005] 5. The server sends a summary of the user's inquiry and an example of an appropriate response to the operator terminal in the billing department, which then displays this information in a pop-up window.

[1006] 6. The server will notify the user by voice, "Your call will now be forwarded to the billing department. Please wait a moment."

[1007] 7. Operators will respond promptly and appropriately based on the summary and standard response examples.

[1008] In this way, the system of the present invention utilizes voice recognition and natural language processing to reduce the burden on users and improve the efficiency of call center operations.

[1009] The processing flow will be explained below.

[1010] Step 1:

[1011] A user calls a call center.

[1012] Step 2:

[1013] The server runs speech recognition software to capture the user's speech in real time.

[1014] Step 3:

[1015] The server temporarily stores the captured audio data as streaming.

[1016] Step 4:

[1017] The server sends the voice data to the voice recognition engine.

[1018] Step 5:

[1019] The server uses a voice recognition engine to convert the voice data into text data.

[1020] Step 6:

[1021] The server receives the converted text data.

[1022] Step 7:

[1023] The server sends the text data to a natural language processing (NLP) engine.

[1024] Step 8:

[1025] The server analyzes the text data using an NLP engine and classifies the consultation content into specific categories.

[1026] Step 9:

[1027] Based on the analysis results, the server automatically determines the appropriate department and specialized operator to handle the situation.

[1028] Step 10:

[1029] The server transmits the consultation content to the determined department.

[1030] Step 11:

[1031] The terminal (operator's terminal) displays a summary of the received consultation and a standard response example in a pop-up format.

[1032] Step 12:

[1033] The server notifies the user of the current processing status using a speech synthesis engine.

[1034] Step 13:

[1035] Users receive notifications and are informed of the progress of the process.

[1036] Step 14:

[1037] The operator will respond quickly based on the summary and standard response examples displayed on the terminal.

[1038] Step 15:

[1039] Once the operator has completed the response, the results are fed back to the server.

[1040] Specific examples

[1041] A detailed process flow when a user contacts a call center complaining about a slow internet connection:

[1042] Step 1:

[1043] A user says, "My internet connection is slow."

[1044] Step 2:

[1045] The server launches the speech recognition software and captures the audio.

[1046] Step 3:

[1047] The server temporarily stores the captured audio as streaming.

[1048] Step 4:

[1049] The server sends the voice data to the voice recognition engine.

[1050] Step 5:

[1051] The server converts the voice data into text data saying "Your internet connection is slow."

[1052] Step 6:

[1053] The server receives the text data.

[1054] Step 7:

[1055] The server sends the text data to the NLP engine.

[1056] Step 8:

[1057] The server analyzes it using an NLP engine and classifies it as "Technical Support."

[1058] Step 9:

[1059] The server decides to assign it to the "Technical Support" department.

[1060] Step 10:

[1061] The server sends the text data to an operator in the technical support department.

[1062] Step 11:

[1063] The device will pop up a summary that reads "Your internet connection is slow" along with a standard example response.

[1064] Step 12:

[1065] The server will notify the user by voice, "Your call will now be transferred to the technical support department. Please wait a moment."

[1066] Step 13:

[1067] Ensure the user is "transferred to technical support."

[1068] Step 14:

[1069] An operator will quickly guide you through the steps to resolving the problem based on a summary and standard answers.

[1070] Step 15:

[1071] Once the operator has completed the response, the results are fed back to the server.

[1072] Example 1

[1073] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1074] Traditional call centers lacked the technology to efficiently handle user inquiries, and in many cases, operators had to respond manually. This resulted in slow response times and lower user satisfaction. Furthermore, there were sometimes delays in assigning inquiries to the appropriate department, further extending the inquiry time. Furthermore, users could not check the current processing status while waiting, which left them with no way to alleviate their anxiety while waiting.

[1075] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1076] In this invention, the server includes means for receiving voice input, means for converting the voice input into text data, means for analyzing the text data and classifying the consultation content, means for automatically determining a department to handle the consultation based on the consultation content, means for transmitting the consultation content to the department, means for displaying a summary of the consultation content and a standard answer example to the department, and means for providing voice feedback on the processing status to the user. This makes it possible to efficiently distribute inquiries from users to the departments to handle them quickly. Furthermore, the feedback to the user can reduce anxiety while waiting and improve user satisfaction.

[1077] "Means for receiving audio input" refers to equipment or software that captures audio uttered by a user through an audio input device such as a telephone or microphone as digital audio data.

[1078] "Means for converting voice input into text data" refers to a voice recognition engine or software that analyzes captured voice data and converts it into corresponding text data.

[1079] "Means for analyzing text data and classifying consultation content" refers to a system or software that uses a natural language processing engine to analyze text data obtained through voice recognition and classify the consultation content into specific categories.

[1080] "Means for automatically determining the appropriate department based on the content of the consultation" refers to a system or algorithm for selecting the appropriate department or specialized operator based on the classified content of the consultation.

[1081] The "means for transmitting the consultation content to the department in charge" refers to a communication means or protocol for electronically transmitting the user's consultation content to the determined department or operator in charge.

[1082] "Means for displaying a summary of the consultation content and a standard response example to the relevant department" refers to software for summarizing the consultation content received and displaying corresponding standard response examples on the operator terminal.

[1083] The "means for providing the user with voice feedback on the processing status" refers to a voice synthesis engine or system that notifies the user of the current processing status while the user is waiting as a voice message.

[1084] The present invention relates to a system for improving the efficiency of inquiries at a call center by using a voice recognition function. This system includes means for receiving voice input, means for converting the voice input into text data, means for analyzing the text data and classifying the inquiry content, means for automatically determining the department to handle the inquiry based on the inquiry content, means for transmitting the inquiry content to the department, means for displaying a summary of the inquiry content and a standard response example to the department, and means for providing voice feedback on the processing status to the user. The following specific hardware and software are used to implement the invention.

[1085] Acquiring and converting voice input

[1086] When a user contacts the call center via telephone or microphone, the server invokes speech recognition software, using a speech recognition engine such as the Google Cloud Speech-to-Text API, to capture the user's voice input in real time and temporarily store it as audio data.

[1087] The server sends the captured voice data to a speech recognition engine, which converts the voice data into corresponding text data, for example, if the user says "My internet connection is slow," the speech recognition engine converts the voice into text data saying "My internet connection is slow."

[1088] Analysis and classification of consultation content

[1089] The converted text data is sent to a natural language processing engine on the server. An NLP engine such as Google Cloud Natural Language API is used here. The NLP engine analyzes the text data and classifies the inquiry into a specific category. For example, text data such as "My internet connection is slow" would be classified into the "technical support" category.

[1090] Allocation to departments

[1091] The server automatically determines the corresponding department and specialized operator based on the category information received from the NLP engine. For example, if the call is classified as "technical support," the server will automatically connect to an operator in the technical support department.

[1092] Operator notification

[1093] The server then sends a summary of the user's inquiry and a standard response to the operator's terminal in the selected department. The user's inquiry is summarized and a standard response is displayed on the operator's terminal. For example, if the user says "my internet connection is slow," connection troubleshooting procedures and general solutions are displayed in a pop-up format on the terminal.

[1094] User Feedback

[1095] While waiting for the call to be connected to the appropriate department, the server will notify the user of the current processing status by voice. This uses speech synthesis software such as Google Cloud Text-to-Speech API. For example, the server may say, "Your call is now being transferred to the technical support department. Please wait a moment." This voice feedback allows the user to understand the processing status and reduces anxiety while waiting.

[1096] Specific examples

[1097] When a user reports an incorrect charge on their credit card, the system works as follows:

[1098] 1. A user calls and says, "There's an incorrect charge on my credit card."

[1099] 2. The server captures the voice data and uses a speech recognition engine to convert it into text data such as "The credit card charge is incorrect."

[1100] 3. The text data is sent to the NLP engine on the server and classified as a "billing inquiry."

[1101] 4. The server automatically connects to an operator in the billing department.

[1102] 5. The server sends a summary of the user's inquiry and an example of an appropriate response to the operator terminal in the billing department, which then displays this information in a pop-up window.

[1103] 6. The server will notify the user by voice, "Your call will now be forwarded to the billing department. Please wait a moment."

[1104] 7. Operators will respond promptly and appropriately based on the summary and standard response examples.

[1105] Prompt Sentence Examples

[1106] Examples of prompts to be input to a generative AI model:

[1107] "Please explain how this system responds when a user contacts us with a credit card billing issue."

[1108] In this way, the system of the present invention utilizes voice recognition and natural language processing to reduce the burden on users and improve the efficiency of call center operations.

[1109] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1110] Step 1:

[1111] When a user contacts the call center via telephone or microphone, the server activates voice recognition software and captures the user's voice input in real time. The input is the user's voice data, and the output is the captured digital voice data. Specifically, the server obtains the voice data using the Google Cloud Speech-to-Text API or similar.

[1112] Step 2:

[1113] The server sends the captured voice data to the speech recognition engine. The input is the captured digital voice data, and the output is text data. Specifically, the server sends the voice data to the Google Cloud Speech-to-Text API via an HTTP request and receives the text data as a response from the API.

[1114] Step 3:

[1115] The converted text data is sent to the natural language processing engine on the server. The input is text data from the speech recognition engine, and the output is category information for the consultation content. Specifically, the server sends the text data to the Google Cloud Natural Language API and receives category information as a response from the API.

[1116] Step 4:

[1117] The server automatically determines the appropriate department and specialized operator based on the category information received from the NLP engine. The input is the category information of the consultation content, and the output is information on the corresponding department and operator. Specifically, the server references an internal database to map the category and the corresponding department.

[1118] Step 5:

[1119] The server then sends a summary of the user's consultation and a standard example response to the operator terminal of the selected department. The input is information about the department and operator and text data about the consultation, and the output is a summary and example response that are displayed on the operator terminal. Specifically, the server sends the consultation data to the operator terminal as an HTTP request, and the terminal displays this data in a pop-up.

[1120] Step 6:

[1121] While waiting for the corresponding department to connect, the server notifies the user of the current processing status by voice. The input is text data about the current processing status, and the output is synthesized voice data. Specifically, the server uses the Google Cloud Text-to-Speech API to convert the text data into voice data and play it back to the user.

[1122] (Application example 1)

[1123] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1124] Virtual stores require a means to quickly and appropriately respond to inquiries and questions users have about a wide variety of products. However, conventional text-based or operator-based responses can be slow, and a lack of human resources can lead to poor user experience. For this reason, there is a need for technology that uses voice input to automatically analyze inquiries and quickly provide appropriate information.

[1125] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1126] In this invention, the server includes means for receiving voice input, means for converting the voice input into text data, and means for analyzing the text data and classifying the consultation content, thereby making it possible to obtain product information from the database based on the classification results and to respond to the user with the obtained product information in voice and text.

[1127] "Voice input" refers to the device capturing voice uttered by the user.

[1128] "Converting to text data" refers to the process of converting voice input into written information.

[1129] "Analyzing text data" means understanding the content of converted text data and identifying its meaning and intent.

[1130] "Classifying the consultation content" means dividing the content into specific categories or groups based on the analyzed text data.

[1131] "Automatically determining the department to handle the inquiry" means automatically selecting the most appropriate department and person in charge based on the classified inquiry content.

[1132] "Sending the consultation content to the department in charge" means transferring the consultation content and its information to the determined department or person in charge.

[1133] "Displaying a summary of the consultation content and a standard response example" means displaying a concise summary of the consultation content and a recommended response on the operator's terminal.

[1134] "Acquiring product information from a database based on the classification result" means searching the database for product information related to the classified consultation content and acquiring it.

[1135] "Responding to the user with the acquired product information by voice and text" means returning the acquired product information to the user as voice and text information.

[1136] The present invention relates to an application for improving the efficiency of customer support within a virtual store. The system can be implemented in the following forms.

[1137] Hardware and Software Configuration

[1138] Hardware: Smartphone (including microphone, speaker, and display)

[1139] Software: Speech recognition engine (Google Speech-to-Text API), natural language processing engine (Google Cloud Natural Language API), database server (Google Cloud Firestore), application (compatible with Android / iOS)

[1140] Program processing flow

[1141] 1. Acquiring voice input

[1142] The user launches the smartphone app and makes a voice inquiry. The smartphone app captures the user's voice through the microphone and temporarily stores it as voice data.

[1143] 2. Audio data conversion

[1144] The server converts the acquired voice data into text data using the Google Speech-to-Text API. For example, if a user asks, "What size is this dress?", the voice data is converted directly into text information.

[1145] 3. Analysis and Classification of Text Data

[1146] The converted text data is analyzed using the Google Cloud Natural Language API on the server and classified according to the inquiry content, for example, "inquiry about size."

[1147] 4. Determining Corresponding Information

[1148] Based on the classified inquiry, the server retrieves related product information from a database (Google Cloud Firestore), such as "This dress comes in sizes S, M, L, and XL."

[1149] 5. User Feedback

[1150] The acquired product information is returned to the user in voice and text format. The smartphone application provides voice feedback through the speaker and displays text on the display. For example, "This dress is available in sizes S, M, L, and XL" is displayed along with a voice response.

[1151] 6. Notify Operator (if necessary)

[1152] If a complex inquiry is made, the server will escalate the request to a specialist operator who can provide more detailed information and provide an appropriate response.

[1153] Specific examples

[1154] If a user asks "What size is this dress?" in a virtual store, the system works as follows:

[1155] 1. The user makes a voice inquiry through a smartphone app asking, "What size is this dress?"

[1156] 2. The server captures the voice data and converts it into text data such as "What size is this dress?" using the Google Speech-to-Text API.

[1157] 3. The server analyzes the text data using the Google Cloud Natural Language API and classifies it as a "size query."

[1158] 4. Based on the classification results, the server retrieves related product information from Google Cloud Firestore, such as "This dress comes in sizes S, M, L, and XL."

[1159] 5. The server responds to the user with the acquired product information in voice and text format. The smartphone app displays the text on the screen and simultaneously responds with voice.

[1160] Example prompts for natural language generation AI models

[1161] markdown

[1162] Prompt: In a virtual shopping assistant, if a user asks "What size is this dress?", generate a prompt to respond with the appropriate size information to provide to the user.

[1163] Examples:

[1164] markdown

[1165] User input: "What size is this dress?"

[1166] Sizing Information: "This dress comes in sizes S, M, L, and XL."

[1167] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1168] Step 1:

[1169] Acquiring voice input

[1170] The user launches the smartphone app and inputs the inquiry by voice. The smartphone is equipped with a microphone, which captures the user's voice. The input voice is temporarily saved in the smartphone application.

[1171] Input: User voice input

[1172] Output: Audio data

[1173] Step 2:

[1174] Audio data conversion

[1175] The server receives voice data sent from the smartphone. The voice data is sent to the Google Speech-to-Text API and converted into text data in real time. For example, voice data such as "What size is this dress?" is converted into text data such as "What size is this dress?"

[1176] Input: Audio data

[1177] Output: Text data

[1178] Step 3:

[1179] Text data analysis and classification

[1180] The server then sends the converted text data to the Google Cloud Natural Language API, which analyzes the text data and classifies the inquiry into a specific category. For example, the text data "What size is this dress?" would be classified as a "size inquiry."

[1181] Input: Text data

[1182] Output: Classification results

[1183] Step 4:

[1184] Determining correspondence information

[1185] The server searches and retrieves related product information from a database (Google Cloud Firestore) based on the classified inquiry. For example, for an inquiry about size, it retrieves the size information of the specified product.

[1186] Input: Classification results

[1187] Output: Product information data

[1188] Step 5:

[1189] User Feedback

[1190] The server generates the acquired product information as a voice response and a text response. The generated voice data is fed back to the user through the smartphone's speaker, and the text data is displayed on the smartphone's display. For example, the server responds by voice, saying, "This dress comes in sizes S, M, L, and XL," and the same content is displayed on the display.

[1191] Input: Product information data

[1192] Output: Audio and text responses

[1193] Step 6:

[1194] Notify the operator (if necessary)

[1195] If the inquiry is complex and requires a detailed response, the server escalates the inquiry to an appropriate operator, who then sends a summary of the inquiry and a recommended response to the operator's terminal.

[1196] Input: Classification results and product information data

[1197] Output: Notification to operator (summary and standard response example)

[1198] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1199] The present invention relates to a system for improving the efficiency of inquiries handled at a call center by using a voice recognition function and an emotion engine. The system can be implemented in the following forms.

[1200] Acquiring and converting voice input

[1201] When a user calls the call center, the server launches speech recognition software to capture the user's voice input in real time. The server temporarily stores this voice data and sends it to the speech recognition engine. The speech recognition engine converts the voice data into text data. In this case, if the user says, "My Internet connection is slow," the speech recognition engine converts this speech into text: "My Internet connection is slow."

[1202] Emotion analysis

[1203] The converted text data is sent to the server's natural language processing (NLP) engine and also to the emotion engine, which analyzes the user's emotions from the voice data and identifies emotional states such as anger, joy, and sadness. This emotion analysis information is used for subsequent processing.

[1204] Analysis and classification of consultation content

[1205] The NLP engine analyzes the text data and classifies the inquiry into a specific category. For example, the text data "My internet connection is slow" will be classified into the "Technical Support" category.

[1206] Allocation to departments

[1207] The server automatically determines the department and specialized operator to respond to the inquiry based on the classified content and the user's emotions obtained from emotion analysis. For example, if the user expresses anger, the server will prioritize connecting the user to an operator who can respond more quickly and professionally.

[1208] Operator notification

[1209] The server then sends a summary of the user's consultation, a standard response example, and the user's emotional state to the operator terminal in the selected department. The terminal displays this information in a pop-up format, allowing the operator to respond quickly and appropriately while taking the user's emotions into consideration. For example, if the message "Internet connection is slow" and the emotional state of "anger" are displayed, the operator can prepare to respond calmly and quickly.

[1210] User Feedback

[1211] While waiting for the connection to the appropriate department, the server will notify the user of the current processing status by voice. For example, it may say, "Your call is now being transferred to the technical support department. Please wait a moment." This voice feedback allows the user to understand the processing status and reduces anxiety while waiting.

[1212] Specific examples

[1213] If a user angrily calls a call center complaining that their credit card was charged incorrectly, the system works like this:

[1214] 1. A user calls and angrily states that there is an error in the charge on their credit card.

[1215] 2. The server captures the voice data and uses a speech recognition engine to convert it into text data such as "The credit card charge is incorrect."

[1216] 3. The text and voice data are sent to the NLP engine and emotion engine on the server. The emotion engine analyzes the user's anger.

[1217] 4. The text data is classified as a "billing inquiry" and the emotional state of "anger" is detected along with the results of the emotion analysis.

[1218] 5. The server automatically connects to an operator in the billing department, and based on the results of emotion analysis, selects an operator who is best able to handle the emotion of "anger."

[1219] 6. The server sends a summary of the user's consultation, their emotional state, and examples of appropriate responses to the operator terminal in the billing department. The terminal displays this information in a pop-up.

[1220] 7. The server will notify the user by voice, "Your call will now be forwarded to the billing department. Please wait."

[1221] 8. The operator responds calmly, quickly, and sensitively based on the summary and the caller's emotional state.

[1222] In this way, the system of the present invention, which combines an emotion engine, not only performs speech recognition and natural language processing but also analyzes and takes into account the user's emotional state, thereby reducing the burden on the user and further improving the efficiency of call center operations.

[1223] The processing flow will be explained below.

[1224] Step 1:

[1225] A user calls a call center.

[1226] Step 2:

[1227] The server runs speech recognition software to capture the user's speech in real time.

[1228] Step 3:

[1229] The server temporarily stores the captured audio data as streaming.

[1230] Step 4:

[1231] The server sends the voice data to the voice recognition engine.

[1232] Step 5:

[1233] The server uses a speech recognition engine to convert the voice data into text data, for example, "Your internet connection is slow."

[1234] Step 6:

[1235] The server then sends the converted text and voice data to an emotion engine, which analyzes the user's emotions. For example, emotions such as "anger" or "dissatisfaction" are detected from voice patterns.

[1236] Step 7:

[1237] At the same time, the server sends the converted text data to a natural language processing (NLP) engine.

[1238] Step 8:

[1239] The server analyzes the text data using an NLP engine and classifies the inquiry into a specific category. For example, the text "My internet connection is slow" is classified into the "Technical Support" category.

[1240] Step 9:

[1241] The server automatically determines the department and specialist operator to respond to the inquiry based on the content of the inquiry and the results of emotion analysis. For example, if emotion analysis detects anger, it will select an operator from the technical support department who can respond quickly.

[1242] Step 10:

[1243] The server sends the consultation content and emotion analysis results to the determined department.

[1244] Step 11:

[1245] The terminal (operator's terminal) displays a summary of the received consultation, a standard response example, and analyzed emotional information in a pop-up format. For example, it displays the summary "My internet connection is slow" and the emotional state "Anger."

[1246] Step 12:

[1247] The server uses a speech synthesis engine to notify the user of the current status of the process, for example, "Your call is being transferred to technical support. Please wait."

[1248] Step 13:

[1249] Users receive notifications and are informed of the progress of the process.

[1250] Step 14:

[1251] The operator responds quickly and appropriately based on the summary and emotional information displayed on the terminal, for example, by speaking calmly and politely to alleviate the user's anger and try to resolve the technical problem.

[1252] Step 15:

[1253] Once the operator has completed the response, the results are fed back to the server.

[1254] Specific examples

[1255] A detailed process flow when a user contacts a call center because they are angry about an incorrect credit card charge:

[1256] Step 1:

[1257] A user says, "There's an incorrect charge on my credit card."

[1258] Step 2:

[1259] The server launches the speech recognition software and captures the audio.

[1260] Step 3:

[1261] The server temporarily stores the captured audio as streaming.

[1262] Step 4:

[1263] The server sends the voice data to the voice recognition engine.

[1264] Step 5:

[1265] The server converts the voice data into text data saying "The credit card billing is incorrect."

[1266] Step 6:

[1267] The server sends the text data and voice data to the emotion engine, which then analyzes the user's emotion of "anger."

[1268] Step 7:

[1269] The server sends the text data to the NLP engine.

[1270] Step 8:

[1271] The server analyzes the text data using an NLP engine and classifies it as a "billing inquiry."

[1272] Step 9:

[1273] The server determines the department to handle the inquiry based on the emotion of "invoice inquiry" and "anger." For example, it may give priority to an operator in the billing department who can respond quickly.

[1274] Step 10:

[1275] The server sends the consultation details and emotion analysis results to the billing department operator.

[1276] Step 11:

[1277] The device will pop up a summary saying "Credit card charge incorrect" and an emotional state of "Anger."

[1278] Step 12:

[1279] The server will announce, "Your call is now being forwarded to the billing department. Please wait a moment."

[1280] Step 13:

[1281] Verify that the user is "transferred to the billing department."

[1282] Step 14:

[1283] The operator responds calmly and quickly based on the summary and emotional information, for example, by first apologizing to calm the user's anger and then confirming the details of the problem.

[1284] Step 15:

[1285] Once the operator has completed the response, the results are fed back to the server.

[1286] In this way, the system of the present invention, which combines an emotion engine, not only performs speech recognition and natural language processing but also analyzes and takes into account the user's emotional state, thereby reducing the burden on the user and further improving the efficiency of call center operations.

[1287] Example 2

[1288] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1289] Handling inquiries at call centers requires listening to voice input, analyzing content, analyzing emotions, assigning calls to the appropriate department, and providing information to operators promptly. Conventional systems have difficulty performing these tasks efficiently, resulting in problems such as reduced user satisfaction and reduced operator work efficiency. The present invention aims to provide a system that solves these problems and enables efficient call center operation.

[1290] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1291] In this invention, the server includes means for receiving voice input, means for converting the voice input into text data, means for analyzing emotions from the text data and voice data, means for analyzing and classifying the consultation content, means for automatically determining a department to handle the consultation based on the consultation content and emotion analysis results, means for sending a summary of the consultation content and a standard response example to the department to handle the consultation, and means for displaying appropriate information on an operator terminal. This makes it possible to analyze the user's voice input in real time and quickly transfer the consultation to the appropriate department or operator.

[1292] A "means for receiving voice input" is a device or program that has the function of capturing voice from a user and recording it as digital voice data.

[1293] The "means for converting voice input into text data" is a device or program that has the function of analyzing captured voice data and converting it into corresponding text data.

[1294] The "means for analyzing text data and classifying consultation contents" refers to a device or program that has the function of analyzing the contents based on the text data and classifying them into specific categories.

[1295] The "means for analyzing emotions" is a device or program that has the function of identifying the emotional state of a user from voice data and text data.

[1296] The "means for automatically determining the department to handle the call" is a device or program that has the function of selecting the appropriate department or operator to handle the call based on the classified content of the call and the results of emotion analysis.

[1297] The "means for transmitting the consultation content to the department" is a device or program having a function for transmitting the consultation content of the user and related information to the automatically determined department.

[1298] The "means for displaying a summary of the consultation content and a standard response example" is a device or program that has the function of displaying a summary of the consultation content and a standard response example on the operator terminal, thereby supporting a prompt and appropriate response.

[1299] This invention is a system that uses voice recognition and emotion analysis functions to improve the efficiency of inquiries handled at call centers. This system converts user voice input into text data, analyzes the content of the inquiry and the user's emotional state, and automatically assigns the inquiry to the appropriate department and operator.

[1300] To implement this system, the following hardware and software are used.

[1301] Hardware

[1302] server

[1303] Operator terminal (e.g. Windows PC)

[1304] software

[1305] Voice recognition software (e.g., Google Cloud Speech-to-Text)

[1306] Natural Language Processing (NLP) engines (e.g., SpaCy)

[1307] Sentiment analysis engine (e.g. IBM Watson Tone Analyzer)

[1308] Program processing

[1309] Acquiring voice input

[1310] When a user calls the call center, the server captures the user's voice in real time through a telephone line interface and stores it as digital voice data.

[1311] Speech recognition and text conversion

[1312] The server sends the stored voice data to a speech recognition engine, which converts it into text data. For example, if a user says, "My internet connection is slow," the voice data will be converted into text data saying, "My internet connection is slow."

[1313] Emotion analysis

[1314] The converted text and voice data is sent to the emotion engine on the server, which analyzes the data and identifies the user's emotion (e.g., anger, joy, sadness).

[1315] Analysis and classification of consultation content

[1316] The server sends the text data to an NLP engine, which analyzes it and classifies it into a specific category, for example, "My internet connection is slow" would be classified into the "tech support" category.

[1317] Allocation to departments

[1318] The server automatically determines the appropriate department and operator based on the category of the consultation content and the results of emotion analysis. If the user expresses anger, the call will be connected preferentially to an operator who can respond quickly and professionally.

[1319] Operator notification

[1320] The server then sends a summary of the user's consultation and their emotional state to the operator terminal in the selected department. The terminal displays this information in a pop-up format, allowing the operator to respond promptly and appropriately.

[1321] User Feedback

[1322] While waiting for the connection to the appropriate department, the server will notify the user of the current processing status by voice. For example, by guiding the user to "Your call is now being transferred to the technical support department. Please wait a moment," the user can understand the processing status and reduce anxiety while waiting.

[1323] Specific examples

[1324] If a user angrily calls a call center complaining that their credit card was charged incorrectly, the system works like this:

[1325] 1. A user calls and angrily states that there is an error in the charge on their credit card.

[1326] 2. The server captures the voice data and uses a speech recognition engine to convert it into text data such as "The credit card charge is incorrect."

[1327] 3. The text and voice data are sent to the NLP engine and emotion engine on the server. The emotion engine analyzes the user's anger.

[1328] 4. The text data is classified as a "billing inquiry" and the emotional state of "anger" is detected along with the results of the emotion analysis.

[1329] 5. The server automatically connects to an operator in the billing department, and based on the results of emotion analysis, selects an operator who is best able to handle the emotion of "anger."

[1330] 6. The server sends a summary of the user's consultation, their emotional state, and examples of appropriate responses to the operator terminal in the billing department. The terminal displays this information in a pop-up.

[1331] 7. The server will notify the user by voice, "Your call will now be forwarded to the billing department. Please wait."

[1332] 8. The operator responds calmly, quickly, and sensitively based on the summary and the caller's emotional state.

[1333] Prompt Sentence Examples

[1334] "A user is expressing anger over a claim that their credit card is incorrectly charged. What's the best approach to calmly respond?"

[1335] "Some users are complaining about slow internet connections. Can you suggest a good solution for this?"

[1336] In this way, the system of the present invention, which combines an emotion engine, not only performs speech recognition and natural language processing but also analyzes and takes into account the user's emotional state, thereby reducing the burden on the user and further improving the efficiency of call center operations.

[1337] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1338] Step 1: Getting voice input

[1339] When the server detects an incoming call from the user, it immediately activates the voice recognition software and captures the user's voice input in real time. The input is an analog voice signal received over the telephone line. The server converts this voice signal into digital voice data and temporarily stores it.

[1340] Specific behavior:

[1341] The user makes a call.

[1342] The telephone line interface implemented in the server detects the incoming call.

[1343] The server receives voice input through a telephone line interface and stores it as digital voice data.

[1344] Input: Analog audio signal obtained through telephone lines.

[1345] Output: Digital audio data.

[1346] Step 2: Speech recognition and text conversion

[1347] The server then sends the stored digital voice data to a speech recognition engine (such as Google Cloud Speech-to-Text) to convert it into text data. The speech recognition engine then analyzes the digital voice data and generates corresponding text data.

[1348] Specific behavior:

[1349] The server calls the API to send the voice data to the voice recognition engine.

[1350] The speech recognition engine analyzes the speech and generates text data.

[1351] The server receives the generated text data.

[1352] Input: Digital audio data.

[1353] Output: Text data.

[1354] Step 3: Sentiment Analysis

[1355] The server sends the generated text data and saved voice data to an emotion engine (such as IBM Watson Tone Analyzer) to analyze the user's emotional state. The emotion engine analyzes the voice and text data to identify the user's emotions (e.g., anger, joy, sadness).

[1356] Specific behavior:

[1357] The server calls the API to send text and voice data to the emotion engine.

[1358] An emotion engine analyzes the data and identifies the user's emotional state.

[1359] The server receives the results of the sentiment analysis.

[1360] Input: text data, audio data.

[1361] Output: Sentiment analysis results (e.g. anger, joy, sadness).

[1362] Step 4: Analysis and classification of consultation content

[1363] The server sends the text data to a natural language processing (NLP) engine, which analyzes it and classifies the query into a specific category. The NLP engine analyzes the text and classifies it into a specific category (e.g., technical support, billing).

[1364] Specific behavior:

[1365] The server calls the API to send the text data to the NLP engine.

[1366] The NLP engine analyzes the text data and identifies categories.

[1367] The server receives the analysis results.

[1368] Input: Text data.

[1369] Output: Classification results (e.g., tech support, billing).

[1370] Step 5: Assign to departments

[1371] The server automatically determines the appropriate department and operator based on the category of the consultation content and the results of emotion analysis. If the user expresses anger, the caller will be connected preferentially to an operator who can provide a prompt and professional response.

[1372] Specific behavior:

[1373] The server refers to the emotion analysis and classification results and runs an algorithm to select the appropriate department or operator.

[1374] The server issues connection instructions to the selected department or operator.

[1375] Input: Sentiment analysis results, classification results.

[1376] Output: Information on the selected response department and operator.

[1377] Step 6: Notify the operator

[1378] The server then sends a summary of the user's consultation and their emotional state to the operator terminal in the selected department. The operator terminal displays this information in a pop-up format, helping to ensure a prompt and appropriate response.

[1379] Specific behavior:

[1380] The server sends a summary of the user's consultation and the results of emotion analysis to the operator terminal.

[1381] The information received by the operator terminal is displayed as a pop-up.

[1382] Input: Summary of consultation content, sentiment analysis results.

[1383] Output: Information displayed on the operator terminal.

[1384] Step 7: User feedback

[1385] While waiting for the call to be connected to the appropriate department, the server will notify the user of the current processing status by voice, for example, "Your call is now being transferred to the technical support department. Please wait a moment."

[1386] Specific behavior:

[1387] The server generates a feedback message for the user.

[1388] The server generates this message as voice data using voice synthesis software and transmits it to the user over a telephone line.

[1389] Input: Current processing status.

[1390] Output: Audio feedback to the user.

[1391] (Application example 2)

[1392] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1393] Customer service in modern brick-and-mortar stores requires a fast and accurate response to customer problems and complaints. However, traditional methods make it difficult to consider customer emotions, making it difficult to increase customer satisfaction. In particular, there is a lack of appropriate responses to emotional customers, which leads to a decline in the service quality of the entire store.

[1394] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1395] In this invention, the server includes a means for receiving voice input, a means for converting the voice input into text data, and a means for analyzing the user's emotional state from the text data and voice data, thereby enabling the server to analyze the customer's emotions in real time, automatically determine an appropriate response method, and notify store staff in a timely manner.

[1396] "Voice input" refers to voice information uttered by a user, and is information acquired by an input device.

[1397] "Text data" is data obtained by converting voice input into text format, and is information expressed as a character string.

[1398] "Consultation content" refers to the subject of the problem or question that the user communicates through voice input.

[1399] The "emotional state" is an emotional state analyzed from the user's voice input, and indicates, for example, anger, joy, sadness, etc.

[1400] The "department" refers to a specific department or person in charge within an organization that is set up to deal with inquiries from users.

[1401] The term "server" refers to a central computer system that receives voice input, converts it into text data, analyzes emotional states, classifies the content of inquiries, and determines which departments should respond.

[1402] "Speech recognition engine" refers to the software modules and algorithms used to convert voice input into text data.

[1403] An "NLP engine" is a software module for natural language processing that analyzes text data to extract meaning and intent.

[1404] "Operator" refers to the staff at the physical store who responds to customers based on the consultation content and emotional state sent from the server.

[1405] "Feedback" refers to the act of notifying the user of the current processing status by voice or the like.

[1406] The present invention relates to a system for improving the efficiency of customer service in a physical store by using a voice recognition function and an emotion engine. This system is implemented in the following manner.

[1407] Acquiring and converting voice input

[1408] The server receives customer voice input directly from the customer service robot in the physical store. This voice input is captured in real time using a microphone connected to the server. Then, the voice data is converted into text data using a speech recognition engine (e.g., Google Speech Recognition API).

[1409] Emotion analysis

[1410] The converted text and voice data are sent to the server's natural language processing (NLP) engine and emotion engine. The emotion engine uses, for example, the Hugging Face Transformers library to analyze the customer's emotions from the voice data into "anger," "happiness," "sadness," etc. This emotion analysis information is used in subsequent processing steps.

[1411] Analysis and classification of consultation content

[1412] The server's NLP engine analyzes the converted text data and classifies the consultation content into specific categories such as "product usage," "product defect," "inventory check," etc. Here, the classification process can be performed using a library such as TextBlob.

[1413] Deciding how to respond

[1414] The server automatically determines how to respond based on the customer's emotional state obtained from emotion analysis and the classification of the inquiry content. For example, if a customer is in a "confused" emotional state and says, "I don't know where the product is," the server will contact the appropriate staff member who can respond quickly.

[1415] Notification to store staff

[1416] The server then sends a summary of the customer's consultation and emotional state, along with examples of appropriate responses, to the terminal of the selected staff member. This information is displayed in a pop-up format on the staff member's terminal, allowing the staff member to respond quickly and appropriately while taking the customer's emotions into consideration.

[1417] User Feedback

[1418] While waiting for the customer to be connected to the staff member, the server notifies the customer of the current processing status by voice. This notification may include, for example, a message such as "We are currently contacting the staff member in charge. Please wait a moment." This allows the customer to understand the processing status and reduces anxiety while waiting.

[1419] Specific examples

[1420] For example, if a customer approaches a customer service robot in a confused manner, saying, "I don't know where my item is," the system works as follows:

[1421] 1. The server captures the voice data and uses a voice recognition engine to convert it into text data such as "I don't know where the product is."

[1422] 2. The text and voice data are sent to the NLP engine and emotion engine on the server. The emotion engine analyzes the customer's "confusion" emotion.

[1423] 3. The text data is classified as a “location inquiry” and the emotional state is detected along with the sentiment analysis results.

[1424] 4. The server automatically selects the staff member most suited to the consultation and sends a notification.

[1425] 5. The server will notify the customer by voice, "We are currently contacting the staff in charge. Please wait a moment."

[1426] 6. Staff respond quickly to distressed customers based on pop-up summary information and emotional state.

[1427] Example prompt sentence:

[1428] For example, for an input such as "I'm having trouble finding the product," sentiment analysis and content classification are performed. In this way, a system that can easily and effectively improve customer service in physical stores is realized.

[1429] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1430] Step 1:

[1431] The server receives customer voice input in real time from the customer service robot in the physical store. Based on this input, the server directly captures voice data using a high-performance microphone. The input is voice data saying, "I don't know where the product is."

[1432] Step 2:

[1433] The server sends the captured voice data to a voice recognition engine and converts it into text data. Here, the Google Speech Recognition API is used. The input is voice data, and the output is text data such as "I don't know where the product is."

[1434] Step 3:

[1435] The server sends the text and voice data to a natural language processing (NLP) engine and an emotion engine. The NLP engine uses a generative AI model (such as the Hugging Face Transformers library) to analyze the text data and classify the inquiry into a specific category. In this case, the input is the text data, and the output is the category "location inquiry."

[1436] Step 4:

[1437] The server uses an emotion engine to analyze the customer's emotional state from the voice data. The emotion engine also uses a generative AI model to analyze the emotion of "confused" from the voice data. The input is the voice data, and the output is the emotional state of "confused."

[1438] Step 5:

[1439] The server automatically determines how to respond based on the classified consultation content and the analyzed emotional state. In this case, a notification is sent to an available staff member for a customer who expresses "confusion" in a "location inquiry." The input is the consultation content and emotional information, and the output is the selection of the most suitable staff member.

[1440] Step 6:

[1441] The server sends a summary of the customer's consultation and emotional state, as well as examples of appropriate responses, to the staff member's terminal. This notification is displayed in a pop-up format on the terminal. The input is customer information and emotional analysis information, and the output is a notification message to the staff member.

[1442] Step 7:

[1443] The server provides the customer with audio feedback on the current processing status. For example, it generates a message such as "We are currently contacting the staff member in charge. Please wait a moment," and notifies the customer via an audio output device. The input is the response decision information, and the output is the audio notification message.

[1444] In this way, a system can be realized in physical stores that analyzes emotions and classifies content from customer voice input, and then quickly responds appropriately.

[1445] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1446] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1447] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1448] [Fourth embodiment]

[1449] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1450] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1451] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1452] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1453] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1454] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1455] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1456] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1457] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1458] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1459] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1460] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1461] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1462] The present invention relates to a system for improving the efficiency of inquiries handled at a call center by using a voice recognition function. The system can be implemented in the following forms.

[1463] Acquiring and converting voice input

[1464] When a user calls the call center, the server launches speech recognition software to capture the user's voice input in real time. The server temporarily stores this voice data and sends it to the speech recognition engine. The speech recognition engine converts the voice data into text data. For example, if a user says, "My Internet connection is slow," the speech recognition engine converts this speech into text: "My Internet connection is slow."

[1465] Analysis and classification of consultation content

[1466] The text data is sent to a natural language processing (NLP) engine in the server. The NLP engine analyzes the text data and classifies the inquiry into a specific category. For example, the text data "My Internet connection is slow" would be classified into the "Technical Support" category.

[1467] Allocation to departments

[1468] Based on the classified content, the server automatically determines the appropriate department and specialized operator. For example, if the call is classified as "technical support," the call will be automatically connected to an operator in the technical support department.

[1469] Operator notification

[1470] The server then sends a summary of the user's inquiry and a standard response to the operator's terminal in the selected department. The terminal displays this information in a pop-up format, allowing the operator to respond quickly and appropriately. For example, in response to a question about a slow internet connection, connection troubleshooting procedures and general solutions are displayed.

[1471] User Feedback

[1472] While waiting for the connection to the appropriate department, the server will notify the user of the current processing status by voice. For example, it may say, "Your call is now being transferred to the technical support department. Please wait a moment." This voice feedback allows the user to understand the processing status and reduces anxiety while waiting.

[1473] Specific examples

[1474] When a user calls a call center complaining of an incorrect credit card charge, the system works as follows:

[1475] 1. A user calls and says, "There's an incorrect charge on my credit card."

[1476] 2. The server captures the voice data and uses a speech recognition engine to convert it into text data such as "The credit card charge is incorrect."

[1477] 3. The text data is sent to the NLP engine on the server and classified as a "billing inquiry."

[1478] 4. The server automatically connects to an operator in the billing department.

[1479] 5. The server sends a summary of the user's inquiry and an example of an appropriate response to the operator terminal in the billing department, which then displays this information in a pop-up window.

[1480] 6. The server will notify the user by voice, "Your call will now be forwarded to the billing department. Please wait a moment."

[1481] 7. Operators will respond promptly and appropriately based on the summary and standard response examples.

[1482] In this way, the system of the present invention utilizes voice recognition and natural language processing to reduce the burden on users and improve the efficiency of call center operations.

[1483] The processing flow will be explained below.

[1484] Step 1:

[1485] A user calls a call center.

[1486] Step 2:

[1487] The server runs speech recognition software to capture the user's speech in real time.

[1488] Step 3:

[1489] The server temporarily stores the captured audio data as streaming.

[1490] Step 4:

[1491] The server sends the voice data to the voice recognition engine.

[1492] Step 5:

[1493] The server uses a voice recognition engine to convert the voice data into text data.

[1494] Step 6:

[1495] The server receives the converted text data.

[1496] Step 7:

[1497] The server sends the text data to a natural language processing (NLP) engine.

[1498] Step 8:

[1499] The server analyzes the text data using an NLP engine and classifies the consultation content into specific categories.

[1500] Step 9:

[1501] Based on the analysis results, the server automatically determines the appropriate department and specialized operator to handle the situation.

[1502] Step 10:

[1503] The server transmits the consultation content to the determined department.

[1504] Step 11:

[1505] The terminal (operator's terminal) displays a summary of the received consultation and a standard response example in a pop-up format.

[1506] Step 12:

[1507] The server notifies the user of the current processing status using a speech synthesis engine.

[1508] Step 13:

[1509] Users receive notifications and are informed of the progress of the process.

[1510] Step 14:

[1511] The operator will respond quickly based on the summary and standard response examples displayed on the terminal.

[1512] Step 15:

[1513] Once the operator has completed the response, the results are fed back to the server.

[1514] Specific examples

[1515] A detailed process flow when a user contacts a call center complaining about a slow internet connection:

[1516] Step 1:

[1517] A user says, "My internet connection is slow."

[1518] Step 2:

[1519] The server launches the speech recognition software and captures the audio.

[1520] Step 3:

[1521] The server temporarily stores the captured audio as streaming.

[1522] Step 4:

[1523] The server sends the voice data to the voice recognition engine.

[1524] Step 5:

[1525] The server converts the voice data into text data saying "Your internet connection is slow."

[1526] Step 6:

[1527] The server receives the text data.

[1528] Step 7:

[1529] The server sends the text data to the NLP engine.

[1530] Step 8:

[1531] The server analyzes it using an NLP engine and classifies it as "Technical Support."

[1532] Step 9:

[1533] The server decides to assign it to the "Technical Support" department.

[1534] Step 10:

[1535] The server sends the text data to an operator in the technical support department.

[1536] Step 11:

[1537] The device will pop up a summary that reads "Your internet connection is slow" along with a standard example response.

[1538] Step 12:

[1539] The server will notify the user by voice, "Your call will now be transferred to the technical support department. Please wait a moment."

[1540] Step 13:

[1541] Ensure the user is "transferred to technical support."

[1542] Step 14:

[1543] An operator will quickly guide you through the steps to resolving the problem based on a summary and standard answers.

[1544] Step 15:

[1545] Once the operator has completed the response, the results are fed back to the server.

[1546] Example 1

[1547] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1548] Traditional call centers lacked the technology to efficiently handle user inquiries, and in many cases, operators had to respond manually. This resulted in slow response times and lower user satisfaction. Furthermore, there were sometimes delays in assigning inquiries to the appropriate department, further extending the inquiry time. Furthermore, users could not check the current processing status while waiting, which left them with no way to alleviate their anxiety while waiting.

[1549] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1550] In this invention, the server includes means for receiving voice input, means for converting the voice input into text data, means for analyzing the text data and classifying the consultation content, means for automatically determining a department to handle the consultation based on the consultation content, means for transmitting the consultation content to the department, means for displaying a summary of the consultation content and a standard answer example to the department, and means for providing voice feedback on the processing status to the user. This makes it possible to efficiently distribute inquiries from users to the departments to handle them quickly. Furthermore, the feedback to the user can reduce anxiety while waiting and improve user satisfaction.

[1551] "Means for receiving audio input" refers to equipment or software that captures audio uttered by a user through an audio input device such as a telephone or microphone as digital audio data.

[1552] "Means for converting voice input into text data" refers to a voice recognition engine or software that analyzes captured voice data and converts it into corresponding text data.

[1553] "Means for analyzing text data and classifying consultation content" refers to a system or software that uses a natural language processing engine to analyze text data obtained through voice recognition and classify the consultation content into specific categories.

[1554] "Means for automatically determining the appropriate department based on the content of the consultation" refers to a system or algorithm for selecting the appropriate department or specialized operator based on the classified content of the consultation.

[1555] The "means for transmitting the consultation content to the department in charge" refers to a communication means or protocol for electronically transmitting the user's consultation content to the determined department or operator in charge.

[1556] "Means for displaying a summary of the consultation content and a standard response example to the relevant department" refers to software for summarizing the consultation content received and displaying corresponding standard response examples on the operator terminal.

[1557] The "means for providing the user with voice feedback on the processing status" refers to a voice synthesis engine or system that notifies the user of the current processing status while the user is waiting as a voice message.

[1558] The present invention relates to a system for improving the efficiency of inquiries at a call center by using a voice recognition function. This system includes means for receiving voice input, means for converting the voice input into text data, means for analyzing the text data and classifying the inquiry content, means for automatically determining the department to handle the inquiry based on the inquiry content, means for transmitting the inquiry content to the department, means for displaying a summary of the inquiry content and a standard response example to the department, and means for providing voice feedback on the processing status to the user. The following specific hardware and software are used to implement the invention.

[1559] Acquiring and converting voice input

[1560] When a user contacts the call center via telephone or microphone, the server invokes speech recognition software, using a speech recognition engine such as the Google Cloud Speech-to-Text API, to capture the user's voice input in real time and temporarily store it as audio data.

[1561] The server sends the captured voice data to a speech recognition engine, which converts the voice data into corresponding text data, for example, if the user says "My internet connection is slow," the speech recognition engine converts the voice into text data saying "My internet connection is slow."

[1562] Analysis and classification of consultation content

[1563] The converted text data is sent to a natural language processing engine on the server. An NLP engine such as Google Cloud Natural Language API is used here. The NLP engine analyzes the text data and classifies the inquiry into a specific category. For example, text data such as "My internet connection is slow" would be classified into the "technical support" category.

[1564] Allocation to departments

[1565] The server automatically determines the corresponding department and specialized operator based on the category information received from the NLP engine. For example, if the call is classified as "technical support," the server will automatically connect to an operator in the technical support department.

[1566] Operator notification

[1567] The server then sends a summary of the user's inquiry and a standard response to the operator's terminal in the selected department. The user's inquiry is summarized and a standard response is displayed on the operator's terminal. For example, if the user says "my internet connection is slow," connection troubleshooting procedures and general solutions are displayed in a pop-up format on the terminal.

[1568] User Feedback

[1569] While waiting for the call to be connected to the appropriate department, the server will notify the user of the current processing status by voice. This uses speech synthesis software such as Google Cloud Text-to-Speech API. For example, the server may say, "Your call is now being transferred to the technical support department. Please wait a moment." This voice feedback allows the user to understand the processing status and reduces anxiety while waiting.

[1570] Specific examples

[1571] When a user reports an incorrect charge on their credit card, the system works as follows:

[1572] 1. A user calls and says, "There's an incorrect charge on my credit card."

[1573] 2. The server captures the voice data and uses a speech recognition engine to convert it into text data such as "The credit card charge is incorrect."

[1574] 3. The text data is sent to the NLP engine on the server and classified as a "billing inquiry."

[1575] 4. The server automatically connects to an operator in the billing department.

[1576] 5. The server sends a summary of the user's inquiry and an example of an appropriate response to the operator terminal in the billing department, which then displays this information in a pop-up window.

[1577] 6. The server will notify the user by voice, "Your call will now be forwarded to the billing department. Please wait a moment."

[1578] 7. Operators will respond promptly and appropriately based on the summary and standard response examples.

[1579] Prompt Sentence Examples

[1580] Examples of prompts to be input to a generative AI model:

[1581] "Please explain how this system responds when a user contacts us with a credit card billing issue."

[1582] In this way, the system of the present invention utilizes voice recognition and natural language processing to reduce the burden on users and improve the efficiency of call center operations.

[1583] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1584] Step 1:

[1585] When a user contacts the call center via telephone or microphone, the server activates voice recognition software and captures the user's voice input in real time. The input is the user's voice data, and the output is the captured digital voice data. Specifically, the server obtains the voice data using the Google Cloud Speech-to-Text API or similar.

[1586] Step 2:

[1587] The server sends the captured voice data to the speech recognition engine. The input is the captured digital voice data, and the output is text data. Specifically, the server sends the voice data to the Google Cloud Speech-to-Text API via an HTTP request and receives the text data as a response from the API.

[1588] Step 3:

[1589] The converted text data is sent to the natural language processing engine on the server. The input is text data from the speech recognition engine, and the output is category information for the consultation content. Specifically, the server sends the text data to the Google Cloud Natural Language API and receives category information as a response from the API.

[1590] Step 4:

[1591] The server automatically determines the appropriate department and specialized operator based on the category information received from the NLP engine. The input is the category information of the consultation content, and the output is information on the corresponding department and operator. Specifically, the server references an internal database to map the category and the corresponding department.

[1592] Step 5:

[1593] The server then sends a summary of the user's consultation and a standard example response to the operator terminal of the selected department. The input is information about the department and operator and text data about the consultation, and the output is a summary and example response that are displayed on the operator terminal. Specifically, the server sends the consultation data to the operator terminal as an HTTP request, and the terminal displays this data in a pop-up.

[1594] Step 6:

[1595] While waiting for the corresponding department to connect, the server notifies the user of the current processing status by voice. The input is text data about the current processing status, and the output is synthesized voice data. Specifically, the server uses the Google Cloud Text-to-Speech API to convert the text data into voice data and play it back to the user.

[1596] (Application example 1)

[1597] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1598] Virtual stores require a means to quickly and appropriately respond to inquiries and questions users have about a wide variety of products. However, conventional text-based or operator-based responses can be slow, and a lack of human resources can lead to poor user experience. For this reason, there is a need for technology that uses voice input to automatically analyze inquiries and quickly provide appropriate information.

[1599] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1600] In this invention, the server includes means for receiving voice input, means for converting the voice input into text data, and means for analyzing the text data and classifying the consultation content, thereby making it possible to obtain product information from the database based on the classification results and to respond to the user with the obtained product information in voice and text.

[1601] "Voice input" refers to the device capturing voice uttered by the user.

[1602] "Converting to text data" refers to the process of converting voice input into written information.

[1603] "Analyzing text data" means understanding the content of converted text data and identifying its meaning and intent.

[1604] "Classifying the consultation content" means dividing the content into specific categories or groups based on the analyzed text data.

[1605] "Automatically determining the department to handle the inquiry" means automatically selecting the most appropriate department and person in charge based on the classified inquiry content.

[1606] "Sending the consultation content to the department in charge" means transferring the consultation content and its information to the determined department or person in charge.

[1607] "Displaying a summary of the consultation content and a standard response example" means displaying a concise summary of the consultation content and a recommended response on the operator's terminal.

[1608] "Acquiring product information from a database based on the classification result" means searching the database for product information related to the classified consultation content and acquiring it.

[1609] "Responding to the user with the acquired product information by voice and text" means returning the acquired product information to the user as voice and text information.

[1610] The present invention relates to an application for improving the efficiency of customer support within a virtual store. The system can be implemented in the following forms.

[1611] Hardware and Software Configuration

[1612] Hardware: Smartphone (including microphone, speaker, and display)

[1613] Software: Speech recognition engine (Google Speech-to-Text API), natural language processing engine (Google Cloud Natural Language API), database server (Google Cloud Firestore), application (compatible with Android / iOS)

[1614] Program processing flow

[1615] 1. Acquiring voice input

[1616] The user launches the smartphone app and makes a voice inquiry. The smartphone app captures the user's voice through the microphone and temporarily stores it as voice data.

[1617] 2. Audio data conversion

[1618] The server converts the acquired voice data into text data using the Google Speech-to-Text API. For example, if a user asks, "What size is this dress?", the voice data is converted directly into text information.

[1619] 3. Analysis and Classification of Text Data

[1620] The converted text data is analyzed using the Google Cloud Natural Language API on the server and classified according to the inquiry content, for example, "inquiry about size."

[1621] 4. Determining Corresponding Information

[1622] Based on the classified inquiry, the server retrieves related product information from a database (Google Cloud Firestore), such as "This dress comes in sizes S, M, L, and XL."

[1623] 5. User Feedback

[1624] The acquired product information is returned to the user in voice and text format. The smartphone application provides voice feedback through the speaker and displays text on the display. For example, "This dress is available in sizes S, M, L, and XL" is displayed along with a voice response.

[1625] 6. Notify Operator (if necessary)

[1626] If a complex inquiry is made, the server will escalate the request to a specialist operator who can provide more detailed information and provide an appropriate response.

[1627] Specific examples

[1628] If a user asks "What size is this dress?" in a virtual store, the system works as follows:

[1629] 1. The user makes a voice inquiry through a smartphone app asking, "What size is this dress?"

[1630] 2. The server captures the voice data and converts it into text data such as "What size is this dress?" using the Google Speech-to-Text API.

[1631] 3. The server analyzes the text data using the Google Cloud Natural Language API and classifies it as a "size query."

[1632] 4. Based on the classification results, the server retrieves related product information from Google Cloud Firestore, such as "This dress comes in sizes S, M, L, and XL."

[1633] 5. The server responds to the user with the acquired product information in voice and text format. The smartphone app displays the text on the screen and simultaneously responds with voice.

[1634] Example prompts for natural language generation AI models

[1635] markdown

[1636] Prompt: In a virtual shopping assistant, if a user asks "What size is this dress?", generate a prompt to respond with the appropriate size information to provide to the user.

[1637] Examples:

[1638] markdown

[1639] User input: "What size is this dress?"

[1640] Sizing Information: "This dress comes in sizes S, M, L, and XL."

[1641] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1642] Step 1:

[1643] Acquiring voice input

[1644] The user launches the smartphone app and inputs the inquiry by voice. The smartphone is equipped with a microphone, which captures the user's voice. The input voice is temporarily saved in the smartphone application.

[1645] Input: User voice input

[1646] Output: Audio data

[1647] Step 2:

[1648] Audio data conversion

[1649] The server receives voice data sent from the smartphone. The voice data is sent to the Google Speech-to-Text API and converted into text data in real time. For example, voice data such as "What size is this dress?" is converted into text data such as "What size is this dress?"

[1650] Input: Audio data

[1651] Output: Text data

[1652] Step 3:

[1653] Text data analysis and classification

[1654] The server then sends the converted text data to the Google Cloud Natural Language API, which analyzes the text data and classifies the inquiry into a specific category. For example, the text data "What size is this dress?" would be classified as a "size inquiry."

[1655] Input: Text data

[1656] Output: Classification results

[1657] Step 4:

[1658] Determining correspondence information

[1659] The server searches and retrieves related product information from a database (Google Cloud Firestore) based on the classified inquiry. For example, for an inquiry about size, it retrieves the size information of the specified product.

[1660] Input: Classification results

[1661] Output: Product information data

[1662] Step 5:

[1663] User Feedback

[1664] The server generates the acquired product information as a voice response and a text response. The generated voice data is fed back to the user through the smartphone's speaker, and the text data is displayed on the smartphone's display. For example, the server responds by voice, saying, "This dress comes in sizes S, M, L, and XL," and the same content is displayed on the display.

[1665] Input: Product information data

[1666] Output: Audio and text responses

[1667] Step 6:

[1668] Notify the operator (if necessary)

[1669] If the inquiry is complex and requires a detailed response, the server escalates the inquiry to an appropriate operator, who then sends a summary of the inquiry and a recommended response to the operator's terminal.

[1670] Input: Classification results and product information data

[1671] Output: Notification to operator (summary and standard response example)

[1672] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1673] The present invention relates to a system for improving the efficiency of inquiries handled at a call center by using a voice recognition function and an emotion engine. The system can be implemented in the following forms.

[1674] Acquiring and converting voice input

[1675] When a user calls the call center, the server launches speech recognition software to capture the user's voice input in real time. The server temporarily stores this voice data and sends it to the speech recognition engine. The speech recognition engine converts the voice data into text data. In this case, if the user says, "My Internet connection is slow," the speech recognition engine converts this speech into text: "My Internet connection is slow."

[1676] Emotion analysis

[1677] The converted text data is sent to the server's natural language processing (NLP) engine and also to the emotion engine, which analyzes the user's emotions from the voice data and identifies emotional states such as anger, joy, and sadness. This emotion analysis information is used for subsequent processing.

[1678] Analysis and classification of consultation content

[1679] The NLP engine analyzes the text data and classifies the inquiry into a specific category. For example, the text data "My internet connection is slow" will be classified into the "Technical Support" category.

[1680] Allocation to departments

[1681] The server automatically determines the department and specialized operator to respond to the inquiry based on the classified content and the user's emotions obtained from emotion analysis. For example, if the user expresses anger, the server will prioritize connecting the user to an operator who can respond more quickly and professionally.

[1682] Operator notification

[1683] The server then sends a summary of the user's consultation, a standard response example, and the user's emotional state to the operator terminal in the selected department. The terminal displays this information in a pop-up format, allowing the operator to respond quickly and appropriately while taking the user's emotions into consideration. For example, if the message "Internet connection is slow" and the emotional state of "anger" are displayed, the operator can prepare to respond calmly and quickly.

[1684] User Feedback

[1685] While waiting for the connection to the appropriate department, the server will notify the user of the current processing status by voice. For example, it may say, "Your call is now being transferred to the technical support department. Please wait a moment." This voice feedback allows the user to understand the processing status and reduces anxiety while waiting.

[1686] Specific examples

[1687] If a user angrily calls a call center complaining that their credit card was charged incorrectly, the system works like this:

[1688] 1. A user calls and angrily states that there is an error in the charge on their credit card.

[1689] 2. The server captures the voice data and uses a speech recognition engine to convert it into text data such as "The credit card charge is incorrect."

[1690] 3. The text and voice data are sent to the NLP engine and emotion engine on the server. The emotion engine analyzes the user's anger.

[1691] 4. The text data is classified as a "billing inquiry" and the emotional state of "anger" is detected along with the results of the emotion analysis.

[1692] 5. The server automatically connects to an operator in the billing department, and based on the results of emotion analysis, selects an operator who is best able to handle the emotion of "anger."

[1693] 6. The server sends a summary of the user's consultation, their emotional state, and examples of appropriate responses to the operator terminal in the billing department. The terminal displays this information in a pop-up.

[1694] 7. The server will notify the user by voice, "Your call will now be forwarded to the billing department. Please wait."

[1695] 8. The operator responds calmly, quickly, and sensitively based on the summary and the caller's emotional state.

[1696] In this way, the system of the present invention, which combines an emotion engine, not only performs speech recognition and natural language processing but also analyzes and takes into account the user's emotional state, thereby reducing the burden on the user and further improving the efficiency of call center operations.

[1697] The processing flow will be explained below.

[1698] Step 1:

[1699] A user calls a call center.

[1700] Step 2:

[1701] The server runs speech recognition software to capture the user's speech in real time.

[1702] Step 3:

[1703] The server temporarily stores the captured audio data as streaming.

[1704] Step 4:

[1705] The server sends the voice data to the voice recognition engine.

[1706] Step 5:

[1707] The server uses a speech recognition engine to convert the voice data into text data, for example, "Your internet connection is slow."

[1708] Step 6:

[1709] The server then sends the converted text and voice data to an emotion engine, which analyzes the user's emotions. For example, emotions such as "anger" or "dissatisfaction" are detected from voice patterns.

[1710] Step 7:

[1711] At the same time, the server sends the converted text data to a natural language processing (NLP) engine.

[1712] Step 8:

[1713] The server analyzes the text data using an NLP engine and classifies the inquiry into a specific category. For example, the text "My internet connection is slow" is classified into the "Technical Support" category.

[1714] Step 9:

[1715] The server automatically determines the department and specialist operator to respond to the inquiry based on the content of the inquiry and the results of emotion analysis. For example, if emotion analysis detects anger, it will select an operator from the technical support department who can respond quickly.

[1716] Step 10:

[1717] The server sends the consultation content and emotion analysis results to the determined department.

[1718] Step 11:

[1719] The terminal (operator's terminal) displays a summary of the received consultation, a standard response example, and analyzed emotional information in a pop-up format. For example, it displays the summary "My internet connection is slow" and the emotional state "Anger."

[1720] Step 12:

[1721] The server uses a speech synthesis engine to notify the user of the current status of the process, for example, "Your call is being transferred to technical support. Please wait."

[1722] Step 13:

[1723] Users receive notifications and are informed of the progress of the process.

[1724] Step 14:

[1725] The operator responds quickly and appropriately based on the summary and emotional information displayed on the terminal, for example, by speaking calmly and politely to alleviate the user's anger and try to resolve the technical problem.

[1726] Step 15:

[1727] Once the operator has completed the response, the results are fed back to the server.

[1728] Specific examples

[1729] A detailed process flow when a user contacts a call center because they are angry about an incorrect credit card charge:

[1730] Step 1:

[1731] A user says, "There's an incorrect charge on my credit card."

[1732] Step 2:

[1733] The server launches the speech recognition software and captures the audio.

[1734] Step 3:

[1735] The server temporarily stores the captured audio as streaming.

[1736] Step 4:

[1737] The server sends the voice data to the voice recognition engine.

[1738] Step 5:

[1739] The server converts the voice data into text data saying "The credit card billing is incorrect."

[1740] Step 6:

[1741] The server sends the text data and voice data to the emotion engine, which then analyzes the user's emotion of "anger."

[1742] Step 7:

[1743] The server sends the text data to the NLP engine.

[1744] Step 8:

[1745] The server analyzes the text data using an NLP engine and classifies it as a "billing inquiry."

[1746] Step 9:

[1747] The server determines the department to handle the inquiry based on the emotion of "invoice inquiry" and "anger." For example, it may give priority to an operator in the billing department who can respond quickly.

[1748] Step 10:

[1749] The server sends the consultation details and emotion analysis results to the billing department operator.

[1750] Step 11:

[1751] The device will pop up a summary saying "Credit card charge incorrect" and an emotional state of "Anger."

[1752] Step 12:

[1753] The server will announce, "Your call is now being forwarded to the billing department. Please wait a moment."

[1754] Step 13:

[1755] Verify that the user is "transferred to the billing department."

[1756] Step 14:

[1757] The operator responds calmly and quickly based on the summary and emotional information, for example, by first apologizing to calm the user's anger and then confirming the details of the problem.

[1758] Step 15:

[1759] Once the operator has completed the response, the results are fed back to the server.

[1760] In this way, the system of the present invention, which combines an emotion engine, not only performs speech recognition and natural language processing but also analyzes and takes into account the user's emotional state, thereby reducing the burden on the user and further improving the efficiency of call center operations.

[1761] Example 2

[1762] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1763] Handling inquiries at call centers requires listening to voice input, analyzing content, analyzing emotions, assigning calls to the appropriate department, and providing information to operators promptly. Conventional systems have difficulty performing these tasks efficiently, resulting in problems such as reduced user satisfaction and reduced operator work efficiency. The present invention aims to provide a system that solves these problems and enables efficient call center operation.

[1764] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1765] In this invention, the server includes means for receiving voice input, means for converting the voice input into text data, means for analyzing emotions from the text data and voice data, means for analyzing and classifying the consultation content, means for automatically determining a department to handle the consultation based on the consultation content and emotion analysis results, means for sending a summary of the consultation content and a standard response example to the department to handle the consultation, and means for displaying appropriate information on an operator terminal. This makes it possible to analyze the user's voice input in real time and quickly transfer the consultation to the appropriate department or operator.

[1766] A "means for receiving voice input" is a device or program that has the function of capturing voice from a user and recording it as digital voice data.

[1767] The "means for converting voice input into text data" is a device or program that has the function of analyzing captured voice data and converting it into corresponding text data.

[1768] The "means for analyzing text data and classifying consultation contents" refers to a device or program that has the function of analyzing the contents based on the text data and classifying them into specific categories.

[1769] The "means for analyzing emotions" is a device or program that has the function of identifying the emotional state of a user from voice data and text data.

[1770] The "means for automatically determining the department to handle the call" is a device or program that has the function of selecting the appropriate department or operator to handle the call based on the classified content of the call and the results of emotion analysis.

[1771] The "means for transmitting the consultation content to the department" is a device or program having a function for transmitting the consultation content of the user and related information to the automatically determined department.

[1772] The "means for displaying a summary of the consultation content and a standard response example" is a device or program that has the function of displaying a summary of the consultation content and a standard response example on the operator terminal, thereby supporting a prompt and appropriate response.

[1773] This invention is a system that uses voice recognition and emotion analysis functions to improve the efficiency of inquiries handled at call centers. This system converts user voice input into text data, analyzes the content of the inquiry and the user's emotional state, and automatically assigns the inquiry to the appropriate department and operator.

[1774] To implement this system, the following hardware and software are used.

[1775] Hardware

[1776] server

[1777] Operator terminal (e.g. Windows PC)

[1778] software

[1779] Voice recognition software (e.g., Google Cloud Speech-to-Text)

[1780] Natural Language Processing (NLP) engines (e.g., SpaCy)

[1781] Sentiment analysis engine (e.g. IBM Watson Tone Analyzer)

[1782] Program processing

[1783] Acquiring voice input

[1784] When a user calls the call center, the server captures the user's voice in real time through a telephone line interface and stores it as digital voice data.

[1785] Speech recognition and text conversion

[1786] The server sends the stored voice data to a speech recognition engine, which converts it into text data. For example, if a user says, "My internet connection is slow," the voice data will be converted into text data saying, "My internet connection is slow."

[1787] Emotion analysis

[1788] The converted text and voice data is sent to the emotion engine on the server, which analyzes the data and identifies the user's emotion (e.g., anger, joy, sadness).

[1789] Analysis and classification of consultation content

[1790] The server sends the text data to an NLP engine, which analyzes it and classifies it into a specific category, for example, "My internet connection is slow" would be classified into the "tech support" category.

[1791] Allocation to departments

[1792] The server automatically determines the appropriate department and operator based on the category of the consultation content and the results of emotion analysis. If the user expresses anger, the call will be connected preferentially to an operator who can respond quickly and professionally.

[1793] Operator notification

[1794] The server then sends a summary of the user's consultation and their emotional state to the operator terminal in the selected department. The terminal displays this information in a pop-up format, allowing the operator to respond promptly and appropriately.

[1795] User Feedback

[1796] While waiting for the connection to the appropriate department, the server will notify the user of the current processing status by voice. For example, by guiding the user to "Your call is now being transferred to the technical support department. Please wait a moment," the user can understand the processing status and reduce anxiety while waiting.

[1797] Specific examples

[1798] If a user angrily calls a call center complaining that their credit card was charged incorrectly, the system works like this:

[1799] 1. A user calls and angrily states that there is an error in the charge on their credit card.

[1800] 2. The server captures the voice data and uses a speech recognition engine to convert it into text data such as "The credit card charge is incorrect."

[1801] 3. The text and voice data are sent to the NLP engine and emotion engine on the server. The emotion engine analyzes the user's anger.

[1802] 4. The text data is classified as a "billing inquiry" and the emotional state of "anger" is detected along with the results of the emotion analysis.

[1803] 5. The server automatically connects to an operator in the billing department, and based on the results of emotion analysis, selects an operator who is best able to handle the emotion of "anger."

[1804] 6. The server sends a summary of the user's consultation, their emotional state, and examples of appropriate responses to the operator terminal in the billing department. The terminal displays this information in a pop-up.

[1805] 7. The server will notify the user by voice, "Your call will now be forwarded to the billing department. Please wait."

[1806] 8. The operator responds calmly, quickly, and sensitively based on the summary and the caller's emotional state.

[1807] Prompt Sentence Examples

[1808] "A user is expressing anger over a claim that their credit card is incorrectly charged. What's the best approach to calmly respond?"

[1809] "Some users are complaining about slow internet connections. Can you suggest a good solution for this?"

[1810] In this way, the system of the present invention, which combines an emotion engine, not only performs speech recognition and natural language processing but also analyzes and takes into account the user's emotional state, thereby reducing the burden on the user and further improving the efficiency of call center operations.

[1811] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1812] Step 1: Getting voice input

[1813] When the server detects an incoming call from the user, it immediately activates the voice recognition software and captures the user's voice input in real time. The input is an analog voice signal received over the telephone line. The server converts this voice signal into digital voice data and temporarily stores it.

[1814] Specific behavior:

[1815] The user makes a call.

[1816] The telephone line interface implemented in the server detects the incoming call.

[1817] The server receives voice input through a telephone line interface and stores it as digital voice data.

[1818] Input: Analog audio signal obtained through telephone lines.

[1819] Output: Digital audio data.

[1820] Step 2: Speech recognition and text conversion

[1821] The server then sends the stored digital voice data to a speech recognition engine (such as Google Cloud Speech-to-Text) to convert it into text data. The speech recognition engine then analyzes the digital voice data and generates corresponding text data.

[1822] Specific behavior:

[1823] The server calls the API to send the voice data to the voice recognition engine.

[1824] The speech recognition engine analyzes the speech and generates text data.

[1825] The server receives the generated text data.

[1826] Input: Digital audio data.

[1827] Output: Text data.

[1828] Step 3: Sentiment Analysis

[1829] The server sends the generated text data and saved voice data to an emotion engine (such as IBM Watson Tone Analyzer) to analyze the user's emotional state. The emotion engine analyzes the voice and text data to identify the user's emotions (e.g., anger, joy, sadness).

[1830] Specific behavior:

[1831] The server calls the API to send text and voice data to the emotion engine.

[1832] An emotion engine analyzes the data and identifies the user's emotional state.

[1833] The server receives the results of the sentiment analysis.

[1834] Input: text data, audio data.

[1835] Output: Sentiment analysis results (e.g. anger, joy, sadness).

[1836] Step 4: Analysis and classification of consultation content

[1837] The server sends the text data to a natural language processing (NLP) engine, which analyzes it and classifies the query into a specific category. The NLP engine analyzes the text and classifies it into a specific category (e.g., technical support, billing).

[1838] Specific behavior:

[1839] The server calls the API to send the text data to the NLP engine.

[1840] The NLP engine analyzes the text data and identifies categories.

[1841] The server receives the analysis results.

[1842] Input: Text data.

[1843] Output: Classification results (e.g., tech support, billing).

[1844] Step 5: Assign to departments

[1845] The server automatically determines the appropriate department and operator based on the category of the consultation content and the results of emotion analysis. If the user expresses anger, the caller will be connected preferentially to an operator who can provide a prompt and professional response.

[1846] Specific behavior:

[1847] The server refers to the emotion analysis and classification results and runs an algorithm to select the appropriate department or operator.

[1848] The server issues connection instructions to the selected department or operator.

[1849] Input: Sentiment analysis results, classification results.

[1850] Output: Information on the selected response department and operator.

[1851] Step 6: Notify the operator

[1852] The server then sends a summary of the user's consultation and their emotional state to the operator terminal in the selected department. The operator terminal displays this information in a pop-up format, helping to ensure a prompt and appropriate response.

[1853] Specific behavior:

[1854] The server sends a summary of the user's consultation and the results of emotion analysis to the operator terminal.

[1855] The information received by the operator terminal is displayed as a pop-up.

[1856] Input: Summary of consultation content, sentiment analysis results.

[1857] Output: Information displayed on the operator terminal.

[1858] Step 7: User feedback

[1859] While waiting for the call to be connected to the appropriate department, the server will notify the user of the current processing status by voice, for example, "Your call is now being transferred to the technical support department. Please wait a moment."

[1860] Specific behavior:

[1861] The server generates a feedback message for the user.

[1862] The server generates this message as voice data using voice synthesis software and transmits it to the user over a telephone line.

[1863] Input: Current processing status.

[1864] Output: Audio feedback to the user.

[1865] (Application example 2)

[1866] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1867] Customer service in modern brick-and-mortar stores requires a fast and accurate response to customer problems and complaints. However, traditional methods make it difficult to consider customer emotions, making it difficult to increase customer satisfaction. In particular, there is a lack of appropriate responses to emotional customers, which leads to a decline in the service quality of the entire store.

[1868] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1869] In this invention, the server includes a means for receiving voice input, a means for converting the voice input into text data, and a means for analyzing the user's emotional state from the text data and voice data, thereby enabling the server to analyze the customer's emotions in real time, automatically determine an appropriate response method, and notify store staff in a timely manner.

[1870] "Voice input" refers to voice information uttered by a user, and is information acquired by an input device.

[1871] "Text data" is data obtained by converting voice input into text format, and is information expressed as a character string.

[1872] "Consultation content" refers to the subject of the problem or question that the user communicates through voice input.

[1873] The "emotional state" is an emotional state analyzed from the user's voice input, and indicates, for example, anger, joy, sadness, etc.

[1874] The "department" refers to a specific department or person in charge within an organization that is set up to deal with inquiries from users.

[1875] The term "server" refers to a central computer system that receives voice input, converts it into text data, analyzes emotional states, classifies the content of inquiries, and determines which departments should respond.

[1876] "Speech recognition engine" refers to the software modules and algorithms used to convert voice input into text data.

[1877] An "NLP engine" is a software module for natural language processing that analyzes text data to extract meaning and intent.

[1878] "Operator" refers to the staff at the physical store who responds to customers based on the consultation content and emotional state sent from the server.

[1879] "Feedback" refers to the act of notifying the user of the current processing status by voice or the like.

[1880] The present invention relates to a system for improving the efficiency of customer service in a physical store by using a voice recognition function and an emotion engine. This system is implemented in the following manner.

[1881] Acquiring and converting voice input

[1882] The server receives customer voice input directly from the customer service robot in the physical store. This voice input is captured in real time using a microphone connected to the server. Then, the voice data is converted into text data using a speech recognition engine (e.g., Google Speech Recognition API).

[1883] Emotion analysis

[1884] The converted text and voice data are sent to the server's natural language processing (NLP) engine and emotion engine. The emotion engine uses, for example, the Hugging Face Transformers library to analyze the customer's emotions from the voice data into "anger," "happiness," "sadness," etc. This emotion analysis information is used in subsequent processing steps.

[1885] Analysis and classification of consultation content

[1886] The server's NLP engine analyzes the converted text data and classifies the consultation content into specific categories such as "product usage," "product defect," "inventory check," etc. Here, the classification process can be performed using a library such as TextBlob.

[1887] Deciding how to respond

[1888] The server automatically determines how to respond based on the customer's emotional state obtained from emotion analysis and the classification of the inquiry content. For example, if a customer is in a "confused" emotional state and says, "I don't know where the product is," the server will contact the appropriate staff member who can respond quickly.

[1889] Notification to store staff

[1890] The server then sends a summary of the customer's consultation and emotional state, along with examples of appropriate responses, to the terminal of the selected staff member. This information is displayed in a pop-up format on the staff member's terminal, allowing the staff member to respond quickly and appropriately while taking the customer's emotions into consideration.

[1891] User Feedback

[1892] While waiting for the customer to be connected to the staff member, the server notifies the customer of the current processing status by voice. This notification may include, for example, a message such as "We are currently contacting the staff member in charge. Please wait a moment." This allows the customer to understand the processing status and reduces anxiety while waiting.

[1893] Specific examples

[1894] For example, if a customer approaches a customer service robot in a confused manner, saying, "I don't know where my item is," the system works as follows:

[1895] 1. The server captures the voice data and uses a voice recognition engine to convert it into text data such as "I don't know where the product is."

[1896] 2. The text and voice data are sent to the NLP engine and emotion engine on the server. The emotion engine analyzes the customer's "confusion" emotion.

[1897] 3. The text data is classified as a “location inquiry” and the emotional state is detected along with the sentiment analysis results.

[1898] 4. The server automatically selects the staff member most suited to the consultation and sends a notification.

[1899] 5. The server will notify the customer by voice, "We are currently contacting the staff in charge. Please wait a moment."

[1900] 6. Staff respond quickly to distressed customers based on pop-up summary information and emotional state.

[1901] Example prompt sentence:

[1902] For example, for an input such as "I'm having trouble finding the product," sentiment analysis and content classification are performed. In this way, a system that can easily and effectively improve customer service in physical stores is realized.

[1903] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1904] Step 1:

[1905] The server receives customer voice input in real time from the customer service robot in the physical store. Based on this input, the server directly captures voice data using a high-performance microphone. The input is voice data saying, "I don't know where the product is."

[1906] Step 2:

[1907] The server sends the captured voice data to a voice recognition engine and converts it into text data. Here, the Google Speech Recognition API is used. The input is voice data, and the output is text data such as "I don't know where the product is."

[1908] Step 3:

[1909] The server sends the text and voice data to a natural language processing (NLP) engine and an emotion engine. The NLP engine uses a generative AI model (such as the Hugging Face Transformers library) to analyze the text data and classify the inquiry into a specific category. In this case, the input is the text data, and the output is the category "location inquiry."

[1910] Step 4:

[1911] The server uses an emotion engine to analyze the customer's emotional state from the voice data. The emotion engine also uses a generative AI model to analyze the emotion of "confused" from the voice data. The input is the voice data, and the output is the emotional state of "confused."

[1912] Step 5:

[1913] The server automatically determines how to respond based on the classified consultation content and the analyzed emotional state. In this case, a notification is sent to an available staff member for a customer who expresses "confusion" in a "location inquiry." The input is the consultation content and emotional information, and the output is the selection of the most suitable staff member.

[1914] Step 6:

[1915] The server sends a summary of the customer's consultation and emotional state, as well as examples of appropriate responses, to the staff member's terminal. This notification is displayed in a pop-up format on the terminal. The input is customer information and emotional analysis information, and the output is a notification message to the staff member.

[1916] Step 7:

[1917] The server provides the customer with audio feedback on the current processing status. For example, it generates a message such as "We are currently contacting the staff member in charge. Please wait a moment," and notifies the customer via an audio output device. The input is the response decision information, and the output is the audio notification message.

[1918] In this way, a system can be realized in physical stores that analyzes emotions and classifies content from customer voice input, and then quickly responds appropriately.

[1919] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1920] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1921] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1922] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1923] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1924] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1925] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1926] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1927] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1928] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1929] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1930] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1931] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1932] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1933] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1934] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1935] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1936] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1937] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1938] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1939] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1940] The following is further disclosed regarding the above embodiment.

[1941] (Claim 1)

[1942] means for receiving audio input;

[1943] means for converting the voice input into text data;

[1944] means for analyzing the text data and classifying the consultation contents;

[1945] means for automatically determining a department to handle the consultation based on the consultation content;

[1946] A means of sending the consultation details to the corresponding department,

[1947] a means for displaying a summary of the consultation content and a standard response example to the department in charge;

[1948] A system including:

[1949] (Claim 2)

[1950] 2. The system according to claim 1, further comprising means for feeding back the consultation content to the user by audio output.

[1951] (Claim 3)

[1952] 2. The system according to claim 1, further comprising means for acquiring the voice input, converting it into text data, and automatically allocating it to a corresponding department in real time.

[1953] "Example 1"

[1954] (Claim 1)

[1955] means for receiving audio input;

[1956] means for converting the voice input into text data;

[1957] means for analyzing the text data and classifying the consultation contents;

[1958] means for automatically determining a department to handle the consultation based on the consultation content;

[1959] A means of sending the consultation details to the corresponding department,

[1960] a means for displaying a summary of the consultation content and a standard response example to the department in charge;

[1961] means for providing audio feedback to the user about the processing status;

[1962] A system including:

[1963] (Claim 2)

[1964] 2. The system according to claim 1, further comprising means for feeding back the consultation content to the user by audio output.

[1965] (Claim 3)

[1966] 2. The system according to claim 1, further comprising means for acquiring the voice input, converting it into text data, and automatically allocating it to a corresponding department in real time.

[1967] "Application Example 1"

[1968] (Claim 1)

[1969] means for receiving audio input;

[1970] means for converting the voice input into text data;

[1971] means for analyzing the text data and classifying the consultation contents;

[1972] means for automatically determining a department to handle the consultation based on the consultation content;

[1973] A means of sending the consultation details to the corresponding department,

[1974] a means for displaying a summary of the consultation content and a standard response example to the department in charge;

[1975] A means for retrieving product information from a database based on the classification result;

[1976] means for responding to the user with the acquired product information by voice and text;

[1977] A system including:

[1978] (Claim 2)

[1979] 2. The system according to claim 1, further comprising means for feeding back the consultation content to the user by audio output.

[1980] (Claim 3)

[1981] 2. The system according to claim 1, further comprising means for acquiring the voice input, converting it into text data, and automatically allocating it to a corresponding department in real time.

[1982] "Example 2: Combining Emotion Engines"

[1983] (Claim 1)

[1984] means for receiving audio input;

[1985] means for converting the voice input into text data;

[1986] means for analyzing the text data and classifying the consultation contents;

[1987] means for analyzing emotions from the consultation content and voice data;

[1988] means for automatically determining a department to handle the consultation based on the emotion analysis result and the consultation content;

[1989] A means of sending the consultation details to the corresponding department,

[1990] a means for displaying a summary of the consultation content and a standard response example to the department in charge;

[1991] A system including:

[1992] (Claim 2)

[1993] 2. The system according to claim 1, further comprising means for feeding back the consultation content to the user by audio output.

[1994] (Claim 3)

[1995] 2. The system according to claim 1, further comprising means for acquiring the voice input, converting it into text data, and automatically allocating it to a corresponding department in real time.

[1996] "Application example 2 when combining emotion engines"

[1997] (Claim 1)

[1998] means for receiving audio input;

[1999] means for converting the voice input into text data;

[2000] means for analyzing the text data and classifying the consultation contents;

[2001] means for analyzing the emotional state of a user from the text data and voice data;

[2002] means for automatically determining a department to handle the consultation based on the consultation content and emotional state;

[2003] A means of sending the consultation details to the corresponding department,

[2004] a means for displaying a summary of the consultation content and emotional state and a standard response example;

[2005] A system including:

[2006] (Claim 2)

[2007] 2. The system according to claim 1, further comprising means for providing feedback to the user on the consultation content and the current processing status by audio output.

[2008] (Claim 3)

[2009] 2. The system according to claim 1, further comprising means for acquiring the voice input, converting it into text data, and automatically allocating it to a corresponding department in real time. [Explanation of symbols]

[2010] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for receiving audio input; means for converting the voice input into text data; means for analyzing the text data and classifying the consultation contents; means for automatically determining a department to handle the consultation based on the consultation content; A means of sending the consultation details to the corresponding department, a means for displaying a summary of the consultation content and a standard response example to the department in charge; A system including:

2. The system according to claim 1 , further comprising means for feeding back the consultation content to the user by audio output.

3. The system according to claim 1 , further comprising means for acquiring the voice input, converting it into text data, and automatically allocating it to a corresponding department in real time.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A