System

The system addresses operator stress by converting voice to text, detecting harassment, and switching to AI support, enhancing efficiency and satisfaction in call centers.

JP2026030454APending Publication Date: 2026-02-20SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024133437
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-08
Publication Date
2026-02-20

AI Technical Summary

Technical Problem

Call center operators experience stress and reduced productivity due to customer complaints and harassment, with conventional solutions lacking immediacy and efficiency.

Method used

A system that acquires voice input from operators, converts it to text in real-time, performs natural language processing to detect harassment, and automatically switches to AI support, while notifying customers about the system's implementation.

Benefits of technology

Reduces the psychological burden on operators, improves work efficiency, and enhances workplace satisfaction by providing quick and appropriate responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026030454000001_ABST
    Figure 2026030454000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: The system according to claim 1, further comprising: means for detecting a customer harassment by performing a natural language process on the basis of the converted text data; means for automatically switching from the operator to a AI response when the customer harassment is detected; and means for notifying the customer of the introduction of the system in advance.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] The purpose of this invention is to solve the problems of stress, reduced productivity, and psychological burden that call center operators experience when faced with customer complaints and customer harassment. Conventional solutions lack immediacy and efficiency, so there is a need to reduce the mental and work burden on operators. [Means for solving the problem]

[0005] The present invention provides a system that acquires voice input from the terminal used by the operator, analyzes the voice data in real time, and converts it into text. Furthermore, natural language processing is performed based on the converted text data to detect customer harassment. If customer harassment is detected, the operator automatically switches to AI support, reducing the operator's burden and improving work efficiency. Furthermore, by notifying customers in advance of the introduction of this system, it is expected that customer harassment will also be deterred. In this way, a system is provided that improves the working environment for operators and increases productivity and workplace satisfaction.

[0006] "Terminal" means a device used by an operator to obtain voice input.

[0007] "Voice data" refers to the voice information exchanged between the customer and the operator obtained from the terminal.

[0008] "Real-time" refers to processing occurring immediately as audio data is acquired.

[0009] "Analysis" refers to the process of converting voice data into text data and subsequent data analysis.

[0010] "Text" refers to the written information obtained by converting voice data.

[0011] "Natural language processing" refers to a set of technologies for analyzing text data to detect whether or not customer harassment has occurred.

[0012] "Customer harassment" refers to inappropriate words, actions, or behavior by customers toward operators.

[0013] "AI-enabled" refers to artificial intelligence responding on behalf of an operator when customer harassment is detected.

[0014] "Means of prior notification" refers to the methods and procedures for notifying customers in advance about the introduction of the system.

[0015] "Harassment detection result" refers to the result of determining that customer harassment has occurred through natural language processing. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] This invention is an AI system for reducing the workload of call center operators and improving productivity and workplace satisfaction.

[0038] composition

[0039] The system consists of the following main components:

[0040] 1. Terminal: A telephone terminal for the operator.

[0041] 2. Server: A computing device that hosts the AI ​​model and speech recognition module.

[0042] 3. Management Console: Operator and administrator interface.

[0043] Program processing

[0044] Acquiring voice input

[0045] The terminal captures the voice conversation between the operator and the customer in real time and sends the voice data to the server, which receives the voice data and transfers it directly to the voice recognition module.

[0046] Voice Recognition

[0047] The server converts the received voice data into text data using a speech recognition module, which is then used for subsequent natural language processing (NLP).

[0048] Real-time analysis of conversation content

[0049] The server inputs the converted text data into an NLP module to detect customer harassment, which uses historical data and keyword filtering for analysis.

[0050] Customer Harassment Detection

[0051] The server receives the analysis results from the NLP module, and if customer harassment is detected, it notifies the operator and automatically switches to the AI-enabled module.

[0052] AI-powered response

[0053] The AI-enabled module takes over the response and sends the generated AI response to the terminal in real time, which relieves the operator of the psychological burden and allows the appropriate response to be taken.

[0054] Advance notice to customers

[0055] The user will notify the customer in advance that this system is being implemented, including the purpose of the AI ​​to prevent harassment.

[0056] Specific examples

[0057] For example, when an operator receives a call from a customer, the voice is captured on the terminal and immediately sent to the server. The server converts the voice into text and analyzes it using the NLP module. If the customer makes a harassing remark such as, "Poor service! Call the manager!", the NLP module detects this and the server switches to the AI-enabled module. The AI ​​generates an appropriate response and responds to the customer via the terminal. This process reduces the burden on the operator and enables efficient business operations.

[0058] In this way, the system reduces the burden on operators and contributes to improving work efficiency and customer satisfaction.

[0059] The processing flow will be explained below.

[0060] Step 1:

[0061] The terminal captures voice inputs exchanged between the operator and the customer in real time, and the captured voice data is immediately sent to the server.

[0062] Step 2:

[0063] The server passes the received voice data to a speech recognition module, which converts the voice into text. This conversion process is performed in real time, and the resulting text data is generated.

[0064] Step 3:

[0065] The server feeds the generated text data into a natural language processing (NLP) module to analyze the conversation, which uses keyword-based filtering and historical data to detect customer harassment.

[0066] Step 4:

[0067] The server checks the analysis results received from the NLP module, and if customer harassment is detected, it notifies an operator and passes the response to the AI-enabled module.

[0068] Step 5:

[0069] The server activates the AI-enabled module and generates an appropriate response, which is then sent to the device in real time.

[0070] Step 6:

[0071] The terminal receives the AI ​​response from the server and forwards it to the customer, which relieves the operator of the psychological burden.

[0072] Step 7:

[0073] When the user installs the system, the user will provide advance notice to the customer, including the fact that the system is being installed for the purpose of preventing customer harassment.

[0074] Example 1

[0075] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0076] In conventional call centers, operators have to deal with harassing comments from customers, which places a heavy psychological burden on them, often resulting in a decline in productivity and workplace satisfaction. Furthermore, if appropriate responses are not made promptly, customer satisfaction also declines, which is an issue.

[0077] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0078] In this invention, the server includes means for acquiring voice input from a terminal used by an operator, means for transmitting the acquired voice data to a computing device in real time, means for the computing device to analyze the transmitted voice data using a voice recognition module, means for converting the voice data analyzed by the voice recognition module into text data, means for inputting the converted text data into a natural language processing module and analyzing the conversation content in real time, means for detecting customer harassment from the analysis results using the natural language processing module, means for automatically switching from the operator to an AI response when customer harassment is detected, means for transmitting the generated AI response to the terminal in real time to respond to the customer, and means for notifying the customer in advance of the introduction of the system. This reduces the psychological burden on the operator and enables quick and appropriate response.

[0079] "Terminal" refers to the communication device used by the operator, which is a hardware device for making calls with customers.

[0080] "Means for acquiring voice input" refers to a device or method that allows the terminal to capture the voice exchanged between the operator and the customer in real time and acquire it as digital data.

[0081] "Means for transmitting voice data in real time" refers to a device or method that has a mechanism for instantly transferring voice data acquired from a terminal to a server.

[0082] A "computing device" is a computer system for analyzing and processing audio data. Also called a server.

[0083] A "voice recognition module" is software or algorithms for converting voice data into text data.

[0084] The "means for converting into text data" is a function that uses a voice recognition module to convert voice data into text information and output it in digital text format.

[0085] A "natural language processing module" is software or algorithms for analyzing text data and understanding and processing its content.

[0086] The "means for detecting customer harassment" is a function that uses a natural language processing module to automatically identify harassing behavior in customer comments.

[0087] The "means of switching to AI response" is a mechanism that automatically switches from an operator to an AI module when customer harassment is detected.

[0088] "Means for generating an AI response" refers to the function of using a generative AI model to create an appropriate response and output the result in digital form.

[0089] "Means for transmitting AI responses to terminals" refers to the function of transferring the generated AI responses to terminals in real time, and having an operator or automated response system provide the responses to customers.

[0090] "Means for notifying customers of the introduction of the system" refers to a method or device for notifying customers in advance that the system will be used in the call center.

[0091] This invention is an AI system that reduces the workload of call center operators and improves productivity and workplace satisfaction. The system consists of terminals, a server, and a management console.

[0092] composition

[0093] Terminal: The telephone terminal used by the operator is a communication device for making calls with customers. As an example, we will use a general IP phone.

[0094] Server: A server is a computing device that hosts the speech recognition module and the natural language processing module. As a concrete example, a typical server computer is used.

[0095] Management Console: The management console is an interface that allows operators and administrators to monitor and manage the state of the system. For example, we will use a web-based interface.

[0096] Program processing

[0097] The server acquires, analyzes, and generates a response based on the following procedure:

[0098] Acquiring voice input

[0099] The terminal captures the voice conversation between the operator and the customer in real time, and then immediately transmits the captured voice data to the server, where it is converted into WAV or MP3 format.

[0100] Voice Recognition

[0101] The server stores the received voice data in memory and converts it into text data using the Google Speech-to-Text API, which is then used for the next natural language processing step.

[0102] Real-time analysis of conversation content

[0103] The server passes the text data to the Google Cloud Natural Language API, which analyzes the conversation in real time. The NLP module uses historical data and keyword filtering to detect harassment.

[0104] Customer Harassment Detection

[0105] The server receives the results of the NLP analysis and checks for flags of customer harassment. If harassment is detected, the server automatically switches to the AI-enabled module.

[0106] AI-powered response

[0107] The server uses the OpenAI GPT-4 API to send prompts to generate appropriate AI responses, which are then sent to the device and provided to the customer via an operator or automated response system.

[0108] Advance notice to customers

[0109] The user will notify the customer in advance that this system is being implemented, and this notification will include the purpose of the AI ​​preventing harassment.

[0110] As a concrete example, let's consider the flow when an operator receives a customer utterance, "Poor service! Call the person in charge!" The device captures the voice and sends it to the server. The server converts the voice into text, and the NLP module detects harassment. The server switches to the AI-enabled module, which generates an appropriate AI response. This response is sent to the device, and the operator or automated response system responds to the customer.

[0111] An example of a prompt is "Operator: What should I do if I receive harassing comments such as, 'Your service is bad! Call the person in charge!'"

[0112] As described above, this system reduces the burden on operators and enables quick and appropriate responses.

[0113] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0114] Step 1: Getting voice input

[0115] The terminal captures the voice exchanged between the operator and the customer as soon as the call begins. The terminal converts the captured voice data into WAV or MP3 format and sends it to the server in real time. The input is voice data, and the output is the converted voice file.

[0116] Step 2: Receiving and storing audio data

[0117] The server receives the audio file sent from the terminal and temporarily stores it in memory. The input here is the audio file sent from the terminal, and the output is the temporarily stored audio data.

[0118] Step 3: Voice Recognition

[0119] The server calls the Google Speech-to-Text API to convert the stored audio data into text data. Specifically, it sends the audio file to the API and receives the returned text data. The input is the audio data, and the output is the converted text data.

[0120] Step 4: Add text data to the queue

[0121] The text data acquired by the server is added to a queue for further processing. The input is the text data obtained by speech recognition, and the output is the text data added to the queue.

[0122] Step 5: Real-time analysis of conversation content

[0123] The server sends the text data to the Google Cloud Natural Language API, which analyzes the conversation in real time. The NLP module uses historical data and keyword filtering to identify harassing behavior in the text data. The input is the text data removed from the queue, and the output is the analysis results.

[0124] Step 6: Detect customer harassment

[0125] The server receives the results of the NLP analysis and determines whether or not customer harassment has occurred. Specifically, it determines this by checking the flags and evaluation values ​​of the analysis results. The input is the analysis results, and the output is a flag or evaluation value indicating whether or not harassment has occurred.

[0126] Step 7: Switch to automatic AI support

[0127] When the server detects customer harassment, it automatically switches from the operator to the AI-enabled module. Specifically, the server generates instructions to hand over the response to the AI. The input is the harassment detection result, and the output is an instruction to start the AI-enabled module.

[0128] Step 8: Generate an AI response

[0129] The server calls the OpenAI GPT-4 API and provides prompts to generate appropriate responses to harassment. Specifically, it sends the prompt sentence as input to the API and obtains the returned response. The input is the prompt sentence, and the output is the generated AI response.

[0130] Step 9: Sending an AI Response

[0131] The server sends the generated AI response to the terminal and provides it to the customer through an operator or an automated response system. The input is the generated AI response, and the output is the AI ​​response sent to the terminal.

[0132] Step 10: Notify customers in advance

[0133] The user notifies the customer in advance that the system is in place and that the AI ​​is intended to prevent harassment. Specifically, voice guidance and pre-recorded messages are used. The input is a notification message, and the output is a notification to the customer.

[0134] (Application example 1)

[0135] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0136] In conventional call center systems, operators are often exposed to harassment and stressful interactions from customers, which increases their workload and reduces workplace satisfaction. In addition, there are also situations in which staff in brick-and-mortar stores feel stressed when interacting with customers, raising concerns that this could lead to a decline in work efficiency. To solve these problems, a system is needed that reduces the stress that occurs during customer interactions and improves work efficiency.

[0137] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0138] In this invention, the server includes means for acquiring voice input from a terminal used by an operator, means for analyzing the acquired voice data in real time, means for converting the analyzed voice data into text, means for performing natural language processing based on the converted text data to detect customer harassment, means for automatically switching from the operator to AI response when customer harassment is detected, means for notifying customers in advance that the system has been introduced, and means for detecting harassment while staff are interacting with customers in a physical store and for AI to take over the response. This makes it possible to detect harassment while serving customers and for AI to respond, reducing the workload of staff and enabling work efficiency and improved customer satisfaction.

[0139] The "terminal used by the operator" refers to an information processing device used by the operator to interact with the customer, and has voice input and output functions.

[0140] "Means for acquiring voice input" refers to a device or system configuration that has the function of capturing voices emitted by operators or customers in real time and acquiring them as digital data.

[0141] "Means for analyzing acquired voice data in real time" refers to a device or program that has the function of processing captured voice data in real time and performing the necessary analysis.

[0142] "Means for converting analyzed voice data into text" refers to a device or system configuration that utilizes voice recognition technology to automatically convert voice data into text data.

[0143] "Means for detecting customer harassment through natural language processing" refers to a device or system configuration that has the function of determining whether or not customer harassment has occurred by utilizing machine learning algorithms and keyword filtering based on text data.

[0144] "Means for automatically switching from human to AI response" refers to a device or system configuration that has the functionality to allow an AI system to automatically take over response when customer harassment is detected.

[0145] "Means of notifying customers in advance of the system implementation" refers to methods and means for communicating information and precautions regarding the system implementation to customers in advance.

[0146] "Means for detecting harassment while staff are interacting with customers in a physical store, and for AI to take over the response" refers to a device or system configuration that has the functionality to enable a system to detect customer harassment while staff are interacting with customers in a physical store, and for AI to take over the interaction.

[0147] This invention is a system that detects harassment that occurs when staff interact with customers in brick-and-mortar stores and transfers the response to an AI system, thereby reducing the workload of staff and improving work efficiency. The specific configuration and operation of the system are described in detail below.

[0148] System configuration

[0149] The system mainly consists of the following components:

[0150] 1. Terminals: Devices such as smart glasses or smartphones used by staff in brick-and-mortar stores.

[0151] 2. Server: A computing device that hosts the speech recognition module and the natural language processing (NLP) module.

[0152] 3. Management Console: The interface for managing and configuring the entire system.

[0153] Hardware and Software

[0154] Hardware: Smart glasses or smartphones are used to capture customer conversations in real time.

[0155] Software: We use the Google Cloud Speech-to-Text API to convert voice data to text, and Hugging Face Transformers for NLP analysis.

[0156] Data processing: Converts voice data into text, detects customer harassment through NLP analysis, and generates appropriate AI responses.

[0157] Acquiring and analyzing voice input

[0158] The device captures the customer's conversation in real time and sends the audio data to the server, which then receives the data and converts it into text using the Google Cloud Speech-to-Text API.

[0159] Natural language processing of text data

[0160] The converted text data is then analyzed by an NLP module, which uses Hugging Face's Transformers library to detect customer harassment. The NLP module performs the analysis using historical data and keyword filtering.

[0161] AI-powered response

[0162] If customer harassment is detected, the server notifies the operator and automatically switches to the AI-enabled module, which generates an appropriate AI response based on the analysis results and sends it to the terminal in real time. This relieves staff from psychological burden and improves work efficiency.

[0163] Advance notice to customers

[0164] Users will notify their customers in advance that the system is in place, which includes AI-based detection of harassing behavior and automatic response measures.

[0165] Specific examples

[0166] For example, if a store staff member is interacting with a customer and the customer says, "I don't like the product! Call the manager!", the system will detect this. The NLP module analyzes the text data, and if it recognizes harassment, the server switches to the AI-enabled module. The AI ​​generates a response such as, "Sir, we'll help you find a quieter location," which is then conveyed to the customer via their device.

[0167] Prompt Sentence Examples

[0168] "How would you respond when a customer says, 'I don't like the product! Call the store manager!' What would be an appropriate response?"

[0169] In this way, the system reduces the burden on staff and contributes to improved work efficiency and customer satisfaction.

[0170] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0171] Step 1:

[0172] The terminal captures the conversation with the customer in real time and sends the voice data to the server. At this time, the terminal's voice input function is used to acquire the input voice data as a digital signal. The input data is then transferred from the terminal to the server.

[0173] Step 2:

[0174] The server converts the received voice data into text data using the Google Cloud Speech-to-Text API. The speech recognition module analyzes the voice data and generates corresponding text data. The server analyzes the input voice data and outputs it as text data.

[0175] Step 3:

[0176] The server sends the converted text data to a natural language processing (NLP) module to analyze the text data. It uses Hugging Face's Transformers library to detect harassing remarks in the text data. For analysis, the NLP module receives the text data as input and outputs the customer harassment detection results.

[0177] Step 4:

[0178] The server receives the analysis results from the NLP module and automatically switches to the AI-enabled module if harassment is detected. The instruction to switch is based on the harassment detection results as input data, and switching to the AI-enabled module as output.

[0179] Step 5:

[0180] The AI-enabled module generates an appropriate AI response based on the analysis results. The generative AI model creates the generative AI response in real time and sends it to the device. It receives the analysis results as input and outputs the appropriate response sentence.

[0181] Step 6:

[0182] The device receives the AI ​​response and conveys it to the customer as a voice output. The AI-generated response is provided to the customer via smart glasses or a smartphone. The AI ​​response that serves as input is converted into voice output and communicated to the customer.

[0183] Step 7:

[0184] The user notifies the customer in advance about the system's implementation and its functions. This notification uses prompts to explain information about the system's implementation and the AI's harassment detection function. The content of the notification is the input data, and the message to the customer is the output.

[0185] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0186] This invention is an AI system that reduces the burden on call center operators and enables efficient responses when they are faced with customer complaints or customer harassment. Furthermore, by combining it with an emotion engine that recognizes user emotions, it improves the quality of customer service.

[0187] composition

[0188] The system consists of the following main components:

[0189] 1. Terminal: A telephone terminal for the operator.

[0190] 2. Server: A computing device that hosts the AI ​​model, speech recognition module, and emotion engine.

[0191] 3. Management Console: Operator and administrator interface.

[0192] Program processing

[0193] Acquiring voice input

[0194] The terminal captures the voice conversation between the operator and the customer in real time and sends the voice data to the server, which receives the voice data and transfers it directly to the voice recognition module.

[0195] Voice Recognition

[0196] The server converts the received voice data into text data using a speech recognition module, which is then used for subsequent natural language processing (NLP) and emotion engines.

[0197] Real-time analysis of conversation content and emotion recognition

[0198] The server inputs the received text data into the NLP module to analyze the conversation content. At the same time, the emotion engine analyzes the emotion based on the user's voice tone, speed, and text content. The analysis results of the emotion engine are integrated with the results of the NLP module.

[0199] Customer Harassment Detection

[0200] The server receives the analysis results from the NLP module, and if customer harassment is detected, it notifies the operator and automatically switches to the AI-enabled module.

[0201] AI-powered response

[0202] The AI-enabled module generates an appropriate response, dynamically adjusting the AI ​​response content based on the analysis results of the emotion engine, and transmitting the generated AI response to the device in real time.

[0203] Advance notice to customers

[0204] When the system is introduced, the user will provide advance notice to customers, including the fact that the system is being introduced to prevent customer harassment and that it also has emotion recognition capabilities.

[0205] Specific examples

[0206] For example, when an operator receives a call from a customer, the voice is captured on the device and immediately sent to the server. The server converts the voice into text and analyzes it using the NLP module, while the emotion engine analyzes the tone and speed of the voice. If a customer makes a harassing remark such as, "Poor service! Call the manager!" and the emotion engine detects a high level of anger, the server will comprehensively assess this, notify the operator, and switch to the AI-enabled module. The AI ​​will then generate an appropriate response based on the emotion, such as, "We apologize for the inconvenience," and respond to the customer via the device. This process reduces the operator's burden and improves the quality of customer service.

[0207] In this way, the system reduces the burden on operators and contributes to improving work efficiency and customer satisfaction.

[0208] The processing flow will be explained below.

[0209] Step 1:

[0210] The terminal captures voice inputs exchanged between the operator and the customer in real time, and the captured voice data is immediately sent to the server.

[0211] Step 2:

[0212] The server passes the received voice data to a speech recognition module, which converts the voice into text. This conversion process is performed in real time, and the resulting text data is generated.

[0213] Step 3:

[0214] The server then inputs the generated text data into a natural language processing (NLP) module to analyze the conversation, but at this stage it does not yet make a judgment as to whether or not customer harassment has occurred.

[0215] Step 4:

[0216] In parallel, the server passes the text data and voice data to the emotion engine, which analyzes the user's emotions based on the tone, speed, and text content of the voice.

[0217] Step 5:

[0218] The server integrates the analysis results from the NLP module and the emotion engine to determine whether customer harassment has occurred and the user's emotional state. If the emotion engine detects high emotions such as "anger" or "irritation," the possibility of harassment increases.

[0219] Step 6:

[0220] If the server detects customer harassment and high emotions, it will notify the operator and automatically switch to the AI-enabled module, including the specific part that was deemed to be harassment.

[0221] Step 7:

[0222] The server then activates an AI-enabled module to generate an appropriate response that takes into account the user's emotional state. For example, if the emotion engine detects "anger," it will generate a sympathetic response such as, "We're sorry. We're very sorry for the inconvenience."

[0223] Step 8:

[0224] The server generates an AI response and sends it to the terminal in real time. The terminal then forwards the AI ​​response to the customer, and the operator responds on their behalf, allowing the customer to continue their conversation.

[0225] Step 9:

[0226] When the system is introduced, the user will provide advance notice to customers about its use, including that the AI ​​will prevent customer harassment and take appropriate action using emotion recognition.

[0227] By doing so, the system of the present invention reduces the workload of operators and realizes more sophisticated customer service. In addition, by taking into account the user's emotions, it is expected that customer satisfaction will also improve.

[0228] Example 2

[0229] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0230] In modern call center operations, operators are burdened with dealing with customer harassment and emotional complaints from customers. This burden increases the operators' mental stress, reduces work efficiency, and ultimately leads to a decline in customer satisfaction. Conventional systems did not provide an effective means to solve these problems.

[0231] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0232] In this invention, the server includes means for acquiring voice input from a terminal used by an operator, means for analyzing the acquired voice data in real time, means for converting the analyzed voice data into text, means for performing natural language processing and emotion recognition based on the converted text data to detect customer harassment, means for automatically switching from the operator to AI response when customer harassment is detected and generating a response based on the emotion recognition results, and means for notifying customers in advance of the introduction of the system. This reduces the burden on operators and enables efficient customer service that takes emotions into consideration.

[0233] "Terminal" refers to the communications equipment used by the operator and is a device for obtaining voice input.

[0234] A "server" is a computing device that performs the main processing of the system, such as analyzing voice data, natural language processing, emotion recognition, and detecting customer harassment.

[0235] "Voice input" refers to the voice data exchanged between the operator and the customer.

[0236] "Means for analyzing in real time" refers to a processing method or system configuration for instantly analyzing voice data acquired from a terminal.

[0237] A "means for converting voice data to text" is a processing method or system configuration for converting voice data to text data using voice recognition technology.

[0238] "Text data" is voice data converted into character information.

[0239] "Natural language processing" is a technical field that understands and analyzes the meaning of text data as human language.

[0240] "Emotion recognition" is a technology for identifying a speaker's emotions based on voice and text data.

[0241] "Customer harassment" refers to inappropriate comments or actions made by customers toward operators.

[0242] "AI-enabled" is the process of using artificial intelligence to automatically generate appropriate responses and assist operators.

[0243] "Harassment detection result" is information that indicates whether customer harassment is occurring based on the results of natural language processing and emotion recognition.

[0244] "Means for notifying customers in advance of the introduction of the system" refers to the methods and processes for informing customers in advance that the system will be used.

[0245] This invention is a system that reduces the burden on call center operators when dealing with customer complaints and harassment, and realizes efficient and emotion-sensitive responses. In particular, by combining it with an emotion engine that recognizes the user's emotions, the quality of customer service is improved.

[0246] System configuration

[0247] Hardware and software used

[0248] 1. Terminal: A communication device for an operator, capable of capturing voice input (e.g., a telephone or headset).

[0249] 2. Server: A computing device that performs the main processing of the system, such as analyzing voice data, natural language processing, emotion recognition, and customer harassment detection. Specific software examples include Google Speech-to-Text API, OpenAI GPT-4, and IBM Watson Tone Analyzer.

[0250] 3. Management console: An interface for operators and administrators, software for configuring and monitoring the system (e.g., a web-based management screen).

[0251] Operation procedures and examples

[0252] Acquiring voice input

[0253] The terminal captures the voices exchanged between the operator and the customer in real time. For example, when an operator receives a call from a customer, the conversation is picked up through the terminal's microphone and immediately sent to the server.

[0254] Voice Recognition

[0255] The server receives the voice data and converts it into text using the Google Speech-to-Text API. The server sends the voice data to the API and receives the text data within a few seconds.

[0256] Real-time analysis of conversation content and emotion recognition

[0257] The server inputs the received text data into a natural language processing module (OpenAI GPT-4) to analyze the content of the conversation. At the same time, an emotion engine (IBM Watson Tone Analyzer) analyzes emotions based on the tone of voice, speed, and text content. The analysis results are integrated and used for subsequent processing.

[0258] Customer Harassment Detection

[0259] The server checks the analysis results from the NLP module, and if it determines that certain words or phrases constitute customer harassment, it immediately notifies the operator and automatically switches to the AI-enabled module. For example, if the text data contains words such as "call the person in charge" or "the service is terrible," the server will detect harassment.

[0260] AI-powered response

[0261] The server activates an AI-enabled module to generate an appropriate response. At this time, a generative AI model (OpenAI GPT-4) dynamically adjusts the response content based on the results of sentiment analysis. For example, if a customer says, "The service was poor," the AI ​​generates a response such as, "We're sorry, we apologize for the inconvenience," and sends it to the device in real time.

[0262] Advance notice to customers

[0263] When introducing the system, the user notifies customers in advance, for example by email or letter, that "this system has been introduced with the aim of preventing customer harassment and is equipped with emotion recognition functionality."

[0264] Prompt Sentence Examples

[0265] The AI ​​model is fed with the following prompt:

[0266] "If the customer is enraged, generate an effective calming response. Their words are, 'Poor service! Bring out the person in charge!'"

[0267] "Please write an appropriate apology for the user's comments. The situation is that a customer is expressing dissatisfaction with a product."

[0268] This allows the system to reduce the burden on operators and provide quick and emotionally sensitive customer service.

[0269] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0270] Step 1: Getting voice input

[0271] The terminal captures the voice exchanged between the operator and the customer in real time. Specifically, the terminal's microphone receives voice input and processes this voice data as a digital signal. The captured voice data is then sent to the server. For example, when an operator says, "Hello, thank you for your inquiry," the voice is immediately converted into digital data by the terminal and sent to the server. The input here is voice data, and the output is digital voice data.

[0272] Step 2: Voice Recognition

[0273] The server receives the voice data sent from the device and passes it to a voice recognition module. The Google Speech-to-Text API is used as the voice recognition module. When voice data is sent to this API, the API converts the voice data into text data, which is then received by the server. The input here is digital voice data, and the output is text data. Specifically, the server posts the voice data to Google's API and receives the text data within a few seconds.

[0274] Step 3: Real-time analysis of conversation content and emotion recognition

[0275] The server inputs the received text data into a natural language processing module (OpenAI GPT-4). The server then sends the text data to GPT-4, which analyzes the content of the dialogue. At the same time, the emotion engine (IBM Watson Tone Analyzer) analyzes the user's emotions based on the tone of voice, speed, and text content. The server then combines the results of these two analyses. The input here is text data, and the output is analysis results based on natural language processing and emotion recognition results. Specifically, the server sends the text data to GPT-4, where it performs natural language processing, while simultaneously sending it to IBM Watson Tone Analyzer for emotion analysis.

[0276] Step 4: Detect customer harassment

[0277] The server checks the analysis results from the natural language processing module and checks whether specific words or phrases are judged to be customer harassment. For example, if the text data contains words such as "Bring the person in charge" or "The service is terrible" and the emotion engine detects a high level of anger, these conditions will trigger the detection of harassment. The input is the results of natural language processing and emotion recognition, and the output is a customer harassment detection flag. In concrete terms, the server filters the analysis results based on the conditions and determines whether harassment has occurred.

[0278] Step 5: AI response

[0279] The server launches an AI-enabled module and generates an appropriate response. At this time, the generative AI model (OpenAI GPT-4) dynamically adjusts the response content based on the results of emotion analysis. For example, if a customer is angry and complains about poor service, the AI ​​generates a response such as "We apologize for the inconvenience." The server then sends the generated AI response to the device in real time. The input is the emotion recognition result and harassment detection flag, and the output is the response message generated by the AI. Specifically, the server sends a prompt to GPT-4, which generates a response and sends the result to the device.

[0280] Step 6: Notify customers in advance

[0281] When the system is introduced, the user notifies customers in advance. For example, by email or letter, the user can explain that "this system has been introduced to prevent customer harassment and is equipped with emotion recognition functionality." The input is the notification content, and the output is the notification status to the customer. In concrete terms, the user references the customer database and sends a mass notification using an email system or postal system.

[0282] This detailed processing step reduces the burden on the operator and ensures a fast and empathetic customer service experience.

[0283] (Application example 2)

[0284] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0285] Traditional customer support and security operators are required to respond quickly and appropriately to harassment comments and emergency situations from customers, but this process is extremely stressful and places a heavy psychological burden on the operators. Furthermore, there is a lack of systems that can properly recognize customer emotions and generate optimal responses, which can lead to inconsistencies in the quality of customer support. There is a need for a system that can solve these issues, reduce the burden on operators, and provide efficient, high-quality customer support.

[0286] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0287] In this invention, the server includes means for acquiring voice input from a terminal used by an operator, means for analyzing the acquired voice data in real time, means for converting the analyzed voice data into text, means for performing natural language processing based on the converted text data to detect customer harassment, means for automatically switching from the operator to AI response when customer harassment is detected, means for notifying the customer in advance of the system implementation, means for recognizing emotions based on the converted text data and tone of voice, means for evaluating the risk level of customer harassment based on the detected emotion data, and means for detecting customer emergencies, identifying emotions of high tension or fear, and generating an appropriate response. This makes it possible to reduce the burden on operators and improve the quality of customer service.

[0288] "Terminal" means a communication device used by an operator to obtain voice input.

[0289] "Audio data" is an audio signal obtained from a terminal expressed as digital information.

[0290] "Real-time" refers to the fact that audio data is processed immediately after it is acquired.

[0291] "Text data" is voice data converted into text information using voice recognition technology.

[0292] "Natural language processing" is a technology that allows computers to understand and analyze human language.

[0293] "Customer harassment" refers to malicious words, actions, and behavior from customers toward operators.

[0294] "AI-enabled" refers to using artificial intelligence to automatically generate responses to customers.

[0295] "Emotion recognition" is a technology that analyzes voice data and text data to estimate the emotional state of a speaker.

[0296] "Risk level" is a numerical or index expression of the danger or urgency of customer harassment.

[0297] An "emergency incident" is a serious event or emergency that requires immediate action.

[0298] "Response generation" is the process of creating an appropriate response based on the analysis results.

[0299] The present invention relates to a system that enables an operator to efficiently process voice data from a customer and generate an appropriate response. The specific configuration and functions of this system are described below.

[0300] System Configuration

[0301] 1. Terminal

[0302] The terminal is a communication device that allows the operator to communicate with the customer and is used to obtain voice input. Specifically, this applies to smartphones and headsets.

[0303] 2. Server

[0304] The server is a high-performance computing device that hosts multiple functional modules, including a speech recognition module, a natural language processing module, an emotion recognition engine, and an AI response generation module.

[0305] 3. Management Console

[0306] The management console is an interface that allows operators and administrators to monitor the status of the system and change its settings.

[0307] Program processing

[0308] Acquiring voice input

[0309] The device captures conversations with customers in real time and sends the audio data to a server. The hardware used is the smartphone's microphone, and the software is a voice capture module.

[0310] Voice Recognition

[0311] The server receives the audio data and converts it into text using the Google Cloud Speech-to-Text API, which is then used for further processing.

[0312] Natural Language Processing and Emotion Recognition

[0313] The server passes the text data obtained from the voice recognition module to the natural language processing module (SpaCy) and analyzes the conversation. At the same time, it uses the emotion recognition engine (Emotion Recognition API) to analyze the customer's emotions based on the tone and speed of their voice. This allows it to identify the customer's emotional state, such as tension or fear.

[0314] Customer Harassment Detection

[0315] A natural language processing module is used to analyze text data to detect customer harassment, along with historical data and keyword filtering.

[0316] Risk Assessment

[0317] The risk level of customer harassment is assessed based on emotional data obtained from the emotion recognition module. If high levels of tension or fear are detected, the risk level is assessed as high.

[0318] AI-powered response generation

[0319] If customer harassment or a high risk level is detected, the server automatically activates the AI ​​response module to generate an appropriate response, which is then sent to the terminal in real time and provided to the customer. This uses OpenAI's GPT-4 model.

[0320] Specific examples

[0321] For example, when a customer reports an emergency situation on a security hotline, such as "Help! My window is broken," the device captures the voice and sends it to a server. The server converts the voice data into text and performs natural language processing and emotion recognition. If the server detects a high level of fear based on the analysis results, it notifies the operator, and the AI ​​immediately generates an appropriate response, such as "We will immediately dispatch a security team. Please evacuate to a safe location."

[0322] Prompt Sentence Examples

[0323] Audio data: "Help! The window is broken."

[0324] Prompt statement:

[0325] "Analyze the following voice input, understand the customer's sentiment, and generate an appropriate response.

[0326] Voice input: "Help! The window is broken."

[0327] Analysis: Customers have high levels of fear.

[0328] Response generated:"

[0329] Example response: 'We will be sending a security team immediately. Please seek safety.'

[0330] This reduces the burden on operators and improves the quality of customer service.

[0331] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0332] Step 1:

[0333] Acquiring voice input

[0334] The device captures the conversation with the customer in real time and sends the voice data to the server. The input is the customer's voice, and the output is digital voice data. In this process, the smartphone or headset acts as a microphone, and the voice capture module converts the voice signal into digital information.

[0335] Step 2:

[0336] Voice Recognition

[0337] The server converts the received voice data into text data using the Google Cloud Speech-to-Text API. The input is digital voice data, and the output is text data. In this process, a speech recognition model analyzes the voice waveform and generates corresponding text.

[0338] Step 3:

[0339] Natural Language Processing and Emotion Recognition

[0340] The server inputs the text data into a natural language processing module such as SpaCy to analyze the conversation content. At the same time, it uses the Emotion Recognition API to analyze the customer's emotional state from the text and voice features. The input is text data and voice feature data, and the output is the analyzed conversation content and emotion data. In this process, the natural language processing module analyzes the meaning of words and phrases, and the emotion recognition engine analyzes the tone and speed of the voice to identify emotions.

[0341] Step 4:

[0342] Customer Harassment Detection

[0343] The server uses the analysis results from the natural language processing module to detect customer harassment. Past data and keyword filtering are also used. The input is the analysis result of the text, and the output is the detection result of customer harassment. This process checks whether specific keywords or phrases are included in the text to determine whether harassing behavior is detected.

[0344] Step 5:

[0345] Risk Assessment

[0346] The server uses the emotion data obtained from the emotion recognition module to evaluate the risk level of customer harassment. The input is the customer's emotion data, and the output is the risk level evaluation result. In this process, if high tension or fear is detected, the risk level is set high.

[0347] Step 6:

[0348] AI-powered response generation

[0349] When the server detects customer harassment or a high risk level, it activates an AI response module to generate an appropriate response. The generated AI response is sent to the terminal in real time and provided to the customer. The input is the risk level and harassment detection result, and the output is the generated appropriate response. This process utilizes OpenAI's GPT-4 model to generate a response with the appropriate context and tone.

[0350] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0351] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0352] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0353] [Second embodiment]

[0354] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0355] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0356] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0357] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0358] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0359] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0360] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0361] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0362] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0363] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0364] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0365] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0366] This invention is an AI system for reducing the workload of call center operators and improving productivity and workplace satisfaction.

[0367] composition

[0368] The system consists of the following main components:

[0369] 1. Terminal: A telephone terminal for the operator.

[0370] 2. Server: A computing device that hosts the AI ​​model and speech recognition module.

[0371] 3. Management Console: Operator and administrator interface.

[0372] Program processing

[0373] Acquiring voice input

[0374] The terminal captures the voice conversation between the operator and the customer in real time and sends the voice data to the server, which receives the voice data and transfers it directly to the voice recognition module.

[0375] Voice Recognition

[0376] The server converts the received voice data into text data using a speech recognition module, which is then used for subsequent natural language processing (NLP).

[0377] Real-time analysis of conversation content

[0378] The server inputs the converted text data into an NLP module to detect customer harassment, which uses historical data and keyword filtering for analysis.

[0379] Customer Harassment Detection

[0380] The server receives the analysis results from the NLP module, and if customer harassment is detected, it notifies the operator and automatically switches to the AI-enabled module.

[0381] AI-powered response

[0382] The AI-enabled module takes over the response and sends the generated AI response to the terminal in real time, which relieves the operator of the psychological burden and allows the appropriate response to be taken.

[0383] Advance notice to customers

[0384] The user will notify the customer in advance that this system is being implemented, including the purpose of the AI ​​to prevent harassment.

[0385] Specific examples

[0386] For example, when an operator receives a call from a customer, the voice is captured on the terminal and immediately sent to the server. The server converts the voice into text and analyzes it using the NLP module. If the customer makes a harassing remark such as, "Poor service! Call the manager!", the NLP module detects this and the server switches to the AI-enabled module. The AI ​​generates an appropriate response and responds to the customer via the terminal. This process reduces the burden on the operator and enables efficient business operations.

[0387] In this way, the system reduces the burden on operators and contributes to improving work efficiency and customer satisfaction.

[0388] The processing flow will be explained below.

[0389] Step 1:

[0390] The terminal captures voice inputs exchanged between the operator and the customer in real time, and the captured voice data is immediately sent to the server.

[0391] Step 2:

[0392] The server passes the received voice data to a speech recognition module, which converts the voice into text. This conversion process is performed in real time, and the resulting text data is generated.

[0393] Step 3:

[0394] The server feeds the generated text data into a natural language processing (NLP) module to analyze the conversation, which uses keyword-based filtering and historical data to detect customer harassment.

[0395] Step 4:

[0396] The server checks the analysis results received from the NLP module, and if customer harassment is detected, it notifies an operator and passes the response to the AI-enabled module.

[0397] Step 5:

[0398] The server activates the AI-enabled module and generates an appropriate response, which is then sent to the device in real time.

[0399] Step 6:

[0400] The terminal receives the AI ​​response from the server and forwards it to the customer, which relieves the operator of the psychological burden.

[0401] Step 7:

[0402] When the user installs the system, the user will provide advance notice to the customer, including the fact that the system is being installed for the purpose of preventing customer harassment.

[0403] Example 1

[0404] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0405] In conventional call centers, operators have to deal with harassing comments from customers, which places a heavy psychological burden on them, often resulting in a decline in productivity and workplace satisfaction. Furthermore, if appropriate responses are not made promptly, customer satisfaction also declines, which is an issue.

[0406] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0407] In this invention, the server includes means for acquiring voice input from a terminal used by an operator, means for transmitting the acquired voice data to a computing device in real time, means for the computing device to analyze the transmitted voice data using a voice recognition module, means for converting the voice data analyzed by the voice recognition module into text data, means for inputting the converted text data into a natural language processing module and analyzing the conversation content in real time, means for detecting customer harassment from the analysis results using the natural language processing module, means for automatically switching from the operator to an AI response when customer harassment is detected, means for transmitting the generated AI response to the terminal in real time to respond to the customer, and means for notifying the customer in advance of the introduction of the system. This reduces the psychological burden on the operator and enables quick and appropriate response.

[0408] "Terminal" refers to the communication device used by the operator, which is a hardware device for making calls with customers.

[0409] "Means for acquiring voice input" refers to a device or method that allows the terminal to capture the voice exchanged between the operator and the customer in real time and acquire it as digital data.

[0410] "Means for transmitting voice data in real time" refers to a device or method that has a mechanism for instantly transferring voice data acquired from a terminal to a server.

[0411] A "computing device" is a computer system for analyzing and processing audio data. Also called a server.

[0412] A "voice recognition module" is software or algorithms for converting voice data into text data.

[0413] The "means for converting into text data" is a function that uses a voice recognition module to convert voice data into text information and output it in digital text format.

[0414] A "natural language processing module" is software or algorithms for analyzing text data and understanding and processing its content.

[0415] The "means for detecting customer harassment" is a function that uses a natural language processing module to automatically identify harassing behavior in customer comments.

[0416] The "means of switching to AI response" is a mechanism that automatically switches from an operator to an AI module when customer harassment is detected.

[0417] "Means for generating an AI response" refers to the function of using a generative AI model to create an appropriate response and output the result in digital form.

[0418] "Means for transmitting AI responses to terminals" refers to the function of transferring the generated AI responses to terminals in real time, and having an operator or automated response system provide the responses to customers.

[0419] "Means for notifying customers of the introduction of the system" refers to a method or device for notifying customers in advance that the system will be used in the call center.

[0420] This invention is an AI system that reduces the workload of call center operators and improves productivity and workplace satisfaction. The system consists of terminals, a server, and a management console.

[0421] composition

[0422] Terminal: The telephone terminal used by the operator is a communication device for making calls with customers. As an example, we will use a general IP phone.

[0423] Server: A server is a computing device that hosts the speech recognition module and the natural language processing module. As a concrete example, a typical server computer is used.

[0424] Management Console: The management console is an interface that allows operators and administrators to monitor and manage the state of the system. For example, we will use a web-based interface.

[0425] Program processing

[0426] The server acquires, analyzes, and generates a response based on the following procedure:

[0427] Acquiring voice input

[0428] The terminal captures the voice conversation between the operator and the customer in real time, and then immediately transmits the captured voice data to the server, where it is converted into WAV or MP3 format.

[0429] Voice Recognition

[0430] The server stores the received voice data in memory and converts it into text data using the Google Speech-to-Text API, which is then used for the next natural language processing step.

[0431] Real-time analysis of conversation content

[0432] The server passes the text data to the Google Cloud Natural Language API, which analyzes the conversation in real time. The NLP module uses historical data and keyword filtering to detect harassment.

[0433] Customer Harassment Detection

[0434] The server receives the results of the NLP analysis and checks for flags of customer harassment. If harassment is detected, the server automatically switches to the AI-enabled module.

[0435] AI-powered response

[0436] The server uses the OpenAI GPT-4 API to send prompts to generate appropriate AI responses, which are then sent to the device and provided to the customer via an operator or automated response system.

[0437] Advance notice to customers

[0438] The user will notify the customer in advance that this system is being implemented, and this notification will include the purpose of the AI ​​preventing harassment.

[0439] As a concrete example, let's consider the flow when an operator receives a customer utterance, "Poor service! Call the person in charge!" The device captures the voice and sends it to the server. The server converts the voice into text, and the NLP module detects harassment. The server switches to the AI-enabled module, which generates an appropriate AI response. This response is sent to the device, and the operator or automated response system responds to the customer.

[0440] An example of a prompt is "Operator: What should I do if I receive harassing comments such as, 'Your service is bad! Call the person in charge!'"

[0441] As described above, this system reduces the burden on operators and enables quick and appropriate responses.

[0442] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0443] Step 1: Getting voice input

[0444] The terminal captures the voice exchanged between the operator and the customer as soon as the call begins. The terminal converts the captured voice data into WAV or MP3 format and sends it to the server in real time. The input is voice data, and the output is the converted voice file.

[0445] Step 2: Receiving and storing audio data

[0446] The server receives the audio file sent from the terminal and temporarily stores it in memory. The input here is the audio file sent from the terminal, and the output is the temporarily stored audio data.

[0447] Step 3: Voice Recognition

[0448] The server calls the Google Speech-to-Text API to convert the stored audio data into text data. Specifically, it sends the audio file to the API and receives the returned text data. The input is the audio data, and the output is the converted text data.

[0449] Step 4: Add text data to the queue

[0450] The text data acquired by the server is added to a queue for further processing. The input is the text data obtained by speech recognition, and the output is the text data added to the queue.

[0451] Step 5: Real-time analysis of conversation content

[0452] The server sends the text data to the Google Cloud Natural Language API, which analyzes the conversation in real time. The NLP module uses historical data and keyword filtering to identify harassing behavior in the text data. The input is the text data removed from the queue, and the output is the analysis results.

[0453] Step 6: Detect customer harassment

[0454] The server receives the results of the NLP analysis and determines whether or not customer harassment has occurred. Specifically, it determines this by checking the flags and evaluation values ​​of the analysis results. The input is the analysis results, and the output is a flag or evaluation value indicating whether or not harassment has occurred.

[0455] Step 7: Switch to automatic AI support

[0456] When the server detects customer harassment, it automatically switches from the operator to the AI-enabled module. Specifically, the server generates instructions to hand over the response to the AI. The input is the harassment detection result, and the output is an instruction to start the AI-enabled module.

[0457] Step 8: Generate an AI response

[0458] The server calls the OpenAI GPT-4 API and provides prompts to generate appropriate responses to harassment. Specifically, it sends the prompt sentence as input to the API and obtains the returned response. The input is the prompt sentence, and the output is the generated AI response.

[0459] Step 9: Sending an AI Response

[0460] The server sends the generated AI response to the terminal and provides it to the customer through an operator or an automated response system. The input is the generated AI response, and the output is the AI ​​response sent to the terminal.

[0461] Step 10: Notify customers in advance

[0462] The user notifies the customer in advance that the system is in place and that the AI ​​is intended to prevent harassment. Specifically, voice guidance and pre-recorded messages are used. The input is a notification message, and the output is a notification to the customer.

[0463] (Application example 1)

[0464] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0465] In conventional call center systems, operators are often exposed to harassment and stressful interactions from customers, which increases their workload and reduces workplace satisfaction. In addition, there are also situations in which staff in brick-and-mortar stores feel stressed when interacting with customers, raising concerns that this could lead to a decline in work efficiency. To solve these problems, a system is needed that reduces the stress that occurs during customer interactions and improves work efficiency.

[0466] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0467] In this invention, the server includes means for acquiring voice input from a terminal used by an operator, means for analyzing the acquired voice data in real time, means for converting the analyzed voice data into text, means for performing natural language processing based on the converted text data to detect customer harassment, means for automatically switching from the operator to AI response when customer harassment is detected, means for notifying customers in advance that the system has been introduced, and means for detecting harassment while staff are interacting with customers in a physical store and for AI to take over the response. This makes it possible to detect harassment while serving customers and for AI to respond, reducing the workload of staff and enabling work efficiency and improved customer satisfaction.

[0468] The "terminal used by the operator" refers to an information processing device used by the operator to interact with the customer, and has voice input and output functions.

[0469] "Means for acquiring voice input" refers to a device or system configuration that has the function of capturing voices emitted by operators or customers in real time and acquiring them as digital data.

[0470] "Means for analyzing acquired voice data in real time" refers to a device or program that has the function of processing captured voice data in real time and performing the necessary analysis.

[0471] "Means for converting analyzed voice data into text" refers to a device or system configuration that utilizes voice recognition technology to automatically convert voice data into text data.

[0472] "Means for detecting customer harassment through natural language processing" refers to a device or system configuration that has the function of determining whether or not customer harassment has occurred by utilizing machine learning algorithms and keyword filtering based on text data.

[0473] "Means for automatically switching from human to AI response" refers to a device or system configuration that has the functionality to allow an AI system to automatically take over response when customer harassment is detected.

[0474] "Means of notifying customers in advance of the system implementation" refers to methods and means for communicating information and precautions regarding the system implementation to customers in advance.

[0475] "Means for detecting harassment while staff are interacting with customers in a physical store, and for AI to take over the response" refers to a device or system configuration that has the functionality to enable a system to detect customer harassment while staff are interacting with customers in a physical store, and for AI to take over the interaction.

[0476] This invention is a system that detects harassment that occurs when staff interact with customers in brick-and-mortar stores and transfers the response to an AI system, thereby reducing the workload of staff and improving work efficiency. The specific configuration and operation of the system are described in detail below.

[0477] System configuration

[0478] The system mainly consists of the following components:

[0479] 1. Terminals: Devices such as smart glasses or smartphones used by staff in brick-and-mortar stores.

[0480] 2. Server: A computing device that hosts the speech recognition module and the natural language processing (NLP) module.

[0481] 3. Management Console: The interface for managing and configuring the entire system.

[0482] Hardware and Software

[0483] Hardware: Smart glasses or smartphones are used to capture customer conversations in real time.

[0484] Software: We use the Google Cloud Speech-to-Text API to convert voice data to text, and Hugging Face Transformers for NLP analysis.

[0485] Data processing: Converts voice data into text, detects customer harassment through NLP analysis, and generates appropriate AI responses.

[0486] Acquiring and analyzing voice input

[0487] The device captures the customer's conversation in real time and sends the audio data to the server, which then receives the data and converts it into text using the Google Cloud Speech-to-Text API.

[0488] Natural language processing of text data

[0489] The converted text data is then analyzed by an NLP module, which uses Hugging Face's Transformers library to detect customer harassment. The NLP module performs the analysis using historical data and keyword filtering.

[0490] AI-powered response

[0491] If customer harassment is detected, the server notifies the operator and automatically switches to the AI-enabled module, which generates an appropriate AI response based on the analysis results and sends it to the terminal in real time. This relieves staff from psychological burden and improves work efficiency.

[0492] Advance notice to customers

[0493] Users will notify their customers in advance that the system is in place, which includes AI-based detection of harassing behavior and automatic response measures.

[0494] Specific examples

[0495] For example, if a store staff member is interacting with a customer and the customer says, "I don't like the product! Call the manager!", the system will detect this. The NLP module analyzes the text data, and if it recognizes harassment, the server switches to the AI-enabled module. The AI ​​generates a response such as, "Sir, we'll help you find a quieter location," which is then conveyed to the customer via their device.

[0496] Prompt Sentence Examples

[0497] "How would you respond when a customer says, 'I don't like the product! Call the store manager!' What would be an appropriate response?"

[0498] In this way, the system reduces the burden on staff and contributes to improved work efficiency and customer satisfaction.

[0499] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0500] Step 1:

[0501] The terminal captures the conversation with the customer in real time and sends the voice data to the server. At this time, the terminal's voice input function is used to acquire the input voice data as a digital signal. The input data is then transferred from the terminal to the server.

[0502] Step 2:

[0503] The server converts the received voice data into text data using the Google Cloud Speech-to-Text API. The speech recognition module analyzes the voice data and generates corresponding text data. The server analyzes the input voice data and outputs it as text data.

[0504] Step 3:

[0505] The server sends the converted text data to a natural language processing (NLP) module to analyze the text data. It uses Hugging Face's Transformers library to detect harassing remarks in the text data. For analysis, the NLP module receives the text data as input and outputs the customer harassment detection results.

[0506] Step 4:

[0507] The server receives the analysis results from the NLP module and automatically switches to the AI-enabled module if harassment is detected. The instruction to switch is based on the harassment detection results as input data, and switching to the AI-enabled module as output.

[0508] Step 5:

[0509] The AI-enabled module generates an appropriate AI response based on the analysis results. The generative AI model creates the generative AI response in real time and sends it to the device. It receives the analysis results as input and outputs the appropriate response sentence.

[0510] Step 6:

[0511] The device receives the AI ​​response and conveys it to the customer as a voice output. The AI-generated response is provided to the customer via smart glasses or a smartphone. The AI ​​response that serves as input is converted into voice output and communicated to the customer.

[0512] Step 7:

[0513] The user notifies the customer in advance about the system's implementation and its functions. This notification uses prompts to explain information about the system's implementation and the AI's harassment detection function. The content of the notification is the input data, and the message to the customer is the output.

[0514] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0515] This invention is an AI system that reduces the burden on call center operators and enables efficient responses when they are faced with customer complaints or customer harassment. Furthermore, by combining it with an emotion engine that recognizes user emotions, it improves the quality of customer service.

[0516] composition

[0517] The system consists of the following main components:

[0518] 1. Terminal: A telephone terminal for the operator.

[0519] 2. Server: A computing device that hosts the AI ​​model, speech recognition module, and emotion engine.

[0520] 3. Management Console: Operator and administrator interface.

[0521] Program processing

[0522] Acquiring voice input

[0523] The terminal captures the voice conversation between the operator and the customer in real time and sends the voice data to the server, which receives the voice data and transfers it directly to the voice recognition module.

[0524] Voice Recognition

[0525] The server converts the received voice data into text data using a speech recognition module, which is then used for subsequent natural language processing (NLP) and emotion engines.

[0526] Real-time analysis of conversation content and emotion recognition

[0527] The server inputs the received text data into the NLP module to analyze the conversation content. At the same time, the emotion engine analyzes the emotion based on the user's voice tone, speed, and text content. The analysis results of the emotion engine are integrated with the results of the NLP module.

[0528] Customer Harassment Detection

[0529] The server receives the analysis results from the NLP module, and if customer harassment is detected, it notifies the operator and automatically switches to the AI-enabled module.

[0530] AI-powered response

[0531] The AI-enabled module generates an appropriate response, dynamically adjusting the AI ​​response content based on the analysis results of the emotion engine, and transmitting the generated AI response to the device in real time.

[0532] Advance notice to customers

[0533] When the system is introduced, the user will provide advance notice to customers, including the fact that the system is being introduced to prevent customer harassment and that it also has emotion recognition capabilities.

[0534] Specific examples

[0535] For example, when an operator receives a call from a customer, the voice is captured on the device and immediately sent to the server. The server converts the voice into text and analyzes it using the NLP module, while the emotion engine analyzes the tone and speed of the voice. If a customer makes a harassing remark such as, "Poor service! Call the manager!" and the emotion engine detects a high level of anger, the server will comprehensively assess this, notify the operator, and switch to the AI-enabled module. The AI ​​will then generate an appropriate response based on the emotion, such as, "We apologize for the inconvenience," and respond to the customer via the device. This process reduces the operator's burden and improves the quality of customer service.

[0536] In this way, the system reduces the burden on operators and contributes to improving work efficiency and customer satisfaction.

[0537] The processing flow will be explained below.

[0538] Step 1:

[0539] The terminal captures voice inputs exchanged between the operator and the customer in real time, and the captured voice data is immediately sent to the server.

[0540] Step 2:

[0541] The server passes the received voice data to a speech recognition module, which converts the voice into text. This conversion process is performed in real time, and the resulting text data is generated.

[0542] Step 3:

[0543] The server then inputs the generated text data into a natural language processing (NLP) module to analyze the conversation, but at this stage it does not yet make a judgment as to whether or not customer harassment has occurred.

[0544] Step 4:

[0545] In parallel, the server passes the text data and voice data to the emotion engine, which analyzes the user's emotions based on the tone, speed, and text content of the voice.

[0546] Step 5:

[0547] The server integrates the analysis results from the NLP module and the emotion engine to determine whether customer harassment has occurred and the user's emotional state. If the emotion engine detects high emotions such as "anger" or "irritation," the possibility of harassment increases.

[0548] Step 6:

[0549] If the server detects customer harassment and high emotions, it will notify the operator and automatically switch to the AI-enabled module, including the specific part that was deemed to be harassment.

[0550] Step 7:

[0551] The server then activates an AI-enabled module to generate an appropriate response that takes into account the user's emotional state. For example, if the emotion engine detects "anger," it will generate a sympathetic response such as, "We're sorry. We're very sorry for the inconvenience."

[0552] Step 8:

[0553] The server generates an AI response and sends it to the terminal in real time. The terminal then forwards the AI ​​response to the customer, and the operator responds on their behalf, allowing the customer to continue their conversation.

[0554] Step 9:

[0555] When the system is introduced, the user will provide advance notice to customers about its use, including that the AI ​​will prevent customer harassment and take appropriate action using emotion recognition.

[0556] By doing so, the system of the present invention reduces the workload of operators and realizes more sophisticated customer service. In addition, by taking into account the user's emotions, it is expected that customer satisfaction will also improve.

[0557] Example 2

[0558] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0559] In modern call center operations, operators are burdened with dealing with customer harassment and emotional complaints from customers. This burden increases the operators' mental stress, reduces work efficiency, and ultimately leads to a decline in customer satisfaction. Conventional systems did not provide an effective means to solve these problems.

[0560] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0561] In this invention, the server includes means for acquiring voice input from a terminal used by an operator, means for analyzing the acquired voice data in real time, means for converting the analyzed voice data into text, means for performing natural language processing and emotion recognition based on the converted text data to detect customer harassment, means for automatically switching from the operator to AI response when customer harassment is detected and generating a response based on the emotion recognition results, and means for notifying customers in advance of the introduction of the system. This reduces the burden on operators and enables efficient customer service that takes emotions into consideration.

[0562] "Terminal" refers to the communications equipment used by the operator and is a device for obtaining voice input.

[0563] A "server" is a computing device that performs the main processing of the system, such as analyzing voice data, natural language processing, emotion recognition, and detecting customer harassment.

[0564] "Voice input" refers to the voice data exchanged between the operator and the customer.

[0565] "Means for analyzing in real time" refers to a processing method or system configuration for instantly analyzing voice data acquired from a terminal.

[0566] A "means for converting voice data to text" is a processing method or system configuration for converting voice data to text data using voice recognition technology.

[0567] "Text data" is voice data converted into character information.

[0568] "Natural language processing" is a technical field that understands and analyzes the meaning of text data as human language.

[0569] "Emotion recognition" is a technology for identifying a speaker's emotions based on voice and text data.

[0570] "Customer harassment" refers to inappropriate comments or actions made by customers toward operators.

[0571] "AI-enabled" is the process of using artificial intelligence to automatically generate appropriate responses and assist operators.

[0572] "Harassment detection result" is information that indicates whether customer harassment is occurring based on the results of natural language processing and emotion recognition.

[0573] "Means for notifying customers in advance of the introduction of the system" refers to the methods and processes for informing customers in advance that the system will be used.

[0574] This invention is a system that reduces the burden on call center operators when dealing with customer complaints and harassment, and realizes efficient and emotion-sensitive responses. In particular, by combining it with an emotion engine that recognizes the user's emotions, the quality of customer service is improved.

[0575] System configuration

[0576] Hardware and software used

[0577] 1. Terminal: A communication device for an operator, capable of capturing voice input (e.g., a telephone or headset).

[0578] 2. Server: A computing device that performs the main processing of the system, such as analyzing voice data, natural language processing, emotion recognition, and customer harassment detection. Specific software examples include Google Speech-to-Text API, OpenAI GPT-4, and IBM Watson Tone Analyzer.

[0579] 3. Management console: An interface for operators and administrators, software for configuring and monitoring the system (e.g., a web-based management screen).

[0580] Operation procedures and examples

[0581] Acquiring voice input

[0582] The terminal captures the voices exchanged between the operator and the customer in real time. For example, when an operator receives a call from a customer, the conversation is picked up through the terminal's microphone and immediately sent to the server.

[0583] Voice Recognition

[0584] The server receives the voice data and converts it into text using the Google Speech-to-Text API. The server sends the voice data to the API and receives the text data within a few seconds.

[0585] Real-time analysis of conversation content and emotion recognition

[0586] The server inputs the received text data into a natural language processing module (OpenAI GPT-4) to analyze the content of the conversation. At the same time, an emotion engine (IBM Watson Tone Analyzer) analyzes emotions based on the tone of voice, speed, and text content. The analysis results are integrated and used for subsequent processing.

[0587] Customer Harassment Detection

[0588] The server checks the analysis results from the NLP module, and if it determines that certain words or phrases constitute customer harassment, it immediately notifies the operator and automatically switches to the AI-enabled module. For example, if the text data contains words such as "call the person in charge" or "the service is terrible," the server will detect harassment.

[0589] AI-powered response

[0590] The server activates an AI-enabled module to generate an appropriate response. At this time, a generative AI model (OpenAI GPT-4) dynamically adjusts the response content based on the results of sentiment analysis. For example, if a customer says, "The service was poor," the AI ​​generates a response such as, "We're sorry, we apologize for the inconvenience," and sends it to the device in real time.

[0591] Advance notice to customers

[0592] When introducing the system, the user notifies customers in advance, for example by email or letter, that "this system has been introduced with the aim of preventing customer harassment and is equipped with emotion recognition functionality."

[0593] Prompt Sentence Examples

[0594] The AI ​​model is fed with the following prompt:

[0595] "If the customer is enraged, generate an effective calming response. Their words are, 'Poor service! Bring out the person in charge!'"

[0596] "Please write an appropriate apology for the user's comments. The situation is that a customer is expressing dissatisfaction with a product."

[0597] This allows the system to reduce the burden on operators and provide quick and emotionally sensitive customer service.

[0598] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0599] Step 1: Getting voice input

[0600] The terminal captures the voice exchanged between the operator and the customer in real time. Specifically, the terminal's microphone receives voice input and processes this voice data as a digital signal. The captured voice data is then sent to the server. For example, when an operator says, "Hello, thank you for your inquiry," the voice is immediately converted into digital data by the terminal and sent to the server. The input here is voice data, and the output is digital voice data.

[0601] Step 2: Voice Recognition

[0602] The server receives the voice data sent from the device and passes it to a voice recognition module. The Google Speech-to-Text API is used as the voice recognition module. When voice data is sent to this API, the API converts the voice data into text data, which is then received by the server. The input here is digital voice data, and the output is text data. Specifically, the server posts the voice data to Google's API and receives the text data within a few seconds.

[0603] Step 3: Real-time analysis of conversation content and emotion recognition

[0604] The server inputs the received text data into a natural language processing module (OpenAI GPT-4). The server then sends the text data to GPT-4, which analyzes the content of the dialogue. At the same time, the emotion engine (IBM Watson Tone Analyzer) analyzes the user's emotions based on the tone of voice, speed, and text content. The server then combines the results of these two analyses. The input here is text data, and the output is analysis results based on natural language processing and emotion recognition results. Specifically, the server sends the text data to GPT-4, where it performs natural language processing, while simultaneously sending it to IBM Watson Tone Analyzer for emotion analysis.

[0605] Step 4: Detect customer harassment

[0606] The server checks the analysis results from the natural language processing module and checks whether specific words or phrases are judged to be customer harassment. For example, if the text data contains words such as "Bring the person in charge" or "The service is terrible" and the emotion engine detects a high level of anger, these conditions will trigger the detection of harassment. The input is the results of natural language processing and emotion recognition, and the output is a customer harassment detection flag. In concrete terms, the server filters the analysis results based on the conditions and determines whether harassment has occurred.

[0607] Step 5: AI response

[0608] The server launches an AI-enabled module and generates an appropriate response. At this time, the generative AI model (OpenAI GPT-4) dynamically adjusts the response content based on the results of emotion analysis. For example, if a customer is angry and complains about poor service, the AI ​​generates a response such as "We apologize for the inconvenience." The server then sends the generated AI response to the device in real time. The input is the emotion recognition result and harassment detection flag, and the output is the response message generated by the AI. Specifically, the server sends a prompt to GPT-4, which generates a response and sends the result to the device.

[0609] Step 6: Notify customers in advance

[0610] When the system is introduced, the user notifies customers in advance. For example, by email or letter, the user can explain that "this system has been introduced to prevent customer harassment and is equipped with emotion recognition functionality." The input is the notification content, and the output is the notification status to the customer. In concrete terms, the user references the customer database and sends a mass notification using an email system or postal system.

[0611] This detailed processing step reduces the burden on the operator and ensures a fast and empathetic customer service experience.

[0612] (Application example 2)

[0613] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0614] Traditional customer support and security operators are required to respond quickly and appropriately to harassment comments and emergency situations from customers, but this process is extremely stressful and places a heavy psychological burden on the operators. Furthermore, there is a lack of systems that can properly recognize customer emotions and generate optimal responses, which can lead to inconsistencies in the quality of customer support. There is a need for a system that can solve these issues, reduce the burden on operators, and provide efficient, high-quality customer support.

[0615] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0616] In this invention, the server includes means for acquiring voice input from a terminal used by an operator, means for analyzing the acquired voice data in real time, means for converting the analyzed voice data into text, means for performing natural language processing based on the converted text data to detect customer harassment, means for automatically switching from the operator to AI response when customer harassment is detected, means for notifying the customer in advance of the system implementation, means for recognizing emotions based on the converted text data and tone of voice, means for evaluating the risk level of customer harassment based on the detected emotion data, and means for detecting customer emergencies, identifying emotions of high tension or fear, and generating an appropriate response. This makes it possible to reduce the burden on operators and improve the quality of customer service.

[0617] "Terminal" means a communication device used by an operator to obtain voice input.

[0618] "Audio data" is an audio signal obtained from a terminal expressed as digital information.

[0619] "Real-time" refers to the fact that audio data is processed immediately after it is acquired.

[0620] "Text data" is voice data converted into text information using voice recognition technology.

[0621] "Natural language processing" is a technology that allows computers to understand and analyze human language.

[0622] "Customer harassment" refers to malicious words, actions, and behavior from customers toward operators.

[0623] "AI-enabled" refers to using artificial intelligence to automatically generate responses to customers.

[0624] "Emotion recognition" is a technology that analyzes voice data and text data to estimate the emotional state of a speaker.

[0625] "Risk level" is a numerical or index expression of the danger or urgency of customer harassment.

[0626] An "emergency incident" is a serious event or emergency that requires immediate action.

[0627] "Response generation" is the process of creating an appropriate response based on the analysis results.

[0628] The present invention relates to a system that enables an operator to efficiently process voice data from a customer and generate an appropriate response. The specific configuration and functions of this system are described below.

[0629] System Configuration

[0630] 1. Terminal

[0631] The terminal is a communication device that allows the operator to communicate with the customer and is used to obtain voice input. Specifically, this applies to smartphones and headsets.

[0632] 2. Server

[0633] The server is a high-performance computing device that hosts multiple functional modules, including a speech recognition module, a natural language processing module, an emotion recognition engine, and an AI response generation module.

[0634] 3. Management Console

[0635] The management console is an interface that allows operators and administrators to monitor the status of the system and change its settings.

[0636] Program processing

[0637] Acquiring voice input

[0638] The device captures conversations with customers in real time and sends the audio data to a server. The hardware used is the smartphone's microphone, and the software is a voice capture module.

[0639] Voice Recognition

[0640] The server receives the audio data and converts it into text using the Google Cloud Speech-to-Text API, which is then used for further processing.

[0641] Natural Language Processing and Emotion Recognition

[0642] The server passes the text data obtained from the voice recognition module to the natural language processing module (SpaCy) and analyzes the conversation. At the same time, it uses the emotion recognition engine (Emotion Recognition API) to analyze the customer's emotions based on the tone and speed of their voice. This allows it to identify the customer's emotional state, such as tension or fear.

[0643] Customer Harassment Detection

[0644] A natural language processing module is used to analyze text data to detect customer harassment, along with historical data and keyword filtering.

[0645] Risk Assessment

[0646] The risk level of customer harassment is assessed based on emotional data obtained from the emotion recognition module. If high levels of tension or fear are detected, the risk level is assessed as high.

[0647] AI-powered response generation

[0648] If customer harassment or a high risk level is detected, the server automatically activates the AI ​​response module to generate an appropriate response, which is then sent to the terminal in real time and provided to the customer. This uses OpenAI's GPT-4 model.

[0649] Specific examples

[0650] For example, when a customer reports an emergency situation on a security hotline, such as "Help! My window is broken," the device captures the voice and sends it to a server. The server converts the voice data into text and performs natural language processing and emotion recognition. If the server detects a high level of fear based on the analysis results, it notifies the operator, and the AI ​​immediately generates an appropriate response, such as "We will immediately dispatch a security team. Please evacuate to a safe location."

[0651] Prompt Sentence Examples

[0652] Audio data: "Help! The window is broken."

[0653] Prompt statement:

[0654] "Analyze the following voice input, understand the customer's sentiment, and generate an appropriate response.

[0655] Voice input: "Help! The window is broken."

[0656] Analysis: Customers have high levels of fear.

[0657] Response generated:"

[0658] Example response: 'We will be sending a security team immediately. Please seek safety.'

[0659] This reduces the burden on operators and improves the quality of customer service.

[0660] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0661] Step 1:

[0662] Acquiring voice input

[0663] The device captures the conversation with the customer in real time and sends the voice data to the server. The input is the customer's voice, and the output is digital voice data. In this process, the smartphone or headset acts as a microphone, and the voice capture module converts the voice signal into digital information.

[0664] Step 2:

[0665] Voice Recognition

[0666] The server converts the received voice data into text data using the Google Cloud Speech-to-Text API. The input is digital voice data, and the output is text data. In this process, a speech recognition model analyzes the voice waveform and generates corresponding text.

[0667] Step 3:

[0668] Natural Language Processing and Emotion Recognition

[0669] The server inputs the text data into a natural language processing module such as SpaCy to analyze the conversation content. At the same time, it uses the Emotion Recognition API to analyze the customer's emotional state from the text and voice features. The input is text data and voice feature data, and the output is the analyzed conversation content and emotion data. In this process, the natural language processing module analyzes the meaning of words and phrases, and the emotion recognition engine analyzes the tone and speed of the voice to identify emotions.

[0670] Step 4:

[0671] Customer Harassment Detection

[0672] The server uses the analysis results from the natural language processing module to detect customer harassment. Past data and keyword filtering are also used. The input is the analysis result of the text, and the output is the detection result of customer harassment. This process checks whether specific keywords or phrases are included in the text to determine whether harassing behavior is detected.

[0673] Step 5:

[0674] Risk Assessment

[0675] The server uses the emotion data obtained from the emotion recognition module to evaluate the risk level of customer harassment. The input is the customer's emotion data, and the output is the risk level evaluation result. In this process, if high tension or fear is detected, the risk level is set high.

[0676] Step 6:

[0677] AI-powered response generation

[0678] When the server detects customer harassment or a high risk level, it activates an AI response module to generate an appropriate response. The generated AI response is sent to the terminal in real time and provided to the customer. The input is the risk level and harassment detection result, and the output is the generated appropriate response. This process utilizes OpenAI's GPT-4 model to generate a response with the appropriate context and tone.

[0679] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0680] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0681] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0682] [Third embodiment]

[0683] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0684] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0685] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0686] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0687] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0688] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0689] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0690] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0691] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0692] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0693] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0694] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0695] This invention is an AI system for reducing the workload of call center operators and improving productivity and workplace satisfaction.

[0696] composition

[0697] The system consists of the following main components:

[0698] 1. Terminal: A telephone terminal for the operator.

[0699] 2. Server: A computing device that hosts the AI ​​model and speech recognition module.

[0700] 3. Management Console: Operator and administrator interface.

[0701] Program processing

[0702] Acquiring voice input

[0703] The terminal captures the voice conversation between the operator and the customer in real time and sends the voice data to the server, which receives the voice data and transfers it directly to the voice recognition module.

[0704] Voice Recognition

[0705] The server converts the received voice data into text data using a speech recognition module, which is then used for subsequent natural language processing (NLP).

[0706] Real-time analysis of conversation content

[0707] The server inputs the converted text data into an NLP module to detect customer harassment, which uses historical data and keyword filtering for analysis.

[0708] Customer Harassment Detection

[0709] The server receives the analysis results from the NLP module, and if customer harassment is detected, it notifies the operator and automatically switches to the AI-enabled module.

[0710] AI-powered response

[0711] The AI-enabled module takes over the response and sends the generated AI response to the terminal in real time, which relieves the operator of the psychological burden and allows the appropriate response to be taken.

[0712] Advance notice to customers

[0713] The user will notify the customer in advance that this system is being implemented, including the purpose of the AI ​​to prevent harassment.

[0714] Specific examples

[0715] For example, when an operator receives a call from a customer, the voice is captured on the terminal and immediately sent to the server. The server converts the voice into text and analyzes it using the NLP module. If the customer makes a harassing remark such as, "Poor service! Call the manager!", the NLP module detects this and the server switches to the AI-enabled module. The AI ​​generates an appropriate response and responds to the customer via the terminal. This process reduces the burden on the operator and enables efficient business operations.

[0716] In this way, the system reduces the burden on operators and contributes to improving work efficiency and customer satisfaction.

[0717] The processing flow will be explained below.

[0718] Step 1:

[0719] The terminal captures voice inputs exchanged between the operator and the customer in real time, and the captured voice data is immediately sent to the server.

[0720] Step 2:

[0721] The server passes the received voice data to a speech recognition module, which converts the voice into text. This conversion process is performed in real time, and the resulting text data is generated.

[0722] Step 3:

[0723] The server feeds the generated text data into a natural language processing (NLP) module to analyze the conversation, which uses keyword-based filtering and historical data to detect customer harassment.

[0724] Step 4:

[0725] The server checks the analysis results received from the NLP module, and if customer harassment is detected, it notifies an operator and passes the response to the AI-enabled module.

[0726] Step 5:

[0727] The server activates the AI-enabled module and generates an appropriate response, which is then sent to the device in real time.

[0728] Step 6:

[0729] The terminal receives the AI ​​response from the server and forwards it to the customer, which relieves the operator of the psychological burden.

[0730] Step 7:

[0731] When the user installs the system, the user will provide advance notice to the customer, including the fact that the system is being installed for the purpose of preventing customer harassment.

[0732] Example 1

[0733] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0734] In conventional call centers, operators have to deal with harassing comments from customers, which places a heavy psychological burden on them, often resulting in a decline in productivity and workplace satisfaction. Furthermore, if appropriate responses are not made promptly, customer satisfaction also declines, which is an issue.

[0735] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0736] In this invention, the server includes means for acquiring voice input from a terminal used by an operator, means for transmitting the acquired voice data to a computing device in real time, means for the computing device to analyze the transmitted voice data using a voice recognition module, means for converting the voice data analyzed by the voice recognition module into text data, means for inputting the converted text data into a natural language processing module and analyzing the conversation content in real time, means for detecting customer harassment from the analysis results using the natural language processing module, means for automatically switching from the operator to an AI response when customer harassment is detected, means for transmitting the generated AI response to the terminal in real time to respond to the customer, and means for notifying the customer in advance of the introduction of the system. This reduces the psychological burden on the operator and enables quick and appropriate response.

[0737] "Terminal" refers to the communication device used by the operator, which is a hardware device for making calls with customers.

[0738] "Means for acquiring voice input" refers to a device or method that allows the terminal to capture the voice exchanged between the operator and the customer in real time and acquire it as digital data.

[0739] "Means for transmitting voice data in real time" refers to a device or method that has a mechanism for instantly transferring voice data acquired from a terminal to a server.

[0740] A "computing device" is a computer system for analyzing and processing audio data. Also called a server.

[0741] A "voice recognition module" is software or algorithms for converting voice data into text data.

[0742] The "means for converting into text data" is a function that uses a voice recognition module to convert voice data into text information and output it in digital text format.

[0743] A "natural language processing module" is software or algorithms for analyzing text data and understanding and processing its content.

[0744] The "means for detecting customer harassment" is a function that uses a natural language processing module to automatically identify harassing behavior in customer comments.

[0745] The "means of switching to AI response" is a mechanism that automatically switches from an operator to an AI module when customer harassment is detected.

[0746] "Means for generating an AI response" refers to the function of using a generative AI model to create an appropriate response and output the result in digital form.

[0747] "Means for transmitting AI responses to terminals" refers to the function of transferring the generated AI responses to terminals in real time, and having an operator or automated response system provide the responses to customers.

[0748] "Means for notifying customers of the introduction of the system" refers to a method or device for notifying customers in advance that the system will be used in the call center.

[0749] This invention is an AI system that reduces the workload of call center operators and improves productivity and workplace satisfaction. The system consists of terminals, a server, and a management console.

[0750] composition

[0751] Terminal: The telephone terminal used by the operator is a communication device for making calls with customers. As an example, we will use a general IP phone.

[0752] Server: A server is a computing device that hosts the speech recognition module and the natural language processing module. As a concrete example, a typical server computer is used.

[0753] Management Console: The management console is an interface that allows operators and administrators to monitor and manage the state of the system. For example, we will use a web-based interface.

[0754] Program processing

[0755] The server acquires, analyzes, and generates a response based on the following procedure:

[0756] Acquiring voice input

[0757] The terminal captures the voice conversation between the operator and the customer in real time, and then immediately transmits the captured voice data to the server, where it is converted into WAV or MP3 format.

[0758] Voice Recognition

[0759] The server stores the received voice data in memory and converts it into text data using the Google Speech-to-Text API, which is then used for the next natural language processing step.

[0760] Real-time analysis of conversation content

[0761] The server passes the text data to the Google Cloud Natural Language API, which analyzes the conversation in real time. The NLP module uses historical data and keyword filtering to detect harassment.

[0762] Customer Harassment Detection

[0763] The server receives the results of the NLP analysis and checks for flags of customer harassment. If harassment is detected, the server automatically switches to the AI-enabled module.

[0764] AI-powered response

[0765] The server uses the OpenAI GPT-4 API to send prompts to generate appropriate AI responses, which are then sent to the device and provided to the customer via an operator or automated response system.

[0766] Advance notice to customers

[0767] The user will notify the customer in advance that this system is being implemented, and this notification will include the purpose of the AI ​​preventing harassment.

[0768] As a concrete example, let's consider the flow when an operator receives a customer utterance, "Poor service! Call the person in charge!" The device captures the voice and sends it to the server. The server converts the voice into text, and the NLP module detects harassment. The server switches to the AI-enabled module, which generates an appropriate AI response. This response is sent to the device, and the operator or automated response system responds to the customer.

[0769] An example of a prompt is "Operator: What should I do if I receive harassing comments such as, 'Your service is bad! Call the person in charge!'"

[0770] As described above, this system reduces the burden on operators and enables quick and appropriate responses.

[0771] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0772] Step 1: Getting voice input

[0773] The terminal captures the voice exchanged between the operator and the customer as soon as the call begins. The terminal converts the captured voice data into WAV or MP3 format and sends it to the server in real time. The input is voice data, and the output is the converted voice file.

[0774] Step 2: Receiving and storing audio data

[0775] The server receives the audio file sent from the terminal and temporarily stores it in memory. The input here is the audio file sent from the terminal, and the output is the temporarily stored audio data.

[0776] Step 3: Voice Recognition

[0777] The server calls the Google Speech-to-Text API to convert the stored audio data into text data. Specifically, it sends the audio file to the API and receives the returned text data. The input is the audio data, and the output is the converted text data.

[0778] Step 4: Add text data to the queue

[0779] The text data acquired by the server is added to a queue for further processing. The input is the text data obtained by speech recognition, and the output is the text data added to the queue.

[0780] Step 5: Real-time analysis of conversation content

[0781] The server sends the text data to the Google Cloud Natural Language API, which analyzes the conversation in real time. The NLP module uses historical data and keyword filtering to identify harassing behavior in the text data. The input is the text data removed from the queue, and the output is the analysis results.

[0782] Step 6: Detect customer harassment

[0783] The server receives the results of the NLP analysis and determines whether or not customer harassment has occurred. Specifically, it determines this by checking the flags and evaluation values ​​of the analysis results. The input is the analysis results, and the output is a flag or evaluation value indicating whether or not harassment has occurred.

[0784] Step 7: Switch to automatic AI support

[0785] When the server detects customer harassment, it automatically switches from the operator to the AI-enabled module. Specifically, the server generates instructions to hand over the response to the AI. The input is the harassment detection result, and the output is an instruction to start the AI-enabled module.

[0786] Step 8: Generate an AI response

[0787] The server calls the OpenAI GPT-4 API and provides prompts to generate appropriate responses to harassment. Specifically, it sends the prompt sentence as input to the API and obtains the returned response. The input is the prompt sentence, and the output is the generated AI response.

[0788] Step 9: Sending an AI Response

[0789] The server sends the generated AI response to the terminal and provides it to the customer through an operator or an automated response system. The input is the generated AI response, and the output is the AI ​​response sent to the terminal.

[0790] Step 10: Notify customers in advance

[0791] The user notifies the customer in advance that the system is in place and that the AI ​​is intended to prevent harassment. Specifically, voice guidance and pre-recorded messages are used. The input is a notification message, and the output is a notification to the customer.

[0792] (Application example 1)

[0793] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0794] In conventional call center systems, operators are often exposed to harassment and stressful interactions from customers, which increases their workload and reduces workplace satisfaction. In addition, there are also situations in which staff in brick-and-mortar stores feel stressed when interacting with customers, raising concerns that this could lead to a decline in work efficiency. To solve these problems, a system is needed that reduces the stress that occurs during customer interactions and improves work efficiency.

[0795] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0796] In this invention, the server includes means for acquiring voice input from a terminal used by an operator, means for analyzing the acquired voice data in real time, means for converting the analyzed voice data into text, means for performing natural language processing based on the converted text data to detect customer harassment, means for automatically switching from the operator to AI response when customer harassment is detected, means for notifying customers in advance that the system has been introduced, and means for detecting harassment while staff are interacting with customers in a physical store and for AI to take over the response. This makes it possible to detect harassment while serving customers and for AI to respond, reducing the workload of staff and enabling work efficiency and improved customer satisfaction.

[0797] The "terminal used by the operator" refers to an information processing device used by the operator to interact with the customer, and has voice input and output functions.

[0798] "Means for acquiring voice input" refers to a device or system configuration that has the function of capturing voices emitted by operators or customers in real time and acquiring them as digital data.

[0799] "Means for analyzing acquired voice data in real time" refers to a device or program that has the function of processing captured voice data in real time and performing the necessary analysis.

[0800] "Means for converting analyzed voice data into text" refers to a device or system configuration that utilizes voice recognition technology to automatically convert voice data into text data.

[0801] "Means for detecting customer harassment through natural language processing" refers to a device or system configuration that has the function of determining whether or not customer harassment has occurred by utilizing machine learning algorithms and keyword filtering based on text data.

[0802] "Means for automatically switching from human to AI response" refers to a device or system configuration that has the functionality to allow an AI system to automatically take over response when customer harassment is detected.

[0803] "Means of notifying customers in advance of the system implementation" refers to methods and means for communicating information and precautions regarding the system implementation to customers in advance.

[0804] "Means for detecting harassment while staff are interacting with customers in a physical store, and for AI to take over the response" refers to a device or system configuration that has the functionality to enable a system to detect customer harassment while staff are interacting with customers in a physical store, and for AI to take over the interaction.

[0805] This invention is a system that detects harassment that occurs when staff interact with customers in brick-and-mortar stores and transfers the response to an AI system, thereby reducing the workload of staff and improving work efficiency. The specific configuration and operation of the system are described in detail below.

[0806] System configuration

[0807] The system mainly consists of the following components:

[0808] 1. Terminals: Devices such as smart glasses or smartphones used by staff in brick-and-mortar stores.

[0809] 2. Server: A computing device that hosts the speech recognition module and the natural language processing (NLP) module.

[0810] 3. Management Console: The interface for managing and configuring the entire system.

[0811] Hardware and Software

[0812] Hardware: Smart glasses or smartphones are used to capture customer conversations in real time.

[0813] Software: We use the Google Cloud Speech-to-Text API to convert voice data to text, and Hugging Face Transformers for NLP analysis.

[0814] Data processing: Converts voice data into text, detects customer harassment through NLP analysis, and generates appropriate AI responses.

[0815] Acquiring and analyzing voice input

[0816] The device captures the customer's conversation in real time and sends the audio data to the server, which then receives the data and converts it into text using the Google Cloud Speech-to-Text API.

[0817] Natural language processing of text data

[0818] The converted text data is then analyzed by an NLP module, which uses Hugging Face's Transformers library to detect customer harassment. The NLP module performs the analysis using historical data and keyword filtering.

[0819] AI-powered response

[0820] If customer harassment is detected, the server notifies the operator and automatically switches to the AI-enabled module, which generates an appropriate AI response based on the analysis results and sends it to the terminal in real time. This relieves staff from psychological burden and improves work efficiency.

[0821] Advance notice to customers

[0822] Users will notify their customers in advance that the system is in place, which includes AI-based detection of harassing behavior and automatic response measures.

[0823] Specific examples

[0824] For example, if a store staff member is interacting with a customer and the customer says, "I don't like the product! Call the manager!", the system will detect this. The NLP module analyzes the text data, and if it recognizes harassment, the server switches to the AI-enabled module. The AI ​​generates a response such as, "Sir, we'll help you find a quieter location," which is then conveyed to the customer via their device.

[0825] Prompt Sentence Examples

[0826] "How would you respond when a customer says, 'I don't like the product! Call the store manager!' What would be an appropriate response?"

[0827] In this way, the system reduces the burden on staff and contributes to improved work efficiency and customer satisfaction.

[0828] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0829] Step 1:

[0830] The terminal captures the conversation with the customer in real time and sends the voice data to the server. At this time, the terminal's voice input function is used to acquire the input voice data as a digital signal. The input data is then transferred from the terminal to the server.

[0831] Step 2:

[0832] The server converts the received voice data into text data using the Google Cloud Speech-to-Text API. The speech recognition module analyzes the voice data and generates corresponding text data. The server analyzes the input voice data and outputs it as text data.

[0833] Step 3:

[0834] The server sends the converted text data to a natural language processing (NLP) module to analyze the text data. It uses Hugging Face's Transformers library to detect harassing remarks in the text data. For analysis, the NLP module receives the text data as input and outputs the customer harassment detection results.

[0835] Step 4:

[0836] The server receives the analysis results from the NLP module and automatically switches to the AI-enabled module if harassment is detected. The instruction to switch is based on the harassment detection results as input data, and switching to the AI-enabled module as output.

[0837] Step 5:

[0838] The AI-enabled module generates an appropriate AI response based on the analysis results. The generative AI model creates the generative AI response in real time and sends it to the device. It receives the analysis results as input and outputs the appropriate response sentence.

[0839] Step 6:

[0840] The device receives the AI ​​response and conveys it to the customer as a voice output. The AI-generated response is provided to the customer via smart glasses or a smartphone. The AI ​​response that serves as input is converted into voice output and communicated to the customer.

[0841] Step 7:

[0842] The user notifies the customer in advance about the system's implementation and its functions. This notification uses prompts to explain information about the system's implementation and the AI's harassment detection function. The content of the notification is the input data, and the message to the customer is the output.

[0843] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0844] This invention is an AI system that reduces the burden on call center operators and enables efficient responses when they are faced with customer complaints or customer harassment. Furthermore, by combining it with an emotion engine that recognizes user emotions, it improves the quality of customer service.

[0845] composition

[0846] The system consists of the following main components:

[0847] 1. Terminal: A telephone terminal for the operator.

[0848] 2. Server: A computing device that hosts the AI ​​model, speech recognition module, and emotion engine.

[0849] 3. Management Console: Operator and administrator interface.

[0850] Program processing

[0851] Acquiring voice input

[0852] The terminal captures the voice conversation between the operator and the customer in real time and sends the voice data to the server, which receives the voice data and transfers it directly to the voice recognition module.

[0853] Voice Recognition

[0854] The server converts the received voice data into text data using a speech recognition module, which is then used for subsequent natural language processing (NLP) and emotion engines.

[0855] Real-time analysis of conversation content and emotion recognition

[0856] The server inputs the received text data into the NLP module to analyze the conversation content. At the same time, the emotion engine analyzes the emotion based on the user's voice tone, speed, and text content. The analysis results of the emotion engine are integrated with the results of the NLP module.

[0857] Customer Harassment Detection

[0858] The server receives the analysis results from the NLP module, and if customer harassment is detected, it notifies the operator and automatically switches to the AI-enabled module.

[0859] AI-powered response

[0860] The AI-enabled module generates an appropriate response, dynamically adjusting the AI ​​response content based on the analysis results of the emotion engine, and transmitting the generated AI response to the device in real time.

[0861] Advance notice to customers

[0862] When the system is introduced, the user will provide advance notice to customers, including the fact that the system is being introduced to prevent customer harassment and that it also has emotion recognition capabilities.

[0863] Specific examples

[0864] For example, when an operator receives a call from a customer, the voice is captured on the device and immediately sent to the server. The server converts the voice into text and analyzes it using the NLP module, while the emotion engine analyzes the tone and speed of the voice. If a customer makes a harassing remark such as, "Poor service! Call the manager!" and the emotion engine detects a high level of anger, the server will comprehensively assess this, notify the operator, and switch to the AI-enabled module. The AI ​​will then generate an appropriate response based on the emotion, such as, "We apologize for the inconvenience," and respond to the customer via the device. This process reduces the operator's burden and improves the quality of customer service.

[0865] In this way, the system reduces the burden on operators and contributes to improving work efficiency and customer satisfaction.

[0866] The processing flow will be explained below.

[0867] Step 1:

[0868] The terminal captures voice inputs exchanged between the operator and the customer in real time, and the captured voice data is immediately sent to the server.

[0869] Step 2:

[0870] The server passes the received voice data to a speech recognition module, which converts the voice into text. This conversion process is performed in real time, and the resulting text data is generated.

[0871] Step 3:

[0872] The server then inputs the generated text data into a natural language processing (NLP) module to analyze the conversation, but at this stage it does not yet make a judgment as to whether or not customer harassment has occurred.

[0873] Step 4:

[0874] In parallel, the server passes the text data and voice data to the emotion engine, which analyzes the user's emotions based on the tone, speed, and text content of the voice.

[0875] Step 5:

[0876] The server integrates the analysis results from the NLP module and the emotion engine to determine whether customer harassment has occurred and the user's emotional state. If the emotion engine detects high emotions such as "anger" or "irritation," the possibility of harassment increases.

[0877] Step 6:

[0878] If the server detects customer harassment and high emotions, it will notify the operator and automatically switch to the AI-enabled module, including the specific part that was deemed to be harassment.

[0879] Step 7:

[0880] The server then activates an AI-enabled module to generate an appropriate response that takes into account the user's emotional state. For example, if the emotion engine detects "anger," it will generate a sympathetic response such as, "We're sorry. We're very sorry for the inconvenience."

[0881] Step 8:

[0882] The server generates an AI response and sends it to the terminal in real time. The terminal then forwards the AI ​​response to the customer, and the operator responds on their behalf, allowing the customer to continue their conversation.

[0883] Step 9:

[0884] When the system is introduced, the user will provide advance notice to customers about its use, including that the AI ​​will prevent customer harassment and take appropriate action using emotion recognition.

[0885] By doing so, the system of the present invention reduces the workload of operators and realizes more sophisticated customer service. In addition, by taking into account the user's emotions, it is expected that customer satisfaction will also improve.

[0886] Example 2

[0887] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0888] In modern call center operations, operators are burdened with dealing with customer harassment and emotional complaints from customers. This burden increases the operators' mental stress, reduces work efficiency, and ultimately leads to a decline in customer satisfaction. Conventional systems did not provide an effective means to solve these problems.

[0889] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0890] In this invention, the server includes means for acquiring voice input from a terminal used by an operator, means for analyzing the acquired voice data in real time, means for converting the analyzed voice data into text, means for performing natural language processing and emotion recognition based on the converted text data to detect customer harassment, means for automatically switching from the operator to AI response when customer harassment is detected and generating a response based on the emotion recognition results, and means for notifying customers in advance of the introduction of the system. This reduces the burden on operators and enables efficient customer service that takes emotions into consideration.

[0891] "Terminal" refers to the communications equipment used by the operator and is a device for obtaining voice input.

[0892] A "server" is a computing device that performs the main processing of the system, such as analyzing voice data, natural language processing, emotion recognition, and detecting customer harassment.

[0893] "Voice input" refers to the voice data exchanged between the operator and the customer.

[0894] "Means for analyzing in real time" refers to a processing method or system configuration for instantly analyzing voice data acquired from a terminal.

[0895] A "means for converting voice data to text" is a processing method or system configuration for converting voice data to text data using voice recognition technology.

[0896] "Text data" is voice data converted into character information.

[0897] "Natural language processing" is a technical field that understands and analyzes the meaning of text data as human language.

[0898] "Emotion recognition" is a technology for identifying a speaker's emotions based on voice and text data.

[0899] "Customer harassment" refers to inappropriate comments or actions made by customers toward operators.

[0900] "AI-enabled" is the process of using artificial intelligence to automatically generate appropriate responses and assist operators.

[0901] "Harassment detection result" is information that indicates whether customer harassment is occurring based on the results of natural language processing and emotion recognition.

[0902] "Means for notifying customers in advance of the introduction of the system" refers to the methods and processes for informing customers in advance that the system will be used.

[0903] This invention is a system that reduces the burden on call center operators when dealing with customer complaints and harassment, and realizes efficient and emotion-sensitive responses. In particular, by combining it with an emotion engine that recognizes the user's emotions, the quality of customer service is improved.

[0904] System configuration

[0905] Hardware and software used

[0906] 1. Terminal: A communication device for an operator, capable of capturing voice input (e.g., a telephone or headset).

[0907] 2. Server: A computing device that performs the main processing of the system, such as analyzing voice data, natural language processing, emotion recognition, and customer harassment detection. Specific software examples include Google Speech-to-Text API, OpenAI GPT-4, and IBM Watson Tone Analyzer.

[0908] 3. Management console: An interface for operators and administrators, software for configuring and monitoring the system (e.g., a web-based management screen).

[0909] Operation procedures and examples

[0910] Acquiring voice input

[0911] The terminal captures the voices exchanged between the operator and the customer in real time. For example, when an operator receives a call from a customer, the conversation is picked up through the terminal's microphone and immediately sent to the server.

[0912] Voice Recognition

[0913] The server receives the voice data and converts it into text using the Google Speech-to-Text API. The server sends the voice data to the API and receives the text data within a few seconds.

[0914] Real-time analysis of conversation content and emotion recognition

[0915] The server inputs the received text data into a natural language processing module (OpenAI GPT-4) to analyze the content of the conversation. At the same time, an emotion engine (IBM Watson Tone Analyzer) analyzes emotions based on the tone of voice, speed, and text content. The analysis results are integrated and used for subsequent processing.

[0916] Customer Harassment Detection

[0917] The server checks the analysis results from the NLP module, and if it determines that certain words or phrases constitute customer harassment, it immediately notifies the operator and automatically switches to the AI-enabled module. For example, if the text data contains words such as "call the person in charge" or "the service is terrible," the server will detect harassment.

[0918] AI-powered response

[0919] The server activates an AI-enabled module to generate an appropriate response. At this time, a generative AI model (OpenAI GPT-4) dynamically adjusts the response content based on the results of sentiment analysis. For example, if a customer says, "The service was poor," the AI ​​generates a response such as, "We're sorry, we apologize for the inconvenience," and sends it to the device in real time.

[0920] Advance notice to customers

[0921] When introducing the system, the user notifies customers in advance, for example by email or letter, that "this system has been introduced with the aim of preventing customer harassment and is equipped with emotion recognition functionality."

[0922] Prompt Sentence Examples

[0923] The AI ​​model is fed with the following prompt:

[0924] "If the customer is enraged, generate an effective calming response. Their words are, 'Poor service! Bring out the person in charge!'"

[0925] "Please write an appropriate apology for the user's comments. The situation is that a customer is expressing dissatisfaction with a product."

[0926] This allows the system to reduce the burden on operators and provide quick and emotionally sensitive customer service.

[0927] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0928] Step 1: Getting voice input

[0929] The terminal captures the voice exchanged between the operator and the customer in real time. Specifically, the terminal's microphone receives voice input and processes this voice data as a digital signal. The captured voice data is then sent to the server. For example, when an operator says, "Hello, thank you for your inquiry," the voice is immediately converted into digital data by the terminal and sent to the server. The input here is voice data, and the output is digital voice data.

[0930] Step 2: Voice Recognition

[0931] The server receives the voice data sent from the device and passes it to a voice recognition module. The Google Speech-to-Text API is used as the voice recognition module. When voice data is sent to this API, the API converts the voice data into text data, which is then received by the server. The input here is digital voice data, and the output is text data. Specifically, the server posts the voice data to Google's API and receives the text data within a few seconds.

[0932] Step 3: Real-time analysis of conversation content and emotion recognition

[0933] The server inputs the received text data into a natural language processing module (OpenAI GPT-4). The server then sends the text data to GPT-4, which analyzes the content of the dialogue. At the same time, the emotion engine (IBM Watson Tone Analyzer) analyzes the user's emotions based on the tone of voice, speed, and text content. The server then combines the results of these two analyses. The input here is text data, and the output is analysis results based on natural language processing and emotion recognition results. Specifically, the server sends the text data to GPT-4, where it performs natural language processing, while simultaneously sending it to IBM Watson Tone Analyzer for emotion analysis.

[0934] Step 4: Detect customer harassment

[0935] The server checks the analysis results from the natural language processing module and checks whether specific words or phrases are judged to be customer harassment. For example, if the text data contains words such as "Bring the person in charge" or "The service is terrible" and the emotion engine detects a high level of anger, these conditions will trigger the detection of harassment. The input is the results of natural language processing and emotion recognition, and the output is a customer harassment detection flag. In concrete terms, the server filters the analysis results based on the conditions and determines whether harassment has occurred.

[0936] Step 5: AI response

[0937] The server launches an AI-enabled module and generates an appropriate response. At this time, the generative AI model (OpenAI GPT-4) dynamically adjusts the response content based on the results of emotion analysis. For example, if a customer is angry and complains about poor service, the AI ​​generates a response such as "We apologize for the inconvenience." The server then sends the generated AI response to the device in real time. The input is the emotion recognition result and harassment detection flag, and the output is the response message generated by the AI. Specifically, the server sends a prompt to GPT-4, which generates a response and sends the result to the device.

[0938] Step 6: Notify customers in advance

[0939] When the system is introduced, the user notifies customers in advance. For example, by email or letter, the user can explain that "this system has been introduced to prevent customer harassment and is equipped with emotion recognition functionality." The input is the notification content, and the output is the notification status to the customer. In concrete terms, the user references the customer database and sends a mass notification using an email system or postal system.

[0940] This detailed processing step reduces the burden on the operator and ensures a fast and empathetic customer service experience.

[0941] (Application example 2)

[0942] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0943] Traditional customer support and security operators are required to respond quickly and appropriately to harassment comments and emergency situations from customers, but this process is extremely stressful and places a heavy psychological burden on the operators. Furthermore, there is a lack of systems that can properly recognize customer emotions and generate optimal responses, which can lead to inconsistencies in the quality of customer support. There is a need for a system that can solve these issues, reduce the burden on operators, and provide efficient, high-quality customer support.

[0944] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0945] In this invention, the server includes means for acquiring voice input from a terminal used by an operator, means for analyzing the acquired voice data in real time, means for converting the analyzed voice data into text, means for performing natural language processing based on the converted text data to detect customer harassment, means for automatically switching from the operator to AI response when customer harassment is detected, means for notifying the customer in advance of the system implementation, means for recognizing emotions based on the converted text data and tone of voice, means for evaluating the risk level of customer harassment based on the detected emotion data, and means for detecting customer emergencies, identifying emotions of high tension or fear, and generating an appropriate response. This makes it possible to reduce the burden on operators and improve the quality of customer service.

[0946] "Terminal" means a communication device used by an operator to obtain voice input.

[0947] "Audio data" is an audio signal obtained from a terminal expressed as digital information.

[0948] "Real-time" refers to the fact that audio data is processed immediately after it is acquired.

[0949] "Text data" is voice data converted into text information using voice recognition technology.

[0950] "Natural language processing" is a technology that allows computers to understand and analyze human language.

[0951] "Customer harassment" refers to malicious words, actions, and behavior from customers toward operators.

[0952] "AI-enabled" refers to using artificial intelligence to automatically generate responses to customers.

[0953] "Emotion recognition" is a technology that analyzes voice data and text data to estimate the emotional state of a speaker.

[0954] "Risk level" is a numerical or index expression of the danger or urgency of customer harassment.

[0955] An "emergency incident" is a serious event or emergency that requires immediate action.

[0956] "Response generation" is the process of creating an appropriate response based on the analysis results.

[0957] The present invention relates to a system that enables an operator to efficiently process voice data from a customer and generate an appropriate response. The specific configuration and functions of this system are described below.

[0958] System Configuration

[0959] 1. Terminal

[0960] The terminal is a communication device that allows the operator to communicate with the customer and is used to obtain voice input. Specifically, this applies to smartphones and headsets.

[0961] 2. Server

[0962] The server is a high-performance computing device that hosts multiple functional modules, including a speech recognition module, a natural language processing module, an emotion recognition engine, and an AI response generation module.

[0963] 3. Management Console

[0964] The management console is an interface that allows operators and administrators to monitor the status of the system and change its settings.

[0965] Program processing

[0966] Acquiring voice input

[0967] The device captures conversations with customers in real time and sends the audio data to a server. The hardware used is the smartphone's microphone, and the software is a voice capture module.

[0968] Voice Recognition

[0969] The server receives the audio data and converts it into text using the Google Cloud Speech-to-Text API, which is then used for further processing.

[0970] Natural Language Processing and Emotion Recognition

[0971] The server passes the text data obtained from the voice recognition module to the natural language processing module (SpaCy) and analyzes the conversation. At the same time, it uses the emotion recognition engine (Emotion Recognition API) to analyze the customer's emotions based on the tone and speed of their voice. This allows it to identify the customer's emotional state, such as tension or fear.

[0972] Customer Harassment Detection

[0973] A natural language processing module is used to analyze text data to detect customer harassment, along with historical data and keyword filtering.

[0974] Risk Assessment

[0975] The risk level of customer harassment is assessed based on emotional data obtained from the emotion recognition module. If high levels of tension or fear are detected, the risk level is assessed as high.

[0976] AI-powered response generation

[0977] If customer harassment or a high risk level is detected, the server automatically activates the AI ​​response module to generate an appropriate response, which is then sent to the terminal in real time and provided to the customer. This uses OpenAI's GPT-4 model.

[0978] Specific examples

[0979] For example, when a customer reports an emergency situation on a security hotline, such as "Help! My window is broken," the device captures the voice and sends it to a server. The server converts the voice data into text and performs natural language processing and emotion recognition. If the server detects a high level of fear based on the analysis results, it notifies the operator, and the AI ​​immediately generates an appropriate response, such as "We will immediately dispatch a security team. Please evacuate to a safe location."

[0980] Prompt Sentence Examples

[0981] Audio data: "Help! The window is broken."

[0982] Prompt statement:

[0983] "Analyze the following voice input, understand the customer's sentiment, and generate an appropriate response.

[0984] Voice input: "Help! The window is broken."

[0985] Analysis: Customers have high levels of fear.

[0986] Response generated:"

[0987] Example response: 'We will be sending a security team immediately. Please seek safety.'

[0988] This reduces the burden on operators and improves the quality of customer service.

[0989] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0990] Step 1:

[0991] Acquiring voice input

[0992] The device captures the conversation with the customer in real time and sends the voice data to the server. The input is the customer's voice, and the output is digital voice data. In this process, the smartphone or headset acts as a microphone, and the voice capture module converts the voice signal into digital information.

[0993] Step 2:

[0994] Voice Recognition

[0995] The server converts the received voice data into text data using the Google Cloud Speech-to-Text API. The input is digital voice data, and the output is text data. In this process, a speech recognition model analyzes the voice waveform and generates corresponding text.

[0996] Step 3:

[0997] Natural Language Processing and Emotion Recognition

[0998] The server inputs the text data into a natural language processing module such as SpaCy to analyze the conversation content. At the same time, it uses the Emotion Recognition API to analyze the customer's emotional state from the text and voice features. The input is text data and voice feature data, and the output is the analyzed conversation content and emotion data. In this process, the natural language processing module analyzes the meaning of words and phrases, and the emotion recognition engine analyzes the tone and speed of the voice to identify emotions.

[0999] Step 4:

[1000] Customer Harassment Detection

[1001] The server uses the analysis results from the natural language processing module to detect customer harassment. Past data and keyword filtering are also used. The input is the analysis result of the text, and the output is the detection result of customer harassment. This process checks whether specific keywords or phrases are included in the text to determine whether harassing behavior is detected.

[1002] Step 5:

[1003] Risk Assessment

[1004] The server uses the emotion data obtained from the emotion recognition module to evaluate the risk level of customer harassment. The input is the customer's emotion data, and the output is the risk level evaluation result. In this process, if high tension or fear is detected, the risk level is set high.

[1005] Step 6:

[1006] AI-powered response generation

[1007] When the server detects customer harassment or a high risk level, it activates an AI response module to generate an appropriate response. The generated AI response is sent to the terminal in real time and provided to the customer. The input is the risk level and harassment detection result, and the output is the generated appropriate response. This process utilizes OpenAI's GPT-4 model to generate a response with the appropriate context and tone.

[1008] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1009] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1010] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1011] [Fourth embodiment]

[1012] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1013] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1014] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1015] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1016] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1017] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1018] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1019] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1020] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1021] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1022] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1023] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1024] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1025] This invention is an AI system for reducing the workload of call center operators and improving productivity and workplace satisfaction.

[1026] composition

[1027] The system consists of the following main components:

[1028] 1. Terminal: A telephone terminal for the operator.

[1029] 2. Server: A computing device that hosts the AI ​​model and speech recognition module.

[1030] 3. Management Console: Operator and administrator interface.

[1031] Program processing

[1032] Acquiring voice input

[1033] The terminal captures the voice conversation between the operator and the customer in real time and sends the voice data to the server, which receives the voice data and transfers it directly to the voice recognition module.

[1034] Voice Recognition

[1035] The server converts the received voice data into text data using a speech recognition module, which is then used for subsequent natural language processing (NLP).

[1036] Real-time analysis of conversation content

[1037] The server inputs the converted text data into an NLP module to detect customer harassment, which uses historical data and keyword filtering for analysis.

[1038] Customer Harassment Detection

[1039] The server receives the analysis results from the NLP module, and if customer harassment is detected, it notifies the operator and automatically switches to the AI-enabled module.

[1040] AI-powered response

[1041] The AI-enabled module takes over the response and sends the generated AI response to the terminal in real time, which relieves the operator of the psychological burden and allows the appropriate response to be taken.

[1042] Advance notice to customers

[1043] The user will notify the customer in advance that this system is being implemented, including the purpose of the AI ​​to prevent harassment.

[1044] Specific examples

[1045] For example, when an operator receives a call from a customer, the voice is captured on the terminal and immediately sent to the server. The server converts the voice into text and analyzes it using the NLP module. If the customer makes a harassing remark such as, "Poor service! Call the manager!", the NLP module detects this and the server switches to the AI-enabled module. The AI ​​generates an appropriate response and responds to the customer via the terminal. This process reduces the burden on the operator and enables efficient business operations.

[1046] In this way, the system reduces the burden on operators and contributes to improving work efficiency and customer satisfaction.

[1047] The processing flow will be explained below.

[1048] Step 1:

[1049] The terminal captures voice inputs exchanged between the operator and the customer in real time, and the captured voice data is immediately sent to the server.

[1050] Step 2:

[1051] The server passes the received voice data to a speech recognition module, which converts the voice into text. This conversion process is performed in real time, and the resulting text data is generated.

[1052] Step 3:

[1053] The server feeds the generated text data into a natural language processing (NLP) module to analyze the conversation, which uses keyword-based filtering and historical data to detect customer harassment.

[1054] Step 4:

[1055] The server checks the analysis results received from the NLP module, and if customer harassment is detected, it notifies an operator and passes the response to the AI-enabled module.

[1056] Step 5:

[1057] The server activates the AI-enabled module and generates an appropriate response, which is then sent to the device in real time.

[1058] Step 6:

[1059] The terminal receives the AI ​​response from the server and forwards it to the customer, which relieves the operator of the psychological burden.

[1060] Step 7:

[1061] When the user installs the system, the user will provide advance notice to the customer, including the fact that the system is being installed for the purpose of preventing customer harassment.

[1062] Example 1

[1063] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1064] In conventional call centers, operators have to deal with harassing comments from customers, which places a heavy psychological burden on them, often resulting in a decline in productivity and workplace satisfaction. Furthermore, if appropriate responses are not made promptly, customer satisfaction also declines, which is an issue.

[1065] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1066] In this invention, the server includes means for acquiring voice input from a terminal used by an operator, means for transmitting the acquired voice data to a computing device in real time, means for the computing device to analyze the transmitted voice data using a voice recognition module, means for converting the voice data analyzed by the voice recognition module into text data, means for inputting the converted text data into a natural language processing module and analyzing the conversation content in real time, means for detecting customer harassment from the analysis results using the natural language processing module, means for automatically switching from the operator to an AI response when customer harassment is detected, means for transmitting the generated AI response to the terminal in real time to respond to the customer, and means for notifying the customer in advance of the introduction of the system. This reduces the psychological burden on the operator and enables quick and appropriate response.

[1067] "Terminal" refers to the communication device used by the operator, which is a hardware device for making calls with customers.

[1068] "Means for acquiring voice input" refers to a device or method that allows the terminal to capture the voice exchanged between the operator and the customer in real time and acquire it as digital data.

[1069] "Means for transmitting voice data in real time" refers to a device or method that has a mechanism for instantly transferring voice data acquired from a terminal to a server.

[1070] A "computing device" is a computer system for analyzing and processing audio data. Also called a server.

[1071] A "voice recognition module" is software or algorithms for converting voice data into text data.

[1072] The "means for converting into text data" is a function that uses a voice recognition module to convert voice data into text information and output it in digital text format.

[1073] A "natural language processing module" is software or algorithms for analyzing text data and understanding and processing its content.

[1074] The "means for detecting customer harassment" is a function that uses a natural language processing module to automatically identify harassing behavior in customer comments.

[1075] The "means of switching to AI response" is a mechanism that automatically switches from an operator to an AI module when customer harassment is detected.

[1076] "Means for generating an AI response" refers to the function of using a generative AI model to create an appropriate response and output the result in digital form.

[1077] "Means for transmitting AI responses to terminals" refers to the function of transferring the generated AI responses to terminals in real time, and having an operator or automated response system provide the responses to customers.

[1078] "Means for notifying customers of the introduction of the system" refers to a method or device for notifying customers in advance that the system will be used in the call center.

[1079] This invention is an AI system that reduces the workload of call center operators and improves productivity and workplace satisfaction. The system consists of terminals, a server, and a management console.

[1080] composition

[1081] Terminal: The telephone terminal used by the operator is a communication device for making calls with customers. As an example, we will use a general IP phone.

[1082] Server: A server is a computing device that hosts the speech recognition module and the natural language processing module. As a concrete example, a typical server computer is used.

[1083] Management Console: The management console is an interface that allows operators and administrators to monitor and manage the state of the system. For example, we will use a web-based interface.

[1084] Program processing

[1085] The server acquires, analyzes, and generates a response based on the following procedure:

[1086] Acquiring voice input

[1087] The terminal captures the voice conversation between the operator and the customer in real time, and then immediately transmits the captured voice data to the server, where it is converted into WAV or MP3 format.

[1088] Voice Recognition

[1089] The server stores the received voice data in memory and converts it into text data using the Google Speech-to-Text API, which is then used for the next natural language processing step.

[1090] Real-time analysis of conversation content

[1091] The server passes the text data to the Google Cloud Natural Language API, which analyzes the conversation in real time. The NLP module uses historical data and keyword filtering to detect harassment.

[1092] Customer Harassment Detection

[1093] The server receives the results of the NLP analysis and checks for flags of customer harassment. If harassment is detected, the server automatically switches to the AI-enabled module.

[1094] AI-powered response

[1095] The server uses the OpenAI GPT-4 API to send prompts to generate appropriate AI responses, which are then sent to the device and provided to the customer via an operator or automated response system.

[1096] Advance notice to customers

[1097] The user will notify the customer in advance that this system is being implemented, and this notification will include the purpose of the AI ​​preventing harassment.

[1098] As a concrete example, let's consider the flow when an operator receives a customer utterance, "Poor service! Call the person in charge!" The device captures the voice and sends it to the server. The server converts the voice into text, and the NLP module detects harassment. The server switches to the AI-enabled module, which generates an appropriate AI response. This response is sent to the device, and the operator or automated response system responds to the customer.

[1099] An example of a prompt is "Operator: What should I do if I receive harassing comments such as, 'Your service is bad! Call the person in charge!'"

[1100] As described above, this system reduces the burden on operators and enables quick and appropriate responses.

[1101] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1102] Step 1: Getting voice input

[1103] The terminal captures the voice exchanged between the operator and the customer as soon as the call begins. The terminal converts the captured voice data into WAV or MP3 format and sends it to the server in real time. The input is voice data, and the output is the converted voice file.

[1104] Step 2: Receiving and storing audio data

[1105] The server receives the audio file sent from the terminal and temporarily stores it in memory. The input here is the audio file sent from the terminal, and the output is the temporarily stored audio data.

[1106] Step 3: Voice Recognition

[1107] The server calls the Google Speech-to-Text API to convert the stored audio data into text data. Specifically, it sends the audio file to the API and receives the returned text data. The input is the audio data, and the output is the converted text data.

[1108] Step 4: Add text data to the queue

[1109] The text data acquired by the server is added to a queue for further processing. The input is the text data obtained by speech recognition, and the output is the text data added to the queue.

[1110] Step 5: Real-time analysis of conversation content

[1111] The server sends the text data to the Google Cloud Natural Language API, which analyzes the conversation in real time. The NLP module uses historical data and keyword filtering to identify harassing behavior in the text data. The input is the text data removed from the queue, and the output is the analysis results.

[1112] Step 6: Detect customer harassment

[1113] The server receives the results of the NLP analysis and determines whether or not customer harassment has occurred. Specifically, it determines this by checking the flags and evaluation values ​​of the analysis results. The input is the analysis results, and the output is a flag or evaluation value indicating whether or not harassment has occurred.

[1114] Step 7: Switch to automatic AI support

[1115] When the server detects customer harassment, it automatically switches from the operator to the AI-enabled module. Specifically, the server generates instructions to hand over the response to the AI. The input is the harassment detection result, and the output is an instruction to start the AI-enabled module.

[1116] Step 8: Generate an AI response

[1117] The server calls the OpenAI GPT-4 API and provides prompts to generate appropriate responses to harassment. Specifically, it sends the prompt sentence as input to the API and obtains the returned response. The input is the prompt sentence, and the output is the generated AI response.

[1118] Step 9: Sending an AI Response

[1119] The server sends the generated AI response to the terminal and provides it to the customer through an operator or an automated response system. The input is the generated AI response, and the output is the AI ​​response sent to the terminal.

[1120] Step 10: Notify customers in advance

[1121] The user notifies the customer in advance that the system is in place and that the AI ​​is intended to prevent harassment. Specifically, voice guidance and pre-recorded messages are used. The input is a notification message, and the output is a notification to the customer.

[1122] (Application example 1)

[1123] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1124] In conventional call center systems, operators are often exposed to harassment and stressful interactions from customers, which increases their workload and reduces workplace satisfaction. In addition, there are also situations in which staff in brick-and-mortar stores feel stressed when interacting with customers, raising concerns that this could lead to a decline in work efficiency. To solve these problems, a system is needed that reduces the stress that occurs during customer interactions and improves work efficiency.

[1125] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1126] In this invention, the server includes means for acquiring voice input from a terminal used by an operator, means for analyzing the acquired voice data in real time, means for converting the analyzed voice data into text, means for performing natural language processing based on the converted text data to detect customer harassment, means for automatically switching from the operator to AI response when customer harassment is detected, means for notifying customers in advance that the system has been introduced, and means for detecting harassment while staff are interacting with customers in a physical store and for AI to take over the response. This makes it possible to detect harassment while serving customers and for AI to respond, reducing the workload of staff and enabling work efficiency and improved customer satisfaction.

[1127] The "terminal used by the operator" refers to an information processing device used by the operator to interact with the customer, and has voice input and output functions.

[1128] "Means for acquiring voice input" refers to a device or system configuration that has the function of capturing voices emitted by operators or customers in real time and acquiring them as digital data.

[1129] "Means for analyzing acquired voice data in real time" refers to a device or program that has the function of processing captured voice data in real time and performing the necessary analysis.

[1130] "Means for converting analyzed voice data into text" refers to a device or system configuration that utilizes voice recognition technology to automatically convert voice data into text data.

[1131] "Means for detecting customer harassment through natural language processing" refers to a device or system configuration that has the function of determining whether or not customer harassment has occurred by utilizing machine learning algorithms and keyword filtering based on text data.

[1132] "Means for automatically switching from human to AI response" refers to a device or system configuration that has the functionality to allow an AI system to automatically take over response when customer harassment is detected.

[1133] "Means of notifying customers in advance of the system implementation" refers to methods and means for communicating information and precautions regarding the system implementation to customers in advance.

[1134] "Means for detecting harassment while staff are interacting with customers in a physical store, and for AI to take over the response" refers to a device or system configuration that has the functionality to enable a system to detect customer harassment while staff are interacting with customers in a physical store, and for AI to take over the interaction.

[1135] This invention is a system that detects harassment that occurs when staff interact with customers in brick-and-mortar stores and transfers the response to an AI system, thereby reducing the workload of staff and improving work efficiency. The specific configuration and operation of the system are described in detail below.

[1136] System configuration

[1137] The system mainly consists of the following components:

[1138] 1. Terminals: Devices such as smart glasses or smartphones used by staff in brick-and-mortar stores.

[1139] 2. Server: A computing device that hosts the speech recognition module and the natural language processing (NLP) module.

[1140] 3. Management Console: The interface for managing and configuring the entire system.

[1141] Hardware and Software

[1142] Hardware: Smart glasses or smartphones are used to capture customer conversations in real time.

[1143] Software: We use the Google Cloud Speech-to-Text API to convert voice data to text, and Hugging Face Transformers for NLP analysis.

[1144] Data processing: Converts voice data into text, detects customer harassment through NLP analysis, and generates appropriate AI responses.

[1145] Acquiring and analyzing voice input

[1146] The device captures the customer's conversation in real time and sends the audio data to the server, which then receives the data and converts it into text using the Google Cloud Speech-to-Text API.

[1147] Natural language processing of text data

[1148] The converted text data is then analyzed by an NLP module, which uses Hugging Face's Transformers library to detect customer harassment. The NLP module performs the analysis using historical data and keyword filtering.

[1149] AI-powered response

[1150] If customer harassment is detected, the server notifies the operator and automatically switches to the AI-enabled module, which generates an appropriate AI response based on the analysis results and sends it to the terminal in real time. This relieves staff from psychological burden and improves work efficiency.

[1151] Advance notice to customers

[1152] Users will notify their customers in advance that the system is in place, which includes AI-based detection of harassing behavior and automatic response measures.

[1153] Specific examples

[1154] For example, if a store staff member is interacting with a customer and the customer says, "I don't like the product! Call the manager!", the system will detect this. The NLP module analyzes the text data, and if it recognizes harassment, the server switches to the AI-enabled module. The AI ​​generates a response such as, "Sir, we'll help you find a quieter location," which is then conveyed to the customer via their device.

[1155] Prompt Sentence Examples

[1156] "How would you respond when a customer says, 'I don't like the product! Call the store manager!' What would be an appropriate response?"

[1157] In this way, the system reduces the burden on staff and contributes to improved work efficiency and customer satisfaction.

[1158] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1159] Step 1:

[1160] The terminal captures the conversation with the customer in real time and sends the voice data to the server. At this time, the terminal's voice input function is used to acquire the input voice data as a digital signal. The input data is then transferred from the terminal to the server.

[1161] Step 2:

[1162] The server converts the received voice data into text data using the Google Cloud Speech-to-Text API. The speech recognition module analyzes the voice data and generates corresponding text data. The server analyzes the input voice data and outputs it as text data.

[1163] Step 3:

[1164] The server sends the converted text data to a natural language processing (NLP) module to analyze the text data. It uses Hugging Face's Transformers library to detect harassing remarks in the text data. For analysis, the NLP module receives the text data as input and outputs the customer harassment detection results.

[1165] Step 4:

[1166] The server receives the analysis results from the NLP module and automatically switches to the AI-enabled module if harassment is detected. The instruction to switch is based on the harassment detection results as input data, and switching to the AI-enabled module as output.

[1167] Step 5:

[1168] The AI-enabled module generates an appropriate AI response based on the analysis results. The generative AI model creates the generative AI response in real time and sends it to the device. It receives the analysis results as input and outputs the appropriate response sentence.

[1169] Step 6:

[1170] The device receives the AI ​​response and conveys it to the customer as a voice output. The AI-generated response is provided to the customer via smart glasses or a smartphone. The AI ​​response that serves as input is converted into voice output and communicated to the customer.

[1171] Step 7:

[1172] The user notifies the customer in advance about the system's implementation and its functions. This notification uses prompts to explain information about the system's implementation and the AI's harassment detection function. The content of the notification is the input data, and the message to the customer is the output.

[1173] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1174] This invention is an AI system that reduces the burden on call center operators and enables efficient responses when they are faced with customer complaints or customer harassment. Furthermore, by combining it with an emotion engine that recognizes user emotions, it improves the quality of customer service.

[1175] composition

[1176] The system consists of the following main components:

[1177] 1. Terminal: A telephone terminal for the operator.

[1178] 2. Server: A computing device that hosts the AI ​​model, speech recognition module, and emotion engine.

[1179] 3. Management Console: Operator and administrator interface.

[1180] Program processing

[1181] Acquiring voice input

[1182] The terminal captures the voice conversation between the operator and the customer in real time and sends the voice data to the server, which receives the voice data and transfers it directly to the voice recognition module.

[1183] Voice Recognition

[1184] The server converts the received voice data into text data using a speech recognition module, which is then used for subsequent natural language processing (NLP) and emotion engines.

[1185] Real-time analysis of conversation content and emotion recognition

[1186] The server inputs the received text data into the NLP module to analyze the conversation content. At the same time, the emotion engine analyzes the emotion based on the user's voice tone, speed, and text content. The analysis results of the emotion engine are integrated with the results of the NLP module.

[1187] Customer Harassment Detection

[1188] The server receives the analysis results from the NLP module, and if customer harassment is detected, it notifies the operator and automatically switches to the AI-enabled module.

[1189] AI-powered response

[1190] The AI-enabled module generates an appropriate response, dynamically adjusting the AI ​​response content based on the analysis results of the emotion engine, and transmitting the generated AI response to the device in real time.

[1191] Advance notice to customers

[1192] When the system is introduced, the user will provide advance notice to customers, including the fact that the system is being introduced to prevent customer harassment and that it also has emotion recognition capabilities.

[1193] Specific examples

[1194] For example, when an operator receives a call from a customer, the voice is captured on the device and immediately sent to the server. The server converts the voice into text and analyzes it using the NLP module, while the emotion engine analyzes the tone and speed of the voice. If a customer makes a harassing remark such as, "Poor service! Call the manager!" and the emotion engine detects a high level of anger, the server will comprehensively assess this, notify the operator, and switch to the AI-enabled module. The AI ​​will then generate an appropriate response based on the emotion, such as, "We apologize for the inconvenience," and respond to the customer via the device. This process reduces the operator's burden and improves the quality of customer service.

[1195] In this way, the system reduces the burden on operators and contributes to improving work efficiency and customer satisfaction.

[1196] The processing flow will be explained below.

[1197] Step 1:

[1198] The terminal captures voice inputs exchanged between the operator and the customer in real time, and the captured voice data is immediately sent to the server.

[1199] Step 2:

[1200] The server passes the received voice data to a speech recognition module, which converts the voice into text. This conversion process is performed in real time, and the resulting text data is generated.

[1201] Step 3:

[1202] The server then inputs the generated text data into a natural language processing (NLP) module to analyze the conversation, but at this stage it does not yet make a judgment as to whether or not customer harassment has occurred.

[1203] Step 4:

[1204] In parallel, the server passes the text data and voice data to the emotion engine, which analyzes the user's emotions based on the tone, speed, and text content of the voice.

[1205] Step 5:

[1206] The server integrates the analysis results from the NLP module and the emotion engine to determine whether customer harassment has occurred and the user's emotional state. If the emotion engine detects high emotions such as "anger" or "irritation," the possibility of harassment increases.

[1207] Step 6:

[1208] If the server detects customer harassment and high emotions, it will notify the operator and automatically switch to the AI-enabled module, including the specific part that was deemed to be harassment.

[1209] Step 7:

[1210] The server then activates an AI-enabled module to generate an appropriate response that takes into account the user's emotional state. For example, if the emotion engine detects "anger," it will generate a sympathetic response such as, "We're sorry. We're very sorry for the inconvenience."

[1211] Step 8:

[1212] The server generates an AI response and sends it to the terminal in real time. The terminal then forwards the AI ​​response to the customer, and the operator responds on their behalf, allowing the customer to continue their conversation.

[1213] Step 9:

[1214] When the system is introduced, the user will provide advance notice to customers about its use, including that the AI ​​will prevent customer harassment and take appropriate action using emotion recognition.

[1215] By doing so, the system of the present invention reduces the workload of operators and realizes more sophisticated customer service. In addition, by taking into account the user's emotions, it is expected that customer satisfaction will also improve.

[1216] Example 2

[1217] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1218] In modern call center operations, operators are burdened with dealing with customer harassment and emotional complaints from customers. This burden increases the operators' mental stress, reduces work efficiency, and ultimately leads to a decline in customer satisfaction. Conventional systems did not provide an effective means to solve these problems.

[1219] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1220] In this invention, the server includes means for acquiring voice input from a terminal used by an operator, means for analyzing the acquired voice data in real time, means for converting the analyzed voice data into text, means for performing natural language processing and emotion recognition based on the converted text data to detect customer harassment, means for automatically switching from the operator to AI response when customer harassment is detected and generating a response based on the emotion recognition results, and means for notifying customers in advance of the introduction of the system. This reduces the burden on operators and enables efficient customer service that takes emotions into consideration.

[1221] "Terminal" refers to the communications equipment used by the operator and is a device for obtaining voice input.

[1222] A "server" is a computing device that performs the main processing of the system, such as analyzing voice data, natural language processing, emotion recognition, and detecting customer harassment.

[1223] "Voice input" refers to the voice data exchanged between the operator and the customer.

[1224] "Means for analyzing in real time" refers to a processing method or system configuration for instantly analyzing voice data acquired from a terminal.

[1225] A "means for converting voice data to text" is a processing method or system configuration for converting voice data to text data using voice recognition technology.

[1226] "Text data" is voice data converted into character information.

[1227] "Natural language processing" is a technical field that understands and analyzes the meaning of text data as human language.

[1228] "Emotion recognition" is a technology for identifying a speaker's emotions based on voice and text data.

[1229] "Customer harassment" refers to inappropriate comments or actions made by customers toward operators.

[1230] "AI-enabled" is the process of using artificial intelligence to automatically generate appropriate responses and assist operators.

[1231] "Harassment detection result" is information that indicates whether customer harassment is occurring based on the results of natural language processing and emotion recognition.

[1232] "Means for notifying customers in advance of the introduction of the system" refers to the methods and processes for informing customers in advance that the system will be used.

[1233] This invention is a system that reduces the burden on call center operators when dealing with customer complaints and harassment, and realizes efficient and emotion-sensitive responses. In particular, by combining it with an emotion engine that recognizes the user's emotions, the quality of customer service is improved.

[1234] System configuration

[1235] Hardware and software used

[1236] 1. Terminal: A communication device for an operator, capable of capturing voice input (e.g., a telephone or headset).

[1237] 2. Server: A computing device that performs the main processing of the system, such as analyzing voice data, natural language processing, emotion recognition, and customer harassment detection. Specific software examples include Google Speech-to-Text API, OpenAI GPT-4, and IBM Watson Tone Analyzer.

[1238] 3. Management console: An interface for operators and administrators, software for configuring and monitoring the system (e.g., a web-based management screen).

[1239] Operation procedures and examples

[1240] Acquiring voice input

[1241] The terminal captures the voices exchanged between the operator and the customer in real time. For example, when an operator receives a call from a customer, the conversation is picked up through the terminal's microphone and immediately sent to the server.

[1242] Voice Recognition

[1243] The server receives the voice data and converts it into text using the Google Speech-to-Text API. The server sends the voice data to the API and receives the text data within a few seconds.

[1244] Real-time analysis of conversation content and emotion recognition

[1245] The server inputs the received text data into a natural language processing module (OpenAI GPT-4) to analyze the content of the conversation. At the same time, an emotion engine (IBM Watson Tone Analyzer) analyzes emotions based on the tone of voice, speed, and text content. The analysis results are integrated and used for subsequent processing.

[1246] Customer Harassment Detection

[1247] The server checks the analysis results from the NLP module, and if it determines that certain words or phrases constitute customer harassment, it immediately notifies the operator and automatically switches to the AI-enabled module. For example, if the text data contains words such as "call the person in charge" or "the service is terrible," the server will detect harassment.

[1248] AI-powered response

[1249] The server activates an AI-enabled module to generate an appropriate response. At this time, a generative AI model (OpenAI GPT-4) dynamically adjusts the response content based on the results of sentiment analysis. For example, if a customer says, "The service was poor," the AI ​​generates a response such as, "We're sorry, we apologize for the inconvenience," and sends it to the device in real time.

[1250] Advance notice to customers

[1251] When introducing the system, the user notifies customers in advance, for example by email or letter, that "this system has been introduced with the aim of preventing customer harassment and is equipped with emotion recognition functionality."

[1252] Prompt Sentence Examples

[1253] The AI ​​model is fed with the following prompt:

[1254] "If the customer is enraged, generate an effective calming response. Their words are, 'Poor service! Bring out the person in charge!'"

[1255] "Please write an appropriate apology for the user's comments. The situation is that a customer is expressing dissatisfaction with a product."

[1256] This allows the system to reduce the burden on operators and provide quick and emotionally sensitive customer service.

[1257] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1258] Step 1: Getting voice input

[1259] The terminal captures the voice exchanged between the operator and the customer in real time. Specifically, the terminal's microphone receives voice input and processes this voice data as a digital signal. The captured voice data is then sent to the server. For example, when an operator says, "Hello, thank you for your inquiry," the voice is immediately converted into digital data by the terminal and sent to the server. The input here is voice data, and the output is digital voice data.

[1260] Step 2: Voice Recognition

[1261] The server receives the voice data sent from the device and passes it to a voice recognition module. The Google Speech-to-Text API is used as the voice recognition module. When voice data is sent to this API, the API converts the voice data into text data, which is then received by the server. The input here is digital voice data, and the output is text data. Specifically, the server posts the voice data to Google's API and receives the text data within a few seconds.

[1262] Step 3: Real-time analysis of conversation content and emotion recognition

[1263] The server inputs the received text data into a natural language processing module (OpenAI GPT-4). The server then sends the text data to GPT-4, which analyzes the content of the dialogue. At the same time, the emotion engine (IBM Watson Tone Analyzer) analyzes the user's emotions based on the tone of voice, speed, and text content. The server then combines the results of these two analyses. The input here is text data, and the output is analysis results based on natural language processing and emotion recognition results. Specifically, the server sends the text data to GPT-4, where it performs natural language processing, while simultaneously sending it to IBM Watson Tone Analyzer for emotion analysis.

[1264] Step 4: Detect customer harassment

[1265] The server checks the analysis results from the natural language processing module and checks whether specific words or phrases are judged to be customer harassment. For example, if the text data contains words such as "Bring the person in charge" or "The service is terrible" and the emotion engine detects a high level of anger, these conditions will trigger the detection of harassment. The input is the results of natural language processing and emotion recognition, and the output is a customer harassment detection flag. In concrete terms, the server filters the analysis results based on the conditions and determines whether harassment has occurred.

[1266] Step 5: AI response

[1267] The server launches an AI-enabled module and generates an appropriate response. At this time, the generative AI model (OpenAI GPT-4) dynamically adjusts the response content based on the results of emotion analysis. For example, if a customer is angry and complains about poor service, the AI ​​generates a response such as "We apologize for the inconvenience." The server then sends the generated AI response to the device in real time. The input is the emotion recognition result and harassment detection flag, and the output is the response message generated by the AI. Specifically, the server sends a prompt to GPT-4, which generates a response and sends the result to the device.

[1268] Step 6: Notify customers in advance

[1269] When the system is introduced, the user notifies customers in advance. For example, by email or letter, the user can explain that "this system has been introduced to prevent customer harassment and is equipped with emotion recognition functionality." The input is the notification content, and the output is the notification status to the customer. In concrete terms, the user references the customer database and sends a mass notification using an email system or postal system.

[1270] This detailed processing step reduces the burden on the operator and ensures a fast and empathetic customer service experience.

[1271] (Application example 2)

[1272] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1273] Traditional customer support and security operators are required to respond quickly and appropriately to harassment comments and emergency situations from customers, but this process is extremely stressful and places a heavy psychological burden on the operators. Furthermore, there is a lack of systems that can properly recognize customer emotions and generate optimal responses, which can lead to inconsistencies in the quality of customer support. There is a need for a system that can solve these issues, reduce the burden on operators, and provide efficient, high-quality customer support.

[1274] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1275] In this invention, the server includes means for acquiring voice input from a terminal used by an operator, means for analyzing the acquired voice data in real time, means for converting the analyzed voice data into text, means for performing natural language processing based on the converted text data to detect customer harassment, means for automatically switching from the operator to AI response when customer harassment is detected, means for notifying the customer in advance of the system implementation, means for recognizing emotions based on the converted text data and tone of voice, means for evaluating the risk level of customer harassment based on the detected emotion data, and means for detecting customer emergencies, identifying emotions of high tension or fear, and generating an appropriate response. This makes it possible to reduce the burden on operators and improve the quality of customer service.

[1276] "Terminal" means a communication device used by an operator to obtain voice input.

[1277] "Audio data" is an audio signal obtained from a terminal expressed as digital information.

[1278] "Real-time" refers to the fact that audio data is processed immediately after it is acquired.

[1279] "Text data" is voice data converted into text information using voice recognition technology.

[1280] "Natural language processing" is a technology that allows computers to understand and analyze human language.

[1281] "Customer harassment" refers to malicious words, actions, and behavior from customers toward operators.

[1282] "AI-enabled" refers to using artificial intelligence to automatically generate responses to customers.

[1283] "Emotion recognition" is a technology that analyzes voice data and text data to estimate the emotional state of a speaker.

[1284] "Risk level" is a numerical or index expression of the danger or urgency of customer harassment.

[1285] An "emergency incident" is a serious event or emergency that requires immediate action.

[1286] "Response generation" is the process of creating an appropriate response based on the analysis results.

[1287] The present invention relates to a system that enables an operator to efficiently process voice data from a customer and generate an appropriate response. The specific configuration and functions of this system are described below.

[1288] System Configuration

[1289] 1. Terminal

[1290] The terminal is a communication device that allows the operator to communicate with the customer and is used to obtain voice input. Specifically, this applies to smartphones and headsets.

[1291] 2. Server

[1292] The server is a high-performance computing device that hosts multiple functional modules, including a speech recognition module, a natural language processing module, an emotion recognition engine, and an AI response generation module.

[1293] 3. Management Console

[1294] The management console is an interface that allows operators and administrators to monitor the status of the system and change its settings.

[1295] Program processing

[1296] Acquiring voice input

[1297] The device captures conversations with customers in real time and sends the audio data to a server. The hardware used is the smartphone's microphone, and the software is a voice capture module.

[1298] Voice Recognition

[1299] The server receives the audio data and converts it into text using the Google Cloud Speech-to-Text API, which is then used for further processing.

[1300] Natural Language Processing and Emotion Recognition

[1301] The server passes the text data obtained from the voice recognition module to the natural language processing module (SpaCy) and analyzes the conversation. At the same time, it uses the emotion recognition engine (Emotion Recognition API) to analyze the customer's emotions based on the tone and speed of their voice. This allows it to identify the customer's emotional state, such as tension or fear.

[1302] Customer Harassment Detection

[1303] A natural language processing module is used to analyze text data to detect customer harassment, along with historical data and keyword filtering.

[1304] Risk Assessment

[1305] The risk level of customer harassment is assessed based on emotional data obtained from the emotion recognition module. If high levels of tension or fear are detected, the risk level is assessed as high.

[1306] AI-powered response generation

[1307] If customer harassment or a high risk level is detected, the server automatically activates the AI ​​response module to generate an appropriate response, which is then sent to the terminal in real time and provided to the customer. This uses OpenAI's GPT-4 model.

[1308] Specific examples

[1309] For example, when a customer reports an emergency situation on a security hotline, such as "Help! My window is broken," the device captures the voice and sends it to a server. The server converts the voice data into text and performs natural language processing and emotion recognition. If the server detects a high level of fear based on the analysis results, it notifies the operator, and the AI ​​immediately generates an appropriate response, such as "We will immediately dispatch a security team. Please evacuate to a safe location."

[1310] Prompt Sentence Examples

[1311] Audio data: "Help! The window is broken."

[1312] Prompt statement:

[1313] "Analyze the following voice input, understand the customer's sentiment, and generate an appropriate response.

[1314] Voice input: "Help! The window is broken."

[1315] Analysis: Customers have high levels of fear.

[1316] Response generated:"

[1317] Example response: 'We will be sending a security team immediately. Please seek safety.'

[1318] This reduces the burden on operators and improves the quality of customer service.

[1319] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1320] Step 1:

[1321] Acquiring voice input

[1322] The device captures the conversation with the customer in real time and sends the voice data to the server. The input is the customer's voice, and the output is digital voice data. In this process, the smartphone or headset acts as a microphone, and the voice capture module converts the voice signal into digital information.

[1323] Step 2:

[1324] Voice Recognition

[1325] The server converts the received voice data into text data using the Google Cloud Speech-to-Text API. The input is digital voice data, and the output is text data. In this process, a speech recognition model analyzes the voice waveform and generates corresponding text.

[1326] Step 3:

[1327] Natural Language Processing and Emotion Recognition

[1328] The server inputs the text data into a natural language processing module such as SpaCy to analyze the conversation content. At the same time, it uses the Emotion Recognition API to analyze the customer's emotional state from the text and voice features. The input is text data and voice feature data, and the output is the analyzed conversation content and emotion data. In this process, the natural language processing module analyzes the meaning of words and phrases, and the emotion recognition engine analyzes the tone and speed of the voice to identify emotions.

[1329] Step 4:

[1330] Customer Harassment Detection

[1331] The server uses the analysis results from the natural language processing module to detect customer harassment. Past data and keyword filtering are also used. The input is the analysis result of the text, and the output is the detection result of customer harassment. This process checks whether specific keywords or phrases are included in the text to determine whether harassing behavior is detected.

[1332] Step 5:

[1333] Risk Assessment

[1334] The server uses the emotion data obtained from the emotion recognition module to evaluate the risk level of customer harassment. The input is the customer's emotion data, and the output is the risk level evaluation result. In this process, if high tension or fear is detected, the risk level is set high.

[1335] Step 6:

[1336] AI-powered response generation

[1337] When the server detects customer harassment or a high risk level, it activates an AI response module to generate an appropriate response. The generated AI response is sent to the terminal in real time and provided to the customer. The input is the risk level and harassment detection result, and the output is the generated appropriate response. This process utilizes OpenAI's GPT-4 model to generate a response with the appropriate context and tone.

[1338] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1339] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1340] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1341] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1342] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1343] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1344] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1345] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1346] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1347] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1348] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1349] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1350] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1351] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1352] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1353] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1354] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1355] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1356] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1357] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1358] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1359] The following is further disclosed regarding the above embodiment.

[1360] (Claim 1)

[1361] a means for obtaining voice input from a terminal used by an operator;

[1362] A means for analyzing the acquired voice data in real time;

[1363] a means for converting the analyzed voice data into text;

[1364] A means for detecting customer harassment by performing natural language processing based on the converted text data;

[1365] A means to automatically switch from an operator to an AI response when customer harassment is detected, and

[1366] A means of notifying customers in advance of the system's implementation;

[1367] A system including:

[1368] (Claim 2)

[1369] 2. The system of claim 1, wherein the AI ​​response means includes means for generating an appropriate AI response based on the harassment detection result.

[1370] (Claim 3)

[1371] 10. The system of claim 1, wherein the natural language processing means includes means for utilizing historical data and keyword filtering for harassment detection.

[1372] "Example 1"

[1373] (Claim 1)

[1374] a means for obtaining voice input from a terminal used by an operator;

[1375] means for transmitting the acquired audio data to a computing device in real time;

[1376] means for the computing device to analyze the transmitted voice data with a voice recognition module;

[1377] A means for converting the voice data analyzed by the voice recognition module into text data;

[1378] A means for inputting the converted text data into a natural language processing module and analyzing the conversation content in real time;

[1379] A means for detecting customer harassment from the analysis results using a natural language processing module;

[1380] A means to automatically switch from an operator to an AI response when customer harassment is detected, and

[1381] A means for sending the generated AI response to the terminal in real time and responding to the customer;

[1382] A means of notifying customers in advance of the system's implementation;

[1383] A system including:

[1384] (Claim 2)

[1385] 10. The system of claim 1, wherein the system uses a generative AI model to generate appropriate AI responses and uses reply prompt sentences.

[1386] (Claim 3)

[1387] 10. The system of claim 1, wherein the natural language processing module utilizes historical data and keyword filtering for harassment detection.

[1388] "Application Example 1"

[1389] (Claim 1)

[1390] a means for obtaining voice input from a terminal used by an operator;

[1391] A means for analyzing the acquired voice data in real time;

[1392] a means for converting the analyzed voice data into text;

[1393] A means for detecting customer harassment by performing natural language processing based on the converted text data;

[1394] A means to automatically switch from an operator to an AI response when customer harassment is detected, and

[1395] A means of notifying customers in advance of the system's implementation;

[1396] A method for detecting harassment while staff are interacting with customers in physical stores and for AI to take over the response.

[1397] A system including:

[1398] (Claim 2)

[1399] 2. The system of claim 1, wherein the AI ​​response means includes means for generating an appropriate AI response based on the harassment detection result.

[1400] (Claim 3)

[1401] 10. The system of claim 1, wherein the natural language processing means includes means for utilizing historical data and keyword filtering for harassment detection.

[1402] "Example 2: Combining Emotion Engines"

[1403] New Claims

[1404] (Claim 1)

[1405] a means for obtaining voice input from a terminal used by an operator;

[1406] A means for analyzing the acquired voice data in real time;

[1407] a means for converting the analyzed voice data into text;

[1408] A means for detecting customer harassment by performing natural language processing and emotion recognition based on the converted text data;

[1409] When customer harassment is detected, the system automatically switches from an operator to an AI response, generating a response based on emotion recognition results.

[1410] A means of notifying customers in advance of the system's implementation;

[1411] A system including:

[1412] (Claim 2)

[1413] 10. The system of claim 1, further comprising means for generating an appropriate AI response based on the harassment detection result and the emotion recognition result.

[1414] (Claim 3)

[1415] 10. The system of claim 1, wherein the natural language processing means includes means for utilizing historical data and keyword filtering for harassment detection.

[1416] "Application example 2 when combining emotion engines"

[1417] (Claim 1)

[1418] a means for obtaining voice input from a terminal used by an operator;

[1419] A means for analyzing the acquired voice data in real time;

[1420] a means for converting the analyzed voice data into text;

[1421] A means for detecting customer harassment by performing natural language processing based on the converted text data;

[1422] A means to automatically switch from an operator to an AI response when customer harassment is detected, and

[1423] A means of notifying customers in advance of the system's implementation;

[1424] means for recognizing emotions based on the converted text data and tone of voice;

[1425] a means for assessing the risk level of customer harassment based on the detected emotion data;

[1426] A system that detects customer emergencies and includes a means for identifying high levels of stress or fear and generating an appropriate response.

[1427] (Claim 2)

[1428] 10. The system of claim 1, further comprising means for generating an appropriate AI response based on the harassment detection result and the emotion data.

[1429] (Claim 3)

[1430] 10. The system of claim 1, further comprising means for utilizing historical data and keyword filtering for natural language processing and emotion recognition. [Explanation of symbols]

[1431] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for obtaining voice input from a terminal used by an operator; A means for analyzing the acquired voice data in real time; a means for converting the analyzed voice data into text; natural language processing means for performing natural language processing based on the converted text data to detect customer harassment; An AI response method that automatically switches from an operator to an AI response when customer harassment is detected, A means of notifying customers in advance of the system's implementation; A system including:

2. 2. The system of claim 1, wherein the AI ​​response means includes means for generating an appropriate AI response based on the harassment detection result.

3. 10. The system of claim 1, wherein the natural language processing means includes means for utilizing historical data and keyword filtering for harassment detection.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A