system
The system addresses the inefficiency in managing customer voice data by converting, classifying, and summarizing it, enabling efficient data analysis and timely business insights.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2026-04-09
AI Technical Summary
Existing contact center systems struggle to efficiently classify, summarize, and analyze customer voice data for effective management decisions due to underutilization of texturized data from voice recognition technology.
A system that collects voice data, converts it into text, classifies and summarizes it using machine learning models, stores the data in a database, and aggregates it within specified periods to facilitate efficient information retrieval and analysis.
Enables automatic classification and summarization of customer inquiries, providing timely and useful information for business decisions by organizing and analyzing voice data effectively.
Smart Images

Figure 2026062306000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In order to efficiently classify and summarize the content of inquiries from customers and make it useful for management, it is necessary to extract and organize useful information from a large amount of data. However, in the current contact center, although there is data that has been texturized using voice recognition technology, the data has not been fully utilized. For this reason, there is a problem that it is difficult to quickly and efficiently obtain the information necessary for management decisions. Specifically, there is a need for a system that classifies, summarizes, stores, and analyzes customer voices.
Means for Solving the Problems
[0005] In order to solve the above problems, the present invention provides the following means.
[0006] The present invention provides a system that includes means for collecting voice data of inquiries from multiple customers, means for converting the voice data into text data, means for classifying the converted text data into a number of predefined categories, means for summarizing the classified text data, means for storing the classified data and the summarized data, and means for aggregating data within a specified period and calculating the number of inquiries for each category.
[0007] Furthermore, the present invention provides a system that includes means for using a machine learning model to classify text data and means for using a machine learning model to summarize the text data.
[0008] The system also includes means for using speech recognition technology to convert speech data into text data, and means for using a database to store and retrieve the text data.
[0009] This makes it possible to automatically classify and summarize customer inquiries, and efficiently acquire information that is useful for business decisions.
[0010] "Inquiry voice data" refers to inquiry information in voice format sent by customers.
[0011] "Converting to text data" refers to the process of converting audio data into written information.
[0012] "Categorizing" refers to dividing text data into multiple groups based on specific criteria or rules.
[0013] "Summarization" refers to extracting and shortening the essential parts of the original text data.
[0014] "Saving data" refers to recording information in a memory area so that it can be retrieved later.
[0015] "Aggregate data within a specified period" refers to performing statistical processing based on data collected within a specific period.
[0016] "Machine learning model" refers to a computational model that learns patterns from input data and performs classification and prediction.
[0017] "Speech recognition technology" refers to a technology for analyzing speech information and converting it into data in text format.
[0018] "Database" refers to a structure for efficiently storing, managing, and retrieving large amounts of data.
Brief Description of Drawings
[0019] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.
Mode for Carrying Out the Invention
[0020] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0021] First, the terms used in the following description will be explained.
[0022] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0023] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0024] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0025] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0027] [First Embodiment]
[0028] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0029] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0030] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0031] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0032] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0034] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0035] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0036] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0037] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0038] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0039] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0040] This invention relates to a system for collecting customer inquiry voice data and efficiently classifying and summarizing that data. This system operates in cooperation with a server, terminals, and users.
[0041] Data collection
[0042] The server first collects inquiry audio data from multiple customers. For example, inquiry audio data is audio files recorded when a customer calls the call center. These audio files are stored in a database.
[0043] Speech recognition and text conversion
[0044] The device converts the collected audio data into text data using speech recognition technology. Specifically, it extracts text information from the audio data using a speech recognition library. This process converts the audio data into text data.
[0045] Text classification and summarization
[0046] The server classifies the converted text data into several predefined categories. Machine learning models are used for classification. For example, categories such as "Delivery," "Returns," and "Payment" are predefined, and the server determines which category each text belongs to. Furthermore, the server summarizes the text data. For summarization, machine learning models are used to extract the important parts of the original text and save them in a shortened format.
[0047] Data storage
[0048] The server stores the categorized and summarized data in a database. This organizes the query results, making them easy to search and analyze later.
[0049] Data aggregation and reporting
[0050] The system aggregates data within a user-specified period and calculates the number of inquiries for each category. This aggregated data is crucial for managers to understand customer feedback and make appropriate business decisions.
[0051] Specific example
[0052] For example, if a customer asks, "Please tell me the delivery status of my order," this audio is first collected and converted to text using speech recognition. The text is categorized under "Delivery," and a summary such as "Inquiry about checking delivery status" is generated. This data is stored in a database, and by aggregating the data for a specified period, managers can identify which inquiries fall under the "Delivery" category.
[0053] In this way, the present invention is a system for efficiently classifying and summarizing customer inquiries and providing information useful for business management.
[0054] The following describes the processing flow.
[0055] Step 1:
[0056] A user makes a phone call to the contact center with an inquiry. The content of the inquiry is recorded as voice data.
[0057] Step 2:
[0058] The server stores the inquiry voice data in a database. The database also includes information such as the time of the inquiry and the caller's information.
[0059] Step 3:
[0060] The device converts stored audio data into text data using speech recognition technology. Specifically, it analyzes the audio data using a speech recognition library and converts it into text information.
[0061] Step 4:
[0062] The server stores the converted text data in a database. This allows for centralized management of audio data and its corresponding text data.
[0063] Step 5:
[0064] The server uses machine learning models to categorize the stored text data. For example, the text is automatically classified into categories such as "delivery," "returns," and "payment."
[0065] Step 6:
[0066] The server summarizes each classified text data. This summarization uses machine learning models to extract key parts of the text and generate new text data in a shortened format.
[0067] Step 7:
[0068] The server stores the classified and summarized data in a new database. This ensures that each query is stored in a structured format.
[0069] Step 8:
[0070] The system aggregates data for a period specified by the user. The server retrieves data from the database for that period and calculates the number of queries for each category.
[0071] Step 9:
[0072] The server generates aggregated results in report format. The report includes the number of inquiries for each category and data showing specific trends.
[0073] Step 10:
[0074] The terminal provides the user with the generated report. The report is displayed through a web interface or dashboard.
[0075] (Example 1)
[0076] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0077] Traditional customer service systems struggled to effectively collect, classify, and summarize customer inquiry voice data. This resulted in challenges in quickly and accurately understanding inquiries and providing information to support appropriate business decisions.
[0078] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0079] In this invention, the server includes means for collecting inquiry voice data from multiple customers, means for converting the voice data into text data, and means for classifying the converted text data into a plurality of predefined categories. This makes it possible to efficiently perform a series of processes within the system, from collecting inquiry voice data to classifying, summarizing, and storing it, as well as aggregating the data.
[0080] A "customer" is a person or organization that receives various services or products, and is the person to whom inquiries about those services or products are made.
[0081] "Inquiry voice data" refers to data recorded in audio format of customer inquiries about services or products.
[0082] "Means of collecting voice data" refers to the technology and equipment used to record and save customer inquiries.
[0083] "Means of converting to text data" refers to speech recognition technologies and software used to convert audio data into text information.
[0084] A "predefined category" is a set of categories that are set in advance to efficiently classify the content of inquiries, and may include, for example, "delivery," "returns," and "payment."
[0085] "Means of classification" refers to techniques and machine learning models used to analyze text data and classify it based on predefined categories.
[0086] "Means of summarization" refers to techniques and generative models that extract important parts from original text data and summarize the content concisely.
[0087] "Means of preservation" refers to technologies and devices for recording classified and summarized data in databases or other storage systems.
[0088] "Means for aggregating and calculating the number of inquiries for each category" refers to technologies and algorithms for analyzing data stored in a database and aggregating the number of inquiries for each category within a specified period.
[0089] A "database" refers to a system or software for efficiently storing and managing audio data, text data, classification data, summary data, and other similar information.
[0090] A "machine learning model" refers to a model that analyzes data and automatically performs classification and summarization based on that analysis. This includes statistical methods and algorithms.
[0091] A "generative model" refers to an AI model or algorithm used to generate new data or summaries from provided data.
[0092] A "prompt statement" is a phrase input into a generative model, and it refers to an instruction that the model uses to generate the desired output.
[0093] This invention is a system for collecting customer inquiry voice data and efficiently classifying and summarizing that data. This system operates through the coordinated efforts of a server, terminals, and users.
[0094] Data collection
[0095] The server first collects voice data from multiple customers. This voice data consists of audio files recorded when customers call the call center. Specifically, it is stored in a database on the server. To collect the voice data, a VoIP system is used to capture voice communications, which are then temporarily stored as audio files.
[0096] Speech recognition and text conversion
[0097] The device retrieves audio data stored on the server and converts it into text data using speech recognition technology. Examples of speech recognition libraries used include Google® Cloud Speech-to-Text API. The device sends the audio data to the API, receives the recognition result, and saves it in text format. The converted text data is then saved to the server's database.
[0098] Text classification and summarization
[0099] The server classifies the converted text data into several predefined categories using a machine learning model. For example, categories such as "Delivery," "Returns," and "Payment" are predefined, and the server automatically determines which category each text belongs to. This process uses a pre-trained classification model with scikit-learn. It also summarizes the text data using a generative AI model (e.g., GPT-3®). The generated summaries are stored in a database.
[0100] Data storage
[0101] The classified text data and summary data are stored in a database by the server. During this process, they are organized into records containing category labels and summary text, and then saved.
[0102] Data aggregation and reporting
[0103] Users request data aggregation for a specified period. Users can enter the specified period through the web application interface. The server queries the database based on the specified period and aggregates the number of queries for each category. This is done using SQL queries. The aggregation results are reported to the user and displayed on the web application's dashboard.
[0104] Specific example
[0105] For example, if a customer asks, "Please tell me the delivery status of my order," the server first collects this audio. Then, the terminal uses speech recognition technology to convert the audio data into text, "Please tell me the delivery status of my order." This text data is then categorized by the server as "Delivery" and further summarized as "Inquiry about checking delivery status" using a generative AI model. This data is stored in a database, and later, the system aggregates the number of inquiries in the "Delivery" category and reports it to the user.
[0106] Example of a prompt
[0107] The following is an example of a prompt to input into a generative AI model:
[0108] "Convert the customer inquiry audio 'Please tell me the delivery status of my item' into text, and categorize that text under 'Delivery.' Furthermore, summarize this text and save it to the database."
[0109] In this way, the present invention is a system for efficiently classifying and summarizing customer inquiries and providing information useful for business management.
[0110] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0111] Step 1:
[0112] The server collects voice data from multiple customers. When a customer calls the call center, the VoIP system captures the voice data and temporarily stores it as an audio file. The input here is the customer's voice, and the output is an audio file. This audio file is stored in a database.
[0113] Step 2:
[0114] The device retrieves an audio file stored on the server. The input is the path to the audio file, and the output is the retrieved audio file itself. Specifically, the audio file is downloaded using a REST API.
[0115] Step 3:
[0116] The device converts the acquired audio file into text data using the Google Cloud Speech-to-Text API. The input is an audio file, and the output is the converted text data. The audio data is sent to the API, and the recognition results are received in text format.
[0117] Step 4:
[0118] The terminal sends the converted text data to the server. The input is text data, and the output is text data stored in the server's database. Specifically, a REST API is used to send text data to the server and store it in the database.
[0119] Step 5:
[0120] The server inputs text data stored in a database into a machine learning model and classifies it into predefined categories. The input is text data, and the output is the classification result of the text. The model used here is a classification model trained with scikit-learn. Specifically, it generates data with category labels attached.
[0121] Step 6:
[0122] The server inputs classified text data into a generating AI model (e.g., GPT-3) to generate a summary. The input is classified text data, and the output is summarized text data. The summary extracts the important parts and puts them into a shorter format.
[0123] Step 7:
[0124] The server stores the classified and summarized data in a database. The input is category labels and summary text data, and the output is the records stored in the database. This ensures that the data is systematically stored, facilitating later searching and analysis.
[0125] Step 8:
[0126] The user requests data aggregation for a specified period. The input is the specified period (e.g., a date range), and the output is the query aggregation results for each category within that period. The user submits the request via a web application.
[0127] Step 9:
[0128] The server queries the database based on a specified period and aggregates the number of queries for each category. The input is period information, and the output is the aggregated result. SQL queries are used to search the database and calculate the number of queries for each category.
[0129] Step 10:
[0130] The server reports the aggregated results to the user. The input is the aggregated results, and the output is a report displayed on the web application's dashboard. The user views this report and uses it to inform business decisions.
[0131] (Application Example 1)
[0132] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0133] Traditional customer inquiry systems had problems such as the time it took to collect, classify, and summarize voice data, making real-time data display and analysis difficult. This reduced the efficiency of customer service and made it difficult for managers to make quick business decisions.
[0134] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0135] In this invention, the server includes means for collecting voice data of inquiries from multiple customers, means for converting the voice data into text data, means for classifying the converted text data into a plurality of predefined categories, means for summarizing the classified text data, means for storing the classified data and the summarized data, means for aggregating data within a specified period and calculating the number of inquiries for each category, and means for collecting and displaying the text data in real time. This improves the efficiency of customer service and enables managers to make quick business decisions.
[0136] "Multiple customers" refers to two or more users who are the entities making inquiries.
[0137] "Inquiry voice data" refers to voice information uttered by customers, which includes questions and requests.
[0138] "Means of collection" refers to systems and devices for recording audio data in digital format and transmitting it to a server.
[0139] "Text data" refers to audio data converted into written information, which can then be processed mechanically.
[0140] "Means of conversion" refer to technologies and systems for converting audio data into text information.
[0141] "Multiple predefined categories" refers to several pre-defined classification items, such as "delivery," "returns," and "payment."
[0142] "Means of classification" refer to technologies and systems for sorting text data into predefined categories.
[0143] "Methods of summarization" refer to technologies and systems for extracting important parts from text data and summarizing them concisely.
[0144] "Means of preservation" refers to systems for storing data in databases or storage devices.
[0145] "Data within a specified period" refers to query data collected within a specific time frame.
[0146] "Means for calculating the number of inquiries per category" refers to technologies or systems for calculating the number of inquiries classified into each category.
[0147] "Means of collecting and displaying data in real time" refers to technologies and systems for instantly collecting and visually displaying inquiry data.
[0148] A "machine learning model" is an algorithm or program that learns from large amounts of data and automatically performs tasks such as data classification and summarization.
[0149] "Speech recognition technology" is a technology that converts speech data into text information.
[0150] A "database" is a system for efficiently storing, searching, and accessing large amounts of data.
[0151] A "visualization algorithm" is a technology that displays collected data in the form of graphs and charts, making it easy to understand visually.
[0152] An "algorithm" is a set of procedures or computational steps for solving a specific problem.
[0153] This invention relates to a system for collecting customer inquiry voice data and efficiently classifying and summarizing that data. The system operates in cooperation with a server, terminals, and users.
[0154] Data collection
[0155] The server collects voice data from multiple customers. This voice data includes, for example, audio files recorded when a customer calls the call center. These audio files are stored in a database.
[0156] Speech recognition and text conversion
[0157] The device converts the collected audio data into text data using speech recognition technology. Specifically, it uses the Google Speech-to-Text API. It extracts text information from the audio data and converts it into text data.
[0158] Text classification and summarization
[0159] The server classifies the converted text data into several predefined categories. This classification uses machine learning models (e.g., TENSORFLOW®, PyTorch). For example, categories such as "Shipping," "Returns," and "Payment" are predefined, and the server determines which category each text belongs to. Furthermore, the server summarizes the text data. For summarization, the server also uses machine learning models to extract the important parts of the original text and saves them in a shortened format.
[0160] Real-time display
[0161] The server implements a visualization algorithm to collect and display the aforementioned text data in real time. This allows inquiry data to be retrieved immediately, enabling stakeholders to understand the situation in real time.
[0162] Data storage
[0163] The server stores categorized and summarized data in a database. This organizes query results, making them easy to search and analyze later. Databases used include Firebase and MySQL®.
[0164] Data aggregation and reporting
[0165] The system aggregates data within a period specified by the user and calculates the number of inquiries for each category. This aggregated result provides important data for managers to understand customer feedback and make appropriate business decisions.
[0166] Specific example
[0167] For example, if a customer asks, "Please tell me the delivery status of my order," the audio is first collected and converted to text using speech recognition. The text is categorized under "Delivery," and a summary such as "Inquiry about checking delivery status" is generated. This data is stored in a database, and by aggregating the data for a specified period, managers can understand which inquiries fall under the "Delivery" category. Furthermore, because the collected information is displayed in real time, it is possible to shorten the time from the moment an inquiry is made until a response is given.
[0168] Example of a prompt
[0169] Use the Google Speech-to-Text API to convert speech to text, input the text into a TensorFlow model, and categorize it into categories such as "Shipping," "Returns," and "Payment." Furthermore, extract and summarize the key points. Save the classification results and summaries to a Firebase database. Finally, aggregate the number of inquiries within a specified period and display it as a report.
[0170] This significantly improves the efficiency of customer service, allowing managers to quickly implement appropriate countermeasures.
[0171] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0172] Step 1: Collect audio data
[0173] The server collects customer inquiry voice data in real time. Specifically, it records the voice call when a customer makes an inquiry and transfers the audio file to the server. The input is the customer's voice data, and the output is an audio file in digital format.
[0174] Step 2: Converting audio data to text
[0175] The device converts the collected audio data into text data. Here, the Google Speech-to-Text API is used to convert the audio data into text information. The input is the audio data collected in step 1, and the output is the converted text data.
[0176] Step 3: Classification of text data
[0177] The server classifies the converted text data into several predefined categories. A machine learning model (e.g., TensorFlow) is used to execute an algorithm that classifies the text data. The input is the text data obtained in step 2, and the output is the text data classified by category.
[0178] Step 4: Textbook Summary
[0179] The server summarizes the classified text data. It uses a machine learning model (e.g., PyTorch) to extract and shorten the important parts. The input is the text data classified in step 3, and the output is the summarized text data.
[0180] Step 5: Save Data
[0181] The server stores the classified and summarized data in a database. Specifically, it uses a database management system such as Firebase or MySQL to store the data. The input is the data obtained in steps 3 and 4, and the output is the state in which it is stored in the database.
[0182] Step 6: Real-time display of data
[0183] The server collects query data in real time and displays it using a visualization algorithm. The input is all the data obtained from step 2 onward, and the output is a visual display format that is updated in real time.
[0184] Step 7: Data aggregation and reporting
[0185] The user aggregates inquiry data within a specified period and calculates the number of inquiries for each category. The server uses an aggregation algorithm to calculate the number of inquiries and displays it as a report. The input is all data stored in the database, and the output is a report of the aggregated results.
[0186] The above steps enable a system that efficiently classifies, summarizes, and displays customer inquiry voice data in real time.
[0187] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0188] This invention combines a system that collects customer inquiry voice data, efficiently classifies and summarizes that data, with an emotion engine that recognizes user emotions. This system operates in cooperation with a server, terminals, and users.
[0189] Data collection
[0190] When a user calls the contact center with an inquiry, the content of the inquiry is recorded as voice data. The server stores this voice data in a database. The database also includes information such as the time of the inquiry and the caller's information.
[0191] Speech recognition and text conversion
[0192] The device converts stored audio data into text data using speech recognition technology. It analyzes the audio data using a speech recognition library and converts it into text information. This converted text data is stored in a database.
[0193] Text classification and summarization
[0194] The server uses a machine learning model to classify the converted text data into specific categories. For example, it might be categorized as "delivery," "returns," or "payment." Another machine learning model is used for summarization, extracting the most important parts from the converted text data and saving them in a shortened format.
[0195] emotion recognition
[0196] A distinctive feature of this invention is the provision of an emotion engine. The server uses this emotion engine to recognize the user's emotions from text data. For example, it determines emotions such as "joy," "anger," and "sadness" from the wording and expressions in the text. This recognized emotion data is also stored in a database.
[0197] Data storage
[0198] The server centrally stores categorized, summarized, and sentiment data in a database. This organizes query results, making them easy to search and analyze later.
[0199] Data aggregation and analysis
[0200] To aggregate data within a user-specified period, the server retrieves data from the database for that period and calculates the number of inquiries for each category. It also analyzes sentiment data to understand user sentiment trends. This analysis provides crucial data for managers to understand customer feedback and their emotions, enabling them to make appropriate business decisions.
[0201] Specific example
[0202] For example, if a customer makes an inquiry such as, "Please tell me the delivery status of my order," this audio is first collected and converted into text using speech recognition. The text is then categorized under "Delivery," and a summary such as "Inquiry about checking delivery status" is generated. Furthermore, an emotion engine recognizes that this inquiry contains the emotion of "anxiety." All of this data is stored in a database, and by aggregating data for a specified period, managers can understand that there are many inquiries in the "Delivery" category, and that many of them contain the emotion of "anxiety."
[0203] Thus, this invention efficiently classifies and summarizes customer inquiries and provides information useful for business management by further analyzing sentiment data.
[0204] The following describes the processing flow.
[0205] Step 1:
[0206] A user makes a phone call to the contact center with an inquiry. The content of the inquiry is recorded as voice data.
[0207] Step 2:
[0208] The server stores the inquiry voice data in a database. The database also includes information such as the time of the inquiry and the caller's information.
[0209] Step 3:
[0210] The device converts stored audio data into text data using speech recognition technology. Specifically, it analyzes the audio data using a speech recognition library and converts it into text information.
[0211] Step 4:
[0212] The server stores the converted text data in a database. This allows for centralized management of audio data and its corresponding text data.
[0213] Step 5:
[0214] The server uses machine learning models to classify stored text data into specific categories. For example, categories such as "delivery," "returns," and "payments" are predefined, and the server determines which category each text belongs to.
[0215] Step 6:
[0216] The server summarizes each classified text data. This summarization uses machine learning models to extract key parts of the text and generate new text data in a shortened format.
[0217] Step 7:
[0218] The server stores the classified and summarized data in a database. This organizes the query content, making it easy to search and analyze later.
[0219] Step 8:
[0220] The server inputs the stored text data into the emotion engine to recognize the user's emotions. The emotion engine determines emotions such as "joy," "anger," and "sadness" from the text data.
[0221] Step 9:
[0222] The server stores the recognized sentiment data in a database. This allows for the management of sentiment data for each query.
[0223] Step 10:
[0224] The system aggregates data for a period specified by the user. The server retrieves data from the database for that period and calculates the number of queries and sentiment trends for each category.
[0225] Step 11:
[0226] The server generates aggregated results in report format. The report includes the number of inquiries for each category and corresponding sentiment data.
[0227] Step 12:
[0228] The terminal provides the user with the generated report. The report is displayed through a web interface or dashboard.
[0229] Specific example
[0230] For example, if a customer makes an inquiry such as, "Please tell me the delivery status of my order," this audio is first collected and converted into text using speech recognition. The text is then categorized under "Delivery," and a summary such as "Inquiry about checking delivery status" is generated. Furthermore, an emotion engine recognizes that this inquiry contains the emotion of "anxiety." All of this data is stored in a database, and by aggregating data for a specified period, managers can understand that there are many inquiries in the "Delivery" category, and that many of them contain the emotion of "anxiety."
[0231] (Example 2)
[0232] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0233] Conventional inquiry management systems struggled to efficiently collect, classify, and summarize customer inquiry voice data. Furthermore, the lack of a means to accurately recognize and analyze the emotions expressed in the inquiries presented posed challenges in providing appropriate responses and service improvements that considered customer feelings.
[0234] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for collecting voice data inquiries from multiple customers, means for converting the voice data into text data, means for classifying the converted text data into a plurality of predefined categories, means for summarizing the classified text data, means for recognizing the user's emotions from the text data, means for storing the classified data, summarized data, and emotion data, and means for aggregating data within a specified period and calculating the number of inquiries and emotion trends for each category. This enables efficient management of customer inquiries and detailed data analysis, including user emotions.
[0235] "Inquiry voice data" refers to voice information collected from customer phone calls, voice messages, and other means.
[0236] "Text data" refers to digital data obtained by converting audio data into written text.
[0237] A "machine learning model" is an algorithm that is trained using large amounts of data and automatically performs specific tasks (such as classification, prediction, and generation).
[0238] "Speech recognition technology" is a technology that analyzes speech data and converts it into text data.
[0239] A "category" is a predefined item or group used to classify the content of an inquiry.
[0240] A "summary" is data that extracts the important parts from long text data and expresses them in a shorter format.
[0241] An "emotion engine" is software or an algorithm that analyzes text data and identifies the emotions contained within it.
[0242] A "database" is a system for efficiently storing, retrieving, and managing data.
[0243] "Data aggregation" is the process of compiling data based on certain criteria and calculating statistics or totals.
[0244] "Emotional trends" refer to data that shows fluctuations and tendencies in users' emotions over a certain period of time.
[0245] This invention is a system that efficiently collects, classifies, and summarizes customer inquiry voice data, and further recognizes the user's emotions. This system operates in cooperation with a server, terminals, and users.
[0246] Data collection
[0247] When a user calls the contact center, the content of their inquiry is recorded as voice data by the server. The server stores the collected voice data, along with the time of the inquiry and the caller's information, in a database.
[0248] Specific example:
[0249] When a user asks, "Please tell me the availability of the product," the voice message is recorded by the server and stored in the database.
[0250] Speech recognition and text conversion
[0251] The audio data stored on the server is transferred to the terminal. The terminal uses a speech recognition library (e.g., Google Cloud Speech-to-Text API) to convert the audio data into text data, and then saves that text data back to the database.
[0252] Specific example:
[0253] The device converts the spoken phrase "Please tell me the product's stock status" into text data using the Google Cloud Speech-to-Text API and saves that text to a database.
[0254] Text classification and summarization
[0255] The server classifies the converted text data into specific categories using a machine learning model (e.g., BERT). It also uses another machine learning model (e.g., GPT-3) for summarization, extracting important parts from the text data and saving them in a shortened format.
[0256] Specific example:
[0257] The server converts the text "Please tell me the product's stock status" into a summary "Check product stock" and categorizes it under "Product Category".
[0258] emotion recognition
[0259] The server uses an emotion engine (e.g., Amazon Comprehend) to recognize the user's emotions from text data. The recognized emotion data is stored in a database.
[0260] Specific example:
[0261] The server recognizes the emotion of "anxiety" from the text "Please tell me the product's stock status" and stores the emotion data in the database.
[0262] Data storage
[0263] The server centrally stores audio data, text data, categorized and summarized data, and sentiment data in a database.
[0264] Specific example:
[0265] The database stores the audio file in the format "Audio file: Audio001.wav, Text: Please tell me the product's stock status, Category: Product category, Summary: Check product stock, Emotion: Anxiety".
[0266] Data aggregation and analysis
[0267] The user (manager) aggregates data for a specified period. The server retrieves this data from the database and calculates the number of inquiries and sentiment trends for each category. The server reports these analysis results to the manager to help them make informed business decisions.
[0268] Specific example:
[0269] The server generates a report stating that "data from 2023-10-01 to 2023-10-10 was compiled, and there were 50 inquiries regarding product categories, 40 of which included feelings of anxiety." The manager then reviews this report and uses it as a guide for business management.
[0270] Example of a prompt
[0271] "Design a system to collect, transcribe, categorize, and summarize customer inquiry audio, and to recognize user sentiment."
[0272] This system efficiently classifies and summarizes customer inquiries and analyzes sentiment data to provide valuable information for business management.
[0273] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0274] Step 1:
[0275] A user calls the contact center. During this call, the user verbally communicates their specific inquiry. For example, "Please tell me the availability of the product."
[0276] Input: User inquiry voice
[0277] Output: Raw audio data sent to the server
[0278] Step 2:
[0279] The server collects voice data from user inquiries in real time and saves this voice data as a file. It also adds information about the time of the inquiry and the caller, and saves this data to a database.
[0280] Input: User inquiry voice
[0281] Output: Audio file (e.g., audio001.wav) and metadata (time, sender information)
[0282] Specific operation:
[0283] The server saves the audio data as "audio001.wav" and simultaneously records it in the database in the form of "2023-10-10 10:00:00, Customer A, audio001.wav".
[0284] Step 3:
[0285] The audio data stored on the server is transferred to the terminal. The terminal converts this audio data into text data using an audio recognition library (e.g., Google Cloud Speech-to-Text API).
[0286] Input: Audio file (audio001.wav)
[0287] Output: Text data (e.g., "Please tell me the inventory status of the product")
[0288] Specific operation:
[0289] The terminal receives the audio file "audio001.wav" and calls the Google Cloud Speech-to-Text API to convert it into text. As a result, the text data "Please tell me the inventory status of the product" is generated.
[0290] Step 4:
[0291] The terminal saves the converted text data in the database. In this way, the text data is permanently recorded and can be searched later.
[0292] Input: Text data (e.g., "Please tell me the inventory status of the product")
[0293] Output: The text data is saved in the database
[0294] Specific operations:
[0295] The text "Please tell me the inventory status of the product" is saved in the database.
[0296] Step 5:
[0297] The server classifies the converted text data into specific categories using a machine learning model (e.g., BERT). For example, categories such as "Delivery", "Return", "Payment".
[0298] Input: Text data (e.g., "Please tell me the inventory status of the product")
[0299] Output: Classified category (e.g., "Product category")
[0300] Specific operations:
[0301] The server adds the label "Product category" to the database.
[0302] Step 6:
[0303] The server uses a summarization model (e.g., GPT-3) to summarize the text data. This summarized data is also saved in the database.
[0304] Input: Text data (e.g., "Please tell me the inventory status of the product")
[0305] Output: Summarized data (e.g., "Inventory confirmation of the product")
[0306] Specific operations:
[0307] The server uses the summarization model to generate a summary of "Inventory confirmation of the product" and saves it in the database.
[0308] Step 7:
[0309] The server uses an emotion engine (e.g., Amazon Comprehend) to recognize the user's emotions from text data. This emotion data is also stored in a database.
[0310] Input: Text data (Example: "Please tell me the stock status of the product.")
[0311] Output: Emotional data (e.g., "anxiety")
[0312] Specific actions:
[0313] The server uses Amazon Comprehend to recognize the emotion "anxiety" from the text "Please tell me the availability of the product" and saves it to the database.
[0314] Step 8:
[0315] The server centrally stores audio data, text data, categorized data, summary data, and sentiment data in a database. This allows for easy searching and analysis later.
[0316] Input: Audio data, text data, classification data, summary data, sentiment data
[0317] Output: Centrally organized query data
[0318] Specific actions:
[0319] The database will store the audio file in the format: "Audio file: Audio001.wav, Text: Please tell me the product's stock status, Category: Product category, Summary: Check product stock, Emotion: Anxiety".
[0320] Step 9:
[0321] The user (manager) has the server aggregate data for a specified period. The server retrieves query data from the database for that period and calculates the number of queries and sentiment trends for each category.
[0322] Input: Aggregation instructions (Example: "From 2023-10-01 to 2023-10-10")
[0323] Output: Aggregated report (e.g., number of inquiries for product categories, anxiety trend)
[0324] Specific actions:
[0325] The server retrieves data from the database for a specified period and generates a report stating that "there were 50 inquiries for the product category, 40 of which included feelings of anxiety."
[0326] Step 10:
[0327] The server provides the generated report to the user (manager). The manager then uses the report to make business decisions.
[0328] Input: Summary Report
[0329] Output: Information for business decision-making
[0330] Specific actions:
[0331] The manager reviews the provided report and understands that there are many inquiries in the "delivery" category, many of which involve feelings of "anxiety," and then considers measures to improve the situation.
[0332] (Application Example 2)
[0333] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0334] Traditional customer inquiry systems often require manual classification and summarization of inquiries, which is time-consuming and labor-intensive. Furthermore, analyzing customer sentiment is difficult, making it challenging to grasp customer satisfaction in real time. Therefore, while rapid responses based on inquiry categories and content are required, comprehensive responses, including sentiment analysis, are currently lacking. In addition, there is a lack of efficient methods for managers to derive business guidance from past inquiry data.
[0335] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0336] In this invention, the server includes means for collecting voice data of inquiries from multiple customers, means for converting the voice data into text data, means for classifying the converted text data into a plurality of predefined categories, means for summarizing the classified text data, means for recognizing customer emotions from the text data, means for storing the classified data, summarized data, and emotion data, and means for aggregating data within a specified period and calculating the number of inquiries and emotion trends for each category. This not only enables efficient classification and summarization of inquiry content, but also allows for real-time recognition of customer emotions and prompt and appropriate responses based on that information. Furthermore, it facilitates management decisions based on the stored data, leading to improved customer satisfaction and optimized operations.
[0337] "Customer" refers to the person who receives a product or service.
[0338] "Inquiry voice data" refers to voice data of questions and requests provided by customers via telephone or voice recording.
[0339] "Text data" refers to the character information obtained by converting inquiry voice data using speech recognition technology.
[0340] A "category" is a predefined group or type used to make text data easier to classify. Examples of categories include "Shipping," "Returns," and "Payment."
[0341] "Summary" means extracting the key points from text data and providing them in a shortened format.
[0342] "Emotion recognition technology" refers to techniques for recognizing a customer's psychological state and emotions from text data. This technology determines emotions such as "joy," "anger," and "sadness."
[0343] A "machine learning model" is an algorithm that uses large amounts of data to learn patterns and perform classification and prediction. It is used for text classification and summarization.
[0344] "Speech recognition technology" refers to the technology that converts speech data into text information. This is a part of speech processing technology.
[0345] A "database" refers to a system for systematically storing large amounts of data and for efficiently processing and retrieving it.
[0346] A "user interface" is an interface that allows a system and a user to interact directly. It facilitates the display and manipulation of data.
[0347] This invention is a system that efficiently collects, classifies, and summarizes customer inquiry voice data, and further recognizes customer emotions. This system operates in cooperation with a server, terminals, and users.
[0348] Data collection
[0349] When a user calls the contact center to make an inquiry, the device records the content of the inquiry as voice data. The server stores this voice data in a central database. The database also includes information such as the time of the inquiry and the caller's information.
[0350] Speech recognition and text conversion
[0351] The device converts stored audio data into text data using speech recognition technology. Specifically, it uses the Google Cloud Speech-to-Text API to analyze the audio data and convert it into text information. This converted text data is then stored in a database.
[0352] Text classification and summarization
[0353] The server classifies the converted text data into specific categories using machine learning models. For example, it uses natural language processing models such as BERT (Bidirectional Encoder Representations from Transformers) to classify the text into categories such as "delivery," "returns," and "payments." Furthermore, another machine learning model, BART (Bidirectional and Auto-Regressive Transformers), is used for summarization, extracting important parts from the converted text data and saving them in a shortened format.
[0354] emotion recognition
[0355] The server uses a natural language processing model to recognize customer emotions from text data. Specifically, it uses Google Cloud Natural Language Sentiment Analysis to determine emotions such as "joy," "anger," and "sadness" from the wording and expressions in the text. This recognized emotion data is also stored in a database.
[0356] Data storage
[0357] The server centrally stores categorized, summarized, and sentiment data in a database. This organizes query results, making them easy to search and analyze later.
[0358] Data aggregation and analysis
[0359] To aggregate data within a user-specified period, the server retrieves data from the database for that period and calculates the number of inquiries for each category. It also analyzes sentiment data to understand customer sentiment trends. This analysis provides crucial data for managers to understand customer feedback and their emotions, enabling them to make appropriate business decisions.
[0360] Specific example
[0361] For example, if a customer makes an inquiry such as, "Please tell me the delivery status of my order," the terminal first collects this audio and uses speech recognition technology to convert it into text, "Please tell me the delivery status of my order." The server categorizes this text into the "Delivery" category and generates a summary such as "Inquiry about checking delivery status." Furthermore, an emotion recognition engine recognizes that this inquiry contains the emotion of "anxiety." All of this data is stored in a database, and by aggregating the data for a specified period, managers can understand that there are many inquiries in the "Delivery" category, and that many of them contain the emotion of "anxiety."
[0362] Example of a prompt
[0363] Voice input: "Please tell me the delivery status of my item. It hasn't arrived yet and I'm worried."
[0364] Prompt: "Analyze the following text: Please tell me the delivery status of my item. I haven't received it yet and I'm worried. Based on this, categorize the inquiry, generate a summary, and analyze the sentiment."
[0365] Output: Category: "Delivery", Summary: "Checking delivery status", Emotion: "Anxiety"
[0366] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0367] Step 1:
[0368] When a user calls the contact center to make an inquiry, the device records the content of the inquiry as voice data. In this operation, the voice input is saved as an audio file. The output is the recorded voice data.
[0369] Step 2:
[0370] The server sends the recorded audio data to a database for storage. The input is the audio data, and the output is the audio data stored in the database. This data also includes information such as the time of the inquiry and the caller's information.
[0371] Step 3:
[0372] The device retrieves the stored audio data and converts it into text data using speech recognition technology. Specifically, it uses the Google Cloud Speech-to-Text API to analyze the audio data and convert it into text information. The input is audio data, and the output is text data.
[0373] Step 4:
[0374] The server receives the converted text data and uses a machine learning model to classify it into specific categories. The BERT model is used to categorize the text data into categories such as "delivery," "returns," and "payments." The input is text data, and the output is the classified categories.
[0375] Step 5:
[0376] The server uses a BART model to summarize classified text data. The input is text data, and the output is summarized text data. The server extracts the most important parts from the transformed text data and saves them in a shortened format.
[0377] Step 6:
[0378] To recognize customer emotions from text data, the server uses Google Cloud Natural Language Sentiment Analysis. The input is text data, and the output is emotion data. Emotions such as "joy," "anger," and "sadness" are determined from the wording and expressions in the text.
[0379] Step 7:
[0380] The server stores classified data, summarized data, and sentiment data in a database. The input consists of classified categories, summarized data, and sentiment data, while the output is the unified data stored in the database.
[0381] Step 8:
[0382] To aggregate data within a period specified by the user, the server retrieves data from the database for that period. The input is the specified aggregation period, and the output is the data within that period. The server calculates the number of inquiries and sentiment trends for each category.
[0383] Step 9:
[0384] Based on aggregated data and sentiment data, the results are displayed through a user interface. The input is aggregated data, and the output is a display on the user interface. This allows managers to understand customer voices and sentiments, enabling them to make appropriate business decisions.
[0385] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0386] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0387] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0388] [Second Embodiment]
[0389] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0390] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0391] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0392] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0393] The microphone 238 receives voice signals from the user 20 and accepts instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0394] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0395] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0396] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0397] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0398] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0399] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0400] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0401] This invention relates to a system for collecting customer inquiry voice data and efficiently classifying and summarizing that data. This system operates in cooperation with a server, terminals, and users.
[0402] Data collection
[0403] The server first collects inquiry audio data from multiple customers. For example, inquiry audio data is audio files recorded when a customer calls the call center. These audio files are stored in a database.
[0404] Speech recognition and text conversion
[0405] The device converts the collected audio data into text data using speech recognition technology. Specifically, it extracts text information from the audio data using a speech recognition library. This process converts the audio data into text data.
[0406] Text classification and summarization
[0407] The server classifies the converted text data into several predefined categories. Machine learning models are used for classification. For example, categories such as "Delivery," "Returns," and "Payment" are predefined, and the server determines which category each text belongs to. Furthermore, the server summarizes the text data. For summarization, machine learning models are used to extract the important parts of the original text and save them in a shortened format.
[0408] Data storage
[0409] The server stores the categorized and summarized data in a database. This organizes the query results, making them easy to search and analyze later.
[0410] Data aggregation and reporting
[0411] The system aggregates data within a user-specified period and calculates the number of inquiries for each category. This aggregated data is crucial for managers to understand customer feedback and make appropriate business decisions.
[0412] Specific example
[0413] For example, if a customer asks, "Please tell me the delivery status of my order," this audio is first collected and converted to text using speech recognition. The text is categorized under "Delivery," and a summary such as "Inquiry about checking delivery status" is generated. This data is stored in a database, and by aggregating the data for a specified period, managers can identify which inquiries fall under the "Delivery" category.
[0414] In this way, the present invention is a system for efficiently classifying and summarizing customer inquiries and providing information useful for business management.
[0415] The following describes the processing flow.
[0416] Step 1:
[0417] A user makes a phone call to the contact center with an inquiry. The content of the inquiry is recorded as voice data.
[0418] Step 2:
[0419] The server stores the inquiry voice data in a database. The database also includes information such as the time of the inquiry and the caller's information.
[0420] Step 3:
[0421] The device converts stored audio data into text data using speech recognition technology. Specifically, it analyzes the audio data using a speech recognition library and converts it into text information.
[0422] Step 4:
[0423] The server stores the converted text data in a database. This allows for centralized management of audio data and its corresponding text data.
[0424] Step 5:
[0425] The server uses machine learning models to categorize the stored text data. For example, the text is automatically classified into categories such as "delivery," "returns," and "payment."
[0426] Step 6:
[0427] The server summarizes each classified text data. This summarization uses machine learning models to extract key parts of the text and generate new text data in a shortened format.
[0428] Step 7:
[0429] The server stores the classified and summarized data in a new database. This ensures that each query is stored in a structured format.
[0430] Step 8:
[0431] The system aggregates data for a period specified by the user. The server retrieves data from the database for that period and calculates the number of queries for each category.
[0432] Step 9:
[0433] The server generates aggregated results in report format. The report includes the number of inquiries for each category and data showing specific trends.
[0434] Step 10:
[0435] The terminal provides the user with the generated report. The report is displayed through a web interface or dashboard.
[0436] (Example 1)
[0437] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0438] Traditional customer service systems struggled to effectively collect, classify, and summarize customer inquiry voice data. This resulted in challenges in quickly and accurately understanding inquiries and providing information to support appropriate business decisions.
[0439] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0440] In this invention, the server includes means for collecting inquiry voice data from multiple customers, means for converting the voice data into text data, and means for classifying the converted text data into a plurality of predefined categories. This makes it possible to efficiently perform a series of processes within the system, from collecting inquiry voice data to classifying, summarizing, and storing it, as well as aggregating the data.
[0441] A "customer" is a person or organization that receives various services or products, and is the person to whom inquiries about those services or products are made.
[0442] "Inquiry voice data" refers to data recorded in audio format of customer inquiries about services or products.
[0443] "Means of collecting voice data" refers to the technology and equipment used to record and save customer inquiries.
[0444] "Means of converting to text data" refers to speech recognition technologies and software used to convert audio data into text information.
[0445] A "predefined category" is a set of categories that are set up in advance to efficiently classify the content of inquiries, and may include, for example, "delivery," "returns," and "payment."
[0446] "Means of classification" refers to techniques and machine learning models used to analyze text data and classify it based on predefined categories.
[0447] "Means of summarization" refers to techniques and generative models that extract important parts from original text data and summarize the content concisely.
[0448] "Means of preservation" refers to technologies and devices for recording classified and summarized data in databases or other storage systems.
[0449] "Means for aggregating and calculating the number of inquiries for each category" refers to technologies and algorithms for analyzing data stored in a database and aggregating the number of inquiries for each category within a specified period.
[0450] A "database" refers to a system or software used to efficiently store and manage audio data, text data, classification data, summary data, and other similar information.
[0451] A "machine learning model" refers to a model that analyzes data and automatically performs classification and summarization based on that analysis. This includes statistical methods and algorithms.
[0452] A "generative model" refers to an AI model or algorithm used to generate new data or summaries from provided data.
[0453] A "prompt statement" is a phrase input into a generative model, and it refers to an instruction that the model uses to generate the desired output.
[0454] This invention is a system for collecting customer inquiry voice data and efficiently classifying and summarizing that data. This system operates through the coordinated efforts of a server, terminals, and users.
[0455] Data collection
[0456] The server first collects voice data from multiple customers. This voice data consists of audio files recorded when customers call the call center. Specifically, it is stored in a database on the server. To collect the voice data, a VoIP system is used to capture voice communications, which are then temporarily stored as audio files.
[0457] Speech recognition and text conversion
[0458] The device retrieves audio data stored on the server and converts it into text data using speech recognition technology. Examples of speech recognition libraries used include the Google Cloud Speech-to-Text API. The device sends the audio data to the API, receives the recognition result, and saves it in text format. The converted text data is then saved to the server's database.
[0459] Text classification and summarization
[0460] The server classifies the converted text data into several predefined categories using a machine learning model. For example, categories such as "Delivery," "Returns," and "Payment" are predefined, and the server automatically determines which category each text belongs to. This process uses a pre-trained classification model with scikit-learn. It also summarizes the text data using a generative AI model (e.g., GPT-3). The generated summaries are stored in a database.
[0461] Data storage
[0462] The classified text data and summary data are stored in a database by the server. During this process, they are organized into records containing category labels and summary text, and then saved.
[0463] Data aggregation and reporting
[0464] Users request data aggregation for a specified period. Users can enter the specified period through the web application interface. The server queries the database based on the specified period and aggregates the number of queries for each category. This is done using SQL queries. The aggregation results are reported to the user and displayed on the web application's dashboard.
[0465] Specific example
[0466] For example, if a customer asks, "Please tell me the delivery status of my order," the server first collects this audio. Then, the terminal uses speech recognition technology to convert the audio data into text, "Please tell me the delivery status of my order." This text data is then categorized by the server as "Delivery" and further summarized as "Inquiry about checking delivery status" using a generative AI model. This data is stored in a database, and later, the system aggregates the number of inquiries in the "Delivery" category and reports it to the user.
[0467] Example of a prompt
[0468] The following is an example of a prompt to input into a generative AI model:
[0469] "Convert the customer inquiry audio 'Please tell me the delivery status of my order' into text, and categorize that text under 'Delivery.' Furthermore, summarize this text and save it to the database."
[0470] In this way, the present invention is a system for efficiently classifying and summarizing customer inquiries and providing information useful for business management.
[0471] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0472] Step 1:
[0473] The server collects voice data from multiple customers. When a customer calls the call center, the VoIP system captures the voice data and temporarily stores it as an audio file. The input here is the customer's voice, and the output is an audio file. This audio file is stored in a database.
[0474] Step 2:
[0475] The device retrieves an audio file stored on the server. The input is the path to the audio file, and the output is the retrieved audio file itself. Specifically, the audio file is downloaded using a REST API.
[0476] Step 3:
[0477] The device converts the acquired audio file into text data using the Google Cloud Speech-to-Text API. The input is an audio file, and the output is the converted text data. The audio data is sent to the API, and the recognition results are received in text format.
[0478] Step 4:
[0479] The terminal sends the converted text data to the server. The input is text data, and the output is text data stored in the server's database. Specifically, a REST API is used to send text data to the server and store it in the database.
[0480] Step 5:
[0481] The server inputs text data stored in a database into a machine learning model and classifies it into predefined categories. The input is text data, and the output is the classification result of the text. The model used here is a classification model trained with scikit-learn. Specifically, it generates data with category labels attached.
[0482] Step 6:
[0483] The server inputs classified text data into a generating AI model (e.g., GPT-3) to generate a summary. The input is classified text data, and the output is summarized text data. The summary extracts the important parts and puts them into a shorter format.
[0484] Step 7:
[0485] The server stores the classified and summarized data in a database. The input is category labels and summary text data, and the output is the records stored in the database. This ensures that the data is systematically stored, facilitating later searching and analysis.
[0486] Step 8:
[0487] The user requests data aggregation for a specified period. The input is the specified period (e.g., a date range), and the output is the query aggregation results for each category within that period. The user submits the request via a web application.
[0488] Step 9:
[0489] The server queries the database based on a specified period and aggregates the number of queries for each category. The input is period information, and the output is the aggregated result. SQL queries are used to search the database and calculate the number of queries for each category.
[0490] Step 10:
[0491] The server reports the aggregated results to the user. The input is the aggregated results, and the output is a report displayed on the web application's dashboard. The user views this report and uses it to inform business decisions.
[0492] (Application Example 1)
[0493] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0494] Traditional customer inquiry systems had problems such as the time it took to collect, classify, and summarize voice data, making real-time data display and analysis difficult. This reduced the efficiency of customer service and made it difficult for managers to make quick business decisions.
[0495] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0496] In this invention, the server includes means for collecting voice data of inquiries from multiple customers, means for converting the voice data into text data, means for classifying the converted text data into a plurality of predefined categories, means for summarizing the classified text data, means for storing the classified data and the summarized data, means for aggregating data within a specified period and calculating the number of inquiries for each category, and means for collecting and displaying the text data in real time. This improves the efficiency of customer service and enables managers to make quick business decisions.
[0497] "Multiple customers" refers to two or more users who are the entities making inquiries.
[0498] "Inquiry voice data" refers to voice information uttered by customers, which includes questions and requests.
[0499] "Means of collection" refers to systems and devices for recording audio data in digital format and transmitting it to a server.
[0500] "Text data" refers to audio data converted into written information, which can then be processed mechanically.
[0501] "Means of conversion" refer to technologies and systems for converting audio data into text information.
[0502] "Multiple predefined categories" refers to several pre-defined classification items, such as "delivery," "returns," and "payment."
[0503] "Means of classification" refer to technologies and systems for sorting text data into predefined categories.
[0504] "Methods of summarization" refer to technologies and systems for extracting important parts from text data and summarizing them concisely.
[0505] "Means of preservation" refers to systems for storing data in databases or storage devices.
[0506] "Data within a specified period" refers to query data collected within a specific time frame.
[0507] "Means for calculating the number of inquiries per category" refers to technologies or systems for calculating the number of inquiries classified into each category.
[0508] "Means of collecting and displaying data in real time" refers to technologies and systems for instantly collecting and visually displaying inquiry data.
[0509] A "machine learning model" is an algorithm or program that learns from large amounts of data and automatically performs tasks such as data classification and summarization.
[0510] "Speech recognition technology" is a technology that converts speech data into text information.
[0511] A "database" is a system for efficiently storing, searching, and accessing large amounts of data.
[0512] A "visualization algorithm" is a technology that displays collected data in the form of graphs and charts, making it easy to understand visually.
[0513] An "algorithm" is a set of procedures or computational steps for solving a specific problem.
[0514] This invention relates to a system for collecting customer inquiry voice data and efficiently classifying and summarizing that data. The system operates in cooperation with a server, terminals, and users.
[0515] Data collection
[0516] The server collects voice data from multiple customers. This voice data includes, for example, audio files recorded when a customer calls the call center. These audio files are stored in a database.
[0517] Speech recognition and text conversion
[0518] The device converts the collected audio data into text data using speech recognition technology. Specifically, it uses the Google Speech-to-Text API. It extracts text information from the audio data and converts it into text data.
[0519] Text classification and summarization
[0520] The server classifies the converted text data into several predefined categories. This classification uses machine learning models (e.g., TensorFlow, PyTorch). For example, categories such as "Delivery," "Returns," and "Payment" are predefined, and the server determines which category each text belongs to. Furthermore, the server summarizes the text data. This summarization also uses machine learning models to extract the most important parts of the original text and saves them in a shortened format.
[0521] Real-time display
[0522] The server implements a visualization algorithm to collect and display the aforementioned text data in real time. This allows inquiry data to be retrieved immediately, enabling stakeholders to understand the situation in real time.
[0523] Data storage
[0524] The server stores categorized and summarized data in a database. This organizes the query results, making them easy to search and analyze later. Databases used include Firebase and MySQL.
[0525] Data aggregation and reporting
[0526] The system aggregates data within a period specified by the user and calculates the number of inquiries for each category. This aggregated result provides important data for managers to understand customer feedback and make appropriate business decisions.
[0527] Specific example
[0528] For example, if a customer asks, "Please tell me the delivery status of my order," the audio is first collected and converted to text using speech recognition. The text is categorized under "Delivery," and a summary such as "Inquiry about checking delivery status" is generated. This data is stored in a database, and by aggregating the data for a specified period, managers can understand which inquiries fall under the "Delivery" category. Furthermore, because the collected information is displayed in real time, it is possible to shorten the time from the moment an inquiry is made until a response is given.
[0529] Example of a prompt
[0530] Use the Google Speech-to-Text API to convert speech to text, input the text into a TensorFlow model, and categorize it into categories such as "Shipping," "Returns," and "Payment." Furthermore, extract and summarize the key points. Save the classification results and summaries to a Firebase database. Finally, aggregate the number of inquiries within a specified period and display it as a report.
[0531] This significantly improves the efficiency of customer service, allowing managers to quickly implement appropriate countermeasures.
[0532] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0533] Step 1: Collect audio data
[0534] The server collects customer inquiry voice data in real time. Specifically, it records the voice call when a customer makes an inquiry and transfers the audio file to the server. The input is the customer's voice data, and the output is an audio file in digital format.
[0535] Step 2: Converting audio data to text
[0536] The device converts the collected audio data into text data. Here, the Google Speech-to-Text API is used to convert the audio data into text information. The input is the audio data collected in step 1, and the output is the converted text data.
[0537] Step 3: Classification of text data
[0538] The server classifies the converted text data into several predefined categories. A machine learning model (e.g., TensorFlow) is used to execute an algorithm that classifies the text data. The input is the text data obtained in step 2, and the output is the text data classified by category.
[0539] Step 4: Textbook Summary
[0540] The server summarizes the classified text data. It uses a machine learning model (e.g., PyTorch) to extract and shorten the important parts. The input is the text data classified in step 3, and the output is the summarized text data.
[0541] Step 5: Save Data
[0542] The server stores the classified and summarized data in a database. Specifically, it uses a database management system such as Firebase or MySQL to store the data. The input is the data obtained in steps 3 and 4, and the output is the state in which it is stored in the database.
[0543] Step 6: Real-time display of data
[0544] The server collects query data in real time and displays it using a visualization algorithm. The input is all the data obtained from step 2 onward, and the output is a visual display format that is updated in real time.
[0545] Step 7: Data aggregation and reporting
[0546] The user aggregates inquiry data within a specified period and calculates the number of inquiries for each category. The server uses an aggregation algorithm to calculate the number of inquiries and displays it as a report. The input is all data stored in the database, and the output is a report of the aggregated results.
[0547] The above steps enable a system that efficiently classifies, summarizes, and displays customer inquiry voice data in real time.
[0548] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0549] This invention combines a system that collects customer inquiry voice data, efficiently classifies and summarizes that data, with an emotion engine that recognizes user emotions. This system operates in cooperation with a server, terminals, and users.
[0550] Data collection
[0551] When a user calls the contact center with an inquiry, the content of the inquiry is recorded as voice data. The server stores this voice data in a database. The database also includes information such as the time of the inquiry and the caller's information.
[0552] Speech recognition and text conversion
[0553] The device converts stored audio data into text data using speech recognition technology. It analyzes the audio data using a speech recognition library and converts it into text information. This converted text data is stored in a database.
[0554] Text classification and summarization
[0555] The server uses a machine learning model to classify the converted text data into specific categories. For example, it might be categorized as "delivery," "returns," or "payment." Another machine learning model is used for summarization, extracting the most important parts from the converted text data and saving them in a shortened format.
[0556] emotion recognition
[0557] A distinctive feature of this invention is the provision of an emotion engine. The server uses this emotion engine to recognize the user's emotions from text data. For example, it determines emotions such as "joy," "anger," and "sadness" from the wording and expressions in the text. This recognized emotion data is also stored in a database.
[0558] Data storage
[0559] The server centrally stores categorized, summarized, and sentiment data in a database. This organizes query results, making them easy to search and analyze later.
[0560] Data aggregation and analysis
[0561] To aggregate data within a user-specified period, the server retrieves data from the database for that period and calculates the number of inquiries for each category. It also analyzes sentiment data to understand user sentiment trends. This analysis provides crucial data for managers to understand customer feedback and their emotions, enabling them to make appropriate business decisions.
[0562] Specific example
[0563] For example, if a customer makes an inquiry such as, "Please tell me the delivery status of my order," this audio is first collected and converted into text using speech recognition. The text is then categorized under "Delivery," and a summary such as "Inquiry about checking delivery status" is generated. Furthermore, an emotion engine recognizes that this inquiry contains the emotion of "anxiety." All of this data is stored in a database, and by aggregating data for a specified period, managers can understand that there are many inquiries in the "Delivery" category, and that many of them contain the emotion of "anxiety."
[0564] Thus, this invention efficiently classifies and summarizes customer inquiries and provides information useful for business management by further analyzing sentiment data.
[0565] The following describes the processing flow.
[0566] Step 1:
[0567] A user makes a phone call to the contact center with an inquiry. The content of the inquiry is recorded as voice data.
[0568] Step 2:
[0569] The server stores the inquiry voice data in a database. The database also includes information such as the time of the inquiry and the caller's information.
[0570] Step 3:
[0571] The device converts stored audio data into text data using speech recognition technology. Specifically, it analyzes the audio data using a speech recognition library and converts it into text information.
[0572] Step 4:
[0573] The server stores the converted text data in a database. This allows for centralized management of audio data and its corresponding text data.
[0574] Step 5:
[0575] The server uses machine learning models to classify stored text data into specific categories. For example, categories such as "delivery," "returns," and "payments" are predefined, and the server determines which category each text belongs to.
[0576] Step 6:
[0577] The server summarizes each classified text data. This summarization uses machine learning models to extract key parts of the text and generate new text data in a shortened format.
[0578] Step 7:
[0579] The server stores the classified and summarized data in a database. This organizes the query content, making it easy to search and analyze later.
[0580] Step 8:
[0581] The server inputs the stored text data into the emotion engine to recognize the user's emotions. The emotion engine determines emotions such as "joy," "anger," and "sadness" from the text data.
[0582] Step 9:
[0583] The server stores the recognized sentiment data in a database. This allows for the management of sentiment data for each query.
[0584] Step 10:
[0585] The system aggregates data for a period specified by the user. The server retrieves data from the database for that period and calculates the number of queries and sentiment trends for each category.
[0586] Step 11:
[0587] The server generates aggregated results in report format. The report includes the number of inquiries for each category and corresponding sentiment data.
[0588] Step 12:
[0589] The terminal provides the user with the generated report. The report is displayed through a web interface or dashboard.
[0590] Specific example
[0591] For example, if a customer makes an inquiry such as, "Please tell me the delivery status of my order," this audio is first collected and converted into text using speech recognition. The text is then categorized under "Delivery," and a summary such as "Inquiry about checking delivery status" is generated. Furthermore, an emotion engine recognizes that this inquiry contains the emotion of "anxiety." All of this data is stored in a database, and by aggregating data for a specified period, managers can understand that there are many inquiries in the "Delivery" category, and that many of them contain the emotion of "anxiety."
[0592] (Example 2)
[0593] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0594] Conventional inquiry management systems have struggled to efficiently collect, classify, and summarize customer inquiry voice data. Furthermore, the lack of a means to accurately recognize and analyze the emotions expressed in the inquiries presented by users meant that providing appropriate responses and service improvements that considered customer feelings was challenging.
[0595] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for collecting voice data inquiries from multiple customers, means for converting the voice data into text data, means for classifying the converted text data into a plurality of predefined categories, means for summarizing the classified text data, means for recognizing the user's emotions from the text data, means for storing the classified data, summarized data, and emotion data, and means for aggregating data within a specified period and calculating the number of inquiries and emotion trends for each category. This enables efficient management of customer inquiries and detailed data analysis, including user emotions.
[0596] "Inquiry voice data" refers to voice information collected from customer phone calls, voice messages, and other means.
[0597] "Text data" refers to digital data obtained by converting audio data into written text.
[0598] A "machine learning model" is an algorithm that is trained using large amounts of data and automatically performs specific tasks (such as classification, prediction, and generation).
[0599] "Speech recognition technology" is a technology that analyzes speech data and converts it into text data.
[0600] A "category" is a predefined item or group used to classify the content of an inquiry.
[0601] A "summary" is data that extracts the important parts from long text data and expresses them in a shorter format.
[0602] An "emotion engine" is software or an algorithm that analyzes text data and identifies the emotions contained within it.
[0603] A "database" is a system for efficiently storing, retrieving, and managing data.
[0604] "Data aggregation" is the process of compiling data based on certain criteria and calculating statistics or totals.
[0605] "Emotional trends" refer to data that shows fluctuations and tendencies in users' emotions over a certain period of time.
[0606] This invention relates to a system that efficiently collects, classifies, and summarizes customer inquiry voice data, and further recognizes the user's emotions. This system operates in cooperation with a server, terminals, and users.
[0607] Data collection
[0608] When a user calls the contact center, the content of their inquiry is recorded as voice data by the server. The server stores the collected voice data, along with the time of the inquiry and the caller's information, in a database.
[0609] Specific example:
[0610] When a user asks, "Please tell me the availability of the product," the voice message is recorded by the server and stored in the database.
[0611] Speech recognition and text conversion
[0612] The audio data stored on the server is transferred to the terminal. The terminal uses a speech recognition library (e.g., Google Cloud Speech-to-Text API) to convert the audio data into text data, and then saves that text data back to the database.
[0613] Specific example:
[0614] The device converts the spoken phrase "Please tell me the product's stock status" into text data using the Google Cloud Speech-to-Text API and saves that text to a database.
[0615] Text classification and summarization
[0616] The server classifies the converted text data into specific categories using a machine learning model (e.g., BERT). It also uses another machine learning model (e.g., GPT-3) for summarization, extracting the most important parts from the text data and saving them in a shortened format.
[0617] Specific example:
[0618] The server converts the text "Please tell me the product's stock status" into a summary "Check product stock" and categorizes it under "Product Category".
[0619] emotion recognition
[0620] The server uses an emotion engine (e.g., Amazon Comprehend) to recognize the user's emotions from text data. The recognized emotion data is stored in a database.
[0621] Specific example:
[0622] The server recognizes the emotion of "anxiety" from the text "Please tell me the product's stock status" and stores the emotion data in the database.
[0623] Data storage
[0624] The server centrally stores audio data, text data, categorized and summarized data, and sentiment data in a database.
[0625] Specific example:
[0626] The database stores the audio file in the format "Audio file: Audio001.wav, Text: Please tell me the product's stock status, Category: Product category, Summary: Check product stock, Emotion: Anxiety".
[0627] Data aggregation and analysis
[0628] The user (manager) aggregates data for a specified period. The server retrieves this data from the database and calculates the number of inquiries and sentiment trends for each category. The server reports these analysis results to the manager to help them make informed business decisions.
[0629] Specific example:
[0630] The server generates a report stating that "data from 2023-10-01 to 2023-10-10 was compiled, and there were 50 inquiries regarding product categories, 40 of which included feelings of anxiety." The manager then reviews this report and uses it as a guide for business management.
[0631] Example of a prompt
[0632] "Design a system to collect, transcribe, categorize, and summarize customer inquiry audio, and to recognize user sentiment."
[0633] This system efficiently classifies and summarizes customer inquiries and analyzes sentiment data to provide valuable information for business management.
[0634] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0635] Step 1:
[0636] A user calls the contact center. During this call, the user verbally communicates their specific inquiry. For example, "Please tell me the availability of the product."
[0637] Input: User inquiry voice
[0638] Output: Raw audio data sent to the server
[0639] Step 2:
[0640] The server collects voice data from user inquiries in real time and saves this voice data as a file. It also adds information about the time of the inquiry and the caller, and saves this data to a database.
[0641] Input: User inquiry voice
[0642] Output: Audio file (e.g., audio001.wav) and metadata (time, caller information)
[0643] Specific actions:
[0644] The server saves the audio data as "Audio001.wav" and simultaneously records it in the database in the format "2023-10-10 10:00:00, Customer A, Audio001.wav".
[0645] Step 3:
[0646] The audio data stored on the server is transferred to the terminal. The terminal then converts this audio data into text data using a speech recognition library (e.g., Google Cloud Speech-to-Text API).
[0647] Input: Audio file (Audio001.wav)
[0648] Output: Text data (Example: "Please tell me the product's stock status")
[0649] Specific actions:
[0650] The device receives the audio file "Audio001.wav" and calls the Google Cloud Speech-to-Text API to convert it to text. As a result, the text data "Please tell me the product's stock status" is generated.
[0651] Step 4:
[0652] The device saves the converted text data to a database. This ensures that the text data is permanently recorded and searchable later.
[0653] Input: Text data (Example: "Please tell me the stock status of the product.")
[0654] Output: Text data is saved to the database.
[0655] Specific actions:
[0656] The database will store the text, "Please tell me the product's stock status."
[0657] Step 5:
[0658] The server uses a machine learning model (e.g., BERT) to classify the converted text data into specific categories, such as "delivery," "returns," and "payment."
[0659] Input: Text data (Example: "Please tell me the stock status of the product.")
[0660] Output: Classified category (e.g., "Product Category")
[0661] Specific actions:
[0662] The server adds a label called "product category" to the database.
[0663] Step 6:
[0664] The server uses a summarization model (e.g., GPT-3) to summarize the text data. This summarized data is also stored in the database.
[0665] Input: Text data (Example: "Please tell me the stock status of the product.")
[0666] Output: Summary data (e.g., "Check product inventory")
[0667] Specific actions:
[0668] The server uses a summarization model to generate a summary such as "Check product inventory" and saves it to the database.
[0669] Step 7:
[0670] The server uses an emotion engine (e.g., Amazon Comprehend) to recognize the user's emotions from text data. This emotion data is also stored in a database.
[0671] Input: Text data (Example: "Please tell me the stock status of the product.")
[0672] Output: Emotional data (e.g., "anxiety")
[0673] Specific actions:
[0674] The server uses Amazon Comprehend to recognize the emotion "anxiety" from the text "Please tell me the availability of the product" and saves it to the database.
[0675] Step 8:
[0676] The server centrally stores audio data, text data, categorized data, summary data, and sentiment data in a database. This allows for easy searching and analysis later.
[0677] Input: Audio data, text data, classification data, summary data, sentiment data
[0678] Output: Centrally organized query data
[0679] Specific actions:
[0680] The database will store the audio file in the format: "Audio file: Audio001.wav, Text: Please tell me the product's stock status, Category: Product category, Summary: Check product stock, Emotion: Anxiety".
[0681] Step 9:
[0682] The user (manager) has the server aggregate data for a specified period. The server retrieves query data from the database for that period and calculates the number of queries and sentiment trends for each category.
[0683] Input: Aggregation instructions (Example: "From 2023-10-01 to 2023-10-10")
[0684] Output: Aggregated report (e.g., number of inquiries for product categories, anxiety trend)
[0685] Specific actions:
[0686] The server retrieves data from the database for a specified period and generates a report stating that "there were 50 inquiries for the product category, 40 of which included feelings of anxiety."
[0687] Step 10:
[0688] The server provides the generated report to the user (manager). The manager then uses the report to make business decisions.
[0689] Input: Summary Report
[0690] Output: Information for business decision-making
[0691] Specific actions:
[0692] The manager reviews the provided report and understands that there are many inquiries in the "delivery" category, many of which involve feelings of "anxiety," and then considers measures to improve the situation.
[0693] (Application Example 2)
[0694] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0695] Traditional customer inquiry systems often require manual classification and summarization of inquiries, which is time-consuming and labor-intensive. Furthermore, analyzing customer sentiment is difficult, making it challenging to grasp customer satisfaction in real time. Therefore, while rapid responses based on inquiry categories and content are required, comprehensive responses, including sentiment analysis, are currently lacking. In addition, there is a lack of efficient methods for managers to derive business guidance from past inquiry data.
[0696] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0697] In this invention, the server includes means for collecting voice data of inquiries from multiple customers, means for converting the voice data into text data, means for classifying the converted text data into a plurality of predefined categories, means for summarizing the classified text data, means for recognizing customer emotions from the text data, means for storing the classified data, summarized data, and emotion data, and means for aggregating data within a specified period and calculating the number of inquiries and emotion trends for each category. This not only enables efficient classification and summarization of inquiry content, but also allows for real-time recognition of customer emotions and prompt and appropriate responses based on that information. Furthermore, it facilitates management decisions based on the stored data, leading to improved customer satisfaction and optimized operations.
[0698] "Customer" refers to the person who receives a product or service.
[0699] "Inquiry voice data" refers to voice data of questions and requests provided by customers via telephone or voice recording.
[0700] "Text data" refers to the character information obtained by converting inquiry voice data using speech recognition technology.
[0701] A "category" is a predefined group or type used to make text data easier to classify. Examples of categories include "Shipping," "Returns," and "Payment."
[0702] "Summary" means extracting the key points from text data and providing them in a shortened format.
[0703] "Emotion recognition technology" refers to techniques for recognizing a customer's psychological state and emotions from text data. This technology determines emotions such as "joy," "anger," and "sadness."
[0704] A "machine learning model" is an algorithm that uses large amounts of data to learn patterns and perform classification and prediction. It is used for text classification and summarization.
[0705] "Speech recognition technology" refers to the technology that converts speech data into text information. This is a part of speech processing technology.
[0706] A "database" refers to a system for systematically storing large amounts of data and for efficiently processing and retrieving it.
[0707] A "user interface" is an interface that allows a system and a user to interact directly. It facilitates the display and manipulation of data.
[0708] This invention is a system that efficiently collects, classifies, and summarizes customer inquiry voice data, and further recognizes customer emotions. This system operates in cooperation with a server, terminals, and users.
[0709] Data collection
[0710] When a user calls the contact center to make an inquiry, the device records the content of the inquiry as voice data. The server stores this voice data in a central database. The database also includes information such as the time of the inquiry and the caller's information.
[0711] Speech recognition and text conversion
[0712] The device converts stored audio data into text data using speech recognition technology. Specifically, it uses the Google Cloud Speech-to-Text API to analyze the audio data and convert it into text information. This converted text data is then stored in a database.
[0713] Text classification and summarization
[0714] The server classifies the converted text data into specific categories using machine learning models. For example, it uses natural language processing models such as BERT (Bidirectional Encoder Representations from Transformers) to classify the text into categories such as "delivery," "returns," and "payments." Furthermore, another machine learning model, BART (Bidirectional and Auto-Regressive Transformers), is used for summarization, extracting important parts from the converted text data and saving them in a shortened format.
[0715] emotion recognition
[0716] The server uses a natural language processing model to recognize customer emotions from text data. Specifically, it uses Google Cloud Natural Language Sentiment Analysis to determine emotions such as "joy," "anger," and "sadness" from the wording and expressions in the text. This recognized emotion data is also stored in a database.
[0717] Data storage
[0718] The server centrally stores categorized, summarized, and sentiment data in a database. This organizes query results, making them easy to search and analyze later.
[0719] Data aggregation and analysis
[0720] To aggregate data within a user-specified period, the server retrieves data from the database for that period and calculates the number of inquiries for each category. It also analyzes sentiment data to understand customer sentiment trends. This analysis provides crucial data for managers to understand customer feedback and their emotions, enabling them to make appropriate business decisions.
[0721] Specific example
[0722] For example, if a customer makes an inquiry such as, "Please tell me the delivery status of my order," the terminal first collects this audio and uses speech recognition technology to convert it into text, "Please tell me the delivery status of my order." The server categorizes this text into the "Delivery" category and generates a summary such as "Inquiry about checking delivery status." Furthermore, an emotion recognition engine recognizes that this inquiry contains the emotion of "anxiety." All of this data is stored in a database, and by aggregating the data for a specified period, managers can understand that there are many inquiries in the "Delivery" category, and that many of them contain the emotion of "anxiety."
[0723] Example of a prompt
[0724] Voice input: "Please tell me the delivery status of my item. It hasn't arrived yet and I'm worried."
[0725] Prompt: "Analyze the following text: Please tell me the delivery status of my item. I haven't received it yet and I'm worried. Based on this, categorize the inquiry, generate a summary, and analyze the sentiment."
[0726] Output: Category: "Delivery", Summary: "Checking delivery status", Emotion: "Anxiety"
[0727] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0728] Step 1:
[0729] When a user calls the contact center to make an inquiry, the device records the content of the inquiry as voice data. In this operation, the voice input is saved as an audio file. The output is the recorded voice data.
[0730] Step 2:
[0731] The server sends the recorded audio data to a database for storage. The input is the audio data, and the output is the audio data stored in the database. This data also includes information such as the time of the inquiry and the caller.
[0732] Step 3:
[0733] The device retrieves the stored audio data and converts it into text data using speech recognition technology. Specifically, it uses the Google Cloud Speech-to-Text API to analyze the audio data and convert it into text information. The input is audio data, and the output is text data.
[0734] Step 4:
[0735] The server receives the converted text data and uses a machine learning model to classify it into specific categories. The BERT model is used to categorize the text data into categories such as "delivery," "returns," and "payments." The input is text data, and the output is the classified categories.
[0736] Step 5:
[0737] The server uses a BART model to summarize classified text data. The input is text data, and the output is summarized text data. The server extracts the most important parts from the transformed text data and saves them in a shortened format.
[0738] Step 6:
[0739] To recognize customer emotions from text data, the server uses Google Cloud Natural Language Sentiment Analysis. The input is text data, and the output is emotion data. Emotions such as "joy," "anger," and "sadness" are determined from the wording and expressions in the text.
[0740] Step 7:
[0741] The server stores classified data, summarized data, and sentiment data in a database. The input consists of classified categories, summarized data, and sentiment data, while the output is the unified data stored in the database.
[0742] Step 8:
[0743] To aggregate data within a period specified by the user, the server retrieves data from the database for that period. The input is the specified aggregation period, and the output is the data within that period. The server calculates the number of inquiries and sentiment trends for each category.
[0744] Step 9:
[0745] Based on aggregated data and sentiment data, the results are displayed through a user interface. The input is aggregated data, and the output is a display on the user interface. This allows managers to understand customer voices and sentiments, enabling them to make appropriate business decisions.
[0746] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0747] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0748] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0749] [Third Embodiment]
[0750] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0751] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0752] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0753] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0754] The microphone 238 receives voice signals from the user 20 and accepts instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0755] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0756] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0757] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0758] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0759] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0760] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0761] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0762] This invention relates to a system for collecting customer inquiry voice data and efficiently classifying and summarizing that data. This system operates in cooperation with a server, terminals, and users.
[0763] Data collection
[0764] The server first collects inquiry audio data from multiple customers. For example, inquiry audio data is audio files recorded when a customer calls the call center. These audio files are stored in a database.
[0765] Speech recognition and text conversion
[0766] The device converts the collected audio data into text data using speech recognition technology. Specifically, it extracts text information from the audio data using a speech recognition library. This process converts the audio data into text data.
[0767] Text classification and summarization
[0768] The server classifies the converted text data into several predefined categories. Machine learning models are used for classification. For example, categories such as "Delivery," "Returns," and "Payment" are predefined, and the server determines which category each text belongs to. Furthermore, the server summarizes the text data. For summarization, machine learning models are used to extract the important parts of the original text and save them in a shortened format.
[0769] Data storage
[0770] The server stores the categorized and summarized data in a database. This organizes the query results, making them easy to search and analyze later.
[0771] Data aggregation and reporting
[0772] The system aggregates data within a user-specified period and calculates the number of inquiries for each category. This aggregated data is crucial for managers to understand customer feedback and make appropriate business decisions.
[0773] Specific example
[0774] For example, if a customer asks, "Please tell me the delivery status of my order," this audio is first collected and converted to text using speech recognition. The text is categorized under "Delivery," and a summary such as "Inquiry about checking delivery status" is generated. This data is stored in a database, and by aggregating the data for a specified period, managers can identify which inquiries fall under the "Delivery" category.
[0775] In this way, the present invention is a system for efficiently classifying and summarizing customer inquiries and providing information useful for business management.
[0776] The following describes the processing flow.
[0777] Step 1:
[0778] A user makes a phone call to the contact center with an inquiry. The content of the inquiry is recorded as voice data.
[0779] Step 2:
[0780] The server stores the inquiry voice data in a database. The database also includes information such as the time of the inquiry and the caller's information.
[0781] Step 3:
[0782] The device converts stored audio data into text data using speech recognition technology. Specifically, it analyzes the audio data using a speech recognition library and converts it into text information.
[0783] Step 4:
[0784] The server stores the converted text data in a database. This allows for centralized management of audio data and its corresponding text data.
[0785] Step 5:
[0786] The server uses machine learning models to categorize the stored text data. For example, the text is automatically classified into categories such as "delivery," "returns," and "payment."
[0787] Step 6:
[0788] The server summarizes each classified text data. This summarization uses machine learning models to extract key parts of the text and generate new text data in a shortened format.
[0789] Step 7:
[0790] The server stores the classified and summarized data in a new database. This ensures that each query is stored in a structured format.
[0791] Step 8:
[0792] The system aggregates data for a period specified by the user. The server retrieves data from the database for that period and calculates the number of queries for each category.
[0793] Step 9:
[0794] The server generates aggregated results in report format. The report includes the number of inquiries for each category and data showing specific trends.
[0795] Step 10:
[0796] The terminal provides the user with the generated report. The report is displayed through a web interface or dashboard.
[0797] (Example 1)
[0798] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0799] Traditional customer service systems struggled to effectively collect, classify, and summarize customer inquiry voice data. This resulted in challenges in quickly and accurately understanding inquiries and providing information to support appropriate business decisions.
[0800] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0801] In this invention, the server includes means for collecting inquiry voice data from multiple customers, means for converting the voice data into text data, and means for classifying the converted text data into a plurality of predefined categories. This makes it possible to efficiently perform a series of processes within the system, from collecting inquiry voice data to classifying, summarizing, and storing it, as well as aggregating the data.
[0802] A "customer" is a person or organization that receives various services or products, and is the person to whom inquiries about those services or products are made.
[0803] "Inquiry voice data" refers to data recorded in audio format of customer inquiries about services or products.
[0804] "Means of collecting voice data" refers to the technology and equipment used to record and save customer inquiries.
[0805] "Means of converting to text data" refers to speech recognition technologies and software used to convert audio data into text information.
[0806] A "predefined category" is a set of categories that are set up in advance to efficiently classify the content of inquiries, and may include, for example, "delivery," "returns," and "payment."
[0807] "Means of classification" refers to techniques and machine learning models used to analyze text data and classify it based on predefined categories.
[0808] "Means of summarization" refers to techniques and generative models that extract important parts from original text data and summarize the content concisely.
[0809] "Means of preservation" refers to technologies and devices for recording classified and summarized data in databases or other storage systems.
[0810] "Means for aggregating and calculating the number of inquiries for each category" refers to technologies and algorithms for analyzing data stored in a database and aggregating the number of inquiries for each category within a specified period.
[0811] A "database" refers to a system or software used to efficiently store and manage audio data, text data, classification data, summary data, and other similar information.
[0812] A "machine learning model" refers to a model that analyzes data and automatically performs classification and summarization based on that analysis. This includes statistical methods and algorithms.
[0813] A "generative model" refers to an AI model or algorithm used to generate new data or summaries from provided data.
[0814] A "prompt statement" is a phrase input into a generative model, and it refers to an instruction that the model uses to generate the desired output.
[0815] This invention is a system for collecting customer inquiry voice data and efficiently classifying and summarizing that data. This system operates through the coordinated efforts of a server, terminals, and users.
[0816] Data collection
[0817] The server first collects voice data from multiple customers. This voice data consists of audio files recorded when customers call the call center. Specifically, it is stored in a database on the server. To collect the voice data, a VoIP system is used to capture voice communications, which are then temporarily stored as audio files.
[0818] Speech recognition and text conversion
[0819] The device retrieves audio data stored on the server and converts it into text data using speech recognition technology. Examples of speech recognition libraries used include the Google Cloud Speech-to-Text API. The device sends the audio data to the API, receives the recognition result, and saves it in text format. The converted text data is then saved to the server's database.
[0820] Text classification and summarization
[0821] The server classifies the converted text data into several predefined categories using a machine learning model. For example, categories such as "Delivery," "Returns," and "Payment" are predefined, and the server automatically determines which category each text belongs to. This process uses a pre-trained classification model with scikit-learn. It also summarizes the text data using a generative AI model (e.g., GPT-3). The generated summaries are stored in a database.
[0822] Data storage
[0823] The classified text data and summary data are stored in a database by the server. During this process, they are organized into records containing category labels and summary text, and then saved.
[0824] Data aggregation and reporting
[0825] Users request data aggregation for a specified period. Users can enter the specified period through the web application interface. The server queries the database based on the specified period and aggregates the number of queries for each category. This is done using SQL queries. The aggregation results are reported to the user and displayed on the web application's dashboard.
[0826] Specific example
[0827] For example, if a customer asks, "Please tell me the delivery status of my order," the server first collects this audio. Then, the terminal uses speech recognition technology to convert the audio data into text, "Please tell me the delivery status of my order." This text data is then categorized by the server as "Delivery" and further summarized as "Inquiry about checking delivery status" using a generative AI model. This data is stored in a database, and later, the system aggregates the number of inquiries in the "Delivery" category and reports it to the user.
[0828] Example of a prompt
[0829] The following is an example of a prompt to input into a generative AI model:
[0830] "Convert the customer inquiry audio 'Please tell me the delivery status of my order' into text, and categorize that text under 'Delivery.' Furthermore, summarize this text and save it to the database."
[0831] In this way, the present invention is a system for efficiently classifying and summarizing customer inquiries and providing information useful for business management.
[0832] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0833] Step 1:
[0834] The server collects voice data from multiple customers. When a customer calls the call center, the VoIP system captures the voice data and temporarily stores it as an audio file. The input here is the customer's voice, and the output is an audio file. This audio file is stored in a database.
[0835] Step 2:
[0836] The device retrieves an audio file stored on the server. The input is the path to the audio file, and the output is the retrieved audio file itself. Specifically, the audio file is downloaded using a REST API.
[0837] Step 3:
[0838] The device converts the acquired audio file into text data using the Google Cloud Speech-to-Text API. The input is an audio file, and the output is the converted text data. The audio data is sent to the API, and the recognition results are received in text format.
[0839] Step 4:
[0840] The terminal sends the converted text data to the server. The input is text data, and the output is text data stored in the server's database. Specifically, a REST API is used to send text data to the server and store it in the database.
[0841] Step 5:
[0842] The server inputs text data stored in a database into a machine learning model and classifies it into predefined categories. The input is text data, and the output is the classification result of the text. The model used here is a classification model trained with scikit-learn. Specifically, it generates data with category labels attached.
[0843] Step 6:
[0844] The server inputs classified text data into a generating AI model (e.g., GPT-3) to generate a summary. The input is classified text data, and the output is summarized text data. The summary extracts the important parts and puts them into a shorter format.
[0845] Step 7:
[0846] The server stores the classified and summarized data in a database. The input is category labels and summary text data, and the output is the records stored in the database. This ensures that the data is systematically stored, facilitating later searching and analysis.
[0847] Step 8:
[0848] The user requests data aggregation for a specified period. The input is the specified period (e.g., a date range), and the output is the query aggregation results for each category within that period. The user submits the request via a web application.
[0849] Step 9:
[0850] The server queries the database based on a specified period and aggregates the number of queries for each category. The input is period information, and the output is the aggregated result. SQL queries are used to search the database and calculate the number of queries for each category.
[0851] Step 10:
[0852] The server reports the aggregated results to the user. The input is the aggregated results, and the output is a report displayed on the web application's dashboard. The user views this report and uses it to inform business decisions.
[0853] (Application Example 1)
[0854] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0855] Traditional customer inquiry systems had problems such as the time it took to collect, classify, and summarize voice data, making real-time data display and analysis difficult. This reduced the efficiency of customer service and made it difficult for managers to make quick business decisions.
[0856] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0857] In this invention, the server includes means for collecting voice data of inquiries from multiple customers, means for converting the voice data into text data, means for classifying the converted text data into a plurality of predefined categories, means for summarizing the classified text data, means for storing the classified data and the summarized data, means for aggregating data within a specified period and calculating the number of inquiries for each category, and means for collecting and displaying the text data in real time. This improves the efficiency of customer service and enables managers to make quick business decisions.
[0858] "Multiple customers" refers to two or more users who are the entities making inquiries.
[0859] "Inquiry voice data" refers to voice information uttered by customers, which includes questions and requests.
[0860] "Means of collection" refers to systems and devices for recording audio data in digital format and transmitting it to a server.
[0861] "Text data" refers to audio data converted into written information, which can then be processed mechanically.
[0862] "Means of conversion" refer to technologies and systems for converting audio data into text information.
[0863] "Multiple predefined categories" refers to several pre-defined classification items, such as "delivery," "returns," and "payment."
[0864] "Means of classification" refer to technologies and systems for sorting text data into predefined categories.
[0865] "Methods of summarization" refer to technologies and systems for extracting important parts from text data and summarizing them concisely.
[0866] "Means of preservation" refers to systems for storing data in databases or storage devices.
[0867] "Data within a specified period" refers to query data collected within a specific time frame.
[0868] "Means for calculating the number of inquiries per category" refers to technologies or systems for calculating the number of inquiries classified into each category.
[0869] "Means of collecting and displaying data in real time" refers to technologies and systems for instantly collecting and visually displaying inquiry data.
[0870] A "machine learning model" is an algorithm or program that learns from large amounts of data and automatically performs tasks such as data classification and summarization.
[0871] "Speech recognition technology" is a technology that converts speech data into text information.
[0872] A "database" is a system for efficiently storing, searching, and accessing large amounts of data.
[0873] A "visualization algorithm" is a technology that displays collected data in the form of graphs and charts, making it easy to understand visually.
[0874] An "algorithm" is a set of procedures or computational steps for solving a specific problem.
[0875] This invention relates to a system for collecting customer inquiry voice data and efficiently classifying and summarizing that data. The system operates in cooperation with a server, terminals, and users.
[0876] Data collection
[0877] The server collects voice data from multiple customers. This voice data includes, for example, audio files recorded when a customer calls the call center. These audio files are stored in a database.
[0878] Speech recognition and text conversion
[0879] The device converts the collected audio data into text data using speech recognition technology. Specifically, it uses the Google Speech-to-Text API. It extracts text information from the audio data and converts it into text data.
[0880] Text classification and summarization
[0881] The server classifies the converted text data into several predefined categories. This classification uses machine learning models (e.g., TensorFlow, PyTorch). For example, categories such as "Delivery," "Returns," and "Payment" are predefined, and the server determines which category each text belongs to. Furthermore, the server summarizes the text data. This summarization also uses machine learning models to extract the most important parts of the original text and saves them in a shortened format.
[0882] Real-time display
[0883] The server implements a visualization algorithm to collect and display the aforementioned text data in real time. This allows inquiry data to be immediately retrieved, enabling stakeholders to understand the situation in real time.
[0884] Data storage
[0885] The server stores categorized and summarized data in a database. This organizes the query results, making them easy to search and analyze later. Databases used include Firebase and MySQL.
[0886] Data aggregation and reporting
[0887] The system aggregates data within a period specified by the user and calculates the number of inquiries for each category. This aggregated result provides important data for managers to understand customer feedback and make appropriate business decisions.
[0888] Specific example
[0889] For example, if a customer asks, "Please tell me the delivery status of my order," the audio is first collected and converted to text using speech recognition. The text is categorized under "Delivery," and a summary such as "Inquiry about checking delivery status" is generated. This data is stored in a database, and by aggregating the data for a specified period, managers can understand which inquiries fall under the "Delivery" category. Furthermore, because the collected information is displayed in real time, it is possible to shorten the time from the moment an inquiry is made until a response is given.
[0890] Example of a prompt
[0891] Use the Google Speech-to-Text API to convert speech to text, input the text into a TensorFlow model, and categorize it into categories such as "Shipping," "Returns," and "Payment." Furthermore, extract and summarize the key points. Save the classification results and summaries to a Firebase database. Finally, aggregate the number of inquiries within a specified period and display it as a report.
[0892] This significantly improves the efficiency of customer service, allowing managers to quickly implement appropriate countermeasures.
[0893] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0894] Step 1: Collect audio data
[0895] The server collects customer inquiry voice data in real time. Specifically, it records the voice call when a customer makes an inquiry and transfers the audio file to the server. The input is the customer's voice data, and the output is an audio file in digital format.
[0896] Step 2: Converting audio data to text
[0897] The device converts the collected audio data into text data. Here, the Google Speech-to-Text API is used to convert the audio data into text information. The input is the audio data collected in step 1, and the output is the converted text data.
[0898] Step 3: Classification of text data
[0899] The server classifies the converted text data into several predefined categories. A machine learning model (e.g., TensorFlow) is used to execute an algorithm that classifies the text data. The input is the text data obtained in step 2, and the output is the text data classified by category.
[0900] Step 4: Textbook Summary
[0901] The server summarizes the classified text data. It uses a machine learning model (e.g., PyTorch) to extract and shorten the important parts. The input is the text data classified in step 3, and the output is the summarized text data.
[0902] Step 5: Save Data
[0903] The server stores the classified and summarized data in a database. Specifically, it uses a database management system such as Firebase or MySQL to store the data. The input is the data obtained in steps 3 and 4, and the output is the state in which it is stored in the database.
[0904] Step 6: Real-time display of data
[0905] The server collects query data in real time and displays it using a visualization algorithm. The input is all the data obtained from step 2 onward, and the output is a visual display format that is updated in real time.
[0906] Step 7: Data aggregation and reporting
[0907] The user aggregates inquiry data within a specified period and calculates the number of inquiries for each category. The server uses an aggregation algorithm to calculate the number of inquiries and displays it as a report. The input is all data stored in the database, and the output is a report of the aggregated results.
[0908] The above steps enable a system that efficiently classifies, summarizes, and displays customer inquiry voice data in real time.
[0909] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0910] This invention combines a system that collects customer inquiry voice data, efficiently classifies and summarizes that data, with an emotion engine that recognizes user emotions. This system operates in cooperation with a server, terminals, and users.
[0911] Data collection
[0912] When a user calls the contact center with an inquiry, the content of the inquiry is recorded as voice data. The server stores this voice data in a database. The database also includes information such as the time of the inquiry and the caller's information.
[0913] Speech recognition and text conversion
[0914] The device converts stored audio data into text data using speech recognition technology. It analyzes the audio data using a speech recognition library and converts it into text information. This converted text data is stored in a database.
[0915] Text classification and summarization
[0916] The server uses a machine learning model to classify the converted text data into specific categories. For example, it might be categorized as "delivery," "returns," or "payment." Another machine learning model is used for summarization, extracting the most important parts from the converted text data and saving them in a shortened format.
[0917] emotion recognition
[0918] A distinctive feature of this invention is the provision of an emotion engine. The server uses this emotion engine to recognize the user's emotions from text data. For example, it determines emotions such as "joy," "anger," and "sadness" from the wording and expressions in the text. This recognized emotion data is also stored in a database.
[0919] Data storage
[0920] The server centrally stores categorized, summarized, and sentiment data in a database. This organizes query results, making them easy to search and analyze later.
[0921] Data aggregation and analysis
[0922] To aggregate data within a user-specified period, the server retrieves data from the database for that period and calculates the number of inquiries for each category. It also analyzes sentiment data to understand user sentiment trends. This analysis provides crucial data for managers to understand customer feedback and their emotions, enabling them to make appropriate business decisions.
[0923] Specific example
[0924] For example, if a customer makes an inquiry such as, "Please tell me the delivery status of my order," this audio is first collected and converted into text using speech recognition. The text is then categorized under "Delivery," and a summary such as "Inquiry about checking delivery status" is generated. Furthermore, an emotion engine recognizes that this inquiry contains the emotion of "anxiety." All of this data is stored in a database, and by aggregating data for a specified period, managers can understand that there are many inquiries in the "Delivery" category, and that many of them contain the emotion of "anxiety."
[0925] Thus, this invention efficiently classifies and summarizes customer inquiries and provides information useful for business management by further analyzing sentiment data.
[0926] The following describes the processing flow.
[0927] Step 1:
[0928] A user makes a phone call to the contact center with an inquiry. The content of the inquiry is recorded as voice data.
[0929] Step 2:
[0930] The server stores the inquiry voice data in a database. The database also includes information such as the time of the inquiry and the caller's information.
[0931] Step 3:
[0932] The device converts stored audio data into text data using speech recognition technology. Specifically, it analyzes the audio data using a speech recognition library and converts it into text information.
[0933] Step 4:
[0934] The server stores the converted text data in a database. This allows for centralized management of audio data and its corresponding text data.
[0935] Step 5:
[0936] The server uses machine learning models to classify stored text data into specific categories. For example, categories such as "delivery," "returns," and "payments" are predefined, and the server determines which category each text belongs to.
[0937] Step 6:
[0938] The server summarizes each classified text data. This summarization uses machine learning models to extract key parts of the text and generate new text data in a shortened format.
[0939] Step 7:
[0940] The server stores the classified and summarized data in a database. This organizes the query content, making it easy to search and analyze later.
[0941] Step 8:
[0942] The server inputs the stored text data into the emotion engine to recognize the user's emotions. The emotion engine determines emotions such as "joy," "anger," and "sadness" from the text data.
[0943] Step 9:
[0944] The server stores the recognized sentiment data in a database. This allows for the management of sentiment data for each query.
[0945] Step 10:
[0946] The system aggregates data for a period specified by the user. The server retrieves data from the database for that period and calculates the number of queries and sentiment trends for each category.
[0947] Step 11:
[0948] The server generates aggregated results in report format. The report includes the number of inquiries for each category and corresponding sentiment data.
[0949] Step 12:
[0950] The terminal provides the user with the generated report. The report is displayed through a web interface or dashboard.
[0951] Specific example
[0952] For example, if a customer makes an inquiry such as, "Please tell me the delivery status of my order," this audio is first collected and converted into text using speech recognition. The text is then categorized under "Delivery," and a summary such as "Inquiry about checking delivery status" is generated. Furthermore, an emotion engine recognizes that this inquiry contains the emotion of "anxiety." All of this data is stored in a database, and by aggregating data for a specified period, managers can understand that there are many inquiries in the "Delivery" category, and that many of them contain the emotion of "anxiety."
[0953] (Example 2)
[0954] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0955] Conventional inquiry management systems struggled to efficiently collect, classify, and summarize customer inquiry voice data. Furthermore, the lack of a means to accurately recognize and analyze the emotions expressed in the inquiries presented posed challenges in providing appropriate responses and service improvements that considered customer feelings.
[0956] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for collecting voice data inquiries from multiple customers, means for converting the voice data into text data, means for classifying the converted text data into a plurality of predefined categories, means for summarizing the classified text data, means for recognizing the user's emotions from the text data, means for storing the classified data, summarized data, and emotion data, and means for aggregating data within a specified period and calculating the number of inquiries and emotion trends for each category. This enables efficient management of customer inquiries and detailed data analysis, including user emotions.
[0957] "Inquiry voice data" refers to voice information collected from customer phone calls, voice messages, and other means.
[0958] "Text data" refers to digital data obtained by converting audio data into written text.
[0959] A "machine learning model" is an algorithm that is trained using large amounts of data and automatically performs specific tasks (such as classification, prediction, and generation).
[0960] "Speech recognition technology" is a technology that analyzes speech data and converts it into text data.
[0961] A "category" is a predefined item or group used to classify the content of an inquiry.
[0962] A "summary" is data that extracts the important parts from long text data and expresses them in a shorter format.
[0963] An "emotion engine" is software or an algorithm that analyzes text data and identifies the emotions contained within it.
[0964] A "database" is a system for efficiently storing, retrieving, and managing data.
[0965] "Data aggregation" is the process of compiling data based on certain criteria and calculating statistics or totals.
[0966] "Emotional trends" refer to data that shows fluctuations and tendencies in users' emotions over a certain period of time.
[0967] This invention is a system that efficiently collects, classifies, and summarizes customer inquiry voice data, and further recognizes the user's emotions. This system operates in cooperation with a server, terminals, and users.
[0968] Data collection
[0969] When a user calls the contact center, the content of their inquiry is recorded as voice data by the server. The server stores the collected voice data, along with the time of the inquiry and the caller's information, in a database.
[0970] Specific example:
[0971] When a user asks, "Please tell me the availability of the product," the voice message is recorded by the server and stored in the database.
[0972] Speech recognition and text conversion
[0973] The audio data stored on the server is transferred to the terminal. The terminal uses a speech recognition library (e.g., Google Cloud Speech-to-Text API) to convert the audio data into text data, and then saves that text data back to the database.
[0974] Specific example:
[0975] The device converts the spoken phrase "Please tell me the product's stock status" into text data using the Google Cloud Speech-to-Text API and saves that text to a database.
[0976] Text classification and summarization
[0977] The server classifies the converted text data into specific categories using a machine learning model (e.g., BERT). It also uses another machine learning model (e.g., GPT-3) for summarization, extracting important parts from the text data and saving them in a shortened format.
[0978] Specific example:
[0979] The server converts the text "Please tell me the product's stock status" into a summary "Check product stock" and categorizes it under "Product Category".
[0980] emotion recognition
[0981] The server uses an emotion engine (e.g., Amazon Comprehend) to recognize the user's emotions from text data. The recognized emotion data is stored in a database.
[0982] Specific example:
[0983] The server recognizes the emotion of "anxiety" from the text "Please tell me the product's stock status" and stores the emotion data in the database.
[0984] Data storage
[0985] The server centrally stores audio data, text data, categorized and summarized data, and sentiment data in a database.
[0986] Specific example:
[0987] The database stores the audio file in the format "Audio file: Audio001.wav, Text: Please tell me the product's stock status, Category: Product category, Summary: Check product stock, Emotion: Anxiety".
[0988] Data aggregation and analysis
[0989] The user (manager) aggregates data for a specified period. The server retrieves this data from the database and calculates the number of inquiries and sentiment trends for each category. The server reports these analysis results to the manager to help them make informed business decisions.
[0990] Specific example:
[0991] The server generates a report stating that "data from 2023-10-01 to 2023-10-10 was compiled, and there were 50 inquiries regarding product categories, 40 of which included feelings of anxiety." The manager then reviews this report and uses it as a guide for business management.
[0992] Example of a prompt
[0993] "Design a system to collect, transcribe, categorize, and summarize customer inquiry audio, and to recognize user sentiment."
[0994] This system efficiently classifies and summarizes customer inquiries and analyzes sentiment data to provide valuable information for business management.
[0995] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0996] Step 1:
[0997] A user calls the contact center. During this call, the user verbally communicates their specific inquiry. For example, "Please tell me the availability of the product."
[0998] Input: User inquiry voice
[0999] Output: Raw audio data sent to the server
[1000] Step 2:
[1001] The server collects voice data from user inquiries in real time and saves this voice data as a file. It also adds information about the time of the inquiry and the caller, and saves this data to a database.
[1002] Input: User inquiry voice
[1003] Output: Audio file (e.g., audio001.wav) and metadata (time, caller information)
[1004] Specific actions:
[1005] The server saves the audio data as "Audio001.wav" and simultaneously records it in the database in the format "2023-10-10 10:00:00, Customer A, Audio001.wav".
[1006] Step 3:
[1007] The audio data stored on the server is transferred to the terminal. The terminal then converts this audio data into text data using a speech recognition library (e.g., Google Cloud Speech-to-Text API).
[1008] Input: Audio file (Audio001.wav)
[1009] Output: Text data (Example: "Please tell me the product's stock status")
[1010] Specific actions:
[1011] The device receives the audio file "Audio001.wav" and calls the Google Cloud Speech-to-Text API to convert it to text. As a result, the text data "Please tell me the product's stock status" is generated.
[1012] Step 4:
[1013] The device saves the converted text data to a database. This ensures that the text data is permanently recorded and searchable later.
[1014] Input: Text data (Example: "Please tell me the stock status of the product.")
[1015] Output: Text data is saved to the database.
[1016] Specific actions:
[1017] The database will store the text, "Please tell me the product's stock status."
[1018] Step 5:
[1019] The server uses a machine learning model (e.g., BERT) to classify the converted text data into specific categories, such as "delivery," "returns," and "payment."
[1020] Input: Text data (Example: "Please tell me the stock status of the product.")
[1021] Output: Classified category (e.g., "Product Category")
[1022] Specific actions:
[1023] The server adds a label called "product category" to the database.
[1024] Step 6:
[1025] The server uses a summarization model (e.g., GPT-3) to summarize the text data. This summarized data is also stored in the database.
[1026] Input: Text data (Example: "Please tell me the stock status of the product.")
[1027] Output: Summary data (e.g., "Check product inventory")
[1028] Specific actions:
[1029] The server uses a summarization model to generate a summary such as "Check product inventory" and saves it to the database.
[1030] Step 7:
[1031] The server uses an emotion engine (e.g., Amazon Comprehend) to recognize the user's emotions from text data. This emotion data is also stored in a database.
[1032] Input: Text data (Example: "Please tell me the stock status of the product.")
[1033] Output: Emotional data (e.g., "anxiety")
[1034] Specific actions:
[1035] The server uses Amazon Comprehend to recognize the emotion "anxiety" from the text "Please tell me the availability of the product" and saves it to the database.
[1036] Step 8:
[1037] The server centrally stores audio data, text data, categorized data, summary data, and sentiment data in a database. This allows for easy searching and analysis later.
[1038] Input: Audio data, text data, classification data, summary data, sentiment data
[1039] Output: Centrally organized query data
[1040] Specific actions:
[1041] The database will store the audio file in the format: "Audio file: Audio001.wav, Text: Please tell me the product's stock status, Category: Product category, Summary: Check product stock, Emotion: Anxiety".
[1042] Step 9:
[1043] The user (manager) has the server aggregate data for a specified period. The server retrieves query data from the database for that period and calculates the number of queries and sentiment trends for each category.
[1044] Input: Aggregation instructions (Example: "From 2023-10-01 to 2023-10-10")
[1045] Output: Aggregated report (e.g., number of inquiries for product categories, anxiety trend)
[1046] Specific actions:
[1047] The server retrieves data from the database for a specified period and generates a report stating that "there were 50 inquiries for the product category, 40 of which included feelings of anxiety."
[1048] Step 10:
[1049] The server provides the generated report to the user (manager). The manager then uses the report to make business decisions.
[1050] Input: Summary Report
[1051] Output: Information for business decision-making
[1052] Specific actions:
[1053] The manager reviews the provided report and understands that there are many inquiries in the "delivery" category, many of which involve feelings of "anxiety," and then considers measures to improve the situation.
[1054] (Application Example 2)
[1055] Next, we will explain Application Example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1056] Traditional customer inquiry systems often require manual classification and summarization of inquiries, which is time-consuming and labor-intensive. Furthermore, analyzing customer sentiment is difficult, making it challenging to grasp customer satisfaction in real time. Therefore, while rapid responses based on inquiry categories and content are required, comprehensive responses, including sentiment analysis, are currently lacking. In addition, there is a lack of efficient methods for managers to derive business guidance from past inquiry data.
[1057] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1058] In this invention, the server includes means for collecting voice data of inquiries from multiple customers, means for converting the voice data into text data, means for classifying the converted text data into a plurality of predefined categories, means for summarizing the classified text data, means for recognizing customer emotions from the text data, means for storing the classified data, summarized data, and emotion data, and means for aggregating data within a specified period and calculating the number of inquiries and emotion trends for each category. This not only enables efficient classification and summarization of inquiry content, but also allows for real-time recognition of customer emotions and prompt and appropriate responses based on that information. Furthermore, it facilitates management decisions based on the stored data, leading to improved customer satisfaction and optimized operations.
[1059] "Customer" refers to the person who receives a product or service.
[1060] "Inquiry voice data" refers to voice data of questions and requests provided by customers via telephone or voice recording.
[1061] "Text data" refers to the character information obtained by converting inquiry voice data using speech recognition technology.
[1062] A "category" is a predefined group or type used to make text data easier to classify. Examples of categories include "Shipping," "Returns," and "Payment."
[1063] "Summary" means extracting the key points from text data and providing them in a shortened format.
[1064] "Emotion recognition technology" refers to techniques for recognizing a customer's psychological state and emotions from text data. This technology determines emotions such as "joy," "anger," and "sadness."
[1065] A "machine learning model" is an algorithm that uses large amounts of data to learn patterns and perform classification and prediction. It is used for text classification and summarization.
[1066] "Speech recognition technology" refers to the technology that converts speech data into text information. This is a part of speech processing technology.
[1067] A "database" refers to a system for systematically storing large amounts of data and for efficiently processing and retrieving it.
[1068] A "user interface" is an interface that allows a system and a user to interact directly. It facilitates the display and manipulation of data.
[1069] This invention is a system that efficiently collects, classifies, and summarizes customer inquiry voice data, and further recognizes customer emotions. This system operates in cooperation with a server, terminals, and users.
[1070] Data collection
[1071] When a user calls the contact center to make an inquiry, the device records the content of the inquiry as voice data. The server stores this voice data in a central database. The database also includes information such as the time of the inquiry and the caller's information.
[1072] Speech recognition and text conversion
[1073] The device converts stored audio data into text data using speech recognition technology. Specifically, it uses the Google Cloud Speech-to-Text API to analyze the audio data and convert it into text information. This converted text data is then stored in a database.
[1074] Text classification and summarization
[1075] The server classifies the converted text data into specific categories using machine learning models. For example, it uses natural language processing models such as BERT (Bidirectional Encoder Representations from Transformers) to classify the text into categories such as "delivery," "returns," and "payments." Furthermore, another machine learning model, BART (Bidirectional and Auto-Regressive Transformers), is used for summarization, extracting important parts from the converted text data and saving them in a shortened format.
[1076] emotion recognition
[1077] The server uses a natural language processing model to recognize customer emotions from text data. Specifically, it uses Google Cloud Natural Language Sentiment Analysis to determine emotions such as "joy," "anger," and "sadness" from the wording and expressions in the text. This recognized emotion data is also stored in a database.
[1078] Data storage
[1079] The server centrally stores categorized, summarized, and sentiment data in a database. This organizes query results, making them easy to search and analyze later.
[1080] Data aggregation and analysis
[1081] To aggregate data within a user-specified period, the server retrieves data from the database for that period and calculates the number of inquiries for each category. It also analyzes sentiment data to understand customer sentiment trends. This analysis provides crucial data for managers to understand customer feedback and their emotions, enabling them to make appropriate business decisions.
[1082] Specific example
[1083] For example, if a customer makes an inquiry such as, "Please tell me the delivery status of my order," the terminal first collects this audio and uses speech recognition technology to convert it into text, "Please tell me the delivery status of my order." The server categorizes this text into the "Delivery" category and generates a summary such as "Inquiry about checking delivery status." Furthermore, an emotion recognition engine recognizes that this inquiry contains the emotion of "anxiety." All of this data is stored in a database, and by aggregating the data for a specified period, managers can understand that there are many inquiries in the "Delivery" category, and that many of them contain the emotion of "anxiety."
[1084] Example of a prompt
[1085] Voice input: "Please tell me the delivery status of my item. It hasn't arrived yet and I'm worried."
[1086] Prompt: "Analyze the following text: Please tell me the delivery status of my item. I haven't received it yet and I'm worried. Based on this, categorize the inquiry, generate a summary, and analyze the sentiment."
[1087] Output: Category: "Delivery", Summary: "Checking delivery status", Emotion: "Anxiety"
[1088] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1089] Step 1:
[1090] When a user calls the contact center to make an inquiry, the device records the content of the inquiry as voice data. In this operation, the voice input is saved as an audio file. The output is the recorded voice data.
[1091] Step 2:
[1092] The server sends the recorded audio data to a database for storage. The input is the audio data, and the output is the audio data stored in the database. This data also includes information such as the time of the inquiry and the caller's information.
[1093] Step 3:
[1094] The device retrieves the stored audio data and converts it into text data using speech recognition technology. Specifically, it uses the Google Cloud Speech-to-Text API to analyze the audio data and convert it into text information. The input is audio data, and the output is text data.
[1095] Step 4:
[1096] The server receives the converted text data and uses a machine learning model to classify it into specific categories. The BERT model is used to categorize the text data into categories such as "delivery," "returns," and "payments." The input is text data, and the output is the classified categories.
[1097] Step 5:
[1098] The server uses a BART model to summarize classified text data. The input is text data, and the output is summarized text data. The server extracts the most important parts from the transformed text data and saves them in a shortened format.
[1099] Step 6:
[1100] To recognize customer emotions from text data, the server uses Google Cloud Natural Language Sentiment Analysis. The input is text data, and the output is emotion data. Emotions such as "joy," "anger," and "sadness" are determined from the wording and expressions in the text.
[1101] Step 7:
[1102] The server stores classified data, summarized data, and sentiment data in a database. The input consists of classified categories, summarized data, and sentiment data, while the output is the unified data stored in the database.
[1103] Step 8:
[1104] To aggregate data within a period specified by the user, the server retrieves data from the database for that period. The input is the specified aggregation period, and the output is the data within that period. The server calculates the number of inquiries and sentiment trends for each category.
[1105] Step 9:
[1106] Based on aggregated data and sentiment data, the results are displayed through a user interface. The input is aggregated data, and the output is a display on the user interface. This allows managers to understand customer voices and sentiments, enabling them to make appropriate business decisions.
[1107] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1108] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1109] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1110] [Fourth Embodiment]
[1111] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1112] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1113] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1114] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1115] The microphone 238 receives voice signals from the user 20 and accepts instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1116] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1117] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1118] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1119] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1120] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1121] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1122] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1123] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1124] This invention relates to a system for collecting customer inquiry voice data and efficiently classifying and summarizing that data. This system operates in cooperation with a server, terminals, and users.
[1125] Data collection
[1126] The server first collects inquiry audio data from multiple customers. For example, inquiry audio data is audio files recorded when a customer calls the call center. These audio files are stored in a database.
[1127] Speech recognition and text conversion
[1128] The device converts the collected audio data into text data using speech recognition technology. Specifically, it extracts text information from the audio data using a speech recognition library. This process converts the audio data into text data.
[1129] Text classification and summarization
[1130] The server classifies the converted text data into several predefined categories. Machine learning models are used for classification. For example, categories such as "Delivery," "Returns," and "Payment" are predefined, and the server determines which category each text belongs to. Furthermore, the server summarizes the text data. For summarization, machine learning models are used to extract the important parts of the original text and save them in a shortened format.
[1131] Data storage
[1132] The server stores the categorized and summarized data in a database. This organizes the query results, making them easy to search and analyze later.
[1133] Data aggregation and reporting
[1134] The system aggregates data within a user-specified period and calculates the number of inquiries for each category. This aggregated data is crucial for managers to understand customer feedback and make appropriate business decisions.
[1135] Specific example
[1136] For example, if a customer asks, "Please tell me the delivery status of my order," this audio is first collected and converted to text using speech recognition. The text is categorized under "Delivery," and a summary such as "Inquiry about checking delivery status" is generated. This data is stored in a database, and by aggregating the data for a specified period, managers can identify which inquiries fall under the "Delivery" category.
[1137] In this way, the present invention is a system for efficiently classifying and summarizing customer inquiries and providing information useful for business management.
[1138] The following describes the processing flow.
[1139] Step 1:
[1140] A user makes a phone call to the contact center with an inquiry. The content of the inquiry is recorded as voice data.
[1141] Step 2:
[1142] The server stores the inquiry voice data in a database. The database also includes information such as the time of the inquiry and the caller's information.
[1143] Step 3:
[1144] The device converts stored audio data into text data using speech recognition technology. Specifically, it analyzes the audio data using a speech recognition library and converts it into text information.
[1145] Step 4:
[1146] The server stores the converted text data in a database. This allows for centralized management of audio data and its corresponding text data.
[1147] Step 5:
[1148] The server uses machine learning models to categorize the stored text data. For example, the text is automatically classified into categories such as "delivery," "returns," and "payment."
[1149] Step 6:
[1150] The server summarizes each classified text data. This summarization uses machine learning models to extract key parts of the text and generate new text data in a shortened format.
[1151] Step 7:
[1152] The server stores the classified and summarized data in a new database. This ensures that each query is stored in a structured format.
[1153] Step 8:
[1154] The system aggregates data for a period specified by the user. The server retrieves data from the database for that period and calculates the number of queries for each category.
[1155] Step 9:
[1156] The server generates aggregated results in report format. The report includes the number of inquiries for each category and data showing specific trends.
[1157] Step 10:
[1158] The terminal provides the user with the generated report. The report is displayed through a web interface or dashboard.
[1159] (Example 1)
[1160] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1161] Traditional customer service systems struggled to effectively collect, classify, and summarize customer inquiry voice data. This resulted in challenges in quickly and accurately understanding inquiries and providing information to support appropriate business decisions.
[1162] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1163] In this invention, the server includes means for collecting inquiry voice data from multiple customers, means for converting the voice data into text data, and means for classifying the converted text data into a plurality of predefined categories. This makes it possible to efficiently perform a series of processes within the system, from collecting inquiry voice data to classifying, summarizing, and storing it, as well as aggregating the data.
[1164] A "customer" is a person or organization that receives various services or products, and is the person to whom inquiries about those services or products are made.
[1165] "Inquiry voice data" refers to data recorded in audio format of customer inquiries about services or products.
[1166] "Means of collecting voice data" refers to the technology and equipment used to record and save customer inquiries.
[1167] "Means of converting to text data" refers to speech recognition technologies and software used to convert audio data into text information.
[1168] A "predefined category" is a set of categories that are set up in advance to efficiently classify the content of inquiries, and may include, for example, "delivery," "returns," and "payment."
[1169] "Means of classification" refers to techniques and machine learning models used to analyze text data and classify it based on predefined categories.
[1170] "Means of summarization" refers to techniques and generative models that extract important parts from original text data and summarize the content concisely.
[1171] "Means of preservation" refers to technologies and devices for recording classified and summarized data in databases or other storage systems.
[1172] "Means for aggregating and calculating the number of inquiries for each category" refers to technologies and algorithms for analyzing data stored in a database and aggregating the number of inquiries for each category within a specified period.
[1173] A "database" refers to a system or software used to efficiently store and manage audio data, text data, classification data, summary data, and other similar information.
[1174] A "machine learning model" refers to a model that analyzes data and automatically performs classification and summarization based on that analysis. This includes statistical methods and algorithms.
[1175] A "generative model" refers to an AI model or algorithm used to generate new data or summaries from provided data.
[1176] A "prompt statement" is a phrase input into a generative model, and it refers to an instruction that the model uses to generate the desired output.
[1177] This invention is a system for collecting customer inquiry voice data and efficiently classifying and summarizing that data. This system operates through the coordinated efforts of a server, terminals, and users.
[1178] Data collection
[1179] The server first collects voice data from multiple customers. This voice data consists of audio files recorded when customers call the call center. Specifically, it is stored in a database on the server. To collect the voice data, a VoIP system is used to capture voice communications, which are then temporarily stored as audio files.
[1180] Speech recognition and text conversion
[1181] The device retrieves audio data stored on the server and converts it into text data using speech recognition technology. Examples of speech recognition libraries used include the Google Cloud Speech-to-Text API. The device sends the audio data to the API, receives the recognition result, and saves it in text format. The converted text data is then saved to the server's database.
[1182] Text classification and summarization
[1183] The server classifies the converted text data into several predefined categories using a machine learning model. For example, categories such as "Delivery," "Returns," and "Payment" are predefined, and the server automatically determines which category each text belongs to. This process uses a pre-trained classification model with scikit-learn. It also summarizes the text data using a generative AI model (e.g., GPT-3). The generated summaries are stored in a database.
[1184] Data storage
[1185] The classified text data and summary data are stored in a database by the server. During this process, they are organized into records containing category labels and summary text, and then saved.
[1186] Data aggregation and reporting
[1187] Users request data aggregation for a specified period. Users can enter the specified period through the web application interface. The server queries the database based on the specified period and aggregates the number of queries for each category. This is done using SQL queries. The aggregation results are reported to the user and displayed on the web application's dashboard.
[1188] Specific example
[1189] For example, if a customer asks, "Please tell me the delivery status of my order," the server first collects this audio. Then, the terminal uses speech recognition technology to convert the audio data into text, "Please tell me the delivery status of my order." This text data is then categorized by the server as "Delivery" and further summarized as "Inquiry about checking delivery status" using a generative AI model. This data is stored in a database, and later, the system aggregates the number of inquiries in the "Delivery" category and reports it to the user.
[1190] Example of a prompt
[1191] The following is an example of a prompt to input into a generative AI model:
[1192] "Convert the customer inquiry audio 'Please tell me the delivery status of my order' into text, and categorize that text under 'Delivery.' Furthermore, summarize this text and save it to the database."
[1193] In this way, the present invention is a system for efficiently classifying and summarizing customer inquiries and providing information useful for business management.
[1194] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1195] Step 1:
[1196] The server collects voice data from multiple customers. When a customer calls the call center, the VoIP system captures the voice data and temporarily stores it as an audio file. The input here is the customer's voice, and the output is an audio file. This audio file is stored in a database.
[1197] Step 2:
[1198] The device retrieves an audio file stored on the server. The input is the path to the audio file, and the output is the retrieved audio file itself. Specifically, the audio file is downloaded using a REST API.
[1199] Step 3:
[1200] The device converts the acquired audio file into text data using the Google Cloud Speech-to-Text API. The input is an audio file, and the output is the converted text data. The audio data is sent to the API, and the recognition results are received in text format.
[1201] Step 4:
[1202] The terminal sends the converted text data to the server. The input is text data, and the output is text data stored in the server's database. Specifically, a REST API is used to send text data to the server and store it in the database.
[1203] Step 5:
[1204] The server inputs text data stored in a database into a machine learning model and classifies it into predefined categories. The input is text data, and the output is the classification result of the text. The model used here is a classification model trained with scikit-learn. Specifically, it generates data with category labels attached.
[1205] Step 6:
[1206] The server inputs classified text data into a generating AI model (e.g., GPT-3) to generate a summary. The input is classified text data, and the output is summarized text data. The summary extracts the important parts and puts them into a shorter format.
[1207] Step 7:
[1208] The server stores the classified and summarized data in a database. The input is category labels and summary text data, and the output is the records stored in the database. This ensures that the data is systematically stored, facilitating later searching and analysis.
[1209] Step 8:
[1210] The user requests data aggregation for a specified period. The input is the specified period (e.g., a date range), and the output is the query aggregation results for each category within that period. The user submits the request via a web application.
[1211] Step 9:
[1212] The server queries the database based on a specified period and aggregates the number of queries for each category. The input is period information, and the output is the aggregated result. SQL queries are used to search the database and calculate the number of queries for each category.
[1213] Step 10:
[1214] The server reports the aggregated results to the user. The input is the aggregated results, and the output is a report displayed on the web application's dashboard. The user views this report and uses it to inform business decisions.
[1215] (Application Example 1)
[1216] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1217] Traditional customer inquiry systems had problems such as the time it took to collect, classify, and summarize voice data, making real-time data display and analysis difficult. This reduced the efficiency of customer service and made it difficult for managers to make quick business decisions.
[1218] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1219] In this invention, the server includes means for collecting voice data of inquiries from multiple customers, means for converting the voice data into text data, means for classifying the converted text data into a plurality of predefined categories, means for summarizing the classified text data, means for storing the classified data and the summarized data, means for aggregating data within a specified period and calculating the number of inquiries for each category, and means for collecting and displaying the text data in real time. This improves the efficiency of customer service and enables managers to make quick business decisions.
[1220] "Multiple customers" refers to two or more users who are the entities making inquiries.
[1221] "Inquiry voice data" refers to voice information uttered by customers, which includes questions and requests.
[1222] "Means of collection" refers to systems and devices for recording audio data in digital format and transmitting it to a server.
[1223] "Text data" refers to audio data converted into written information, which can then be processed mechanically.
[1224] "Means of conversion" refer to technologies and systems for converting audio data into text information.
[1225] "Multiple predefined categories" refers to several pre-defined classification items, such as "delivery," "returns," and "payment."
[1226] "Means of classification" refer to technologies and systems for sorting text data into predefined categories.
[1227] "Methods of summarization" refer to technologies and systems for extracting important parts from text data and summarizing them concisely.
[1228] "Means of preservation" refers to systems for storing data in databases or storage devices.
[1229] "Data within a specified period" refers to query data collected within a specific time frame.
[1230] "Means for calculating the number of inquiries per category" refers to technologies or systems for calculating the number of inquiries classified into each category.
[1231] "Means of collecting and displaying data in real time" refers to technologies and systems for instantly collecting and visually displaying inquiry data.
[1232] A "machine learning model" is an algorithm or program that learns from large amounts of data and automatically performs tasks such as data classification and summarization.
[1233] "Speech recognition technology" is a technology that converts speech data into text information.
[1234] A "database" is a system for efficiently storing, searching, and accessing large amounts of data.
[1235] A "visualization algorithm" is a technology that displays collected data in the form of graphs and charts, making it easy to understand visually.
[1236] An "algorithm" is a set of procedures or computational steps for solving a specific problem.
[1237] This invention relates to a system for collecting customer inquiry voice data and efficiently classifying and summarizing that data. The system operates in cooperation with a server, terminals, and users.
[1238] Data collection
[1239] The server collects voice data from multiple customers. This voice data includes, for example, audio files recorded when a customer calls the call center. These audio files are stored in a database.
[1240] Speech recognition and text conversion
[1241] The device converts the collected audio data into text data using speech recognition technology. Specifically, it uses the Google Speech-to-Text API. It extracts text information from the audio data and converts it into text data.
[1242] Text classification and summarization
[1243] The server classifies the converted text data into several predefined categories. This classification uses machine learning models (e.g., TensorFlow, PyTorch). For example, categories such as "Delivery," "Returns," and "Payment" are predefined, and the server determines which category each text belongs to. Furthermore, the server summarizes the text data. This summarization also uses machine learning models to extract the most important parts of the original text and saves them in a shortened format.
[1244] Real-time display
[1245] The server implements a visualization algorithm to collect and display the aforementioned text data in real time. This allows inquiry data to be immediately retrieved, enabling stakeholders to understand the situation in real time.
[1246] Data storage
[1247] The server stores categorized and summarized data in a database. This organizes the query results, making them easy to search and analyze later. Databases used include Firebase and MySQL.
[1248] Data aggregation and reporting
[1249] The system aggregates data within a period specified by the user and calculates the number of inquiries for each category. This aggregated result provides important data for managers to understand customer feedback and make appropriate business decisions.
[1250] Specific example
[1251] For example, if a customer asks, "Please tell me the delivery status of my order," the audio is first collected and converted to text using speech recognition. The text is categorized under "Delivery," and a summary such as "Inquiry about checking delivery status" is generated. This data is stored in a database, and by aggregating the data for a specified period, managers can understand which inquiries fall under the "Delivery" category. Furthermore, because the collected information is displayed in real time, it is possible to shorten the time from the moment an inquiry is made until a response is given.
[1252] Example of a prompt
[1253] Use the Google Speech-to-Text API to convert speech to text, input the text into a TensorFlow model, and categorize it into categories such as "Shipping," "Returns," and "Payment." Furthermore, extract and summarize the key points. Save the classification results and summaries to a Firebase database. Finally, aggregate the number of inquiries within a specified period and display it as a report.
[1254] This significantly improves the efficiency of customer service, allowing managers to quickly implement appropriate countermeasures.
[1255] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1256] Step 1: Collect audio data
[1257] The server collects customer inquiry voice data in real time. Specifically, it records the voice call when a customer makes an inquiry and transfers the audio file to the server. The input is the customer's voice data, and the output is an audio file in digital format.
[1258] Step 2: Converting audio data to text
[1259] The device converts the collected audio data into text data. Here, the Google Speech-to-Text API is used to convert the audio data into text information. The input is the audio data collected in step 1, and the output is the converted text data.
[1260] Step 3: Classification of text data
[1261] The server classifies the converted text data into several predefined categories. A machine learning model (e.g., TensorFlow) is used to execute an algorithm that classifies the text data. The input is the text data obtained in step 2, and the output is the text data classified by category.
[1262] Step 4: Textbook Summary
[1263] The server summarizes the classified text data. It uses a machine learning model (e.g., PyTorch) to extract and shorten the important parts. The input is the text data classified in step 3, and the output is the summarized text data.
[1264] Step 5: Save Data
[1265] The server stores the classified and summarized data in a database. Specifically, it uses a database management system such as Firebase or MySQL to store the data. The input is the data obtained in steps 3 and 4, and the output is the state in which it is stored in the database.
[1266] Step 6: Real-time display of data
[1267] The server collects query data in real time and displays it using a visualization algorithm. The input is all the data obtained from step 2 onward, and the output is a visual display format that is updated in real time.
[1268] Step 7: Data aggregation and reporting
[1269] The user aggregates inquiry data within a specified period and calculates the number of inquiries for each category. The server uses an aggregation algorithm to calculate the number of inquiries and displays it as a report. The input is all data stored in the database, and the output is a report of the aggregated results.
[1270] The above steps enable a system that efficiently classifies, summarizes, and displays customer inquiry voice data in real time.
[1271] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1272] This invention combines a system that collects customer inquiry voice data, efficiently classifies and summarizes that data, with an emotion engine that recognizes user emotions. This system operates in cooperation with a server, terminals, and users.
[1273] Data collection
[1274] When a user calls the contact center with an inquiry, the content of the inquiry is recorded as voice data. The server stores this voice data in a database. The database also includes information such as the time of the inquiry and the caller's information.
[1275] Speech recognition and text conversion
[1276] The device converts stored audio data into text data using speech recognition technology. It analyzes the audio data using a speech recognition library and converts it into text information. This converted text data is stored in a database.
[1277] Text classification and summarization
[1278] The server uses a machine learning model to classify the converted text data into specific categories. For example, it might be categorized as "delivery," "returns," or "payment." Another machine learning model is used for summarization, extracting the most important parts from the converted text data and saving them in a shortened format.
[1279] emotion recognition
[1280] A distinctive feature of this invention is the provision of an emotion engine. The server uses this emotion engine to recognize the user's emotions from text data. For example, it determines emotions such as "joy," "anger," and "sadness" from the wording and expressions in the text. This recognized emotion data is also stored in a database.
[1281] Data storage
[1282] The server centrally stores categorized, summarized, and sentiment data in a database. This organizes query results, making them easy to search and analyze later.
[1283] Data aggregation and analysis
[1284] To aggregate data within a user-specified period, the server retrieves data from the database for that period and calculates the number of inquiries for each category. It also analyzes sentiment data to understand user sentiment trends. This analysis provides crucial data for managers to understand customer feedback and their emotions, enabling them to make appropriate business decisions.
[1285] Specific example
[1286] For example, if a customer makes an inquiry such as, "Please tell me the delivery status of my order," this audio is first collected and converted into text using speech recognition. The text is then categorized under "Delivery," and a summary such as "Inquiry about checking delivery status" is generated. Furthermore, an emotion engine recognizes that this inquiry contains the emotion of "anxiety." All of this data is stored in a database, and by aggregating data for a specified period, managers can understand that there are many inquiries in the "Delivery" category, and that many of them contain the emotion of "anxiety."
[1287] Thus, this invention efficiently classifies and summarizes customer inquiries and provides information useful for business management by further analyzing sentiment data.
[1288] The following describes the processing flow.
[1289] Step 1:
[1290] A user makes a phone call to the contact center with an inquiry. The content of the inquiry is recorded as voice data.
[1291] Step 2:
[1292] The server stores the inquiry voice data in a database. The database also includes information such as the time of the inquiry and the caller's information.
[1293] Step 3:
[1294] The device converts stored audio data into text data using speech recognition technology. Specifically, it analyzes the audio data using a speech recognition library and converts it into text information.
[1295] Step 4:
[1296] The server stores the converted text data in a database. This allows for centralized management of audio data and its corresponding text data.
[1297] Step 5:
[1298] The server uses machine learning models to classify stored text data into specific categories. For example, categories such as "delivery," "returns," and "payments" are predefined, and the server determines which category each text belongs to.
[1299] Step 6:
[1300] The server summarizes each classified text data. This summarization uses machine learning models to extract key parts of the text and generate new text data in a shortened format.
[1301] Step 7:
[1302] The server stores the classified and summarized data in a database. This organizes the query content, making it easy to search and analyze later.
[1303] Step 8:
[1304] The server inputs the stored text data into the emotion engine to recognize the user's emotions. The emotion engine determines emotions such as "joy," "anger," and "sadness" from the text data.
[1305] Step 9:
[1306] The server stores the recognized sentiment data in a database. This allows for the management of sentiment data for each query.
[1307] Step 10:
[1308] The system aggregates data for a period specified by the user. The server retrieves data from the database for that period and calculates the number of queries and sentiment trends for each category.
[1309] Step 11:
[1310] The server generates aggregated results in report format. The report includes the number of inquiries for each category and corresponding sentiment data.
[1311] Step 12:
[1312] The terminal provides the user with the generated report. The report is displayed through a web interface or dashboard.
[1313] Specific example
[1314] For example, if a customer makes an inquiry such as, "Please tell me the delivery status of my order," this audio is first collected and converted into text using speech recognition. The text is then categorized under "Delivery," and a summary such as "Inquiry about checking delivery status" is generated. Furthermore, an emotion engine recognizes that this inquiry contains the emotion of "anxiety." All of this data is stored in a database, and by aggregating data for a specified period, managers can understand that there are many inquiries in the "Delivery" category, and that many of them contain the emotion of "anxiety."
[1315] (Example 2)
[1316] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1317] Conventional inquiry management systems struggled to efficiently collect, classify, and summarize customer inquiry voice data. Furthermore, the lack of a means to accurately recognize and analyze the emotions expressed in the inquiries presented posed challenges in providing appropriate responses and service improvements that considered customer feelings.
[1318] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for collecting voice data inquiries from multiple customers, means for converting the voice data into text data, means for classifying the converted text data into a plurality of predefined categories, means for summarizing the classified text data, means for recognizing the user's emotions from the text data, means for storing the classified data, summarized data, and emotion data, and means for aggregating data within a specified period and calculating the number of inquiries and emotion trends for each category. This enables efficient management of customer inquiries and detailed data analysis, including user emotions.
[1319] "Inquiry voice data" refers to voice information collected from customer phone calls, voice messages, and other means.
[1320] "Text data" refers to digital data obtained by converting audio data into written text.
[1321] A "machine learning model" is an algorithm that is trained using large amounts of data and automatically performs specific tasks (such as classification, prediction, and generation).
[1322] "Speech recognition technology" is a technology that analyzes speech data and converts it into text data.
[1323] A "category" is a predefined item or group used to classify the content of an inquiry.
[1324] A "summary" is data that extracts the important parts from long text data and expresses them in a shorter format.
[1325] An "emotion engine" is software or an algorithm that analyzes text data and identifies the emotions contained within it.
[1326] A "database" is a system for efficiently storing, retrieving, and managing data.
[1327] "Data aggregation" is the process of compiling data based on certain criteria and calculating statistics or totals.
[1328] "Emotional trends" refer to data that shows fluctuations and tendencies in users' emotions over a certain period of time.
[1329] This invention is a system that efficiently collects, classifies, and summarizes customer inquiry voice data, and further recognizes the user's emotions. This system operates in cooperation with a server, terminals, and users.
[1330] Data collection
[1331] When a user calls the contact center, the content of their inquiry is recorded as voice data by the server. The server stores the collected voice data, along with the time of the inquiry and the caller's information, in a database.
[1332] Specific example:
[1333] When a user asks, "Please tell me the availability of the product," the voice message is recorded by the server and stored in the database.
[1334] Speech recognition and text conversion
[1335] The audio data stored on the server is transferred to the terminal. The terminal uses a speech recognition library (e.g., Google Cloud Speech-to-Text API) to convert the audio data into text data, and then saves that text data back to the database.
[1336] Specific example:
[1337] The device converts the spoken phrase "Please tell me the product's stock status" into text data using the Google Cloud Speech-to-Text API and saves that text to a database.
[1338] Text classification and summarization
[1339] The server classifies the converted text data into specific categories using a machine learning model (e.g., BERT). It also uses another machine learning model (e.g., GPT-3) for summarization, extracting important parts from the text data and saving them in a shortened format.
[1340] Specific example:
[1341] The server converts the text "Please tell me the product's stock status" into a summary "Check product stock" and categorizes it under "Product Category".
[1342] emotion recognition
[1343] The server uses an emotion engine (e.g., Amazon Comprehend) to recognize the user's emotions from text data. The recognized emotion data is stored in a database.
[1344] Specific example:
[1345] The server recognizes the emotion of "anxiety" from the text "Please tell me the product's stock status" and stores the emotion data in the database.
[1346] Data storage
[1347] The server centrally stores audio data, text data, categorized and summarized data, and sentiment data in a database.
[1348] Specific example:
[1349] The database stores the audio file in the format "Audio file: Audio001.wav, Text: Please tell me the product's stock status, Category: Product category, Summary: Check product stock, Emotion: Anxiety".
[1350] Data aggregation and analysis
[1351] The user (manager) aggregates data for a specified period. The server retrieves this data from the database and calculates the number of inquiries and sentiment trends for each category. The server reports these analysis results to the manager to help them make informed business decisions.
[1352] Specific example:
[1353] The server generates a report stating that "data from 2023-10-01 to 2023-10-10 was compiled, and there were 50 inquiries regarding product categories, 40 of which included feelings of anxiety." The manager then reviews this report and uses it as a guide for business management.
[1354] Example of a prompt
[1355] "Design a system to collect, transcribe, categorize, and summarize customer inquiry audio, and to recognize user sentiment."
[1356] This system efficiently classifies and summarizes customer inquiries and analyzes sentiment data to provide valuable information for business management.
[1357] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1358] Step 1:
[1359] A user calls the contact center. During this call, the user verbally communicates their specific inquiry. For example, "Please tell me the availability of the product."
[1360] Input: User inquiry voice
[1361] Output: Raw audio data sent to the server
[1362] Step 2:
[1363] The server collects voice data from user inquiries in real time and saves this voice data as a file. It also adds information about the time of the inquiry and the caller, and saves this data to a database.
[1364] Input: User inquiry voice
[1365] Output: Audio file (e.g., audio001.wav) and metadata (time, caller information)
[1366] Specific actions:
[1367] The server saves the audio data as "Audio001.wav" and simultaneously records it in the database in the format "2023-10-10 10:00:00, Customer A, Audio001.wav".
[1368] Step 3:
[1369] The audio data stored on the server is transferred to the terminal. The terminal then converts this audio data into text data using a speech recognition library (e.g., Google Cloud Speech-to-Text API).
[1370] Input: Audio file (Audio001.wav)
[1371] Output: Text data (Example: "Please tell me the product's stock status")
[1372] Specific actions:
[1373] The device receives the audio file "Audio001.wav" and calls the Google Cloud Speech-to-Text API to convert it to text. As a result, the text data "Please tell me the product's stock status" is generated.
[1374] Step 4:
[1375] The device saves the converted text data to a database. This ensures that the text data is permanently recorded and searchable later.
[1376] Input: Text data (Example: "Please tell me the stock status of the product.")
[1377] Output: Text data is saved to the database.
[1378] Specific actions:
[1379] The database will store the text, "Please tell me the product's stock status."
[1380] Step 5:
[1381] The server uses a machine learning model (e.g., BERT) to classify the converted text data into specific categories, such as "delivery," "returns," and "payment."
[1382] Input: Text data (Example: "Please tell me the stock status of the product.")
[1383] Output: Classified category (e.g., "Product Category")
[1384] Specific actions:
[1385] The server adds a label called "product category" to the database.
[1386] Step 6:
[1387] The server uses a summarization model (e.g., GPT-3) to summarize the text data. This summarized data is also stored in the database.
[1388] Input: Text data (Example: "Please tell me the stock status of the product.")
[1389] Output: Summary data (e.g., "Check product inventory")
[1390] Specific actions:
[1391] The server uses a summarization model to generate a summary such as "Check product inventory" and saves it to the database.
[1392] Step 7:
[1393] The server uses an emotion engine (e.g., Amazon Comprehend) to recognize the user's emotions from text data. This emotion data is also stored in a database.
[1394] Input: Text data (Example: "Please tell me the stock status of the product.")
[1395] Output: Emotional data (e.g., "anxiety")
[1396] Specific actions:
[1397] The server uses Amazon Comprehend to recognize the emotion "anxiety" from the text "Please tell me the availability of the product" and saves it to the database.
[1398] Step 8:
[1399] The server centrally stores audio data, text data, categorized data, summary data, and sentiment data in a database. This allows for easy searching and analysis later.
[1400] Input: Audio data, text data, classification data, summary data, sentiment data
[1401] Output: Centrally organized query data
[1402] Specific actions:
[1403] The database will store the audio file in the format: "Audio file: Audio001.wav, Text: Please tell me the product's stock status, Category: Product category, Summary: Check product stock, Emotion: Anxiety".
[1404] Step 9:
[1405] The user (manager) has the server aggregate data for a specified period. The server retrieves query data from the database for that period and calculates the number of queries and sentiment trends for each category.
[1406] Input: Aggregation instructions (Example: "From 2023-10-01 to 2023-10-10")
[1407] Output: Aggregated report (e.g., number of inquiries for product categories, anxiety trend)
[1408] Specific actions:
[1409] The server retrieves data from the database for a specified period and generates a report stating that "there were 50 inquiries for the product category, 40 of which included feelings of anxiety."
[1410] Step 10:
[1411] The server provides the generated report to the user (manager). The manager then uses the report to make business decisions.
[1412] Input: Summary Report
[1413] Output: Information for business decision-making
[1414] Specific actions:
[1415] The manager reviews the provided report and understands that there are many inquiries in the "delivery" category, many of which involve feelings of "anxiety," and then considers measures to improve the situation.
[1416] (Application Example 2)
[1417] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1418] Traditional customer inquiry systems often require manual classification and summarization of inquiries, which is time-consuming and labor-intensive. Furthermore, analyzing customer sentiment is difficult, making it challenging to grasp customer satisfaction in real time. Therefore, while rapid responses based on inquiry categories and content are required, comprehensive responses, including sentiment analysis, are currently lacking. In addition, there is a lack of efficient methods for managers to derive business guidance from past inquiry data.
[1419] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1420] In this invention, the server includes means for collecting voice data of inquiries from multiple customers, means for converting the voice data into text data, means for classifying the converted text data into a plurality of predefined categories, means for summarizing the classified text data, means for recognizing customer emotions from the text data, means for storing the classified data, summarized data, and emotion data, and means for aggregating data within a specified period and calculating the number of inquiries and emotion trends for each category. This not only enables efficient classification and summarization of inquiry content, but also allows for real-time recognition of customer emotions and prompt and appropriate responses based on that information. Furthermore, it facilitates management decisions based on the stored data, leading to improved customer satisfaction and optimized operations.
[1421] "Customer" refers to the person who receives a product or service.
[1422] "Inquiry voice data" refers to voice data of questions and requests provided by customers via telephone or voice recording.
[1423] "Text data" refers to the character information obtained by converting inquiry voice data using speech recognition technology.
[1424] A "category" is a predefined group or type used to make text data easier to classify. Examples of categories include "Shipping," "Returns," and "Payment."
[1425] "Summary" means extracting the key points from text data and providing them in a shortened format.
[1426] "Emotion recognition technology" refers to techniques for recognizing a customer's psychological state and emotions from text data. This technology determines emotions such as "joy," "anger," and "sadness."
[1427] A "machine learning model" is an algorithm that uses large amounts of data to learn patterns and perform classification and prediction. It is used for text classification and summarization.
[1428] "Speech recognition technology" refers to the technology that converts speech data into text information. This is a part of speech processing technology.
[1429] A "database" refers to a system for systematically storing large amounts of data and for efficiently processing and retrieving it.
[1430] A "user interface" is an interface that allows a system and a user to interact directly. It facilitates the display and manipulation of data.
[1431] This invention is a system that efficiently collects, classifies, and summarizes customer inquiry voice data, and further recognizes customer emotions. This system operates in cooperation with a server, terminals, and users.
[1432] Data collection
[1433] When a user calls the contact center to make an inquiry, the device records the content of the inquiry as voice data. The server stores this voice data in a central database. The database also includes information such as the time of the inquiry and the caller's information.
[1434] Speech recognition and text conversion
[1435] The device converts stored audio data into text data using speech recognition technology. Specifically, it uses the Google Cloud Speech-to-Text API to analyze the audio data and convert it into text information. This converted text data is then stored in a database.
[1436] Text classification and summarization
[1437] The server classifies the converted text data into specific categories using machine learning models. For example, it uses natural language processing models such as BERT (Bidirectional Encoder Representations from Transformers) to classify the text into categories such as "delivery," "returns," and "payments." Furthermore, another machine learning model, BART (Bidirectional and Auto-Regressive Transformers), is used for summarization, extracting important parts from the converted text data and saving them in a shortened format.
[1438] emotion recognition
[1439] The server uses a natural language processing model to recognize customer emotions from text data. Specifically, it uses Google Cloud Natural Language Sentiment Analysis to determine emotions such as "joy," "anger," and "sadness" from the wording and expressions in the text. This recognized emotion data is also stored in a database.
[1440] Data storage
[1441] The server centrally stores categorized, summarized, and sentiment data in a database. This organizes query results, making them easy to search and analyze later.
[1442] Data aggregation and analysis
[1443] To aggregate data within a user-specified period, the server retrieves data from the database for that period and calculates the number of inquiries for each category. It also analyzes sentiment data to understand customer sentiment trends. This analysis provides crucial data for managers to understand customer feedback and their emotions, enabling them to make appropriate business decisions.
[1444] Specific example
[1445] For example, if a customer makes an inquiry such as, "Please tell me the delivery status of my order," the terminal first collects this audio and uses speech recognition technology to convert it into text, "Please tell me the delivery status of my order." The server categorizes this text into the "Delivery" category and generates a summary such as "Inquiry about checking delivery status." Furthermore, an emotion recognition engine recognizes that this inquiry contains the emotion of "anxiety." All of this data is stored in a database, and by aggregating the data for a specified period, managers can understand that there are many inquiries in the "Delivery" category, and that many of them contain the emotion of "anxiety."
[1446] Example of a prompt
[1447] Voice input: "Please tell me the delivery status of my item. It hasn't arrived yet and I'm worried."
[1448] Prompt: "Analyze the following text: Please tell me the delivery status of my item. I haven't received it yet and I'm worried. Based on this, categorize the inquiry, generate a summary, and analyze the sentiment."
[1449] Output: Category: "Delivery", Summary: "Checking delivery status", Emotion: "Anxiety"
[1450] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1451] Step 1:
[1452] When a user calls the contact center to make an inquiry, the device records the content of the inquiry as voice data. In this operation, the voice input is saved as an audio file. The output is the recorded voice data.
[1453] Step 2:
[1454] The server sends the recorded audio data to a database for storage. The input is the audio data, and the output is the audio data stored in the database. This data also includes information such as the time of the inquiry and the caller's information.
[1455] Step 3:
[1456] The device retrieves the stored audio data and converts it into text data using speech recognition technology. Specifically, it uses the Google Cloud Speech-to-Text API to analyze the audio data and convert it into text information. The input is audio data, and the output is text data.
[1457] Step 4:
[1458] The server receives the converted text data and uses a machine learning model to classify it into specific categories. The BERT model is used to categorize the text data into categories such as "delivery," "returns," and "payments." The input is text data, and the output is the classified categories.
[1459] Step 5:
[1460] The server uses a BART model to summarize classified text data. The input is text data, and the output is summarized text data. The server extracts the most important parts from the transformed text data and saves them in a shortened format.
[1461] Step 6:
[1462] To recognize customer emotions from text data, the server uses Google Cloud Natural Language Sentiment Analysis. The input is text data, and the output is emotion data. Emotions such as "joy," "anger," and "sadness" are determined from the wording and expressions in the text.
[1463] Step 7:
[1464] The server stores classified data, summarized data, and sentiment data in a database. The input consists of classified categories, summarized data, and sentiment data, while the output is the unified data stored in the database.
[1465] Step 8:
[1466] To aggregate data within a period specified by the user, the server retrieves data from the database for that period. The input is the specified aggregation period, and the output is the data within that period. The server calculates the number of inquiries and sentiment trends for each category.
[1467] Step 9:
[1468] Based on aggregated data and sentiment data, the results are displayed through a user interface. The input is aggregated data, and the output is a display on the user interface. This allows managers to understand customer voices and sentiments, enabling them to make appropriate business decisions.
[1469] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1470] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1471] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1472] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1473] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1474] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1475] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1476] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1477] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1478] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1479] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1480] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1481] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1482] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1483] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1484] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1485] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1486] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1487] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1488] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[1489] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[1490] The following is further disclosed regarding the embodiments described above.
[1491] (Claim 1)
[1492] A means of collecting voice data from inquiries from multiple customers,
[1493] Means for converting the aforementioned audio data into text data,
[1494] Means for classifying the converted text data into a plurality of predefined categories,
[1495] Means for summarizing the classified text data,
[1496] Means for storing the classified data and summarized data,
[1497] A means for aggregating data within a specified period and calculating the number of inquiries for each category,
[1498] A system that includes this.
[1499] (Claim 2)
[1500] Methods for using machine learning models to classify text data,
[1501] Means for using a machine learning model to summarize the aforementioned text data,
[1502] The system according to claim 1, further comprising:
[1503] (Claim 3)
[1504] A means of using speech recognition technology to convert audio data into text data,
[1505] Means for using a database to store and retrieve the aforementioned text data,
[1506] The system according to claim 1, further comprising:
[1507] "Example 1"
[1508] (Claim 1)
[1509] A means of collecting voice data from inquiries from multiple customers,
[1510] Means for converting the aforementioned audio data into text data,
[1511] Means for classifying the converted text data into a plurality of predefined categories,
[1512] Means for summarizing the classified text data,
[1513] Means for storing the classified data and summarized data,
[1514] A means for aggregating data within a specified period and calculating the number of inquiries for each category,
[1515] Means for using a database to store and retrieve audio data,
[1516] A system that includes this.
[1517] (Claim 2)
[1518] Methods for using machine learning models to classify text data,
[1519] Means for using a generative model to summarize the aforementioned text data,
[1520] The system according to claim 1, further comprising:
[1521] (Claim 3)
[1522] A means of using speech recognition technology to convert audio data into text data,
[1523] A means of generating and inputting prompt statements to summarize text data,
[1524] The system according to claim 1, further comprising:
[1525] "Application Example 1"
[1526] (Claim 1)
[1527] A means of collecting voice data from inquiries from multiple customers,
[1528] Means for converting the aforementioned audio data into text data,
[1529] Means for classifying the converted text data into a plurality of predefined categories,
[1530] Means for summarizing the classified text data,
[1531] Means for storing the classified data and summarized data,
[1532] A means for aggregating data within a specified period and calculating the number of inquiries for each category,
[1533] A means for collecting and displaying the aforementioned text data in real time,
[1534] A system that includes this.
[1535] (Claim 2)
[1536] Methods for using machine learning models to classify text data,
[1537] Means for using a machine learning model to summarize the aforementioned text data,
[1538] Means for using a visualization algorithm to display the aforementioned text data in real time,
[1539] The system according to claim 1, further comprising:
[1540] (Claim 3)
[1541] A means of using speech recognition technology to convert audio data into text data,
[1542] Means for using a database to store and retrieve the aforementioned text data,
[1543] A means for using an algorithm to analyze the aforementioned audio data in real time,
[1544] The system according to claim 1, further comprising:
[1545] "Example 2 of combining an emotion engine"
[1546] (Claim 1)
[1547] A means of collecting voice data from inquiries from multiple customers,
[1548] Means for converting the aforementioned audio data into text data,
[1549] Means for classifying the converted text data into a plurality of predefined categories,
[1550] Means for summarizing the classified text data,
[1551] A means for recognizing the user's emotions from the aforementioned text data,
[1552] Means for storing the classified data, summarized data and sentiment data,
[1553] A means for aggregating data within a specified period and calculating the number of inquiries and sentiment trends for each category,
[1554] A system that includes this.
[1555] (Claim 2)
[1556] Methods for using machine learning models to classify text data,
[1557] Means for using a machine learning model to summarize the aforementioned text data,
[1558] Means for using an emotion engine to recognize emotions from the aforementioned text data,
[1559] The system according to claim 1, further comprising:
[1560] (Claim 3)
[1561] A means of using speech recognition technology to convert audio data into text data,
[1562] Means for using a database to store and retrieve the aforementioned text data,
[1563] An analytical means for visualizing the number of inquiries and sentiment trends for each of the aforementioned categories,
[1564] The system according to claim 1, further comprising:
[1565] "Application example 2 when combining with an emotional engine"
[1566] (Claim 1)
[1567] A means of collecting voice data from inquiries from multiple customers,
[1568] Means for converting the aforementioned audio data into text data,
[1569] Means for classifying the converted text data into a plurality of predefined categories,
[1570] Means for summarizing the classified text data,
[1571] A means of recognizing customer emotions from text data,
[1572] Means for storing the classified data, summarized data and sentiment data,
[1573] A means for aggregating data within a specified period and calculating the number of inquiries and sentiment trends for each category,
[1574] A system that includes this.
[1575] (Claim 2)
[1576] Methods for using machine learning models to classify text data,
[1577] Means for using a machine learning model to summarize the aforementioned text data,
[1578] Means for using a natural language processing model to recognize emotions from the aforementioned text data,
[1579] The system according to claim 1, further comprising:
[1580] (Claim 3)
[1581] A means of using speech recognition technology to convert audio data into text data,
[1582] Means for using a database to store and retrieve the aforementioned text data and sentiment data,
[1583] Means comprising a user interface for displaying the classified data and sentiment data,
[1584] The system according to claim 1, further comprising: [Explanation of symbols]
[1585] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of collecting voice data from inquiries from multiple customers, Means for converting the aforementioned audio data into text data, Means for classifying the converted text data into a plurality of predefined categories, Means for summarizing the classified text data, Means for storing the classified data and summarized data, A means for aggregating data within a specified period and calculating the number of inquiries for each category, A system that includes this.
2. Methods for using machine learning models to classify text data, Means for using a machine learning model to summarize the aforementioned text data, The system according to claim 1, further comprising:
3. A means of using speech recognition technology to convert audio data into text data, Means for using a database to store and retrieve the aforementioned text data, The system according to claim 1, further comprising:
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A