system

The system addresses the inefficiency in analyzing customer service voice data by automatically recording, analyzing, and generating statistical data to improve business operations through a recording unit, analysis unit, and guideline provision unit, thereby enhancing customer service quality and efficiency.

JP2026038587APending Publication Date: 2026-03-06SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024142110
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-23
Publication Date
2026-03-06

AI Technical Summary

Technical Problem

Conventional technologies have not efficiently collected and analyzed customer service voice data to improve business operations.

Method used

A system comprising a recording unit, analysis unit, and guideline provision unit that automatically records, analyzes, and generates statistical data from customer service voice data using a generation AI to provide guidelines for business improvement.

Benefits of technology

The system effectively evaluates customer service quality and provides actionable guidelines for improvement by analyzing voice data, enhancing customer service efficiency and business operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026038587000001_ABST
    Figure 2026038587000001_ABST
Patent Text Reader

Abstract

The system according to the embodiment aims to analyze voice data of customer service and provide guidelines for business improvement. [Solution] A system according to an embodiment includes a recording unit, an analysis unit, a statistics generation unit, and a guideline provision unit. The recording unit automatically records the voice of customers. The analysis unit analyzes the voice data recorded by the recording unit. The statistics generation unit generates statistical data based on the data analyzed by the analysis unit. The guideline provision unit provides guidelines for business improvement based on the statistical data generated by the statistics generation unit.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional technologies have not been able to efficiently collect and analyze customer service voice data and utilize it to improve business operations, so there is room for improvement.

[0005] The system according to the embodiment aims to analyze voice data of customer service and provide guidelines for business improvement. [Means for solving the problem]

[0006] The system according to the embodiment includes a recording unit, an analysis unit, a statistics generation unit, and a guideline provision unit. The recording unit automatically records the voice of customers. The analysis unit analyzes the voice data recorded by the recording unit. The statistics generation unit generates statistical data based on the data analyzed by the analysis unit. The guideline provision unit provides guidelines for business improvement based on the statistical data generated by the statistics generation unit. [Effects of the Invention]

[0007] The system according to the embodiment can analyze voice data of customer service and provide guidelines for business improvement. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. DETAILED DESCRIPTION OF THE INVENTION

[0009] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0010] First, the terms used in the following description will be explained.

[0011] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, the processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), or a TPU (Tensor Processing Unit).

[0012] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0013] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0014] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), and Bluetooth (registered trademark).

[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0016] [First embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0017] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0019] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0020] The reception device 38 includes a touch panel 38A and a microphone 38B, and receives user input. The touch panel 38A detects contact with a pointer (for example, a pen or a finger) to receive user input by the touch of the pointer. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 (see FIG. 2) acquires the data indicating the user input.

[0021] Output device 40 includes a display 40A and a speaker 40B, and presents data to a user by outputting the data in a form of expression that the user can perceive (e.g., audio and / or text). Display 40A displays visible information such as text and images in accordance with instructions from processor 46. Speaker 40B outputs audio in accordance with instructions from processor 46. Camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0022] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0023] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0024] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0025] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate a user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotion, including estimation and prediction of the user's emotion, but is not limited to these examples. Furthermore, the estimation and prediction of emotion also includes, for example, emotion analysis.

[0026] In the smart device 14, the specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used together with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as the control unit 46A in accordance with the specific processing program 60 executed on the RAM 48. Note that the smart device 14 has a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59, and can also perform processing similar to that of the specific processing unit 290 using these models.

[0027] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains a processing result (prediction result, etc.) using the data generation model 58 by communicating with the server device having the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device owned by a user (e.g., a mobile phone, a robot, a home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.

[0028] (Example 1) An automatic audio recording analysis system according to an embodiment of the present invention automatically records the audio of customer service, analyzes it using a generation AI, generates statistical data, and provides guidelines for business improvement. The automatic audio recording analysis system automatically records the audio of customer service, analyzes it using a generation AI, and generates statistical data to evaluate the quality of customer service and provide guidelines for business improvement. For example, the automatic audio recording analysis system automatically records the audio of customer service. For example, the automatic audio recording analysis system uses a recording device installed in a store to automatically start recording each time a customer service session is held and stop recording when the session ends. In this way, all audio data of customer service sessions is collected. Next, the automatic audio recording analysis system analyzes the recorded audio data using a generation AI. The input to the generation AI is the recorded audio data itself, and the generation AI performs analysis based on the content. For example, the generation AI receives a prompt such as "Please analyze the content of this audio data" and analyzes the content of the customer service session. Next, the automatic audio recording analysis system generates statistical data based on the data analyzed by the generation AI. Natural language processing technology is used to generate the statistical data. For example, the generation AI generates statistical data such as the most frequently used words and phrases during customer service, the most frequently asked questions, and the most frequently given answers. Next, the automatic audio recording analysis system provides guidelines for business improvement based on the generated statistical data. For example, the generation AI evaluates the quality of customer service based on the statistical data and provides specific guidelines for business improvement. This allows the automatic audio recording analysis system to evaluate the quality of customer service and provide guidelines for business improvement. For example, the automatic audio recording analysis system can analyze the words and phrases frequently used during customer service and suggest more effective ways to serve customers. It can also analyze the most frequently asked questions and answers to understand customer needs. This can improve the quality of customer service and increase business efficiency.

[0029] An automatic audio recording analysis system according to an embodiment includes a recording unit, an analysis unit, a statistics generation unit, and a guideline provision unit. The recording unit automatically records audio recorded during customer service. Examples of audio recorded during customer service include, but are not limited to, conversations in a storefront or telephone conversations. The recording unit, for example, uses a recording device installed in a storefront to automatically start recording each time a customer service session is held and stop recording when the session ends. The recording unit can also start recording when a specific keyword is detected. For example, the recording unit starts recording when a customer service session begins and stops recording when the session ends. The analysis unit uses a generation AI to analyze the audio data recorded by the recording unit. Analysis can be performed using, for example, speech recognition technology or emotion analysis technology, but is not limited to, examples. For example, the generation AI analyzes the audio data to understand the content of the customer service session. The analysis unit can also use the generation AI to analyze what words were used in the audio data, what questions were asked, and what answers were given. For example, the generation AI analyzes the audio data to extract the most frequently used words and phrases during customer service. The statistics generation unit generates statistical data based on the data analyzed by the analysis unit. The statistical data includes, for example, the most frequently used words and phrases during customer service, the most frequently asked questions, and the most frequently given answers, but is not limited to these examples. For example, the statistics generation unit uses a generation AI to tally the most frequently used words and phrases during customer service. The statistics generation unit can also use a generation AI to tally the most frequently asked questions and answers during customer service. For example, the statistics generation unit uses a generation AI to generate statistical data such as the most frequently used words and phrases during customer service, the most frequently asked questions, and the most frequently given answers. The guideline provision unit provides guidelines for business improvement based on the statistical data generated by the statistics generation unit. The guidelines include, for example, an evaluation of the quality of customer service and specific suggestions for business improvement, but are not limited to these examples. For example, the guideline provision unit uses a generation AI to evaluate the quality of customer service based on the statistical data and provide specific guidelines for business improvement. The guideline provision unit can also use a generation AI to provide guidelines including evaluation indicators such as customer satisfaction and customer service time.For example, the guideline providing unit uses a generation AI to provide guidelines including evaluation indicators such as customer satisfaction and customer service time, etc. This allows the automatic audio recording analysis system according to the embodiment to evaluate the quality of customer service and provide guidelines for business improvement.

[0030] The analysis unit may include a specific algorithm for analyzing the voice data. Specific algorithms include, but are not limited to, a voice recognition algorithm and an emotion analysis algorithm. For example, the analysis unit converts the voice data into text data using a voice recognition algorithm. The analysis unit may also estimate emotions from the voice data using an emotion analysis algorithm. For example, the analysis unit may analyze the tone and pitch of the voice data to estimate emotions. The analysis unit may also analyze the content of the voice data using a generation AI. For example, the analysis unit may use the generation AI to analyze what words were used in the voice data, what questions were asked, what answers were given, etc. This improves the accuracy of the voice data analysis. Some or all of the above-described processing in the analysis unit may be performed using, or without, the generation AI. For example, the analysis unit may input voice data to the generation AI and have the generation AI output the analysis results.

[0031] The statistics generation unit can generate statistical data on the most frequently used words or phrases, the most frequently asked questions, and the most frequently given answers during customer service. Methods for generating statistical data include, but are not limited to, frequency analysis and trend analysis. For example, the statistics generation unit can use frequency analysis to extract the most frequently used words or phrases during customer service. The statistics generation unit can also use trend analysis to extract the most frequently asked questions and answers during customer service. For example, the statistics generation unit can use a generation AI to generate statistical data on the most frequently used words or phrases, the most frequently asked questions, and the most frequently given answers during customer service. This generates detailed statistical data for evaluating the quality of customer service. Some or all of the above-described processing in the statistics generation unit can be performed using, for example, the generation AI, or can be performed without using the generation AI. For example, the statistics generation unit can input data analyzed by the analysis unit into the generation AI and have the generation AI output statistical data.

[0032] The guideline providing unit can evaluate the quality of customer service based on the generated statistical data and provide guidelines for business improvement. The guideline providing unit can evaluate the quality of customer service based on the statistical data, for example, using a generation AI. The guideline providing unit can also provide specific guidelines for business improvement using the generation AI. For example, the guideline providing unit can use the generation AI to analyze words and phrases frequently used in customer service and suggest more effective ways to provide customer service. The guideline providing unit can also use the generation AI to analyze the most frequently asked questions and answers and understand customer needs. For example, the guideline providing unit can use the generation AI to provide guidelines including evaluation indicators such as customer satisfaction and customer service time. This evaluates the quality of customer service and provides specific guidelines for business improvement. Some or all of the above-mentioned processing in the guideline providing unit can be performed, for example, using the generation AI, or can be performed without using the generation AI. For example, the guideline providing unit can input the statistical data generated by the statistics generation unit to the generation AI and cause the generation AI to output guidelines for business improvement.

[0033] The guideline providing unit may include an evaluation index for customer satisfaction or customer service time. Examples of evaluation indexes include, but are not limited to, customer satisfaction and response time. For example, the guideline providing unit uses survey results or feedback ratings to evaluate customer satisfaction. The guideline providing unit may also use average response time or longest response time to evaluate response time. For example, the guideline providing unit uses a generation AI to provide a guideline including evaluation indexes such as customer satisfaction and customer service time. This provides a guideline that takes into account evaluation indexes such as customer satisfaction and customer service time. Some or all of the above-described processing in the guideline providing unit may be performed using, for example, the generation AI, or may be performed without using the generation AI. For example, the guideline providing unit may input evaluation indexes to the generation AI and cause the generation AI to output a guideline.

[0034] The recording unit can add a function to automatically filter background noise during recording. For example, the recording unit can use a noise reduction algorithm to detect ambient noise in real time during recording and filter it using noise cancellation technology. The recording unit can also analyze background noise after recording and extract only important audio. For example, the recording unit can automatically remove noise in a specific frequency band during recording. This automatically filters out background noise, improving the quality of the recording. Some or all of the above-described processing in the recording unit can be performed using, or without, a generation AI. For example, the recording unit can input audio data acquired during recording to a generation AI and have the generation AI perform noise filtering.

[0035] The recording unit can simultaneously record multiple audio channels during recording, allowing for later separation and analysis. The recording unit, for example, uses multiple microphones to simultaneously record audio from different directions. The recording unit can also analyze the audio of each channel individually after recording and separate the audio for each speaker. For example, the recording unit separates the audio of each channel in real time during recording and saves it separately. This allows for simultaneous recording of multiple audio channels and subsequent separation and analysis. Some or all of the above-described processing in the recording unit may be performed using, or without, a generation AI. For example, the recording unit can input data of multiple audio channels into a generation AI and have the generation AI perform separation and analysis.

[0036] The recording unit can encode audio data in real time during recording, improving storage efficiency. For example, the recording unit encodes audio data in a compressed format (e.g., MP3) in real time. The recording unit can also upload audio data to cloud storage in real time during recording. For example, the recording unit divides and saves audio data during recording, improving storage efficiency. This improves storage efficiency by encoding audio data in real time. Some or all of the above-mentioned processing in the recording unit may be performed using, or without, a generation AI. For example, the recording unit can input audio data acquired during recording to a generation AI and have the generation AI perform encoding in real time.

[0037] The recording unit can add a function to emphasize a recording when a specific keyword is detected during recording. For example, if a specific keyword (e.g., "important" or "problem") is detected during recording, the recording unit emphasizes that portion and saves it. The recording unit can also automatically extract portions containing specific keywords after recording and save them as separate files. For example, if a specific keyword is detected during recording, the recording unit automatically increases the volume of that portion. In this way, by emphasizing the recording when a specific keyword is detected, important information can be emphasized and saved. Some or all of the above-described processing in the recording unit may be performed using, or without, a generation AI. For example, the recording unit can input audio data acquired during recording into a generation AI and have the generation AI detect and emphasize specific keywords.

[0038] The recording unit can be added with a function to automatically upload audio data to the cloud during recording. For example, the recording unit uploads audio data to cloud storage in real time while recording. The recording unit can also automatically back up audio data to the cloud after recording. For example, the recording unit divides audio data during recording and uploads it to the cloud to improve storage efficiency. This makes it easier to back up data by automatically uploading audio data to the cloud. Some or all of the above-mentioned processing in the recording unit may be performed using, or without, a generation AI. For example, the recording unit can input audio data acquired during recording into the generation AI and have the generation AI upload the data to the cloud.

[0039] The recording unit can be added with a function to encrypt and store audio data during recording. For example, the recording unit can encrypt audio data in real time during recording to enhance security. The recording unit can also encrypt audio data after recording and store it in cloud storage. For example, the recording unit can divide and encrypt audio data during recording to achieve both security and storage efficiency. By encrypting and storing audio data in this way, data security is improved. Some or all of the above-mentioned processing in the recording unit may be performed using, or without, a generation AI. For example, the recording unit can input audio data acquired during recording to a generation AI and have the generation AI perform encryption.

[0040] During analysis, the analysis unit understands the context of the voice data and can provide more accurate analysis results. For example, the analysis unit has the generation AI analyze the context of the voice data and understand the context. The analysis unit can also analyze the intention of the speaker of the voice data and provide analysis results based on the context. For example, the analysis unit has the generation AI analyze the context of the voice data and interpret ambiguous expressions. This understanding of the context of the voice data improves the accuracy of the analysis results. Some or all of the above-mentioned processing in the analysis unit may be performed using, or without, the generation AI. For example, the analysis unit can input voice data to the generation AI and have the generation AI understand the context and generate analysis results.

[0041] During analysis, the analysis unit can identify the speaker of the voice data and generate individual analysis results. The analysis unit identifies the speaker of the voice data, for example, by extracting voice features or using a speaker recognition algorithm. The analysis unit can also generate analysis results for each speaker using a generation AI. For example, the analysis unit identifies the speaker of the voice data using a generation AI and generates analysis results for each speaker. In this way, individual analysis results are generated by identifying the speaker of the voice data. Some or all of the above-mentioned processing in the analysis unit may be performed using, for example, a generation AI, or may be performed without using a generation AI. For example, the analysis unit can input voice data to a generation AI and have the generation AI identify the speaker and generate individual analysis results.

[0042] During analysis, the analysis unit can analyze the emotional tone of the voice data and detect changes in emotion. The analysis unit analyzes the emotional tone of the voice data, for example, by extracting features of the voice tone or using an emotion classification algorithm. The analysis unit can also detect changes in emotion using a generation AI. For example, the analysis unit analyzes the emotional tone of the voice data and detects changes in emotion using a generation AI. In this way, changes in emotion can be detected by analyzing the emotional tone of the voice data. Some or all of the above-described processing in the analysis unit may be performed using, for example, the generation AI, or may be performed without using the generation AI. For example, the analysis unit can input voice data to the generation AI and have the generation AI analyze the emotional tone and detect changes in emotion.

[0043] During analysis, the analysis unit can automatically detect the language of the voice data and apply an appropriate language model. The analysis unit can automatically detect the language of the voice data using, for example, a language identification algorithm. The analysis unit can also apply an appropriate language model using a generation AI. For example, the analysis unit can automatically detect the language of the voice data using a generation AI and apply an appropriate language model. By automatically detecting the language of the voice data, an appropriate language model is applied, improving analysis accuracy. Some or all of the above-mentioned processing in the analysis unit can be performed using, for example, a generation AI, or can be performed without using a generation AI. For example, the analysis unit can input voice data to a generation AI and have the generation AI detect the language and apply an appropriate language model.

[0044] The analysis unit can adjust the speed of the audio data during analysis to improve the accuracy of the analysis. The analysis unit can adjust the speed of the audio data using, for example, a speed adjustment algorithm. The analysis unit can also perform analysis at an appropriate speed using a generation AI. For example, the analysis unit can adjust the speed of the audio data using a generation AI to improve the accuracy of the analysis. By adjusting the speed of the audio data, the accuracy of the analysis is improved. Some or all of the above-mentioned processing in the analysis unit can be performed using, for example, the generation AI, or can be performed without using the generation AI. For example, the analysis unit can input audio data to the generation AI and have the generation AI perform speed adjustment and analysis.

[0045] During analysis, the analysis unit can analyze background sounds of the audio data and provide environmental information. The analysis unit analyzes the background sounds of the audio data, for example, by extracting features of the background sounds or using an environmental sound analysis algorithm. The analysis unit can also provide environmental information using a generation AI. For example, the analysis unit analyzes the background sounds of the audio data and provides environmental information using a generation AI. In this way, environmental information is provided by analyzing the background sounds of the audio data. Some or all of the above-described processing in the analysis unit may be performed using, for example, the generation AI, or may be performed without using the generation AI. For example, the analysis unit can input audio data to the generation AI and cause the generation AI to analyze the background sounds and provide environmental information.

[0046] The statistics generation unit can generate statistical data for different time periods and days of the week when generating statistics. The statistics generation unit generates statistical data for different time periods and days of the week, for example, using a time period or day of the week classification method or a data aggregation method. The statistics generation unit can also use a generation AI to analyze customer service data for different time periods and days of the week and generate statistical data. For example, the statistics generation unit can use a generation AI to analyze customer service patterns for different time periods and days of the week and provide statistical data. This enables detailed analysis by generating statistical data for different time periods and days of the week. Some or all of the above-described processing in the statistics generation unit may be performed using, or without, the generation AI. For example, the statistics generation unit can input customer service data into the generation AI and have the generation AI generate statistical data for each time period and day of the week.

[0047] The statistics generation unit can generate detailed statistical data including the success rate and failure rate of customer service when generating statistics. The statistics generation unit evaluates the success rate and failure rate of customer service, for example, using definitions of success and failure and evaluation methods. The statistics generation unit can also use a generation AI to analyze the success rate and failure rate of customer service and generate statistical data. For example, the statistics generation unit can use a generation AI to compare the success rate and failure rate of customer service and provide detailed statistical data. This allows the quality of customer service to be evaluated by generating detailed statistical data including the success rate and failure rate of customer service. Some or all of the above-mentioned processing in the statistics generation unit can be performed, for example, using the generation AI, or can be performed without using the generation AI. For example, the statistics generation unit can input customer service data into the generation AI and have the generation AI generate statistical data on the success rate and failure rate.

[0048] The statistics generation unit can generate statistical data based on customer attribute information (age, gender). For example, the statistics generation unit generates statistical data that takes customer attribute information into consideration using a specific type and collection method of customer attribute information. The statistics generation unit can also use a generation AI to analyze customer age information and gender information and generate statistical data. For example, the statistics generation unit can use a generation AI to comprehensively analyze customer attribute information and provide detailed statistical data. This enables more detailed analysis by generating statistical data that takes customer attribute information into consideration. Some or all of the above-mentioned processing in the statistics generation unit can be performed using, for example, a generation AI, or can be performed without using a generation AI. For example, the statistics generation unit can input customer attribute information into the generation AI and have the generation AI generate statistical data.

[0049] The statistics generation unit can generate comparative statistical data between different stores when generating statistics. The statistics generation unit generates comparative statistical data between different stores, for example, using data collection methods and comparison criteria for each store. The statistics generation unit can also use a generation AI to analyze customer service data from different stores and generate comparative statistical data. For example, the statistics generation unit uses a generation AI to compare the customer service performance of different stores and provide statistical data. This allows the performance of each store to be evaluated by generating comparative statistical data between different stores. Some or all of the above-mentioned processing in the statistics generation unit may be performed using, for example, a generation AI, or may be performed without using a generation AI. For example, the statistics generation unit can input customer service data into the generation AI and have the generation AI generate comparative statistical data between stores.

[0050] The statistics generation unit can analyze customer service trends and generate future forecast data when generating statistics. The statistics generation unit can analyze customer service trends using, for example, a trend analysis algorithm or a data collection period. The statistics generation unit can also use a generation AI to analyze past customer service data and identify trends. For example, the statistics generation unit can use a generation AI to generate future forecast data based on customer service trends. This allows future demand to be predicted by analyzing customer service trends and generating future forecast data. Some or all of the above-mentioned processing in the statistics generation unit can be performed using, for example, a generation AI, or can be performed without using a generation AI. For example, the statistics generation unit can input customer service data into the generation AI and have the generation AI analyze trends and generate future forecast data.

[0051] When generating statistics, the statistic generation unit can generate statistical data that associates a customer's purchase history with customer service data. The statistic generation unit associates a customer's purchase history with customer service data, for example, using a data matching method or an association algorithm. The statistic generation unit can also use a generation AI to analyze a customer's purchase history and generate statistical data associated with the customer service data. For example, the statistic generation unit can use a generation AI to analyze a customer's purchasing pattern and provide statistical data associated with the customer service data. This allows for more detailed analysis of customer behavior by generating statistical data that associates a customer's purchase history with the customer service data. Some or all of the above-described processing in the statistic generation unit may be performed, for example, using a generation AI, or may be performed without using a generation AI. For example, the statistic generation unit can input a customer's purchase history and customer service data into the generation AI and have the generation AI generate the associated statistical data.

[0052] When providing a guideline, the guideline providing unit can provide an optimal guideline by referring to past guideline provision history. The guideline providing unit references past guideline provision history, for example, using a history data storage method or a reference algorithm. The guideline providing unit can also use a generation AI to analyze past guideline provision history and provide an optimal guideline. For example, the guideline providing unit uses a generation AI to provide a guideline for a similar situation based on past guideline provision history. In this way, an optimal guideline is provided by referring to the past guideline provision history. Some or all of the above-described processing in the guideline providing unit may be performed using, for example, the generation AI, or may be performed without using the generation AI. For example, the guideline providing unit can input past guideline provision history into the generation AI and cause the generation AI to provide an optimal guideline.

[0053] When providing a guideline, the guideline providing unit can provide the guideline taking into consideration evaluation indexes of customer service. The guideline providing unit provides the guideline using evaluation indexes such as customer satisfaction and response time. The guideline providing unit can also use a generation AI to analyze customer satisfaction and provide optimal guidelines. For example, the guideline providing unit uses a generation AI to analyze customer service time and provide efficient guidelines. In this way, more effective guidelines are provided by taking into consideration evaluation indexes of customer service. Some or all of the above-mentioned processing in the guideline providing unit may be performed using, for example, the generation AI, or may be performed without using the generation AI. For example, the guideline providing unit can input evaluation indexes into the generation AI and have the generation AI provide the guideline.

[0054] When providing the guidelines, the guideline providing unit can provide detailed guidelines that indicate specific areas for improvement in customer service. The guideline providing unit indicates specific areas for improvement in customer service using, for example, a method for extracting and presenting areas for improvement. The guideline providing unit can also use a generation AI to analyze specific areas for improvement in customer service and provide detailed guidelines. For example, the guideline providing unit uses the generation AI to identify areas for improvement in customer service and provide a specific action plan. This enables more effective improvements by indicating specific areas for improvement in customer service. Some or all of the above-mentioned processing in the guideline providing unit may be performed using, for example, the generation AI, or may be performed without using the generation AI. For example, the guideline providing unit can input customer service data into the generation AI and cause the generation AI to extract areas for improvement and provide detailed guidelines.

[0055] When providing the guidelines, the guideline providing unit can provide customization guidelines according to different industries or business formats. The guideline providing unit provides the customization guidelines using, for example, the content and provision method of the guidelines according to the industry or business format. The guideline providing unit can also use the generation AI to analyze the characteristics of each industry and provide the customization guidelines. For example, the guideline providing unit uses the generation AI to analyze the characteristics of each business format and provide optimal guidelines. This enables more effective improvements by providing customization guidelines according to different industries or business formats. Some or all of the above-mentioned processing in the guideline providing unit may be performed using, for example, the generation AI, or may be performed without using the generation AI. For example, the guideline providing unit can input data related to the industry or business format into the generation AI and cause the generation AI to generate the customization guidelines.

[0056] The guideline providing unit can propose a customer service training program when providing the guideline. The guideline providing unit proposes a customer service training program using, for example, the content and implementation method of the training. The guideline providing unit can also use the generation AI to analyze the customer service training program and propose an optimal program. For example, the guideline providing unit uses the generation AI to analyze customer service training needs and provide a customized program. In this way, by proposing a customer service training program, customer service skills can be improved. Some or all of the above-mentioned processing in the guideline providing unit may be performed using, for example, the generation AI, or may be performed without using the generation AI. For example, the guideline providing unit can input data related to the training program into the generation AI and have the generation AI execute the optimal program proposal.

[0057] The guideline providing unit can provide guidelines that reflect customer feedback when providing guidelines. The guideline providing unit reflects customer feedback using, for example, a feedback collection method or a reflection method. The guideline providing unit can also use a generation AI to analyze customer feedback and provide optimal guidelines. For example, the guideline providing unit uses the generation AI to identify areas for improvement based on customer feedback and provide guidelines. In this way, more effective guidelines are provided by reflecting customer feedback. Some or all of the above-mentioned processing in the guideline providing unit may be performed using, for example, the generation AI, or may be performed without using the generation AI. For example, the guideline providing unit can input customer feedback data into the generation AI and cause the generation AI to reflect the feedback and provide guidelines.

[0058] The system according to the embodiment is not limited to the above-described example, and various modifications are possible, for example, as follows.

[0059] In addition to analyzing the audio data, the analysis unit can analyze background sounds in the audio data and provide environmental information. For example, the analysis unit can detect specific background sounds (e.g., car sounds, birds chirping) in the audio data and include the environmental information in the analysis results. The analysis unit can also analyze background sounds occurring in specific time periods in the audio data and provide environmental information for each time period. Furthermore, the analysis unit can analyze background sounds occurring in specific locations in the audio data and provide environmental information for each location. In this way, more detailed environmental information can be provided by analyzing the background sounds in the audio data.

[0060] In addition to analyzing the voice data, the analysis unit can identify the speaker of the voice data and generate individual analysis results. For example, the analysis unit can identify a specific speaker (e.g., a customer or a store clerk) in the voice data and provide analysis results for that speaker. The analysis unit can also analyze the emotion of a specific speaker in the voice data and provide analysis results based on that emotion. Furthermore, the analysis unit can analyze the content of a specific speaker's speech in the voice data and provide analysis results based on that content. In this way, by identifying the speaker of the voice data, more detailed analysis results can be provided.

[0061] In addition to generating statistical data, the statistics generation unit can generate statistical data for different time periods and days of the week. For example, the statistics generation unit can compile the most frequently used words and phrases during customer service by time period and provide statistical data for each time period. The statistics generation unit can also compile the most frequently asked questions and answers during customer service by day of the week and provide statistical data for each day of the week. Furthermore, the statistics generation unit can compile statistical data such as the most frequently used words and phrases, the most frequently asked questions, and the most frequently answered answers during customer service by time period and day of the week and provide detailed statistical data. In this way, generating statistical data for different time periods and days of the week enables more detailed analysis.

[0062] The guideline providing unit can evaluate the quality of customer service based on the generated statistical data and provide guidelines for business improvement. For example, the guideline providing unit uses a generation AI to evaluate the quality of customer service based on the statistical data. The guideline providing unit can also use the generation AI to provide specific guidelines for business improvement. For example, the guideline providing unit can use the generation AI to analyze words and phrases frequently used in customer service and suggest more effective ways to provide customer service. The guideline providing unit can also use the generation AI to analyze the most frequently asked questions and answers to understand customer needs. For example, the guideline providing unit uses the generation AI to provide guidelines including evaluation indicators such as customer satisfaction and customer service time. This evaluates the quality of customer service and provides specific guidelines for business improvement. Some or all of the above-mentioned processing in the guideline providing unit may be performed, for example, using the generation AI, or may be performed without using the generation AI. For example, the guideline providing unit can input the statistical data generated by the statistics generation unit to the generation AI and cause the generation AI to output guidelines for business improvement.

[0063] The guideline providing unit may include an evaluation index for customer satisfaction or customer service time. Examples of evaluation indexes include, but are not limited to, customer satisfaction and response time. For example, the guideline providing unit uses survey results or feedback ratings to evaluate customer satisfaction. The guideline providing unit may also use average response time or longest response time to evaluate response time. For example, the guideline providing unit uses a generation AI to provide a guideline including evaluation indexes such as customer satisfaction and customer service time. This provides a guideline that takes into account evaluation indexes such as customer satisfaction and customer service time. Some or all of the above-described processing in the guideline providing unit may be performed using, for example, the generation AI, or may be performed without using the generation AI. For example, the guideline providing unit may input evaluation indexes to the generation AI and cause the generation AI to output a guideline.

[0064] The recording unit can add a function to automatically filter background noise during recording. For example, the recording unit can use a noise reduction algorithm to detect ambient noise in real time during recording and filter it using noise cancellation technology. The recording unit can also analyze background noise after recording and extract only important audio. For example, the recording unit can automatically remove noise in a specific frequency band during recording. This automatically filters out background noise, improving the quality of the recording. Some or all of the above-described processing in the recording unit can be performed using, or without, a generation AI. For example, the recording unit can input audio data acquired during recording to a generation AI and have the generation AI perform noise filtering.

[0065] The recording unit can simultaneously record multiple audio channels during recording, allowing for later separation and analysis. The recording unit, for example, uses multiple microphones to simultaneously record audio from different directions. The recording unit can also analyze the audio of each channel individually after recording and separate the audio for each speaker. For example, the recording unit separates the audio of each channel in real time during recording and saves it separately. This allows for simultaneous recording of multiple audio channels and subsequent separation and analysis. Some or all of the above-described processing in the recording unit may be performed using, or without, a generation AI. For example, the recording unit can input data of multiple audio channels into a generation AI and have the generation AI perform separation and analysis.

[0066] The recording unit can encode audio data in real time during recording, improving storage efficiency. For example, the recording unit encodes audio data in a compressed format (e.g., MP3) in real time. The recording unit can also upload audio data to cloud storage in real time during recording. For example, the recording unit divides and saves audio data during recording, improving storage efficiency. This improves storage efficiency by encoding audio data in real time. Some or all of the above-mentioned processing in the recording unit may be performed using, or without, a generation AI. For example, the recording unit can input audio data acquired during recording to a generation AI and have the generation AI perform encoding in real time.

[0067] The recording unit can add a function to emphasize a recording when a specific keyword is detected during recording. For example, if a specific keyword (e.g., "important" or "problem") is detected during recording, the recording unit emphasizes that portion and saves it. The recording unit can also automatically extract portions containing specific keywords after recording and save them as separate files. For example, if a specific keyword is detected during recording, the recording unit automatically increases the volume of that portion. In this way, by emphasizing the recording when a specific keyword is detected, important information can be emphasized and saved. Some or all of the above-described processing in the recording unit may be performed using, or without, a generation AI. For example, the recording unit can input audio data acquired during recording into a generation AI and have the generation AI detect and emphasize specific keywords.

[0068] The processing flow of the first embodiment will be briefly explained below.

[0069] Step 1: The recording unit automatically records the audio of customer service. This includes conversations in the store and answering the phone. The recording unit uses a recording device installed in the store to automatically start recording each time a customer is served and stop recording when the customer service ends. It can also start recording when a specific keyword is detected. Step 2: The analysis unit uses the generation AI to analyze the voice data recorded by the recording unit. The analysis is carried out using voice recognition technology and emotion analysis technology. The analysis unit analyzes the voice data to understand the content of the customer service. It also analyzes what words were used in the voice data, what questions were asked, and what answers were given. Step 3: The statistics generation unit generates statistical data based on the data analyzed by the analysis unit. This statistical data includes the most frequently used words and phrases during customer service, the most frequently asked questions, and the most frequently given answers. The statistics generation unit aggregates this data using generation AI. Step 4: The guideline provider provides guidelines for business improvement based on the statistical data generated by the statistics generator. The guidelines include an evaluation of the quality of customer service and specific proposals for business improvement. Using a generation AI, the guideline provider provides guidelines that include evaluation indicators such as customer satisfaction and customer service time.

[0070] (Example 2) An automatic audio recording analysis system according to an embodiment of the present invention automatically records the audio of customer service, analyzes it using a generation AI, generates statistical data, and provides guidelines for business improvement. The automatic audio recording analysis system automatically records the audio of customer service, analyzes it using a generation AI, and generates statistical data to evaluate the quality of customer service and provide guidelines for business improvement. For example, the automatic audio recording analysis system automatically records the audio of customer service. For example, the automatic audio recording analysis system uses a recording device installed in a store to automatically start recording each time a customer service session is held and stop recording when the session ends. In this way, all audio data of customer service sessions is collected. Next, the automatic audio recording analysis system analyzes the recorded audio data using a generation AI. The input to the generation AI is the recorded audio data itself, and the generation AI performs analysis based on the content. For example, the generation AI receives a prompt such as "Please analyze the content of this audio data" and analyzes the content of the customer service session. Next, the automatic audio recording analysis system generates statistical data based on the data analyzed by the generation AI. Natural language processing technology is used to generate the statistical data. For example, the generation AI generates statistical data such as the most frequently used words and phrases during customer service, the most frequently asked questions, and the most frequently given answers. Next, the automatic audio recording analysis system provides guidelines for business improvement based on the generated statistical data. For example, the generation AI evaluates the quality of customer service based on the statistical data and provides specific guidelines for business improvement. This allows the automatic audio recording analysis system to evaluate the quality of customer service and provide guidelines for business improvement. For example, the automatic audio recording analysis system can analyze the words and phrases frequently used during customer service and suggest more effective ways to serve customers. It can also analyze the most frequently asked questions and answers to understand customer needs. This can improve the quality of customer service and increase business efficiency.

[0071] An automatic audio recording analysis system according to an embodiment includes a recording unit, an analysis unit, a statistics generation unit, and a guideline provision unit. The recording unit automatically records audio recorded during customer service. Examples of audio recorded during customer service include, but are not limited to, conversations in a storefront or telephone conversations. The recording unit, for example, uses a recording device installed in a storefront to automatically start recording each time a customer service session is held and stop recording when the session ends. The recording unit can also start recording when a specific keyword is detected. For example, the recording unit starts recording when a customer service session begins and stops recording when the session ends. The analysis unit uses a generation AI to analyze the audio data recorded by the recording unit. Analysis can be performed using, for example, speech recognition technology or emotion analysis technology, but is not limited to, examples. For example, the generation AI analyzes the audio data to understand the content of the customer service session. The analysis unit can also use the generation AI to analyze what words were used in the audio data, what questions were asked, and what answers were given. For example, the generation AI analyzes the audio data to extract the most frequently used words and phrases during customer service. The statistics generation unit generates statistical data based on the data analyzed by the analysis unit. The statistical data includes, for example, the most frequently used words and phrases during customer service, the most frequently asked questions, and the most frequently given answers, but is not limited to these examples. For example, the statistics generation unit uses a generation AI to tally the most frequently used words and phrases during customer service. The statistics generation unit can also use a generation AI to tally the most frequently asked questions and answers during customer service. For example, the statistics generation unit uses a generation AI to generate statistical data such as the most frequently used words and phrases during customer service, the most frequently asked questions, and the most frequently given answers. The guideline provision unit provides guidelines for business improvement based on the statistical data generated by the statistics generation unit. The guidelines include, for example, an evaluation of the quality of customer service and specific suggestions for business improvement, but are not limited to these examples. For example, the guideline provision unit uses a generation AI to evaluate the quality of customer service based on the statistical data and provide specific guidelines for business improvement. The guideline provision unit can also use a generation AI to provide guidelines including evaluation indicators such as customer satisfaction and customer service time.For example, the guideline providing unit uses a generation AI to provide guidelines including evaluation indicators such as customer satisfaction and customer service time, etc. This allows the automatic audio recording analysis system according to the embodiment to evaluate the quality of customer service and provide guidelines for business improvement.

[0072] The analysis unit may include a specific algorithm for analyzing the voice data. Specific algorithms include, but are not limited to, a voice recognition algorithm and an emotion analysis algorithm. For example, the analysis unit converts the voice data into text data using a voice recognition algorithm. The analysis unit may also estimate emotions from the voice data using an emotion analysis algorithm. For example, the analysis unit may analyze the tone and pitch of the voice data to estimate emotions. The analysis unit may also analyze the content of the voice data using a generation AI. For example, the analysis unit may use the generation AI to analyze what words were used in the voice data, what questions were asked, what answers were given, etc. This improves the accuracy of the voice data analysis. Some or all of the above-described processing in the analysis unit may be performed using, or without, the generation AI. For example, the analysis unit may input voice data to the generation AI and have the generation AI output the analysis results.

[0073] The statistics generation unit can generate statistical data on the most frequently used words or phrases, the most frequently asked questions, and the most frequently given answers during customer service. Methods for generating statistical data include, but are not limited to, frequency analysis and trend analysis. For example, the statistics generation unit can use frequency analysis to extract the most frequently used words or phrases during customer service. The statistics generation unit can also use trend analysis to extract the most frequently asked questions and answers during customer service. For example, the statistics generation unit can use a generation AI to generate statistical data on the most frequently used words or phrases, the most frequently asked questions, and the most frequently given answers during customer service. This generates detailed statistical data for evaluating the quality of customer service. Some or all of the above-described processing in the statistics generation unit can be performed using, for example, the generation AI, or can be performed without using the generation AI. For example, the statistics generation unit can input data analyzed by the analysis unit into the generation AI and have the generation AI output statistical data.

[0074] The guideline providing unit can evaluate the quality of customer service based on the generated statistical data and provide guidelines for business improvement. The guideline providing unit can evaluate the quality of customer service based on the statistical data, for example, using a generation AI. The guideline providing unit can also provide specific guidelines for business improvement using the generation AI. For example, the guideline providing unit can use the generation AI to analyze words and phrases frequently used in customer service and suggest more effective ways to provide customer service. The guideline providing unit can also use the generation AI to analyze the most frequently asked questions and answers and understand customer needs. For example, the guideline providing unit can use the generation AI to provide guidelines including evaluation indicators such as customer satisfaction and customer service time. This evaluates the quality of customer service and provides specific guidelines for business improvement. Some or all of the above-mentioned processing in the guideline providing unit can be performed, for example, using the generation AI, or can be performed without using the generation AI. For example, the guideline providing unit can input the statistical data generated by the statistics generation unit to the generation AI and cause the generation AI to output guidelines for business improvement.

[0075] The guideline providing unit may include an evaluation index for customer satisfaction or customer service time. Examples of evaluation indexes include, but are not limited to, customer satisfaction and response time. For example, the guideline providing unit uses survey results or feedback ratings to evaluate customer satisfaction. The guideline providing unit may also use average response time or longest response time to evaluate response time. For example, the guideline providing unit uses a generation AI to provide a guideline including evaluation indexes such as customer satisfaction and customer service time. This provides a guideline that takes into account evaluation indexes such as customer satisfaction and customer service time. Some or all of the above-described processing in the guideline providing unit may be performed using, for example, the generation AI, or may be performed without using the generation AI. For example, the guideline providing unit may input evaluation indexes to the generation AI and cause the generation AI to output a guideline.

[0076] The recording unit can estimate the user's emotions and adjust the start timing of recording based on the estimated user emotions. The recording unit estimates the user's emotions using, for example, voice tone analysis or facial expression recognition. For example, if the user is nervous, the recording unit may not start recording until the user is relaxed. Furthermore, if the user is excited, the recording unit may start recording immediately to avoid missing important information. For example, if the user is calm, the recording unit may start recording in accordance with the natural flow of conversation. This allows for more appropriate recording by adjusting the start timing of recording according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, an emotion engine or a generation AI. The generation AI may be, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above-described processing in the recording unit may be performed using, for example, the generation AI. For example, the recording unit may input the user's voice data into the generation AI and have the generation AI perform emotion estimation.

[0077] The recording unit can add a function to automatically filter background noise during recording. For example, the recording unit can use a noise reduction algorithm to detect ambient noise in real time during recording and filter it using noise cancellation technology. The recording unit can also analyze background noise after recording and extract only important audio. For example, the recording unit can automatically remove noise in a specific frequency band during recording. This automatically filters out background noise, improving the quality of the recording. Some or all of the above-described processing in the recording unit can be performed using, or without, a generation AI. For example, the recording unit can input audio data acquired during recording to a generation AI and have the generation AI perform noise filtering.

[0078] The recording unit can simultaneously record multiple audio channels during recording, allowing for later separation and analysis. The recording unit, for example, uses multiple microphones to simultaneously record audio from different directions. The recording unit can also analyze the audio of each channel individually after recording and separate the audio for each speaker. For example, the recording unit separates the audio of each channel in real time during recording and saves it separately. This allows for simultaneous recording of multiple audio channels and subsequent separation and analysis. Some or all of the above-described processing in the recording unit may be performed using, or without, a generation AI. For example, the recording unit can input data of multiple audio channels into a generation AI and have the generation AI perform separation and analysis.

[0079] The recording unit can encode audio data in real time during recording, improving storage efficiency. For example, the recording unit encodes audio data in a compressed format (e.g., MP3) in real time. The recording unit can also upload audio data to cloud storage in real time during recording. For example, the recording unit divides and saves audio data during recording, improving storage efficiency. This improves storage efficiency by encoding audio data in real time. Some or all of the above-mentioned processing in the recording unit may be performed using, or without, a generation AI. For example, the recording unit can input audio data acquired during recording to a generation AI and have the generation AI perform encoding in real time.

[0080] The recording unit can estimate the user's emotions and adjust the timing to stop recording based on the estimated user emotions. The recording unit estimates the user's emotions using, for example, voice tone analysis or facial expression recognition. For example, the recording unit continues recording after the user finishes speaking until the user's emotions calm down. Furthermore, if the user is excited, the recording unit can continue recording until important information is complete. For example, if the user is relaxed, the recording unit stops recording at the end of a natural conversation. This allows for more appropriate recording by adjusting the timing to stop recording according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, an emotion engine or a generation AI. The generation AI can be, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above-described processing in the recording unit can be performed using, for example, the generation AI. For example, the recording unit can input the user's voice data into the generation AI and have the generation AI perform emotion estimation.

[0081] The recording unit can add a function to emphasize a recording when a specific keyword is detected during recording. For example, if a specific keyword (e.g., "important" or "problem") is detected during recording, the recording unit emphasizes that portion and saves it. The recording unit can also automatically extract portions containing specific keywords after recording and save them as separate files. For example, if a specific keyword is detected during recording, the recording unit automatically increases the volume of that portion. In this way, by emphasizing the recording when a specific keyword is detected, important information can be emphasized and saved. Some or all of the above-described processing in the recording unit may be performed using, or without, a generation AI. For example, the recording unit can input audio data acquired during recording into a generation AI and have the generation AI detect and emphasize specific keywords.

[0082] The recording unit can be added with a function to automatically upload audio data to the cloud during recording. For example, the recording unit uploads audio data to cloud storage in real time while recording. The recording unit can also automatically back up audio data to the cloud after recording. For example, the recording unit divides audio data during recording and uploads it to the cloud to improve storage efficiency. This makes it easier to back up data by automatically uploading audio data to the cloud. Some or all of the above-mentioned processing in the recording unit may be performed using, or without, a generation AI. For example, the recording unit can input audio data acquired during recording into the generation AI and have the generation AI upload the data to the cloud.

[0083] The recording unit can be added with a function to encrypt and store audio data during recording. For example, the recording unit can encrypt audio data in real time during recording to enhance security. The recording unit can also encrypt audio data after recording and store it in cloud storage. For example, the recording unit can divide and encrypt audio data during recording to achieve both security and storage efficiency. By encrypting and storing audio data in this way, data security is improved. Some or all of the above-mentioned processing in the recording unit may be performed using, or without, a generation AI. For example, the recording unit can input audio data acquired during recording to a generation AI and have the generation AI perform encryption.

[0084] The analysis unit can estimate the user's emotions and adjust the analysis algorithm based on the estimated user emotions. The analysis unit estimates the user's emotions using, for example, voice tone analysis or facial expression recognition. For example, if the user is nervous, the analysis unit performs analysis that emphasizes emotional changes. The analysis unit can also perform detailed content analysis if the user is relaxed. For example, if the user is excited, the analysis unit prioritizes detecting important keywords. This enables more appropriate analysis by adjusting the analysis algorithm according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, an emotion engine or a generation AI. The generation AI can be, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above-mentioned processing in the analysis unit can be performed using, for example, the generation AI, or without the generation AI. For example, the analysis unit can input the user's voice data into the generation AI and have the generation AI estimate emotions and adjust the analysis algorithm.

[0085] During analysis, the analysis unit understands the context of the voice data and can provide more accurate analysis results. For example, the analysis unit has the generation AI analyze the context of the voice data and understand the context. The analysis unit can also analyze the intention of the speaker of the voice data and provide analysis results based on the context. For example, the analysis unit has the generation AI analyze the context of the voice data and interpret ambiguous expressions. This understanding of the context of the voice data improves the accuracy of the analysis results. Some or all of the above-mentioned processing in the analysis unit may be performed using, or without, the generation AI. For example, the analysis unit can input voice data to the generation AI and have the generation AI understand the context and generate analysis results.

[0086] During analysis, the analysis unit can identify the speaker of the voice data and generate individual analysis results. The analysis unit identifies the speaker of the voice data, for example, by extracting voice features or using a speaker recognition algorithm. The analysis unit can also generate analysis results for each speaker using a generation AI. For example, the analysis unit identifies the speaker of the voice data using a generation AI and generates analysis results for each speaker. In this way, individual analysis results are generated by identifying the speaker of the voice data. Some or all of the above-mentioned processing in the analysis unit may be performed using, for example, a generation AI, or may be performed without using a generation AI. For example, the analysis unit can input voice data to a generation AI and have the generation AI identify the speaker and generate individual analysis results.

[0087] During analysis, the analysis unit can analyze the emotional tone of the voice data and detect changes in emotion. The analysis unit analyzes the emotional tone of the voice data, for example, by extracting features of the voice tone or using an emotion classification algorithm. The analysis unit can also detect changes in emotion using a generation AI. For example, the analysis unit analyzes the emotional tone of the voice data and detects changes in emotion using a generation AI. In this way, changes in emotion can be detected by analyzing the emotional tone of the voice data. Some or all of the above-described processing in the analysis unit may be performed using, for example, the generation AI, or may be performed without using the generation AI. For example, the analysis unit can input voice data to the generation AI and have the generation AI analyze the emotional tone and detect changes in emotion.

[0088] The analysis unit can estimate the user's emotions and adjust the display method of the analysis results based on the estimated user emotions. The analysis unit estimates the user's emotions using, for example, voice tone analysis or facial expression recognition. For example, if the user is nervous, the analysis unit provides a simple, highly visible display method. Furthermore, if the user is relaxed, the analysis unit can provide a display method that includes detailed information. For example, if the user is in a hurry, the analysis unit provides a display method that focuses on the main points. This allows for more appropriate display by adjusting the display method of the analysis results according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, an emotion engine or a generation AI. The generation AI can be, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above-described processing in the analysis unit can be performed using, for example, the generation AI. For example, the analysis unit can input the user's voice data into the generation AI and have the generation AI estimate the emotion and adjust the display method.

[0089] During analysis, the analysis unit can automatically detect the language of the voice data and apply an appropriate language model. The analysis unit can automatically detect the language of the voice data using, for example, a language identification algorithm. The analysis unit can also apply an appropriate language model using a generation AI. For example, the analysis unit can automatically detect the language of the voice data using a generation AI and apply an appropriate language model. By automatically detecting the language of the voice data, an appropriate language model is applied, improving analysis accuracy. Some or all of the above-mentioned processing in the analysis unit can be performed using, for example, a generation AI, or can be performed without using a generation AI. For example, the analysis unit can input voice data to a generation AI and have the generation AI detect the language and apply an appropriate language model.

[0090] The analysis unit can adjust the speed of the audio data during analysis to improve the accuracy of the analysis. The analysis unit can adjust the speed of the audio data using, for example, a speed adjustment algorithm. The analysis unit can also perform analysis at an appropriate speed using a generation AI. For example, the analysis unit can adjust the speed of the audio data using a generation AI to improve the accuracy of the analysis. By adjusting the speed of the audio data, the accuracy of the analysis is improved. Some or all of the above-mentioned processing in the analysis unit can be performed using, for example, the generation AI, or can be performed without using the generation AI. For example, the analysis unit can input audio data to the generation AI and have the generation AI perform speed adjustment and analysis.

[0091] During analysis, the analysis unit can analyze background sounds of the audio data and provide environmental information. The analysis unit analyzes the background sounds of the audio data, for example, by extracting features of the background sounds or using an environmental sound analysis algorithm. The analysis unit can also provide environmental information using a generation AI. For example, the analysis unit analyzes the background sounds of the audio data and provides environmental information using a generation AI. In this way, environmental information is provided by analyzing the background sounds of the audio data. Some or all of the above-described processing in the analysis unit may be performed using, for example, the generation AI, or may be performed without using the generation AI. For example, the analysis unit can input audio data to the generation AI and cause the generation AI to analyze the background sounds and provide environmental information.

[0092] The statistics generation unit can estimate the user's emotions and adjust the statistical data generation method based on the estimated user's emotions. The statistics generation unit can estimate the user's emotions using, for example, voice tone analysis or facial expression recognition. The statistics generation unit can also adjust the statistical data generation method according to changes in emotions using a generation AI. For example, if the user is nervous, the statistics generation unit can generate statistical data that emphasizes emotional changes using the generation AI. This adjusts the statistical data generation method according to the user's emotions, thereby generating more appropriate statistical data. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generation AI. The generation AI can be, for example, a text generation AI (e.g., LLM) or a multimodal generation AI, but is not limited to such examples. Some or all of the above-mentioned processing in the statistics generation unit can be performed using, for example, the generation AI, or without the generation AI. For example, the statistics generation unit can input the user's voice data into the generation AI and have the generation AI estimate emotions and adjust the statistical data generation method.

[0093] The statistics generation unit can generate statistical data for different time periods and days of the week when generating statistics. The statistics generation unit generates statistical data for different time periods and days of the week, for example, using a time period or day of the week classification method or a data aggregation method. The statistics generation unit can also use a generation AI to analyze customer service data for different time periods and days of the week and generate statistical data. For example, the statistics generation unit can use a generation AI to analyze customer service patterns for different time periods and days of the week and provide statistical data. This enables detailed analysis by generating statistical data for different time periods and days of the week. Some or all of the above-described processing in the statistics generation unit may be performed using, or without, the generation AI. For example, the statistics generation unit can input customer service data into the generation AI and have the generation AI generate statistical data for each time period and day of the week.

[0094] The statistics generation unit can generate detailed statistical data including the success rate and failure rate of customer service when generating statistics. The statistics generation unit evaluates the success rate and failure rate of customer service, for example, using definitions of success and failure and evaluation methods. The statistics generation unit can also use a generation AI to analyze the success rate and failure rate of customer service and generate statistical data. For example, the statistics generation unit can use a generation AI to compare the success rate and failure rate of customer service and provide detailed statistical data. This allows the quality of customer service to be evaluated by generating detailed statistical data including the success rate and failure rate of customer service. Some or all of the above-mentioned processing in the statistics generation unit can be performed, for example, using the generation AI, or can be performed without using the generation AI. For example, the statistics generation unit can input customer service data into the generation AI and have the generation AI generate statistical data on the success rate and failure rate.

[0095] The statistics generation unit can generate statistical data based on customer attribute information (age, gender). For example, the statistics generation unit generates statistical data that takes customer attribute information into consideration using a specific type and collection method of customer attribute information. The statistics generation unit can also use a generation AI to analyze customer age information and gender information and generate statistical data. For example, the statistics generation unit can use a generation AI to comprehensively analyze customer attribute information and provide detailed statistical data. This enables more detailed analysis by generating statistical data that takes customer attribute information into consideration. Some or all of the above-mentioned processing in the statistics generation unit can be performed using, for example, a generation AI, or can be performed without using a generation AI. For example, the statistics generation unit can input customer attribute information into the generation AI and have the generation AI generate statistical data.

[0096] The statistics generation unit can estimate the user's emotions and adjust the display method of statistical data based on the estimated user emotions. The statistics generation unit can estimate the user's emotions using, for example, voice tone analysis or facial expression recognition. The statistics generation unit can also adjust the display method of statistical data according to changes in emotions using a generation AI. For example, if the user is nervous, the statistics generation unit can provide a simple, highly visible display method using the generation AI. This allows for more appropriate display by adjusting the display method of statistical data according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generation AI. The generation AI can be, for example, a text generation AI (e.g., LLM) or a multimodal generation AI, but is not limited to such examples. Some or all of the above-mentioned processing in the statistics generation unit can be performed using, for example, the generation AI, or without the generation AI. For example, the statistics generation unit can input the user's voice data into the generation AI and have the generation AI estimate emotions and adjust the display method.

[0097] The statistics generation unit can generate comparative statistical data between different stores when generating statistics. The statistics generation unit generates comparative statistical data between different stores, for example, using data collection methods and comparison criteria for each store. The statistics generation unit can also use a generation AI to analyze customer service data from different stores and generate comparative statistical data. For example, the statistics generation unit uses a generation AI to compare the customer service performance of different stores and provide statistical data. This allows the performance of each store to be evaluated by generating comparative statistical data between different stores. Some or all of the above-mentioned processing in the statistics generation unit may be performed using, for example, a generation AI, or may be performed without using a generation AI. For example, the statistics generation unit can input customer service data into the generation AI and have the generation AI generate comparative statistical data between stores.

[0098] The statistics generation unit can analyze customer service trends and generate future forecast data when generating statistics. The statistics generation unit can analyze customer service trends using, for example, a trend analysis algorithm or a data collection period. The statistics generation unit can also use a generation AI to analyze past customer service data and identify trends. For example, the statistics generation unit can use a generation AI to generate future forecast data based on customer service trends. This allows future demand to be predicted by analyzing customer service trends and generating future forecast data. Some or all of the above-mentioned processing in the statistics generation unit can be performed using, for example, a generation AI, or can be performed without using a generation AI. For example, the statistics generation unit can input customer service data into the generation AI and have the generation AI analyze trends and generate future forecast data.

[0099] When generating statistics, the statistic generation unit can generate statistical data that associates a customer's purchase history with customer service data. The statistic generation unit associates a customer's purchase history with customer service data, for example, using a data matching method or an association algorithm. The statistic generation unit can also use a generation AI to analyze a customer's purchase history and generate statistical data associated with the customer service data. For example, the statistic generation unit can use a generation AI to analyze a customer's purchasing pattern and provide statistical data associated with the customer service data. This allows for more detailed analysis of customer behavior by generating statistical data that associates a customer's purchase history with the customer service data. Some or all of the above-described processing in the statistic generation unit may be performed, for example, using a generation AI, or may be performed without using a generation AI. For example, the statistic generation unit can input a customer's purchase history and customer service data into the generation AI and have the generation AI generate the associated statistical data.

[0100] The guideline providing unit can estimate the user's emotion and adjust the method of providing the guideline based on the estimated user's emotion. The guideline providing unit estimates the user's emotion using, for example, voice tone analysis or facial expression recognition. The guideline providing unit can also adjust the method of providing the guideline in response to changes in emotion using a generation AI. For example, if the user is nervous, the guideline providing unit uses the generation AI to provide a simple, highly visible guideline. This adjusts the method of providing the guideline in response to the user's emotion, thereby providing a more appropriate guideline. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generation AI. The generation AI may be, for example, a text generation AI (e.g., LLM) or a multimodal generation AI, but is not limited to such examples. Some or all of the above-described processing in the guideline providing unit may be performed using, for example, the generation AI, or may be performed without using the generation AI. For example, the guideline providing unit can input the user's voice data into the generation AI and cause the generation AI to estimate the emotion and adjust the method of providing the guideline.

[0101] When providing a guideline, the guideline providing unit can provide an optimal guideline by referring to past guideline provision history. The guideline providing unit references past guideline provision history, for example, using a history data storage method or a reference algorithm. The guideline providing unit can also use a generation AI to analyze past guideline provision history and provide an optimal guideline. For example, the guideline providing unit uses a generation AI to provide a guideline for a similar situation based on past guideline provision history. In this way, an optimal guideline is provided by referring to the past guideline provision history. Some or all of the above-described processing in the guideline providing unit may be performed using, for example, the generation AI, or may be performed without using the generation AI. For example, the guideline providing unit can input past guideline provision history into the generation AI and cause the generation AI to provide an optimal guideline.

[0102] When providing a guideline, the guideline providing unit can provide the guideline taking into consideration evaluation indexes of customer service. The guideline providing unit provides the guideline using evaluation indexes such as customer satisfaction and response time. The guideline providing unit can also use a generation AI to analyze customer satisfaction and provide optimal guidelines. For example, the guideline providing unit uses a generation AI to analyze customer service time and provide efficient guidelines. In this way, more effective guidelines are provided by taking into consideration evaluation indexes of customer service. Some or all of the above-mentioned processing in the guideline providing unit may be performed using, for example, the generation AI, or may be performed without using the generation AI. For example, the guideline providing unit can input evaluation indexes into the generation AI and have the generation AI provide the guideline.

[0103] When providing the guidelines, the guideline providing unit can provide detailed guidelines that indicate specific areas for improvement in customer service. The guideline providing unit indicates specific areas for improvement in customer service using, for example, a method for extracting and presenting areas for improvement. The guideline providing unit can also use a generation AI to analyze specific areas for improvement in customer service and provide detailed guidelines. For example, the guideline providing unit uses the generation AI to identify areas for improvement in customer service and provide a specific action plan. This enables more effective improvements by indicating specific areas for improvement in customer service. Some or all of the above-mentioned processing in the guideline providing unit may be performed using, for example, the generation AI, or may be performed without using the generation AI. For example, the guideline providing unit can input customer service data into the generation AI and cause the generation AI to extract areas for improvement and provide detailed guidelines.

[0104] The guideline providing unit can estimate the user's emotions and determine the priority of guidelines based on the estimated user emotions. The guideline providing unit estimates the user's emotions using, for example, voice tone analysis or facial expression recognition. The guideline providing unit can also determine the priority of guidelines according to changes in emotions using a generation AI. For example, when the user is nervous, the guideline providing unit uses the generation AI to provide important guidelines with priority. This determines the priority of guidelines according to the user's emotions, thereby providing more effective guidelines. Emotion estimation is realized using an emotion estimation function, for example, using an emotion engine or a generation AI. The generation AI may be, for example, a text generation AI (e.g., LLM) or a multimodal generation AI, but is not limited to such examples. Some or all of the above-described processing in the guideline providing unit may be performed using, for example, the generation AI, or may be performed without using the generation AI. For example, the guideline providing unit can input the user's voice data into the generation AI and cause the generation AI to estimate emotions and determine the priority of guidelines.

[0105] When providing the guidelines, the guideline providing unit can provide customization guidelines according to different industries or business formats. The guideline providing unit provides the customization guidelines using, for example, the content and provision method of the guidelines according to the industry or business format. The guideline providing unit can also use the generation AI to analyze the characteristics of each industry and provide the customization guidelines. For example, the guideline providing unit uses the generation AI to analyze the characteristics of each business format and provide optimal guidelines. This enables more effective improvements by providing customization guidelines according to different industries or business formats. Some or all of the above-mentioned processing in the guideline providing unit may be performed using, for example, the generation AI, or may be performed without using the generation AI. For example, the guideline providing unit can input data related to the industry or business format into the generation AI and cause the generation AI to generate the customization guidelines.

[0106] The guideline providing unit can propose a customer service training program when providing the guideline. The guideline providing unit proposes a customer service training program using, for example, the content and implementation method of the training. The guideline providing unit can also use the generation AI to analyze the customer service training program and propose an optimal program. For example, the guideline providing unit uses the generation AI to analyze customer service training needs and provide a customized program. In this way, by proposing a customer service training program, customer service skills can be improved. Some or all of the above-mentioned processing in the guideline providing unit may be performed using, for example, the generation AI, or may be performed without using the generation AI. For example, the guideline providing unit can input data related to the training program into the generation AI and have the generation AI execute the optimal program proposal.

[0107] The guideline providing unit can provide guidelines that reflect customer feedback when providing guidelines. The guideline providing unit reflects customer feedback using, for example, a feedback collection method or a reflection method. The guideline providing unit can also use a generation AI to analyze customer feedback and provide optimal guidelines. For example, the guideline providing unit uses the generation AI to identify areas for improvement based on customer feedback and provide guidelines. In this way, more effective guidelines are provided by reflecting customer feedback. Some or all of the above-mentioned processing in the guideline providing unit may be performed using, for example, the generation AI, or may be performed without using the generation AI. For example, the guideline providing unit can input customer feedback data into the generation AI and cause the generation AI to reflect the feedback and provide guidelines. === Hard Collateral 1-1 === Each of the multiple elements including the recording unit, analysis unit, statistics generation unit, and guideline provision unit described above is realized, for example, by at least one of the smart device 14 and the data processing device 12. For example, the recording unit automatically records the voice of the customer using the microphone 38B of the smart device 14. The analysis unit is realized by the specific processing unit 290 of the data processing device 12 and analyzes the recorded voice data using a generation AI. The statistics generation unit is realized by the specific processing unit 290 of the data processing device 12 and generates statistical data based on the analyzed data. The guideline provision unit is realized by the specific processing unit 290 of the data processing device 12 and provides guidelines for business improvement based on the generated statistical data. === Hard Collateral 1-2 === Each of the multiple elements including the recording unit, analysis unit, statistics generation unit, and guideline provision unit described above is realized, for example, by at least one of the smart glasses 214 and the data processing device 12. For example, the recording unit automatically records the voice of the customer using the microphone 238 of the smart glasses 214. The analysis unit is realized by the specific processing unit 290 of the data processing device 12 and analyzes the recorded voice data using a generation AI. The statistics generation unit is realized by the specific processing unit 290 of the data processing device 12 and generates statistical data based on the analyzed data. The guideline provision unit is realized by the specific processing unit 290 of the data processing device 12 and provides guidelines for business improvement based on the generated statistical data. === Hard Collateral 1-3 === Each of the multiple elements including the recording unit, analysis unit, statistics generation unit, and guideline provision unit described above is realized, for example, by at least one of the headset-type terminal 314 and the data processing device 12. For example, the recording unit automatically records the voice of the customer using the microphone 238 of the headset-type terminal 314. The analysis unit is realized by the specific processing unit 290 of the data processing device 12 and analyzes the recorded voice data using a generation AI. The statistics generation unit is realized by the specific processing unit 290 of the data processing device 12 and generates statistical data based on the analyzed data. The guideline provision unit is realized by the specific processing unit 290 of the data processing device 12 and provides guidelines for business improvement based on the generated statistical data. === Hard Collateral 1-4 === Each of the multiple elements including the recording unit, analysis unit, statistics generation unit, and guideline provision unit described above is realized, for example, by at least one of the robot 414 and the data processing device 12. For example, the recording unit automatically records the voice of the customer service using the microphone 238 of the robot 414. The analysis unit is realized by the specific processing unit 290 of the data processing device 12 and analyzes the recorded voice data using a generation AI. The statistics generation unit is realized by the specific processing unit 290 of the data processing device 12 and generates statistical data based on the analyzed data. The guideline provision unit is realized by the specific processing unit 290 of the data processing device 12 and provides guidelines for business improvement based on the generated statistical data.

[0108] The system according to the embodiment is not limited to the above-described example, and various modifications are possible, for example, as follows.

[0109] In addition to analyzing the audio data, the analysis unit can analyze background sounds in the audio data and provide environmental information. For example, the analysis unit can detect specific background sounds (e.g., car sounds, birds chirping) in the audio data and include the environmental information in the analysis results. The analysis unit can also analyze background sounds occurring in specific time periods in the audio data and provide environmental information for each time period. Furthermore, the analysis unit can analyze background sounds occurring in specific locations in the audio data and provide environmental information for each location. In this way, more detailed environmental information can be provided by analyzing the background sounds in the audio data.

[0110] In addition to analyzing the voice data, the analysis unit can identify the speaker of the voice data and generate individual analysis results. For example, the analysis unit can identify a specific speaker (e.g., a customer or a store clerk) in the voice data and provide analysis results for that speaker. The analysis unit can also analyze the emotion of a specific speaker in the voice data and provide analysis results based on that emotion. Furthermore, the analysis unit can analyze the content of a specific speaker's speech in the voice data and provide analysis results based on that content. In this way, by identifying the speaker of the voice data, more detailed analysis results can be provided.

[0111] In addition to generating statistical data, the statistics generation unit can generate statistical data for different time periods and days of the week. For example, the statistics generation unit can compile the most frequently used words and phrases during customer service by time period and provide statistical data for each time period. The statistics generation unit can also compile the most frequently asked questions and answers during customer service by day of the week and provide statistical data for each day of the week. Furthermore, the statistics generation unit can compile statistical data such as the most frequently used words and phrases, the most frequently asked questions, and the most frequently answered answers during customer service by time period and day of the week and provide detailed statistical data. In this way, generating statistical data for different time periods and days of the week enables more detailed analysis.

[0112] The guideline providing unit can evaluate the quality of customer service based on the generated statistical data and provide guidelines for business improvement. For example, the guideline providing unit uses a generation AI to evaluate the quality of customer service based on the statistical data. The guideline providing unit can also use the generation AI to provide specific guidelines for business improvement. For example, the guideline providing unit can use the generation AI to analyze words and phrases frequently used in customer service and suggest more effective ways to provide customer service. The guideline providing unit can also use the generation AI to analyze the most frequently asked questions and answers to understand customer needs. For example, the guideline providing unit uses the generation AI to provide guidelines including evaluation indicators such as customer satisfaction and customer service time. This evaluates the quality of customer service and provides specific guidelines for business improvement. Some or all of the above-mentioned processing in the guideline providing unit may be performed, for example, using the generation AI, or may be performed without using the generation AI. For example, the guideline providing unit can input the statistical data generated by the statistics generation unit to the generation AI and cause the generation AI to output guidelines for business improvement.

[0113] The guideline providing unit may include an evaluation index for customer satisfaction or customer service time. Examples of evaluation indexes include, but are not limited to, customer satisfaction and response time. For example, the guideline providing unit uses survey results or feedback ratings to evaluate customer satisfaction. The guideline providing unit may also use average response time or longest response time to evaluate response time. For example, the guideline providing unit uses a generation AI to provide a guideline including evaluation indexes such as customer satisfaction and customer service time. This provides a guideline that takes into account evaluation indexes such as customer satisfaction and customer service time. Some or all of the above-described processing in the guideline providing unit may be performed using, for example, the generation AI, or may be performed without using the generation AI. For example, the guideline providing unit may input evaluation indexes to the generation AI and cause the generation AI to output a guideline.

[0114] The recording unit can estimate the user's emotions and adjust the start timing of recording based on the estimated user emotions. The recording unit estimates the user's emotions using, for example, voice tone analysis or facial expression recognition. For example, if the user is nervous, the recording unit may not start recording until the user is relaxed. Furthermore, if the user is excited, the recording unit may start recording immediately to avoid missing important information. For example, if the user is calm, the recording unit may start recording in accordance with the natural flow of conversation. This allows for more appropriate recording by adjusting the start timing of recording according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, an emotion engine or a generation AI. The generation AI may be, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above-described processing in the recording unit may be performed using, for example, the generation AI. For example, the recording unit may input the user's voice data into the generation AI and have the generation AI perform emotion estimation.

[0115] The recording unit can add a function to automatically filter background noise during recording. For example, the recording unit can use a noise reduction algorithm to detect ambient noise in real time during recording and filter it using noise cancellation technology. The recording unit can also analyze background noise after recording and extract only important audio. For example, the recording unit can automatically remove noise in a specific frequency band during recording. This automatically filters out background noise, improving the quality of the recording. Some or all of the above-described processing in the recording unit can be performed using, or without, a generation AI. For example, the recording unit can input audio data acquired during recording to a generation AI and have the generation AI perform noise filtering.

[0116] The recording unit can simultaneously record multiple audio channels during recording, allowing for later separation and analysis. The recording unit, for example, uses multiple microphones to simultaneously record audio from different directions. The recording unit can also analyze the audio of each channel individually after recording and separate the audio for each speaker. For example, the recording unit separates the audio of each channel in real time during recording and saves it separately. This allows for simultaneous recording of multiple audio channels and subsequent separation and analysis. Some or all of the above-described processing in the recording unit may be performed using, or without, a generation AI. For example, the recording unit can input data of multiple audio channels into a generation AI and have the generation AI perform separation and analysis.

[0117] The recording unit can encode audio data in real time during recording, improving storage efficiency. For example, the recording unit encodes audio data in a compressed format (e.g., MP3) in real time. The recording unit can also upload audio data to cloud storage in real time during recording. For example, the recording unit divides and saves audio data during recording, improving storage efficiency. This improves storage efficiency by encoding audio data in real time. Some or all of the above-mentioned processing in the recording unit may be performed using, or without, a generation AI. For example, the recording unit can input audio data acquired during recording to a generation AI and have the generation AI perform encoding in real time.

[0118] The recording unit can estimate the user's emotions and adjust the timing to stop recording based on the estimated user emotions. The recording unit estimates the user's emotions using, for example, voice tone analysis or facial expression recognition. For example, the recording unit continues recording after the user finishes speaking until the user's emotions calm down. Furthermore, if the user is excited, the recording unit can continue recording until important information is complete. For example, if the user is relaxed, the recording unit stops recording at the end of a natural conversation. This allows for more appropriate recording by adjusting the timing to stop recording according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, an emotion engine or a generation AI. The generation AI can be, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above-described processing in the recording unit can be performed using, for example, the generation AI. For example, the recording unit can input the user's voice data into the generation AI and have the generation AI perform emotion estimation.

[0119] The recording unit can add a function to emphasize a recording when a specific keyword is detected during recording. For example, if a specific keyword (e.g., "important" or "problem") is detected during recording, the recording unit emphasizes that portion and saves it. The recording unit can also automatically extract portions containing specific keywords after recording and save them as separate files. For example, if a specific keyword is detected during recording, the recording unit automatically increases the volume of that portion. In this way, by emphasizing the recording when a specific keyword is detected, important information can be emphasized and saved. Some or all of the above-described processing in the recording unit may be performed using, or without, a generation AI. For example, the recording unit can input audio data acquired during recording into a generation AI and have the generation AI detect and emphasize specific keywords.

[0120] The processing flow of the second embodiment will be briefly explained below.

[0121] Step 1: The recording unit automatically records the audio of customer service. This includes conversations in the store and answering the phone. The recording unit uses a recording device installed in the store to automatically start recording each time a customer is served and stop recording when the customer service ends. It can also start recording when a specific keyword is detected. Step 2: The analysis unit uses the generation AI to analyze the voice data recorded by the recording unit. The analysis is carried out using voice recognition technology and emotion analysis technology. The analysis unit analyzes the voice data to understand the content of the customer service. It also analyzes what words were used in the voice data, what questions were asked, and what answers were given. Step 3: The statistics generation unit generates statistical data based on the data analyzed by the analysis unit. This statistical data includes the most frequently used words and phrases during customer service, the most frequently asked questions, and the most frequently given answers. The statistics generation unit aggregates this data using generation AI. Step 4: The guideline provider provides guidelines for business improvement based on the statistical data generated by the statistics generator. The guidelines include an evaluation of the quality of customer service and specific proposals for business improvement. Using a generation AI, the guideline provider provides guidelines that include evaluation indicators such as customer satisfaction and customer service time.

[0122] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0123] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AIs include the data generation model 58, such as a neural network model (e.g., a neural network model), and a neural network model (e.g., a neural network model). The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating speech, text data indicating text, and image data indicating an image is also input to the data generation model 58. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specification processing unit 290 performs the above-mentioned specification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AIs other than the generative AI. The AI ​​other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. The AI ​​may also be an AI agent. When the processes of each of the above-mentioned parts are performed by AI, the processes may be performed in part or entirely by AI, but are not limited to these examples. The processes performed by AI, including the generative AI, may be replaced with rule-based processes.

[0124] Furthermore, the processing by the data processing system 10 described above is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0125] The correspondence between each part and the device or control part is not limited to the example described above, and various modifications are possible.

[0126] [Second embodiment] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0127] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0128] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN and / or a LAN.

[0129] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0130] The microphone 238 receives instructions and the like from the user by receiving voice uttered by the user. The microphone 238 captures the voice uttered by the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to instructions from the processor 46.

[0131] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the user's surroundings (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0132] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0133] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0134] The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0135] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate a user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotion, including estimation and prediction of the user's emotion, but is not limited to these examples. Furthermore, the estimation and prediction of emotion also includes, for example, emotion analysis.

[0136] In the smart glasses 214, the specific processing is performed by the processor 46. A specific processing program 60 is stored in the storage 50. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as the control unit 46A in accordance with the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0137] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 communicates with the server device having the data generation model 58 to obtain a processing result (such as a prediction result) using the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device (for example, a mobile phone, a robot, a home appliance, etc.) owned by a user.

[0138] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0139] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 receives a prompt containing an instruction, as well as inference data such as voice data representing speech, text data representing text, and image data representing an image. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The identification processing unit 290 performs the above-mentioned identification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AIs other than the generative AI. The AI ​​other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. The AI ​​may also be an AI agent. When the processes of each of the above-mentioned parts are performed by AI, the processes may be performed in part or entirely by AI, but are not limited to these examples. The processes performed by AI, including the generative AI, may be replaced with rule-based processes.

[0140] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or an external device, etc., and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or an external device, etc.

[0141] The correspondence between each part and the device or control part is not limited to the example described above, and various modifications are possible.

[0142] [Third embodiment] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0143] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0144] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN and / or a LAN.

[0145] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0146] The microphone 238 receives instructions and the like from the user by receiving voice uttered by the user. The microphone 238 captures the voice uttered by the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to instructions from the processor 46.

[0147] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the user's surroundings (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0148] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0149] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0150] The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0151] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate a user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotion, including estimation and prediction of the user's emotion, but is not limited to these examples. Furthermore, the estimation and prediction of emotion also includes, for example, emotion analysis.

[0152] In the headset type terminal 314, the identification process is performed by the processor 46. A identification program 60 is stored in the storage 50. The processor 46 reads the identification program 60 from the storage 50 and executes the read identification program 60 on the RAM 48. The identification process is realized by the processor 46 operating as a control unit 46A in accordance with the identification program 60 executed on the RAM 48. Note that the headset type terminal 314 has a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59, and can also perform processing similar to that of the identification processing unit 290 using these models.

[0153] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 communicates with the server device having the data generation model 58 to obtain a processing result (such as a prediction result) using the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device (for example, a mobile phone, a robot, a home appliance, etc.) owned by a user.

[0154] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0155] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 receives a prompt containing an instruction, as well as inference data such as voice data representing speech, text data representing text, and image data representing an image. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The identification processing unit 290 performs the above-mentioned identification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AIs other than the generative AI. The AI ​​other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. The AI ​​may also be an AI agent. When the processes of each of the above-mentioned parts are performed by AI, the processes may be performed in part or entirely by AI, but are not limited to these examples. The processes performed by AI, including the generative AI, may be replaced with rule-based processes.

[0156] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset type terminal 314, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset type terminal 314. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the headset type terminal 314 or an external device, etc., and the headset type terminal 314 acquires or collects information required for processing from the data processing device 12 or an external device, etc.

[0157] The correspondence between each part and the device or control part is not limited to the example described above, and various modifications are possible.

[0158] [Fourth embodiment] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[0159] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0160] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN and / or a LAN.

[0161] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[0162] The microphone 238 receives instructions and the like from the user by receiving voice uttered by the user. The microphone 238 captures the voice uttered by the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to instructions from the processor 46.

[0163] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS image sensor or a CCD image sensor, and captures images of the user's surroundings (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0164] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0165] The control object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[0166] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0167] The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0168] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate a user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotion, including estimation and prediction of the user's emotion, but is not limited to these examples. Furthermore, the estimation and prediction of emotion also includes, for example, emotion analysis.

[0169] In the robot 414, the processor 46 performs the identification process. The storage 50 stores the identification program 60. The processor 46 reads the identification program 60 from the storage 50 and executes the read identification program 60 on the RAM 48. The identification process is realized by the processor 46 operating as the control unit 46A in accordance with the identification program 60 executed on the RAM 48. The robot 414 also has a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59, and can perform the same process as the identification processing unit 290 using these models.

[0170] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 communicates with the server device having the data generation model 58 to obtain a processing result (such as a prediction result) using the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device (for example, a mobile phone, a robot, a home appliance, etc.) owned by a user.

[0171] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0172] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 receives a prompt containing an instruction, as well as inference data such as voice data representing speech, text data representing text, and image data representing an image. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The identification processing unit 290 performs the above-mentioned identification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AIs other than the generative AI. The AI ​​other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. The AI ​​may also be an AI agent. When the processes of each of the above-mentioned parts are performed by AI, the processes may be performed in part or entirely by AI, but are not limited to these examples. The processes performed by AI, including the generative AI, may be replaced with rule-based processes.

[0173] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or an external device, etc., and the robot 414 acquires or collects information required for processing from the data processing device 12 or an external device, etc.

[0174] The correspondence between each part and the device or control part is not limited to the example described above, and various modifications are possible.

[0175] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0176] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion encompasses both emotions and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[0177] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[0178] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[0179] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is expressed, and when they approach the ideal, a state of pleasure is expressed. Emotions can also be created for robots, cars, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is expressed, and when they approach the ideal, a state of pleasure is expressed. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on speech emotion recognition and brain physiological signal analysis systems for emotions, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the area called "reaction," where sensation is dominant. The right half of the emotion map lists emotions belonging to the area called "situation," where situational awareness is dominant.

[0180] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[0181] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[0182] In the above embodiment, an example was given in which a specific process is performed by one computer 22, but the technology disclosed herein is not limited to this, and distributed processing of the specific process may be performed by multiple computers including computer 22.

[0183] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[0184] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0185] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[0186] The hardware resource for executing a specific process can be any of the following types of processors: A processor, for example, is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. A processor also includes a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[0187] The hardware resource that executes the specific process may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific process may be a single processor.

[0188] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[0189] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[0190] In the above example, the first to fourth embodiments have been described separately, but some or all of these embodiments may be combined. The smart device 14, smart glasses 214, headset terminal 314, and robot 414 are merely examples, and they may be combined, or other devices may be used. In the above example, the first and second embodiments have been described separately, but they may be combined.

[0191] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[0192] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[0193] [Explanation of symbols]

[0194] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot

Claims

1. A recording unit that automatically records the voice of customers, an analysis unit that analyzes the voice data recorded by the recording unit; a statistics generation unit that generates statistical data based on the data analyzed by the analysis unit; a guideline providing unit that provides a guideline for business improvement based on the statistical data generated by the statistics generating unit. A system characterized by:

2. The analysis unit Contains specific algorithms for analyzing audio data 2. The system of claim 1.

3. The statistics generation unit Generate statistics on the most frequently used words or phrases, most frequently asked questions, and most frequently given answers during customer interactions 2. The system of claim 1.

4. The guideline providing unit Based on the generated statistical data, the quality of customer service is evaluated and guidelines for business improvement are provided.

2. The system of claim 1.

5. The guideline providing unit Includes metrics for customer satisfaction or service time 2. The system of claim 1.

6. The recording unit Estimate the user's emotion and adjust the start timing of recording based on the estimated user emotion.

2. The system of claim 1.

7. The recording unit Add a feature to automatically filter background noise when recording.

2. The system of claim 1.

8. The recording unit When recording, multiple audio channels can be recorded simultaneously for later analysis.

2. The system of claim 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A