system

The system addresses inconsistencies in customer service quality by analyzing audio and video data to generate training materials and feedback, enhancing staff skills and paving the way for unmanned service systems.

JP2026070906APending Publication Date: 2026-04-28SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-16
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Conventional customer service operations face challenges in maintaining consistent quality due to insufficient recording and analysis of service content, varying training among staff, and lack of specific feedback for improving customer satisfaction, hindering efficient service improvement.

Method used

A system that collects audio and video data during customer service, analyzing it using natural language processing and computer vision to calculate a customer satisfaction score, automatically extracts successful and unsuccessful examples, and generates training materials to improve service quality.

Benefits of technology

Enables objective evaluation of customer satisfaction and efficient improvement of customer service skills by providing real-time feedback and generating tailored training materials, supporting the development of unmanned customer service systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026070906000001_ABST
    Figure 2026070906000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] Using a device to collect audio and video data during customer service, A means of converting audio data from customer conversations into text using natural language processing, A method for analyzing video data to quantify customer facial expressions and movements, A means of integrating these audio and video analysis results to evaluate customer service, A means of automatically generating and providing training materials based on evaluation results, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In conventional customer service operations, there are problems that the recording and analysis of customer service content are insufficient, resulting in variations in the quality of customer response. Also, training in customer service operations varies among staff, making it difficult to maintain a consistent level of customer service. Furthermore, due to the lack of specific feedback for improving customer satisfaction, there is a problem that efficient service improvement is difficult to carry out.

Means for Solving the Problems

[0005] This invention provides a system that collects audio and video data during customer service and analyzes it using natural language processing and computer vision technologies to calculate a customer satisfaction score. Furthermore, it automatically extracts successful and unsuccessful examples of customer service based on the analysis results and generates training materials to improve the quality of customer service. This system is also intended for future application in unmanned customer service systems, enabling the generation of automated response scripts based on the collected data.

[0006] "Customer service" refers to the direct interaction and response that takes place when providing goods or services to customers.

[0007] "Audio data" refers to information recorded in digital format from conversations and other audio interactions during customer service.

[0008] "Video data" refers to information saved in digital format from visual data captured during customer service interactions.

[0009] "Natural language processing" refers to the technology that enables computers to understand, analyze, and generate human language.

[0010] "Converting to text" refers to the process of converting non-text data, such as audio, into written text.

[0011] "Quantification" means converting qualitative data into a quantitative format.

[0012] "Evaluation" refers to judging the quality of customer service and customer satisfaction based on numerical values ​​and criteria.

[0013] "Training materials" refer to learning materials and content provided for the purpose of educating staff and improving their skills.

[0014] "Automatic generation" refers to the mechanical creation of digital content based on pre-set rules and algorithms.

[0015] "Provide" means that the system delivers the information and content it generates to the user in a state where it can be utilized.

Brief Explanation of Drawings

[0016] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.

Embodiments for Carrying Out the Invention

[0017] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0020] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0021] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0024] [First Embodiment]

[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0037] This invention relates to a system that collects audio and video data in real time during customer service and analyzes it to help with staff training and improving customer satisfaction. A terminal simultaneously collects audio and video using a camera-equipped AI microphone installed in the store. When a user begins receiving customer service, the terminal automatically starts responding, recording the conversation with high accuracy and capturing the customer's facial expressions as video.

[0038] The server receives audio data transmitted from the terminal and converts it to text in real time using a natural language processing engine. The converted text is used to analyze important interactions in customer-staff conversations and extract meaning. Meanwhile, the server analyzes video data and uses computer vision algorithms to quantify customer emotions and reactions.

[0039] The server integrates the results of analyzing the collected audio and video data to calculate a customer satisfaction score. This score is used to identify areas for improvement in staff service, and further detailed analysis is performed by comparing it with past success stories. Based on this, the server automatically generates training materials and provides them to the staff. For example, if a staff member smiles infrequently, a smile training video tailored to that staff member is selected, and a link is sent to that staff member.

[0040] This system allows users to receive direct feedback, enabling them to effectively improve their customer service skills. Furthermore, in the long term, the accumulated data is expected to be used in the design and improvement of unmanned customer service systems. The collected data will also serve as a valuable resource for expanding the system to other stores and entering new markets.

[0041] The following describes the processing flow.

[0042] Step 1:

[0043] The terminal activates the camera-equipped AI microphone installed in the store and begins real-time collection of audio and video data at the customer service counter. When the sensor detects that a customer has entered the customer service area, recording and video recording automatically begin.

[0044] Step 2:

[0045] The terminal sends the collected audio data to the server. The server uses speech recognition software to convert the audio into text data. This text data is used as basic data for analyzing the content of the conversation.

[0046] Step 3:

[0047] The server receives video data transmitted from the terminal and analyzes it using computer vision algorithms. This allows it to recognize the customer's emotional state from their facial expressions and gestures, quantify them, and store them in a database.

[0048] Step 4:

[0049] The server integrates the analyzed audio and video data to calculate a customer satisfaction score. This score is used to evaluate the quality of staff service and to indicate which elements contributed to customer satisfaction.

[0050] Step 5:

[0051] The server generates specific feedback for improving customer service based on customer satisfaction scores and analysis results. Furthermore, training materials to address areas that need improvement are automatically generated and provided to staff.

[0052] Step 6:

[0053] Users will access training materials and feedback provided by the server and utilize them as a means to improve their customer service skills. They are expected to periodically attempt to improve their customer service based on the feedback they receive.

[0054] Step 7:

[0055] The server analyzes data collected over the long term to develop unmanned customer service systems and apply them to new business models. Based on the accumulated data, it designs the scenarios and automated response systems necessary for unmanned customer service.

[0056] (Example 1)

[0057] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0058] Improving the quality of traditional customer service relies on the abilities of the staff, making it difficult to identify specific areas for improvement and provide effective feedback. Furthermore, the lack of objective methods for evaluating customer satisfaction hindered efficient improvement of customer service skills. Additionally, the insufficient accumulation and utilization of data necessary for designing and improving unmanned customer service systems made the development of innovative customer service methods challenging.

[0059] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0060] In this invention, the server includes means for using equipment for collecting audio and video data to convert audio data from customer interactions into text using natural language processing, means for analyzing video data to quantify customer emotions and reactions, means for integrating these audio and video analysis results to perform customer satisfaction evaluations, means for automatically generating and providing training materials based on identified areas for improvement, and means for providing feedback for improving customer service skills and accumulating and utilizing that data. This makes it possible to objectively evaluate customer satisfaction and efficiently improve customer service skills.

[0061] "Audio and video data" refers to digital information that records the content of conversations between customers and staff during customer service, as well as their facial expressions and actions at the time.

[0062] "Natural language processing" is a technology that analyzes audio data as text information and processes human language using computers.

[0063] "Quantification" is a method of converting customer emotions and reactions obtained from video data into quantitative numerical values.

[0064] "Customer satisfaction evaluation" is a process of measuring and evaluating customer satisfaction based on the results of analyzing collected audio and video data.

[0065] "Training materials" are learning resources created based on analyzed customer service data, with the aim of improving staff customer service skills.

[0066] "Feedback" refers to information provided to staff based on customer satisfaction ratings, used to improve customer service.

[0067] This invention is a system for collecting and analyzing audio and video data with the aim of improving customer service operations. The terminal, equipped with a camera and AI microphone installed in the store, simultaneously collects audio and video during interactions with customers. When a user begins serving a customer, the terminal automatically records audio and captures the customer's facial expressions and actions. The audio data is sent to a server and converted into text data in real time using a natural language processing engine (e.g., speech recognition API).

[0068] The server analyzes the converted text data and extracts key conversation points. Video data is also sent to the server, where computer vision algorithms (e.g., image analysis software) quantify customer emotions and reactions. Integrating these analysis results, the server calculates a customer satisfaction score. This score identifies areas for improvement in staff service and allows for detailed analysis by comparing it to past success stories.

[0069] Based on the analysis results, the server automatically generates and provides training materials for staff. For example, if it is determined that a staff member is not smiling enough, a link to a smile training video tailored to that staff member will be sent. Users can receive this feedback and effectively improve their customer service skills. Furthermore, in the long term, the accumulated data can be used to design unmanned customer service systems and support their deployment to other stores.

[0070] As a concrete example, let's consider a customer service scenario in a restaurant. By executing the prompt, "Explain how AI analyzes audio and video in a restaurant customer service scenario and provides feedback to staff," we can model the system's operation and verify its effectiveness. In this way, efficient improvements to customer service operations using AI models can be achieved.

[0071] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0072] Step 1:

[0073] The terminal uses a camera-equipped AI microphone installed in the store to simultaneously collect audio and video of users and customers. Inputs include real-time audio and video data. This data is captured with high precision by the AI ​​microphone and stored in the terminal's recording device. The output consists of collected audio and video files.

[0074] Step 2:

[0075] The terminal sends the collected audio data to the server. The transmitted audio data serves as the input. This data is transferred to the server in real time using a secure protocol. The output is the audio data stored on the server.

[0076] Step 3:

[0077] The server receives the audio data and converts it to text using a natural language processing engine. The input is the audio data sent in step 2. A speech recognition algorithm is used during the text conversion process, and natural language text is obtained as output.

[0078] Step 4:

[0079] The server analyzes the converted text and extracts key conversation points. The input is the text obtained in step 3. A text analysis algorithm is applied to identify customer requests and positive elements. The output is data of the identified key conversation points.

[0080] Step 5:

[0081] The terminal transmits the collected video data to the server. The input is real-time video data. This data, like the audio data, is transmitted securely. The output is video data stored on the server.

[0082] Step 6:

[0083] The server inputs video data into a computer vision algorithm to quantify customer emotions and reactions. The input is the video data transmitted in step 5. Through image analysis, customer smiles and gestures are quantified, and the output is quantified emotion data.

[0084] Step 7:

[0085] The server integrates the audio and video analysis results to calculate a customer satisfaction score. The input is the analysis data obtained in steps 4 and 6. The integration algorithm links the audio and video results to output a single satisfaction score.

[0086] Step 8:

[0087] The server analyzes customer satisfaction scores and identifies areas for improvement in customer service. The input is the customer satisfaction score calculated in step 7. By comparing this score with past success data, specific areas for improvement are identified. The output provides information on these improvement points.

[0088] Step 9:

[0089] The server generates and provides training materials for staff based on the identified areas for improvement. The input is the information on the improvement points extracted in step 8. The training materials are automatically created using a generation AI model, and the learning materials provided to staff are linked as output.

[0090] Step 10:

[0091] Users receive feedback from the server and improve their customer service skills. The input is the training material provided in step 9. Based on the feedback, users implement specific skill improvement measures, resulting in improved customer service abilities.

[0092] (Application Example 1)

[0093] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0094] Traditional customer service often lacked immediate feedback to improve service quality, resulting in a slow development of staff skills. Furthermore, limited means of accurately understanding customer emotions made it difficult to immediately assess the appropriateness of service. Additionally, the lack of real-time feedback for staff after service interactions hindered training efficiency.

[0095] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0096] In this invention, the server includes means for converting voice information into text using natural language processing with a wearable device worn by staff interacting with customers in the store; means for analyzing video information to quantify customer emotions and reactions; means for integrating these voice and video analysis results to evaluate customer service behavior; and means for automatically generating and providing individualized training materials for staff use based on the evaluation results. This enables real-time feedback during customer service, allowing staff to immediately identify areas for improvement and efficiently improve their customer service skills.

[0097] A "wearable device" is a device that staff can wear and use, and that has the function of acquiring audio and video information.

[0098] "Audio information" refers to the content of conversations recorded during interactions between customers and staff, which are then converted into text data using natural language processing.

[0099] "Video information" refers to recording customers' facial expressions and movements, which serves as the basis for quantifying their emotions and reactions.

[0100] "Natural language processing" is a technology that converts spoken information into text information and analyzes the content of conversations with customers.

[0101] "Quantification" refers to expressing customer emotions and reactions as quantitative data based on video information.

[0102] "Evaluating customer service behavior" involves assessing staff members' customer service skills and the quality of their responses based on analyzed audio and video information.

[0103] "Training materials" refer to educational materials and information provided to staff to improve their customer service skills, and are generated based on individual areas for improvement.

[0104] "Real-time" refers to the process of collecting information and then immediately processing or providing feedback with virtually no delay.

[0105] The system for realizing this invention primarily includes a server, a wearable device worn by staff, and software for recording customer-staff interactions. The wearable device has a built-in camera and microphone to collect customer conversations and video information in real time. Furthermore, the device is designed to be lightweight and comfortable to wear, taking into account movement within the shop.

[0106] The server receives audio information transmitted from wearable devices and converts the audio into text data using a speech recognition API (e.g., Google® Cloud Speech-to-Text). This process analyzes the audio content as linguistic information and extracts key points. Simultaneously, video information is analyzed using a computer vision library (such as OpenCV), and customer expressions and behaviors are quantified as emotional scores. These analysis results are integrated and stored on the server as an evaluation of customer service behavior.

[0107] Based on the evaluation results, the server automatically generates the necessary training materials for each staff member. For example, if the analysis indicates that a staff member smiles infrequently, a video link specifically for smile training will be sent to the staff member. This feedback is displayed in real time on the staff member's smartphone app, allowing for immediate improvement. As a concrete example, if customer reactions when introducing seasonal products are analyzed, training may be provided to deepen knowledge about those products.

[0108] Furthermore, a generative AI model is utilized, with examples of prompts such as, "Please suggest what kind of training program would be effective in improving real-time responses to customer questions." This prompt allows staff to learn more advanced responses and gain specific insights to improve their customer service skills.

[0109] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0110] Step 1:

[0111] The terminal collects audio and video information in real time during interactions with customers through a wearable device worn by the staff. This information is input using the camera and microphone within the wearable device, and the data is transmitted to a server.

[0112] Step 2:

[0113] The server receives audio information from a wearable device as input and converts it into text data using a speech recognition API. The speech recognition API analyzes the audio waveform, extracts linguistic content based on this analysis, and generates output in text format.

[0114] Step 3:

[0115] The server uses a computer vision library to analyze the received video information. Taking the video information as input, it detects the customer's facial expressions and movements within each frame and outputs them as numerical data representing their emotions. It then scores whether the customer is smiling or showing interest.

[0116] Step 4:

[0117] The server integrates analysis results obtained from audio and video information to evaluate customer service behavior. Using the integrated data as input, it calculates and outputs an overall customer service evaluation score based on key interaction points and emotional scores. This score indicates which aspects of the customer service were successful and which need improvement.

[0118] Step 5:

[0119] The server generates individual training materials for each staff member based on their evaluation score. It selects the necessary educational content to address specific challenges during customer service and outputs it as a link to the staff member's terminal. This output could include, for example, videos on smile training or product description improvement.

[0120] Step 6:

[0121] Users (staff) receive and review training materials provided by the server on their smartphones or tablets. They can immediately begin learning by inputting feedback and training links. Users perform actions to solve specific problems and improve their customer service skills.

[0122] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0123] This invention provides a system that combines an emotion engine to enable a more detailed understanding of customer emotions during customer service, thereby achieving more effective customer service and staff training. This system uses a terminal equipped with a camera and AI microphone installed in the store to collect audio and video data in real time. When a customer receives customer service, the terminal automatically starts recording and video and transmits the data to a server.

[0124] The server converts audio data into text using natural language processing technology and analyzes the content of the conversation. Simultaneously, video data is analyzed using computer vision algorithms, and the customer's facial expressions and movements are quantified. Here, an emotion engine is utilized to identify the customer's emotional state based on the information obtained from their facial expressions and movements. This goes beyond mere data collection, enabling high-quality customer service evaluation that captures the customer's emotions.

[0125] The server integrates the analyzed data and performs a customer service evaluation that includes emotional data, along with a customer satisfaction score. This evaluation is used to understand the strengths and weaknesses of staff members in customer service. For example, if a particular customer was unsatisfactory, it can analyze in detail which emotions caused the dissatisfaction. Based on this, customized training materials are automatically generated for each staff member and provided to the user.

[0126] This system allows users to receive specific and emotionally sensitive feedback, which directly leads to improvements in staff customer service skills. Long-term data accumulation strengthens the foundation for developing unmanned customer service systems and designing new services. In summary, this invention contributes to improving the quality and operational efficiency of customer service operations.

[0127] The following describes the processing flow.

[0128] Step 1:

[0129] The terminal activates a camera-equipped AI microphone installed at the customer service counter, automatically starting to collect audio and video data when a customer enters the service area. Sensors detect customer movement, seamlessly recording and videotaping.

[0130] Step 2:

[0131] The terminal transfers the collected audio data to the server. The server converts the audio into text using natural language processing algorithms and analyzes the conversation content in real time. This text data is used as foundational information to extract important points of the dialogue.

[0132] Step 3:

[0133] The server receives video data transmitted from the terminal and analyzes it using computer vision technology. Specifically, it detects and quantifies the customer's facial expressions and gestures, and models the customer's behavioral patterns.

[0134] Step 4:

[0135] The server utilizes an emotion engine to identify the customer's emotional state from quantified facial expressions and gestures. This allows it to extract the customer's feelings of pleasure or displeasure during customer service and record them as data.

[0136] Step 5:

[0137] The server integrates voice-to-text, digitized facial expression data, and emotion data to calculate a customer satisfaction score. This score is used as an indicator to evaluate the quality of service and is utilized for more detailed service analysis.

[0138] Step 6:

[0139] The server automatically generates training materials based on evaluation results. It provides feedback tailored to the customer's emotional state and satisfaction level, and creates and delivers training content that includes individualized areas for improvement for staff.

[0140] Step 7:

[0141] Users receive training materials and feedback from the server and attempt to improve their customer service skills based on that information. They conduct training and adjust their customer service methods as needed.

[0142] Step 8:

[0143] The server analyzes all the data accumulated over the long term and uses it to design unmanned customer service systems and develop new customer service models. This will facilitate the implementation of advanced, data-driven customer service systems.

[0144] (Example 2)

[0145] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0146] Traditional customer service evaluation systems have struggled to accurately capture customer emotions and expressions and provide individually tailored feedback. This has made it difficult to identify specific areas for improvement in customer service quality, hindering overall customer satisfaction.

[0147] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0148] In this invention, the server includes means for using a camera to collect audio and image data, a device for converting audio data into text using natural language processing from customer interactions, a device for analyzing image data and converting customer facial expressions and actions into numerical values, a device for integrating the analysis results of these audio and image data to evaluate customer service, a device for automatically generating and providing personalized training materials based on the evaluation results, and a device for identifying the customer's emotional state using an emotion determination device. This enables detailed customer service evaluation that takes customer emotions into account and optimization of staff training based on that evaluation.

[0149] "Audio data" refers to sound information recorded from conversations and interactions with customers.

[0150] "Image data" refers to visual information that records a customer's facial expressions, actions, and other details.

[0151] "Recording device" refers to a device used to collect audio and image data.

[0152] "Natural language processing" refers to the technology of converting audio data into text and analyzing its content.

[0153] "Converting to numerical values" refers to the process of quantifying a customer's facial expressions and actions and transforming them into an evaluable format.

[0154] "Integrating analysis results" refers to combining information obtained from audio and image data to conduct a comprehensive customer service evaluation.

[0155] "Individualized training materials" refer to educational content optimized for each staff member based on evaluation results.

[0156] An "emotion determination device" refers to a device that identifies a customer's emotional state from their facial expressions and actions.

[0157] This invention provides a system that enables effective customer service and training by gaining a detailed understanding of customer emotions. The terminal uses a camera-equipped AI microphone installed in the store to collect audio and video data in real time. This device automatically starts recording and video when a customer is being served and transmits the data to a server.

[0158] The server converts received audio data into text using natural language processing (NLP) technology and analyzes the content of the conversation in detail. It also applies computer vision algorithms to video data to quantify customer facial expressions and movements. This process utilizes an emotion recognition device to identify the customer's emotional state. This series of data analyses enables high-quality customer service evaluations that reflect customer emotions.

[0159] The server integrates the analyzed data, calculates customer satisfaction scores, and performs service evaluations that include emotional data. This evaluation is used to automatically generate customized training materials for each staff member. For example, if a particular customer interaction was unsatisfactory, the server analyzes the cause in detail to identify which emotions triggered that reaction. Based on these results, personalized training materials are provided to the staff, serving as user feedback.

[0160] For example, by inputting a prompt such as "Analyze the customer's emotional state and generate customized feedback" into the AI ​​model, actual training material is created. This system allows users to receive specific and emotionally sensitive feedback, which helps improve staff customer service skills. Furthermore, long-term data accumulation strengthens the foundation for the development of unmanned customer service systems and new service designs, contributing to improvements in the overall quality and efficiency of customer service operations.

[0161] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0162] Step 1:

[0163] The terminal uses a camera-equipped AI microphone installed in the store to collect customer audio and video data in real time. Specifically, the terminal automatically starts recording and video as soon as a customer receives service. The input is the customer's audio and video, and the output is the transmission of this data to the server.

[0164] Step 2:

[0165] The server receives audio data sent from the terminal and converts it into text using natural language processing techniques. Here, a speech recognition algorithm is used, with audio data as input and text data as output. This text data forms the basis for analyzing the content of the conversation.

[0166] Step 3:

[0167] The server applies computer vision algorithms to video data to quantify the customer's facial expressions and movements. Specifically, it analyzes facial feature points and body movement patterns through image processing. The input is video data, and the output is quantified facial expression and movement data.

[0168] Step 4:

[0169] The server passes quantified facial and movement data to an emotion recognition device to identify the customer's emotional state. This process involves analysis to assign specific emotion labels (e.g., "joy" or "anxiety") to the data. The input is quantified data, and the output is emotion-labeled data.

[0170] Step 5:

[0171] The server integrates the analysis results, calculates a customer satisfaction score, and performs a customer service evaluation incorporating emotional data. Statistical methods are used to convert various data points into an overall evaluation value. Inputs are text data and emotionally labeled data, and output is a customer service evaluation score.

[0172] Step 6:

[0173] The server generates customized training materials for each staff member based on customer service evaluations and provides them to the user. In this process, a generative AI model is used to create specific training content based on prompts (for example, "Generate feedback that takes the customer's emotional state into consideration"). The input is the customer service evaluation score, and the output is the training material.

[0174] (Application Example 2)

[0175] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0176] In today's commercial environment, the quality of customer service is a crucial factor influencing customer satisfaction. However, traditional methods make it difficult to accurately grasp customer emotions and improve the immediate response capabilities of store staff. Therefore, there is a need to analyze customer emotions in real time and improve the quality of customer service based on that analysis.

[0177] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0178] In this invention, the server includes means for using a device for collecting audio and video data during customer service; means for converting audio data from customer interactions into text using natural language processing; means for analyzing video data to quantify customer expressions and actions; means for integrating these audio and video analysis results to evaluate customer service; means for automatically generating and providing training materials based on the evaluation results; and means for store employees to grasp customer emotions in real time using an intuitive information display device. This enables store employees to instantly understand customer emotions and provide appropriate customer service.

[0179] "Devices for collecting audio and video data" refer to devices installed in customer service settings to capture interactions and actions with customers, and include audio microphones and cameras.

[0180] "Means of converting to text using natural language processing" refers to software and technology for analyzing collected audio data and converting it into text format.

[0181] "Methods for analyzing video data to quantify customer expressions and movements" refers to software and algorithms that use computer vision technology to analyze video data and quantify emotions based on the customer's facial expressions and body movements.

[0182] "Means for evaluating customer service" refers to an analytical system that uses the results of integrating audio and video data to evaluate the emotional state of customers and quantify the quality of customer service.

[0183] "Means for automatically generating and providing training materials" refers to a program and process for automatically creating and providing educational materials aimed at improving the skills of store employees based on customer service evaluation results.

[0184] An "intuitive information display device" is a tool that directly visualizes information, such as smart glasses, enabling store employees to instantly grasp a customer's emotional state during service.

[0185] To implement this invention, it is necessary to build a system specifically designed for customer service. This system utilizes devices for collecting audio and video data, enabling real-time analysis of customer interactions. The server converts audio data into text using natural language processing technology, analyzes video data using computer vision, and quantifies emotional states. This allows store employees to understand customer emotions in real time and respond appropriately.

[0186] The specific system configuration involves installing devices such as cameras and microphones and connecting them to a server to collect video and audio data. The server analyzes this data and uses an advanced sentiment analysis engine to identify the customer's emotional state.

[0187] By using smart glasses as an intuitive information display device, store staff can visually confirm customers' emotions. For example, if a customer shows signs of anxiety, the smart glasses' display will show "Relaxing customer service recommended." This information provides direct guidance for staff to provide more supportive customer service.

[0188] An example of a prompt message is, "Analyze this customer's emotions and provide appropriate service during the interaction. Use the emotional data obtained from the conversation and facial expressions to adjust the service method in real time." This is how you can set instructions for the AI ​​system.

[0189] In this way, by linking the server and terminals, a system is built that enables high-quality customer service and improves customer satisfaction.

[0190] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0191] Step 1:

[0192] The terminal uses a microphone and camera installed during customer service to collect audio and video data of the customer. The collected data is sent to a server for real-time analysis. In this step, the input is the customer's audio and video, and the output is digital data sent to the server.

[0193] Step 2:

[0194] The server inputs the received audio data into a natural language processing engine and converts it into text format. This process analyzes the audio signal and outputs the transcribed conversation. The input for this step is audio data, and the output is text data.

[0195] Step 3:

[0196] The server analyzes the video data using computer vision algorithms to quantify the customer's facial expressions and movements. For example, it processes facial feature points to obtain emotion labels. The input for this step is video data, and the output is quantified emotion data.

[0197] Step 4:

[0198] The server integrates the data from steps 2 and 3 and uses an emotion engine to assess the customer's overall emotional state. This gives an understanding of how happy or anxious the customer is feeling. The inputs to this step are text data and emotion data, and the output is a detailed emotion rating score.

[0199] Step 5:

[0200] Based on the evaluation results, the server displays intuitive feedback on the employee's smart glasses. For example, it might say, "The customer needs to relax." The input for this step is the emotional evaluation score, and the output is the feedback message displayed on the smart glasses.

[0201] Step 6:

[0202] The server automatically generates training materials for store employees using a generative AI model based on long-term data accumulation. In this step, training scenarios are created based on the analysis of the accumulated data. The inputs to this step are evaluation results and historical data, and the output is the automatically generated training materials.

[0203] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0204] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search)<url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0205] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0206] [Second Embodiment]

[0207] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0208] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0209] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0210] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0211] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0212] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0213] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0214] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0215] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0216] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0217] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0218] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0219] This invention relates to a system that collects audio and video data in real time during customer service and analyzes it to help with staff training and improving customer satisfaction. A terminal simultaneously collects audio and video using a camera-equipped AI microphone installed in the store. When a user begins receiving customer service, the terminal automatically starts responding, recording the conversation with high accuracy and capturing the customer's facial expressions as video.

[0220] The server receives audio data transmitted from the terminal and converts it to text in real time using a natural language processing engine. The converted text is used to analyze important interactions in customer-staff conversations and extract meaning. Meanwhile, the server analyzes video data and uses computer vision algorithms to quantify customer emotions and reactions.

[0221] The server integrates the results of analyzing the collected audio and video data to calculate a customer satisfaction score. This score is used to identify areas for improvement in staff service, and further detailed analysis is performed by comparing it with past success stories. Based on this, the server automatically generates training materials and provides them to the staff. For example, if a staff member smiles infrequently, a smile training video tailored to that staff member is selected, and a link is sent to that staff member.

[0222] This system allows users to receive direct feedback, enabling them to effectively improve their customer service skills. Furthermore, in the long term, the accumulated data is expected to be used in the design and improvement of unmanned customer service systems. The collected data will also serve as a valuable resource for expanding the system to other stores and entering new markets.

[0223] The following describes the processing flow.

[0224] Step 1:

[0225] The terminal activates the camera-equipped AI microphone installed in the store and begins real-time collection of audio and video data at the customer service counter. When the sensor detects that a customer has entered the customer service area, recording and video recording automatically begin.

[0226] Step 2:

[0227] The terminal sends the collected audio data to the server. The server uses speech recognition software to convert the audio into text data. This text data is used as basic data for analyzing the content of the conversation.

[0228] Step 3:

[0229] The server receives video data transmitted from the terminal and analyzes it using computer vision algorithms. This allows it to recognize the customer's emotional state from their facial expressions and gestures, quantify them, and store them in a database.

[0230] Step 4:

[0231] The server integrates the analyzed audio and video data to calculate a customer satisfaction score. This score is used to evaluate the quality of staff service and to indicate which elements contributed to customer satisfaction.

[0232] Step 5:

[0233] The server generates specific feedback for improving customer service based on customer satisfaction scores and analysis results. Furthermore, training materials to address areas that need improvement are automatically generated and provided to staff.

[0234] Step 6:

[0235] Users will access training materials and feedback provided by the server and utilize them as a means to improve their customer service skills. They are expected to periodically attempt to improve their customer service based on the feedback they receive.

[0236] Step 7:

[0237] The server analyzes data collected over the long term to develop unmanned customer service systems and apply them to new business models. Based on the accumulated data, it designs the scenarios and automated response systems necessary for unmanned customer service.

[0238] (Example 1)

[0239] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0240] Improving the quality of traditional customer service relies on the abilities of the staff, making it difficult to identify specific areas for improvement and provide effective feedback. Furthermore, the lack of objective methods for evaluating customer satisfaction hindered efficient improvement of customer service skills. Additionally, the insufficient accumulation and utilization of data necessary for designing and improving unmanned customer service systems made the development of innovative customer service methods challenging.

[0241] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0242] In this invention, the server includes means for using equipment for collecting audio and video data to convert audio data from customer interactions into text using natural language processing, means for analyzing video data to quantify customer emotions and reactions, means for integrating these audio and video analysis results to perform customer satisfaction evaluations, means for automatically generating and providing training materials based on identified areas for improvement, and means for providing feedback for improving customer service skills and accumulating and utilizing that data. This makes it possible to objectively evaluate customer satisfaction and efficiently improve customer service skills.

[0243] "Audio and video data" refers to digital information that records the content of conversations between customers and staff during customer service, as well as their facial expressions and actions at the time.

[0244] "Natural language processing" is a technology that analyzes speech data as text information and processes human language using computers.

[0245] "Quantification" is a method of converting customer emotions and reactions obtained from video data into quantitative numerical values.

[0246] "Customer satisfaction evaluation" is a process of measuring and evaluating customer satisfaction based on the results of analyzing collected audio and video data.

[0247] "Training materials" are learning resources created based on analyzed customer service data, with the aim of improving staff customer service skills.

[0248] "Feedback" refers to information provided to staff based on customer satisfaction ratings, used to improve customer service.

[0249] This invention is a system for collecting and analyzing audio and video data with the aim of improving customer service operations. The terminal, equipped with a camera and AI microphone installed in the store, simultaneously collects audio and video during interactions with customers. When a user begins interacting with a customer, the terminal automatically records audio and captures the customer's facial expressions and actions. The audio data is sent to a server and converted into text data in real time using a natural language processing engine (e.g., speech recognition API).

[0250] The server analyzes the converted text data and extracts key conversation points. Video data is also sent to the server, where computer vision algorithms (e.g., image analysis software) quantify customer emotions and reactions. Integrating these analysis results, the server calculates a customer satisfaction score. This score identifies areas for improvement in staff service and allows for detailed analysis by comparing it to past success stories.

[0251] Based on the analysis results, the server automatically generates and provides training materials for staff. For example, if it is determined that a staff member is not smiling enough, a link to a smile training video tailored to that staff member will be sent. Users can receive this feedback and effectively improve their customer service skills. Furthermore, in the long term, the accumulated data can be used to design unmanned customer service systems and support their deployment to other stores.

[0252] As a concrete example, let's consider a customer service scenario in a restaurant. By executing the prompt, "Explain how AI analyzes audio and video in a restaurant customer service scenario and provides feedback to staff," we can model the system's operation and verify its effectiveness. In this way, efficient improvements to customer service operations using AI models can be achieved.

[0253] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0254] Step 1:

[0255] The terminal uses a camera-equipped AI microphone installed in the store to simultaneously collect audio and video of users and customers. Inputs include real-time audio and video data. This data is captured with high precision by the AI ​​microphone and stored in the terminal's recording device. The output consists of collected audio and video files.

[0256] Step 2:

[0257] The terminal sends the collected audio data to the server. The transmitted audio data becomes the input. This data is transferred to the server in real time using a secure protocol. The output is the audio data stored on the server.

[0258] Step 3:

[0259] The server receives the audio data and converts it to text using a natural language processing engine. The input is the audio data sent in step 2. A speech recognition algorithm is used during the text conversion process, and natural language text is obtained as output.

[0260] Step 4:

[0261] The server analyzes the converted text and extracts key conversation points. The input is the text obtained in step 3. A text analysis algorithm is applied to identify customer requests and positive elements. The output is data of the identified key conversation points.

[0262] Step 5:

[0263] The terminal transmits the collected video data to the server. The input is real-time video data. This data, like the audio data, is transmitted securely. The output is video data stored on the server.

[0264] Step 6:

[0265] The server inputs video data into a computer vision algorithm to quantify customer emotions and reactions. The input is the video data transmitted in step 5. Through image analysis, customer smiles and gestures are quantified, and the output is quantified emotion data.

[0266] Step 7:

[0267] The server integrates the audio and video analysis results to calculate a customer satisfaction score. The input is the analysis data obtained in steps 4 and 6. The integration algorithm links the audio and video results to output a single satisfaction score.

[0268] Step 8:

[0269] The server analyzes customer satisfaction scores and identifies areas for improvement in customer service. The input is the customer satisfaction score calculated in step 7. By comparing this score with past success data, specific areas for improvement are identified. The output provides information on these improvement points.

[0270] Step 9:

[0271] The server generates and provides training materials for staff based on the identified areas for improvement. The input is the information on the improvement points extracted in step 8. The training materials are automatically created using a generation AI model, and the learning materials provided to staff are linked as output.

[0272] Step 10:

[0273] Users receive feedback from the server and improve their customer service skills. The input is the training material provided in step 9. Based on the feedback, users implement specific skill improvement measures, resulting in improved customer service abilities.

[0274] (Application Example 1)

[0275] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0276] Traditional customer service often lacked immediate feedback to improve service quality, resulting in a slow development of staff skills. Furthermore, limited means of accurately understanding customer emotions made it difficult to immediately assess the appropriateness of service. Additionally, the lack of real-time feedback for staff after service interactions hindered training efficiency.

[0277] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0278] In this invention, the server includes means for converting voice information into text by natural language processing using a wearable device worn by staff who interact with customers in the store, means for analyzing video information to quantify the emotions and reactions of customers, means for integrating the results of such voice and video analysis to evaluate customer service behavior, and means for automatically generating and providing individual training materials used by the staff based on the evaluation results. This enables real-time feedback during customer service, allowing the staff to immediately know the necessary improvement points and efficiently improve their customer service skills.

[0279] A "wearable device" is a device that can be worn and used by staff and has the function of acquiring voice and video information.

[0280] "Voice information" refers to the content of conversations recorded during interactions between customers and staff, and is converted into text data by natural language processing.

[0281] "Video information" records the expressions and actions of customers and serves as a basis for quantifying the emotions and reactions of customers.

[0282] "Natural language processing" is a technology that converts voice information into text information and analyzes the content of conversations with customers.

[0283] "Quantification" means expressing the emotions and reactions of customers as quantitative data based on video information.

[0284] "Evaluation of customer service behavior" means evaluating the customer service skills and the quality of responses of staff based on the analyzed voice and video information.

[0285] "Training materials" refer to teaching materials and information provided to staff to improve their customer service skills, and are generated based on individual improvement points.

[0286] "Real-time" means that information is processed and feedback is provided almost immediately after being collected, with little delay.

[0287] The system for realizing this invention mainly includes a server, wearable devices worn by staff, and software for recording the interactions between customers and staff. The wearable devices are equipped with cameras and microphones to collect conversations and video information with customers in real time. Also, this device is designed to be lightweight and not affect the wearing sensation, taking into account movement within the store.

[0288] The server receives the voice information transmitted from the wearable device and converts the voice into text data using a voice recognition API (e.g., Google Cloud Speech-to-Text). In this process, the content of the voice is analyzed as language information, and important points are extracted. In parallel, the video information is analyzed using a computer vision library (such as OpenCV), and the expressions and actions of the customers are quantified as emotion scores. These analysis results are integrated and stored in the server as an evaluation of customer service behavior.

[0289] Based on the evaluation results, the server automatically generates the necessary training materials for each staff member. For example, if it is analyzed that there are few smiles, a video link specialized in smile training is sent to the staff. This feedback is displayed in real time on the staff's smartphone app and can lead to immediate improvement. As a specific example, there may be cases where the reactions of customers when introducing seasonal limited products are analyzed, and training is provided to deepen knowledge about those products.

[0290] Furthermore, a generative AI model is utilized, and examples of prompt texts include "Please propose what kind of educational programs are effective to improve real-time responses to customer questions." With this prompt, staff can learn more advanced responses and obtain specific insights for improving their customer service skills.

[0291] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0292] Step 1:

[0293] The terminal collects audio and video information in real time during interactions with customers through a wearable device worn by the staff. This information is input using the camera and microphone within the wearable device, and the data is transmitted to a server.

[0294] Step 2:

[0295] The server receives audio information from a wearable device as input and converts it into text data using a speech recognition API. The speech recognition API analyzes the audio waveform, extracts linguistic content based on this analysis, and generates output in text format.

[0296] Step 3:

[0297] The server uses a computer vision library to analyze the received video information. Taking the video information as input, it detects the customer's facial expressions and movements within each frame and outputs them as numerical data representing their emotions. It then scores whether the customer is smiling or showing interest.

[0298] Step 4:

[0299] The server integrates analysis results obtained from audio and video information to evaluate customer service behavior. Using the integrated data as input, it calculates and outputs an overall customer service evaluation score based on key interaction points and emotional scores. This score indicates which aspects of the customer service were successful and which need improvement.

[0300] Step 5:

[0301] The server generates individual training materials for each staff member based on their evaluation score. It selects the necessary educational content to address specific challenges during customer service and outputs it as a link to the staff member's terminal. This output could include, for example, videos on smile training or product description improvement.

[0302] Step 6:

[0303] Users (staff) receive and review training materials provided by the server on their smartphones or tablets. They can immediately begin learning by inputting feedback and training links. Users perform actions to solve specific problems and improve their customer service skills.

[0304] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0305] This invention provides a system that combines an emotion engine to enable a more detailed understanding of customer emotions during customer service, thereby achieving more effective customer service and staff training. This system uses a terminal equipped with a camera and AI microphone installed in the store to collect audio and video data in real time. When a customer receives customer service, the terminal automatically starts recording and video and transmits the data to a server.

[0306] The server converts audio data into text using natural language processing technology and analyzes the content of the conversation. Simultaneously, video data is analyzed using computer vision algorithms, and the customer's facial expressions and movements are quantified. Here, an emotion engine is utilized to identify the customer's emotional state based on the information obtained from their facial expressions and movements. This goes beyond mere data collection, enabling high-quality customer service evaluation that captures the customer's emotions.

[0307] The server integrates the analyzed data and conducts a customer service evaluation that includes sentiment data, along with the customer satisfaction score. This evaluation is used to understand the strengths and weaknesses of the staff in customer service. For example, if the response to a specific customer is unsatisfactory, it is possible to analyze in detail which sentiment caused the problem. Based on this, training materials customized for each staff member are automatically generated and provided to the user.

[0308] With this system, the user can receive specific and sentiment - considerate feedback, which directly leads to the improvement of the staff's customer service skills. The long - term data accumulation strengthens the foundation for the development of unmanned customer service systems and new service designs. With this content, the present invention contributes to the improvement of the quality of customer service operations and operational efficiency.

[0309] The following describes the processing flow.

[0310] Step 1:

[0311] The terminal activates the camera - equipped AI microphone installed at the customer service counter and automatically starts collecting audio and video data when a customer enters the customer service area. The sensor detects the movement of the customer and seamlessly records the audio and video.

[0312] Step 2:

[0313] The terminal transfers the collected audio data to the server. The server converts the audio into text using natural language processing algorithms and analyzes the conversation content in real - time. This text data is used as basic information for extracting important points of the conversation.

[0314] Step 3:

[0315] The server receives video data transmitted from the terminal and analyzes it using computer vision technology. Specifically, it detects and quantifies the customer's facial expressions and gestures, and models the customer's behavioral patterns.

[0316] Step 4:

[0317] The server utilizes an emotion engine to identify the customer's emotional state from quantified facial expressions and gestures. This allows it to extract the customer's feelings of pleasure or displeasure during customer service and record them as data.

[0318] Step 5:

[0319] The server integrates voice-to-text, digitized facial expression data, and emotion data to calculate a customer satisfaction score. This score is used as an indicator to evaluate the quality of service and is utilized for more detailed service analysis.

[0320] Step 6:

[0321] The server automatically generates training materials based on evaluation results. It provides feedback tailored to the customer's emotional state and satisfaction level, and creates and delivers training content that includes individualized areas for improvement for staff.

[0322] Step 7:

[0323] Users receive training materials and feedback from the server and attempt to improve their customer service skills based on that information. They conduct training and adjust their customer service methods as needed.

[0324] Step 8:

[0325] The server analyzes all the data accumulated over the long term and uses it to design unmanned customer service systems and develop new customer service models. This will facilitate the implementation of advanced, data-driven customer service systems.

[0326] (Example 2)

[0327] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0328] Traditional customer service evaluation systems have struggled to accurately capture customer emotions and expressions and provide individually tailored feedback. This has made it difficult to identify specific areas for improvement in customer service quality, hindering overall customer satisfaction.

[0329] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0330] In this invention, the server includes means for using a camera to collect audio and image data, a device for converting audio data into text using natural language processing from customer interactions, a device for analyzing image data and converting customer facial expressions and actions into numerical values, a device for integrating the analysis results of these audio and image data to evaluate customer service, a device for automatically generating and providing personalized training materials based on the evaluation results, and a device for identifying the customer's emotional state using an emotion determination device. This enables detailed customer service evaluation that takes customer emotions into account and optimization of staff training based on that evaluation.

[0331] "Audio data" refers to sound information recorded from conversations and interactions with customers.

[0332] "Image data" refers to visual information that records a customer's facial expressions, actions, and other details.

[0333] "Recording device" refers to a device used to collect audio and image data.

[0334] "Natural language processing" refers to the technology of converting audio data into text and analyzing its content.

[0335] "Converting to numerical values" refers to the process of quantifying a customer's facial expressions and actions and transforming them into an evaluable format.

[0336] "Integrating analysis results" refers to combining information obtained from audio and image data to conduct a comprehensive customer service evaluation.

[0337] "Individualized training materials" refer to educational content optimized for each staff member based on evaluation results.

[0338] An "emotion determination device" refers to a device that identifies a customer's emotional state from their facial expressions and actions.

[0339] This invention provides a system that enables effective customer service and training by gaining a detailed understanding of customer emotions. The terminal uses a camera-equipped AI microphone installed in the store to collect audio and video data in real time. This device automatically starts recording and video when a customer is being served and transmits the data to a server.

[0340] The server converts received audio data into text using natural language processing (NLP) technology and analyzes the content of the conversation in detail. It also applies computer vision algorithms to video data to quantify customer facial expressions and movements. This process utilizes an emotion recognition device to identify the customer's emotional state. This series of data analyses enables high-quality customer service evaluations that reflect customer emotions.

[0341] The server integrates the analyzed data, calculates customer satisfaction scores, and performs service evaluations that include emotional data. This evaluation is used to automatically generate customized training materials for each staff member. For example, if a particular customer interaction was unsatisfactory, the server analyzes the cause in detail to identify which emotions triggered that reaction. Based on these results, personalized training materials are provided to the staff, serving as user feedback.

[0342] For example, by inputting a prompt such as "Analyze the customer's emotional state and generate customized feedback" into the AI ​​model, actual training material is created. This system allows users to receive specific and emotionally sensitive feedback, which helps improve staff customer service skills. Furthermore, long-term data accumulation strengthens the foundation for the development of unmanned customer service systems and new service designs, contributing to improvements in the overall quality and efficiency of customer service operations.

[0343] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0344] Step 1:

[0345] The terminal uses a camera-equipped AI microphone installed in the store to collect customer audio and video data in real time. Specifically, the terminal automatically starts recording and video as soon as a customer receives service. The input is the customer's audio and video, and the output is the transmission of this data to the server.

[0346] Step 2:

[0347] The server receives audio data sent from the terminal and converts it into text using natural language processing techniques. Here, a speech recognition algorithm is used, with audio data as input and text data as output. This text data forms the basis for analyzing the content of the conversation.

[0348] Step 3:

[0349] The server applies computer vision algorithms to video data to quantify the customer's facial expressions and movements. Specifically, it analyzes facial feature points and body movement patterns through image processing. The input is video data, and the output is quantified facial expression and movement data.

[0350] Step 4:

[0351] The server passes quantified facial and movement data to an emotion recognition device to identify the customer's emotional state. This process involves analysis to assign specific emotion labels (e.g., "joy" or "anxiety") to the data. The input is quantified data, and the output is emotion-labeled data.

[0352] Step 5:

[0353] The server integrates the analysis results, calculates a customer satisfaction score, and performs a customer service evaluation incorporating emotional data. Statistical methods are used to convert various data points into an overall evaluation value. Inputs are text data and emotionally labeled data, and output is a customer service evaluation score.

[0354] Step 6:

[0355] The server generates customized training materials for each staff member based on customer service evaluations and provides them to the user. In this process, a generative AI model is used to create specific training content based on prompts (for example, "Generate feedback that takes the customer's emotional state into consideration"). The input is the customer service evaluation score, and the output is the training material.

[0356] (Application Example 2)

[0357] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0358] In today's commercial environment, the quality of customer service is a crucial factor influencing customer satisfaction. However, traditional methods make it difficult to accurately grasp customer emotions and improve the immediate response capabilities of store staff. Therefore, there is a need to analyze customer emotions in real time and improve the quality of customer service based on that analysis.

[0359] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0360] In this invention, the server includes means for using a device for collecting audio and video data during customer service; means for converting audio data from customer interactions into text using natural language processing; means for analyzing video data to quantify customer expressions and actions; means for integrating these audio and video analysis results to evaluate customer service; means for automatically generating and providing training materials based on the evaluation results; and means for store employees to grasp customer emotions in real time using an intuitive information display device. This enables store employees to instantly understand customer emotions and provide appropriate customer service.

[0361] "Devices for collecting audio and video data" refer to devices installed in customer service settings to capture interactions and actions with customers, and include audio microphones and cameras.

[0362] "Means of converting to text using natural language processing" refers to software and technology for analyzing collected audio data and converting it into text format.

[0363] "Methods for analyzing video data to quantify customer expressions and movements" refers to software and algorithms that use computer vision technology to analyze video data and quantify emotions based on the customer's facial expressions and body movements.

[0364] "Means for evaluating customer service" refers to an analytical system that uses the results of integrating audio and video data to evaluate the emotional state of customers and quantify the quality of customer service.

[0365] "Means for automatically generating and providing training materials" refers to a program and process for automatically creating and providing educational materials aimed at improving the skills of store employees based on customer service evaluation results.

[0366] An "intuitive information display device" is a tool that directly visualizes information, such as smart glasses, enabling store employees to instantly grasp a customer's emotional state during service.

[0367] To implement this invention, it is necessary to build a system specifically designed for customer service. This system utilizes devices for collecting audio and video data, enabling real-time analysis of customer interactions. The server converts audio data into text using natural language processing technology, analyzes video data using computer vision, and quantifies emotional states. This allows store employees to understand customer emotions in real time and respond appropriately.

[0368] The specific system configuration involves installing devices such as cameras and microphones and connecting them to a server to collect video and audio data. The server analyzes this data and uses an advanced sentiment analysis engine to identify the customer's emotional state.

[0369] By using smart glasses as an intuitive information display device, store staff can visually confirm customers' emotions. For example, if a customer shows signs of anxiety, the smart glasses' display will show "Relaxing customer service recommended." This information provides direct guidance for staff to provide more supportive customer service.

[0370] An example of a prompt message is, "Analyze this customer's emotions and provide appropriate service during the interaction. Use the emotional data obtained from the conversation and facial expressions to adjust the service method in real time." This is how you can set instructions for the AI ​​system.

[0371] In this way, by linking the server and terminals, a system is built that enables high-quality customer service and improves customer satisfaction.

[0372] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0373] Step 1:

[0374] The terminal uses a microphone and camera installed during customer service to collect audio and video data of the customer. The collected data is sent to a server for real-time analysis. In this step, the input is the customer's audio and video, and the output is digital data sent to the server.

[0375] Step 2:

[0376] The server inputs the received audio data into a natural language processing engine and converts it into text format. This process analyzes the audio signal and outputs the transcribed conversation. The input for this step is audio data, and the output is text data.

[0377] Step 3:

[0378] The server analyzes the video data using computer vision algorithms to quantify the customer's facial expressions and movements. For example, it processes facial feature points to obtain emotion labels. The input for this step is video data, and the output is quantified emotion data.

[0379] Step 4:

[0380] The server integrates the data from steps 2 and 3 and uses an emotion engine to assess the customer's overall emotional state. This gives an understanding of how happy or anxious the customer is feeling. The inputs to this step are text data and emotion data, and the output is a detailed emotion rating score.

[0381] Step 5:

[0382] Based on the evaluation results, the server displays intuitive feedback on the employee's smart glasses. For example, it might say, "The customer needs to relax." The input for this step is the emotional evaluation score, and the output is the feedback message displayed on the smart glasses.

[0383] Step 6:

[0384] The server automatically generates training materials for store employees using a generative AI model based on long-term data accumulation. In this step, training scenarios are created based on the analysis of the accumulated data. The inputs to this step are evaluation results and historical data, and the output is the automatically generated training materials.

[0385] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0386] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0387] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0388] [Third Embodiment]

[0389] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0390] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0391] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0392] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0393] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0394] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0395] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0396] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0397] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0398] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0399] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0400] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0401] This invention relates to a system that collects audio and video data in real time during customer service and analyzes it to help with staff training and improving customer satisfaction. A terminal simultaneously collects audio and video using a camera-equipped AI microphone installed in the store. When a user begins receiving customer service, the terminal automatically starts responding, recording the conversation with high accuracy and capturing the customer's facial expressions as video.

[0402] The server receives audio data transmitted from the terminal and converts it to text in real time using a natural language processing engine. The converted text is used to analyze important interactions in customer-staff conversations and extract meaning. Meanwhile, the server analyzes video data and uses computer vision algorithms to quantify customer emotions and reactions.

[0403] The server integrates the results of analyzing the collected audio and video data to calculate a customer satisfaction score. This score is used to identify areas for improvement in staff service, and further detailed analysis is performed by comparing it with past success stories. Based on this, the server automatically generates training materials and provides them to the staff. For example, if a staff member smiles infrequently, a smile training video tailored to that staff member is selected, and a link is sent to that staff member.

[0404] This system allows users to receive direct feedback, enabling them to effectively improve their customer service skills. Furthermore, in the long term, the accumulated data is expected to be used in the design and improvement of unmanned customer service systems. The collected data will also serve as a valuable resource for expanding the system to other stores and entering new markets.

[0405] The following describes the processing flow.

[0406] Step 1:

[0407] The terminal activates the camera-equipped AI microphone installed in the store and begins real-time collection of audio and video data at the customer service counter. When the sensor detects that a customer has entered the customer service area, recording and video recording automatically begin.

[0408] Step 2:

[0409] The terminal sends the collected audio data to the server. The server uses speech recognition software to convert the audio into text data. This text data is used as basic data for analyzing the content of the conversation.

[0410] Step 3:

[0411] The server receives video data transmitted from the terminal and analyzes it using computer vision algorithms. This allows it to recognize the customer's emotional state from their facial expressions and gestures, quantify them, and store them in a database.

[0412] Step 4:

[0413] The server integrates the analyzed audio and video data to calculate a customer satisfaction score. This score is used to evaluate the quality of staff service and to indicate which elements contributed to customer satisfaction.

[0414] Step 5:

[0415] The server generates specific feedback for improving customer service based on customer satisfaction scores and analysis results. Furthermore, training materials to address areas that need improvement are automatically generated and provided to staff.

[0416] Step 6:

[0417] Users will access training materials and feedback provided by the server and utilize them as a means to improve their customer service skills. They are expected to periodically attempt to improve their customer service based on the feedback they receive.

[0418] Step 7:

[0419] The server analyzes data collected over the long term to develop unmanned customer service systems and apply them to new business models. Based on the accumulated data, it designs the scenarios and automated response systems necessary for unmanned customer service.

[0420] (Example 1)

[0421] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0422] Improving the quality of traditional customer service relies on the abilities of the staff, making it difficult to identify specific areas for improvement and provide effective feedback. Furthermore, the lack of objective methods for evaluating customer satisfaction hindered efficient improvement of customer service skills. Additionally, the insufficient accumulation and utilization of data necessary for designing and improving unmanned customer service systems made the development of innovative customer service methods challenging.

[0423] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0424] In this invention, the server includes means for using equipment for collecting audio and video data to convert audio data from customer interactions into text using natural language processing, means for analyzing video data to quantify customer emotions and reactions, means for integrating these audio and video analysis results to perform customer satisfaction evaluations, means for automatically generating and providing training materials based on identified areas for improvement, and means for providing feedback for improving customer service skills and accumulating and utilizing that data. This makes it possible to objectively evaluate customer satisfaction and efficiently improve customer service skills.

[0425] "Audio and video data" refers to digital information that records the content of conversations between customers and staff during customer service, as well as their facial expressions and actions at the time.

[0426] "Natural language processing" is a technology that analyzes speech data as text information and processes human language using computers.

[0427] "Quantification" is a method of converting customer emotions and reactions obtained from video data into quantitative numerical values.

[0428] "Customer satisfaction evaluation" is a process of measuring and evaluating customer satisfaction based on the results of analyzing collected audio and video data.

[0429] "Training materials" are learning resources created based on analyzed customer service data, with the aim of improving staff customer service skills.

[0430] "Feedback" refers to information provided to staff based on customer satisfaction ratings, used to improve customer service.

[0431] This invention is a system for collecting and analyzing audio and video data with the aim of improving customer service operations. The terminal, equipped with a camera and AI microphone installed in the store, simultaneously collects audio and video during interactions with customers. When a user begins interacting with a customer, the terminal automatically records audio and captures the customer's facial expressions and actions. The audio data is sent to a server and converted into text data in real time using a natural language processing engine (e.g., speech recognition API).

[0432] The server analyzes the converted text data and extracts key conversation points. Video data is also sent to the server, where computer vision algorithms (e.g., image analysis software) quantify customer emotions and reactions. Integrating these analysis results, the server calculates a customer satisfaction score. This score identifies areas for improvement in staff service and allows for detailed analysis by comparing it to past success stories.

[0433] Based on the analysis results, the server automatically generates and provides training materials for staff. For example, if it is determined that a staff member is not smiling enough, a link to a smile training video tailored to that staff member will be sent. Users can receive this feedback and effectively improve their customer service skills. Furthermore, in the long term, the accumulated data can be used to design unmanned customer service systems and support their deployment to other stores.

[0434] As a concrete example, let's consider a customer service scenario in a restaurant. By executing the prompt, "Explain how AI analyzes audio and video in a restaurant customer service scenario and provides feedback to staff," we can model the system's operation and verify its effectiveness. In this way, efficient improvements to customer service operations using AI models can be achieved.

[0435] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0436] Step 1:

[0437] The terminal uses a camera-equipped AI microphone installed in the store to simultaneously collect audio and video of users and customers. Inputs include real-time audio and video data. This data is captured with high precision by the AI ​​microphone and stored in the terminal's recording device. The output consists of collected audio and video files.

[0438] Step 2:

[0439] The terminal sends the collected audio data to the server. The transmitted audio data becomes the input. This data is transferred to the server in real time using a secure protocol. The output is the audio data stored on the server.

[0440] Step 3:

[0441] The server receives the audio data and converts it to text using a natural language processing engine. The input is the audio data sent in step 2. A speech recognition algorithm is used during the text conversion process, and natural language text is obtained as output.

[0442] Step 4:

[0443] The server analyzes the converted text and extracts key conversation points. The input is the text obtained in step 3. A text analysis algorithm is applied to identify customer requests and positive elements. The output is data of the identified key conversation points.

[0444] Step 5:

[0445] The terminal transmits the collected video data to the server. The input is real-time video data. This data, like the audio data, is transmitted securely. The output is video data stored on the server.

[0446] Step 6:

[0447] The server inputs video data into a computer vision algorithm to quantify customer emotions and reactions. The input is the video data transmitted in step 5. Through image analysis, customer smiles and gestures are quantified, and the output is quantified emotion data.

[0448] Step 7:

[0449] The server integrates the audio and video analysis results to calculate a customer satisfaction score. The input is the analysis data obtained in steps 4 and 6. The integration algorithm links the audio and video results to output a single satisfaction score.

[0450] Step 8:

[0451] The server analyzes customer satisfaction scores and identifies areas for improvement in customer service. The input is the customer satisfaction score calculated in step 7. By comparing this score with past success data, specific areas for improvement are identified. The output provides information on these improvement points.

[0452] Step 9:

[0453] The server generates and provides training materials for staff based on the identified areas for improvement. The input is the information on the improvement points extracted in step 8. The training materials are automatically created using a generation AI model, and the learning materials provided to staff are linked as output.

[0454] Step 10:

[0455] Users receive feedback from the server and improve their customer service skills. The input is the training material provided in step 9. Based on the feedback, users implement specific skill improvement measures, resulting in improved customer service abilities.

[0456] (Application Example 1)

[0457] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0458] Traditional customer service often lacked immediate feedback to improve service quality, resulting in a slow development of staff skills. Furthermore, limited means of accurately understanding customer emotions made it difficult to immediately assess the appropriateness of service. Additionally, the lack of real-time feedback for staff after service interactions hindered training efficiency.

[0459] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0460] In this invention, the server includes means for converting voice information into text using natural language processing with a wearable device worn by staff interacting with customers in the store; means for analyzing video information to quantify customer emotions and reactions; means for integrating these voice and video analysis results to evaluate customer service behavior; and means for automatically generating and providing individualized training materials for staff use based on the evaluation results. This enables real-time feedback during customer service, allowing staff to immediately identify areas for improvement and efficiently improve their customer service skills.

[0461] A "wearable device" is a device that staff can wear and use, and that has the function of acquiring audio and video information.

[0462] "Audio information" refers to the content of conversations recorded during interactions between customers and staff, which are then converted into text data using natural language processing.

[0463] "Video information" refers to recording customers' facial expressions and movements, which serves as the basis for quantifying their emotions and reactions.

[0464] "Natural language processing" is a technology that converts spoken information into text information and analyzes the content of conversations with customers.

[0465] "Quantification" refers to expressing customer emotions and reactions as quantitative data based on video information.

[0466] "Evaluating customer service behavior" involves assessing staff members' customer service skills and the quality of their responses based on analyzed audio and video information.

[0467] "Training materials" refer to educational materials and information provided to staff to improve their customer service skills, and are generated based on individual areas for improvement.

[0468] "Real-time" refers to the process of collecting information and then immediately processing or providing feedback with virtually no delay.

[0469] The system for realizing this invention primarily includes a server, a wearable device worn by staff, and software for recording customer-staff interactions. The wearable device has a built-in camera and microphone to collect customer conversations and video information in real time. Furthermore, the device is designed to be lightweight and comfortable to wear, taking into account movement within the shop.

[0470] The server receives audio information transmitted from wearable devices and converts it into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text). This process analyzes the audio content as linguistic information and extracts key points. Simultaneously, video information is analyzed using a computer vision library (such as OpenCV), and customer expressions and behaviors are quantified as emotion scores. These analysis results are integrated and stored on the server as an evaluation of customer service behavior.

[0471] Based on the evaluation results, the server automatically generates the necessary training materials for each staff member. For example, if the analysis indicates that a staff member smiles infrequently, a video link specifically for smile training will be sent to the staff member. This feedback is displayed in real time on the staff member's smartphone app, allowing for immediate improvement. As a concrete example, if customer reactions when introducing seasonal products are analyzed, training may be provided to deepen knowledge about those products.

[0472] Furthermore, a generative AI model is utilized, with examples of prompts such as, "Please suggest what kind of training program would be effective in improving real-time responses to customer questions." This prompt allows staff to learn more advanced responses and gain specific insights to improve their customer service skills.

[0473] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0474] Step 1:

[0475] The terminal collects audio and video information in real time during interactions with customers through a wearable device worn by the staff. This information is input using the camera and microphone within the wearable device, and the data is transmitted to a server.

[0476] Step 2:

[0477] The server receives audio information from a wearable device as input and converts it into text data using a speech recognition API. The speech recognition API analyzes the audio waveform, extracts linguistic content based on this analysis, and generates output in text format.

[0478] Step 3:

[0479] The server uses a computer vision library to analyze the received video information. Taking the video information as input, it detects the customer's facial expressions and movements within each frame and outputs them as numerical data representing their emotions. It then scores whether the customer is smiling or showing interest.

[0480] Step 4:

[0481] The server integrates analysis results obtained from audio and video information to evaluate customer service behavior. Using the integrated data as input, it calculates and outputs an overall customer service evaluation score based on key interaction points and emotional scores. This score indicates which aspects of the customer service were successful and which need improvement.

[0482] Step 5:

[0483] The server generates individual training materials for each staff member based on their evaluation score. It selects the necessary educational content to address specific challenges during customer service and outputs it as a link to the staff member's terminal. This output could include, for example, videos on smile training or product description improvement.

[0484] Step 6:

[0485] Users (staff) receive and review training materials provided by the server on their smartphones or tablets. They can immediately begin learning by inputting feedback and training links. Users perform actions to solve specific problems and improve their customer service skills.

[0486] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0487] This invention provides a system that combines an emotion engine to enable a more detailed understanding of customer emotions during customer service, thereby achieving more effective customer service and staff training. This system uses a terminal equipped with a camera and AI microphone installed in the store to collect audio and video data in real time. When a customer receives customer service, the terminal automatically starts recording and video and transmits the data to a server.

[0488] The server converts audio data into text using natural language processing technology and analyzes the content of the conversation. Simultaneously, video data is analyzed using computer vision algorithms, and the customer's facial expressions and movements are quantified. Here, an emotion engine is utilized to identify the customer's emotional state based on the information obtained from their facial expressions and movements. This goes beyond mere data collection, enabling high-quality customer service evaluation that captures the customer's emotions.

[0489] The server integrates the analyzed data and performs a customer service evaluation that includes emotional data, along with a customer satisfaction score. This evaluation is used to understand the strengths and weaknesses of staff members in customer service. For example, if a particular customer was unsatisfactory, it can analyze in detail which emotions caused the dissatisfaction. Based on this, customized training materials are automatically generated for each staff member and provided to the user.

[0490] This system allows users to receive specific and emotionally sensitive feedback, which directly leads to improvements in staff customer service skills. Long-term data accumulation strengthens the foundation for developing unmanned customer service systems and designing new services. In summary, this invention contributes to improving the quality and operational efficiency of customer service operations.

[0491] The following describes the processing flow.

[0492] Step 1:

[0493] The terminal activates a camera-equipped AI microphone installed at the customer service counter, automatically starting to collect audio and video data when a customer enters the service area. Sensors detect customer movement, seamlessly recording and videotaping.

[0494] Step 2:

[0495] The terminal transfers the collected audio data to the server. The server converts the audio into text using natural language processing algorithms and analyzes the conversation content in real time. This text data is used as foundational information to extract important points of the dialogue.

[0496] Step 3:

[0497] The server receives video data transmitted from the terminal and analyzes it using computer vision technology. Specifically, it detects and quantifies the customer's facial expressions and gestures, and models the customer's behavioral patterns.

[0498] Step 4:

[0499] The server utilizes an emotion engine to identify the customer's emotional state from quantified facial expressions and gestures. This allows it to extract the customer's feelings of pleasure or displeasure during service and record them as data.

[0500] Step 5:

[0501] The server integrates voice-to-text, digitized facial expression data, and emotion data to calculate a customer satisfaction score. This score is used as an indicator to evaluate the quality of service and is utilized for more detailed service analysis.

[0502] Step 6:

[0503] The server automatically generates training materials based on evaluation results. It provides feedback tailored to the customer's emotional state and satisfaction level, and creates and delivers training content that includes individualized areas for improvement for the staff.

[0504] Step 7:

[0505] Users receive training materials and feedback from the server and attempt to improve their customer service skills based on that information. They conduct training and adjust their customer service methods as needed.

[0506] Step 8:

[0507] The server analyzes all the data accumulated over the long term and uses it to design unmanned customer service systems and develop new customer service models. This will facilitate the implementation of advanced, data-driven customer service systems.

[0508] (Example 2)

[0509] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0510] Traditional customer service evaluation systems have struggled to accurately capture customer emotions and expressions and provide individually tailored feedback. This has made it difficult to identify specific areas for improvement in customer service quality, hindering overall customer satisfaction.

[0511] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0512] In this invention, the server includes means for using a camera to collect audio and image data, a device for converting audio data into text using natural language processing from conversations with customers, a device for analyzing image data and converting customer facial expressions and actions into numerical values, a device for integrating the analysis results of these audio and image data to evaluate customer service, a device for automatically generating and providing personalized training materials based on the evaluation results, and a device for identifying the customer's emotional state using an emotion determination device. This enables detailed customer service evaluation that takes customer emotions into account and optimization of staff training based on that evaluation.

[0513] "Audio data" refers to sound information recorded from conversations and interactions with customers.

[0514] "Image data" refers to visual information that records a customer's facial expressions, actions, and other details.

[0515] "Recording device" refers to a device used to collect audio and image data.

[0516] "Natural language processing" refers to the technology of converting audio data into text and analyzing its content.

[0517] "Converting to numerical values" refers to the process of quantifying a customer's facial expressions and actions and transforming them into an evaluable format.

[0518] "Integrating analysis results" refers to combining information obtained from audio and image data to conduct a comprehensive customer service evaluation.

[0519] "Individualized training materials" refer to educational content optimized for each staff member based on evaluation results.

[0520] An "emotion determination device" refers to a device that identifies a customer's emotional state from their facial expressions and actions.

[0521] This invention provides a system that enables effective customer service and training by gaining a detailed understanding of customer emotions. The terminal uses a camera-equipped AI microphone installed in the store to collect audio and video data in real time. This device automatically starts recording and video when a customer is being served and transmits the data to a server.

[0522] The server converts received audio data into text using natural language processing (NLP) technology and analyzes the content of the conversation in detail. It also applies computer vision algorithms to video data to quantify customer facial expressions and movements. This process utilizes an emotion detection device to identify the customer's emotional state. This series of data analyses enables high-quality customer service evaluations that reflect customer emotions.

[0523] The server integrates the analyzed data, calculates customer satisfaction scores, and performs customer service evaluations that include emotional data. This evaluation is used to automatically generate customized training materials for each staff member. For example, if a particular customer interaction was unsatisfactory, the server analyzes the cause in detail to identify which emotions triggered that reaction. Based on these results, personalized training materials are provided to the staff, serving as user feedback.

[0524] For example, by inputting a prompt such as "Analyze the customer's emotional state and generate customized feedback" into the AI ​​model, actual training material is created. This system allows users to receive specific and emotionally sensitive feedback, which helps improve staff customer service skills. Furthermore, long-term data accumulation strengthens the foundation for the development of unmanned customer service systems and new service designs, contributing to improvements in the overall quality and efficiency of customer service operations.

[0525] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0526] Step 1:

[0527] The terminal uses a camera-equipped AI microphone installed in the store to collect customer audio and video data in real time. Specifically, the terminal automatically starts recording and video as soon as a customer receives service. The input is the customer's audio and video, and the output is the transmission of this data to the server.

[0528] Step 2:

[0529] The server receives audio data sent from the terminal and converts it into text using natural language processing techniques. Here, a speech recognition algorithm is used, with audio data as input and text data as output. This text data forms the basis for analyzing the content of the conversation.

[0530] Step 3:

[0531] The server applies computer vision algorithms to video data to quantify the customer's facial expressions and movements. Specifically, it analyzes facial feature points and body movement patterns through image processing. The input is video data, and the output is quantified facial expression and movement data.

[0532] Step 4:

[0533] The server passes quantified facial and movement data to an emotion recognition device to identify the customer's emotional state. This process involves analysis to assign specific emotion labels (e.g., "joy" or "anxiety") to the data. The input is quantified data, and the output is emotion-labeled data.

[0534] Step 5:

[0535] The server integrates the analysis results, calculates a customer satisfaction score, and performs a customer service evaluation incorporating emotional data. Statistical methods are used to convert various data points into a comprehensive evaluation value. Inputs are text data and emotionally labeled data, and output is a customer service evaluation score.

[0536] Step 6:

[0537] The server generates customized training materials for each staff member based on customer service evaluations and provides them to the user. In this process, a generative AI model is used to create specific training content based on prompts (for example, "Generate feedback that takes the customer's emotional state into consideration"). The input is the customer service evaluation score, and the output is the training material.

[0538] (Application Example 2)

[0539] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0540] In today's commercial environment, the quality of customer service is a crucial factor influencing customer satisfaction. However, traditional methods make it difficult to accurately grasp customer emotions and improve the immediate response capabilities of store staff. Therefore, there is a need to analyze customer emotions in real time and improve the quality of customer service based on that analysis.

[0541] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0542] In this invention, the server includes means for using a device for collecting audio and video data during customer service; means for converting audio data from customer interactions into text using natural language processing; means for analyzing video data to quantify customer expressions and actions; means for integrating these audio and video analysis results to evaluate customer service; means for automatically generating and providing training materials based on the evaluation results; and means for store employees to grasp customer emotions in real time using an intuitive information display device. This enables store employees to instantly understand customer emotions and provide appropriate customer service.

[0543] "Devices for collecting audio and video data" refer to devices installed in customer service settings to capture interactions and actions with customers, and include audio microphones and cameras.

[0544] "Means of converting to text using natural language processing" refers to software and technology for analyzing collected audio data and converting it into text format.

[0545] "Methods for analyzing video data to quantify customer expressions and movements" refers to software and algorithms that use computer vision technology to analyze video data and quantify emotions based on the customer's facial expressions and body movements.

[0546] "Means for evaluating customer service" refers to an analytical system that uses the results of integrating audio and video data to evaluate the emotional state of customers and quantify the quality of customer service.

[0547] "Means for automatically generating and providing training materials" refers to a program and process for automatically creating and providing educational materials aimed at improving the skills of store employees based on customer service evaluation results.

[0548] An "intuitive information display device" is a tool that directly visualizes information, such as smart glasses, enabling store employees to instantly grasp a customer's emotional state during service.

[0549] To implement this invention, it is necessary to build a system specifically designed for customer service. This system utilizes devices for collecting audio and video data, enabling real-time analysis of customer interactions. The server converts audio data into text using natural language processing technology, analyzes video data using computer vision, and quantifies emotional states. This allows store employees to understand customer emotions in real time and respond appropriately.

[0550] The specific system configuration involves installing devices such as cameras and microphones and connecting them to a server to collect video and audio data. The server analyzes this data and uses an advanced sentiment analysis engine to identify the customer's emotional state.

[0551] By using smart glasses as an intuitive information display device, store staff can visually confirm customers' emotions. For example, if a customer shows signs of anxiety, the smart glasses' display will show "Relaxing customer service recommended." This information provides direct guidance for staff to provide more supportive customer service.

[0552] An example of a prompt message is, "Analyze this customer's emotions and provide appropriate service during the interaction. Use the emotional data obtained from the conversation and facial expressions to adjust the service method in real time." This is how you can set instructions for the AI ​​system.

[0553] In this way, by linking the server and terminals, a system is built that enables high-quality customer service and improves customer satisfaction.

[0554] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0555] Step 1:

[0556] The terminal uses a microphone and camera installed during customer service to collect audio and video data of the customer. The collected data is sent to a server for real-time analysis. In this step, the input is the customer's audio and video, and the output is digital data sent to the server.

[0557] Step 2:

[0558] The server inputs the received audio data into a natural language processing engine and converts it into text format. This process analyzes the audio signal and outputs the transcribed conversation. The input for this step is audio data, and the output is text data.

[0559] Step 3:

[0560] The server analyzes the video data using computer vision algorithms to quantify the customer's facial expressions and movements. For example, it processes facial feature points to obtain emotion labels. The input for this step is video data, and the output is quantified emotion data.

[0561] Step 4:

[0562] The server integrates the data from steps 2 and 3 and uses an emotion engine to assess the customer's overall emotional state. This gives an understanding of how happy or anxious the customer is feeling. The inputs to this step are text data and emotion data, and the output is a detailed emotion rating score.

[0563] Step 5:

[0564] Based on the evaluation results, the server displays intuitive feedback on the employee's smart glasses. For example, it might say, "The customer needs to relax." The input for this step is the emotional evaluation score, and the output is the feedback message displayed on the smart glasses.

[0565] Step 6:

[0566] The server automatically generates training materials for store employees using a generative AI model based on long-term data accumulation. In this step, training scenarios are created based on the analysis of the accumulated data. The inputs to this step are evaluation results and historical data, and the output is the automatically generated training materials.

[0567] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0568] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0569] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0570] [Fourth Embodiment]

[0571] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0572] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0573] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0574] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0575] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0576] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0577] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0578] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0579] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0580] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0581] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0582] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0583] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0584] This invention relates to a system that collects audio and video data in real time during customer service and analyzes it to help with staff training and improving customer satisfaction. A terminal simultaneously collects audio and video using a camera-equipped AI microphone installed in the store. When a user begins receiving customer service, the terminal automatically starts responding, recording the conversation with high accuracy and capturing the customer's facial expressions as video.

[0585] The server receives audio data transmitted from the terminal and converts it to text in real time using a natural language processing engine. The converted text is used to analyze important interactions in customer-staff conversations and extract meaning. Meanwhile, the server analyzes video data and uses computer vision algorithms to quantify customer emotions and reactions.

[0586] The server integrates the results of analyzing the collected audio and video data to calculate a customer satisfaction score. This score is used to identify areas for improvement in staff service, and further detailed analysis is performed by comparing it with past success stories. Based on this, the server automatically generates training materials and provides them to the staff. For example, if a staff member smiles infrequently, a smile training video tailored to that staff member is selected, and a link is sent to that staff member.

[0587] This system allows users to receive direct feedback, enabling them to effectively improve their customer service skills. Furthermore, in the long term, the accumulated data is expected to be used in the design and improvement of unmanned customer service systems. The collected data will also serve as a valuable resource for expanding the system to other stores and entering new markets.

[0588] The following describes the processing flow.

[0589] Step 1:

[0590] The terminal activates the camera-equipped AI microphone installed in the store and begins real-time collection of audio and video data at the customer service counter. When the sensor detects that a customer has entered the customer service area, recording and video recording automatically begin.

[0591] Step 2:

[0592] The terminal sends the collected audio data to the server. The server uses speech recognition software to convert the audio into text data. This text data is used as basic data for analyzing the content of the conversation.

[0593] Step 3:

[0594] The server receives video data transmitted from the terminal and analyzes it using computer vision algorithms. This allows it to recognize the customer's emotional state from their facial expressions and gestures, quantify them, and store them in a database.

[0595] Step 4:

[0596] The server integrates the analyzed audio and video data to calculate a customer satisfaction score. This score is used to evaluate the quality of staff service and to indicate which elements contributed to customer satisfaction.

[0597] Step 5:

[0598] The server generates specific feedback for improving customer service based on customer satisfaction scores and analysis results. Furthermore, training materials to address areas that need improvement are automatically generated and provided to staff.

[0599] Step 6:

[0600] Users will access training materials and feedback provided by the server and utilize them as a means to improve their customer service skills. They are expected to periodically attempt to improve their customer service based on the feedback they receive.

[0601] Step 7:

[0602] The server analyzes data collected over the long term to develop unmanned customer service systems and apply them to new business models. Based on the accumulated data, it designs the scenarios and automated response systems necessary for unmanned customer service.

[0603] (Example 1)

[0604] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0605] Improving the quality of traditional customer service relies on the abilities of the staff, making it difficult to identify specific areas for improvement and provide effective feedback. Furthermore, the lack of objective methods for evaluating customer satisfaction hindered efficient improvement of customer service skills. Additionally, the insufficient accumulation and utilization of data necessary for designing and improving unmanned customer service systems made the development of innovative customer service methods challenging.

[0606] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0607] In this invention, the server includes means for using equipment for collecting audio and video data to convert audio data from customer interactions into text using natural language processing, means for analyzing video data to quantify customer emotions and reactions, means for integrating these audio and video analysis results to perform customer satisfaction evaluations, means for automatically generating and providing training materials based on identified areas for improvement, and means for providing feedback for improving customer service skills and accumulating and utilizing that data. This makes it possible to objectively evaluate customer satisfaction and efficiently improve customer service skills.

[0608] "Audio and video data" refers to digital information that records the content of conversations between customers and staff during customer service, as well as their facial expressions and actions at the time.

[0609] "Natural language processing" is a technology that analyzes speech data as text information and processes human language using computers.

[0610] "Quantification" is a method of converting customer emotions and reactions obtained from video data into quantitative numerical values.

[0611] "Customer satisfaction evaluation" is a process of measuring and evaluating customer satisfaction based on the results of analyzing collected audio and video data.

[0612] "Training materials" are learning resources created based on analyzed customer service data, with the aim of improving staff customer service skills.

[0613] "Feedback" refers to information provided to staff based on customer satisfaction ratings, used to improve customer service.

[0614] This invention is a system for collecting and analyzing audio and video data with the aim of improving customer service operations. The terminal, equipped with a camera and AI microphone installed in the store, simultaneously collects audio and video during interactions with customers. When a user begins interacting with a customer, the terminal automatically records audio and captures the customer's facial expressions and actions. The audio data is sent to a server and converted into text data in real time using a natural language processing engine (e.g., speech recognition API).

[0615] The server analyzes the converted text data and extracts key conversation points. Video data is also sent to the server, where computer vision algorithms (e.g., image analysis software) quantify customer emotions and reactions. Integrating these analysis results, the server calculates a customer satisfaction score. This score identifies areas for improvement in staff service and allows for detailed analysis by comparing it to past success stories.

[0616] Based on the analysis results, the server automatically generates and provides training materials for staff. For example, if it is determined that a staff member is not smiling enough, a link to a smile training video tailored to that staff member will be sent. Users can receive this feedback and effectively improve their customer service skills. Furthermore, in the long term, the accumulated data can be used to design unmanned customer service systems and support their deployment to other stores.

[0617] As a concrete example, let's consider a customer service scenario in a restaurant. By executing the prompt, "Explain how AI analyzes audio and video in a restaurant customer service scenario and provides feedback to staff," we can model the system's operation and verify its effectiveness. In this way, efficient improvements to customer service operations using AI models can be achieved.

[0618] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0619] Step 1:

[0620] The terminal uses a camera-equipped AI microphone installed in the store to simultaneously collect audio and video of users and customers. Inputs include real-time audio and video data. This data is captured with high precision by the AI ​​microphone and stored in the terminal's recording device. The output consists of collected audio and video files.

[0621] Step 2:

[0622] The terminal sends the collected audio data to the server. The transmitted audio data becomes the input. This data is transferred to the server in real time using a secure protocol. The output is the audio data stored on the server.

[0623] Step 3:

[0624] The server receives the audio data and converts it to text using a natural language processing engine. The input is the audio data sent in step 2. A speech recognition algorithm is used during the text conversion process, and natural language text is obtained as output.

[0625] Step 4:

[0626] The server analyzes the converted text and extracts key conversation points. The input is the text obtained in step 3. A text analysis algorithm is applied to identify customer requests and positive elements. The output is data of the identified key conversation points.

[0627] Step 5:

[0628] The terminal transmits the collected video data to the server. The input is real-time video data. This data, like the audio data, is transmitted securely. The output is video data stored on the server.

[0629] Step 6:

[0630] The server inputs video data into a computer vision algorithm to quantify customer emotions and reactions. The input is the video data transmitted in step 5. Through image analysis, customer smiles and gestures are quantified, and the output is quantified emotion data.

[0631] Step 7:

[0632] The server integrates the audio and video analysis results to calculate a customer satisfaction score. The input is the analysis data obtained in steps 4 and 6. The integration algorithm links the audio and video results to output a single satisfaction score.

[0633] Step 8:

[0634] The server analyzes customer satisfaction scores and identifies areas for improvement in customer service. The input is the customer satisfaction score calculated in step 7. By comparing this score with past success data, specific areas for improvement are identified. The output provides information on these improvement points.

[0635] Step 9:

[0636] The server generates and provides training materials for staff based on the identified areas for improvement. The input is the information on the improvement points extracted in step 8. The training materials are automatically created using a generation AI model, and the learning materials provided to staff are linked as output.

[0637] Step 10:

[0638] Users receive feedback from the server and improve their customer service skills. The input is the training material provided in step 9. Based on the feedback, users implement specific skill improvement measures, resulting in improved customer service abilities.

[0639] (Application Example 1)

[0640] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0641] Traditional customer service often lacked immediate feedback to improve service quality, resulting in a slow development of staff skills. Furthermore, limited means of accurately understanding customer emotions made it difficult to immediately assess the appropriateness of service. Additionally, the lack of real-time feedback for staff after service interactions hindered training efficiency.

[0642] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0643] In this invention, the server includes means for converting voice information into text using natural language processing with a wearable device worn by staff interacting with customers in the store; means for analyzing video information to quantify customer emotions and reactions; means for integrating these voice and video analysis results to evaluate customer service behavior; and means for automatically generating and providing individualized training materials for staff use based on the evaluation results. This enables real-time feedback during customer service, allowing staff to immediately identify areas for improvement and efficiently improve their customer service skills.

[0644] A "wearable device" is a device that staff can wear and use, and that has the function of acquiring audio and video information.

[0645] "Audio information" refers to the content of conversations recorded during interactions between customers and staff, which are then converted into text data using natural language processing.

[0646] "Video information" refers to recording customers' facial expressions and movements, which serves as the basis for quantifying their emotions and reactions.

[0647] "Natural language processing" is a technology that converts spoken information into text information and analyzes the content of conversations with customers.

[0648] "Quantification" refers to expressing customer emotions and reactions as quantitative data based on video information.

[0649] "Evaluating customer service behavior" involves assessing staff members' customer service skills and the quality of their responses based on analyzed audio and video information.

[0650] "Training materials" refer to educational materials and information provided to staff to improve their customer service skills, and are generated based on individual areas for improvement.

[0651] "Real-time" refers to the process of collecting information and then immediately processing or providing feedback with virtually no delay.

[0652] The system for realizing this invention primarily includes a server, a wearable device worn by staff, and software for recording customer-staff interactions. The wearable device has a built-in camera and microphone to collect customer conversations and video information in real time. Furthermore, the device is designed to be lightweight and comfortable to wear, taking into account movement within the shop.

[0653] The server receives audio information transmitted from wearable devices and converts it into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text). This process analyzes the audio content as linguistic information and extracts key points. Simultaneously, video information is analyzed using a computer vision library (such as OpenCV), and customer expressions and behaviors are quantified as emotion scores. These analysis results are integrated and stored on the server as an evaluation of customer service behavior.

[0654] Based on the evaluation results, the server automatically generates the necessary training materials for each staff member. For example, if the analysis indicates that a staff member smiles infrequently, a video link specifically for smile training will be sent to the staff member. This feedback is displayed in real time on the staff member's smartphone app, allowing for immediate improvement. As a concrete example, if customer reactions when introducing seasonal products are analyzed, training may be provided to deepen knowledge about those products.

[0655] Furthermore, a generative AI model is utilized, with examples of prompts such as, "Please suggest what kind of training program would be effective in improving real-time responses to customer questions." This prompt allows staff to learn more advanced responses and gain specific insights to improve their customer service skills.

[0656] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0657] Step 1:

[0658] The terminal collects audio and video information in real time during interactions with customers through a wearable device worn by the staff. This information is input using the camera and microphone within the wearable device, and the data is transmitted to a server.

[0659] Step 2:

[0660] The server receives audio information from a wearable device as input and converts it into text data using a speech recognition API. The speech recognition API analyzes the audio waveform, extracts linguistic content based on this analysis, and generates output in text format.

[0661] Step 3:

[0662] The server uses a computer vision library to analyze the received video information. Taking the video information as input, it detects the customer's facial expressions and movements within each frame and outputs them as numerical data representing their emotions. It then scores whether the customer is smiling or showing interest.

[0663] Step 4:

[0664] The server integrates analysis results obtained from audio and video information to evaluate customer service behavior. Using the integrated data as input, it calculates and outputs an overall customer service evaluation score based on key interaction points and emotional scores. This score indicates which aspects of the customer service were successful and which need improvement.

[0665] Step 5:

[0666] The server generates individual training materials for each staff member based on their evaluation score. It selects the necessary educational content to address specific challenges during customer service and outputs it as a link to the staff member's terminal. This output could include, for example, videos on smile training or product description improvement.

[0667] Step 6:

[0668] Users (staff) receive and review training materials provided by the server on their smartphones or tablets. They can immediately begin learning by inputting feedback and training links. Users perform actions to solve specific problems and improve their customer service skills.

[0669] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0670] This invention provides a system that combines an emotion engine to enable a more detailed understanding of customer emotions during customer service, thereby achieving more effective customer service and staff training. This system uses a terminal equipped with a camera and AI microphone installed in the store to collect audio and video data in real time. When a customer receives customer service, the terminal automatically starts recording and video and transmits the data to a server.

[0671] The server converts audio data into text using natural language processing technology and analyzes the content of the conversation. Simultaneously, video data is analyzed using computer vision algorithms, and the customer's facial expressions and movements are quantified. Here, an emotion engine is utilized to identify the customer's emotional state based on the information obtained from their facial expressions and movements. This goes beyond mere data collection, enabling high-quality customer service evaluation that captures the customer's emotions.

[0672] The server integrates the analyzed data and performs a customer service evaluation that includes emotional data, along with a customer satisfaction score. This evaluation is used to understand the strengths and weaknesses of staff members in customer service. For example, if a particular customer was unsatisfactory, it can analyze in detail which emotions caused the dissatisfaction. Based on this, customized training materials are automatically generated for each staff member and provided to the user.

[0673] This system allows users to receive specific and emotionally sensitive feedback, which directly leads to improvements in staff customer service skills. Long-term data accumulation strengthens the foundation for developing unmanned customer service systems and designing new services. In summary, this invention contributes to improving the quality and operational efficiency of customer service operations.

[0674] The following describes the processing flow.

[0675] Step 1:

[0676] The terminal activates a camera-equipped AI microphone installed at the customer service counter, automatically starting to collect audio and video data when a customer enters the service area. Sensors detect customer movement, seamlessly recording and videotaping.

[0677] Step 2:

[0678] The terminal transfers the collected audio data to the server. The server converts the audio into text using natural language processing algorithms and analyzes the conversation content in real time. This text data is used as foundational information to extract important points of the dialogue.

[0679] Step 3:

[0680] The server receives video data transmitted from the terminal and analyzes it using computer vision technology. Specifically, it detects and quantifies the customer's facial expressions and gestures, and models the customer's behavioral patterns.

[0681] Step 4:

[0682] The server utilizes an emotion engine to identify the customer's emotional state from quantified facial expressions and gestures. This allows it to extract the customer's feelings of pleasure or displeasure during service and record them as data.

[0683] Step 5:

[0684] The server integrates voice-to-text, digitized facial expression data, and emotion data to calculate a customer satisfaction score. This score is used as an indicator to evaluate the quality of service and is utilized for more detailed service analysis.

[0685] Step 6:

[0686] The server automatically generates training materials based on evaluation results. It provides feedback tailored to the customer's emotional state and satisfaction level, and creates and delivers training content that includes individualized areas for improvement for the staff.

[0687] Step 7:

[0688] Users receive training materials and feedback from the server and attempt to improve their customer service skills based on that information. They conduct training and adjust their customer service methods as needed.

[0689] Step 8:

[0690] The server analyzes all the data accumulated over the long term and uses it to design unmanned customer service systems and develop new customer service models. This will facilitate the implementation of advanced, data-driven customer service systems.

[0691] (Example 2)

[0692] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0693] Traditional customer service evaluation systems have struggled to accurately capture customer emotions and expressions and provide individually tailored feedback. This has made it difficult to identify specific areas for improvement in customer service quality, hindering overall customer satisfaction.

[0694] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0695] In this invention, the server includes means for using a camera to collect audio and image data, a device for converting audio data into text using natural language processing from conversations with customers, a device for analyzing image data and converting customer facial expressions and actions into numerical values, a device for integrating the analysis results of these audio and image data to evaluate customer service, a device for automatically generating and providing personalized training materials based on the evaluation results, and a device for identifying the customer's emotional state using an emotion determination device. This enables detailed customer service evaluation that takes customer emotions into account and optimization of staff training based on that evaluation.

[0696] "Audio data" refers to sound information recorded from conversations and interactions with customers.

[0697] "Image data" refers to visual information that records a customer's facial expressions, actions, and other details.

[0698] "Recording device" refers to a device used to collect audio and image data.

[0699] "Natural language processing" refers to the technology of converting audio data into text and analyzing its content.

[0700] "Converting to numerical values" refers to the process of quantifying a customer's facial expressions and actions and transforming them into an evaluable format.

[0701] "Integrating analysis results" refers to combining information obtained from audio and image data to conduct a comprehensive customer service evaluation.

[0702] "Individualized training materials" refer to educational content optimized for each staff member based on evaluation results.

[0703] An "emotion determination device" refers to a device that identifies a customer's emotional state from their facial expressions and actions.

[0704] This invention provides a system that enables effective customer service and training by gaining a detailed understanding of customer emotions. The terminal uses a camera-equipped AI microphone installed in the store to collect audio and video data in real time. This device automatically starts recording and video when a customer is being served and transmits the data to a server.

[0705] The server converts received audio data into text using natural language processing (NLP) technology and analyzes the content of the conversation in detail. It also applies computer vision algorithms to video data to quantify customer facial expressions and movements. This process utilizes an emotion detection device to identify the customer's emotional state. This series of data analyses enables high-quality customer service evaluations that reflect customer emotions.

[0706] The server integrates the analyzed data, calculates customer satisfaction scores, and performs customer service evaluations that include emotional data. This evaluation is used to automatically generate customized training materials for each staff member. For example, if a particular customer interaction was unsatisfactory, the server analyzes the cause in detail to identify which emotions triggered that reaction. Based on these results, personalized training materials are provided to the staff, serving as user feedback.

[0707] For example, by inputting a prompt such as "Analyze the customer's emotional state and generate customized feedback" into the AI ​​model, actual training material is created. This system allows users to receive specific and emotionally sensitive feedback, which helps improve staff customer service skills. Furthermore, long-term data accumulation strengthens the foundation for the development of unmanned customer service systems and new service designs, contributing to improvements in the overall quality and efficiency of customer service operations.

[0708] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0709] Step 1:

[0710] The terminal uses a camera-equipped AI microphone installed in the store to collect customer audio and video data in real time. Specifically, the terminal automatically starts recording and video as soon as a customer receives service. The input is the customer's audio and video, and the output is the transmission of this data to the server.

[0711] Step 2:

[0712] The server receives audio data sent from the terminal and converts it into text using natural language processing techniques. Here, a speech recognition algorithm is used, with audio data as input and text data as output. This text data forms the basis for analyzing the content of the conversation.

[0713] Step 3:

[0714] The server applies computer vision algorithms to video data to quantify the customer's facial expressions and movements. Specifically, it analyzes facial feature points and body movement patterns through image processing. The input is video data, and the output is quantified facial expression and movement data.

[0715] Step 4:

[0716] The server passes quantified facial and movement data to an emotion recognition device to identify the customer's emotional state. This process involves analysis to assign specific emotion labels (e.g., "joy" or "anxiety") to the data. The input is quantified data, and the output is emotion-labeled data.

[0717] Step 5:

[0718] The server integrates the analysis results, calculates a customer satisfaction score, and performs a customer service evaluation incorporating emotional data. Statistical methods are used to convert various data points into a comprehensive evaluation value. Inputs are text data and emotionally labeled data, and output is a customer service evaluation score.

[0719] Step 6:

[0720] The server generates customized training materials for each staff member based on customer service evaluations and provides them to the user. In this process, a generative AI model is used to create specific training content based on prompts (for example, "Generate feedback that takes the customer's emotional state into consideration"). The input is the customer service evaluation score, and the output is the training material.

[0721] (Application Example 2)

[0722] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0723] In today's commercial environment, the quality of customer service is a crucial factor influencing customer satisfaction. However, traditional methods make it difficult to accurately grasp customer emotions and improve the immediate response capabilities of store staff. Therefore, there is a need to analyze customer emotions in real time and improve the quality of customer service based on that analysis.

[0724] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0725] In this invention, the server includes means for using a device for collecting audio and video data during customer service; means for converting audio data from customer interactions into text using natural language processing; means for analyzing video data to quantify customer expressions and actions; means for integrating these audio and video analysis results to evaluate customer service; means for automatically generating and providing training materials based on the evaluation results; and means for store employees to grasp customer emotions in real time using an intuitive information display device. This enables store employees to instantly understand customer emotions and provide appropriate customer service.

[0726] "Devices for collecting audio and video data" refer to devices installed in customer service settings to capture interactions and actions with customers, and include audio microphones and cameras.

[0727] "Means of converting to text using natural language processing" refers to software and technology for analyzing collected audio data and converting it into text format.

[0728] "Methods for analyzing video data to quantify customer expressions and movements" refers to software and algorithms that use computer vision technology to analyze video data and quantify emotions based on the customer's facial expressions and body movements.

[0729] "Means for evaluating customer service" refers to an analytical system that uses the results of integrating audio and video data to evaluate the emotional state of customers and quantify the quality of customer service.

[0730] "Means for automatically generating and providing training materials" refers to a program and process for automatically creating and providing educational materials aimed at improving the skills of store employees based on customer service evaluation results.

[0731] An "intuitive information display device" is a tool that directly visualizes information, such as smart glasses, enabling store employees to instantly grasp a customer's emotional state during service.

[0732] To implement this invention, it is necessary to build a system specifically designed for customer service. This system utilizes devices for collecting audio and video data, enabling real-time analysis of customer interactions. The server converts audio data into text using natural language processing technology, analyzes video data using computer vision, and quantifies emotional states. This allows store employees to understand customer emotions in real time and respond appropriately.

[0733] The specific system configuration involves installing devices such as cameras and microphones and connecting them to a server to collect video and audio data. The server analyzes this data and uses an advanced sentiment analysis engine to identify the customer's emotional state.

[0734] By using smart glasses as an intuitive information display device, store staff can visually confirm customers' emotions. For example, if a customer shows signs of anxiety, the smart glasses' display will show "Relaxing customer service recommended." This information provides direct guidance for staff to provide more supportive customer service.

[0735] An example of a prompt message is, "Analyze this customer's emotions and provide appropriate service during the interaction. Use the emotional data obtained from the conversation and facial expressions to adjust the service method in real time." This is how you can set instructions for the AI ​​system.

[0736] In this way, by linking the server and terminals, a system is built that enables high-quality customer service and improves customer satisfaction.

[0737] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0738] Step 1:

[0739] The terminal uses a microphone and camera installed during customer service to collect audio and video data of the customer. The collected data is sent to a server for real-time analysis. In this step, the input is the customer's audio and video, and the output is digital data sent to the server.

[0740] Step 2:

[0741] The server inputs the received audio data into a natural language processing engine and converts it into text format. This process analyzes the audio signal and outputs the transcribed conversation. The input for this step is audio data, and the output is text data.

[0742] Step 3:

[0743] The server analyzes the video data using computer vision algorithms to quantify the customer's facial expressions and movements. For example, it processes facial feature points to obtain emotion labels. The input for this step is video data, and the output is quantified emotion data.

[0744] Step 4:

[0745] The server integrates the data from steps 2 and 3 and uses an emotion engine to assess the customer's overall emotional state. This gives an understanding of how happy or anxious the customer is feeling. The inputs to this step are text data and emotion data, and the output is a detailed emotion rating score.

[0746] Step 5:

[0747] Based on the evaluation results, the server displays intuitive feedback on the employee's smart glasses. For example, it might say, "The customer needs to relax." The input for this step is the emotional evaluation score, and the output is the feedback message displayed on the smart glasses.

[0748] Step 6:

[0749] The server automatically generates training materials for store employees using a generative AI model based on long-term data accumulation. In this step, training scenarios are created based on the analysis of the accumulated data. The inputs to this step are evaluation results and historical data, and the output is the automatically generated training materials.

[0750] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0751] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0752] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0753] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0754] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. In the upper and lower directions of the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. Also, the upper side of the concentric circles is where "pleasant" emotions are located, and the lower side is where "unpleasant" emotions are located. In this way, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0755] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0756] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0757] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0758] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0759] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0760] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0761] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0762] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0763] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0764] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0765] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0766] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0767] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0768] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0769] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0770] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0771] The following is further disclosed regarding the embodiments described above.

[0772] (Claim 1)

[0773] Using a device to collect audio and video data during customer service,

[0774] A means of converting audio data from customer conversations into text using natural language processing,

[0775] A method for analyzing video data to quantify customer facial expressions and movements,

[0776] A means of integrating these audio and video analysis results to evaluate customer service,

[0777] A means of automatically generating and providing training materials based on evaluation results,

[0778] A system that includes this.

[0779] (Claim 2)

[0780] The system according to claim 1, which acquires and analyzes customer service voice data in real time.

[0781] (Claim 3)

[0782] The system according to claim 1, which generates an automated response script from accumulated evaluation data.

[0783] "Example 1"

[0784] (Claim 1)

[0785] Using equipment to collect audio and video data,

[0786] A means of converting audio data from customer conversations into text using natural language processing,

[0787] A method for analyzing video data to quantify customer emotions and reactions,

[0788] A means of integrating these audio and video analysis results to evaluate customer satisfaction,

[0789] A method for identifying areas for improvement in customer service by comparing evaluation results with past success stories,

[0790] A means of automatically generating and providing training materials based on identified areas for improvement,

[0791] A means of providing feedback to improve customer service skills, and accumulating and utilizing that data,

[0792] A system that includes this.

[0793] (Claim 2)

[0794] The system according to claim 1, which acquires and analyzes customer service audio and video data in real time.

[0795] (Claim 3)

[0796] The system according to claim 1, which provides a foundation for designing and improving an unmanned customer service system based on accumulated customer satisfaction evaluation data.

[0797] "Application Example 1"

[0798] (Claim 1)

[0799] Using wearable devices worn by staff who interact with customers in the store,

[0800] A means of converting voice information from customer conversations into text using natural language processing,

[0801] A method for analyzing video information to quantify customer emotions and reactions,

[0802] A means of integrating these audio and video analysis results to evaluate customer service behavior,

[0803] Based on the evaluation results, a means of automatically generating and providing individual training materials for staff use,

[0804] A system that includes this.

[0805] (Claim 2)

[0806] The system according to claim 1, wherein the staff providing customer service use a display device to provide real-time feedback.

[0807] (Claim 3)

[0808] The system according to claim 1, which uses accumulated evaluation data to improve an unmanned customer service system.

[0809] "Example 2 of combining an emotion engine"

[0810] (Claim 1)

[0811] Using a camera for collecting audio and image data,

[0812] A device that converts voice data from customer conversations into text using natural language processing,

[0813] A device that analyzes image data and converts customer facial expressions and movements into numerical values,

[0814] A device that integrates the results of these voice and image analyses to evaluate customer service,

[0815] A device that automatically generates and provides individualized training materials based on evaluation results,

[0816] A device that identifies a customer's emotional state using an emotion assessment device,

[0817] A system that includes this.

[0818] (Claim 2)

[0819] The system according to claim 1, which acquires and analyzes customer service voice data in real time.

[0820] (Claim 3)

[0821] The system according to claim 1, which generates an automatically generated response format from accumulated evaluation data.

[0822] "Application example 2 when combining with an emotional engine"

[0823] (Claim 1)

[0824] Using equipment to collect audio and video data during customer service,

[0825] A means of converting audio data from customer conversations into text using natural language processing,

[0826] A method for analyzing video data to quantify customer facial expressions and movements,

[0827] A means of integrating these audio and video analysis results to evaluate customer service,

[0828] A means of automatically generating and providing training materials based on evaluation results,

[0829] A means for store staff to grasp customer emotions in real time using an intuitive information display device,

[0830] A system that includes this.

[0831] (Claim 2)

[0832] The system according to claim 1, which acquires and analyzes customer service voice data in real time.

[0833] (Claim 3)

[0834] The system according to claim 1, which generates an automated response script from accumulated evaluation data. [Explanation of symbols]

[0835] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. Using a device to collect audio and video data during customer service, A means of converting audio data from customer conversations into text using natural language processing, A method for analyzing video data to quantify customer facial expressions and movements, A means of integrating these audio and video analysis results to evaluate customer service, A means of automatically generating and providing training materials based on evaluation results, A system that includes this.

2. The system according to claim 1, which acquires and analyzes customer service voice data in real time.

3. The system according to claim 1, which generates an automated response script from accumulated evaluation data.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A