system

A system that converts still image and audio data into moving image and text data to generate customer service scripts and evaluate satisfaction in real-time addresses the challenge of inconsistent service quality, enhancing customer satisfaction and feedback utilization.

JP2026062228APending Publication Date: 2026-04-09SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

In modern retail and service industries, there is a challenge in maintaining consistent customer service quality due to variations among employees' skills, and there are limited means to collect and utilize real-time customer feedback and evaluations.

Method used

A system that converts still image and audio data into moving image and text data to generate customer service scripts, collects conversation data in real-time, and calculates evaluation scores using a scoring algorithm to provide consistent customer service and feedback.

Benefits of technology

Enables consistent maintenance of employee customer service quality, increases customer satisfaction, and allows for the collection and utilization of real-time customer feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026062228000001_ABST
    Figure 2026062228000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] In order to convert still image data into video data, a means for receiving the still image data and audio data as input, A means for generating video data based on the input still image data and audio data, Means for transmitting the generated video data to a terminal, The terminal includes means for displaying the transmitted video data, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In modern retail and service industries, customer service skills for improving customer satisfaction are very important. However, not all employees have advanced customer service skills, and it takes time and cost to acquire customer service skills. As a result, there are variations in the quality of customer service among stores, making it difficult to consistently maintain customer satisfaction. Furthermore, it is also difficult to collect and utilize evaluations of customer service staff and customer feedback in real time. To solve these problems, it is necessary to introduce an AI system with excellent customer service skills.

Means for Solving the Problems

[0005] This invention provides a system for converting still image data into moving image data. Specifically, it includes means for generating moving image data based on still image data and audio data, and means for transmitting the generated moving image data to a terminal and displaying it on the terminal. Simultaneously, it includes means for converting customer service audio data into text data, generating a customer service script based on this, and transmitting it to the terminal to interact with the user. Furthermore, it includes means for collecting and transcribing conversation data in real time, calculating an evaluation score using a scoring algorithm, and transmitting and displaying this score on the terminal. This enables consistent maintenance of employee customer service quality, increases customer satisfaction, and allows for the collection and utilization of real-time customer feedback.

[0006] "Still image data" refers to digital image information saved in a format where the image does not move.

[0007] "Motion image data" refers to digital data that forms a moving image by displaying multiple still images in sequence.

[0008] "Audio data" refers to data used to store and transmit sound information in digital format.

[0009] A "terminal" refers to an electronic device used by a user, which has network connectivity and is capable of receiving and displaying data.

[0010] "Customer service audio data" refers to data that stores audio information recorded during customer service interactions in digital format.

[0011] "Character data" refers to linguistic information expressed in text format.

[0012] A "customer service script" refers to a written document containing the phrases and responses used during customer service interactions.

[0013] "Conversation data" refers to data used to digitally store and process audio recordings of conversations between users and customers.

[0014] "Scoring algorithm" refers to a calculation method or program for calculating an evaluation score based on the input data.

[0015] "Evaluation score" refers to the numerical value of the result evaluated based on certain criteria.

Brief Explanation of Drawings

[0016] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13]It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.

Embodiments for Carrying Out the Invention

[0017] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), etc.

[0020] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0021] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0024] [First Embodiment]

[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0037] This invention relates to a system that generates moving image data based on still image data and audio data to provide consistent customer service. It also includes collecting and analyzing conversation data during customer service in real time to evaluate customer satisfaction. Specific embodiments of this system are described below.

[0038] 1. Generating the avatar chat screen

[0039] Server Processing

[0040] The server retrieves pre-prepared still images and audio data of employees. Next, it sends this data via an API (Application Programming Interface) to request the generation of video data. The video data returned from the API is received by the server, which then sends the video data to the terminal.

[0041] Terminal processing

[0042] The terminal receives video data transmitted from the server. After receiving the data, the terminal plays it back and displays it to the user.

[0043] Specific example

[0044] For example, if a server has still images and audio data of employee A, the server sends these to an API to generate video data. Upon receiving the video data as a response from the API, the server sends it to a terminal. The terminal plays this data, providing the user with a video that makes it appear as if employee A is speaking.

[0045] 2. Generating and using customer service scripts

[0046] Server Processing

[0047] The server collects voice data of excellent crew members' customer service interactions. This collected voice data is converted into text data using a speech recognition API. Next, the server feeds this text data into a GPT model (Greater Global Pattern Testing) to generate appropriate customer service scripts. The generated scripts are then sent from the server to the terminals.

[0048] Terminal processing

[0049] The terminal receives a customer service script sent from the server. Based on this script, the terminal interacts with the user.

[0050] Specific example

[0051] If Crew B possesses excellent customer service skills, the server records Crew B's customer service audio. This audio data is transcribed and fed into a GPT model to generate a customer service script. The terminal uses this script to provide appropriate responses to the user's questions.

[0052] 3. Talk Cancellation Scoring

[0053] Server Processing

[0054] The server collects conversation data with users in real time. The collected conversation data is converted into text data via a speech recognition API and then input into a scoring algorithm. This algorithm analyzes the conversation data and calculates a score to evaluate customer satisfaction and willingness to continue using the service. Finally, this evaluation score is sent from the server to the terminal.

[0055] Terminal processing

[0056] The device receives evaluation scores sent from the server and displays them to the user in real time.

[0057] Specific example

[0058] While a conversation with user C is in progress, the terminal collects the conversation content in real time and sends it to the server. The server transcribes this audio data and calculates an evaluation score using a scoring algorithm. By displaying this score on the terminal, it becomes possible to continuously evaluate and improve the quality of the service.

[0059] In this way, the present invention can consistently improve customer satisfaction through an AI system with advanced customer service skills.

[0060] The following describes the processing flow.

[0061] 1. Generating the avatar chat screen

[0062] Step 1:

[0063] The server retrieves pre-prepared still images and audio data of employees.

[0064] Step 2:

[0065] The server creates and sends an HTTP POST request to send the acquired still image and audio data to the API.

[0066] Step 3:

[0067] The server waits for a response from the API and receives the generated video data. This video data is generated from still images and audio data.

[0068] Step 4:

[0069] The server sends the received video data to the terminal. It uses an appropriate communication protocol (e.g., HTTP, WebSocket) for this purpose.

[0070] Step 5:

[0071] The terminal receives video data transmitted from the server. After receiving the data, the terminal plays it back and displays it to the user.

[0072] Specific example

[0073] Still images and audio data of employee A are stored on the server. The server sends this data to an API, receives video data generated by the API, and sends it to the terminal. The terminal receives the video data and displays to the user an image that makes it appear as if employee A is speaking.

[0074] 2. Generating and using customer service scripts

[0075] Step 1:

[0076] The server collects audio data of excellent customer service from top-performing crew members. This audio data is pre-recorded and stored in a database.

[0077] Step 2:

[0078] The server sends the collected customer service voice data to a speech recognition API, where it is converted into text data. The voice data is sent to the API via an HTTP POST request.

[0079] Step 3:

[0080] The server receives the text data returned from the speech recognition API.

[0081] Step 4:

[0082] The server inputs the received text data into a GPT model to generate a customer service script. The GPT model is pre-trained and can automatically generate appropriate responses.

[0083] Step 5:

[0084] The server sends the generated customer service script to the terminal, using the appropriate communication protocol.

[0085] Step 6:

[0086] The terminal receives customer service scripts sent from the server and provides appropriate responses based on the user's questions and requests.

[0087] Specific example

[0088] Crew B's excellent customer service voice is stored on the server. The server sends this voice data to a speech recognition API to transcribe it and input it into a GPT model. The generated script is sent to the terminal, which uses it to respond to the user.

[0089] 3. Talk Cancellation Scoring

[0090] Step 1:

[0091] The device collects conversation data with the user in real time. This collection is done through a voice input device.

[0092] Step 2:

[0093] The terminal sends the collected audio data to the server. An appropriate communication protocol (e.g., HTTP, WebSocket) is used.

[0094] Step 3:

[0095] The server sends the received audio data to a speech recognition API, where it is converted into text data.

[0096] Step 4:

[0097] The server receives the text data returned from the speech recognition API.

[0098] Step 5:

[0099] The server inputs the received text data into a scoring algorithm to calculate a score for evaluating customer satisfaction and willingness to continue using the service.

[0100] Step 6:

[0101] The server sends the calculated evaluation score to the terminal.

[0102] Step 7:

[0103] The device receives evaluation scores sent from the server and displays them to the user in real time.

[0104] Specific example

[0105] As a conversation with user C progresses, the terminal collects conversation data in real time and sends it to the server. The server sends the data to a speech recognition API to transcribe it and calculates an evaluation score using a scoring algorithm. This score is sent to the terminal, which then displays it to the user in real time.

[0106] In this way, the system of the present invention can provide advanced customer service skills and real-time customer evaluation, thereby consistently improving customer satisfaction.

[0107] (Example 1)

[0108] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0109] In modern society, consistent customer service is required to enhance customer satisfaction. Especially in online customer service, direct interaction with customers is difficult, and service often depends on the individual skills of employees. This can lead to inconsistencies in service quality, potentially resulting in decreased customer satisfaction. Furthermore, the lack of adequate means to evaluate customer satisfaction in real time and improve service quality makes prompt responses difficult. A system is needed to address these challenges and improve customer satisfaction.

[0110] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0111] In this invention, the server includes means for receiving still image data and audio data as input, means for generating video data based on the input still image data and audio data, means for transmitting the generated video data to a terminal, means for converting customer service audio data into text data and generating a customer service script using a generation AI model, means for transmitting the generated customer service script to a terminal, means for converting conversation data into text data and calculating an evaluation score using a scoring algorithm, and means for transmitting the calculated evaluation score to a terminal. As a result, the server generates video data and transmits it to a terminal to achieve consistent customer service, and it becomes possible to generate scripts that utilize the customer service skills of crew members with excellent customer service skills and use them on the terminal. Furthermore, it becomes possible to perform real-time customer satisfaction evaluation and continuously improve the quality of customer service.

[0112] "Still image data" refers to digital data containing image information captured at a specific point in time.

[0113] "Audio data" refers to data that records human voices in digital format.

[0114] "Motion image data" refers to digital data in video format that includes a sequence of images and synchronized audio.

[0115] A "terminal" is an electronic device that a user can directly operate.

[0116] "Customer service audio data" refers to data that digitally records the voices spoken by employees during customer service interactions.

[0117] "Text data" refers to data that visualizes audio data as text using speech recognition technology.

[0118] A "generative AI model" is an artificial intelligence model trained to generate appropriate responses or scripts from specific input data.

[0119] A "customer service script" is text data containing a series of sentences and response instructions used during customer service.

[0120] "Conversation data" refers to digital data that records voice interactions between a user and a system.

[0121] A "scoring algorithm" is a calculation method used to analyze input data and calculate an evaluation score based on specific criteria.

[0122] An "evaluation score" is a numerical evaluation result calculated by a scoring algorithm.

[0123] This invention relates to a system that generates moving image data based on still image data and audio data to provide consistent customer service. It also includes a function to collect and analyze conversation data during customer service in real time and evaluate customer satisfaction. Specific embodiments of this invention will now be described.

[0124] First, the server retrieves pre-prepared still image data (JPEG format) and audio data (WAV format) of the employee. This data is sent via an HTTP POST request using an API such as Microsoft® Azure® Face API. The response returned from the API contains the generated video data (MP4 format), which the server receives and temporarily stores. Subsequently, the server sends the stored video data to the terminal as an HTTP response.

[0125] Next, the terminal receives video data sent from the server. After receiving the data, the terminal initializes a media player to play it and displays the video to the user. This series of processes allows the terminal to play a video that makes it appear as if an employee is speaking.

[0126] Next, to collect excellent customer service voice data (WAV format) from crew members, the server uses the Google® Cloud Speech-to-Text API to convert the voice data into text data. Then, the text data is fed into a generative AI model such as the GPT-3® model to generate customer service scripts. The generated scripts are temporarily stored on the server and sent to the terminal. The terminal receives the customer service scripts sent from the server and uses them to interact with the user.

[0127] Next, the server collects conversation data (in WAV format) from the user in real time. The conversation data is converted into text data using the Google Cloud Speech-to-Text API and then fed into a scoring algorithm. This algorithm analyzes the conversation data and calculates a score that evaluates customer satisfaction and willingness to continue using the service. This evaluation score is sent from the server to the terminal and displayed to the user in real time, enabling continuous evaluation and improvement of service quality.

[0128] Specific example

[0129] 1. The server sends still images and audio data of employee A to the API to generate video data. The terminal receives this data and displays a video to the user that makes it appear as if employee A is speaking.

[0130] 2. The customer service voice of Crew B is transcribed and input into the GPT-3 model to generate a customer service script. The terminal uses this script to respond to the user's questions.

[0131] 3. The conversation with User C is collected, and a customer satisfaction evaluation score is calculated using a scoring algorithm. The terminal displays this score and evaluates the service quality.

[0132] Example of a prompt

[0133] 1. "Please generate video data based on a still image of employee A and the following audio data."

[0134] 2. "Please generate a customer service script based on the text data of this customer service audio."

[0135] 3. "Based on the conversation data, calculate a score to evaluate customer satisfaction."

[0136] In this way, the present invention provides a system that generates moving image data using still image data and audio data, evaluates customer satisfaction in real time, and realizes consistently high-quality customer service.

[0137] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0138] Generating an avatar chat screen

[0139] Server Processing

[0140] Step 1:

[0141] The server retrieves pre-prepared still image data (e.g., JPEG format) and audio data (e.g., WAV format) of employees. The server receives the still image and audio data as input. This data serves as the basic information for generating video data in subsequent processing.

[0142] Step 2:

[0143] The server sends the acquired still image and audio data to an API (e.g., Microsoft Azure's Face API). It makes an HTTP POST request to the API endpoint, sending the still image and audio data. Based on the input data, the API executes a process and generates a video image data by combining the still image and audio data.

[0144] Step 3:

[0145] The server receives video data (e.g., in MP4 format) returned from the API. It temporarily stores the video data obtained as an HTTP response and prepares for the subsequent transmission process.

[0146] Step 4:

[0147] The server sends the generated video data to the terminal as an HTTP response. This allows the video data to be delivered to the terminal via the internet.

[0148] Terminal processing

[0149] Step 1:

[0150] The terminal receives video data sent from the server as an HTTP response. It stores the received data in temporary storage. It then parses and saves the URL or binary data used as input data.

[0151] Step 2:

[0152] The device initializes its media player for playing saved video data. It internally loads the video data and prepares it for display to the user.

[0153] Step 3:

[0154] The device plays video and displays it to the user. It uses a media player to display the video along a timeline and outputs it in a way that the user can perceive.

[0155] Generating and using customer service scripts

[0156] Server Processing

[0157] Step 1:

[0158] The server collects voice data (e.g., in WAV format) of high-performing crew members providing excellent customer service. The collected voice data is temporarily stored to prepare for the subsequent speech recognition process.

[0159] Step 2:

[0160] The server sends the collected audio data to the Google Cloud Speech-to-Text API, where it is converted into text data. The audio data is sent as input to the speech recognition API, and an HTTP response containing text data is received. This converts the audio data into text data from start to finish.

[0161] Step 3:

[0162] The server inputs the obtained character data into a GPT-3 model to generate customer service scripts. Alternatively, it inputs the character data into an AI model to generate customer service scripts as corresponding text data.

[0163] Step 4:

[0164] The server temporarily stores the generated customer service script and then sends it to the terminal as an HTTP response. This delivers the generated script to the user's terminal.

[0165] Terminal processing

[0166] Step 1:

[0167] The terminal receives the customer service script sent from the server. It retrieves the script information as an HTTP response and saves it to temporary storage. It then reads the contents of the script as input data and saves it.

[0168] Step 2:

[0169] The terminal handles customer service interactions based on saved customer service scripts. It dynamically references the script in response to user input and generates appropriate responses. Specifically, it responds to user inquiries and requests according to the instructions in the script.

[0170] Talk cancellation scoring

[0171] Server Processing

[0172] Step 1:

[0173] The server collects conversation data with users (e.g., in WAV format) in real time. Each time a conversation occurs, it is captured as audio data and stored on the server. This data serves as foundational information for analyzing customer satisfaction in subsequent processing.

[0174] Step 2:

[0175] The server sends the collected audio data to the Google Cloud Speech-to-Text API, where it is converted into text data. The audio data is sent to the speech recognition API, and an HTTP response is received as text data. This converts the conversation data into text data.

[0176] Step 3:

[0177] The server inputs the obtained character data into a scoring algorithm to calculate an evaluation score. The character data is input to the scoring algorithm, and a score is calculated based on the evaluation criteria.

[0178] Step 4:

[0179] The server temporarily stores the calculated evaluation score and then sends it to the terminal as an HTTP response. This delivers the generated evaluation score to the user's terminal.

[0180] Terminal processing

[0181] Step 1:

[0182] The terminal receives the evaluation score sent from the server. It retrieves the score information as an HTTP response and saves it to temporary storage. It then reads the score content as input data and saves it.

[0183] Step 2:

[0184] The device displays saved evaluation scores to the user in real time. The interface for displaying evaluation scores is initialized, and customer satisfaction scores are displayed dynamically. This enables real-time evaluation of service quality.

[0185] Through these steps, the system generates video data using still image and audio data, enabling consistent customer service and real-time customer satisfaction evaluation.

[0186] (Application Example 1)

[0187] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0188] In traditional brick-and-mortar stores, inadequate customer service was sometimes difficult due to staff absences or insufficient manpower. Furthermore, maintaining consistent service quality was challenging, and there were limited means to evaluate and improve customer satisfaction in real time. This invention aims to solve these problems by using an AI system with advanced customer service skills to consistently improve customer satisfaction.

[0189] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0190] In this invention, the server includes means for receiving still image data and audio data as input in order to convert still image data into video data; means for generating video data based on the input still image data and audio data; means for transmitting the generated video data to a terminal; means for displaying the transmitted video data on the terminal; means for converting customer service audio data into speech recognition data; means for generating a customer service script based on the speech recognition data; means for evaluating conversation data in real time and calculating a customer satisfaction evaluation score; and means for transmitting the evaluation score to a terminal and displaying it on the terminal. This enables virtual customer service by an avatar host using a smartphone application, allowing for consistent, high-quality customer service even when staff are absent in a physical store. Furthermore, it makes it possible to evaluate customer satisfaction in real time and continuously improve the quality of service.

[0191] "Still image data" refers to data of a static image that does not move.

[0192] "Motion data" refers to video data generated using a series of still images.

[0193] "Audio data" refers to data that records a person's voice in digital format.

[0194] "Speech recognition data" refers to data obtained by analyzing speech data and representing its content as text data.

[0195] A "customer service script" is text data that describes the content of conversations and response methods during customer service, and is used to instruct and guide customer service operations.

[0196] "Conversation data" refers to language-based information data exchanged between customers and systems or staff.

[0197] An "evaluation score" is a numerical value calculated based on conversation data and customer responses, and is an indicator used to evaluate customer satisfaction and service quality.

[0198] A "terminal" refers to an electronic device used to receive and display data transmitted from a server.

[0199] A "scoring algorithm" is a set of computational methods and rules used to analyze conversational data and other input data and calculate an evaluation score.

[0200] "Motion image generation means" refers to a method or apparatus for creating motion image data based on still image data and audio data.

[0201] This invention provides a specific method for constructing a virtual customer service system using a smartphone application in a physical store. This system generates moving image data based on still image data and audio data to provide consistent customer service to customers. It also collects and analyzes conversation data during customer service in real time to evaluate customer satisfaction.

[0202] Hardware and software configuration

[0203] Servers and smartphones will be used as the primary hardware. Specifically, the following:

[0204] server:

[0205] A cloud server for storing and processing still image and audio data.

[0206] Speech-to-Text is used to convert speech data into text data.

[0207] A video image generation API (e.g., DeepMotion) is used to generate video image data from still images and audio.

[0208] Using a GPT model (e.g., OpenAI's GPT-4), customer service scripts are generated from text data.

[0209] A scoring algorithm is used to evaluate customer satisfaction in real time.

[0210] Smartphone device:

[0211] A device for receiving, displaying, and playing back data transmitted from a server.

[0212] Collect user interactions and send them to the server.

[0213] Data processing and data calculation

[0214] Video generation:

[0215] The server accepts still image data and audio data as input. This data is passed to a video generation API to generate video data. The generated video data is then sent from the server to the terminal. The terminal displays this data, providing the user with an avatar image.

[0216] Customer service script generation:

[0217] The server converts customer service voice data into text data using a speech recognition API. The text data is input into a GPT model to generate an appropriate customer service script. This script is sent from the server to the terminal, which then uses this script to interact with the customer.

[0218] Customer satisfaction rating:

[0219] The server collects conversation data in real time and converts it into text data using a speech recognition API. The converted text data is analyzed by a scoring algorithm to calculate an evaluation score. This evaluation score is sent to the terminal and displayed to the user in real time.

[0220] Specific example

[0221] For example, in a clothing store, if a customer asks for trousers to match a jacket when no staff are present, this system would function effectively. If the customer asks, "Please recommend trousers that would go with this jacket," the system generates an avatar, which is displayed as video data. Furthermore, based on a customer service script, the avatar can provide an appropriate response and address the customer's questions. It is possible to collect data during the conversation in real time and display customer satisfaction as an evaluation score.

[0222] Example of a prompt

[0223] Please generate a customer service script for recommending trousers to match a jacket to a customer visiting a clothing store. Additionally, design an algorithm to evaluate customer satisfaction in real time during the interaction.

[0224] Thus, the system of the present invention can consistently improve customer satisfaction in physical stores by using an AI system with advanced customer service skills.

[0225] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0226] Step 1:

[0227] The server receives still image and audio data. When a user uploads still images and audio of employees during the store's preparation phase, the server retrieves them. The input data consists of still image and audio data, while the output data is a dataset to be passed to the video generation API.

[0228] Step 2:

[0229] The server uses the received still image and audio data to call a video generation API and generate video data. By providing still image and audio data as input to this video generation API (e.g., DeepMotion), animated video of employees is generated. The input data consists of image and audio files for the API call, and the output data is the generated video data.

[0230] Step 3:

[0231] The server sends the generated video data to the terminal. The server receives the video data and sends it to the terminal using a communication protocol. The input data is the video data, and the output data is the confirmation message sent to the terminal.

[0232] Step 4:

[0233] The terminal receives and displays video data transmitted from the server. The user operates the terminal to play the video data, which then displays an avatar. The input data is video data, and the output data is the display of the avatar image.

[0234] Step 5:

[0235] The server receives audio data during customer service interactions and converts it into text data using a speech recognition API. It collects the audio of the user-avatar conversation in real time and sends it to a speech recognition API such as Google Cloud Speech-to-Text to obtain text data. The input data is audio data, and the output data is the converted text data.

[0236] Step 6:

[0237] The server inputs text data into a GPT model and generates a customer service script. This text data is then input into an OpenAI GPT-4 model, and a process is executed to generate an appropriate customer service script. The input data is text data, and the output data is the customer service script.

[0238] Step 7:

[0239] The server sends the generated customer service script to the terminal. Sending the generated customer service script to the terminal enables real-time interaction with the user on the terminal. The input data is the customer service script, and the output data is the confirmation message sent to the terminal.

[0240] Step 8:

[0241] The terminal interacts with the user based on the received customer service script. The script is displayed on the terminal, and the avatar responds accordingly. The input data is the customer service script, and the output data is the result of the conversation with the user.

[0242] Step 9:

[0243] The server collects conversation data with the user in real time and converts it into text data using a speech recognition API. The input data is the audio data of the conversation, and the output data is the converted text data.

[0244] Step 10:

[0245] The server inputs the converted character data into a scoring algorithm and calculates an evaluation score. The input data is the converted character data, and the output data is the evaluation score.

[0246] Step 11:

[0247] The server sends the calculated evaluation score to the terminal, which then displays the evaluation score. The input data is the evaluation score, and the output data is the customer satisfaction evaluation result displayed on the terminal.

[0248] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0249] This invention relates to a system that recognizes user emotions and provides dynamic customer service responses based on those emotions. This system generates dynamic image data from still image data and audio data, and further generates customer service scripts by converting customer service audio data into text data. In addition, by combining this with an emotion engine that evaluates conversation data and recognizes user emotions, it achieves more sophisticated customer service.

[0250] 1. Generating the avatar chat screen

[0251] Server Processing

[0252] The server retrieves pre-prepared still images and audio data of employees. This data is sent to the API, which requests the generation of video data. The video data returned from the API is received by the server and then sent to the terminal.

[0253] Terminal processing

[0254] The terminal receives video data transmitted from the server, plays this data, and displays it to the user.

[0255] Specific example

[0256] For example, if a server has still images and audio data of employee A, it sends this data to an API to generate video data. Upon receiving the video data as a response from the API, the server sends it to a terminal. The terminal plays this video data, providing the user with a video that makes it appear as if employee A is speaking.

[0257] 2. Generating and using customer service scripts

[0258] Server Processing

[0259] The server collects voice data of high-performing crew members' customer service interactions. This voice data is converted into text data using a speech recognition API. The text data is input into a GPT model to generate customer service scripts. The generated scripts are sent from the server to the terminals.

[0260] Terminal processing

[0261] The terminal receives customer service scripts sent from the server and uses these scripts to interact with the user.

[0262] Specific example

[0263] The voice data of Crew B's customer service is stored on the server. The server sends this voice data to a speech recognition API to transcribe it and input it into a GPT model. The generated script is sent to the terminal, which then provides an appropriate response based on the user's questions and requests.

[0264] 3. Talk Cancellation Scoring

[0265] Server Processing

[0266] The server collects conversation data with the user in real time. This conversation data is converted into text data via a speech recognition API and input into a scoring algorithm. The algorithm calculates an evaluation score, which is then sent from the server to the terminal.

[0267] Terminal processing

[0268] The device receives evaluation scores sent from the server and displays them to the user in real time.

[0269] Specific example

[0270] If a conversation with user C is in progress, the terminal collects conversation data in real time and sends it to the server. The server transcribes the audio data and calculates an evaluation score using a scoring algorithm. This score is sent to the terminal, which displays it to the user in real time.

[0271] 4. Adding an emotion engine

[0272] Server Processing

[0273] The server sends video and audio data to the emotion engine. This engine analyzes the user's facial expressions and tone of voice to generate emotion information. This emotion information is sent back to the server, which is then instructed to dynamically modify the customer service script based on the user's emotions.

[0274] Terminal processing

[0275] The device receives emotional information and adjusts customer service scripts based on it. It also uses emotional information to provide more appropriate responses.

[0276] Specific example

[0277] When user D is feeling stressed, the emotion engine analyzes this and transmits the stress information to the server. Based on this information, the server generates a dedicated customer service script and sends it to the terminal. The terminal uses this script to provide a more relaxing response to user D.

[0278] The system of this invention can enhance customer satisfaction by analyzing the user's emotions in real-time and dynamically changing the customer service response based on that.

[0279] The processing flow will be described below.

[0280] Processing flow of the system combined with the emotion engine

[0281] 1. Generation of the avatar talk screen

[0282] Step 1:

[0283] The server acquires the prepared still images and voice data of the employees.

[0284] Step 2:

[0285] The server creates and sends an HTTP POST request to transmit the acquired still images and voice data to the API.

[0286] Step 3:

[0287] The server waits for the response from the API and receives the generated moving image data. This moving image data is generated based on the still images and voice data.

[0288] Step 4:

[0289] The server transmits the received moving image data to the terminal. At this time, an appropriate communication protocol (e.g., HTTP, WebSocket, etc.) is used.

[0290] Step 5:

[0291] The terminal receives video data transmitted from the server. After receiving the data, the terminal plays it back and displays it to the user.

[0292] 2. Generating and using customer service scripts

[0293] Step 1:

[0294] The server collects audio data of excellent customer service from top-performing crew members. This audio data is pre-recorded and stored in a database.

[0295] Step 2:

[0296] The server sends the collected customer service voice data to a speech recognition API, where it is converted into text data. The voice data is sent to the API via an HTTP POST request.

[0297] Step 3:

[0298] The server receives the text data returned from the speech recognition API.

[0299] Step 4:

[0300] The server inputs the received text data into a GPT model to generate a customer service script. The GPT model is pre-trained and can automatically generate appropriate responses.

[0301] Step 5:

[0302] The server sends the generated customer service script to the terminal, using the appropriate communication protocol.

[0303] Step 6:

[0304] The terminal receives customer service scripts sent from the server and provides appropriate responses based on the user's questions and requests.

[0305] 3. Conversation Termination Scoring

[0306] Step 1:

[0307] The terminal collects the conversation data with the user in real time. This collection is carried out through the voice input device.

[0308] Step 2:

[0309] The terminal sends the collected voice data to the server. An appropriate communication protocol (e.g., HTTP, WebSocket, etc.) is used.

[0310] Step 3:

[0311] The server sends the received voice data to the speech recognition API and converts it into character data.

[0312] Step 4:

[0313] The server receives the character data returned from the speech recognition API.

[0314] Step 5:

[0315] The server inputs the received character data into the scoring algorithm and calculates a score for evaluating customer satisfaction and willingness to continue using.

[0316] Step 6:

[0317] The server sends the calculated evaluation score to the terminal.

[0318] Step 7:

[0319] The terminal receives the evaluation score sent from the server and displays it to the user in real time.

[0320] 4. Addition of Emotional Engine

[0321] Step 1:

[0322] The server sends video and audio data to the emotion engine. The emotion engine analyzes the user's emotions based on their facial expressions and tone of voice.

[0323] Step 2:

[0324] The emotion engine analyzes the user's emotions and sends that emotional information back to the server.

[0325] Step 3:

[0326] The server dynamically modifies customer service scripts based on emotional information received from the emotion engine. For example, if a user is dissatisfied, it generates a more courteous response.

[0327] Step 4:

[0328] The server sends the modified customer service script to the terminal, using the appropriate communication protocol.

[0329] Step 5:

[0330] The terminal receives emotional information and customer service scripts sent from the server and responds appropriately to the user.

[0331] Specific example

[0332] If user D is experiencing stress, the emotion engine analyzes this and sends this information to the server. Based on the emotion information, the server generates a dedicated customer service script and sends it to the terminal. The terminal uses this script to provide a relaxed interaction with user D.

[0333] Thus, the system of the present invention provides a mechanism that can improve customer satisfaction by analyzing the user's emotions in real time through each processing step and dynamically changing customer service responses based on those emotions.

[0334] (Example 2)

[0335] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0336] Traditional customer service systems struggled to analyze user emotions in real time based on facial expressions and tone of voice, and to dynamically adjust customer service accordingly. As a result, they were unable to enhance user satisfaction and could only provide standardized responses, failing to offer optimal service to individual users.

[0337] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0338] In this invention, the server includes means for receiving still image data and audio data as input in order to convert still image data into video data; means for generating video data based on the input still image data and audio data; means for transmitting the generated video data to a terminal; means for receiving customer service audio data as input in order to transcribe customer service audio data; means for converting the input customer service audio data into text data; means for generating a customer service script based on the text data; means for transmitting the generated customer service script to a terminal; means for receiving conversation data as input in order to evaluate conversation data; means for converting the input conversation data into text data; means for inputting the text data into a scoring algorithm to calculate an evaluation score; means for receiving video data and audio data as input and analyzing the user's emotions; and means for dynamically changing the customer service script based on the analyzed emotion information. This enables dynamic responses that respond to the user's emotions.

[0339] 1. "Still image data" refers to a single, still image or photograph that is saved as an image file.

[0340] 2. "Motion image data" is a series of images consisting of multiple frames, and is visual information that changes over time.

[0341] 3. "Audio data" refers to data that stores audio in digital format, and includes information such as spoken language and sounds.

[0342] 4. "Customer service audio data" refers to data that records audio from customer service situations.

[0343] 5. "Text data" refers to data stored in text format, which is information expressed as words or sentences.

[0344] 6. A "customer service script" is a series of text messages generated to assist in interactions with users.

[0345] 7. "Conversation data" refers to a record of a conversation, either in audio or text format, between a user and a customer service representative.

[0346] 8. A "scoring algorithm" is a calculation method for determining an evaluation score based on input data.

[0347] 9. The "evaluation score" is a numerical value calculated by a scoring algorithm and is used to evaluate the quality of conversation and emotional state.

[0348] 10. An "emotion engine" is software that analyzes video and audio data to determine the user's emotional state.

[0349] 11. "Analyzed emotional information" refers to data about the user's emotional state obtained as a result of analysis by the emotion engine.

[0350] 12. A "server" is a computer system used to store, process, and distribute data over a network.

[0351] 13. A "terminal" is a device that a user directly operates and has the function of communicating with a server to send and receive data.

[0352] 14. A "generative AI model" is an artificial intelligence model used to learn from large amounts of data and generate new text or responses.

[0353] This invention relates to a system that recognizes user emotions and provides dynamic customer service responses based on those emotions. This system consists of a server and a terminal, each performing the following functions.

[0354] The server first retrieves pre-prepared still image and audio data of employees from the database. This still image and audio data is sent to the video generation API and converted into video data. The converted video data is sent back to the server, which then sends it to the terminal. The terminal plays the received video data and displays it to the user.

[0355] Next, the server collects voice data of high-performing crew members and converts this voice data into text data using a speech recognition API. The converted text data is input into a generative AI model (e.g., GPT-4) to generate a customer service script. The generated script is sent from the server to the terminal, which uses this script to interact with the user.

[0356] The server further collects conversation data with the user in real time and converts this data into text data via a speech recognition API. The converted text data is input into a scoring algorithm, which calculates an evaluation score. This evaluation score is sent from the server to the terminal, which displays the score to the user in real time.

[0357] The server also sends video and audio data to the emotion engine. This emotion engine analyzes the user's facial expressions and tone of voice and generates emotion information. The generated emotion information is sent back to the server, which dynamically modifies the customer service script based on this information. The terminal receives this modified script and provides a more appropriate response.

[0358] For example, consider a case where a server holds still images and audio data of employee A. This data is sent to an API to generate video data. The generated video data is sent to a terminal via the server, and the terminal displays it, providing the user with a video that makes it appear as if employee A is actually speaking.

[0359] Furthermore, if Crew B's customer service voice data is stored on the server, the server sends this data to a speech recognition API and converts it into text data. Next, this text data is input into a generation AI model to generate a customer service script. This script is sent to the terminal, which then provides an appropriate response based on the user's questions and requests.

[0360] Furthermore, if a conversation with user C is ongoing, the terminal collects conversation data in real time and sends it to the server. The server converts the audio data into text data and calculates an evaluation score using a scoring algorithm. This score is sent to the terminal, which then displays it to the user in real time.

[0361] Finally, if user D appears to be experiencing stress, the emotion engine analyzes this and sends stress information to the server. Based on this information, the server generates a customized customer service script and sends it to the terminal. The terminal uses this script to provide a more relaxed interaction with user D.

[0362] Examples of prompt messages include the following:

[0363] "Could you tell me about your recent orders?"

[0364] "I'd like to learn more about premium membership."

[0365] "Please tell me how to return an item."

[0366] This system allows for real-time analysis of user emotions and dynamic adjustments to customer service responses based on those analyses, thereby increasing customer satisfaction.

[0367] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0368] Step 1:

[0369] Acquisition of still image data and audio data

[0370] The server retrieves pre-prepared still image and audio data of employees from the database. The input is the still image and audio data in the database, and the output is the retrieved still image and audio data. This data is then ready to be sent to the API.

[0371] Step 2:

[0372] Sending data to the API

[0373] The server sends the acquired still image and audio data to the video generation API. The input is the acquired still image and audio data, and the output is the request to send to the API. The server constructs the appropriate request to the API and sends the data.

[0374] Step 3:

[0375] Receiving data from API

[0376] The server receives video data sent back from the API. The input is the response data from the API, and the output is the received video data. The server receives this video data and prepares to send it to the terminal.

[0377] Step 4:

[0378] Sending video data to the terminal

[0379] The server sends the received video data to the terminal. The input is the video data from the API, and the output is the video data sent to the terminal. The server sends data to the terminal's address.

[0380] Step 5:

[0381] Receiving and displaying video data.

[0382] The terminal receives video data transmitted from the server. The input is video data from the server, and the output is video data for display. The terminal plays this video data and displays it to the user.

[0383] Step 6:

[0384] Collection and transcription of customer service voice data

[0385] The server collects customer service voice data from excellent crew members and converts this voice data into text data using a speech recognition API. The input is customer service voice data, and the output is text data. By converting to text data, the server prepares it for feeding into a generative AI model.

[0386] Step 7:

[0387] Input to the Generative AI Model

[0388] The server inputs the converted character data into a generating AI model (e.g., GPT-4) to generate a customer service script. The input is character data, and the output is a customer service script. The generating AI model uses the character data to create a script suitable for the service.

[0389] Step 8:

[0390] Sending the script to the terminal

[0391] The server sends the generated script to the terminal. The input is the generated customer service script, and the output is the script sent to the terminal. The server sends the script to the terminal's address.

[0392] Step 9:

[0393] Use of the submitted script

[0394] The terminal receives customer service scripts sent from the server. The input is the customer service script from the server, and the output is the response based on the script. The terminal uses the script to interact with the user.

[0395] Step 10:

[0396] Collection and transcription of conversation data

[0397] The server collects conversation data with the user and converts this data into text data via a speech recognition API. The input is conversation data, and the output is text data. The server then inputs the converted text data into a scoring algorithm.

[0398] Step 11:

[0399] Calculation of evaluation score

[0400] The server inputs text data into a scoring algorithm and calculates an evaluation score. The input is text data, and the output is the evaluation score. The algorithm works to evaluate the quality and emotional state of the conversation.

[0401] Step 12:

[0402] Sending evaluation scores to the device

[0403] The server sends the calculated evaluation score to the terminal. The input is the evaluation score, and the output is the score sent to the terminal. The server sends the score to the terminal's address.

[0404] Step 13:

[0405] Display of evaluation score

[0406] The terminal receives evaluation scores sent from the server and displays them to the user in real time. The input is the evaluation score from the server, and the output is the score information for display. The user can check this score.

[0407] Step 14:

[0408] Sending and analyzing emotional data

[0409] The server sends video and audio data to the emotion engine. The input is video and audio data, and the output is the data sent to the emotion engine. The emotion engine analyzes this data and generates emotion information.

[0410] Step 15:

[0411] Changes to scripts based on emotional information

[0412] The server receives the emotional information analyzed by the emotion engine and dynamically modifies the customer service script based on this information. The input is emotional information, and the output is the dynamically modified customer service script. The server generates a new script and sends it to the terminal.

[0413] Step 16:

[0414] Using the modified script

[0415] The terminal receives a newly transmitted customer service script and interacts with the user based on it. The input is the new script from the server, and the output is the interaction based on the new script. The terminal provides an interaction that reflects the script.

[0416] (Application Example 2)

[0417] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0418] In modern retail customer service, understanding and responding appropriately to customer emotions is crucial for improving customer satisfaction. However, traditional customer service systems have struggled to analyze customer emotional information in real time and dynamically adjust customer service scripts based on that analysis. This can lead to one-way communication with customers and potentially decrease customer satisfaction.

[0419] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for receiving still image data and audio data as input, means for generating moving image data based on the input still image data and audio data, means for transmitting the generated moving image data to a terminal, means for displaying the transmitted moving image data on the terminal, means for receiving image data and audio data to analyze the user's emotions, means for dynamically adjusting a customer service script based on the analyzed emotion information, means for transmitting the dynamically adjusted customer service script to a terminal, and means for interacting with the user using the transmitted customer service script on the terminal. This makes it possible to analyze the customer's emotions in real time and dynamically adjust customer service responses based on that analysis.

[0420] "Still image data" refers to individual frames that make up a video.

[0421] "Audio data" refers to data that records audio information digitally or in analog format.

[0422] "Motion data" refers to a video that is formed by sequentially combining multiple still images over time.

[0423] "Terminal" refers to an electronic device used by a user, and specifically includes personal computers, smartphones, smart glasses, etc.

[0424] "Emotional information" refers to data that indicates a user's psychological state and mood, analyzed from their facial expressions and tone of voice.

[0425] A "customer service script" refers to pre-designed text data of responses used in interactions with users.

[0426] A "generative AI model" refers to an artificial intelligence algorithm used to generate new data or text based on existing data.

[0427] A "prompt" is a text input to a generative AI model that instructs the model to produce a specific output.

[0428] This invention is a system that improves customer service in physical stores by analyzing user emotions in real time and dynamically adjusting customer service based on that analysis. This system operates with a server and terminals working together and consists of the following main components.

[0429] First, the server accepts still image data and audio data as input. This data, collected by hardware such as cameras and microphones, is sent to the server. The server uses this still image data and audio data to generate video data. Specifically, video data is generated using APIs and machine learning algorithms. After generation, this video data is sent to the terminal.

[0430] The terminal receives video and image data transmitted from the server and displays it to the user. Devices such as smartphones and smart glasses are used as terminals. This terminal functions as an interface with the user and plays the video and image data.

[0431] Next, to analyze the user's emotions, the server sends image and audio data to an emotion analysis engine. This emotion analysis engine identifies the user's psychological state from their facial expressions and tone of voice and generates emotion information. This emotion information is then sent back to the server.

[0432] The server dynamically adjusts the customer service script based on this emotional information. A generative AI model is used to create prompts containing emotional information and generate an appropriate customer service script. OpenAI's GPT model is used as the generative AI model in this process. The generated customer service script is sent to the terminal, which then interacts with the user based on it.

[0433] As a concrete example, consider a scenario where a customer near a fitting room asks for feedback on a product they are trying on. In this case, the staff member performs real-time sentiment analysis via smart glasses and provides appropriate feedback based on the generated customer service script. For example, if the customer asks, "Does this jacket suit me?", the generating AI model would receive a prompt like this:

[0434] "When a customer asks, 'Does this jacket suit me?', generate the optimal customer service script. The customer's voice tone is relaxed, and their facial expression is smiling."

[0435] This system allows for real-time analysis of customer emotions and dynamic adjustments to customer service based on that analysis, which is expected to improve customer satisfaction.

[0436] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0437] Step 1:

[0438] The server accepts still image and audio data as input from the camera and microphone. This data is collected in real time to mimic human customer service staff. Still images captured by the camera and audio recorded by the microphone are sent to the server.

[0439] Step 2:

[0440] The server generates video data based on the received still image and audio data. Here, a deep learning model is used to generate the video data. This deep learning model, for example, is one that has been pre-trained on employee movements and speech patterns. The generated video data is temporarily stored on the server.

[0441] Step 3:

[0442] The server sends the generated video data to the terminal. The terminal consists of devices such as smart glasses or smartphones. The video data received by the terminal is played back for the customer to see.

[0443] Step 4:

[0444] The server sends video and audio data to the emotion analysis engine to analyze the user's emotions. The emotion analysis engine generates emotional information from the user's facial expressions and tone of voice. For example, it can determine whether the user is relaxed based on their smile or tone of voice. The analysis results are then returned to the server.

[0445] Step 5:

[0446] The server creates a prompt sentence based on the emotion information returned from the emotion analysis engine and inputs it into the generative AI model. The generative AI model uses OpenAI's GPT model to generate a customer service script based on this prompt sentence. Because the prompt sentence includes the user's voice content and emotion information, a very natural conversation is possible.

[0447] Step 6:

[0448] The server sends the generated customer service script to the terminal. The terminal uses this script to interact with the customer. For example, if a customer in a fitting room is asking for product feedback, the terminal provides an appropriate response.

[0449] Step 7:

[0450] The device collects new data based on user interaction and sends it to the server. This allows the system to continuously improve, enabling more natural and effective responses.

[0451] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0452] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0453] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0454] [Second Embodiment]

[0455] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0456] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0457] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0458] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0459] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0460] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0461] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0462] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0463] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0464] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0465] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0466] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0467] This invention relates to a system that generates moving image data based on still image data and audio data to provide consistent customer service. It also includes collecting and analyzing conversation data during customer service in real time to evaluate customer satisfaction. Specific embodiments of this system are described below.

[0468] 1. Generating the avatar chat screen

[0469] Server Processing

[0470] The server retrieves pre-prepared still images and audio data of employees. Next, it sends this data via an API (Application Programming Interface) to request the generation of video data. The video data returned from the API is received by the server, which then sends the video data to the terminal.

[0471] Terminal processing

[0472] The terminal receives video data transmitted from the server. After receiving the data, the terminal plays it back and displays it to the user.

[0473] Specific example

[0474] For example, if a server has still images and audio data of employee A, the server sends these to an API to generate video data. Upon receiving the video data as a response from the API, the server sends it to a terminal. The terminal plays this data, providing the user with a video that makes it appear as if employee A is speaking.

[0475] 2. Generating and using customer service scripts

[0476] Server Processing

[0477] The server collects voice data of excellent crew members' customer service interactions. This collected voice data is converted into text data using a speech recognition API. Next, the server feeds this text data into a GPT model (Greater Global Pattern Testing) to generate appropriate customer service scripts. The generated scripts are then sent from the server to the terminals.

[0478] Terminal processing

[0479] The terminal receives a customer service script sent from the server. Based on this script, the terminal interacts with the user.

[0480] Specific example

[0481] If Crew B possesses excellent customer service skills, the server records Crew B's customer service audio. This audio data is transcribed and fed into a GPT model to generate a customer service script. The terminal uses this script to provide appropriate responses to the user's questions.

[0482] 3. Talk Cancellation Scoring

[0483] Server Processing

[0484] The server collects conversation data with users in real time. The collected conversation data is converted into text data via a speech recognition API and then input into a scoring algorithm. This algorithm analyzes the conversation data and calculates a score to evaluate customer satisfaction and willingness to continue using the service. Finally, this evaluation score is sent from the server to the terminal.

[0485] Terminal processing

[0486] The device receives evaluation scores sent from the server and displays them to the user in real time.

[0487] Specific example

[0488] While a conversation with user C is in progress, the terminal collects the conversation content in real time and sends it to the server. The server transcribes this audio data and calculates an evaluation score using a scoring algorithm. By displaying this score on the terminal, it becomes possible to continuously evaluate and improve the quality of the service.

[0489] In this way, the present invention can consistently improve customer satisfaction through an AI system with advanced customer service skills.

[0490] The following describes the processing flow.

[0491] 1. Generating the avatar chat screen

[0492] Step 1:

[0493] The server retrieves pre-prepared still images and audio data of employees.

[0494] Step 2:

[0495] The server creates and sends an HTTP POST request to send the acquired still image and audio data to the API.

[0496] Step 3:

[0497] The server waits for a response from the API and receives the generated video data. This video data is generated from still images and audio data.

[0498] Step 4:

[0499] The server sends the received video data to the terminal. It uses an appropriate communication protocol (e.g., HTTP, WebSocket) for this purpose.

[0500] Step 5:

[0501] The terminal receives video data transmitted from the server. After receiving the data, the terminal plays it back and displays it to the user.

[0502] Specific example

[0503] Still images and audio data of employee A are stored on the server. The server sends this data to an API, receives video data generated by the API, and sends it to the terminal. The terminal receives the video data and displays to the user an image that makes it appear as if employee A is speaking.

[0504] 2. Generating and using customer service scripts

[0505] Step 1:

[0506] The server collects audio data of excellent customer service from top-performing crew members. This audio data is pre-recorded and stored in a database.

[0507] Step 2:

[0508] The server sends the collected customer service voice data to a speech recognition API, where it is converted into text data. The voice data is sent to the API via an HTTP POST request.

[0509] Step 3:

[0510] The server receives the text data returned from the speech recognition API.

[0511] Step 4:

[0512] The server inputs the received text data into a GPT model to generate a customer service script. The GPT model is pre-trained and can automatically generate appropriate responses.

[0513] Step 5:

[0514] The server sends the generated customer service script to the terminal, using the appropriate communication protocol.

[0515] Step 6:

[0516] The terminal receives customer service scripts sent from the server and provides appropriate responses based on the user's questions and requests.

[0517] Specific example

[0518] Crew B's excellent customer service voice is stored on the server. The server sends this voice data to a speech recognition API to transcribe it and input it into a GPT model. The generated script is sent to the terminal, which uses it to respond to the user.

[0519] 3. Talk Cancellation Scoring

[0520] Step 1:

[0521] The device collects conversation data with the user in real time. This collection is done through a voice input device.

[0522] Step 2:

[0523] The terminal sends the collected audio data to the server. An appropriate communication protocol (e.g., HTTP, WebSocket) is used.

[0524] Step 3:

[0525] The server sends the received audio data to a speech recognition API, where it is converted into text data.

[0526] Step 4:

[0527] The server receives the text data returned from the speech recognition API.

[0528] Step 5:

[0529] The server inputs the received text data into a scoring algorithm to calculate a score for evaluating customer satisfaction and willingness to continue using the service.

[0530] Step 6:

[0531] The server sends the calculated evaluation score to the terminal.

[0532] Step 7:

[0533] The device receives evaluation scores sent from the server and displays them to the user in real time.

[0534] Specific example

[0535] As a conversation with user C progresses, the terminal collects conversation data in real time and sends it to the server. The server sends the data to a speech recognition API to transcribe it and calculates an evaluation score using a scoring algorithm. This score is sent to the terminal, which then displays it to the user in real time.

[0536] In this way, the system of the present invention can provide advanced customer service skills and real-time customer evaluation, thereby consistently improving customer satisfaction.

[0537] (Example 1)

[0538] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0539] In modern society, consistent customer service is required to enhance customer satisfaction. Especially in online customer service, direct interaction with customers is difficult, and service often depends on the individual skills of employees. This can lead to inconsistencies in service quality, potentially resulting in decreased customer satisfaction. Furthermore, the lack of adequate means to evaluate customer satisfaction in real time and improve service quality makes prompt responses difficult. A system is needed to address these challenges and improve customer satisfaction.

[0540] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0541] In this invention, the server includes means for receiving still image data and audio data as input, means for generating video data based on the input still image data and audio data, means for transmitting the generated video data to a terminal, means for converting customer service audio data into text data and generating a customer service script using a generation AI model, means for transmitting the generated customer service script to a terminal, means for converting conversation data into text data and calculating an evaluation score using a scoring algorithm, and means for transmitting the calculated evaluation score to a terminal. As a result, the server generates video data and transmits it to a terminal to achieve consistent customer service, and it becomes possible to generate scripts that utilize the customer service skills of crew members with excellent customer service skills and use them on the terminal. Furthermore, it becomes possible to perform real-time customer satisfaction evaluation and continuously improve the quality of customer service.

[0542] "Still image data" refers to digital data containing image information captured at a specific point in time.

[0543] "Audio data" refers to data that records human voices in digital format.

[0544] "Motion image data" refers to digital data in video format that includes a sequence of images and synchronized audio.

[0545] A "terminal" is an electronic device that a user can directly operate.

[0546] "Customer service audio data" refers to data that digitally records the voices spoken by employees during customer service interactions.

[0547] "Text data" refers to data that visualizes audio data as text using speech recognition technology.

[0548] A "generative AI model" is an artificial intelligence model trained to generate appropriate responses or scripts from specific input data.

[0549] A "customer service script" is text data containing a series of sentences and response instructions used during customer service.

[0550] "Conversation data" refers to digital data that records voice interactions between a user and a system.

[0551] A "scoring algorithm" is a calculation method used to analyze input data and calculate an evaluation score based on specific criteria.

[0552] An "evaluation score" is a numerical evaluation result calculated by a scoring algorithm.

[0553] This invention relates to a system that generates moving image data based on still image data and audio data to provide consistent customer service. It also includes a function to collect and analyze conversation data during customer service in real time and evaluate customer satisfaction. Specific embodiments of this invention will now be described.

[0554] First, the server retrieves pre-prepared still image data (JPEG format) and audio data (WAV format) of the employee. This data is sent via an HTTP POST request using an API such as Microsoft Azure's Face API. The response returned from the API contains the generated video data (MP4 format), which the server receives and temporarily stores. Subsequently, the server sends the stored video data to the terminal as an HTTP response.

[0555] Next, the terminal receives video data sent from the server. After receiving the data, the terminal initializes a media player to play it and displays the video to the user. This series of processes allows the terminal to play a video that makes it appear as if an employee is speaking.

[0556] Next, to collect excellent customer service audio data (in WAV format) from crew members, the server uses the Google Cloud Speech-to-Text API to convert the audio data into text data. Then, the text data is fed into a generative AI model such as the GPT-3 model to generate customer service scripts. The generated scripts are temporarily stored on the server and sent to the terminal. The terminal receives the customer service scripts sent from the server and uses them to interact with the user.

[0557] Next, the server collects conversation data (in WAV format) from the user in real time. The conversation data is converted into text data using the Google Cloud Speech-to-Text API and then fed into a scoring algorithm. This algorithm analyzes the conversation data and calculates a score that evaluates customer satisfaction and willingness to continue using the service. This evaluation score is sent from the server to the terminal and displayed to the user in real time, enabling continuous evaluation and improvement of service quality.

[0558] Specific example

[0559] 1. The server sends still images and audio data of employee A to the API to generate video data. The terminal receives this data and displays a video to the user that makes it appear as if employee A is speaking.

[0560] 2. The customer service voice of Crew B is transcribed and input into the GPT-3 model to generate a customer service script. The terminal uses this script to respond to the user's questions.

[0561] 3. The conversation with User C is collected, and a customer satisfaction evaluation score is calculated using a scoring algorithm. The terminal displays this score and evaluates the service quality.

[0562] Example of a prompt

[0563] 1. "Please generate video data based on a still image of employee A and the following audio data."

[0564] 2. "Please generate a customer service script based on the text data of this customer service audio."

[0565] 3. "Based on the conversation data, calculate a score to evaluate customer satisfaction."

[0566] In this way, the present invention provides a system that generates moving image data using still image data and audio data, evaluates customer satisfaction in real time, and realizes consistently high-quality customer service.

[0567] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0568] Generating an avatar chat screen

[0569] Server Processing

[0570] Step 1:

[0571] The server retrieves pre-prepared still image data (e.g., JPEG format) and audio data (e.g., WAV format) of employees. The server receives the still image and audio data as input. This data serves as the basic information for generating video data in subsequent processing.

[0572] Step 2:

[0573] The server sends the acquired still image and audio data to an API (e.g., Microsoft Azure's Face API). It makes an HTTP POST request to the API endpoint, sending the still image and audio data. Based on the input data, the API executes a process and generates a video image data by combining the still image and audio data.

[0574] Step 3:

[0575] The server receives video data (e.g., in MP4 format) returned from the API. It temporarily stores the video data obtained as an HTTP response and prepares for the subsequent transmission process.

[0576] Step 4:

[0577] The server sends the generated video data to the terminal as an HTTP response. This allows the video data to be delivered to the terminal via the internet.

[0578] Terminal processing

[0579] Step 1:

[0580] The terminal receives video data sent from the server as an HTTP response. It stores the received data in temporary storage. It then parses and stores the URL or binary data used as input data.

[0581] Step 2:

[0582] The device initializes its media player for playing saved video data. It internally loads the video data and prepares it for display to the user.

[0583] Step 3:

[0584] The device plays video and displays it to the user. It uses a media player to display the video along a timeline and outputs it in a way that the user can perceive.

[0585] Generating and using customer service scripts

[0586] Server Processing

[0587] Step 1:

[0588] The server collects voice data (e.g., in WAV format) of high-performing crew members providing excellent customer service. The collected voice data is temporarily stored to prepare for the subsequent speech recognition process.

[0589] Step 2:

[0590] The server sends the collected audio data to the Google Cloud Speech-to-Text API, where it is converted into text data. The audio data is then sent as input to the speech recognition API, and an HTTP response containing the text data is received. This converts the audio data into text data from start to finish.

[0591] Step 3:

[0592] The server inputs the obtained character data into a GPT-3 model to generate customer service scripts. The character data is input into a generation AI model to generate customer service scripts as corresponding text data.

[0593] Step 4:

[0594] The server temporarily stores the generated customer service script and then sends it to the terminal as an HTTP response. This delivers the generated script to the user's terminal.

[0595] Terminal processing

[0596] Step 1:

[0597] The terminal receives the customer service script sent from the server. It retrieves the script information as an HTTP response and saves it to temporary storage. It then reads the contents of the script as input data and saves it.

[0598] Step 2:

[0599] The terminal handles customer service based on saved customer service scripts. It dynamically references the script in response to user input and generates appropriate responses. Specifically, it responds to user inquiries and requests according to the instructions in the script.

[0600] Talk cancellation scoring

[0601] Server Processing

[0602] Step 1:

[0603] The server collects conversation data with users (e.g., in WAV format) in real time. Each time a conversation occurs, it is captured as audio data and stored on the server. This data serves as foundational information for analyzing customer satisfaction in subsequent processing.

[0604] Step 2:

[0605] The server sends the collected audio data to the Google Cloud Speech-to-Text API, where it is converted into text data. The audio data is sent to the speech recognition API, and an HTTP response is received as text data. This converts the conversation data into text data.

[0606] Step 3:

[0607] The server inputs the obtained character data into a scoring algorithm to calculate an evaluation score. The character data is input to the scoring algorithm, and a score is calculated based on the evaluation criteria.

[0608] Step 4:

[0609] The server temporarily stores the calculated evaluation score and then sends it to the terminal as an HTTP response. This delivers the generated evaluation score to the user's terminal.

[0610] Terminal processing

[0611] Step 1:

[0612] The terminal receives the evaluation score sent from the server. It retrieves the score information as an HTTP response and saves it to temporary storage. It then reads the score content as input data and saves it.

[0613] Step 2:

[0614] The device displays saved evaluation scores to the user in real time. The interface for displaying evaluation scores is initialized, and customer satisfaction scores are displayed dynamically. This enables real-time evaluation of service quality.

[0615] Through these steps, the system generates video data using still image and audio data, enabling consistent customer service and real-time customer satisfaction evaluation.

[0616] (Application Example 1)

[0617] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0618] In traditional brick-and-mortar stores, inadequate customer service was sometimes difficult due to staff absences or insufficient manpower. Furthermore, maintaining consistent service quality was challenging, and there were limited means to evaluate and improve customer satisfaction in real time. This invention aims to solve these problems by using an AI system with advanced customer service skills to consistently improve customer satisfaction.

[0619] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0620] In this invention, the server includes means for receiving still image data and audio data as input in order to convert still image data into video data; means for generating video data based on the input still image data and audio data; means for transmitting the generated video data to a terminal; means for displaying the transmitted video data on the terminal; means for converting customer service audio data into speech recognition data; means for generating a customer service script based on the speech recognition data; means for evaluating conversation data in real time and calculating a customer satisfaction evaluation score; and means for transmitting the evaluation score to a terminal and displaying it on the terminal. This enables virtual customer service by an avatar host using a smartphone application, allowing for consistent, high-quality customer service even when staff are absent in a physical store. Furthermore, it makes it possible to evaluate customer satisfaction in real time and continuously improve the quality of service.

[0621] "Still image data" refers to data of a static image that does not move.

[0622] "Motion data" refers to video data generated using a series of still images.

[0623] "Audio data" refers to data that records a person's voice in digital format.

[0624] "Speech recognition data" refers to data obtained by analyzing speech data and representing its content as text data.

[0625] A "customer service script" is text data that describes the content of conversations and methods of responding during customer service, and is used to instruct and guide customer service operations.

[0626] "Conversation data" refers to language-based information data exchanged between customers and systems or staff.

[0627] An "evaluation score" is a numerical value calculated based on conversation data and customer responses, and is an indicator used to evaluate customer satisfaction and service quality.

[0628] A "terminal" refers to an electronic device used to receive and display data transmitted from a server.

[0629] A "scoring algorithm" is a set of computational methods and rules used to analyze conversational data and other input data and calculate an evaluation score.

[0630] "Motion image generation means" refers to a method or apparatus for creating motion image data based on still image data and audio data.

[0631] This invention provides a specific method for constructing a virtual customer service system using a smartphone application in a physical store. This system generates moving image data based on still image data and audio data to provide consistent customer service to customers. It also collects and analyzes conversation data during customer service in real time to evaluate customer satisfaction.

[0632] Hardware and software configuration

[0633] Servers and smartphones will be used as the primary hardware. Specifically, the following:

[0634] server:

[0635] A cloud server for storing and processing still image and audio data.

[0636] Speech-to-Text is used to convert speech data into text data.

[0637] A video image generation API (e.g., DeepMotion) is used to generate video image data from still images and audio.

[0638] A GPT model (e.g., OpenAI's GPT-4) is used to generate customer service scripts from text data.

[0639] A scoring algorithm is used to evaluate customer satisfaction in real time.

[0640] Smartphone device:

[0641] A device for receiving, displaying, and playing back data transmitted from a server.

[0642] Collect user interactions and send them to the server.

[0643] Data processing and data calculation

[0644] Video generation:

[0645] The server accepts still image data and audio data as input. This data is passed to a video generation API to generate video data. The generated video data is then sent from the server to the terminal. The terminal displays this data, providing the user with an avatar image.

[0646] Customer service script generation:

[0647] The server converts customer service voice data into text data using a speech recognition API. The text data is input into a GPT model to generate an appropriate customer service script. This script is sent from the server to the terminal, which then uses the script to interact with the customer.

[0648] Customer satisfaction rating:

[0649] The server collects conversation data in real time and converts it into text data using a speech recognition API. The converted text data is analyzed by a scoring algorithm to calculate an evaluation score. This evaluation score is sent to the terminal and displayed to the user in real time.

[0650] Specific example

[0651] For example, in a clothing store, if a customer asks for trousers to match a jacket when no staff are present, this system would function effectively. If the customer asks, "Please recommend trousers that would go with this jacket," the system generates an avatar, which is displayed as video data. Furthermore, based on a customer service script, the avatar can provide an appropriate response and address the customer's questions. It is possible to collect data during the conversation in real time and display customer satisfaction as an evaluation score.

[0652] Example of a prompt

[0653] Please generate a customer service script for recommending trousers to match a jacket to a customer visiting a clothing store. Additionally, design an algorithm to evaluate customer satisfaction in real time during the interaction.

[0654] Thus, the system of the present invention can consistently improve customer satisfaction in physical stores by using an AI system with advanced customer service skills.

[0655] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0656] Step 1:

[0657] The server receives still image and audio data. When a user uploads still images and audio of employees during the store's preparation phase, the server retrieves them. The input data consists of still image and audio data, while the output data is a dataset to be passed to the video generation API.

[0658] Step 2:

[0659] The server uses the received still image and audio data to call a video generation API and generate video data. By providing still image and audio data as input to this video generation API (e.g., DeepMotion), animated video of employees is generated. The input data consists of image and audio files for the API call, and the output data is the generated video data.

[0660] Step 3:

[0661] The server sends the generated video data to the terminal. The server receives the video data and sends it to the terminal using a communication protocol. The input data is the video data, and the output data is the confirmation message sent to the terminal.

[0662] Step 4:

[0663] The terminal receives and displays video data transmitted from the server. The user operates the terminal to play the video data, which then displays an avatar. The input data is video data, and the output data is the display of the avatar image.

[0664] Step 5:

[0665] The server receives audio data during customer service interactions and converts it into text data using a speech recognition API. It collects the audio of the user-avatar conversation in real time and sends it to a speech recognition API such as Google Cloud Speech-to-Text to obtain text data. The input data is audio data, and the output data is the converted text data.

[0666] Step 6:

[0667] The server inputs text data into a GPT model and generates a customer service script. This text data is then input into an OpenAI GPT-4 model, and the process of generating an appropriate customer service script is executed. The input data is text data, and the output data is the customer service script.

[0668] Step 7:

[0669] The server sends the generated customer service script to the terminal. Sending the generated customer service script to the terminal enables real-time interaction with the user on the terminal. The input data is the customer service script, and the output data is the confirmation message sent to the terminal.

[0670] Step 8:

[0671] The terminal interacts with the user based on the received customer service script. The script is displayed on the terminal, and the avatar responds accordingly. The input data is the customer service script, and the output data is the result of the conversation with the user.

[0672] Step 9:

[0673] The server collects conversation data with the user in real time and converts it into text data using a speech recognition API. The input data is the audio data of the conversation, and the output data is the converted text data.

[0674] Step 10:

[0675] The server inputs the converted character data into a scoring algorithm to calculate an evaluation score. The input data is the converted character data, and the output data is the evaluation score.

[0676] Step 11:

[0677] The server sends the calculated evaluation score to the terminal, which then displays the evaluation score. The input data is the evaluation score, and the output data is the customer satisfaction evaluation result displayed on the terminal.

[0678] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0679] This invention relates to a system that recognizes user emotions and provides dynamic customer service responses based on those emotions. This system generates dynamic image data from still image data and audio data, and further generates customer service scripts by converting customer service audio data into text data. In addition, by combining this with an emotion engine that evaluates conversation data and recognizes user emotions, it achieves more sophisticated customer service.

[0680] 1. Generating the avatar chat screen

[0681] Server Processing

[0682] The server retrieves pre-prepared still images and audio data of employees. This data is sent to the API, which requests the generation of video data. The video data returned from the API is received by the server and then sent to the terminal.

[0683] Terminal processing

[0684] The terminal receives video data transmitted from the server, plays this data, and displays it to the user.

[0685] Specific example

[0686] For example, if a server has still images and audio data of employee A, it sends this data to an API to generate video data. Upon receiving the video data as a response from the API, the server sends it to a terminal. The terminal plays this video data, providing the user with a video that makes it appear as if employee A is speaking.

[0687] 2. Generating and using customer service scripts

[0688] Server Processing

[0689] The server collects voice data of high-performing crew members' customer service interactions. This voice data is converted into text data using a speech recognition API. The text data is input into a GPT model to generate customer service scripts. The generated scripts are sent from the server to the terminals.

[0690] Terminal processing

[0691] The terminal receives customer service scripts sent from the server and uses these scripts to interact with the user.

[0692] Specific example

[0693] Crew B's customer service voice data is stored on the server. The server sends this voice data to a speech recognition API to transcribe it and input it into a GPT model. The generated script is sent to the terminal, which then provides an appropriate response based on the user's questions and requests.

[0694] 3. Talk Cancellation Scoring

[0695] Server Processing

[0696] The server collects conversation data with the user in real time. This conversation data is converted into text data via a speech recognition API and input into a scoring algorithm. The algorithm calculates an evaluation score, which is then sent from the server to the terminal.

[0697] Terminal processing

[0698] The device receives evaluation scores sent from the server and displays them to the user in real time.

[0699] Specific example

[0700] If a conversation with user C is in progress, the terminal collects conversation data in real time and sends it to the server. The server transcribes the audio data and calculates an evaluation score using a scoring algorithm. This score is sent to the terminal, which displays it to the user in real time.

[0701] 4. Adding an emotion engine

[0702] Server Processing

[0703] The server sends video and audio data to the emotion engine. This engine analyzes the user's facial expressions and tone of voice to generate emotion information. This emotion information is sent back to the server, which is then instructed to dynamically modify the customer service script based on the user's emotions.

[0704] Terminal processing

[0705] The device receives emotional information and adjusts customer service scripts based on it. It also uses emotional information to provide more appropriate responses.

[0706] Specific example

[0707] If user D is experiencing stress, the emotion engine analyzes this and sends stress information to the server. Based on this information, the server generates a dedicated customer service script and sends it to the terminal. The terminal uses this script to provide a more relaxed interaction with user D.

[0708] The system of this invention can improve customer satisfaction by analyzing user emotions in real time and dynamically changing customer service responses based on those emotions.

[0709] The following describes the processing flow.

[0710] Processing flow of a system combining an emotion engine

[0711] 1. Generating the avatar chat screen

[0712] Step 1:

[0713] The server retrieves pre-prepared still images and audio data of employees.

[0714] Step 2:

[0715] The server creates and sends an HTTP POST request to the API to send the acquired still image and audio data.

[0716] Step 3:

[0717] The server waits for a response from the API and receives the generated video data. This video data is generated from still images and audio data.

[0718] Step 4:

[0719] The server sends the received video data to the terminal. It uses an appropriate communication protocol (e.g., HTTP, WebSocket) for this purpose.

[0720] Step 5:

[0721] The terminal receives video data transmitted from the server. After receiving the data, the terminal plays it back and displays it to the user.

[0722] 2. Generating and using customer service scripts

[0723] Step 1:

[0724] The server collects audio data of excellent customer service from top-performing crew members. This audio data is pre-recorded and stored in a database.

[0725] Step 2:

[0726] The server sends the collected customer service voice data to a speech recognition API, where it is converted into text data. The voice data is sent to the API via an HTTP POST request.

[0727] Step 3:

[0728] The server receives the text data returned from the speech recognition API.

[0729] Step 4:

[0730] The server inputs the received text data into a GPT model to generate a customer service script. The GPT model is pre-trained and can automatically generate appropriate responses.

[0731] Step 5:

[0732] The server sends the generated customer service script to the terminal, using the appropriate communication protocol.

[0733] Step 6:

[0734] The terminal receives customer service scripts sent from the server and provides appropriate responses based on the user's questions and requests.

[0735] 3. Talk Cancellation Scoring

[0736] Step 1:

[0737] The device collects conversation data with the user in real time. This collection is done through a voice input device.

[0738] Step 2:

[0739] The terminal sends the collected audio data to the server. An appropriate communication protocol (e.g., HTTP, WebSocket) is used.

[0740] Step 3:

[0741] The server sends the received audio data to a speech recognition API, where it is converted into text data.

[0742] Step 4:

[0743] The server receives the text data returned from the speech recognition API.

[0744] Step 5:

[0745] The server inputs the received text data into a scoring algorithm to calculate a score for evaluating customer satisfaction and willingness to continue using the service.

[0746] Step 6:

[0747] The server sends the calculated evaluation score to the terminal.

[0748] Step 7:

[0749] The device receives evaluation scores sent from the server and displays them to the user in real time.

[0750] 4. Adding an emotion engine

[0751] Step 1:

[0752] The server sends video and audio data to the emotion engine. The emotion engine analyzes the user's emotions based on their facial expressions and tone of voice.

[0753] Step 2:

[0754] The emotion engine analyzes the user's emotions and sends that emotional information back to the server.

[0755] Step 3:

[0756] The server dynamically modifies customer service scripts based on emotional information received from the emotion engine. For example, if a user is dissatisfied, it generates a more courteous response.

[0757] Step 4:

[0758] The server sends the modified customer service script to the terminal, using the appropriate communication protocol.

[0759] Step 5:

[0760] The terminal receives emotional information and customer service scripts sent from the server and responds appropriately to the user.

[0761] Specific example

[0762] If user D is experiencing stress, the emotion engine analyzes this and sends this information to the server. Based on the emotion information, the server generates a dedicated customer service script and sends it to the terminal. The terminal uses this script to provide a relaxed interaction with user D.

[0763] In this way, the system of the present invention analyzes the user's emotions in real time through each processing step and dynamically changes customer service responses based on that analysis, thereby providing a mechanism that can improve customer satisfaction.

[0764] (Example 2)

[0765] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0766] Traditional customer service systems struggled to analyze user emotions in real time based on facial expressions and tone of voice, and to dynamically adjust customer service accordingly. As a result, they were unable to enhance user satisfaction and could only provide standardized responses, failing to offer optimal service to individual users.

[0767] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0768] In this invention, the server includes means for receiving still image data and audio data as input in order to convert still image data into video data; means for generating video data based on the input still image data and audio data; means for transmitting the generated video data to a terminal; means for receiving customer service audio data as input in order to transcribe customer service audio data; means for converting the input customer service audio data into text data; means for generating a customer service script based on the text data; means for transmitting the generated customer service script to a terminal; means for receiving conversation data as input in order to evaluate conversation data; means for converting the input conversation data into text data; means for inputting the text data into a scoring algorithm to calculate an evaluation score; means for receiving video data and audio data as input and analyzing the user's emotions; and means for dynamically changing the customer service script based on the analyzed emotion information. This enables dynamic responses that respond to the user's emotions.

[0769] 1. "Still image data" refers to a single, still image or photograph that is saved as an image file.

[0770] 2. "Motion image data" is a series of images consisting of multiple frames, and is visual information that changes over time.

[0771] 3. "Audio data" refers to data that stores audio in digital format, and includes information such as spoken language and sounds.

[0772] 4. "Customer service audio data" refers to data that records audio from customer service situations.

[0773] 5. "Text data" refers to data stored in text format, which is information expressed as words or sentences.

[0774] 6. A "customer service script" is a series of text messages generated to assist in interactions with users.

[0775] 7. "Conversation data" refers to a record of a conversation, either in audio or text format, that took place between a user and a customer service representative.

[0776] 8. A "scoring algorithm" is a calculation method for determining an evaluation score based on input data.

[0777] 9. The "evaluation score" is a numerical value calculated by a scoring algorithm and is used to evaluate the quality of conversation and emotional state.

[0778] 10. An "emotion engine" is software that analyzes video and audio data to determine the user's emotional state.

[0779] 11. "Analyzed emotional information" refers to data about the user's emotional state obtained as a result of analysis by the emotion engine.

[0780] 12. A "server" is a computer system used to store, process, and distribute data over a network.

[0781] 13. A "terminal" is a device that a user directly operates and has the function of communicating with a server to send and receive data.

[0782] 14. A "generative AI model" is an artificial intelligence model used to learn from large amounts of data and generate new text or responses.

[0783] This invention relates to a system that recognizes user emotions and provides dynamic customer service responses based on those emotions. This system consists of a server and a terminal, each performing the following functions.

[0784] The server first retrieves pre-prepared still image and audio data of employees from the database. This still image and audio data is sent to the video generation API and converted into video data. The converted video data is sent back to the server, which then sends it to the terminal. The terminal plays the received video data and displays it to the user.

[0785] Next, the server collects voice data of high-performing crew members and converts this voice data into text data using a speech recognition API. The converted text data is input into a generative AI model (e.g., GPT-4) to generate a customer service script. The generated script is sent from the server to the terminal, which uses this script to interact with the user.

[0786] The server further collects conversation data with the user in real time and converts this data into text data via a speech recognition API. The converted text data is input into a scoring algorithm, which calculates an evaluation score. This evaluation score is sent from the server to the terminal, which displays the score to the user in real time.

[0787] The server also sends video and audio data to the emotion engine. This emotion engine analyzes the user's facial expressions and tone of voice and generates emotion information. The generated emotion information is sent back to the server, which dynamically modifies the customer service script based on this information. The terminal receives this modified script and provides a more appropriate response.

[0788] For example, consider a case where a server holds still images and audio data of employee A. This data is sent to an API to generate video data. The generated video data is sent to a terminal via the server, and the terminal displays it, providing the user with a video that makes it appear as if employee A is actually speaking.

[0789] Furthermore, if Crew B's customer service voice data is stored on the server, the server sends this data to a speech recognition API and converts it into text data. Next, this text data is input into a generation AI model to generate a customer service script. This script is sent to the terminal, which then provides an appropriate response based on the user's questions and requests.

[0790] Furthermore, if a conversation with user C is ongoing, the terminal collects conversation data in real time and sends it to the server. The server converts the audio data into text data and calculates an evaluation score using a scoring algorithm. This score is sent to the terminal, which then displays it to the user in real time.

[0791] Finally, if user D appears to be experiencing stress, the emotion engine analyzes this and sends stress information to the server. Based on this information, the server generates a custom customer service script and sends it to the terminal. The terminal uses this script to provide a more relaxed interaction with user D.

[0792] Examples of prompt messages include the following:

[0793] "Could you tell me about your recent orders?"

[0794] "I'd like to learn more about premium membership."

[0795] "Please tell me how to return an item."

[0796] This system allows for real-time analysis of user emotions and dynamic adjustments to customer service responses based on those analyses, thereby increasing customer satisfaction.

[0797] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0798] Step 1:

[0799] Acquisition of still image data and audio data

[0800] The server retrieves pre-prepared still image and audio data of employees from the database. The input is the still image and audio data in the database, and the output is the retrieved still image and audio data. This data is then ready to be sent to the API.

[0801] Step 2:

[0802] Sending data to the API

[0803] The server sends the acquired still image and audio data to the video generation API. The input is the acquired still image and audio data, and the output is the request to send to the API. The server constructs the appropriate request to the API and sends the data.

[0804] Step 3:

[0805] Receiving data from API

[0806] The server receives video data sent back from the API. The input is the response data from the API, and the output is the received video data. The server receives this video data and prepares to send it to the terminal.

[0807] Step 4:

[0808] Sending video data to the terminal

[0809] The server sends the received video data to the terminal. The input is the video data from the API, and the output is the video data sent to the terminal. The server sends data to the terminal's address.

[0810] Step 5:

[0811] Receiving and displaying video data.

[0812] The terminal receives video data transmitted from the server. The input is video data from the server, and the output is video data for display. The terminal plays this video data and displays it to the user.

[0813] Step 6:

[0814] Collection and transcription of customer service voice data

[0815] The server collects customer service voice data from excellent crew members and converts this voice data into text data using a speech recognition API. The input is customer service voice data, and the output is text data. By converting to text data, the server prepares it for supply to a generative AI model.

[0816] Step 7:

[0817] Input to the Generative AI Model

[0818] The server inputs the converted character data into a generating AI model (e.g., GPT-4) to generate a customer service script. The input is character data, and the output is a customer service script. The generating AI model uses the character data to generate a script suitable for the service.

[0819] Step 8:

[0820] Sending the script to the terminal

[0821] The server sends the generated script to the terminal. The input is the generated customer service script, and the output is the script sent to the terminal. The server sends the script to the terminal's address.

[0822] Step 9:

[0823] Use of the submitted script

[0824] The terminal receives customer service scripts sent from the server. The input is the customer service script from the server, and the output is the response based on the script. The terminal uses the script to interact with the user.

[0825] Step 10:

[0826] Collection and transcription of conversation data

[0827] The server collects conversation data with the user and converts this data into text data via a speech recognition API. The input is conversation data, and the output is text data. The server then inputs the converted text data into a scoring algorithm.

[0828] Step 11:

[0829] Calculation of evaluation score

[0830] The server inputs text data into a scoring algorithm and calculates an evaluation score. The input is text data, and the output is the evaluation score. The algorithm works to evaluate the quality and emotional state of the conversation.

[0831] Step 12:

[0832] Sending evaluation scores to the device

[0833] The server sends the calculated evaluation score to the terminal. The input is the evaluation score, and the output is the score sent to the terminal. The server sends the score to the terminal's address.

[0834] Step 13:

[0835] Display of evaluation score

[0836] The terminal receives evaluation scores sent from the server and displays them to the user in real time. The input is the evaluation score from the server, and the output is the score information for display. The user can check this score.

[0837] Step 14:

[0838] Sending and analyzing emotional data

[0839] The server sends video and audio data to the emotion engine. The input is video and audio data, and the output is the data sent to the emotion engine. The emotion engine analyzes this data and generates emotion information.

[0840] Step 15:

[0841] Changes to scripts based on emotional information

[0842] The server receives the emotional information analyzed by the emotion engine and dynamically modifies the customer service script based on this information. The input is emotional information, and the output is the dynamically modified customer service script. The server generates a new script and sends it to the terminal.

[0843] Step 16:

[0844] Using the modified script

[0845] The terminal receives a newly transmitted customer service script and interacts with the user based on it. The input is the new script from the server, and the output is the interaction based on the new script. The terminal provides an interaction that reflects the script.

[0846] (Application Example 2)

[0847] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0848] In modern retail customer service, understanding and responding appropriately to customer emotions is crucial for improving customer satisfaction. However, traditional customer service systems have struggled to analyze customer emotional information in real time and dynamically adjust customer service scripts based on that analysis. This can lead to one-way communication with customers and potentially decrease customer satisfaction.

[0849] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for receiving still image data and audio data as input, means for generating moving image data based on the input still image data and audio data, means for transmitting the generated moving image data to a terminal, means for displaying the transmitted moving image data on the terminal, means for receiving image data and audio data to analyze the user's emotions, means for dynamically adjusting a customer service script based on the analyzed emotion information, means for transmitting the dynamically adjusted customer service script to a terminal, and means for interacting with the user using the transmitted customer service script on the terminal. This makes it possible to analyze the customer's emotions in real time and dynamically adjust customer service responses based on that analysis.

[0850] "Still image data" refers to individual frames that make up a video.

[0851] "Audio data" refers to data that records audio information digitally or in analog format.

[0852] "Motion data" refers to a video that is formed by sequentially combining multiple still images over time.

[0853] "Terminal" refers to an electronic device used by a user, and specifically includes personal computers, smartphones, smart glasses, etc.

[0854] "Emotional information" refers to data that indicates a user's psychological state and mood, analyzed from their facial expressions and tone of voice.

[0855] A "customer service script" refers to pre-designed text data of responses used in interactions with users.

[0856] A "generative AI model" refers to an artificial intelligence algorithm used to generate new data or text based on existing data.

[0857] A "prompt" is a text input to a generative AI model that instructs the model to produce a specific output.

[0858] This invention is a system that analyzes user emotions in real time and dynamically adjusts customer service based on that analysis, in order to improve customer service in physical stores. This system operates with a server and terminals working together and consists of the following main components.

[0859] First, the server accepts still image data and audio data as input. This data, collected by hardware such as cameras and microphones, is sent to the server. The server uses this still image data and audio data to generate video data. Specifically, video data is generated using APIs and machine learning algorithms. After generation, this video data is sent to the terminal.

[0860] The terminal receives video and image data transmitted from the server and displays it to the user. Devices such as smartphones and smart glasses are used as terminals. This terminal functions as an interface with the user and plays the video and image data.

[0861] Next, to analyze the user's emotions, the server sends image and audio data to an emotion analysis engine. This emotion analysis engine identifies the user's psychological state from their facial expressions and tone of voice and generates emotion information. This emotion information is then sent back to the server.

[0862] The server dynamically adjusts the customer service script based on this emotional information. A generative AI model is used to create prompts containing emotional information and generate an appropriate customer service script. OpenAI's GPT model is used as the generative AI model in this process. The generated customer service script is sent to the terminal, which then interacts with the user based on it.

[0863] As a concrete example, consider a scenario where a customer near a fitting room asks for feedback on a product they are trying on. In this case, the staff member performs real-time sentiment analysis via smart glasses and provides appropriate feedback based on the generated customer service script. For example, if the customer asks, "Does this jacket suit me?", the generating AI model would receive a prompt like this:

[0864] "When a customer asks, 'Does this jacket suit me?', generate the optimal customer service script. The customer's voice tone is relaxed, and their facial expression is smiling."

[0865] This system allows for real-time analysis of customer emotions and dynamic adjustments to customer service based on that analysis, which is expected to improve customer satisfaction.

[0866] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0867] Step 1:

[0868] The server accepts still image and audio data as input from the camera and microphone. This data is collected in real time to mimic human customer service staff. Still images captured by the camera and audio recorded by the microphone are sent to the server.

[0869] Step 2:

[0870] The server generates video data based on the received still image and audio data. Here, a deep learning model is used to generate the video data. This deep learning model, for example, is one that has been pre-trained on employee movements and speech patterns. The generated video data is temporarily stored on the server.

[0871] Step 3:

[0872] The server sends the generated video data to the terminal. The terminal consists of devices such as smart glasses or smartphones. The video data received by the terminal is played back for the customer to see.

[0873] Step 4:

[0874] The server sends video and audio data to the emotion analysis engine to analyze the user's emotions. The emotion analysis engine generates emotional information from the user's facial expressions and tone of voice. For example, it can determine whether the user is relaxed based on their smile or tone of voice. The analysis results are then returned to the server.

[0875] Step 5:

[0876] The server creates a prompt sentence based on the emotion information returned from the emotion analysis engine and inputs it into the generative AI model. The generative AI model uses OpenAI's GPT model to generate a customer service script based on this prompt sentence. Because the prompt sentence includes the user's voice content and emotion information, a very natural conversation is possible.

[0877] Step 6:

[0878] The server sends the generated customer service script to the terminal. The terminal uses this script to interact with the customer. For example, if a customer in a fitting room is asking for product feedback, the terminal provides an appropriate response.

[0879] Step 7:

[0880] The device collects new data based on user interaction and sends it to the server. This allows the system to continuously improve, enabling more natural and effective responses.

[0881] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0882] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0883] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0884] [Third Embodiment]

[0885] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0886] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0887] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0888] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0889] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0890] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0891] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0892] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0893] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0894] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0895] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0896] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0897] This invention relates to a system that generates moving image data based on still image data and audio data to provide consistent customer service. It also includes collecting and analyzing conversation data during customer service in real time to evaluate customer satisfaction. Specific embodiments of this system are described below.

[0898] 1. Generating the avatar chat screen

[0899] Server Processing

[0900] The server retrieves pre-prepared still images and audio data of employees. Next, it sends this data via an API (Application Programming Interface) to request the generation of video data. The video data returned from the API is received by the server, which then sends the video data to the terminal.

[0901] Terminal processing

[0902] The terminal receives video data transmitted from the server. After receiving the data, the terminal plays it back and displays it to the user.

[0903] Specific example

[0904] For example, if a server has still images and audio data of employee A, the server sends these to an API to generate video data. Upon receiving the video data as a response from the API, the server sends it to a terminal. The terminal plays this data, providing the user with a video that makes it appear as if employee A is speaking.

[0905] 2. Generating and using customer service scripts

[0906] Server Processing

[0907] The server collects voice data of excellent crew members' customer service interactions. This collected voice data is converted into text data using a speech recognition API. Next, the server feeds this text data into a GPT model (Greater Global Pattern Testing) to generate appropriate customer service scripts. The generated scripts are then sent from the server to the terminals.

[0908] Terminal processing

[0909] The terminal receives a customer service script sent from the server. Based on this script, the terminal interacts with the user.

[0910] Specific example

[0911] If Crew B possesses excellent customer service skills, the server records Crew B's customer service audio. This audio data is transcribed and fed into a GPT model to generate a customer service script. The terminal uses this script to provide appropriate responses to the user's questions.

[0912] 3. Talk Cancellation Scoring

[0913] Server Processing

[0914] The server collects conversation data with users in real time. The collected conversation data is converted into text data via a speech recognition API and then input into a scoring algorithm. This algorithm analyzes the conversation data and calculates a score to evaluate customer satisfaction and willingness to continue using the service. Finally, this evaluation score is sent from the server to the terminal.

[0915] Terminal processing

[0916] The device receives evaluation scores sent from the server and displays them to the user in real time.

[0917] Specific example

[0918] While a conversation with user C is in progress, the terminal collects the conversation content in real time and sends it to the server. The server transcribes this audio data and calculates an evaluation score using a scoring algorithm. By displaying this score on the terminal, it becomes possible to continuously evaluate and improve the quality of the service.

[0919] In this way, the present invention can consistently improve customer satisfaction through an AI system with advanced customer service skills.

[0920] The following describes the processing flow.

[0921] 1. Generating the avatar chat screen

[0922] Step 1:

[0923] The server retrieves pre-prepared still images and audio data of employees.

[0924] Step 2:

[0925] The server creates and sends an HTTP POST request to send the acquired still image and audio data to the API.

[0926] Step 3:

[0927] The server waits for a response from the API and receives the generated video data. This video data is generated from still images and audio data.

[0928] Step 4:

[0929] The server sends the received video data to the terminal. It uses an appropriate communication protocol (e.g., HTTP, WebSocket) for this purpose.

[0930] Step 5:

[0931] The terminal receives video data transmitted from the server. After receiving the data, the terminal plays it back and displays it to the user.

[0932] Specific example

[0933] Still images and audio data of employee A are stored on the server. The server sends this data to an API, receives video data generated by the API, and sends it to the terminal. The terminal receives the video data and displays to the user an image that makes it appear as if employee A is speaking.

[0934] 2. Generating and using customer service scripts

[0935] Step 1:

[0936] The server collects audio data of excellent customer service from top-performing crew members. This audio data is pre-recorded and stored in a database.

[0937] Step 2:

[0938] The server sends the collected customer service voice data to a speech recognition API, where it is converted into text data. The voice data is sent to the API via an HTTP POST request.

[0939] Step 3:

[0940] The server receives the text data returned from the speech recognition API.

[0941] Step 4:

[0942] The server inputs the received text data into a GPT model to generate a customer service script. The GPT model is pre-trained and can automatically generate appropriate responses.

[0943] Step 5:

[0944] The server sends the generated customer service script to the terminal, using the appropriate communication protocol.

[0945] Step 6:

[0946] The terminal receives customer service scripts sent from the server and provides appropriate responses based on the user's questions and requests.

[0947] Specific example

[0948] Crew B's excellent customer service voice is stored on the server. The server sends this voice data to a speech recognition API to transcribe it and input it into a GPT model. The generated script is sent to the terminal, which uses it to respond to the user.

[0949] 3. Talk Cancellation Scoring

[0950] Step 1:

[0951] The device collects conversation data with the user in real time. This collection is done through a voice input device.

[0952] Step 2:

[0953] The terminal sends the collected audio data to the server. An appropriate communication protocol (e.g., HTTP, WebSocket) is used.

[0954] Step 3:

[0955] The server sends the received audio data to a speech recognition API, where it is converted into text data.

[0956] Step 4:

[0957] The server receives the text data returned from the speech recognition API.

[0958] Step 5:

[0959] The server inputs the received text data into a scoring algorithm to calculate a score for evaluating customer satisfaction and willingness to continue using the service.

[0960] Step 6:

[0961] The server sends the calculated evaluation score to the terminal.

[0962] Step 7:

[0963] The device receives evaluation scores sent from the server and displays them to the user in real time.

[0964] Specific example

[0965] As a conversation with user C progresses, the terminal collects conversation data in real time and sends it to the server. The server sends the data to a speech recognition API to transcribe it and calculates an evaluation score using a scoring algorithm. This score is sent to the terminal, which then displays it to the user in real time.

[0966] In this way, the system of the present invention can provide advanced customer service skills and real-time customer evaluation, thereby consistently improving customer satisfaction.

[0967] (Example 1)

[0968] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0969] In modern society, consistent customer service is required to enhance customer satisfaction. Especially in online customer service, direct interaction with customers is difficult, and service often depends on the individual skills of employees. This can lead to inconsistencies in service quality, potentially resulting in decreased customer satisfaction. Furthermore, the lack of adequate means to evaluate customer satisfaction in real time and improve service quality makes prompt responses difficult. A system is needed to address these challenges and improve customer satisfaction.

[0970] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0971] In this invention, the server includes means for receiving still image data and audio data as input, means for generating video data based on the input still image data and audio data, means for transmitting the generated video data to a terminal, means for converting customer service audio data into text data and generating a customer service script using a generation AI model, means for transmitting the generated customer service script to a terminal, means for converting conversation data into text data and calculating an evaluation score using a scoring algorithm, and means for transmitting the calculated evaluation score to a terminal. As a result, the server generates video data and transmits it to a terminal to achieve consistent customer service, and it becomes possible to generate scripts that utilize the customer service skills of crew members with excellent customer service skills and use them on the terminal. Furthermore, it becomes possible to perform real-time customer satisfaction evaluation and continuously improve the quality of customer service.

[0972] "Still image data" refers to digital data containing image information captured at a specific point in time.

[0973] "Audio data" refers to data that records human voices in digital format.

[0974] "Motion image data" refers to digital data in video format that includes a sequence of images and synchronized audio.

[0975] A "terminal" is an electronic device that a user can directly operate.

[0976] "Customer service audio data" refers to data that digitally records the voices spoken by employees during customer service interactions.

[0977] "Text data" refers to data that visualizes audio data as text using speech recognition technology.

[0978] A "generative AI model" is an artificial intelligence model trained to generate appropriate responses or scripts from specific input data.

[0979] A "customer service script" is text data containing a series of sentences and response instructions used during customer service.

[0980] "Conversation data" refers to digital data that records voice interactions between a user and a system.

[0981] A "scoring algorithm" is a calculation method used to analyze input data and calculate an evaluation score based on specific criteria.

[0982] An "evaluation score" is a numerical evaluation result calculated by a scoring algorithm.

[0983] This invention relates to a system that generates moving image data based on still image data and audio data to provide consistent customer service. It also includes a function to collect and analyze conversation data during customer service in real time and evaluate customer satisfaction. Specific embodiments of this invention will now be described.

[0984] First, the server retrieves pre-prepared still image data (JPEG format) and audio data (WAV format) of the employee. This data is sent via an HTTP POST request using an API such as Microsoft Azure's Face API. The response returned from the API contains the generated video data (MP4 format), which the server receives and temporarily stores. Subsequently, the server sends the stored video data to the terminal as an HTTP response.

[0985] Next, the terminal receives video data sent from the server. After receiving the data, the terminal initializes a media player to play it and displays the video to the user. This series of processes allows the terminal to play a video that makes it appear as if an employee is speaking.

[0986] Next, to collect excellent customer service audio data (in WAV format) from crew members, the server uses the Google Cloud Speech-to-Text API to convert the audio data into text data. Then, the text data is fed into a generative AI model such as the GPT-3 model to generate customer service scripts. The generated scripts are temporarily stored on the server and sent to the terminal. The terminal receives the customer service scripts sent from the server and uses them to interact with the user.

[0987] Next, the server collects conversation data (in WAV format) from the user in real time. The conversation data is converted into text data using the Google Cloud Speech-to-Text API and then fed into a scoring algorithm. This algorithm analyzes the conversation data and calculates a score that evaluates customer satisfaction and willingness to continue using the service. This evaluation score is sent from the server to the terminal and displayed to the user in real time, enabling continuous evaluation and improvement of service quality.

[0988] Specific example

[0989] 1. The server sends still images and audio data of employee A to the API to generate video data. The terminal receives this data and displays a video to the user that makes it appear as if employee A is speaking.

[0990] 2. The customer service voice of Crew B is transcribed and input into the GPT-3 model to generate a customer service script. The terminal uses this script to respond to the user's questions.

[0991] 3. The conversation with User C is collected, and a customer satisfaction evaluation score is calculated using a scoring algorithm. The terminal displays this score and evaluates the service quality.

[0992] Example of a prompt

[0993] 1. "Please generate video data based on a still image of employee A and the following audio data."

[0994] 2. "Please generate a customer service script based on the text data of this customer service audio."

[0995] 3. "Based on the conversation data, calculate a score to evaluate customer satisfaction."

[0996] In this way, the present invention provides a system that generates moving image data using still image data and audio data, evaluates customer satisfaction in real time, and realizes consistently high-quality customer service.

[0997] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0998] Generating an avatar chat screen

[0999] Server Processing

[1000] Step 1:

[1001] The server retrieves pre-prepared still image data (e.g., JPEG format) and audio data (e.g., WAV format) of employees. The server receives the still image and audio data as input. This data serves as the basic information for generating video data in subsequent processing.

[1002] Step 2:

[1003] The server sends the acquired still image and audio data to an API (e.g., Microsoft Azure's Face API). It makes an HTTP POST request to the API endpoint, sending the still image and audio data. Based on the input data, the API executes a process and generates a video image data by combining the still image and audio data.

[1004] Step 3:

[1005] The server receives video data (e.g., in MP4 format) returned from the API. It temporarily stores the video data obtained as an HTTP response and prepares for the subsequent transmission process.

[1006] Step 4:

[1007] The server sends the generated video data to the terminal as an HTTP response. This allows the video data to be delivered to the terminal via the internet.

[1008] Terminal processing

[1009] Step 1:

[1010] The terminal receives video data sent from the server as an HTTP response. It stores the received data in temporary storage. It then parses and stores the URL or binary data used as input data.

[1011] Step 2:

[1012] The device initializes its media player for playing saved video data. It internally loads the video data and prepares it for display to the user.

[1013] Step 3:

[1014] The device plays video and displays it to the user. It uses a media player to display the video along a timeline and outputs it in a way that the user can perceive.

[1015] Generating and using customer service scripts

[1016] Server Processing

[1017] Step 1:

[1018] The server collects voice data (e.g., in WAV format) of high-performing crew members providing excellent customer service. The collected voice data is temporarily stored to prepare for the subsequent speech recognition process.

[1019] Step 2:

[1020] The server sends the collected audio data to the Google Cloud Speech-to-Text API, where it is converted into text data. The audio data is then sent as input to the speech recognition API, and an HTTP response containing the text data is received. This converts the audio data into text data from start to finish.

[1021] Step 3:

[1022] The server inputs the obtained character data into a GPT-3 model to generate customer service scripts. The character data is input into a generation AI model to generate customer service scripts as corresponding text data.

[1023] Step 4:

[1024] The server temporarily stores the generated customer service script and then sends it to the terminal as an HTTP response. This delivers the generated script to the user's terminal.

[1025] Terminal processing

[1026] Step 1:

[1027] The terminal receives the customer service script sent from the server. It retrieves the script information as an HTTP response and saves it to temporary storage. It then reads the contents of the script as input data and saves it.

[1028] Step 2:

[1029] The terminal handles customer service interactions based on saved customer service scripts. It dynamically references the script in response to user input and generates appropriate responses. Specifically, it responds to user inquiries and requests according to the instructions in the script.

[1030] Talk cancellation scoring

[1031] Server Processing

[1032] Step 1:

[1033] The server collects conversation data with users (e.g., in WAV format) in real time. Each time a conversation occurs, it is captured as audio data and stored on the server. This data serves as foundational information for analyzing customer satisfaction in subsequent processing.

[1034] Step 2:

[1035] The server sends the collected audio data to the Google Cloud Speech-to-Text API, where it is converted into text data. The audio data is sent to the speech recognition API, and an HTTP response is received as text data. This converts the conversation data into text data.

[1036] Step 3:

[1037] The server inputs the obtained character data into a scoring algorithm to calculate an evaluation score. The character data is input to the scoring algorithm, and a score is calculated based on the evaluation criteria.

[1038] Step 4:

[1039] The server temporarily stores the calculated evaluation score and then sends it to the terminal as an HTTP response. This delivers the generated evaluation score to the user's terminal.

[1040] Terminal processing

[1041] Step 1:

[1042] The terminal receives the evaluation score sent from the server. It retrieves the score information as an HTTP response and saves it to temporary storage. It then reads the score content as input data and saves it.

[1043] Step 2:

[1044] The device displays saved evaluation scores to the user in real time. The interface for displaying evaluation scores is initialized, and customer satisfaction scores are displayed dynamically. This enables real-time evaluation of service quality.

[1045] Through these steps, the system generates video data using still image and audio data, enabling consistent customer service and real-time customer satisfaction evaluation.

[1046] (Application Example 1)

[1047] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1048] In traditional brick-and-mortar stores, inadequate customer service was sometimes difficult due to staff absences or insufficient manpower. Furthermore, maintaining consistent service quality was challenging, and there were limited means to evaluate and improve customer satisfaction in real time. This invention aims to solve these problems by using an AI system with advanced customer service skills to consistently improve customer satisfaction.

[1049] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1050] In this invention, the server includes means for receiving still image data and audio data as input in order to convert still image data into video data; means for generating video data based on the input still image data and audio data; means for transmitting the generated video data to a terminal; means for displaying the transmitted video data on the terminal; means for converting customer service audio data into speech recognition data; means for generating a customer service script based on the speech recognition data; means for evaluating conversation data in real time and calculating a customer satisfaction evaluation score; and means for transmitting the evaluation score to a terminal and displaying it on the terminal. This enables virtual customer service by an avatar host using a smartphone application, allowing for consistent, high-quality customer service even when staff are absent in a physical store. Furthermore, it makes it possible to evaluate customer satisfaction in real time and continuously improve the quality of service.

[1051] "Still image data" refers to data of a static image that does not move.

[1052] "Motion data" refers to video data generated using a series of still images.

[1053] "Audio data" refers to data that records a person's voice in digital format.

[1054] "Speech recognition data" refers to data obtained by analyzing speech data and representing its content as text data.

[1055] A "customer service script" is text data that describes the content of conversations and response methods during customer service, and is used to instruct and guide customer service operations.

[1056] "Conversation data" refers to language-based information data exchanged between customers and systems or staff.

[1057] An "evaluation score" is a numerical value calculated based on conversation data and customer responses, and is an indicator used to evaluate customer satisfaction and service quality.

[1058] A "terminal" refers to an electronic device used to receive and display data transmitted from a server.

[1059] A "scoring algorithm" is a set of computational methods and rules used to analyze conversational data and other input data and calculate an evaluation score.

[1060] "Motion image generation means" refers to a method or apparatus for creating motion image data based on still image data and audio data.

[1061] This invention provides a specific method for constructing a virtual customer service system using a smartphone application in a physical store. This system generates moving image data based on still image data and audio data to provide consistent customer service to customers. It also collects and analyzes conversation data during customer service in real time to evaluate customer satisfaction.

[1062] Hardware and software configuration

[1063] Servers and smartphones will be used as the primary hardware. Specifically, the following:

[1064] server:

[1065] A cloud server for storing and processing still image and audio data.

[1066] Speech-to-Text is used to convert speech data into text data.

[1067] A video image generation API (e.g., DeepMotion) is used to generate video image data from still images and audio.

[1068] A GPT model (e.g., OpenAI's GPT-4) is used to generate customer service scripts from text data.

[1069] A scoring algorithm is used to evaluate customer satisfaction in real time.

[1070] Smartphone device:

[1071] A device for receiving, displaying, and playing back data transmitted from a server.

[1072] Collect user interactions and send them to the server.

[1073] Data processing and data calculation

[1074] Video generation:

[1075] The server accepts still image data and audio data as input. This data is passed to a video generation API to generate video data. The generated video data is then sent from the server to the terminal. The terminal displays this data, providing the user with an avatar image.

[1076] Customer service script generation:

[1077] The server converts customer service voice data into text data using a speech recognition API. The text data is input into a GPT model to generate an appropriate customer service script. This script is sent from the server to the terminal, which then uses the script to interact with the customer.

[1078] Customer satisfaction rating:

[1079] The server collects conversation data in real time and converts it into text data using a speech recognition API. The converted text data is analyzed by a scoring algorithm to calculate an evaluation score. This evaluation score is sent to the terminal and displayed to the user in real time.

[1080] Specific example

[1081] For example, in a clothing store, if a customer asks for trousers to match a jacket when no staff are present, this system would function effectively. If the customer asks, "Please recommend trousers that would go with this jacket," the system generates an avatar, which is displayed as video data. Furthermore, based on a customer service script, the avatar can provide an appropriate response and address the customer's questions. It is possible to collect data during the conversation in real time and display customer satisfaction as an evaluation score.

[1082] Example of a prompt

[1083] Please generate a customer service script for recommending trousers to match a jacket to a customer visiting a clothing store. Additionally, design an algorithm to evaluate customer satisfaction in real time during the interaction.

[1084] Thus, the system of the present invention can consistently improve customer satisfaction in physical stores by using an AI system with advanced customer service skills.

[1085] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1086] Step 1:

[1087] The server receives still image and audio data. When a user uploads still images and audio of employees during the store's preparation phase, the server retrieves them. The input data consists of still image and audio data, while the output data is a dataset to be passed to the video generation API.

[1088] Step 2:

[1089] The server uses the received still image and audio data to call a video generation API and generate video data. By providing still image and audio data as input to this video generation API (e.g., DeepMotion), animated video of employees is generated. The input data consists of image and audio files for the API call, and the output data is the generated video data.

[1090] Step 3:

[1091] The server sends the generated video data to the terminal. The server receives the video data and sends it to the terminal using a communication protocol. The input data is the video data, and the output data is the confirmation message sent to the terminal.

[1092] Step 4:

[1093] The terminal receives and displays video data transmitted from the server. The user operates the terminal to play the video data, which then displays an avatar. The input data is video data, and the output data is the display of the avatar image.

[1094] Step 5:

[1095] The server receives audio data during customer service interactions and converts it into text data using a speech recognition API. It collects the audio of the user-avatar conversation in real time and sends it to a speech recognition API such as Google Cloud Speech-to-Text to obtain text data. The input data is audio data, and the output data is the converted text data.

[1096] Step 6:

[1097] The server inputs text data into a GPT model and generates a customer service script. This text data is then input into an OpenAI GPT-4 model, and the process of generating an appropriate customer service script is executed. The input data is text data, and the output data is the customer service script.

[1098] Step 7:

[1099] The server sends the generated customer service script to the terminal. Sending the generated customer service script to the terminal enables real-time interaction with the user on the terminal. The input data is the customer service script, and the output data is the confirmation message sent to the terminal.

[1100] Step 8:

[1101] The terminal interacts with the user based on the received customer service script. The script is displayed on the terminal, and the avatar responds accordingly. The input data is the customer service script, and the output data is the result of the conversation with the user.

[1102] Step 9:

[1103] The server collects conversation data with the user in real time and converts it into text data using a speech recognition API. The input data is the audio data of the conversation, and the output data is the converted text data.

[1104] Step 10:

[1105] The server inputs the converted character data into a scoring algorithm to calculate an evaluation score. The input data is the converted character data, and the output data is the evaluation score.

[1106] Step 11:

[1107] The server sends the calculated evaluation score to the terminal, which then displays the evaluation score. The input data is the evaluation score, and the output data is the customer satisfaction evaluation result displayed on the terminal.

[1108] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1109] This invention relates to a system that recognizes user emotions and provides dynamic customer service responses based on those emotions. This system generates dynamic image data from still image data and audio data, and further generates customer service scripts by converting customer service audio data into text data. In addition, by combining this with an emotion engine that evaluates conversation data and recognizes user emotions, it achieves more sophisticated customer service.

[1110] 1. Generating the avatar chat screen

[1111] Server Processing

[1112] The server retrieves pre-prepared still images and audio data of employees. This data is sent to the API, which requests the generation of video data. The video data returned from the API is received by the server and then sent to the terminal.

[1113] Terminal processing

[1114] The terminal receives video data transmitted from the server, plays this data, and displays it to the user.

[1115] Specific example

[1116] For example, if a server has still images and audio data of employee A, it sends this data to an API to generate video data. Upon receiving the video data as a response from the API, the server sends it to a terminal. The terminal plays this video data, providing the user with a video that makes it appear as if employee A is speaking.

[1117] 2. Generating and using customer service scripts

[1118] Server Processing

[1119] The server collects voice data of high-performing crew members' customer service interactions. This voice data is converted into text data using a speech recognition API. The text data is input into a GPT model to generate customer service scripts. The generated scripts are sent from the server to the terminals.

[1120] Terminal processing

[1121] The terminal receives customer service scripts sent from the server and uses these scripts to interact with the user.

[1122] Specific example

[1123] Crew B's customer service voice data is stored on the server. The server sends this voice data to a speech recognition API to transcribe it and input it into a GPT model. The generated script is sent to the terminal, which then provides an appropriate response based on the user's questions and requests.

[1124] 3. Talk Cancellation Scoring

[1125] Server Processing

[1126] The server collects conversation data with the user in real time. This conversation data is converted into text data via a speech recognition API and input into a scoring algorithm. The algorithm calculates an evaluation score, which is then sent from the server to the terminal.

[1127] Terminal processing

[1128] The device receives evaluation scores sent from the server and displays them to the user in real time.

[1129] Specific example

[1130] If a conversation with user C is in progress, the terminal collects conversation data in real time and sends it to the server. The server transcribes the audio data and calculates an evaluation score using a scoring algorithm. This score is sent to the terminal, which displays it to the user in real time.

[1131] 4. Adding an emotion engine

[1132] Server Processing

[1133] The server sends video and audio data to the emotion engine. This engine analyzes the user's facial expressions and tone of voice to generate emotion information. This emotion information is sent back to the server, which is then instructed to dynamically modify the customer service script based on the user's emotions.

[1134] Terminal processing

[1135] The device receives emotional information and adjusts customer service scripts based on it. It also uses emotional information to provide more appropriate responses.

[1136] Specific example

[1137] If user D is experiencing stress, the emotion engine analyzes this and sends stress information to the server. Based on this information, the server generates a dedicated customer service script and sends it to the terminal. The terminal uses this script to provide a more relaxed interaction with user D.

[1138] The system of this invention can improve customer satisfaction by analyzing user emotions in real time and dynamically changing customer service responses based on those emotions.

[1139] The following describes the processing flow.

[1140] Processing flow of a system combining an emotion engine

[1141] 1. Generating the avatar chat screen

[1142] Step 1:

[1143] The server retrieves pre-prepared still images and audio data of employees.

[1144] Step 2:

[1145] The server creates and sends an HTTP POST request to the API to send the acquired still image and audio data.

[1146] Step 3:

[1147] The server waits for a response from the API and receives the generated video data. This video data is generated from still images and audio data.

[1148] Step 4:

[1149] The server sends the received video data to the terminal. It uses an appropriate communication protocol (e.g., HTTP, WebSocket) for this purpose.

[1150] Step 5:

[1151] The terminal receives video data transmitted from the server. After receiving the data, the terminal plays it back and displays it to the user.

[1152] 2. Generating and using customer service scripts

[1153] Step 1:

[1154] The server collects audio data of excellent customer service from top-performing crew members. This audio data is pre-recorded and stored in a database.

[1155] Step 2:

[1156] The server sends the collected customer service voice data to a speech recognition API, where it is converted into text data. The voice data is sent to the API via an HTTP POST request.

[1157] Step 3:

[1158] The server receives the text data returned from the speech recognition API.

[1159] Step 4:

[1160] The server inputs the received text data into a GPT model to generate a customer service script. The GPT model is pre-trained and can automatically generate appropriate responses.

[1161] Step 5:

[1162] The server sends the generated customer service script to the terminal, using the appropriate communication protocol.

[1163] Step 6:

[1164] The terminal receives customer service scripts sent from the server and provides appropriate responses based on the user's questions and requests.

[1165] 3. Talk Cancellation Scoring

[1166] Step 1:

[1167] The device collects conversation data with the user in real time. This collection is done through a voice input device.

[1168] Step 2:

[1169] The terminal sends the collected audio data to the server. An appropriate communication protocol (e.g., HTTP, WebSocket) is used.

[1170] Step 3:

[1171] The server sends the received audio data to a speech recognition API, where it is converted into text data.

[1172] Step 4:

[1173] The server receives the text data returned from the speech recognition API.

[1174] Step 5:

[1175] The server inputs the received text data into a scoring algorithm to calculate a score for evaluating customer satisfaction and willingness to continue using the service.

[1176] Step 6:

[1177] The server sends the calculated evaluation score to the terminal.

[1178] Step 7:

[1179] The device receives evaluation scores sent from the server and displays them to the user in real time.

[1180] 4. Adding an emotion engine

[1181] Step 1:

[1182] The server sends video and audio data to the emotion engine. The emotion engine analyzes the user's emotions based on their facial expressions and tone of voice.

[1183] Step 2:

[1184] The emotion engine analyzes the user's emotions and sends that emotional information back to the server.

[1185] Step 3:

[1186] The server dynamically modifies customer service scripts based on emotional information received from the emotion engine. For example, if a user is dissatisfied, it generates a more courteous response.

[1187] Step 4:

[1188] The server sends the modified customer service script to the terminal, using the appropriate communication protocol.

[1189] Step 5:

[1190] The terminal receives emotional information and customer service scripts sent from the server and responds appropriately to the user.

[1191] Specific example

[1192] If user D is experiencing stress, the emotion engine analyzes this and sends this information to the server. Based on the emotion information, the server generates a dedicated customer service script and sends it to the terminal. The terminal uses this script to provide a relaxed interaction with user D.

[1193] In this way, the system of the present invention analyzes the user's emotions in real time through each processing step and dynamically changes customer service responses based on that analysis, thereby providing a mechanism that can improve customer satisfaction.

[1194] (Example 2)

[1195] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1196] Traditional customer service systems struggled to analyze user emotions in real time based on facial expressions and tone of voice, and to dynamically adjust customer service accordingly. As a result, they were unable to enhance user satisfaction and could only provide standardized responses, failing to offer optimal service to individual users.

[1197] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1198] In this invention, the server includes means for receiving still image data and audio data as input in order to convert still image data into video data; means for generating video data based on the input still image data and audio data; means for transmitting the generated video data to a terminal; means for receiving customer service audio data as input in order to transcribe customer service audio data; means for converting the input customer service audio data into text data; means for generating a customer service script based on the text data; means for transmitting the generated customer service script to a terminal; means for receiving conversation data as input in order to evaluate conversation data; means for converting the input conversation data into text data; means for inputting the text data into a scoring algorithm to calculate an evaluation score; means for receiving video data and audio data as input and analyzing the user's emotions; and means for dynamically changing the customer service script based on the analyzed emotion information. This enables dynamic responses that respond to the user's emotions.

[1199] 1. "Still image data" refers to a single, still image or photograph that is saved as an image file.

[1200] 2. "Motion image data" is a series of images consisting of multiple frames, and is visual information that changes over time.

[1201] 3. "Audio data" refers to data that stores audio in digital format, and includes information such as spoken language and sounds.

[1202] 4. "Customer service audio data" refers to data that records audio from customer service situations.

[1203] 5. "Text data" refers to data stored in text format, which is information expressed as words or sentences.

[1204] 6. A "customer service script" is a series of text messages generated to assist in interactions with users.

[1205] 7. "Conversation data" refers to a record of a conversation, either in audio or text format, between a user and a customer service representative.

[1206] 8. A "scoring algorithm" is a calculation method for determining an evaluation score based on input data.

[1207] 9. The "evaluation score" is a numerical value calculated by a scoring algorithm and is used to evaluate the quality of conversation and emotional state.

[1208] 10. An "emotion engine" is software that analyzes video and audio data to determine the user's emotional state.

[1209] 11. "Analyzed emotional information" refers to data about the user's emotional state obtained as a result of analysis by the emotion engine.

[1210] 12. A "server" is a computer system used to store, process, and distribute data over a network.

[1211] 13. A "terminal" is a device that a user directly operates and has the function of communicating with a server to send and receive data.

[1212] 14. A "generative AI model" is an artificial intelligence model used to learn from large amounts of data and generate new text or responses.

[1213] This invention relates to a system that recognizes user emotions and provides dynamic customer service responses based on those emotions. This system consists of a server and a terminal, each performing the following functions.

[1214] The server first retrieves pre-prepared still image and audio data of employees from the database. This still image and audio data is sent to the video generation API and converted into video data. The converted video data is sent back to the server, which then sends it to the terminal. The terminal plays the received video data and displays it to the user.

[1215] Next, the server collects voice data of high-performing crew members and converts this voice data into text data using a speech recognition API. The converted text data is input into a generative AI model (e.g., GPT-4) to generate a customer service script. The generated script is sent from the server to the terminal, which uses this script to interact with the user.

[1216] The server further collects conversation data with the user in real time and converts this data into text data via a speech recognition API. The converted text data is input into a scoring algorithm, which calculates an evaluation score. This evaluation score is sent from the server to the terminal, which displays the score to the user in real time.

[1217] The server also sends video and audio data to the emotion engine. This emotion engine analyzes the user's facial expressions and tone of voice and generates emotion information. The generated emotion information is sent back to the server, which dynamically modifies the customer service script based on this information. The terminal receives this modified script and provides a more appropriate response.

[1218] For example, consider a case where a server holds still images and audio data of employee A. This data is sent to an API to generate video data. The generated video data is sent to a terminal via the server, and the terminal displays it, providing the user with a video that makes it appear as if employee A is actually speaking.

[1219] Furthermore, if Crew B's customer service voice data is stored on the server, the server sends this data to a speech recognition API and converts it into text data. Next, this text data is input into a generation AI model to generate a customer service script. This script is sent to the terminal, which then provides an appropriate response based on the user's questions and requests.

[1220] Furthermore, if a conversation with user C is ongoing, the terminal collects conversation data in real time and sends it to the server. The server converts the audio data into text data and calculates an evaluation score using a scoring algorithm. This score is sent to the terminal, which then displays it to the user in real time.

[1221] Finally, if user D appears to be experiencing stress, the emotion engine analyzes this and sends stress information to the server. Based on this information, the server generates a custom customer service script and sends it to the terminal. The terminal uses this script to provide a more relaxed interaction with user D.

[1222] Examples of prompt messages include the following:

[1223] "Could you tell me about your recent orders?"

[1224] "I'd like to learn more about premium membership."

[1225] "Please tell me how to return an item."

[1226] This system allows for real-time analysis of user emotions and dynamic adjustments to customer service responses based on those analyses, thereby increasing customer satisfaction.

[1227] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1228] Step 1:

[1229] Acquisition of still image data and audio data

[1230] The server retrieves pre-prepared still image and audio data of employees from the database. The input is the still image and audio data in the database, and the output is the retrieved still image and audio data. This data is then ready to be sent to the API.

[1231] Step 2:

[1232] Sending data to the API

[1233] The server sends the acquired still image and audio data to the video generation API. The input is the acquired still image and audio data, and the output is the request to send to the API. The server constructs the appropriate request to the API and sends the data.

[1234] Step 3:

[1235] Receiving data from API

[1236] The server receives video data sent back from the API. The input is the response data from the API, and the output is the received video data. The server receives this video data and prepares to send it to the terminal.

[1237] Step 4:

[1238] Sending video data to the terminal

[1239] The server sends the received video data to the terminal. The input is the video data from the API, and the output is the video data sent to the terminal. The server sends data to the terminal's address.

[1240] Step 5:

[1241] Receiving and displaying video data.

[1242] The terminal receives video data transmitted from the server. The input is video data from the server, and the output is video data for display. The terminal plays this video data and displays it to the user.

[1243] Step 6:

[1244] Collection and transcription of customer service voice data

[1245] The server collects customer service voice data from excellent crew members and converts this voice data into text data using a speech recognition API. The input is customer service voice data, and the output is text data. By converting to text data, the server prepares it for feeding into a generative AI model.

[1246] Step 7:

[1247] Input to the Generative AI Model

[1248] The server inputs the converted character data into a generating AI model (e.g., GPT-4) to generate a customer service script. The input is character data, and the output is a customer service script. The generating AI model uses the character data to create a script suitable for the service.

[1249] Step 8:

[1250] Sending the script to the terminal

[1251] The server sends the generated script to the terminal. The input is the generated customer service script, and the output is the script sent to the terminal. The server sends the script to the terminal's address.

[1252] Step 9:

[1253] Use of the submitted script

[1254] The terminal receives customer service scripts sent from the server. The input is the customer service script from the server, and the output is the response based on the script. The terminal uses the script to interact with the user.

[1255] Step 10:

[1256] Collection and transcription of conversation data

[1257] The server collects conversation data with the user and converts this data into text data via a speech recognition API. The input is conversation data, and the output is text data. The server then inputs the converted text data into a scoring algorithm.

[1258] Step 11:

[1259] Calculation of evaluation score

[1260] The server inputs text data into a scoring algorithm and calculates an evaluation score. The input is text data, and the output is the evaluation score. The algorithm works to evaluate the quality and emotional state of the conversation.

[1261] Step 12:

[1262] Sending evaluation scores to the device

[1263] The server sends the calculated evaluation score to the terminal. The input is the evaluation score, and the output is the score sent to the terminal. The server sends the score to the terminal's address.

[1264] Step 13:

[1265] Display of evaluation score

[1266] The terminal receives evaluation scores sent from the server and displays them to the user in real time. The input is the evaluation score from the server, and the output is the score information for display. The user can check this score.

[1267] Step 14:

[1268] Sending and analyzing emotional data

[1269] The server sends video and audio data to the emotion engine. The input is video and audio data, and the output is the data sent to the emotion engine. The emotion engine analyzes this data and generates emotion information.

[1270] Step 15:

[1271] Changes to scripts based on emotional information

[1272] The server receives the emotional information analyzed by the emotion engine and dynamically modifies the customer service script based on this information. The input is emotional information, and the output is the dynamically modified customer service script. The server generates a new script and sends it to the terminal.

[1273] Step 16:

[1274] Using the modified script

[1275] The terminal receives a newly transmitted customer service script and interacts with the user based on it. The input is the new script from the server, and the output is the interaction based on the new script. The terminal provides an interaction that reflects the script.

[1276] (Application Example 2)

[1277] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1278] In modern retail customer service, understanding and responding appropriately to customer emotions is crucial for improving customer satisfaction. However, traditional customer service systems have struggled to analyze customer emotional information in real time and dynamically adjust customer service scripts based on that analysis. This can lead to one-way communication with customers and potentially decrease customer satisfaction.

[1279] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for receiving still image data and audio data as input, means for generating moving image data based on the input still image data and audio data, means for transmitting the generated moving image data to a terminal, means for displaying the transmitted moving image data on the terminal, means for receiving image data and audio data to analyze the user's emotions, means for dynamically adjusting a customer service script based on the analyzed emotion information, means for transmitting the dynamically adjusted customer service script to a terminal, and means for interacting with the user using the transmitted customer service script on the terminal. This makes it possible to analyze the customer's emotions in real time and dynamically adjust customer service responses based on that analysis.

[1280] "Still image data" refers to individual frames that make up a video.

[1281] "Audio data" refers to data that records audio information digitally or in analog format.

[1282] "Motion data" refers to a video that is formed by sequentially combining multiple still images over time.

[1283] "Terminal" refers to an electronic device used by a user, and specifically includes personal computers, smartphones, smart glasses, etc.

[1284] "Emotional information" refers to data that indicates a user's psychological state and mood, analyzed from their facial expressions and tone of voice.

[1285] A "customer service script" refers to pre-designed text data of responses used in interactions with users.

[1286] A "generative AI model" refers to an artificial intelligence algorithm used to generate new data or text based on existing data.

[1287] A "prompt" is a text input to a generative AI model that instructs the model to produce a specific output.

[1288] This invention is a system that analyzes user emotions in real time and dynamically adjusts customer service based on that analysis, in order to improve customer service in physical stores. This system operates with a server and terminals working together and consists of the following main components.

[1289] First, the server accepts still image data and audio data as input. This data, collected by hardware such as cameras and microphones, is sent to the server. The server uses this still image data and audio data to generate video data. Specifically, video data is generated using APIs and machine learning algorithms. After generation, this video data is sent to the terminal.

[1290] The terminal receives video and image data transmitted from the server and displays it to the user. Devices such as smartphones and smart glasses are used as terminals. This terminal functions as an interface with the user and plays the video and image data.

[1291] Next, to analyze the user's emotions, the server sends image and audio data to an emotion analysis engine. This emotion analysis engine identifies the user's psychological state from their facial expressions and tone of voice and generates emotion information. This emotion information is then sent back to the server.

[1292] The server dynamically adjusts the customer service script based on this emotional information. A generative AI model is used to create prompts containing emotional information and generate an appropriate customer service script. OpenAI's GPT model is used as the generative AI model in this process. The generated customer service script is sent to the terminal, which then interacts with the user based on it.

[1293] As a concrete example, consider a scenario where a customer near a fitting room asks for feedback on a product they are trying on. In this case, the staff member performs real-time sentiment analysis via smart glasses and provides appropriate feedback based on the generated customer service script. For example, if the customer asks, "Does this jacket suit me?", the generating AI model would receive a prompt like this:

[1294] "When a customer asks, 'Does this jacket suit me?', generate the optimal customer service script. The customer's voice tone is relaxed, and their facial expression is smiling."

[1295] This system allows for real-time analysis of customer emotions and dynamic adjustments to customer service based on that analysis, which is expected to improve customer satisfaction.

[1296] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1297] Step 1:

[1298] The server accepts still image and audio data as input from the camera and microphone. This data is collected in real time to mimic human customer service staff. Still images captured by the camera and audio recorded by the microphone are sent to the server.

[1299] Step 2:

[1300] The server generates video data based on the received still image and audio data. Here, a deep learning model is used to generate the video data. This deep learning model, for example, is one that has been pre-trained on employee movements and speech patterns. The generated video data is temporarily stored on the server.

[1301] Step 3:

[1302] The server sends the generated video data to the terminal. The terminal consists of devices such as smart glasses or smartphones. The video data received by the terminal is played back for the customer to see.

[1303] Step 4:

[1304] The server sends video and audio data to the emotion analysis engine to analyze the user's emotions. The emotion analysis engine generates emotional information from the user's facial expressions and tone of voice. For example, it can determine whether the user is relaxed based on their smile or tone of voice. The analysis results are then returned to the server.

[1305] Step 5:

[1306] The server creates a prompt sentence based on the emotion information returned from the emotion analysis engine and inputs it into the generative AI model. The generative AI model uses OpenAI's GPT model to generate a customer service script based on this prompt sentence. Because the prompt sentence includes the user's voice content and emotion information, a very natural conversation is possible.

[1307] Step 6:

[1308] The server sends the generated customer service script to the terminal. The terminal uses this script to interact with the customer. For example, if a customer in a fitting room is asking for product feedback, the terminal provides an appropriate response.

[1309] Step 7:

[1310] The device collects new data based on user interaction and sends it to the server. This allows the system to continuously improve, enabling more natural and effective responses.

[1311] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1312] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1313] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1314] [Fourth Embodiment]

[1315] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1316] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1317] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1318] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1319] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1320] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1321] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1322] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1323] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1324] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1325] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1326] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1327] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1328] This invention relates to a system that generates moving image data based on still image data and audio data to provide consistent customer service. It also includes collecting and analyzing conversation data during customer service in real time to evaluate customer satisfaction. Specific embodiments of this system are described below.

[1329] 1. Generating the avatar chat screen

[1330] Server Processing

[1331] The server retrieves pre-prepared still images and audio data of employees. Next, it sends this data via an API (Application Programming Interface) to request the generation of video data. The video data returned from the API is received by the server, which then sends the video data to the terminal.

[1332] Terminal processing

[1333] The terminal receives video data transmitted from the server. After receiving the data, the terminal plays it back and displays it to the user.

[1334] Specific example

[1335] For example, if a server has still images and audio data of employee A, the server sends these to an API to generate video data. Upon receiving the video data as a response from the API, the server sends it to a terminal. The terminal plays this data, providing the user with a video that makes it appear as if employee A is speaking.

[1336] 2. Generating and using customer service scripts

[1337] Server Processing

[1338] The server collects voice data of excellent crew members' customer service interactions. This collected voice data is converted into text data using a speech recognition API. Next, the server feeds this text data into a GPT model (Greater Global Pattern Testing) to generate appropriate customer service scripts. The generated scripts are then sent from the server to the terminals.

[1339] Terminal processing

[1340] The terminal receives a customer service script sent from the server. Based on this script, the terminal interacts with the user.

[1341] Specific example

[1342] If Crew B possesses excellent customer service skills, the server records Crew B's customer service audio. This audio data is transcribed and fed into a GPT model to generate a customer service script. The terminal uses this script to provide appropriate responses to the user's questions.

[1343] 3. Talk Cancellation Scoring

[1344] Server Processing

[1345] The server collects conversation data with users in real time. The collected conversation data is converted into text data via a speech recognition API and then input into a scoring algorithm. This algorithm analyzes the conversation data and calculates a score to evaluate customer satisfaction and willingness to continue using the service. Finally, this evaluation score is sent from the server to the terminal.

[1346] Terminal processing

[1347] The device receives evaluation scores sent from the server and displays them to the user in real time.

[1348] Specific example

[1349] While a conversation with user C is in progress, the terminal collects the conversation content in real time and sends it to the server. The server transcribes this audio data and calculates an evaluation score using a scoring algorithm. By displaying this score on the terminal, it becomes possible to continuously evaluate and improve the quality of the service.

[1350] In this way, the present invention can consistently improve customer satisfaction through an AI system with advanced customer service skills.

[1351] The following describes the processing flow.

[1352] 1. Generating the avatar chat screen

[1353] Step 1:

[1354] The server retrieves pre-prepared still images and audio data of employees.

[1355] Step 2:

[1356] The server creates and sends an HTTP POST request to send the acquired still image and audio data to the API.

[1357] Step 3:

[1358] The server waits for a response from the API and receives the generated video data. This video data is generated from still images and audio data.

[1359] Step 4:

[1360] The server sends the received video data to the terminal. It uses an appropriate communication protocol (e.g., HTTP, WebSocket) for this purpose.

[1361] Step 5:

[1362] The terminal receives video data transmitted from the server. After receiving the data, the terminal plays it back and displays it to the user.

[1363] Specific example

[1364] Still images and audio data of employee A are stored on the server. The server sends this data to an API, receives video data generated by the API, and sends it to the terminal. The terminal receives the video data and displays to the user an image that makes it appear as if employee A is speaking.

[1365] 2. Generating and using customer service scripts

[1366] Step 1:

[1367] The server collects audio data of excellent customer service from top-performing crew members. This audio data is pre-recorded and stored in a database.

[1368] Step 2:

[1369] The server sends the collected customer service voice data to a speech recognition API, where it is converted into text data. The voice data is sent to the API via an HTTP POST request.

[1370] Step 3:

[1371] The server receives the text data returned from the speech recognition API.

[1372] Step 4:

[1373] The server inputs the received text data into a GPT model to generate a customer service script. The GPT model is pre-trained and can automatically generate appropriate responses.

[1374] Step 5:

[1375] The server sends the generated customer service script to the terminal, using the appropriate communication protocol.

[1376] Step 6:

[1377] The terminal receives customer service scripts sent from the server and provides appropriate responses based on the user's questions and requests.

[1378] Specific example

[1379] Crew B's excellent customer service voice is stored on the server. The server sends this voice data to a speech recognition API to transcribe it and input it into a GPT model. The generated script is sent to the terminal, which uses it to respond to the user.

[1380] 3. Talk Cancellation Scoring

[1381] Step 1:

[1382] The device collects conversation data with the user in real time. This collection is done through a voice input device.

[1383] Step 2:

[1384] The terminal sends the collected audio data to the server. An appropriate communication protocol (e.g., HTTP, WebSocket) is used.

[1385] Step 3:

[1386] The server sends the received audio data to a speech recognition API, where it is converted into text data.

[1387] Step 4:

[1388] The server receives the text data returned from the speech recognition API.

[1389] Step 5:

[1390] The server inputs the received text data into a scoring algorithm to calculate a score for evaluating customer satisfaction and willingness to continue using the service.

[1391] Step 6:

[1392] The server sends the calculated evaluation score to the terminal.

[1393] Step 7:

[1394] The device receives evaluation scores sent from the server and displays them to the user in real time.

[1395] Specific example

[1396] As a conversation with user C progresses, the terminal collects conversation data in real time and sends it to the server. The server sends the data to a speech recognition API to transcribe it and calculates an evaluation score using a scoring algorithm. This score is sent to the terminal, which then displays it to the user in real time.

[1397] In this way, the system of the present invention can provide advanced customer service skills and real-time customer evaluation, thereby consistently improving customer satisfaction.

[1398] (Example 1)

[1399] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1400] In modern society, consistent customer service is required to enhance customer satisfaction. Especially in online customer service, direct interaction with customers is difficult, and service often depends on the individual skills of employees. This can lead to inconsistencies in service quality, potentially resulting in decreased customer satisfaction. Furthermore, the lack of adequate means to evaluate customer satisfaction in real time and improve service quality makes prompt responses difficult. A system is needed to address these challenges and improve customer satisfaction.

[1401] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1402] In this invention, the server includes means for receiving still image data and audio data as input, means for generating video data based on the input still image data and audio data, means for transmitting the generated video data to a terminal, means for converting customer service audio data into text data and generating a customer service script using a generation AI model, means for transmitting the generated customer service script to a terminal, means for converting conversation data into text data and calculating an evaluation score using a scoring algorithm, and means for transmitting the calculated evaluation score to a terminal. As a result, the server generates video data and transmits it to a terminal to achieve consistent customer service, and it becomes possible to generate scripts that utilize the customer service skills of crew members with excellent customer service skills and use them on the terminal. Furthermore, it becomes possible to perform real-time customer satisfaction evaluation and continuously improve the quality of customer service.

[1403] "Still image data" refers to digital data containing image information captured at a specific point in time.

[1404] "Audio data" refers to data that records human voices in digital format.

[1405] "Motion image data" refers to digital data in video format that includes a sequence of images and synchronized audio.

[1406] A "terminal" is an electronic device that a user can directly operate.

[1407] "Customer service audio data" refers to data that digitally records the voices spoken by employees during customer service interactions.

[1408] "Text data" refers to data that visualizes audio data as text using speech recognition technology.

[1409] A "generative AI model" is an artificial intelligence model trained to generate appropriate responses or scripts from specific input data.

[1410] A "customer service script" is text data containing a series of sentences and response instructions used during customer service.

[1411] "Conversation data" refers to digital data that records voice interactions between a user and a system.

[1412] A "scoring algorithm" is a calculation method used to analyze input data and calculate an evaluation score based on specific criteria.

[1413] An "evaluation score" is a numerical evaluation result calculated by a scoring algorithm.

[1414] This invention relates to a system that generates moving image data based on still image data and audio data to provide consistent customer service. It also includes a function to collect and analyze conversation data during customer service in real time and evaluate customer satisfaction. Specific embodiments of this invention will now be described.

[1415] First, the server retrieves pre-prepared still image data (JPEG format) and audio data (WAV format) of the employee. This data is sent via an HTTP POST request using an API such as Microsoft Azure's Face API. The response returned from the API contains the generated video data (MP4 format), which the server receives and temporarily stores. Subsequently, the server sends the stored video data to the terminal as an HTTP response.

[1416] Next, the terminal receives video data sent from the server. After receiving the data, the terminal initializes a media player to play it and displays the video to the user. This series of processes allows the terminal to play a video that makes it appear as if an employee is speaking.

[1417] Next, to collect excellent customer service audio data (in WAV format) from crew members, the server uses the Google Cloud Speech-to-Text API to convert the audio data into text data. Then, the text data is fed into a generative AI model such as the GPT-3 model to generate customer service scripts. The generated scripts are temporarily stored on the server and sent to the terminal. The terminal receives the customer service scripts sent from the server and uses them to interact with the user.

[1418] Next, the server collects conversation data (in WAV format) from the user in real time. The conversation data is converted into text data using the Google Cloud Speech-to-Text API and then fed into a scoring algorithm. This algorithm analyzes the conversation data and calculates a score that evaluates customer satisfaction and willingness to continue using the service. This evaluation score is sent from the server to the terminal and displayed to the user in real time, enabling continuous evaluation and improvement of service quality.

[1419] Specific example

[1420] 1. The server sends still images and audio data of employee A to the API to generate video data. The terminal receives this data and displays a video to the user that makes it appear as if employee A is speaking.

[1421] 2. The customer service voice of Crew B is transcribed and input into the GPT-3 model to generate a customer service script. The terminal uses this script to respond to the user's questions.

[1422] 3. The conversation with User C is collected, and a customer satisfaction evaluation score is calculated using a scoring algorithm. The terminal displays this score and evaluates the service quality.

[1423] Example of a prompt

[1424] 1. "Please generate video data based on a still image of employee A and the following audio data."

[1425] 2. "Please generate a customer service script based on the text data of this customer service audio."

[1426] 3. "Based on the conversation data, calculate a score to evaluate customer satisfaction."

[1427] In this way, the present invention provides a system that generates moving image data using still image data and audio data, evaluates customer satisfaction in real time, and realizes consistently high-quality customer service.

[1428] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1429] Generating an avatar chat screen

[1430] Server Processing

[1431] Step 1:

[1432] The server retrieves pre-prepared still image data (e.g., JPEG format) and audio data (e.g., WAV format) of employees. The server receives the still image and audio data as input. This data serves as the basic information for generating video data in subsequent processing.

[1433] Step 2:

[1434] The server sends the acquired still image and audio data to an API (e.g., Microsoft Azure's Face API). It makes an HTTP POST request to the API endpoint, sending the still image and audio data. Based on the input data, the API executes a process and generates a video image data by combining the still image and audio data.

[1435] Step 3:

[1436] The server receives video data (e.g., in MP4 format) returned from the API. It temporarily stores the video data obtained as an HTTP response and prepares for the subsequent transmission process.

[1437] Step 4:

[1438] The server sends the generated video data to the terminal as an HTTP response. This allows the video data to be delivered to the terminal via the internet.

[1439] Terminal processing

[1440] Step 1:

[1441] The terminal receives video data sent from the server as an HTTP response. It stores the received data in temporary storage. It then parses and stores the URL or binary data used as input data.

[1442] Step 2:

[1443] The device initializes its media player for playing saved video data. It internally loads the video data and prepares it for display to the user.

[1444] Step 3:

[1445] The device plays video and displays it to the user. It uses a media player to display the video along a timeline and outputs it in a way that the user can perceive.

[1446] Generating and using customer service scripts

[1447] Server Processing

[1448] Step 1:

[1449] The server collects voice data (e.g., in WAV format) of high-performing crew members providing excellent customer service. The collected voice data is temporarily stored to prepare for the subsequent speech recognition process.

[1450] Step 2:

[1451] The server sends the collected audio data to the Google Cloud Speech-to-Text API, where it is converted into text data. The audio data is then sent as input to the speech recognition API, and an HTTP response containing the text data is received. This converts the audio data into text data from start to finish.

[1452] Step 3:

[1453] The server inputs the obtained character data into a GPT-3 model to generate customer service scripts. The character data is input into a generation AI model to generate customer service scripts as corresponding text data.

[1454] Step 4:

[1455] The server temporarily stores the generated customer service script and then sends it to the terminal as an HTTP response. This delivers the generated script to the user's terminal.

[1456] Terminal processing

[1457] Step 1:

[1458] The terminal receives the customer service script sent from the server. It retrieves the script information as an HTTP response and saves it to temporary storage. It then reads the contents of the script as input data and saves it.

[1459] Step 2:

[1460] The terminal handles customer service interactions based on saved customer service scripts. It dynamically references the script in response to user input and generates appropriate responses. Specifically, it responds to user inquiries and requests according to the instructions in the script.

[1461] Talk cancellation scoring

[1462] Server Processing

[1463] Step 1:

[1464] The server collects conversation data with users (e.g., in WAV format) in real time. Each time a conversation occurs, it is captured as audio data and stored on the server. This data serves as foundational information for analyzing customer satisfaction in subsequent processing.

[1465] Step 2:

[1466] The server sends the collected audio data to the Google Cloud Speech-to-Text API, where it is converted into text data. The audio data is sent to the speech recognition API, and an HTTP response is received as text data. This converts the conversation data into text data.

[1467] Step 3:

[1468] The server inputs the obtained character data into a scoring algorithm to calculate an evaluation score. The character data is input to the scoring algorithm, and a score is calculated based on the evaluation criteria.

[1469] Step 4:

[1470] The server temporarily stores the calculated evaluation score and then sends it to the terminal as an HTTP response. This delivers the generated evaluation score to the user's terminal.

[1471] Terminal processing

[1472] Step 1:

[1473] The terminal receives the evaluation score sent from the server. It retrieves the score information as an HTTP response and saves it to temporary storage. It then reads the score content as input data and saves it.

[1474] Step 2:

[1475] The device displays saved evaluation scores to the user in real time. The interface for displaying evaluation scores is initialized, and customer satisfaction scores are displayed dynamically. This enables real-time evaluation of service quality.

[1476] Through these steps, the system generates video data using still image and audio data, enabling consistent customer service and real-time customer satisfaction evaluation.

[1477] (Application Example 1)

[1478] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1479] In traditional brick-and-mortar stores, inadequate customer service was sometimes difficult due to staff absences or insufficient manpower. Furthermore, maintaining consistent service quality was challenging, and there were limited means to evaluate and improve customer satisfaction in real time. This invention aims to solve these problems by using an AI system with advanced customer service skills to consistently improve customer satisfaction.

[1480] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1481] In this invention, the server includes means for receiving still image data and audio data as input in order to convert still image data into video data; means for generating video data based on the input still image data and audio data; means for transmitting the generated video data to a terminal; means for displaying the transmitted video data on the terminal; means for converting customer service audio data into speech recognition data; means for generating a customer service script based on the speech recognition data; means for evaluating conversation data in real time and calculating a customer satisfaction evaluation score; and means for transmitting the evaluation score to a terminal and displaying it on the terminal. This enables virtual customer service by an avatar host using a smartphone application, allowing for consistent, high-quality customer service even when staff are absent in a physical store. Furthermore, it makes it possible to evaluate customer satisfaction in real time and continuously improve the quality of service.

[1482] "Still image data" refers to data of a static image that does not move.

[1483] "Motion data" refers to video data generated using a series of still images.

[1484] "Audio data" refers to data that records a person's voice in digital format.

[1485] "Speech recognition data" refers to data obtained by analyzing speech data and representing its content as text data.

[1486] A "customer service script" is text data that describes the content of conversations and response methods during customer service, and is used to instruct and guide customer service operations.

[1487] "Conversation data" refers to language-based information data exchanged between customers and systems or staff.

[1488] An "evaluation score" is a numerical value calculated based on conversation data and customer responses, and is an indicator used to evaluate customer satisfaction and service quality.

[1489] A "terminal" refers to an electronic device used to receive and display data transmitted from a server.

[1490] A "scoring algorithm" is a set of computational methods and rules used to analyze conversational data and other input data and calculate an evaluation score.

[1491] "Motion image generation means" refers to a method or apparatus for creating motion image data based on still image data and audio data.

[1492] This invention provides a specific method for constructing a virtual customer service system using a smartphone application in a physical store. This system generates moving image data based on still image data and audio data to provide consistent customer service to customers. It also collects and analyzes conversation data during customer service in real time to evaluate customer satisfaction.

[1493] Hardware and software configuration

[1494] Servers and smartphones will be used as the primary hardware. Specifically, the following:

[1495] server:

[1496] A cloud server for storing and processing still image and audio data.

[1497] Speech-to-Text is used to convert speech data into text data.

[1498] A video image generation API (e.g., DeepMotion) is used to generate video image data from still images and audio.

[1499] A GPT model (e.g., OpenAI's GPT-4) is used to generate customer service scripts from text data.

[1500] A scoring algorithm is used to evaluate customer satisfaction in real time.

[1501] Smartphone device:

[1502] A device for receiving, displaying, and playing back data transmitted from a server.

[1503] Collect user interactions and send them to the server.

[1504] Data processing and data calculation

[1505] Video generation:

[1506] The server accepts still image data and audio data as input. This data is passed to a video generation API to generate video data. The generated video data is then sent from the server to the terminal. The terminal displays this data, providing the user with an avatar image.

[1507] Customer service script generation:

[1508] The server converts customer service voice data into text data using a speech recognition API. The text data is input into a GPT model to generate an appropriate customer service script. This script is sent from the server to the terminal, which then uses this script to interact with the customer.

[1509] Customer satisfaction rating:

[1510] The server collects conversation data in real time and converts it into text data using a speech recognition API. The converted text data is analyzed by a scoring algorithm to calculate an evaluation score. This evaluation score is sent to the terminal and displayed to the user in real time.

[1511] Specific example

[1512] For example, in a clothing store, if a customer asks for trousers to match a jacket when no staff are present, this system would function effectively. If the customer asks, "Please recommend trousers that would go with this jacket," the system generates an avatar, which is displayed as video data. Furthermore, based on a customer service script, the avatar can provide an appropriate response and address the customer's questions. It is possible to collect data during the conversation in real time and display customer satisfaction as an evaluation score.

[1513] Example of a prompt

[1514] Please generate a customer service script for recommending trousers to match a jacket to a customer visiting a clothing store. Additionally, design an algorithm to evaluate customer satisfaction in real time during the interaction.

[1515] Thus, the system of the present invention can consistently improve customer satisfaction in physical stores by using an AI system with advanced customer service skills.

[1516] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1517] Step 1:

[1518] The server receives still image and audio data. When a user uploads still images and audio of employees during the store's preparation phase, the server retrieves them. The input data consists of still image and audio data, while the output data is a dataset to be passed to the video generation API.

[1519] Step 2:

[1520] The server uses the received still image and audio data to call a video generation API and generate video data. By providing still image and audio data as input to this video generation API (e.g., DeepMotion), animated video of employees is generated. The input data consists of image and audio files for the API call, and the output data is the generated video data.

[1521] Step 3:

[1522] The server sends the generated video data to the terminal. The server receives the video data and sends it to the terminal using a communication protocol. The input data is the video data, and the output data is the confirmation message sent to the terminal.

[1523] Step 4:

[1524] The terminal receives and displays video data transmitted from the server. The user operates the terminal to play the video data, which then displays an avatar. The input data is video data, and the output data is the display of the avatar image.

[1525] Step 5:

[1526] The server receives audio data during customer service interactions and converts it into text data using a speech recognition API. It collects the audio of the user-avatar conversation in real time and sends it to a speech recognition API such as Google Cloud Speech-to-Text to obtain text data. The input data is audio data, and the output data is the converted text data.

[1527] Step 6:

[1528] The server inputs text data into a GPT model and generates a customer service script. This text data is then input into an OpenAI GPT-4 model, and a process is executed to generate an appropriate customer service script. The input data is text data, and the output data is the customer service script.

[1529] Step 7:

[1530] The server sends the generated customer service script to the terminal. Sending the generated customer service script to the terminal enables real-time interaction with the user on the terminal. The input data is the customer service script, and the output data is the confirmation message sent to the terminal.

[1531] Step 8:

[1532] The terminal interacts with the user based on the received customer service script. The script is displayed on the terminal, and the avatar responds accordingly. The input data is the customer service script, and the output data is the result of the conversation with the user.

[1533] Step 9:

[1534] The server collects conversation data with the user in real time and converts it into text data using a speech recognition API. The input data is the audio data of the conversation, and the output data is the converted text data.

[1535] Step 10:

[1536] The server inputs the converted character data into a scoring algorithm to calculate an evaluation score. The input data is the converted character data, and the output data is the evaluation score.

[1537] Step 11:

[1538] The server sends the calculated evaluation score to the terminal, which then displays the evaluation score. The input data is the evaluation score, and the output data is the customer satisfaction evaluation result displayed on the terminal.

[1539] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1540] This invention relates to a system that recognizes user emotions and provides dynamic customer service responses based on those emotions. This system generates dynamic image data from still image data and audio data, and further generates customer service scripts by converting customer service audio data into text data. In addition, by combining this with an emotion engine that evaluates conversation data and recognizes user emotions, it achieves more sophisticated customer service.

[1541] 1. Generating the avatar chat screen

[1542] Server Processing

[1543] The server retrieves pre-prepared still images and audio data of employees. This data is sent to the API, which requests the generation of video data. The video data returned from the API is received by the server and then sent to the terminal.

[1544] Terminal processing

[1545] The terminal receives video data transmitted from the server, plays this data, and displays it to the user.

[1546] Specific example

[1547] For example, if a server has still images and audio data of employee A, it sends this data to an API to generate video data. Upon receiving the video data as a response from the API, the server sends it to a terminal. The terminal plays this video data, providing the user with a video that makes it appear as if employee A is speaking.

[1548] 2. Generating and using customer service scripts

[1549] Server Processing

[1550] The server collects voice data of high-performing crew members' customer service interactions. This voice data is converted into text data using a speech recognition API. The text data is input into a GPT model to generate customer service scripts. The generated scripts are sent from the server to the terminals.

[1551] Terminal processing

[1552] The terminal receives customer service scripts sent from the server and uses these scripts to interact with the user.

[1553] Specific example

[1554] Crew B's customer service voice data is stored on the server. The server sends this voice data to a speech recognition API to transcribe it and input it into a GPT model. The generated script is sent to the terminal, which then provides an appropriate response based on the user's questions and requests.

[1555] 3. Talk Cancellation Scoring

[1556] Server Processing

[1557] The server collects conversation data with the user in real time. This conversation data is converted into text data via a speech recognition API and input into a scoring algorithm. The algorithm calculates an evaluation score, which is then sent from the server to the terminal.

[1558] Terminal processing

[1559] The device receives evaluation scores sent from the server and displays them to the user in real time.

[1560] Specific example

[1561] If a conversation with user C is in progress, the terminal collects conversation data in real time and sends it to the server. The server transcribes the audio data and calculates an evaluation score using a scoring algorithm. This score is sent to the terminal, which displays it to the user in real time.

[1562] 4. Adding an emotion engine

[1563] Server Processing

[1564] The server sends video and audio data to the emotion engine. This engine analyzes the user's facial expressions and tone of voice to generate emotion information. This emotion information is sent back to the server, which is then instructed to dynamically modify the customer service script based on the user's emotions.

[1565] Terminal processing

[1566] The device receives emotional information and adjusts customer service scripts based on it. It also uses emotional information to provide more appropriate responses.

[1567] Specific example

[1568] If user D is experiencing stress, the emotion engine analyzes this and sends stress information to the server. Based on this information, the server generates a dedicated customer service script and sends it to the terminal. The terminal uses this script to provide a more relaxed interaction with user D.

[1569] The system of this invention can improve customer satisfaction by analyzing user emotions in real time and dynamically changing customer service responses based on those emotions.

[1570] The following describes the processing flow.

[1571] Processing flow of a system combining an emotion engine

[1572] 1. Generating the avatar chat screen

[1573] Step 1:

[1574] The server retrieves pre-prepared still images and audio data of employees.

[1575] Step 2:

[1576] The server creates and sends an HTTP POST request to the API to send the acquired still image and audio data.

[1577] Step 3:

[1578] The server waits for a response from the API and receives the generated video data. This video data is generated from still images and audio data.

[1579] Step 4:

[1580] The server sends the received video data to the terminal. It uses an appropriate communication protocol (e.g., HTTP, WebSocket) for this purpose.

[1581] Step 5:

[1582] The terminal receives video data transmitted from the server. After receiving the data, the terminal plays it back and displays it to the user.

[1583] 2. Generating and using customer service scripts

[1584] Step 1:

[1585] The server collects audio data of excellent customer service from top-performing crew members. This audio data is pre-recorded and stored in a database.

[1586] Step 2:

[1587] The server sends the collected customer service voice data to a speech recognition API, where it is converted into text data. The voice data is sent to the API via an HTTP POST request.

[1588] Step 3:

[1589] The server receives the text data returned from the speech recognition API.

[1590] Step 4:

[1591] The server inputs the received text data into a GPT model to generate a customer service script. The GPT model is pre-trained and can automatically generate appropriate responses.

[1592] Step 5:

[1593] The server sends the generated customer service script to the terminal, using the appropriate communication protocol.

[1594] Step 6:

[1595] The terminal receives customer service scripts sent from the server and provides appropriate responses based on the user's questions and requests.

[1596] 3. Talk Cancellation Scoring

[1597] Step 1:

[1598] The device collects conversation data with the user in real time. This collection is done through a voice input device.

[1599] Step 2:

[1600] The terminal sends the collected audio data to the server. An appropriate communication protocol (e.g., HTTP, WebSocket) is used.

[1601] Step 3:

[1602] The server sends the received audio data to a speech recognition API, where it is converted into text data.

[1603] Step 4:

[1604] The server receives the text data returned from the speech recognition API.

[1605] Step 5:

[1606] The server inputs the received text data into a scoring algorithm to calculate a score for evaluating customer satisfaction and willingness to continue using the service.

[1607] Step 6:

[1608] The server sends the calculated evaluation score to the terminal.

[1609] Step 7:

[1610] The device receives evaluation scores sent from the server and displays them to the user in real time.

[1611] 4. Adding an emotion engine

[1612] Step 1:

[1613] The server sends video and audio data to the emotion engine. The emotion engine analyzes the user's emotions based on their facial expressions and tone of voice.

[1614] Step 2:

[1615] The emotion engine analyzes the user's emotions and sends that emotional information back to the server.

[1616] Step 3:

[1617] The server dynamically modifies customer service scripts based on emotional information received from the emotion engine. For example, if a user is dissatisfied, it generates a more courteous response.

[1618] Step 4:

[1619] The server sends the modified customer service script to the terminal, using the appropriate communication protocol.

[1620] Step 5:

[1621] The terminal receives emotional information and customer service scripts sent from the server and responds appropriately to the user.

[1622] Specific example

[1623] If user D is experiencing stress, the emotion engine analyzes this and sends this information to the server. Based on the emotion information, the server generates a dedicated customer service script and sends it to the terminal. The terminal uses this script to provide a relaxed interaction with user D.

[1624] Thus, the system of the present invention provides a mechanism that can improve customer satisfaction by analyzing the user's emotions in real time through each processing step and dynamically changing customer service responses based on those emotions.

[1625] (Example 2)

[1626] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1627] Traditional customer service systems struggled to analyze user emotions in real time based on facial expressions and tone of voice, and to dynamically adjust customer service accordingly. As a result, they were unable to enhance user satisfaction and could only provide standardized responses, failing to offer optimal service to individual users.

[1628] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1629] In this invention, the server includes means for receiving still image data and audio data as input in order to convert still image data into video data; means for generating video data based on the input still image data and audio data; means for transmitting the generated video data to a terminal; means for receiving customer service audio data as input in order to transcribe customer service audio data; means for converting the input customer service audio data into text data; means for generating a customer service script based on the text data; means for transmitting the generated customer service script to a terminal; means for receiving conversation data as input in order to evaluate conversation data; means for converting the input conversation data into text data; means for inputting the text data into a scoring algorithm to calculate an evaluation score; means for receiving video data and audio data as input and analyzing the user's emotions; and means for dynamically changing the customer service script based on the analyzed emotion information. This enables dynamic responses that respond to the user's emotions.

[1630] 1. "Still image data" refers to a single, still image or photograph that is saved as an image file.

[1631] 2. "Motion image data" is a series of images consisting of multiple frames, and is visual information that changes over time.

[1632] 3. "Audio data" refers to data that stores audio in digital format, and includes information such as spoken language and sounds.

[1633] 4. "Customer service audio data" refers to data that records audio from customer service situations.

[1634] 5. "Text data" refers to data stored in text format, which is information expressed as words or sentences.

[1635] 6. A "customer service script" is a series of text messages generated to assist in interactions with users.

[1636] 7. "Conversation data" refers to a record of a conversation, either in audio or text format, between a user and a customer service representative.

[1637] 8. A "scoring algorithm" is a calculation method for determining an evaluation score based on input data.

[1638] 9. The "evaluation score" is a numerical value calculated by a scoring algorithm and is used to evaluate the quality of conversation and emotional state.

[1639] 10. An "emotion engine" is software that analyzes video and audio data to determine the user's emotional state.

[1640] 11. "Analyzed emotional information" refers to data about the user's emotional state obtained as a result of analysis by the emotion engine.

[1641] 12. A "server" is a computer system used to store, process, and distribute data over a network.

[1642] 13. A "terminal" is a device that a user directly operates and has the function of communicating with a server to send and receive data.

[1643] 14. A "generative AI model" is an artificial intelligence model used to learn from large amounts of data and generate new text or responses.

[1644] This invention relates to a system that recognizes user emotions and provides dynamic customer service responses based on those emotions. This system consists of a server and a terminal, each performing the following functions.

[1645] The server first retrieves pre-prepared still image and audio data of employees from the database. This still image and audio data is sent to the video generation API and converted into video data. The converted video data is sent back to the server, which then sends it to the terminal. The terminal plays the received video data and displays it to the user.

[1646] Next, the server collects voice data of high-performing crew members and converts this voice data into text data using a speech recognition API. The converted text data is input into a generative AI model (e.g., GPT-4) to generate a customer service script. The generated script is sent from the server to the terminal, which uses this script to interact with the user.

[1647] The server further collects conversation data with the user in real time and converts this data into text data via a speech recognition API. The converted text data is input into a scoring algorithm, which calculates an evaluation score. This evaluation score is sent from the server to the terminal, which displays the score to the user in real time.

[1648] The server also sends video and audio data to the emotion engine. This emotion engine analyzes the user's facial expressions and tone of voice and generates emotion information. The generated emotion information is sent back to the server, which dynamically modifies the customer service script based on this information. The terminal receives this modified script and provides a more appropriate response.

[1649] For example, consider a case where a server holds still images and audio data of employee A. This data is sent to an API to generate video data. The generated video data is sent to a terminal via the server, and the terminal displays it, providing the user with a video that makes it appear as if employee A is actually speaking.

[1650] Furthermore, if Crew B's customer service voice data is stored on the server, the server sends this data to a speech recognition API and converts it into text data. Next, this text data is input into a generation AI model to generate a customer service script. This script is sent to the terminal, which then provides an appropriate response based on the user's questions and requests.

[1651] Furthermore, if a conversation with user C is ongoing, the terminal collects conversation data in real time and sends it to the server. The server converts the audio data into text data and calculates an evaluation score using a scoring algorithm. This score is sent to the terminal, which then displays it to the user in real time.

[1652] Finally, if user D appears to be experiencing stress, the emotion engine analyzes this and sends stress information to the server. Based on this information, the server generates a customized customer service script and sends it to the terminal. The terminal uses this script to provide a more relaxed interaction with user D.

[1653] Examples of prompt messages include the following:

[1654] "Could you tell me about your recent orders?"

[1655] "I'd like to learn more about premium membership."

[1656] "Please tell me how to return an item."

[1657] This system allows for real-time analysis of user emotions and dynamic adjustments to customer service responses based on those analyses, thereby increasing customer satisfaction.

[1658] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1659] Step 1:

[1660] Acquisition of still image data and audio data

[1661] The server retrieves pre-prepared still image and audio data of employees from the database. The input is the still image and audio data in the database, and the output is the retrieved still image and audio data. This data is then ready to be sent to the API.

[1662] Step 2:

[1663] Sending data to the API

[1664] The server sends the acquired still image and audio data to the video generation API. The input is the acquired still image and audio data, and the output is the request to send to the API. The server constructs the appropriate request to the API and sends the data.

[1665] Step 3:

[1666] Receiving data from API

[1667] The server receives video data sent back from the API. The input is the response data from the API, and the output is the received video data. The server receives this video data and prepares to send it to the terminal.

[1668] Step 4:

[1669] Sending video data to the terminal

[1670] The server sends the received video data to the terminal. The input is the video data from the API, and the output is the video data sent to the terminal. The server sends data to the terminal's address.

[1671] Step 5:

[1672] Receiving and displaying video data.

[1673] The terminal receives video data transmitted from the server. The input is video data from the server, and the output is video data for display. The terminal plays this video data and displays it to the user.

[1674] Step 6:

[1675] Collection and transcription of customer service voice data

[1676] The server collects customer service voice data from excellent crew members and converts this voice data into text data using a speech recognition API. The input is customer service voice data, and the output is text data. By converting to text data, the server prepares it for feeding into a generative AI model.

[1677] Step 7:

[1678] Input to the Generative AI Model

[1679] The server inputs the converted character data into a generating AI model (e.g., GPT-4) to generate a customer service script. The input is character data, and the output is a customer service script. The generating AI model uses the character data to create a script suitable for the service.

[1680] Step 8:

[1681] Sending the script to the terminal

[1682] The server sends the generated script to the terminal. The input is the generated customer service script, and the output is the script sent to the terminal. The server sends the script to the terminal's address.

[1683] Step 9:

[1684] Use of the submitted script

[1685] The terminal receives customer service scripts sent from the server. The input is the customer service script from the server, and the output is the response based on the script. The terminal uses the script to interact with the user.

[1686] Step 10:

[1687] Collection and transcription of conversation data

[1688] The server collects conversation data with the user and converts this data into text data via a speech recognition API. The input is conversation data, and the output is text data. The server then inputs the converted text data into a scoring algorithm.

[1689] Step 11:

[1690] Calculation of evaluation score

[1691] The server inputs text data into a scoring algorithm and calculates an evaluation score. The input is text data, and the output is the evaluation score. The algorithm works to evaluate the quality and emotional state of the conversation.

[1692] Step 12:

[1693] Sending evaluation scores to the device

[1694] The server sends the calculated evaluation score to the terminal. The input is the evaluation score, and the output is the score sent to the terminal. The server sends the score to the terminal's address.

[1695] Step 13:

[1696] Display of evaluation score

[1697] The terminal receives evaluation scores sent from the server and displays them to the user in real time. The input is the evaluation score from the server, and the output is the score information for display. The user can check this score.

[1698] Step 14:

[1699] Sending and analyzing emotional data

[1700] The server sends video and audio data to the emotion engine. The input is video and audio data, and the output is the data sent to the emotion engine. The emotion engine analyzes this data and generates emotion information.

[1701] Step 15:

[1702] Changes to scripts based on emotional information

[1703] The server receives the emotional information analyzed by the emotion engine and dynamically modifies the customer service script based on this information. The input is emotional information, and the output is the dynamically modified customer service script. The server generates a new script and sends it to the terminal.

[1704] Step 16:

[1705] Using the modified script

[1706] The terminal receives a newly transmitted customer service script and interacts with the user based on it. The input is the new script from the server, and the output is the interaction based on the new script. The terminal provides an interaction that reflects the script.

[1707] (Application Example 2)

[1708] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1709] In modern retail customer service, understanding and responding appropriately to customer emotions is crucial for improving customer satisfaction. However, traditional customer service systems have struggled to analyze customer emotional information in real time and dynamically adjust customer service scripts based on that analysis. This can lead to one-way communication with customers and potentially decrease customer satisfaction.

[1710] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for receiving still image data and audio data as input, means for generating moving image data based on the input still image data and audio data, means for transmitting the generated moving image data to a terminal, means for displaying the transmitted moving image data on the terminal, means for receiving image data and audio data to analyze the user's emotions, means for dynamically adjusting a customer service script based on the analyzed emotion information, means for transmitting the dynamically adjusted customer service script to a terminal, and means for interacting with the user using the transmitted customer service script on the terminal. This makes it possible to analyze the customer's emotions in real time and dynamically adjust customer service responses based on that analysis.

[1711] "Still image data" refers to individual frames that make up a video.

[1712] "Audio data" refers to data that records audio information digitally or in analog format.

[1713] "Motion data" refers to a video that is formed by sequentially combining multiple still images over time.

[1714] "Terminal" refers to an electronic device used by a user, and specifically includes personal computers, smartphones, smart glasses, etc.

[1715] "Emotional information" refers to data that indicates a user's psychological state and mood, analyzed from their facial expressions and tone of voice.

[1716] A "customer service script" refers to pre-designed text data of responses used in interactions with users.

[1717] A "generative AI model" refers to an artificial intelligence algorithm used to generate new data or text based on existing data.

[1718] A "prompt" is a text input to a generative AI model that instructs the model to produce a specific output.

[1719] This invention is a system that improves customer service in physical stores by analyzing user emotions in real time and dynamically adjusting customer service based on that analysis. This system operates with a server and terminals working together and consists of the following main components.

[1720] First, the server accepts still image data and audio data as input. This data, collected by hardware such as cameras and microphones, is sent to the server. The server uses this still image data and audio data to generate video data. Specifically, video data is generated using APIs and machine learning algorithms. After generation, this video data is sent to the terminal.

[1721] The terminal receives video and image data transmitted from the server and displays it to the user. Devices such as smartphones and smart glasses are used as terminals. This terminal functions as an interface with the user and plays the video and image data.

[1722] Next, to analyze the user's emotions, the server sends image and audio data to an emotion analysis engine. This emotion analysis engine identifies the user's psychological state from their facial expressions and tone of voice and generates emotion information. This emotion information is then sent back to the server.

[1723] The server dynamically adjusts the customer service script based on this emotional information. A generative AI model is used to create prompts containing emotional information and generate an appropriate customer service script. OpenAI's GPT model is used as the generative AI model in this process. The generated customer service script is sent to the terminal, which then interacts with the user based on it.

[1724] As a concrete example, consider a scenario where a customer near a fitting room asks for feedback on a product they are trying on. In this case, the staff member performs real-time sentiment analysis via smart glasses and provides appropriate feedback based on the generated customer service script. For example, if the customer asks, "Does this jacket suit me?", the generating AI model would receive a prompt like this:

[1725] "When a customer asks, 'Does this jacket suit me?', generate the optimal customer service script. The customer's voice tone is relaxed, and their facial expression is smiling."

[1726] This system allows for real-time analysis of customer emotions and dynamic adjustments to customer service based on that analysis, which is expected to improve customer satisfaction.

[1727] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1728] Step 1:

[1729] The server accepts still image and audio data as input from the camera and microphone. This data is collected in real time to mimic human customer service staff. Still images captured by the camera and audio recorded by the microphone are sent to the server.

[1730] Step 2:

[1731] The server generates video data based on the received still image and audio data. Here, a deep learning model is used to generate the video data. This deep learning model, for example, is one that has been pre-trained on employee movements and speech patterns. The generated video data is temporarily stored on the server.

[1732] Step 3:

[1733] The server sends the generated video data to the terminal. The terminal consists of devices such as smart glasses or smartphones. The video data received by the terminal is played back for the customer to see.

[1734] Step 4:

[1735] The server sends video and audio data to the emotion analysis engine to analyze the user's emotions. The emotion analysis engine generates emotional information from the user's facial expressions and tone of voice. For example, it can determine whether the user is relaxed based on their smile or tone of voice. The analysis results are then returned to the server.

[1736] Step 5:

[1737] The server creates a prompt sentence based on the emotion information returned from the emotion analysis engine and inputs it into the generative AI model. The generative AI model uses OpenAI's GPT model to generate a customer service script based on this prompt sentence. Because the prompt sentence includes the user's voice content and emotion information, a very natural conversation is possible.

[1738] Step 6:

[1739] The server sends the generated customer service script to the terminal. The terminal uses this script to interact with the customer. For example, if a customer in a fitting room is asking for product feedback, the terminal provides an appropriate response.

[1740] Step 7:

[1741] The device collects new data based on user interaction and sends it to the server. This allows the system to continuously improve, enabling more natural and effective responses.

[1742] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1743] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1744] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[1745] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1746] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[1747] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[1748] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[1749] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[1750] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[1751] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[1752] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[1753] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[1754] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[1755] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1756] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[1757] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[1758] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[1759] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[1760] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[1761] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[1762] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[1763] The following is further disclosed regarding the embodiments described above.

[1764] (Claim 1)

[1765] In order to convert still image data into video data, a means for receiving the still image data and audio data as input,

[1766] A means for generating video data based on the input still image data and audio data,

[1767] Means for transmitting the generated video data to a terminal,

[1768] The terminal includes means for displaying the transmitted video data,

[1769] A system that includes this.

[1770] (Claim 2)

[1771] In order to transcribe customer service voice data into text, a means for receiving the customer service voice data as input,

[1772] A means for converting the input customer service voice data into text data,

[1773] Means for generating a customer service script based on the aforementioned text data,

[1774] A means for sending the generated customer service script to a terminal,

[1775] The terminal provides a means for interacting with a user using the transmitted customer service script,

[1776] A system that includes this.

[1777] (Claim 3)

[1778] In order to evaluate the conversation data, a means for receiving the conversation data as input,

[1779] Means for converting the input conversation data into text data,

[1780] A means for inputting the aforementioned character data into a scoring algorithm to calculate an evaluation score,

[1781] A means for transmitting the calculated evaluation score to the terminal,

[1782] The terminal includes means for displaying the transmitted evaluation score,

[1783] A system that includes this.

[1784] "Example 1"

[1785] (Claim 1)

[1786] In order to convert still image data into video data, a means for receiving the still image data and audio data as input,

[1787] A means for generating video data based on the input still image data and audio data,

[1788] Means for transmitting the generated video data to a terminal,

[1789] The terminal includes means for displaying the transmitted video data,

[1790] In order to transcribe customer service voice data into text, a means for receiving the customer service voice data as input,

[1791] A means for converting the input customer service voice data into text data,

[1792] Means for generating a customer service script based on the aforementioned text data,

[1793] A means for sending the generated customer service script to a terminal,

[1794] The terminal provides a means for interacting with a user using the transmitted customer service script,

[1795] In order to evaluate the conversation data, a means for receiving the conversation data as input,

[1796] Means for converting the input conversation data into text data,

[1797] A means for inputting the aforementioned character data into a scoring algorithm to calculate an evaluation score,

[1798] A means for transmitting the calculated evaluation score to the terminal,

[1799] The terminal includes means for displaying the transmitted evaluation score,

[1800] A system that includes this.

[1801] (Claim 2)

[1802] A means for generating a customer service script using a generation AI model, based on customer service voice data of crew members with excellent customer service skills, said customer service voice data, said customer service voice data, said customer service voice data, said customer service voice data, said customer service script,

[1803] The system according to claim 1, wherein the generated customer service script is used in responding to users.

[1804] (Claim 3)

[1805] A means for converting conversation data with users into text data using speech recognition technology and calculating a customer satisfaction evaluation score using a scoring algorithm,

[1806] The system according to claim 1, which displays the calculated evaluation score on a terminal in real time.

[1807] "Application Example 1"

[1808] (Claim 1)

[1809] In order to convert still image data into video data, a means for receiving the still image data and audio data as input,

[1810] A means for generating video data based on the input still image data and audio data,

[1811] Means for transmitting the generated video data to a terminal,

[1812] The terminal includes means for displaying the transmitted video data,

[1813] A means of converting customer service voice data into speech recognition data,

[1814] Means for generating a customer service script based on the aforementioned speech recognition data,

[1815] A method for evaluating conversation data in real time and calculating a customer satisfaction evaluation score,

[1816] A means for transmitting the aforementioned evaluation score to a terminal and displaying it on the terminal,

[1817] A system that includes this.

[1818] (Claim 2)

[1819] In order to transcribe customer service voice data into text, a means for receiving the customer service voice data as input,

[1820] A means for converting the input customer service voice data into text data,

[1821] Means for generating a customer service script based on the aforementioned text data,

[1822] A means for sending the generated customer service script to a terminal,

[1823] The terminal provides a means for interacting with a user using the transmitted customer service script,

[1824] The system according to claim 1, including the following:

[1825] (Claim 3)

[1826] In order to evaluate the conversation data, a means for receiving the conversation data as input,

[1827] Means for converting the input conversation data into text data,

[1828] A means for inputting the aforementioned character data into a scoring algorithm to calculate an evaluation score,

[1829] A means for transmitting the calculated evaluation score to the terminal,

[1830] The terminal includes means for displaying the transmitted evaluation score,

[1831] The system according to claim 1, including the following:

[1832] "Example 2 of combining an emotion engine"

[1833] (Claim 1)

[1834] In order to convert still image data into video data, a means for receiving the still image data and audio data as input,

[1835] A means for generating video data based on the input still image data and audio data,

[1836] Means for transmitting the generated video data to a terminal,

[1837] The terminal includes means for displaying the transmitted video data,

[1838] In order to transcribe customer service voice data into text, a means for receiving the customer service voice data as input,

[1839] A means for converting the input customer service voice data into text data,

[1840] Means for generating a customer service script based on the aforementioned text data,

[1841] A means for sending the generated customer service script to a terminal,

[1842] The terminal provides a means for interacting with a user using the transmitted customer service script,

[1843] In order to evaluate the conversation data, a means for receiving the conversation data as input,

[1844] Means for converting the input conversation data into text data,

[1845] A means for inputting the aforementioned character data into a scoring algorithm to calculate an evaluation score,

[1846] A means for transmitting the calculated evaluation score to the terminal,

[1847] The terminal includes means for displaying the transmitted evaluation score,

[1848] A means for receiving video and audio data as input and analyzing the user's emotions,

[1849] Means for dynamically changing the customer service script based on the analyzed emotional information,

[1850] Means for sending the dynamically modified customer service script to the terminal,

[1851] The terminal provides a means for interacting with a user using the transmitted customer service script,

[1852] A system that includes this.

[1853] (Claim 2)

[1854] The system according to claim 1, comprising means for collecting customer service voice data, converting the customer service voice data into text data, inputting it into a generation AI model to generate a customer service script, transmitting the generated customer service script to a terminal, and using the customer service script to interact with a user at the terminal.

[1855] (Claim 3)

[1856] The system according to claim 1, comprising means for collecting conversation data with a user, converting the conversation data into text data using a speech recognition API, calculating an evaluation score using a scoring algorithm, transmitting the evaluation score to a terminal, and displaying the score in real time on the terminal.

[1857] "Application example 2 when combining with an emotional engine"

[1858] (Claim 1)

[1859] In order to convert still image data into video data, a means for receiving the still image data and audio data as input,

[1860] A means for generating video data based on the input still image data and audio data,

[1861] Means for transmitting the generated video data to a terminal,

[1862] The terminal includes means for displaying the transmitted video data,

[1863] A means for receiving image data and audio data in order to analyze the user's emotions,

[1864] A means for dynamically adjusting the customer service script based on the analyzed emotional information,

[1865] A means for sending the dynamically adjusted customer service script to the terminal,

[1866] The terminal provides a means for interacting with a user using the transmitted customer service script,

[1867] A system that includes this.

[1868] (Claim 2)

[1869] In order to transcribe customer service voice data into text, a means for receiving the customer service voice data as input,

[1870] A means for converting the input customer service voice data into text data,

[1871] Means for generating a customer service script based on the aforementioned text data,

[1872] A means for sending the generated customer service script to a terminal,

[1873] The terminal provides a means for interacting with a user using the transmitted customer service script,

[1874] A means for using a generated AI model together with prompt statements to generate the aforementioned customer service script,

[1875] The system according to claim 1, including the following:

[1876] (Claim 3)

[1877] In order to evaluate the conversation data, a means for receiving the conversation data as input,

[1878] Means for converting the input conversation data into text data,

[1879] A means for inputting the aforementioned character data into a scoring algorithm to calculate an evaluation score,

[1880] A means for transmitting the calculated evaluation score to the terminal,

[1881] The terminal includes means for displaying the transmitted evaluation score,

[1882] A means for dynamically generating the following customer service script based on the aforementioned evaluation score,

[1883] The system according to claim 1, including the following: [Explanation of Symbols]

[1884] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. In order to convert still image data into video data, a means for receiving the still image data and audio data as input, A means for generating video data based on the input still image data and audio data, Means for transmitting the generated video data to a terminal, The terminal includes means for displaying the transmitted video data, A system that includes this.

2. In order to transcribe customer service voice data into text, a means for receiving the customer service voice data as input, A means for converting the input customer service voice data into text data, Means for generating a customer service script based on the aforementioned text data, A means for sending the generated customer service script to a terminal, The terminal provides a means for interacting with a user using the transmitted customer service script, A system that includes this.

3. In order to evaluate the conversation data, a means for receiving the conversation data as input, Means for converting the input conversation data into text data, A means for inputting the aforementioned character data into a scoring algorithm to calculate an evaluation score, A means for transmitting the calculated evaluation score to the terminal, The terminal includes means for displaying the transmitted evaluation score, A system that includes this.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A